Graph data processing method, device, system, electronic device and storage medium
By combining offline data and real-time data processing, graph data is generated and indexed, the problem of insufficient real-time and accuracy of data in graph data processing is solved, maintenance costs are reduced, data integrity and analysis effect are improved.
Patent Information
- Application Number
- CN202310293916.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-23
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2043-03-23
AI Technical Summary
In the prior art, when processing graph data, all data cannot be used in real time, resulting in incomplete real-time and accuracy of the data and high maintenance costs.
By obtaining the data source information of offline data sources and real-time data sources, the graph data structure of graph instances, and processing configuration information, reading and merging data streams, generating graph data matching graph instances, and storing indexes to support fast querying.
It improves the comprehensiveness and integrity of graph data analysis results, reduces the maintenance cost of graph data processing system, and ensures data integrity during graph analysis and traversal.
Smart Images

Figure CN116450890B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of data processing technology, and in particular to a graph data processing method, device, system, electronic device and computer-readable storage medium. Background Art
[0002] Graph data processing refers to processing data sources and outputting them in the form of graphs. In the prior art, when processing graph data, the storage and use of data adopt the classic lambda architecture (a commonly used big data architecture), and offline data and real-time data are processed and used separately. That is, the real-time processing method is combined with the real-time data processing engine to read streaming data from the message queue, and the real-time data processing results are stored in the real-time library to provide real-time graph data services. Offline data processing mainly uses the method of batch loading data files, and then batches of offline data are imported into the offline library at batch time points, and the original graph data is expanded and analyzed in the next step to provide graph data services. When performing graph data retrieval and analysis, whether it is performed on offline data or real-time data, all data cannot be used in real time, resulting in the real-time and accuracy of the data not being the most perfect. In addition, the prior art uses two sets of data processing logic for graph data processing, which has a high maintenance cost. Summary of the invention
[0003] The embodiments of the present application provide a graph data processing method, device, system, electronic device and computer-readable storage medium, which help to improve the comprehensiveness and completeness of the output graph data analysis results and reduce the maintenance cost of the graph data processing system.
[0004] In a first aspect, an embodiment of the present application discloses a graph data processing method, including:
[0005] According to the configuration operation of the configuration personnel, the data source information of the offline data source and the data source information of the real-time data source used as the input data source when processing the graph data, the graph data structure of the graph instance, and the processing configuration information associated with the input data source and the graph data structure are obtained;
[0006] The offline data source is read according to the data source information of the offline data source to obtain a first data stream, and the real-time data source is read according to the data source information of the real-time data source to obtain a second data stream;
[0007] According to the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain a merge processing result;
[0008] Decomposing the merged processing result according to the processing configuration information to generate graph data matching the graph instance;
[0009] Storing the graph data and the index corresponding to the graph instance;
[0010] In response to a query request triggered by a user, the stored graph data is queried based on the index to obtain a query result corresponding to the query request.
[0011] In a second aspect, an embodiment of the present application discloses a graph data processing device, the device comprising:
[0012] A configuration management module, used to obtain, according to the configuration operation of the configuration personnel, the data source information of the offline data source and the data source information of the real-time data source used as the input data source when processing the graph data, the graph data structure of the graph instance, and the processing configuration information associated with the input data source and the graph data structure;
[0013] A data stream acquisition module, configured to read the offline data source according to the data source information of the offline data source to obtain a first data stream, and to read the real-time data source according to the data source information of the real-time data source to obtain a second data stream;
[0014] A data stream merging module, configured to associate and merge the first data stream and the second data stream according to the association relationship between the fields in the first data stream and the second data stream, to obtain a merging result;
[0015] A graph data generation module, configured to perform a disassembly process on the merged processing result according to the processing configuration information, and generate graph data matching the graph instance;
[0016] A graph data and index storage module, used to store the graph data and the index corresponding to the graph instance;
[0017] The query output module is used to query the stored graph data based on the index in response to a query request triggered by a user to obtain a query result corresponding to the query request.
[0018] In a third aspect, an embodiment of the present application discloses a graph data processing system, the system comprising: a data source management engine, a graph database processing engine, a graph data storage backend, an index storage backend, and a query application, wherein:
[0019] The data source management engine is used to manage the data source information of the offline data source and the data source information of the real-time data source used as the input data source when processing graph data, the graph data structure of the pre-configured graph instance, and the processing configuration information associated with the input data source and the graph data structure;
[0020] The graph database processing engine is used to read the offline data source according to the data source information of the offline data source to obtain a first data stream, and to read the real-time data source according to the data source information of the real-time data source to obtain a second data stream;
[0021] The graph database processing engine is further used to merge the first data stream and the second data stream to obtain a merged processing result;
[0022] The graph database processing engine is further used to disassemble the merge processing result according to the processing configuration information to generate graph data matching the graph instance;
[0023] The graph data storage backend is used to store the graph data;
[0024] The index storage backend is used to store the index corresponding to the graph instance;
[0025] The query application is used to query the graph data storage backend based on a query request triggered by a user to obtain a query result corresponding to the query request;
[0026] The graph data storage backend is further used to respond to the query of the query application and retrieve the query result that meets the query request in the locally stored graph data according to the index stored in the index storage backend.
[0027] In a fourth aspect, an embodiment of the present application further discloses an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the graph data processing method described in the embodiment of the present application when executing the computer program.
[0028] In a fifth aspect, an embodiment of the present application discloses a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, performs the steps of the graph data processing method disclosed in the embodiment of the present application.
[0029] The graph data processing system disclosed in the embodiment of the present application includes: a data source management engine, a graph database processing engine, a graph data storage backend, an index storage backend, and a query application. The system manages configuration information through the data source management engine, including: data source information, graph data structure and constructed graph instance, input data source and field mapping configuration information of the graph data structure of the graph instance. After that, the graph database processing engine can collect offline data and real-time data according to the data source information, and merge them into a data stream, and then convert them into graph data and store them in the graph data storage backend. In this way, the query application can query the graph data storage backend according to the received query statement, and the graph data storage backend can return the query result based on the graph data generated by merging the offline data and the real-time data, thereby improving the comprehensiveness and completeness of the output graph data query result. Since the offline data and the real-time data use a common storage location, complete data can always be obtained when performing graph analysis and traversal, which is conducive to global traversal query.
[0030] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0032] Figure 1 It is a schematic diagram of the structure of the graph data processing system disclosed in the embodiment of the present application;
[0033] Figure 2 It is a flowchart of a graph data processing method disclosed in an embodiment of the present application;
[0034] Figure 3 It is a data flow graph in the graph data processing system disclosed in the embodiment of the present application;
[0035] Figure 4 It is a schematic diagram of the structure of a graph data processing device disclosed in an embodiment of the present application;
[0036] Figure 5 A block diagram schematically shows an electronic device for executing the method according to the present application; and
[0037] Figure 6 A storage unit for holding or carrying a program code for implementing the method according to the present application is schematically shown. DETAILED DESCRIPTION
[0038] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0039] The present application embodiment discloses a graph data processing system, such as Figure 1 As shown, the system includes: a data source management engine 110 , a graph database processing engine 120 , a graph data storage backend 130 , an index storage backend 140 , and a query application 150 .
[0040] The specific implementation methods of each component of the graph data processing system are further illustrated below.
[0041] The data source management engine 110 is used to manage the data source information of offline data sources and real-time data sources used as input data sources when processing graph data, the graph data structure of the pre-configured graph instance, and the processing configuration information associated with the input data source and the graph data structure.
[0042] In some embodiments of the present application, the data source management engine 110 is used to manage the information of the data input and output targets of the entire graph data processing, including but not limited to basic metadata such as data format, data type, schema (structure), data source connection address, driver class, constraints, table association relationships, field mapping relationships, etc. Among them, the source of graph data as the input data source can be a relational database, a data file, or other graph database. The metadata in the data source management engine is provided to the graph database processing engine for query and use, and is used for operations such as obtaining data source information, obtaining data constraints and mapping relationships, and normalizing data fields. Optionally, the data source management engine 110 may include: an offline data management module and a real-time data management module.
[0043] In some embodiments of the present application, the data source management engine 110 is also used to: obtain the graph data structure input by the configuration personnel through the configuration interface of the graph data processing system, wherein the graph data structure includes: nodes, edges, and attributes of the nodes and the attributes of the edges; and construct the graph instance according to the graph data structure. For example, when defining the graph model of the device fraud network graph, the device's MAC address, IP address, GPS location, IMEI number, ID number, corporate social credit code and other attributes can be defined as nodes, and the company and MAC address, company and IP address, company and IMEI number, company and GPS location, individual and IMEI number, individual and MAC address, individual and GPS location, etc. can be defined as edges, and the starting point, end point and direction of the edge are defined, as well as the data type, attribute name and other graph data structures of each attribute.
[0044] According to the defined graph data structure, the graph data processing system can create a graph instance.
[0045] In some embodiments of the present application, the data source management engine 110 is also used to: store data source information of the input data source according to the input data source configuration operation of the target graph instance performed by the configuration personnel through the configuration interface of the graph data processing system. For example, the configuration personnel can configure the database type, data source connection address, table association relationship and other information of the input data source of a certain graph instance in the configuration interface of the graph data processing system. Accordingly, after the configuration personnel confirms the configuration, the data source management engine 110 stores the information related to the input data source configured by the configuration personnel.
[0046] Correspondingly, the data source management engine 110 is also used to, in the processing information configuration stage, determine the input data source corresponding to the target graph instance in response to the configuration personnel's selection operation on the input data source; store field mapping configuration information based on the mapping relationship between the data fields of the input data source configured by the configuration personnel and the graph data structure of the target graph instance; and generate processing configuration information based on the input data source corresponding to the target graph instance and the stored field mapping configuration information.
[0047] For example, in the information processing configuration stage, the data source management engine 110 can display an input data source selection interface, and the configuration personnel can select the input data source of the currently configured graph instance through the above selection interface, and select the field name of the input data source through the configuration interface, and map the field of the name to the corresponding graph data structure (i.e., Schema), thereby determining the mapping relationship of field->point or edge attributes, and can configure whether the field needs to be converted, whether the field is indexed, whether it is sorted, etc. Afterwards, the data source management engine 110 stores the mapping relationship of the graph data structure configured by the configuration personnel as field mapping configuration information, and stores it as processing configuration information based on the input data source selected by the configuration personnel and the field mapping configuration information.
[0048] Optionally, the graph database processing engine 120 can read the data source information and process the configuration information through the data source management engine 110. After that, the graph database processing engine 120 performs confluence processing on the data read from the input data source, maps and converts it into a data format adapted to the graph model, and loads and writes it into the graph data storage backend (such as the graph database). The following is an example of the function and implementation of the graph database processing engine 120.
[0049] The graph database processing engine 120 is used to read the offline data source according to the data source information of the offline data source to obtain a first data stream, and to read the real-time data source according to the data source information of the real-time data source to obtain a second data stream.
[0050] Optionally, the offline data source includes: a first database, and the reading of the offline data source according to the data source information of the offline data source to obtain a first data stream includes: writing the full data in the first database corresponding to the data source information of the offline data source into the first message queue by changing the data acquisition tool to obtain the first data stream; or, reading the full data in the first database corresponding to the data source information of the offline data source through a database connection, and writing the read full data into the first message queue to obtain the first data stream.
[0051] In an embodiment of the present application, the graph database processing engine 120 can establish a database connection with the offline data source according to the data source connection address of the offline data source described in the data source information, and then write the offline data in the offline data source into the message queue of the graph data processing system through the CDC (Change Data Capture) tool in full, and use the data in the message queue as a batch write data stream. For example, the FLINK-CDC tool can be used to obtain the full amount of offline data in the offline data source corresponding to the data source information, and write the obtained data into a specified message queue (referred to as the "first message queue" in this article).
[0052] Optionally, for database types that do not support the Flink CDC tool, the database can be read through JDBC (Java Database Connectivity), and then the read data can be written into a message queue (referred to as the "first message queue" in this article) by calling the message queue application interface.
[0053] Optionally, the first message queue is a message queue corresponding to offline data, that is, a message queue corresponding to a batch data stream.
[0054] Optionally, the real-time data source includes: a second database and a message queue, and the real-time data source is read according to the data source information of the real-time data source to obtain the second data stream, including: obtaining the change data in the second database corresponding to the data source information of the real-time data source through a change data acquisition tool to obtain the second data stream; or, reading the real-time data in the message queue specified by the data source information of the real-time data source to obtain the second data stream.
[0055] In some embodiments of the present application, the source of real-time data has multiple channels. For example, the graph database processing engine 120 can obtain real-time data written by an external system through a message queue (such as Kafka) channel as a second data stream. For another example, some real-time data is not written to the message queue in real time, but is directly stored in the database. At this time, the graph database processing engine 120 can use the CDC data collection tool to capture the changed data in the database and obtain the second data stream.
[0056] CDC (change data capture) is a way to back up a database and is often used to back up large amounts of data. The log-based CDC of the MySQL database is to enable the MySQL binary log. To obtain the changes in the data source of the map in the source database in real time, CDC technology can be used. At present, various types of databases in the prior art basically have many mature CDC technologies and corresponding components. In the embodiment of the present application, a CDC data acquisition tool is used to collect data changes in a relational database (RDBMS), which are passed into a real-time computing framework, and Flink real-time computing is used to enter the data into the map in real time.
[0057] The graph database processing engine 120 is further used to merge the first data stream and the second data stream to obtain a merged processing result.
[0058] Optionally, the merging the first data stream and the second data stream to obtain a merged processing result includes: associating and merging the first data stream and the second data stream according to an association relationship between fields in the first data stream and the second data stream to obtain a merged processing result.
[0059] For example, the data stream merging operator calculates data from different data sources according to mapping rules and conversion rules, and converts and encapsulates field data in the source data according to the configured rules. Optionally, the mapping rules and conversion rules can be association methods specified between multiple data streams, for example, the mapping rules can be field mapping relationships specified in field mapping configuration information, and the conversion rules can be field matching, i.e. conversion.
[0060] Optionally, according to the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain a merge processing result, including: determining the first field in the first data stream that matches the field mapping configuration information, and determining the second field in the second data stream that matches the field mapping configuration information; merging the data corresponding to the first field in the first data stream with the data corresponding to the second field in the second data stream to obtain a merge result. For example: fields A1, A2 of stream A and fields B1, B2 of stream B are mapped to the same edge in the field mapping configuration information, then stream A and stream B can be merged and converted into stream C of a specified format, where stream C will carry the field information of stream A and stream B.
[0061] Optionally, the stream data in the first data stream and the second data stream that have a field mapping relationship with the graph data structure may be merged.
[0062] Optionally, according to the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain a merge processing result, including: calling a data stream merging operator (such as the FlinkStream API) of Flink (an open source stream processing framework) to merge the first data stream and the second data stream to obtain a real-time data stream as the merge processing result.
[0063] In one embodiment, the Join class operator of the FluxStream API may be used to merge the first data stream and the second data stream.
[0064] In an embodiment of the present application, the graph database processing engine 120 is further used to decompose the merge processing result according to the processing configuration information to generate graph data matching the graph instance.
[0065] As mentioned above, the processing configuration information includes: field mapping configuration information between the fields of the offline data source and the real-time data source and the points and edges of the graph instance. Optionally, the merging processing result is disassembled according to the processing configuration information to generate graph data matching the graph instance, including: disassembling the merging processing result according to the field mapping configuration information, splitting the merging processing result into data corresponding to the points and edges respectively; writing the data corresponding to the points and edges respectively into the graph instance to generate data corresponding to the points and edges respectively.
[0066] In some embodiments of the present application, the merge processing result is real-time streaming data, including data of fields of interest in the graph instance. In order to facilitate subsequent retrieval and analysis of graph data, the graph data stored in the graph data storage backend is data indexed based on points or edges. Therefore, the graph database processing engine is required to convert the merge processing result into a data format corresponding to the graph data storage backend.
[0067] Optionally, the graph database processing engine needs to perform data conversion according to the specified conversion rules based on the configured functions. For example, according to the field mapping configuration information, the field information that has a mapping relationship with the points, edges, and attributes of the graph instance is extracted, and the converted attribute value of a single edge or a single node is output. When writing to the graph data storage backend, the data of the point or edge will be encapsulated into point and edge objects. For example, the data structure stored in the graph data storage backend is as follows: point V1 (attribute 1 = value 1, attribute 2 = value 2, attribute 3 = value 3, ...), point V2 (attribute 1 = value 1, attribute 2 = value 2, attribute 3 = value 3, ...), edge E1 (identification of the starting point V1, identification of the end point V2, attribute 4 = value 4, attribute 5 = value 5, attribute 6 = value 6, ...), etc. Among them, both point V1 and V2 edge E1 will store globally unique identifiers in the graph data storage backend. The attributes of points and edges are pre-defined attributes for the graph instance. For example, in the device fraud graph model, when edge E1 corresponds to an individual and an IP address, the identifier of the starting point V1 can be the ID number of individual 1, and the identifier of the end point V2 can be the IP address of individual 2. The attributes can be: transaction time, transaction location, transaction amount and other pre-defined attributes.
[0068] In the embodiment of the present application, the graph data storage backend 130 is used to store the graph data.
[0069] Optionally, the graph data storage backend 130 may be a database. The graph data storage backend 130 is used to store data of the graph database, including nodes, edges, and attributes of nodes and edges. Different graph databases provide different graph data storage methods.
[0070] In some embodiments of the present application, the index storage backend 140 is used to store the index corresponding to the graph instance.
[0071] Optionally, the index is generated according to the graph data structure, including: an index of points, an index of edges, an index of point and edge attributes, and an index of the center point of a super point. The index storage backend is used to store indexes, and the index data is independent of the graph data storage backend 130 and provides services for the graph data storage backend 130. Among them, the index stored in the index storage backend 140 is index data generated based on the data stored in the graph data storage backend 130 and the data structure and field mapping configuration information of the graph instance, and its main function is to accelerate queries and graph traversal.
[0072] In database query applications, the index directly determines whether the query statement can be executed efficiently and the result can be calculated. If there is no index, the backend will run the query statement inefficiently, or even be blocked from traversing the result until it times out. In the embodiment of the present application, the index storage backend 140 can be implemented using the mainstream index backend supported by the map database in the prior art, such as ElasticSeach, Apache Solr, and Apache Lucene.
[0073] In an embodiment of the present application, the query application 150 is used to query the graph data storage backend 130 based on a query request triggered by a user to obtain a query result corresponding to the query request.
[0074] The query application 150 includes: a query front end and a query server, and the query request based on the user-triggered query request queries the graph data storage back end to obtain the query result corresponding to the query request, including: the query front end sends a query request to the query server based on the query statement edited by the user, wherein the query request carries the query statement, and the query statement is generated according to the query syntax selected by the user and the graph instance to be queried; the query server performs syntax conversion processing on the query statement carried in the query request according to the query syntax supported by the graph data storage back end to obtain the target query statement supported by the graph data storage back end; the query server queries the graph data storage back end based on the target query statement, and obtains the query result output by the graph data storage back end.
[0075] For example, the query server receives graph query syntax from the query front end, such as Greml in or Cypher syntax, and converts Greml in or Cypher syntax into syntax supported by the graph data storage back end according to the implementation and supported query syntax of the graph data storage back end, and executes the corresponding graph query statement, or executes the graph traversal algorithm, and returns the query result to the query front end. Through syntax conversion processing, it can be compatible with more query syntaxes, making it convenient for users to query graph data.
[0076] In one embodiment of the present application, a user can select a query syntax of the user's preference, or a query syntax supported by the graph data storage backend, on the front-end page of the query application, select a graph instance to be queried, edit the query statement, and pass it to the query server for processing via an http request. The query server determines whether syntax conversion is required based on the query syntax supported by the graph data storage backend. If conversion is required, the preset syntax conversion code is called to convert the query statement sent by the query frontend into a statement executable on the graph data storage backend. Afterwards, the converted query statement is initialized according to the graph instance name (the first query requires obtaining a database connection for the first time, and the connection can be obtained directly from the cache later), and the query statement is executed according to the query application interface provided by the graph data storage backend, waiting for the query result to be returned, and the query result is fed back to the query frontend.
[0077] Optionally, the query server can encapsulate the query results into the format required by the component when the front-end page renders the visual graphics, and then return it to the query front-end.
[0078] The query front end renders the query results into visual graphics using the visualization method selected by the user or the system default.
[0079] The graph data storage backend 130 is further used to respond to the query of the query application and retrieve the query result that meets the query request in the locally stored graph data according to the index stored in the index storage backend.
[0080] For example, the graph data storage backend 130 obtains the query parameters carried in the query application interface in response to the query application interface being called; then, based on the query parameters, the index storage backend 140 performs an index acceleration query to obtain index information; finally, based on the index information, the graph data stored in the graph data storage backend 130 is retrieved to obtain the query result that satisfies the query request.
[0081] The graph data processing system disclosed in the embodiments of the present application is configured to manage the data source information of the offline data source and the data source information of the real-time data source used as the input data source when processing graph data, the graph data structure of the pre-configured graph instance, and the processing configuration information associated with the input data source and the graph data structure by setting a data source management engine; setting a graph database processing engine to read the offline data source according to the data source information of the offline data source to obtain a first data stream, and to read the real-time data source according to the data source information of the real-time data source to obtain a second data stream; and merging the first data stream and the second data stream to obtain a merged processing result, and then disassembling the merged processing result according to the processing configuration information to generate graph data matching the graph instance; setting a graph data storage backend to store the graph data; setting an index storage backend to store the index corresponding to the graph instance; when a user queries graph data through a query application, the query result can be returned based on the offline data and the real-time data, thereby improving the comprehensiveness and completeness of the output graph data query result. Since offline data and real-time data use the same storage location, complete data can always be obtained when performing graph analysis and traversal, which is conducive to global traversal queries.
[0082] In the prior art, the offline data processing architecture is: offline data is uniformly extracted the next day, and the newly added data of the day is uniformly processed in batches and then written into the database. Since the data is updated every day, only yesterday's data can be found when querying, and the real-time data of the day is still in the source database and has not been processed into the graph database. Offline data and real-time data need to be stored separately and cannot be uniformly entered into the graph. The graph data processing system disclosed in the embodiment of the present application merges offline data and real-time data into a data stream for unified storage, thereby achieving the same entry into the graph and improving data integrity and comprehensiveness.
[0083] On the other hand, since the offline data and real-time data are first merged through the graph database processing engine, a set of code logic is used to subsequently process the offline data and real-time data, which can reduce the maintenance cost of the graph data processing system.
[0084] On the other hand, the embodiments of the present application also disclose a method for processing graph data. Figure 2 The flowchart and Figure 3 The data flow diagram shown in the figure illustrates an implementation method of the graph data processing method. Figure 3 yes Figure 1 The data flow diagram of the data processing system shown in the figure.
[0085] like Figure 2As shown, the graph data processing method disclosed in the embodiment of the present application includes: steps 210 to 260.
[0086] Step 210, based on the configuration operations of the configuration personnel, obtain the data source information of the offline data source and the data source information of the real-time data source used as the input data source when processing the graph data, the graph data structure of the graph instance, and the processing configuration information associated with the input data source and the graph data structure.
[0087] like Figure 3 As shown, the data source management engine 110 has built-in functional modules such as offline data configuration management, real-time data configuration management, and graph instance configuration management, which are respectively used for offline data source configuration management, real-time data source configuration management, graph data structure definition, graph instance construction, and field mapping configuration.
[0088] For example, the data source management engine 110 of the graph data processing system may store data source information of the input data source according to the input data source configuration operation of the target graph instance performed by the configuration personnel through the configuration interface of the graph data processing system.
[0089] For another example, in the information processing configuration stage, the data source management engine 110 can determine the input data source corresponding to the target graph instance in response to the configuration personnel's selection operation on the input data source; store field mapping configuration information based on the mapping relationship between the data fields of the input data source configured by the configuration personnel and the graph data structure of the target graph instance; and generate processing configuration information based on the input data source corresponding to the target graph instance and the stored field mapping configuration information.
[0090] According to the configuration operations of the configuration personnel, the data source information of the offline data source and the real-time data source used as the input data source when processing the graph data, the graph data structure of the graph instance, and the specific implementation of the processing configuration information associated with the input data source and the graph data structure are obtained. Please refer to the relevant description of the specific implementation of the data source management engine in the previous text, which will not be repeated here.
[0091] Step 220, reading the offline data source according to the data source information of the offline data source to obtain a first data stream, and reading the real-time data source according to the data source information of the real-time data source to obtain a second data stream.
[0092] After completing the basic configuration, the data source management engine 110 stores and manages the data source information, the created graph instances, and the field mapping configuration information configured for each graph instance, and provides them to the graph database processing engine 120 and the index storage backend 140 for use. Figure 3As shown, the data source management engine 110 includes functional modules such as offline data reading and conversion, real-time data reading and conversion, and confluence calculation, which are respectively used to perform operations such as first data stream collection, second data stream collection, and data stream confluence processing.
[0093] When graph data needs to be generated, the graph database processing engine 120 reads the processing configuration information from the data source management engine 110, and connects the offline data source and the real-time data source according to the processing configuration information to collect data. The graph database processing engine 120 collects the offline data to obtain a first data stream, and collects the real-time data source to obtain a second data stream.
[0094] Optionally, the offline data source includes: a first database, and the database is read according to the data source information of the offline data source to obtain a first data stream, including: writing the full data in the first database corresponding to the data source information of the offline data source into the first message queue by changing the data acquisition tool to obtain the first data stream; or, reading the full data in the first database corresponding to the data source information of the offline data source through a database connection, and writing the read full data into the first message queue to obtain the first data stream.
[0095] Optionally, the real-time data source includes: a second database and a message queue, and the real-time data source is read according to the data source information of the real-time data source to obtain the second data stream, including: obtaining the change data in the second database corresponding to the data source information of the real-time data source through a change data acquisition tool to obtain the second data stream; or, reading the real-time data in the message queue specified by the data source information of the real-time data source to obtain the second data stream.
[0096] The graph database processing engine 120 reads the offline data source according to the data source information of the offline data source to obtain a first data stream, and reads the real-time data source according to the data source information of the real-time data source to obtain a second data stream. For specific implementation methods, please refer to the previous description of the specific implementation methods of the graph database processing engine 120, which will not be repeated here.
[0097] Step 230: According to the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain a merge processing result.
[0098] Next, the graph database processing engine 120 merges the first data stream and the second data stream to merge the offline data and the real-time data into one data stream for unified processing.
[0099] Optionally, based on the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain a merge processing result, including: determining the first field in the first data stream that matches the field mapping configuration information, and determining the second field in the second data stream that matches the field mapping configuration information; merging the data corresponding to the first field in the first data stream with the data corresponding to the second field in the second data stream to obtain a merge result.
[0100] According to the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain the specific implementation method of the merged processing result. Please refer to the previous description of the specific implementation method of the graph database processing engine 120, which will not be repeated here.
[0101] Step 240: Decompose the merged processing result according to the processing configuration information to generate graph data matching the graph instance.
[0102] Afterwards, the graph database processing engine 120 formats the data stream obtained by merging the offline data and the real-time data, and converts it into the data format required by the graph data storage backend 130, that is, the graph data format.
[0103] Optionally, the processing configuration information includes: field mapping configuration information between fields of the offline data source and the real-time data source and the points and edges of the graph instance; and the decomposing the merged processing result according to the processing configuration information to generate graph data matching the graph instance includes: decomposing the merged processing result according to the field mapping configuration information, splitting the merged processing result into data corresponding to the points and edges respectively; writing the data corresponding to the points and edges respectively into the graph instance to generate data corresponding to the points and edges respectively.
[0104] The merge processing result is disassembled according to the processing configuration information to generate a specific implementation method of the graph data matching the graph instance. Please refer to the previous description of the specific implementation method of the graph database processing engine 120, which will not be repeated here.
[0105] Step 250, store the graph data and the index corresponding to the graph instance.
[0106] The graph database processing engine 120 sends the converted graph data to the graph data storage backend 130 for storage.
[0107] The generation and storage method of the index corresponding to the graph instance can be found in the description of the specific implementation of the index storage backend 140 in the previous text, which will not be repeated here.
[0108] At this point, the graph data storage backend 130 stores the data of each vertex and each edge of the graph instance, and the index storage backend 140 stores the indexes of the vertex and edge of the graph instance.
[0109] Step 260, in response to a query request triggered by a user, query the stored graph data based on the index to obtain a query result corresponding to the query request.
[0110] When a user inputs a query statement through the query application 150 of the graph data processing system to query information in a target graph instance, the query application 150 will query the stored graph data based on the index to obtain a query result corresponding to the query request.
[0111] Optionally, the query application includes: a query front end and a query server, which responds to a user triggering a query request and queries the stored graph data based on the index to obtain a query result corresponding to the query request, including: the query front end sends a query request to the query server based on a query statement edited by the user, wherein the query request carries the query statement, and the query statement is generated according to the query syntax selected by the user and the graph instance to be queried; the query server performs syntax conversion on the query statement carried in the query request according to the query syntax supported by the graph data storage backend to obtain a target query statement supported by the graph data storage backend; the query server queries the graph data storage backend based on the target query statement, and obtains the query result output by the graph data storage backend.
[0112] The query application 150 responds to the user triggering a query request and queries the stored graph data based on the index to obtain the specific implementation of the query result corresponding to the query request. Please refer to the relevant description of the specific implementation of the query application 150 in the previous text, which will not be repeated here.
[0113] The graph data processing method disclosed in the embodiment of the present application obtains the data source information of the offline data source and the real-time data source as the input data source when processing the graph data, the graph data structure of the graph instance, and the processing configuration information associated with the input data source and the graph data structure according to the configuration operation of the configuration personnel; reads the offline data source according to the data source information of the offline data source to obtain a first data stream, and reads the real-time data source according to the data source information to obtain a second data stream; associates and merges the first data stream and the second data stream according to the association relationship between the fields in the first data stream and the second data stream to obtain a merged processing result; disassembles the merged processing result according to the processing configuration information to generate graph data matching the graph instance; then, stores the graph data and the index corresponding to the graph instance; when a user triggers a query request, queries the stored graph data containing offline data and real-time data based on the index to obtain the query result corresponding to the query request, thereby improving the comprehensiveness and completeness of the output graph data query result.
[0114] Since offline data and real-time data use the same storage location, complete data can always be obtained when performing graph analysis and traversal, which is conducive to global traversal queries.
[0115] In the prior art, the offline data processing architecture is: offline data is uniformly extracted the next day, and the newly added data of the day is uniformly processed in batches and then written into the database. Since the data is updated every day, only yesterday's data can be found when querying, and the real-time data of the day is still in the source database and has not been processed into the graph database. Offline data and real-time data need to be stored separately and cannot be uniformly entered into the graph. The graph data processing method disclosed in the embodiment of the present application merges offline data and real-time data into a data stream for unified storage, thereby achieving the same entry into the graph and improving data integrity and comprehensiveness.
[0116] On the other hand, since the offline data and real-time data are first merged through the graph database processing engine, a set of code logic is used to subsequently process the offline data and real-time data, which can reduce the maintenance cost of the graph data processing system.
[0117] Correspondingly, the present application also discloses a graph data processing device, such as Figure 4 As shown, the device comprises:
[0118] Configuration management module 410, used to obtain data source information of an offline data source and a real-time data source as input data sources when processing graph data, a graph data structure of a graph instance, and processing configuration information associated with the input data source and the graph data structure according to configuration operations of a configuration personnel;
[0119] The data stream acquisition module 420 is used to read the offline data source according to the data source information of the offline data source to obtain a first data stream, and to read the real-time data source according to the data source information of the real-time data source to obtain a second data stream;
[0120] A data stream merging module 430, configured to associate and merge the first data stream and the second data stream according to the association relationship between the fields in the first data stream and the second data stream to obtain a merging result;
[0121] A graph data generation module 440, configured to perform a disassembly process on the merge processing result according to the processing configuration information to generate graph data matching the graph instance;
[0122] A graph data and index storage module 450, used to store the graph data and the index corresponding to the graph instance;
[0123] The query output module 460 is used to query the stored graph data based on the index in response to a query request triggered by a user to obtain a query result corresponding to the query request.
[0124] Optionally, the data stream merging module 430 is further configured to:
[0125] According to the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain a merge processing result.
[0126] Optionally, the processing configuration information includes: field mapping configuration information between fields of the offline data source and the real-time data source and points and edges of the graph instance, and the graph data generation module is further used to:
[0127] Decomposing the merge processing result according to the field mapping configuration information, and decomposing the merge processing result into data corresponding to points and edges respectively;
[0128] The data corresponding to the points and edges are written into the graph instance to generate data corresponding to the points and edges.
[0129] Optionally, the offline data source includes: a first database, and the reading of the offline data source according to the data source information of the offline data source to obtain the first data stream includes:
[0130] Write the full amount of data in the first database corresponding to the data source information of the offline data source into the first message queue by changing the data acquisition tool to obtain a first data stream; or
[0131] The full amount of data in the first database corresponding to the data source information of the offline data source is read through a database connection, and the read full amount of data is written into a first message queue to obtain a first data stream.
[0132] Optionally, the real-time data source includes: a second database and a message queue, and the real-time data source is read according to the data source information of the real-time data source to obtain the second data stream, including:
[0133] Acquire the changed data in the second database corresponding to the data source information of the real-time data source through a change data acquisition tool to obtain a second data stream; or
[0134] The real-time data in the message queue specified by the data source information of the real-time data source is read to obtain a second data stream.
[0135] Optionally, the query application includes: a query front end and a query server, and the query output module 460 is further used to:
[0136] The query front end sends a query request to the query server based on the query statement edited by the user, wherein the query request carries the query statement, and the query statement is generated according to the query syntax selected by the user and the graph instance to be queried;
[0137] The query server performs syntax conversion processing on the query statement carried in the query request according to the query syntax supported by the graph data storage backend to obtain the target query statement supported by the graph data storage backend;
[0138] The query server queries the graph data storage backend based on the target query statement, and obtains the query result output by the graph data storage backend.
[0139] The graph data processing device disclosed in the embodiments of the present application is used to implement the graph data processing method described in the embodiments of the present application. The specific implementation methods of each module of the device will not be repeated here, and reference can be made to the specific implementation methods of the corresponding steps in the method embodiments.
[0140] The present application discloses a graph data processing device, which obtains data source information of an offline data source and data source information of a real-time data source as input data sources when processing graph data, a graph data structure of a graph instance, and processing configuration information of the input data source associated with the graph data structure according to configuration operations performed by configuration personnel; reads the offline data source according to the data source information of the offline data source to obtain a first data stream, and reads the real-time data source according to the data source information of the real-time data source to obtain a second data stream; associates and merges the first data stream and the second data stream according to the association relationship between the fields in the first data stream and the second data stream to obtain a merged processing result; disassembles the merged processing result according to the processing configuration information to generate graph data matching the graph instance; then, stores the graph data and the index corresponding to the graph instance; when a user triggers a query request, queries the stored graph data containing offline data and real-time data based on the index to obtain a query result corresponding to the query request, thereby improving the comprehensiveness and completeness of the output graph data query result.
[0141] Since offline data and real-time data use the same storage location, complete data can always be obtained when performing graph analysis and traversal, which is conducive to global traversal queries.
[0142] On the other hand, since the offline data and real-time data are first merged through the graph database processing engine, a set of code logic is used to subsequently process the offline data and real-time data, which can reduce the maintenance cost of the graph data processing system.
[0143] Each embodiment in this specification is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.
[0144] The above is a detailed introduction to a graph data processing method and device provided by the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and a core idea. At the same time, for general technical personnel in this field, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.
[0145] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0146] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or digital signal processor (DSP) can be used in practice to implement some or all functions of some or all components in the electronic device according to the embodiment of the present application. The present application can also be implemented as a device or apparatus program (e.g., computer program and computer program product) for executing part or all of the methods described herein. Such a program implementing the present application can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0147] For example, Figure 5 An electronic device that can implement the method according to the present application is shown. The electronic device can be a PC, a mobile terminal, a personal digital assistant, a tablet computer, etc. The electronic device traditionally includes a processor 510 and a memory 520 and a program code 530 stored on the memory 520 and can be run on the processor 510, and the processor 510 implements the method described in the above embodiment when executing the program code 530. The memory 520 can be a computer program product or a computer-readable medium. The memory 520 can be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. The memory 520 has a storage space 5201 for the program code 530 of the computer program for executing any method step in the above method. For example, the storage space 5201 for the program code 530 can include various computer programs for implementing the various steps in the above method respectively. The program code 530 is a computer-readable code. These computer programs can be read from one or more computer program products or written into the one or more computer program products. These computer program products include program code carriers such as hard disks, compact disks (CDs), memory cards or floppy disks. The computer program includes a computer readable code, and when the computer readable code is executed on an electronic device, the electronic device is caused to execute the method according to the above embodiment.
[0148] The embodiment of the present application also discloses a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the steps of the graph data processing method described in the embodiment of the present application are implemented.
[0149] Such a computer program product may be a computer-readable storage medium, which may have a computer-readable storage medium having a Figure 5 The memory 520 in the electronic device shown in the figure is similar to the storage segment, storage space, etc. The program code can be compressed and stored in the computer-readable storage medium in an appropriate form. The computer-readable storage medium is generally as shown in the reference Figure 6 The portable or fixed storage unit. Generally, the storage unit includes computer readable code 530', which is a code read by a processor. When the code is executed by the processor, each step in the method described above is implemented.
[0150] The term "one embodiment", "embodiment" or "one or more embodiments" herein means that a particular feature, structure or characteristic described in conjunction with the embodiment is included in at least one embodiment of the present application. In addition, please note that the examples of the term "in one embodiment" here do not necessarily all refer to the same embodiment.
[0151] In the description provided herein, a large number of specific details are described. However, it is understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.
[0152] In the claims, any reference signs placed between brackets shall not be construed as limiting the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several different elements and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by the same item of hardware. The use of the words first, second, and third etc. does not indicate any order. These words may be interpreted as names.
[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A graph data processing system, It is characterized in that The system includes: a data source management engine, a graph database processing engine, a graph data storage backend, an index storage backend, and a query application, wherein: The data source management engine is used to manage the data source information of the offline data source and the data source information of the real-time data source used as the input data source when processing graph data, the graph data structure of the pre-configured graph instance, and the processing configuration information associated with the input data source and the graph data structure; The graph database processing engine is used to read the offline data source according to the data source information of the offline data source to obtain a first data stream, and to read the real-time data source according to the data source information of the real-time data source to obtain a second data stream; The graph database processing engine is further used to merge the first data stream and the second data stream to obtain a merged processing result; The graph database processing engine is further used to disassemble the merge processing result according to the processing configuration information to generate graph data matching the graph instance; The graph data storage backend is used to store the graph data; The index storage backend is used to store the index corresponding to the graph instance; The query application is used to query the graph data storage backend based on a query request triggered by a user to obtain a query result corresponding to the query request; The graph data storage backend is further used to respond to the query of the query application and retrieve the query result that meets the query request in the locally stored graph data according to the index stored in the index storage backend.
2. The system according to claim 1, It is characterized in that The merging the first data stream and the second data stream to obtain a merging result includes: According to the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain a merge processing result.
3. The system according to claim 1, It is characterized in that The processing configuration information includes: field mapping configuration information between the fields of the offline data source and the real-time data source and the points and edges of the graph instance, and the disassembling processing of the merge processing result according to the processing configuration information to generate graph data matching the graph instance includes: Decomposing the merge processing result according to the field mapping configuration information, and decomposing the merge processing result into data corresponding to points and edges respectively; The data corresponding to the points and edges are written into the graph instance to generate data corresponding to the points and edges.
4. The system according to claim 1, It is characterized in that The offline data source includes: a first database, and the offline data source is read according to the data source information of the offline data source to obtain a first data stream, including: Write all the data in the first database corresponding to the data source information of the offline data source into the first message queue by changing the data acquisition tool to obtain a first data stream; or The full amount of data in the first database corresponding to the data source information of the offline data source is read through a database connection, and the read full amount of data is written into a first message queue to obtain a first data stream.
5. The system according to claim 1, It is characterized in that The real-time data source includes: a second database and a message queue, and the real-time data source is read according to the data source information of the real-time data source to obtain a second data stream, including: Acquire the changed data in the second database corresponding to the data source information of the real-time data source through a change data acquisition tool to obtain a second data stream; or The real-time data in the message queue specified by the data source information of the real-time data source is read to obtain a second data stream.
6. The system according to claim 1, It is characterized in that The query application includes: a query front end and a query server. The query request triggered by the user queries the graph data storage back end to obtain the query result corresponding to the query request, including: The query front end sends a query request to the query server based on the query statement edited by the user, wherein the query request carries the query statement, and the query statement is generated according to the query syntax selected by the user and the graph instance to be queried; The query server performs syntax conversion processing on the query statement carried in the query request according to the query syntax supported by the graph data storage backend to obtain the target query statement supported by the graph data storage backend; The query server queries the graph data storage backend based on the target query statement, and obtains the query result output by the graph data storage backend.
7. A graph data processing method, It is characterized in that The method comprises: According to the configuration operation of the configuration personnel, the data source information of the offline data source and the data source information of the real-time data source used as the input data source when processing the graph data, the graph data structure of the graph instance, and the processing configuration information associated with the input data source and the graph data structure are obtained; The offline data source is read according to the data source information of the offline data source to obtain a first data stream, and the real-time data source is read according to the data source information of the real-time data source to obtain a second data stream; According to the association relationship between the fields in the first data stream and the second data stream, the first data stream and the second data stream are associated and merged to obtain a merge processing result; Decomposing the merged processing result according to the processing configuration information to generate graph data matching the graph instance; Storing the graph data and the index corresponding to the graph instance; In response to a query request triggered by a user, the stored graph data is queried based on the index to obtain a query result corresponding to the query request.
8. A graph data processing device, It is characterized in that The device comprises: A configuration management module, used to obtain, according to the configuration operation of the configuration personnel, the data source information of the offline data source and the data source information of the real-time data source used as the input data source when processing the graph data, the graph data structure of the graph instance, and the processing configuration information associated with the input data source and the graph data structure; A data stream acquisition module, configured to read the offline data source according to the data source information of the offline data source to obtain a first data stream, and to read the real-time data source according to the data source information of the real-time data source to obtain a second data stream; A data stream merging module, configured to associate and merge the first data stream and the second data stream according to the association relationship between the fields in the first data stream and the second data stream, to obtain a merging processing result; A graph data generation module, configured to perform a disassembly process on the merged processing result according to the processing configuration information, and generate graph data matching the graph instance; A graph data and index storage module, used to store the graph data and the index corresponding to the graph instance; The query output module is used to query the stored graph data based on the index in response to a query request triggered by a user to obtain a query result corresponding to the query request.
9. An electronic device comprising a memory, a processor, and a program code stored in the memory and executable on the processor, It is characterized in that When the processor executes the program code, the graph data processing method described in claim 7 is implemented.
10. A computer-readable storage medium having program code stored thereon, It is characterized in that When the program code is executed by a processor, the steps of the graph data processing method described in claim 7 are implemented.
Citation Information
Patent Citations
Scientific and technological resource integration system based on multi-source database
CN113312342A
Field blood relationship processing method and system based on graph database
CN115080570A