Method and device for processing data based on graph database
By using interactive graphs and statistical records in the graph database, the interactive event data is deduplicated and updated, and the problem of difficulty in efficiently deduplication and accumulation statistics in the existing technology is solved, and more accurate statistical analysis results and higher data processing efficiency are achieved.
Patent Information
- Application Number
- CN202110891437.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-04
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2041-08-04
AI Technical Summary
In business prediction analysis based on interactive events, it is difficult for the prior art to efficiently deduplicate and accumulate statistics on indicators or features, resulting in inaccurate statistical results.
By using interactive graphs and statistical records in the graph database, the accumulated values of the record items in the statistical record are deduplicated and updated based on the edge properties between the subject identification and the interactive object identification in the interactive graph, so that the accumulated values in the record items are deduplicated and updated values.
The deduplication and update of the accumulated values is realized, ensuring that the accumulated values obtained during statistical analysis are the deduplication data, improving the accuracy of the statistical results, and saving storage space and calculation time.
Smart Images

Figure CN113672617B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and in particular, to a method and device for processing data based on a graph database. Background Art
[0002] With the development of computer technology, machine learning has been applied to various technical fields to analyze and predict various business data. For example, there are many types of interactive events in the Internet environment, such as purchase events, click events, payment events, etc. In many scenarios, it is necessary to analyze and process various interactive events to make business predictions. For example, the risk level of user operation behavior can be evaluated based on the history of interactive events to prevent and control risks; or user preferences can be evaluated based on historical events to better provide users with personalized services.
[0003] In scenarios where business predictions are made based on interactive events, it is often necessary to use cumulative indicators to characterize the characteristics of the interactive subject. For example, in the merchant risk control scenario, the cumulative indicators of the merchant subject in different dimensions such as interactive users, interactive locations, and interactive product categories, especially the cumulative indicators after deduplication, are very important features for characterizing merchant risks. Compared with simple cumulative values without deduplication, the cumulative statistical values after deduplication can eliminate the impact of multiple interactions of a single user and have better stability.
[0004] Therefore, we hope to have an improved solution to more efficiently deduplicate and accumulate statistics of indicators or features for subsequent business forecasting and analysis. Summary of the invention
[0005] The embodiments of this specification describe a method and device for processing data based on a graph database. When the method uses interactive event data to update statistical records in the graph database, based on the edge attributes between the subject identifier and the interactive object identifier in the interactive graph, the cumulative values of the record items in the statistical records are deduplicated, so that the cumulative values in the record items are the deduplicated updated values, so that the cumulative values of each time period obtained when performing statistical analysis based on the graph database are the deduplicated data, so the statistical results are more accurate. In addition, the graph database of this embodiment stores the interactive graph and statistical records, and does not need to store the detailed data of the interactive events, thus saving storage space, reducing the time for data search and calculation, and improving data processing efficiency.
[0006] According to a first aspect, a method for processing data based on a graph database is provided, comprising: obtaining interaction event data, including a subject identifier of an interaction subject, interaction time, and several interaction object identifiers corresponding to target indicators to be accumulated and counted; reading an interaction graph and a target statistical record of the target indicator for the subject identifier from a graph database, wherein in the interaction graph, interaction subjects and interaction objects having interaction history are connected by edges, and the corresponding edge attributes include the most recent interaction time; the target statistical record includes multiple record items corresponding to multiple time periods, respectively, and a single record item includes each cumulative value of the target indicator in each time period of a predetermined number N of time periods traced back from the corresponding time period as the starting point; for each interaction object identifier among the several interaction object identifiers, performing an update operation on the graph database, wherein the update operation includes: determining a target record item corresponding to the interaction time in the target statistical record; and performing deduplication update on the target record item according to the edge attribute between the subject identifier and the interaction object identifier in the interaction graph.
[0007] In one embodiment, the above-mentioned obtaining of interaction event data includes: obtaining event data corresponding to a single interaction event as the above-mentioned interaction event data.
[0008] In one embodiment, the above-mentioned acquisition of interaction event data includes: sending event data corresponding to the interaction event generated by streaming to a streaming computing engine, the above-mentioned streaming computing engine aggregating the incoming event data according to the subject identifier at a preset time interval to obtain at least one aggregation result; obtaining each aggregation result from the above-mentioned streaming computing engine as interaction event data.
[0009] In one embodiment, determining the target record item corresponding to the above-mentioned interaction time in the above-mentioned target statistical record includes: determining whether there is a record item in the above-mentioned target statistical record whose starting time period includes the above-mentioned interaction time; if so, determining the existing record item as the target record item; if not, adding a new record item in the above-mentioned target statistical record as the target record item, wherein the newly added record item is used to record, starting from the time period corresponding to the above-mentioned interaction time, and tracing back to each cumulative value of the above-mentioned target indicator in each time period of N time periods.
[0010] In one embodiment, the above-mentioned adding a new record item as the target record item in the above-mentioned target statistical record includes: generating a new record item to be filled; in the above-mentioned new record item, initializing the cumulative value corresponding to the starting time period to 0; in the latest record item existing before the above-mentioned new record item, determining the N cumulative values corresponding to the previous N time periods, and copying them to the above-mentioned new record item to obtain the above-mentioned target record item.
[0011] In one embodiment, the above-mentioned target record item is deduplicated and updated according to the edge attribute between the above-mentioned subject identifier and the interactive object identifier in the above-mentioned interaction graph, including: adding 1 to the cumulative value corresponding to the starting time period of the above-mentioned target record item; determining whether there is an edge between the above-mentioned subject identifier and the interactive object identifier in the above-mentioned interaction graph; if the edge exists, subtracting 1 from the cumulative value in the time period corresponding to the most recent interaction time included in the edge attribute, and updating the most recent interaction time of the edge attribute to the above-mentioned interaction time; if the edge does not exist, adding the edge formed between the above-mentioned subject identifier and the interactive object identifier in the above-mentioned interaction graph, wherein the edge attribute of the added edge includes the above-mentioned interaction time.
[0012] In one embodiment, the above operation further includes: determining whether the number of record items included in the deduplicated updated statistical record exceeds a maximum number threshold, and if so, deleting a record item with the earliest starting time period in the updated statistical record.
[0013] In one embodiment, the method further includes: in response to determining that the updating of the plurality of interaction object identifiers is completed, writing the updated interaction graph and target statistical records back to the graph database.
[0014] In one embodiment, the target statistical record is a numerical matrix, and a single record item corresponds to a row of the matrix.
[0015] In one embodiment, the method further includes: determining data statistical parameters, wherein the data statistical parameters include: an identifier of an interactive subject to be analyzed, an indicator to be analyzed, and a time window for statistical analysis; reading statistical records from the graph database as statistical records to be analyzed based on the identifier of the interactive subject to be analyzed and the indicator to be analyzed; determining record items to be analyzed from the statistical records to be analyzed based on the time window; reading data corresponding to the time window from the record items to be analyzed, and performing statistical analysis on the data to obtain statistical analysis results.
[0016] In one embodiment, determining the record items to be analyzed from the statistical records to be analyzed based on the time window includes: determining whether there is a record item with an end time period as a starting point in the statistical records to be analyzed, wherein the end time period is a time period corresponding to the end time of the time window; if so, reading the record item with the end time period as a starting point from the statistical records to be analyzed as the record item to be analyzed; if not, reading a record item with a starting time period greater than the end time period from the statistical records to be analyzed as the record item to be analyzed.
[0017] In one embodiment, data corresponding to the time window is read from the record item to be analyzed, and statistical analysis is performed on the data to obtain statistical analysis results, including: reading from the record item to be analyzed, each cumulative value within each time period from the start time period corresponding to the time window to its end time period, and accumulating the above cumulative values to obtain the above statistical analysis results.
[0018] In one embodiment, the method further comprises: providing feedback of the statistical analysis results.
[0019] According to a second aspect, a device for processing data based on a graph database is provided, comprising: an acquisition unit, configured to acquire interaction event data, including a subject identifier of an interaction subject, an interaction time, and several interaction object identifiers corresponding to the target indicator to be accumulated and counted; a reading unit, configured to read an interaction graph and a target statistical record of the above-mentioned subject identifier for the above-mentioned target indicator from the graph database, wherein in the above-mentioned interaction graph, the interaction subject and the interaction object with an interaction history are connected by an edge, and the corresponding edge attribute includes the most recent interaction time; the above-mentioned target statistical record includes a plurality of record items corresponding to a plurality of time periods respectively, and a single record item includes each accumulated value of the above-mentioned target indicator in each time period of a predetermined number N of time periods traced back forward with the corresponding time period as the starting point; an operation unit, configured to perform an update operation on the above-mentioned graph database for each interaction object identifier among the above-mentioned several interaction object identifiers, wherein the above-mentioned operation unit includes: a determination module, configured to determine the target record item corresponding to the above-mentioned interaction time in the above-mentioned target statistical record; a deduplication update module, configured to perform deduplication update on the above-mentioned target record item according to the edge attribute between the above-mentioned subject identifier and the interaction object identifier in the above-mentioned interaction graph.
[0020] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the method described in any implementation manner in the first aspect.
[0021] According to a fourth aspect, a computing device is provided, comprising a memory and a processor, wherein executable code is stored in the memory, and when the processor executes the executable code, the method described in any implementation manner in the first aspect is implemented.
[0022] According to the method and device for processing data based on a graph database provided in the embodiments of this specification, the interactive event data is first obtained, and then the interactive graph and the target statistical record of the subject identification for the target indicator are read from the graph database, and then, for each interactive object identification in the several interactive object identifications in the interactive event data, the target statistical record in the graph database is deduplicated and updated. When the method uses the interactive event data to update the statistical record, based on the edge attributes between the subject identification and the interactive object identification in the interactive graph, the cumulative value of the record item in the statistical record is deduplicated, so that the cumulative value in the record item is the deduplicated updated value, so that the cumulative value of each time period obtained when performing statistical analysis based on the graph database is the deduplicated data, so the statistical result is more accurate. In addition, the graph database of this embodiment stores the interactive graph and the statistical record, and there is no need to store the detailed data of the interactive event, thereby saving storage space, reducing the time for data search and calculation, and improving data processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A schematic diagram showing application scenarios in which the embodiments of this specification can be applied.
[0024] Figure 2 A flowchart of a method for processing data based on a graph database according to an embodiment is shown;
[0025] Figure 3 A schematic diagram showing an example of using a streaming computing engine to generate aggregate results;
[0026] Figure 4 A schematic diagram showing an interaction graph in an example;
[0027] Figure 5 A schematic diagram showing the steps of performing deduplication update on a target record item in one embodiment is shown;
[0028] Figure 6 A flowchart showing the steps of performing statistical analysis on data based on a graph database according to one embodiment;
[0029] Figure 7 A schematic block diagram of an apparatus for processing data based on a graph database according to an embodiment is shown. DETAILED DESCRIPTION
[0030] The technical solution provided in this specification is further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant inventions, rather than to limit the inventions. It should also be noted that, for ease of description, only the parts related to the relevant inventions are shown in the accompanying drawings. It should be noted that, in the absence of conflict, the embodiments of this specification and the features in the embodiments can be combined with each other.
[0031] As mentioned above, cumulative indicators, especially deduplicated cumulative indicators, are important features for characterizing the characteristics of interactive entities and conducting business forecasting and analysis. For example, for a merchant, it is very important to count the number of different users who have conducted transactions and the number of different commodities sold within a period of time (for example, within the past month) to obtain deduplicated cumulative user volume, commodity category quantity and other cumulative indicators for characterizing the characteristics of the merchant and conducting business analysis.
[0032] In order to calculate the deduplicated cumulative statistics, in one solution, the detailed data can be directly stored. Then, by querying the detailed data, the deduplicated cumulative statistics are calculated in memory. However, this method requires a large amount of data storage and calculation, especially on ultra-large data sets. It is often impossible to directly calculate or the calculation time is very long due to the large amount of data. According to another solution, after storing the daily interaction detailed data, the data of one day can be stored in the daily account data according to certain accumulation rules, and then the cumulative results are calculated based on the daily account and the details of the day. This method requires relying on a scheduled task to trigger the daily account accumulation during the off-peak period, and the daily accumulation has a certain lag. At the same time, this method needs to store the detailed data of one day. When encountering a large amount of hot data, multiple calculations between different daily accounts and some repeated data in the detailed data will lead to a loss of accuracy.
[0033] To this end, the embodiments of this specification provide a solution for data preparatory processing based on a graph database, so as to efficiently and accurately obtain cumulative statistical values. For example, the interaction event is a transaction event, and the cumulative indicator is the number of user IDs that trade with the merchant. Figure 1 Schematic diagram showing application scenarios in which the embodiments of this specification can be applied. Figure 1As shown, the graph database of the embodiment of the present specification includes an interaction graph 101 and a statistical record 102. As an example, merchants M1, M2... and users U1, U2, U3... in the interaction graph 101 are connected by edges, and the statistical record 102 includes multiple record items. After a transaction event occurs between a merchant and a user, the graph database is updated based on the deduplication of the transaction event data. When it is necessary to generate a cumulative value of the number of users who have transactions with a merchant (for example, merchant M1) within a certain time period [T1, T2], the graph database can be queried for each deduplication cumulative value within the time period for statistical analysis, thereby obtaining a deduplication cumulative statistical value.
[0034] Continue to see Figure 2 , Figure 2 The flowchart of a method for processing data based on a graph database according to an embodiment is shown. It can be understood that the method can be executed by any device, equipment, platform, or device cluster with computing and processing capabilities. Figure 2 As shown, the method for processing data based on a graph database may include the following steps:
[0035] Step 201, obtaining interaction event data.
[0036] In this embodiment, the execution subject of the method for executing the data processing based on the graph database can obtain the interaction event data. Here, the interaction event may refer to various events occurring on the online platform, including but not limited to product browsing events, transaction events, etc. The interaction event data may refer to various data related to the interaction event, including: the subject identification of the interaction subject, the interaction time, and several interaction object identifications corresponding to the target indicator to be accumulated and counted, etc. The target indicator to be accumulated and counted may refer to the type of data that you want to accumulate and count on the interaction subject, for example, the data of the user ID that interacts with the interaction subject, the number of network MAC addresses that interact with the interaction subject, and so on.
[0037] Taking the interactive event as a transaction event as an example, the interactive subject may be a merchant, and the subject identifier of the interactive subject may refer to the merchant's identifier, which may include but is not limited to: name, number, account number, etc. The target indicator to be accumulated and counted may refer to the type of data that the merchant wants to accumulate and count, for example, the number of user IDs transacted with the merchant, the number of network MAC addresses transacted with the merchant, etc. Taking the target indicator to be accumulated and counted as the number of user IDs transacted with the merchant as an example, the interactive object identifier corresponding to the target indicator may refer to the user ID.
[0038] In some optional implementations of this embodiment, the above-mentioned obtaining of interaction event data may specifically include: obtaining event data corresponding to a single interaction event as the interaction event data.
[0039] In this implementation, the execution subject can obtain the event data corresponding to a single interaction event as the interaction event data. As an example, the execution subject can first obtain various data related to a single interaction event, and then extract the required data from the various data obtained as the interaction event data. In practice, when the interaction event occurs infrequently, for example, when the number of occurrences per unit time does not exceed a preset number threshold, the event data corresponding to a single interaction event can be obtained each time as the interaction event data, which is used to deduplicate and update the statistical records in the graph database.
[0040] In some optional implementations of this embodiment, the acquisition of interaction event data may also be specifically performed as follows:
[0041] First, the event data corresponding to the interactive event generated by the stream is sent to the stream computing engine, and the stream computing engine aggregates the incoming event data according to the subject identifier at a preset time interval to obtain at least one aggregation result.
[0042] In this implementation, the execution subject can extract the required event data from various data of each interactive event obtained in real time, for example, extract the subject identification of the interactive subject of the interactive event, the interaction time in seconds, and the interactive object identification corresponding to the target indicator to be accumulated and counted. The event data corresponding to each interactive event extracted in real time can form streaming data. At this time, the event data corresponding to the interactive event generated by the stream can be sent to the streaming computing engine, which can aggregate the incoming event data according to the subject identification according to the preset time interval (for example, 1 minute) to obtain at least one aggregation result. Specifically, multiple event data containing the same subject identification in the same time interval are aggregated together to form an aggregation result. The aggregation result may include the subject identification of the interactive subject, the interaction time, and several interactive object identifications corresponding to the target indicator to be accumulated and counted. It can be understood that the level of the interaction time in the aggregation result can be determined according to the time interval. For example, when the time interval is 1 minute, the interaction time in the aggregation result is the interaction time at the minute level. When the time interval is 1 hour, the interaction time in the aggregation result is the interaction time at the hour level.
[0043] Then, each aggregation result is obtained from the stream computing engine as interaction event data.
[0044] Please refer to Figure 3 , Figure 3 A schematic diagram showing an example of using a streaming computing engine to generate aggregate results. Figure 3In the example, the event data corresponding to the interactive events generated by streaming include (U1, M1, 9:10:12), (U2, M1, 9:10:13), (U3, M1, 9:10:13), (U4, M2, 9:10:13), (U4, M2, 9:11:10), (U5, M2, 9:11:14), etc., where U1, U2, ... represent user IDs, M1, M2 represent merchant IDs, and the aggregation time interval is 1 minute. In this example, the aggregation results include {M1, [U1, U2, U3], 9:10:00}, {M2, [U4], 9:10:00}, {M2, [U4, U5], 9:11:00}. It can be understood that Figure 3 The event data, time intervals, etc. aggregated in the example are exemplary, and are not intended to limit the event data, time intervals, etc. It can be understood that the time in this example is Beijing time. In addition, the time in this example can also be expressed in other ways such as UNIX timestamps as needed.
[0045] In this implementation, the execution subject can obtain each aggregation result from the streaming computing engine as the interactive event data. Here, the streaming computing engine is used to aggregate the event data based on the subject identifier and then update the statistical records in the graph database, which can effectively solve the frequent concurrent writing problem in the Internet big data scenario. At the same time, by performing aggregation calculations at preset time intervals, it can be ensured that the cumulative value of the target indicator to be accumulated in the graph database can be updated at the time interval (for example, minutes) level, ensuring the timeliness of the update.
[0046] Step 202: Read the target statistical record of the interaction graph and subject identification for the target indicator from the graph database.
[0047] In this embodiment, an interaction graph may be stored in the graph database. In the interaction graph, the interaction subjects and interaction objects with interaction history are connected by edges, and the corresponding edge attributes include the most recent interaction time. Taking the interaction event as a transaction event as an example, the interaction subjects may include merchants M1, M2, and the interaction objects may include transaction users U1, U2, U3, U4, etc. The corresponding edge attributes may include the most recent transaction time t1, t2, t3, t4, etc., that is, the transaction time of the last transaction. Merchants M1, M2 and transaction users U1, U2, U3, U4, etc. with transaction history can form the following: Figure 4 The interaction diagram shown in the figure has the merchant identifiers M1, M2 and the transaction user identifiers U1, U2, U3, U4, ... as vertices. It can be understood that Figure 4The interaction graph in is only used to describe the structure of the interaction graph, rather than to limit the data content contained in the interaction graph. For example, according to actual needs, the attributes of the edge can also include data such as transaction amount, transaction location (for example, shipping address, geographical location of the user when placing an order, etc.).
[0048] The graph database may also store target statistical records for target indicators with subject identifiers. The target statistical records may include multiple record items corresponding to multiple time periods. A single record item may include the cumulative values of the target indicators in each time period of a predetermined number N time periods, starting from the corresponding time period. Here, the length of the time period can be set according to actual statistical needs. For example, the time period can be set to 1 minute, 1 hour, 1 day, 1 week, and so on. Taking the interactive event as a transaction event, the subject identifier as M, the target indicator as the number of transaction users, and the time period as 1 hour as an example, the following target statistical record can be formed:
[0049] Th V(Th)V(Th-1)V(Th-2)…………V(Th-N);
[0050] Th-1 V(Th-1)V(Th-2)V(Th-3)…………V(Th-N-1);
[0051] Th-2 V(Th-2)V(Th-3)V(Th-4)…………V(Th-N-2);
[0052] Th-3 V(Th-3)V(Th-4)V(Th-5)…………V(Th-N-3)
[0053] …
[0054] …….
[0055] Each row of data in the above example can represent a record item. Among them, Th represents the hourly time period, which can be generated according to the timestamp of the interaction time of the interaction event. For example, the current time is "2021-06-1714:12:07" (UNIX timestamp is 1623910327), then the hourly time period Th in which the current time is located is Th=abs(1623910327 / 3600)=451086. Among them, the abs function is a function used to find the absolute value of data. Th-1 represents the hourly time period one hour before Th, Th-2 represents the hourly time period two hours before Th, and Th-3 and so on. V(Th) represents the cumulative value of the number of trading users in the Th time period, that is, how many trading users have traded with M in the Th time period. V(Th-1) represents the cumulative value of the number of trading users in the Th-1 time period. V(Th-2), V(Th-3), and so on. It can be understood that the time period preceding each record item in the above target statistical record is the time period of the starting point of the record item, which can be used to read data from the graph database later. For example, it can be used as part of the row key of the record item.
[0056] It can be understood that the larger the number of record items in the target statistical record and the number N of time periods traced back in each record item, the more cumulative values can be stored, and the larger the statistical time window that can be selected for statistical analysis, but the larger the storage space occupied. In practice, the number of record items and N can be set according to actual statistical needs. For example, a maximum value can be set for the number of record items and N. For example, the maximum value of the number of record items can be set to 720, that is, 30 (days) * 24 = 720 (hours). The value of N can be set to 719, so that, plus the time period of the starting point, a single record can include the cumulative value of 720 time periods.
[0057] In some optional implementations, the target statistical record may be in the form of a numerical matrix, wherein a single record item in the target statistical record corresponds to a row of the matrix.
[0058] In one embodiment, the graph database may adopt distributed storage. For example, RocksDB, HBase, GeaBase and other distributed storage systems are used for storage. Therefore, in order to facilitate distributed storage, the statistical records in the graph database are generally stored in multiple rows, and each row represents a record item. In order to facilitate data search, each record item in the statistical records stored in the graph database may correspond to a row key, and the row key may include a subject identifier, an indicator, a starting time period, and the like. In this way, the execution subject may read the record items whose row keys match the obtained subject identifier and target indicator from the graph database according to the subject identifier and target indicator of the interactive subject included in the interactive event data obtained in step 201 to form a target statistical record.
[0059] Step 203 , for each interactive object identifier among the plurality of interactive object identifiers, an update operation is performed on the graph database, and the update operation specifically includes the following steps 2031 and 2032 .
[0060] Specifically, in step 2031, a target record item corresponding to the interaction time is determined in the target statistical record.
[0061] In this embodiment, according to the time period of the starting point of each record item in the target statistical record, a record item whose starting point time period corresponds to the interaction time of the interaction event data is determined as the target record item. Specifically:
[0062] First, it may be determined whether there is a record item in the target statistical record whose time period of the starting point includes the interaction time.
[0063] If it exists, the existing record item is determined as the target record item.
[0064] If it does not exist, a new record item is added to the target statistical record as the target record item, wherein the newly added record item is used to record, starting from the time period corresponding to the above interaction time, and tracing back to each cumulative value of the above target indicator in each time period of N time periods. Through this implementation method, the target record item can be determined. At the same time, when there is no record item containing the interaction time in the time period of the starting point in the target statistical record, a new record item can be added to realize the generation of the record item.
[0065] In one embodiment, the above step of adding a new record item in the target statistical record as the target record item may be specifically performed as follows:
[0066] First, a new record item to be filled is generated. As an example, the new record item can be a blank row of data.
[0067] Secondly, in the newly created record item, the cumulative value corresponding to the time period of the starting point is initialized to 0.
[0068] Finally, in the latest record item that exists before the new record item is created, the N cumulative values corresponding to the previous N time periods are determined, and the accumulated values are copied to the new record item to obtain the target record item. Through this embodiment, when there is no record item in the target statistical record whose time period of the starting point contains the interaction time of the interaction event data, a new record item can be added to the target statistical record to realize the generation of the record item.
[0069] Step 2032: perform deduplication update on the target record item according to the edge attribute between the subject identifier and the interaction object identifier in the interaction graph.
[0070] In this embodiment, the execution subject can perform deduplication update on the target record item according to the edge attribute between the subject identifier and the interaction corresponding identifier in the above interaction graph, so that the cumulative value of the target indicator stored in each time period in the target record item is the deduplication value.
[0071] In one embodiment, Figure 5 As shown, the above step of performing deduplication and updating of the target record item according to the edge attribute between the subject identifier and the interaction object identifier in the interaction graph can be specifically performed as follows:
[0072] Step 501, adding 1 to the cumulative value corresponding to the starting time period of the target record item.
[0073] Step 502: determine whether there is an edge with the subject identifier and the interaction object identifier in the interaction graph.
[0074] Step 503: If the edge of the subject identifier and the interactive object identifier exists, the cumulative value in the time period corresponding to the most recent interaction time included in the edge attribute is reduced by 1, and the most recent interaction time of the edge attribute is updated to the above interaction time. This achieves deduplication of the cumulative value in the target record item and update of the most recent interaction time in the edge attribute of the subject identifier and the interactive object identifier.
[0075] Step 504: If the edge between the subject identifier and the interaction object identifier does not exist, then an edge formed between the subject identifier and the interaction object identifier is added to the interaction graph, wherein the edge attribute of the added edge includes the interaction time. Thus, when the edge between the subject identifier and the interaction object identifier does not exist in the interaction graph, an edge between the subject identifier and the interaction object identifier is added to the interaction graph.
[0076] For example, assuming that the interaction event data is {M1, [U1, U2, U3], 451086: 01}, the current target statistics record is as follows:
[0077] 451085 V(451085:00)V(451084:00)V(451083:00)…V(451085:00-N);
[0078] 451084 V(451084:00)V(451083:00)……………………V(451084:00-N);
[0079] 451082 V(451082:00)…………………………………………V(451082:00-N).
[0080] For U1, in the target statistics record, in step 2031, no target record item is found whose starting time period includes the interaction time 451086:01. Here, the time in the target record item is generated based on the UNIX timestamp. Therefore, a new target record item is added. After the addition, the target statistics record is as follows:
[0081] 451086 V(451086:00)V(451085:00)V(451084:00)…V(451086:00-N);
[0082] 451085 V(451085:00)V(451084:00)V(451083:00)…V(451085:00-N);
[0083] 451084 V(451084:00)V(451083:00)……………………V(451084:00-N);
[0084] 451082 V(451082:00)…………………………………………V(451082:00-N).
[0085] In step 2032, V(451086:00) in the target record item "V(451086:00)V(451085:00)V(451084:00) ... V(451086:00-N)" is increased by 1. In addition, assuming that there is an interaction history between U1 and M in the interaction graph, and the most recent interaction time occurs at 451084:30, the value of V(451084:00) in the target record item is reduced by 1. At this time, the target record item is as follows:
[0086] V(451086:00)+1 V(451085:00)V(451084:00)-1…V(451086:00-N).
[0087] In addition, the most recent interaction time between U1 and M in the interaction graph is updated to 451086:01.
[0088] Next, for U2, in step 2031, there is a target record item (the newly added one), and in step 2032, V (451086:00) is continued to be increased by 1. Assuming that there is no interaction history between U2 and M, an edge is added to the interaction graph, and the edge attribute is recorded as 451086:01.
[0089] Then for U3, in step 2031, there is a target record item, and in step 2032, V(451086:00) is continued to be added by 1. Assuming that there is an interaction history between U3 and M, and the most recent interaction time occurs at 451084:28, the value of V(451084:00) in the target record item is reduced by 1, and the most recent interaction time between U3 and M in the interaction graph is updated to 451086:01. The target record item is finally obtained as:
[0090] V(451086:00)+3 V(451085:00)V(451084:00)-2…V(451086:00-N).
[0091] In one embodiment, the above update operation performed on the graph database may also include: Figure 2 Step 2033 is not shown in FIG.
[0092] Step 2033, determine whether the number of record items included in the statistical record after deduplication update exceeds the maximum number threshold, if exceeded, delete the record item with the earliest starting time period in the updated statistical record.
[0093] In this embodiment, since the storage space is limited and new record items may be created during the deduplication update process, in order to ensure that the number of record items in the statistical record is not too large, it can be determined after the deduplication update whether the number of record items included in the statistical record after the deduplication update exceeds the maximum number threshold. If it exceeds, the earliest record item in the time period of the starting point of the updated statistical record is deleted. It can be understood that the above-mentioned maximum number threshold can be determined according to the size of the storage space.
[0094] In one embodiment, the above method for processing data based on a graph database may further include: in response to determining that the update of the above several interactive object identifiers is completed, writing the updated interaction graph and target statistical records back to the above graph database.
[0095] In this embodiment, if it is determined that the updating of several interaction object identifiers in the interaction event data is complete, the updated interaction graph and target statistical records can be written back to the above graph database, thereby realizing the updating of the graph database.
[0096] The above-mentioned embodiment of this specification provides a method for processing data based on a graph database. When using interactive event data to update statistical records, the cumulative values of record items in the statistical records are deduplicated based on the edge attributes between the subject identifier and the interactive object identifier in the interactive graph, so that the cumulative values in the record items are the deduplicated updated values, so that the cumulative values of each time period obtained when performing statistical analysis based on the graph database are the deduplicated data, so that the statistical results are more accurate. In addition, the graph database of this embodiment stores the interactive graph and statistical records, and does not need to store the detailed data of the interactive events, thereby saving storage space, reducing the time for data search and calculation, and improving data processing efficiency.
[0097] Continue to see Figure 6 The method for processing data based on a graph database in the embodiment of the present application may further include the following step of performing statistical analysis on the data based on the graph database:
[0098] Step 601, determining data statistical parameters.
[0099] In this embodiment, the execution subject may receive a data statistics request input by other devices or users, and determine data statistics parameters. Here, the data statistics parameters may include an identifier of the interactive subject to be analyzed, an indicator to be analyzed, and a time window to be statistically analyzed. As an example, the time window to be statistically analyzed may include a start time and an end time. It can be understood that the time level of the start time and the end time is not greater than the level of the time period in the record item. Taking the level of the time period in the record item as hours as an example, the time level of the start time and the end time may be seconds, minutes, hours, etc.
[0100] Step 602: Read statistical records from the graph database as statistical records to be analyzed according to the identifier of the interaction subject to be analyzed and the indicator to be analyzed.
[0101] In this embodiment, the execution subject may read the statistical records of the interaction subject identifier to be analyzed for the indicator to be analyzed from the graph database as the statistical records to be analyzed according to the interaction subject identifier to be analyzed and the indicator to be analyzed included in the data statistical parameters.
[0102] Step 603: Determine the record items to be analyzed from the statistical records to be analyzed according to the time window.
[0103] In this embodiment, a record item can be determined from the above-mentioned statistical record to be analyzed as a record item to be analyzed according to the time window in the data statistical parameter. As an example, the record item to be analyzed can be determined from the statistical record to be analyzed according to the start time and the end time of the time window. For example, a record item whose start time period is greater than the end time period can be read from the statistical record to be analyzed as the record item to be analyzed. Here, the end time period can refer to the time period corresponding to the end time of the above-mentioned time window.
[0104] In some optional implementations, the above step 603 may be specifically performed as follows:
[0105] First, determine whether there is a record item with the end time period as the starting point in the statistical record to be analyzed.
[0106] If there is no record item with the end time period as the starting point, a record item with a starting time period greater than the end time period is read from the above-mentioned statistical record to be analyzed as the record item to be analyzed. As an example, a record item with a starting time period greater than the end time period and a minimum starting time period can be read from the statistical record to be analyzed as the record item to be analyzed.
[0107] If there is a record item with the end time period as the starting point, the record item with the end time period as the starting point is read from the above-mentioned statistical record to be analyzed as the record item to be analyzed. Through this implementation method, it can be ensured that the starting point of the determined record item to be analyzed is equal to or greater than the end time period, thereby ensuring that the data corresponding to the end time period is stored in the record item to be analyzed.
[0108] Step 604: read the data corresponding to the time window from the record item to be analyzed, and perform statistical analysis on it to obtain a statistical analysis result.
[0109] In this embodiment, data corresponding to the time window can be read from the record item to be analyzed, and various statistical analyses can be performed on it, such as summing, averaging, median, etc., to obtain statistical analysis results.
[0110] In some optional implementations, the above step 604 can be specifically performed as follows: read from the above record items to be analyzed, the various cumulative values in each time period from the start time period corresponding to the above time window to its end time period, accumulate the above cumulative values, and obtain the above statistical analysis results.
[0111] For example, assuming that the time window to be statistically analyzed is [451084:00, 451085:00], and according to the interaction subject identifier to be analyzed and the indicator to be analyzed in the data statistical parameters, the following statistical records are read from the graph database as the statistical records to be analyzed:
[0112] 451086 V(451086:00)V(451085:00)V(451084:00)…V(451086:00-N);
[0113] 451084 V(451084:00)V(451083:00)……………………V(451084:00-N);
[0114] 451082 V(451082:00)…………………………………………V(451082:00-N).
[0115] Since there is no record item with the end time period 451085 as the starting point in the statistical record to be analyzed, the record item "V(451086:00)V(451085:00)V(451084:00)...V(451086:00-N)" with a starting time period greater than the end time period 451085 is read as the record item to be analyzed. The data V(451085:00) and V(451084:00) corresponding to the time window are read from the record item to be analyzed, and statistical analysis is performed on them to obtain statistical analysis results.
[0116] In this implementation, the above-mentioned starting time period may refer to the time period corresponding to the starting time of the time window. In this way, the cumulative values in each time period from the starting time period corresponding to the time window to its ending time period are read from the above-mentioned record item to be analyzed, and the read cumulative values are accumulated to obtain the summed statistical analysis result. Since the cumulative values stored in the record item to be analyzed are deduplicated values, the summed statistical analysis result obtained in this implementation is a deduplicated result, which is more accurate.
[0117] In some optional implementations, the above step of performing statistical analysis on the data based on the graph database may also include: Figure 6 Step 605 is not shown in FIG.
[0118] Step 605: Feedback the statistical analysis results.
[0119] In this implementation, the statistical analysis results generated in step 604 can be fed back. For example, the statistical analysis results can be fed back to a pre-specified device, or fed back to a device that sends a data statistics request, so that the user can view the statistical analysis results. Through this implementation, the statistical analysis results can be fed back so that the user can view the statistical analysis results.
[0120] In this embodiment, since the cumulative value of the record items stored in the graph database is the deduplicated updated value, the cumulative value obtained when performing statistical analysis based on the graph database is the deduplicated data, so the obtained statistical result is more accurate.
[0121] According to another embodiment, a device for processing data based on a graph database is provided. The device for processing data based on a graph database can be deployed in any device, platform or device cluster with computing and processing capabilities.
[0122] Figure 7 FIG. 1 is a schematic block diagram of an apparatus for processing data based on a graph database according to an embodiment. Figure 7 As shown, the device 700 for processing data based on the graph database includes: an acquisition unit 701, configured to acquire interaction event data, including a subject identifier of an interaction subject, an interaction time, and a plurality of interaction object identifiers corresponding to the target indicator to be accumulated and counted; a reading unit 702, configured to read an interaction graph and a target statistical record of the target indicator for the subject identifier from the graph database, in which the interaction subject and the interaction object with an interaction history are connected by edges, and the corresponding edge attributes include the most recent interaction time; the target statistical record includes a plurality of record items corresponding to a plurality of time periods, and a single The record items include the cumulative values of the above-mentioned target indicators in each time period of a predetermined number N time periods traced back forward with the corresponding time period as the starting point; the operation unit 703 is configured to perform an update operation on the above-mentioned graph database for each interactive object identifier among the above-mentioned several interactive object identifiers, wherein the above-mentioned operation unit 703 includes: a determination module 7031, configured to determine the target record item corresponding to the above-mentioned interaction time in the above-mentioned target statistical record; a deduplication update module 7032, configured to perform deduplication update on the above-mentioned target record item according to the edge attribute between the above-mentioned subject identifier and the interactive object identifier in the above-mentioned interaction graph.
[0123] In some optional implementations of this embodiment, the acquisition unit 701 is further configured to: acquire event data corresponding to a single interaction event as the interaction event data.
[0124] In some optional implementations of this embodiment, the acquisition unit 701 is further configured to: send the event data corresponding to the interactive event generated by the streaming to the streaming computing engine, the streaming computing engine aggregates the incoming event data according to the subject identifier at a preset time interval to obtain at least one aggregation result; and obtain each aggregation result from the streaming computing engine as the interactive event data.
[0125] In some optional implementations of the present embodiment, the above-mentioned determination module 7031 includes: a first determination submodule (not shown in the figure), configured to determine whether there is a record item in the above-mentioned target statistical record that contains the above-mentioned interaction time in the starting time period; a new addition submodule (not shown in the figure), configured to add a new record item as the target record item in the above-mentioned target statistical record if it does not exist, wherein the newly added record item is used to record, with the time period corresponding to the above-mentioned interaction time as the starting point, and each cumulative value of the above-mentioned target indicator in each time period of N time periods backtracked; a second determination submodule (not shown in the figure), configured to determine the existing record item as the target record item if it exists.
[0126] In some optional implementations of this embodiment, the above-mentioned new sub-module is further configured to: generate a new record item to be filled; in the above-mentioned new record item, initialize the cumulative value corresponding to the starting time period to 0; in the latest record item that exists before the above-mentioned new record item, determine the N cumulative values corresponding to the first N time periods, and copy them to the above-mentioned new record item to obtain the above-mentioned target record item.
[0127] In some optional implementations of the present embodiment, the deduplication update module 7032 is further configured to: add 1 to the cumulative value corresponding to the starting time period of the target record item; determine whether there is an edge between the subject identifier and the interactive object identifier in the interaction graph; if the edge exists, subtract 1 from the cumulative value in the time period corresponding to the most recent interaction time included in the edge attribute; if the edge does not exist, add an edge formed between the subject identifier and the interactive object identifier in the interaction graph, wherein the edge attributes of the added edge include the interaction time.
[0128] In some optional implementations of this embodiment, the above-mentioned operation unit 703 also includes: a deletion module (not shown in the figure), configured to determine whether the number of record items included in the statistical record after deduplication update exceeds the maximum number threshold, and if so, delete the earliest record item in the starting time period of the updated statistical record.
[0129] In some optional implementations of this embodiment, the above-mentioned device 700 also includes: a write-back unit (not shown in the figure), configured to write the updated interaction graph and target statistical records back to the above-mentioned graph database in response to determining that the update of the above-mentioned several interactive object identifiers is completed.
[0130] In some optional implementations of this embodiment, the above-mentioned target statistical record is a numerical matrix, and a single record item corresponds to a row of the matrix.
[0131] In some optional implementations of the present embodiment, the above-mentioned device 700 also includes: a parameter determination unit (not shown in the figure), configured to determine data statistical parameters, wherein the above-mentioned data statistical parameters include: an identifier of the interactive subject to be analyzed, an indicator to be analyzed and a time window for statistical analysis; a record reading unit (not shown in the figure), configured to read statistical records from the above-mentioned graph database as statistical records to be analyzed based on the above-mentioned identifier of the interactive subject to be analyzed and the above-mentioned indicator to be analyzed; a record item determination unit to be analyzed (not shown in the figure), configured to determine the record items to be analyzed from the above-mentioned statistical records to be analyzed based on the above-mentioned time window; a statistical analysis unit (not shown in the figure), configured to read data corresponding to the above-mentioned time window from the above-mentioned record items to be analyzed, and perform statistical analysis on the data to obtain statistical analysis results.
[0132] In some optional implementations of the present embodiment, the above-mentioned record item determination unit is further configured to: determine whether there is a record item with an end time period as a starting point in the above-mentioned statistical record to be analyzed, wherein the above-mentioned end time period is a time period corresponding to the end time of the above-mentioned time window; if not, read a record item with a starting time period greater than the above-mentioned end time period from the above-mentioned statistical record to be analyzed as the record item to be analyzed; if exists, read a record item with the above-mentioned end time period as a starting point from the above-mentioned statistical record to be analyzed as the record item to be analyzed.
[0133] In some optional implementations of the present embodiment, the statistical analysis unit is further configured to: read from the record items to be analyzed, the cumulative values within each time period from the start time period corresponding to the time window to its end time period, add up the cumulative values, and obtain the statistical analysis results.
[0134] In some optional implementations of this embodiment, the apparatus 700 further includes: a feedback unit (not shown in the figure), configured to feed back the statistical analysis results.
[0135] According to another embodiment, there is also provided a computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute Figure 2 The method described.
[0136] According to another embodiment, there is also provided a computing device, comprising a memory and a processor, wherein the memory stores an executable code, and when the processor executes the executable code, Figure 2 The method described.
[0137] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0138] Those of ordinary skill in the art should further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in terms of function in the above description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those of ordinary skill in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0139] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented by hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0140] The specific implementation methods described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific implementation method of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for processing data based on a graph database, comprising: Acquire interaction event data, including the subject identifier of the interaction subject, the interaction time, and several interaction object identifiers corresponding to the target indicator to be accumulated and counted; Reading an interaction graph and a target statistical record of the subject identifier for the target indicator from a graph database, wherein in the interaction graph, an interaction subject and an interaction object having an interaction history are connected by an edge, and the corresponding edge attribute includes the most recent interaction time; the target statistical record includes a plurality of record items corresponding to a plurality of time periods, and a single record item includes each cumulative value of the target indicator in each time period of a predetermined number N of time periods traced back from the corresponding time period as a starting point; For each interactive object identifier among the plurality of interactive object identifiers, performing an update operation on the graph database, the update operation comprising: determining a target record item corresponding to the interaction time in the target statistical record; The target record item is updated to deduplicate according to the edge attribute between the subject identifier and the interaction object identifier in the interaction graph.
2. The method according to claim 1, wherein: The obtaining of interaction event data includes: Event data corresponding to a single interaction event is obtained as the interaction event data.
3. The method according to claim 1, wherein: The obtaining of interaction event data includes: Sending event data corresponding to the interactive event generated by the streaming to the streaming computing engine, the streaming computing engine aggregates the incoming event data according to the subject identifier at a preset time interval to obtain at least one aggregation result; Each aggregation result is obtained from the stream computing engine as interaction event data.
4. The method according to claim 1, wherein: The determining of the target record item corresponding to the interaction time in the target statistical record includes: Determine whether there is a record item in the target statistical record whose time period of the starting point includes the interaction time; If it exists, the existing record item is determined as the target record item; If it does not exist, a new record item is added to the target statistical record as the target record item, wherein the newly added record item is used to record the cumulative values of the target indicators in each time period of N time periods, starting from the time period corresponding to the interaction time.
5. The method according to claim 4, wherein: The adding a new record item in the target statistical record as a target record item includes: Generate a new record item to be filled; In the newly created record item, the accumulated value corresponding to the time period of the starting point is initialized to 0; In the latest record item that exists before the newly created record item, N accumulated values corresponding to the previous N time periods are determined, and the accumulated values are copied to the newly created record item to obtain the target record item.
6. The method according to claim 1, wherein the step of performing deduplication updating on the target record item according to the edge attribute between the subject identifier and the interaction object identifier in the interaction graph comprises: Add 1 to the cumulative value corresponding to the starting time period of the target record item; Determine whether there is an edge between the subject identifier and the interaction object identifier in the interaction graph; If the edge exists, the accumulated value in the time period corresponding to the most recent interaction time included in the edge attribute is reduced by 1, and the most recent interaction time of the edge attribute is updated to the interaction time; If the edge does not exist, an edge formed between the subject identifier and the interaction object identifier is added to the interaction graph, wherein the edge attribute of the added edge includes the interaction time.
7. The method according to claim 1, wherein: The updating operation further includes: Determine whether the number of record items included in the statistical record after deduplication update exceeds the maximum number threshold. If it exceeds, delete the earliest record item in the starting time period of the updated statistical record.
8. The method according to claim 1, wherein: The method further comprises: In response to determining that the updating of the plurality of interaction object identifiers is completed, the updated interaction graph and target statistical records are written back to the graph database.
9. The method according to claim 1, wherein: The target statistical record is a numerical matrix, and a single record item corresponds to a row of the matrix.
10. The method according to claim 1, wherein: The method further comprises: Determining data statistical parameters, wherein the data statistical parameters include: an identifier of an interaction subject to be analyzed, an indicator to be analyzed, and a time window for statistical analysis; Reading statistical records from the graph database as statistical records to be analyzed according to the interaction subject identifier to be analyzed and the indicator to be analyzed; Determining the record items to be analyzed from the statistical records to be analyzed according to the time window; The data corresponding to the time window is read from the record item to be analyzed, and statistical analysis is performed on the data to obtain a statistical analysis result.
11. The method according to claim 10, wherein: The step of determining the record items to be analyzed from the statistical records to be analyzed according to the time window includes: Determine whether there is a record item starting from an end time period in the statistical record to be analyzed, wherein the end time period is a time period corresponding to the end time of the time window; If it exists, read the record item starting from the end time period from the statistical record to be analyzed as the record item to be analyzed; If it does not exist, a record item whose starting time period is greater than the ending time period is read from the statistical record to be analyzed as the record item to be analyzed.
12. The method according to claim 10, wherein: Reading data corresponding to the time window from the record item to be analyzed, and performing statistical analysis on the data to obtain a statistical analysis result, including: The accumulated values in each time period from the start time period to the end time period corresponding to the time window are read from the record item to be analyzed, and the accumulated values are accumulated to obtain the statistical analysis result.
13. The method according to claim 10, wherein: The method further comprises: The statistical analysis results are fed back.
14. A device for processing data based on a graph database, comprising: An acquisition unit is configured to acquire interaction event data, including a subject identifier of an interaction subject, interaction time, and a plurality of interaction object identifiers corresponding to a target indicator to be accumulated and counted; A reading unit is configured to read an interaction graph and a target statistical record of the subject identification for the target indicator from a graph database, wherein in the interaction graph, an interaction subject and an interaction object having an interaction history are connected by an edge, and the corresponding edge attribute includes a most recent interaction time; the target statistical record includes a plurality of record items corresponding to a plurality of time periods, and a single record item includes each cumulative value of the target indicator in each time period of a predetermined number N of time periods traced back from the corresponding time period as a starting point; An operation unit is configured to perform an update operation on the graph database for each interactive object identifier among the several interactive object identifiers, wherein the operation unit includes: a determination module, configured to determine a target record item corresponding to the interaction time in the target statistical record; and a deduplication update module, configured to perform deduplication update on the target record item according to an edge attribute between the subject identifier and the interactive object identifier in the interaction graph.
15. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 13.
16. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 13 is implemented.
Citation Information
Patent Citations
Service index obtaining method and device, server and computer readable storage medium
CN109800225A
Incremental data processing
US20150347439A1