Graph database updating method, program product, electronic equipment and storage medium
By obtaining and comparing the version information of the data to be updated in the graph database, the problems of data coverage and inconsistency are solved, and higher data consistency and refined update management are achieved.
Patent Information
- Application Number
- CN202510140953.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-08
- Publication Date
- 2025-05-27
AI Technical Summary
Existing graph database update methods may cause the later updated data to be overwritten by the first updated data, resulting in data inconsistency and overwriting problems.
By obtaining the identification information and version information of the data to be updated, querying historical version information, and updating the graph database according to the order of version information, ensuring data consistency and reducing coverage and conflicts.
Improve data consistency in the graph database, reduce data coverage and conflicts, and reduce data errors and inconsistencies caused by updates.
Smart Images

Figure CN120045571A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of graph databases. Specifically, it relates to a graph database update method, a program product, an electronic device, and a storage medium. Background Art
[0002] A graph database is a database management system that uses a graph structure to store and query data. It provides a flexible data model and a powerful query language, and can efficiently execute complex relational queries. In the actual application process, as the business develops or the relationships between entities change, it may be necessary to update the data in the database. The current update method can record the content of each update of the graph database, but this data update method may cause the problem that the data updated later is overwritten by the data updated earlier. Summary of the Invention
[0003] The purpose of the embodiments of this application is to provide a graph database update method, a program product, an electronic device, and a storage medium to improve the above technical problems.
[0004] In a first aspect, the embodiments of this application provide a graph database update method, including: obtaining the identification information and the version information to be updated of the data to be updated; the data to be updated includes the vertex data, edge data, or attribute data to be updated in the graph database; the version information to be updated is used to identify the update order of the data to be updated; querying the historical version information of the data to be updated according to the identification information; the historical version information is used to represent the versions that the data to be updated has already stored in the graph database; comparing the version information to be updated with the historical version information, and updating the graph database according to the comparison result.
[0005] In the above implementation process, obtaining the version information to be updated of the vertex data, edge data, or attribute data, and first comparing the version information and then updating, improves the data consistency in the graph database, reduces data overwriting and conflicts, and reduces data errors and inconsistencies caused by updates. And setting version information separately for each attribute field, rather than only controlling the version information of the entire vertex (node) or edge (relationship), realizes a finer-grained control of the update of the graph database. For example, if only one field changes, only the version number of this field will increase, and there is no need to update the version number of the entire node or edge, thereby reducing the amount of update operations.
[0006] Optionally, in the embodiments of the present application, obtaining the identification information and the version information to be updated of the data to be updated includes: receiving the identification information and the version information to be updated of multiple pieces of data to be updated; according to the identification information of the data to be updated, writing the multiple pieces of data to be updated into the corresponding partitions of the message queue respectively; the message queue is partitioned in advance using the identification information of the data to be updated as the partition key value, so that the data to be updated with the same identification information is written into the same partition; obtaining the identification information and the version information to be updated of the data to be updated from the corresponding partitions of the message queue.
[0007] In the above implementation process, by partitioning the message queue and using the identification information as the partition key, the data to be updated with the same identification information can be divided into the same partition for processing in order, reducing the problem of data inconsistency caused by concurrent updates when modifying the data corresponding to the same identification information at the same time. And after using partitioning, the message queue can handle high-concurrency data update requests and adapt to the growing data update requirements.
[0008] Optionally, in the embodiments of the present application, according to the identification information of the data to be updated, writing the multiple pieces of data to be updated into the corresponding partitions of the message queue respectively includes: inputting the identification information of the data to be updated into a pre-trained version conflict recognition model to obtain the version conflict probability of the data to be updated; the version conflict probability is used to characterize the possibility of version conflict of the data to be updated; the version conflict recognition model is used to learn according to the identification information, obtain the version features related to the identification information, and obtain the version conflict probability of the data to be updated based on the version features; if the version conflict probability is greater than the first preset threshold, write the data to be updated into the corresponding partition of the message queue.
[0009] In the above implementation process, writing the high-risk data to be updated into the corresponding partition of the message queue, so that these high-risk data to be updated can be processed in the order of entering the message queue, avoiding the problem of high concurrency. For other data that is not high-risk, it can be directly updated to reasonably use the resources of the message queue, reduce resource consumption, and improve resource utilization. Different-risk data are processed respectively, realizing more refined management of data updates, reducing the problem of version conflicts while increasing the throughput of data processing.
[0010] Optionally, in the embodiments of the present application, before inputting the identification information of the data to be updated into a pre-trained version conflict recognition model, the method further includes: obtaining the historical update data of the graph database; the historical update data includes the identification information and the update timestamp of the historically updated graph data; using a preset time window to extract the historical version features for influencing version conflict from the historical update data; training the machine learning initial model using the historical version features to obtain the version conflict recognition model.
[0011] In the above implementation process, the solution of using machine learning to predict version conflicts can help identify and prevent potential version conflict problems in advance, reduce update anomalies, and improve data consistency and business continuity. The time window used during the training process of the machine learning model can help the model capture time-related historical version features, improving the prediction accuracy and practicality of the model. The model effectively utilizes the identification information during training, and internally searches for or calculates related features. When the probability of predicting version conflicts is accurate and reliable, the model input is simplified and the efficiency is improved.
[0012] Optionally, in the embodiments of the present application, obtaining the identification information and the version information to be updated of the data to be updated from the partitions corresponding to the message queue includes: for different partitions in the message queue, concurrently executing the step of obtaining the identification information and the version information to be updated of the data to be updated; for the same partition in the message queue, sequentially executing the step of obtaining the identification information and the version information to be updated of the data to be updated in the order in which the data to be updated enters the corresponding partition of the message queue.
[0013] In the above implementation process, the partitions of the message queue can improve the scalability and parallel processing ability of the message queue. By dispersing the data into different partitions, data that will not cause concurrent problems such as version conflicts can work in parallel; while the data with the same identification information is processed in order, reducing the possibility of version conflicts, achieving efficient processing of a large number of messages, and improving the consistency and order of the data at the same time.
[0014] Optionally, in the embodiments of the present application, comparing the version information to be updated with the historical version information and updating the graph database according to the comparison result includes: if the identification information of the version information to be updated does not exist in the graph database, performing an update operation and writing the data to be updated into the graph database; if the identification information of the version information to be updated exists in the graph database, comparing the version information to be updated with the historical version information, and if the time order of the historical version information is before the version information to be updated, performing an update operation and updating the graph database with the data to be updated; if the time order of the historical version information is after the version information to be updated, discarding the data to be updated.
[0015] In the above implementation process, by comparing the version information to be updated with the historical version information, the data consistency and integrity of the graph database are maintained, so that the data stored in the database is the latest, realizing fine-grained management of the point data, edge data or attribute data in the graph database. And by the method of updating after comparison, only valid data (the data of the latest version) needs to be written, reducing the amount of data written and the update cost.
[0016] Optionally, in the embodiments of the present application, updating the graph database according to the comparison result includes: if the comparison result indicates that an update operation needs to be performed, obtaining the original data corresponding to the identification information from the graph database according to the identification information, and using the data to be updated to overwrite the original data to implement the update of the graph data.
[0017] In the above implementation process, the overwrite operation makes the data stored in the graph database always up-to-date, and can reduce data redundancy and inconsistency, improving the accuracy and reliability of the data. There is no need to store multiple versions of the same data, thus optimizing the use of storage space, and the overwrite operation simplifies the management complexity of the graph database.
[0018] In a second aspect, the embodiments of the present application further provide a graph database update device, including: an acquisition data module, configured to acquire the identification information and the version information to be updated of the data to be updated; the data to be updated includes the point data, edge data or attribute data to be updated in the graph database; the version information to be updated is used to identify the update order of the data to be updated; a version information acquisition module, configured to query the historical version information of the data to be updated according to the identification information; the historical version information is used to represent the version of the data to be updated that has been stored in the graph database; a comparison and update module, configured to compare the version information to be updated with the historical version information, and update the graph database according to the comparison result.
[0019] In a third aspect, the embodiments of the present application further provide a computer program product, including computer program instructions, and when the computer program instructions are run by a processor, the method provided in the first aspect or any one implementation manner of the first aspect is executed.
[0020] In a fourth aspect, the embodiments of the present application further provide an electronic device, including: a processor and a memory, the memory stores computer program instructions, and when the computer program instructions are run by the processor, the method provided in the first aspect or any one implementation manner of the first aspect is executed.
[0021] In a fifth aspect, the embodiments of the present application further provide a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are run by a processor, the method provided in the first aspect or any one implementation manner of the first aspect is executed.
[0022] By adopting a graph database update method, program product, electronic device, and storage medium provided by this application, the to-be-updated version information of vertex data, edge data, or attribute data is obtained. By comparing the version information first and then performing the update, the data consistency in the graph database is improved, data overwriting and conflicts are reduced, and data errors and inconsistencies caused by updates are reduced. Moreover, version information is set separately for each attribute field instead of only controlling the version information of the entire vertex (node) or edge (relationship), enabling more fine-grained control over the update of the graph database. For example, if only one field changes, only the version number of this field will increase, and there is no need to update the version number of the entire node or edge, thereby reducing the amount of update operations. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions of the embodiments of this application, the following will briefly introduce the drawings required to be used in the embodiments of this application. It should be understood that the following drawings only show some embodiments of this application and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 A flowchart of a graph database update method provided by an embodiment of this application;
[0025] Figure 2 A structural diagram of a graph database update device provided by an embodiment of this application;
[0026] Figure 3 A structural diagram of an electronic device provided by an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0027] The following will describe in detail the embodiments of the technical solutions of this application with reference to the drawings. The following embodiments are only used to more clearly illustrate the technical solutions of this application and thus are only examples and should not be used to limit the protection scope of this application.
[0028] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs; the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit this application.
[0029] In the description of the embodiments of this application, technical terms such as "first" and "second" are only used to distinguish different objects and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity, specific order, or primary-secondary relationship of the indicated technical features. In the description of the embodiments of this application, "a plurality of" means two or more unless otherwise specifically defined.
[0030] In the current graph database Nebula, version control for a specific vertex or edge is not supported. Instead, the logic of "the later update takes effect" is adopted. Without version control, subsequent updates may overwrite previous data, resulting in the loss of important information. For example, if two users update the attributes of the same node almost simultaneously, the later update will overwrite the previous one, potentially leading to data inconsistency. Without version control, it is impossible to track the historical changes of data, making it difficult to audit and trace the evolution process of data, and the update management of the graph database is not refined enough.
[0031] The graph database update method provided by the embodiments of this application obtains the version information to be updated for vertex data, edge data, or attribute data, improves data consistency in the graph database, reduces data overwriting and conflicts, and reduces data errors and inconsistencies caused by updates. And when setting data versions in the graph database, a version number is set separately for each attribute field, rather than only controlling the version number of the entire vertex (node) or edge (relationship), achieving finer-grained control over the updates of the graph database. For example, if only one field changes, only the version number of this field will increase, without the need to update the version number of the entire node or edge, thereby reducing the amount of update operations.
[0032] Please refer to Figure 1 The flowchart showing a graph database update method provided by the embodiments of this application. The graph database update method provided by the embodiments of this application can be applied to an electronic device. The electronic device can include physical devices such as a server, a PC, a tablet computer, or a smart phone, or can also be a virtual device such as a virtual machine or a container. The electronic device can be a single device, or a combination of multiple devices or a cluster of a large number of devices. The graph database update method can include:
[0033] Step S110: Obtain the identification information and the version information to be updated of the data to be updated; the data to be updated includes vertex data, edge data, or attribute data to be updated in the graph database; the version information to be updated is used to identify the update order of the data to be updated.
[0034] Step S120: Query the historical version information of the data to be updated according to the identification information; the historical version information is used to represent the versions that the data to be updated has stored in the graph database.
[0035] Step S130: Compare the version information to be updated with the historical version information, and update the graph database according to the comparison result.
[0036] In step S110: In the graph database, the data to be updated refers to those data entities that are about to be modified. These entities can be vertex (node) data, edge (relationship) data, or the attribute data contained in these entities. That is, the attribute data refers to the attributes of vertex data or edge data. Attributes can be used to store information about entities or relationships.
[0037] For example, a node can represent an entity, such as a person, a place, an item, etc. An edge represents the connection relationship between two nodes, such as "friend", "colleague", or "distance", etc. The attributes of vertex data, for example, the name and age of a person; the attributes of edge data, for example, the distance between two places.
[0038] When setting the data version of the graph database, it is necessary to set version information not only for vertex data and edge data, but also for each (vertex data or edge data's) attribute data. For example, in the "Zhang San" node, the included attribute data has fields such as name, age, and sex. Among them, the version information of the name field is 1241213, the version information of the age field is 1231312, and the version information of the sex field is 1231231.
[0039] If the data to be updated is vertex data, the identification information can be the ID of the node or other identifiers of the node, and the version information to be updated can be the version information that the node is about to be updated; if the data to be updated is edge data, the identification information can be the ID of the edge or other identifiers of the node, and the version information to be updated can be the version information that the edge is about to be updated.
[0040] It can be understood that if the data to be updated is attribute information, then its identification information should be the identification information of the corresponding vertex data or edge data of the attribute information. For example, if the data to be updated is the first attribute data of vertex data (such as node A), the identification information of the data to be updated can be the ID of the above node (A) or other identifiers of node A, and the version information to be updated can be the version information to be updated of the first attribute data.
[0041] If the data to be updated is the second attribute data of edge data (such as edge w), the identification information of the data to be updated can be the ID of the above edge w data or other identifiers, and the version information to be updated can be the version information to be updated of the second attribute data.
[0042] When only a certain attribute of a node or an edge needs to be updated, only that specific field can be operated on without affecting other fields. This makes the business logic more flexible and able to respond to rapidly changing business requirements.
[0043] The version information to be updated is used to identify the update order of the data to be updated. For example, the version information to be updated can be a version number, and the form can be a number or a timestamp. For example, the update with version number 2 occurs after the update with version number 1.
[0044] In step S120: Query the historical version information of the data to be updated according to the identification information. For example, it can be retrieved through an SQL query or obtained by calling an interface. The historical version information is used to represent the version of the data to be updated that has been stored in the graph database, that is, the version after the last successful update of the entity corresponding to the identification information.
[0045] In step S130: For example, if the version information to be updated (such as the version number) is higher than the historical version information, it indicates that the data to be updated is the latest and can be updated. If the version information to be updated is lower than or equal to the historical version information, it indicates that the data to be updated is not the latest and may have been overwritten by other update operations. If the version information to be updated is the latest, the system will execute the update operation and write the new data to be updated and the version information to be updated into the graph database. In the next update, this version information to be updated is compared with the new version information to be updated as the historical version information.
[0046] In the implementation process of the above embodiments: Obtain the version information to be updated of the point data, edge data or attribute data, and improve the data consistency in the graph database by comparing the version information first and then updating, reduce data overwriting and conflicts, and reduce data errors and inconsistencies caused by updates. And set version information for each attribute field separately, rather than only controlling the version information of the entire point (node) or edge (relationship), to achieve finer-grained control over the update of the graph database. For example, if only one field changes, only the version number of this field will increase, and there is no need to update the version number of the entire node or edge, thereby reducing the amount of update operations. When only a certain attribute of a node or edge needs to be updated, only that specific field can be operated on without affecting other fields. This makes the business logic more flexible and able to respond to rapidly changing business requirements.
[0047] Optionally, in the embodiments of the present application, obtaining the identification information and the version information to be updated of the data to be updated includes:
[0048] Receive the identification information and the version information to be updated of multiple data to be updated. For example, the identification information and the version information to be updated of the data to be updated can be collected from various data sources (such as user input, external API calls, sensors, etc.). The collected data can be verified and preprocessed to ensure the integrity and consistency of the data, such as checking the format and range of the version information to be updated. If the data to be updated has a preset format, the data to be updated can also be checked.
[0049] Write multiple data to be updated into the corresponding partitions of the message queue respectively according to the identification information of the data to be updated; the message queue is partitioned in advance using the identification information of the data to be updated as the partition key value, so that the data to be updated with the same identification information is written into the same partition.
[0050] The message queue can be Kafka, etc., and the partitioning strategy can use the identification information of the data to be updated as the partition key value. For example, in Kafka, the node ID or edge ID can be set as the partition key. Encapsulate the data to be updated into messages respectively, and send the data to be updated to the corresponding partition of the message queue according to the identification information of the data to be updated.
[0051] Obtain the identification information and the version information to be updated of the data to be updated from the corresponding partition of the message queue. Configure the consumer of the message queue, which can be an application, a server, a microservice instance, etc. The consumer pulls messages from the partition and parses the message content to obtain the identification information and version information of the data to be updated.
[0052] After the message is successfully processed, the consumer can confirm to the message queue that the message has been processed to avoid duplicate processing of the message.
[0053] In the implementation process of the above embodiments: by partitioning the message queue and using the identification information as the partition key, the data to be updated with the same identification information can be divided into the same partition for sequential processing, reducing the modification of the data corresponding to the same identification information at the same time, resulting in data inconsistency problems caused by concurrent updates. And after using partitioning, the message queue can handle high-concurrency data update requests and adapt to the growing data update requirements.
[0054] Optionally, in the embodiments of the present application, writing multiple data to be updated into the corresponding partitions of the message queue respectively according to the identification information of the data to be updated includes:
[0055] Input the identification information of the data to be updated into a pre-trained version conflict recognition model to obtain the version conflict probability of the data to be updated; the version conflict probability is used to characterize the possibility of version conflict of the data to be updated.
[0056] Version conflict refers to the situation where when multiple users or processes concurrently modify the same data entity (such as a file, a record, etc.), these modifications are incompatible or overwrite each other, resulting in an uncertain or inconsistent final state of the data.
[0057] The version conflict probability is used to characterize the likelihood of a version conflict occurring when the data to be updated is updated in the future. The version conflict probability can be a value between 0 and 1. Based on historical data and learned patterns, it predicts upcoming events. A value close to 1 indicates a high risk, that is, a version conflict is very likely to occur during the next update. A value close to 0 indicates a low risk, that is, the possibility of a version conflict is very small. The version conflict probability can also be classified as high, medium, or low, etc.
[0058] The version conflict probability provides a quantitative metric for assessing the risk level of the data to be updated. Based on the version conflict probability, it can be decided whether to continue with the update operation or take other measures, such as adjusting the strategy for partitioning the message queue or adjusting the strategy for using the message queue.
[0059] The version conflict recognition model is used to learn based on the identification information, obtain version features related to the identification information, and obtain the version conflict probability of the data to be updated based on the version features. Among them, the version features can include update frequency, historical conflict records, the complexity of data changes, etc. The obtaining methods include, for example, querying data to query version features related to the identification information. This database is generated from the historical update data of the identification information, and it can also be obtained by calling an interface and learned from the historical update data of the identification information. The training process of the version conflict recognition model will be described later.
[0060] If the version conflict probability of the data to be updated is greater than the first preset threshold, the data to be updated is written into the corresponding partition of the message queue. If the version conflict probability of the data to be updated is greater than the first preset threshold, it indicates that the data to be updated is likely to be modified at the same time in the future, resulting in a greater possibility of a version conflict. For example, if certain fields or point data are often modified and the modification frequency is high, then the version conflict probability of this field or point data may be greater than the first preset threshold.
[0061] For fields with a relatively high predicted version conflict probability, preventive and optimization measures can be taken in advance to improve data consistency, system stability, and performance. For example, in the embodiments of this application, the number of partitions can be determined for such data with a relatively high version conflict probability, and the partition mechanism of the message queue can be used to improve the problem of high concurrency.
[0062] Exemplarily, when partitioning the message queue, it can be determined based on the number of data to be updated with a version conflict probability greater than a first preset threshold. Suppose there are 10 vertex data or edge data with a version conflict probability greater than the first preset threshold. Then, the identification information of these 10 vertex data or edge data can be used as the partitioning key to partition the message queue. For other low-risk data (version conflict probability not greater than the first preset threshold) in the graph database, there is no need to establish partitions. Since one of the functions of the message queue is to ensure that update operations are executed in sequence to reduce version conflicts. Low-risk data has a relatively low probability of version conflict. Using the message queue method for such data may consume certain resources and reduce processing efficiency. Therefore, such low-risk data can be directly updated, reducing the steps of writing to the message queue, which not only reduces version conflicts but also improves processing efficiency and reduces resource consumption.
[0063] In the implementation process of the above embodiment: Write the high-risk data to be updated into the corresponding partition of the message queue. In this way, these high-risk data to be updated can be processed in the order of entering the message queue, avoiding the problem of high concurrency. For other data that is not high-risk, it can be directly updated to rationally use the resources of the message queue.
[0064] Reduce resource consumption and improve resource utilization. Process different-risk data respectively with corresponding methods, achieving more refined management of data updates, reducing the problem of version conflicts while increasing the throughput of data processing.
[0065] Optionally, in the embodiment of the present application, before inputting the identification information of the data to be updated into the pre-trained version conflict recognition model, the method further includes:
[0066] Obtain the historical update data of the graph database; the historical update data includes the identification information and update timestamps of the historically updated graph data. For example, the identification information and update timestamps of vertex data, the identification information and update timestamps of edge data, the update timestamps of the attribute data of vertex data, and the update timestamps of the attribute data of edge data. It can be understood that the identification information corresponding to the attribute data of vertex data is the identification information of the vertex data, and the identification information corresponding to the attribute data of edge data is the identification information of the edge data.
[0067] Among them, the update timestamp can be the update timestamp at each update. Taking vertex data as an example, the historical update data can record the update timestamp of each of the multiple historical modifications of the vertex data, enabling the model to learn features such as the update frequency and historical conflict times of the vertex data based on this data.
[0068] As an implementation, to enrich the dataset and enable the model to learn more useful features, the historical update data can also include the type of operation (such as insert, delete, modify); system performance metrics can also be collected, such as the time taken to process each update operation, system load, the number of concurrent requests, etc. Data features that may affect version conflicts can be extracted from these data to train the model, improving the generalization ability and robustness of the model.
[0069] Use a preset time window to extract historical version features for influencing version conflicts from the historical update data.
[0070] The time window can be a sliding window or a fixed window. For example, the sliding window technique can be used to extract features within a specific time range, such as the number of updates in the past 1 hour, the maximum number of concurrent updates in the past 24 hours, etc. A fixed window creates a new window at each given time point. For example, the number of updates in a specific time period each day, etc.
[0071] The size of the time window can be determined based on business logic and data characteristics. For example, if version conflicts usually occur within a few minutes after data updates, then the time window may be set from a few minutes to a few hours. The determination of the time window can also consider the periodicity of business operations. For example, during peak hours of the day, this time period of each cycle can be determined as a fixed window.
[0072] The way to extract historical version features for influencing version conflicts is, for example, within the time window, perform an aggregation operation on the historical update data according to the identification information, such as calculating the number of updates of a certain identification information, the average update request volume, detecting outliers, etc. Statistical features within the time window can also be extracted, such as maximum value, minimum value, average value, median, standard deviation, etc.
[0073] The historical version features can also be preprocessed, including cleaning (removing noise and outliers), normalization (making the features have the same scale), and encoding (such as converting categorical features to numerical features).
[0074] Use the historical version features to train an initial machine learning model to obtain a version conflict recognition model. Select a suitable machine learning algorithm, such as logistic regression, random forest, gradient boosting tree, neural network. Use the historical version features to train the determined initial machine learning model and optimize the performance of the model by adjusting the model parameters (hyperparameter tuning). Use techniques such as cross-validation to evaluate the generalization ability of the model to ensure that the model can also have good prediction effects on unseen data.
[0075] In an alternative embodiment, the data input when using the model can be simplified to improve the efficiency of the update operation. By inputting identification information (such as a point ID or the starting point ID of an edge), the version conflict probability of this field can be obtained. During the model training phase, all the data corresponding to the identification information in the graph database can be used to train the model, enabling the model to fully learn the relationship between each piece of identification information in the graph database and the historical version features. The model trained in this way can infer the version features related to the identification information from the identification information. For example, the model can learn how to predict the conflict probability based on the association between the point ID or the starting point ID of the edge and the historical data. In the training dataset, each sample contains an identification information field for subsequent feature inference.
[0076] For example, a lookup table or database query can be used inside the model to obtain the features related to the input identification information. For example, after inputting the point ID, the model can query the database to obtain features such as the update history and the number of concurrent updates of this point. Or an automatically generated feature module can be added to the model structure. After receiving the input identification information, the required features can be automatically generated according to predefined rules. An Embedding layer can be considered to process these identification information and learn their potential representations. This method simplifies the user input. The user only needs to provide the identification information, and the model can automatically handle the rest, eliminating the process of inputting complex information when using the model and improving the user experience.
[0077] In the implementation process of the above embodiment: The machine learning prediction version conflict solution can help identify and prevent potential version conflict problems in advance, reduce update anomalies, and improve data consistency and business continuity. The time window used during the machine learning model training process can help the model capture the time-related historical version features, improving the prediction accuracy and practicality of the model. The model effectively utilizes the identification information during training and looks up or calculates the related features internally. When the prediction of the version conflict probability is accurate and reliable, the model input is simplified and the efficiency is improved.
[0078] Optionally, in the embodiment of the present application, obtaining the identification information and the to-be-updated version information of the data to be updated from the corresponding partition of the message queue includes:
[0079] A partition is a data storage unit in the message queue. It can be divided into multiple partitions. Each partition is an ordered and immutable message sequence, and the messages in each partition are processed in order. Each partition can be distributed on different Brokers, which can make full use of the storage capacity of the cluster.
[0080] For different partitions in the message queue, concurrently execute the steps of obtaining the identification information of the data to be updated and the version information to be updated. Multiple consumers in the consumer group can process the data in different partitions in parallel, improving throughput.
[0081] For the same partition in the message queue, execute in sequence according to the order in which the data to be updated enters the corresponding partition of the message queue: the steps of obtaining the identification information of the data to be updated and the version information to be updated.
[0082] In the implementation process of the above embodiments: The partitions of the message queue can improve the scalability and parallel processing ability of the message queue. By dispersing the data into different partitions, data that will not cause concurrent problems such as version conflicts can work in parallel; while data with the same identification information is processed in order, reducing the possibility of version conflicts, achieving efficient processing of a large number of messages, and at the same time improving data consistency and orderliness.
[0083] Optionally, in the embodiments of the present application, compare the version information to be updated with the historical version information, and update the graph database according to the comparison result, including:
[0084] If the identification information of the version information to be updated does not exist in the graph database, write the data to be updated into the graph database.
[0085] First, it is necessary to check whether there is identification information corresponding to the data to be updated in the graph database. If it does not exist, it means that if the identification information of the data to be updated does not exist in the graph database, it means that there is no such node or edge in the graph database originally, and the insert operation can be executed to write the data to be updated as new data into the graph database.
[0086] If the identification information of the version information to be updated exists in the graph database, compare the version information to be updated with the historical version information. If the time order of the historical version information is before the version information to be updated, perform the update operation and update the graph database with the data to be updated; if the time order of the historical version information is after the version information to be updated, discard the data to be updated.
[0087] If the identification information of the data to be updated exists in the graph database, next, it is necessary to compare the version information to be updated with the historical version information stored in the graph database. For example, compare fields that can represent the time order such as the timestamp or version number in the version information.
[0088] If the time order of the historical version information is before the version information to be updated, it means that the data to be updated is the newer data. In this case, the update operation can be executed to update the corresponding record in the graph database with the data to be updated.
[0089] If the chronological order of the historical version information is after the version information to be updated, it indicates that the data to be updated is not the latest version, and the data to be updated can be discarded without processing it to ensure that the data stored in the graph database is the latest.
[0090] In the implementation process of the above embodiments: By comparing the version information to be updated with the historical version information, the data consistency and integrity of the graph database are maintained, so that the data stored in the database is the latest, and refined management of the vertex data, edge data, or attribute data in the graph database is realized. And by the method of updating after comparison, only valid data (the data of the latest version) needs to be written, reducing the amount of data written and the update cost.
[0091] Optionally, in the embodiments of the present application, updating the graph database according to the comparison result includes: If the comparison result indicates that an update operation needs to be performed, the original data corresponding to the identification information is obtained from the graph database according to the identification information, and the original data is overwritten with the data to be updated to realize the update of the graph data.
[0092] Continuing with the above embodiments, if the chronological order of the historical version information is before the version information to be updated, an update operation is performed, and the original data is overwritten with the data to be updated, which can be realized, for example, through the update interface or query language of the graph database.
[0093] In the implementation process of the above embodiments: The overwrite operation makes the data stored in the graph database always the latest, and can reduce data redundancy and inconsistency, improving the accuracy and reliability of the data. There is no need to store multiple versions of the same data, thus optimizing the use of storage space, and the overwrite operation simplifies the management complexity of the graph database.
[0094] Please refer to Figure 2 the structural schematic diagram of the graph database update device provided by the embodiments of the present application shown; The embodiments of the present application provide a graph database update device 200, including:
[0095] A data acquisition module 210, configured to acquire the identification information of the data to be updated and the version information to be updated; the data to be updated includes the vertex data, edge data, or attribute data to be updated in the graph database; the version information to be updated is used to identify the update order of the data to be updated;
[0096] A version information acquisition module 220, configured to query the historical version information of the data to be updated according to the identification information; the historical version information is used to represent the version of the data to be updated that has been stored in the graph database;
[0097] A comparison and update module 230, configured to compare the version information to be updated with the historical version information, and update the graph database according to the comparison result.
[0098] Optionally, in the embodiments of the present application, the graph database update device 200, the data acquisition module 210 is configured to receive the identification information and the to-be-updated version information of multiple to-be-updated data; write the multiple to-be-updated data into the corresponding partitions of the message queue respectively according to the identification information of the to-be-updated data; the message queue is partitioned in advance using the identification information of the to-be-updated data as the partition key value, so that the to-be-updated data with the same identification information is written into the same partition; obtain the identification information and the to-be-updated version information of the to-be-updated data from the corresponding partitions of the message queue.
[0099] Optionally, in the embodiments of the present application, the graph database update device 200, the data acquisition module 210 is further configured to input the identification information of the to-be-updated data into a pre-trained version conflict recognition model to obtain the version conflict probability of the to-be-updated data; the version conflict probability is used to characterize the possibility of version conflict of the to-be-updated data; the version conflict recognition model is used to learn according to the identification information, obtain the version features related to the identification information, and obtain the version conflict probability of the to-be-updated data based on the version features; if the version conflict probability is greater than the first preset threshold, write the to-be-updated data into the corresponding partition of the message queue.
[0100] Optionally, in the embodiments of the present application, the graph database update device 200, the probability model training module is configured to obtain the historical update data of the graph database; the historical update data includes the identification information and the update timestamp of the historically updated graph data; use a preset time window to extract the historical version features for affecting version conflict from the historical update data; train a machine learning initial model using the historical version features to obtain a version conflict recognition model.
[0101] Optionally, in the embodiments of the present application, the graph database update device 200, the data acquisition module 210 is further configured to, for different partitions in the message queue, perform concurrently: the step of obtaining the identification information and the to-be-updated version information of the to-be-updated data; for the same partition in the message queue, sequentially perform according to the order in which the to-be-updated data enters the corresponding partition of the message queue: the step of obtaining the identification information and the to-be-updated version information of the to-be-updated data.
[0102] Optionally, in the embodiments of the present application, the graph database update device 200, the comparison and update module 230 is specifically configured to, if the identification information of the to-be-updated version information does not exist in the graph database, perform an update operation and write the to-be-updated data into the graph database; if the identification information of the to-be-updated version information exists in the graph database, compare the to-be-updated version information with the historical version information, and if the time order of the historical version information is before the to-be-updated version information, perform an update operation and update the graph database with the to-be-updated data; if the time order of the historical version information is after the to-be-updated version information, discard the to-be-updated data.
[0103] Optionally, in an embodiment of the present application, the graph database updating device 200 and the comparison update module 230 are specifically used to obtain the original data corresponding to the identification information from the graph database according to the identification information if the comparison result indicates that an update operation needs to be performed, and use the data to be updated to overwrite the original data to achieve the update of the graph data.
[0104] It should be understood that the device corresponds to the above-mentioned graph database update method embodiment and can execute the various steps involved in the above-mentioned method embodiment. The specific functions of the device can be found in the above description. To avoid repetition, the detailed description is appropriately omitted here. The device includes at least one software function module that can be stored in the memory in the form of software or firmware or solidified in the operating system (OS) of the device.
[0105] See also Figure 3 The electronic device 300 provided in the embodiment of the present application includes: a processor 310 and a memory 320, wherein the memory 320 stores machine-readable instructions executable by the processor 310, and when the machine-readable instructions are executed by the processor 310, the above method is executed.
[0106] Figure 3 Each component shown in can be implemented by hardware, software or a combination thereof. The electronic device 300 may be a physical device, such as a server, a PC, etc., or a virtual device, such as a virtual machine, a virtualized container, etc. Moreover, the electronic device 300 is not limited to a single device, but may also be a combination of multiple devices or a cluster consisting of a large number of devices.
[0107] An embodiment of the present application further provides a storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the above method is executed.
[0108] Among them, the storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0109] In several embodiments provided by the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and block diagrams in the drawings show the possible architectures, functions, and operations of the devices, methods, and computer program products according to multiple embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of the code, and the module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the drawings. For example, two consecutive blocks may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0110] In addition, in each embodiment of the embodiments of the present application, the various functional modules may be integrated together to form an independent part, or each module may exist separately, or two or more modules may be integrated to form an independent part.
[0111] The above description is only an optional implementation manner of the embodiments of the present application, but the protection scope of the embodiments of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the embodiments of the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the embodiments of the present application.
Claims
1. A graph database updating method, characterized in that: include: Acquire identification information and version information of data to be updated; the data to be updated includes point data, edge data or attribute data to be updated in the graph database; The version information to be updated is used to identify the update order of the data to be updated; Querying historical version information of the data to be updated according to the identification information; the historical version information is used to represent the version of the data to be updated that has been stored in the graph database; The version information to be updated is compared with the historical version information, and the graph database is updated according to the comparison result.
2. The method according to claim 1, characterized in that The obtaining of identification information of the data to be updated and version information to be updated includes: Receiving identification information of a plurality of data to be updated and version information to be updated; According to the identification information of the data to be updated, the plurality of data to be updated are respectively written into the partitions corresponding to the message queue; the message queue is partitioned in advance by using the identification information of the data to be updated as a partition key value, so that the data to be updated with the same identification information are written into the same partition; The identification information of the data to be updated and the version information to be updated are obtained from the partition corresponding to the message queue.
3. The method according to claim 2, characterized in that According to the identification information of the data to be updated, the plurality of data to be updated are respectively written into the partitions corresponding to the message queue, including: Inputting the identification information of the data to be updated into a pre-trained version conflict identification model to obtain the version conflict probability of the data to be updated; the version conflict probability is used to characterize the possibility of version conflict occurring in the data to be updated; the version conflict identification model is used to learn according to the identification information, obtain version features related to the identification information, and obtain the version conflict probability of the data to be updated based on the version features; If the version conflict probability is greater than a first preset threshold, the data to be updated is written into the partition corresponding to the message queue.
4. The method according to claim 3, characterized in that Before inputting the identification information of the data to be updated into the pre-trained version conflict identification model, the method further includes: Acquire historical update data of the graph database; the historical update data includes identification information and update timestamp of historically updated graph data; Using a preset time window, extracting historical version features that affect version conflicts from the historical update data; The historical version features are used to train the machine learning initial model to obtain the version conflict identification model.
5. The method according to claim 2, characterized in that: Acquiring identification information of the data to be updated and version information to be updated from the partition corresponding to the message queue includes: For different partitions in the message queue, concurrently executing: the step of obtaining the identification information of the data to be updated and the version information to be updated; For the same partition in the message queue, the steps of obtaining the identification information of the data to be updated and the version information to be updated are performed in sequence according to the order in which the data to be updated enters the corresponding partition of the message queue.
6. The method according to claim 1, characterized in that Comparing the version information to be updated with the historical version information, and updating the graph database according to the comparison result, includes: If the identification information of the version information to be updated does not exist in the graph database, writing the data to be updated into the graph database; If the identification information of the version information to be updated exists in the graph database, the version information to be updated is compared with the historical version information. If the time sequence of the historical version information is before the version information to be updated, an update operation is performed to update the graph database using the data to be updated; if the time sequence of the historical version information is after the version information to be updated, the data to be updated is discarded.
7. The method according to any one of claims 1 to 6, characterized in that: The updating of the graph database according to the comparison result includes: If the comparison result indicates that an update operation needs to be performed, the original data corresponding to the identification information is obtained from the graph database according to the identification information, and the original data is overwritten with the data to be updated to implement the update of the graph data.
8. A computer program product, characterized in that The method comprises computer program instructions, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is executed.
9. An electronic device, characterized in that: include: A processor and a memory, wherein the memory stores computer program instructions, and when the computer program instructions are executed by the processor, the method according to any one of claims 1 to 7 is executed.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer program instructions, and when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is executed.