Data consistency processing method, distributed system and computer storage medium
By generating data versions in a distributed storage system and periodic conflict detection, combining node priority management and version broadcasting, the problem of inconsistent data versions in a distributed storage system is solved, and data consistency and system stability are achieved.
Patent Information
- Application Number
- CN202510526654.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-05-27
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In a distributed storage system, when multiple nodes update the same target data, it may cause inconsistent data versions, causing read errors, business logic exceptions, and affecting system stability.
The data version is generated when the node triggers an update operation and periodically performs conflict detection on all versions of the target data within the preset time window. When a conflicting version is detected, determine the node priority that stores the conflicting version, select the node version with the highest priority as the target version, and broadcast it to all relevant nodes.
Ensure consistency of data in distributed storage systems, and even if concurrent updates occur, conflicts will be resolved through priority management and version synchronization, thereby improving system stability and reliability.
Smart Images

Figure CN120045573A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present application relate to the field of distributed technologies, and in particular, to a data consistency processing method, a distributed system, and a computer storage medium. Background Art
[0002] In a distributed storage system, multiple nodes may simultaneously update the same target data, resulting in inconsistent versions stored on different nodes. This problem of data inconsistency may cause read errors, business logic exceptions, and even affect the overall stability of the system.
[0003] Traditional distributed consistency processing methods mainly rely on strong consistency protocols (such as Paxos, Raft) or simple timestamp comparison strategies, but these methods have certain limitations. For example, strong consistency protocols usually require high communication overhead and coordination costs, reducing the concurrency of the system, while simple timestamp comparison strategies may lead to data overwrite problems, causing the newer but incorrect version to overwrite the correct version, thus affecting the reliability of the data.
[0004] In addition, in some distributed systems, the detection and processing of data conflicts usually rely on manual intervention or specific business logic. This method not only increases the maintenance cost of the system but also easily causes data recovery delays, affecting the overall data consistency. Summary of the Invention
[0005] Embodiments of the present application provide a data consistency processing method, a distributed system, and a computer storage medium, which can ensure the data consistency of a distributed storage system.
[0006] In a first aspect of the embodiments of the present application, a data consistency processing method is provided, which is applied to a distributed storage system. The method includes: When a node triggers an update operation on target data, generate a data version; Periodically perform conflict detection on all data versions of the target data within a preset time window; When a conflict version is detected, determine the priorities of the nodes storing the conflict version, and determine the highest priority; Determine the target version through the node with the highest priority; Broadcast the target version.
[0007] Optionally, the periodically performing conflict detection on all data versions of the target data within a preset time window includes: When the detection period arrives, respectively obtain the latest data versions from each node storing the target data; Perform conflict detection on the obtained multiple latest data versions.
[0008] Optionally, the conflict detection for the obtained multiple latest data versions includes: Calculating the hash values of the multiple latest data versions respectively; Judging whether there are conflict versions according to the calculated hash values.
[0009] Optionally, before determining the priorities of the nodes storing the conflict versions, the method further includes: Determining a conflict resolution strategy, where the conflict resolution strategy includes latest version first, node reliability first, and conflict merging; The determining the priorities of the nodes storing the conflict versions includes: Determining the priorities of the nodes storing the conflict versions according to the conflict resolution strategy.
[0010] Optionally, the determining the conflict resolution strategy includes: Determining the system state of the distributed storage system; Determining the conflict type of the conflict versions; Determining the conflict resolution strategy according to the system state and / or the conflict type.
[0011] Optionally, the determining the conflict type of the conflict versions includes: Obtaining the number of conflict nodes and the distribution of conflict nodes of the conflict versions; Determining the conflict type according to the number of conflict nodes and the distribution of conflict nodes.
[0012] Optionally, the determining the priorities of the nodes storing the conflict versions according to the conflict resolution strategy includes: When the conflict resolution strategy is latest version first, determining the priorities of the nodes storing the conflict versions according to the time stamps; When the conflict resolution strategy is node reliability first, calculating the reliability of the nodes storing the conflict versions according to network latency, processor load, and data access frequency, and determining the priorities of the nodes storing the conflict versions according to the reliability; When the conflict resolution strategy is conflict merging, determining that the priorities of the nodes storing the conflict versions are the same.
[0013] The second aspect of the embodiments of the present application provides a distributed storage system, including: A generating unit, configured to generate a data version when a node triggers an update operation on target data; A detecting unit, configured to periodically perform conflict detection on all data versions of the target data within a preset time window; The first determination unit is configured to determine the priority of the node storing the conflict version when a conflict version is detected, and determine the highest priority; The second determination unit is configured to determine the target version through the node with the highest priority; The broadcast unit is configured to broadcast the target version.
[0014] Optionally, the detection unit includes: An acquisition module, configured to respectively acquire the latest data versions from each node storing the target data when the detection period arrives; A detection module, configured to perform conflict detection on the multiple acquired latest data versions.
[0015] Optionally, the detection module is specifically configured to: Calculate the hash values of the multiple latest data versions respectively; Judge whether there is a conflict version according to the calculated hash values.
[0016] Optionally, the distributed storage system further includes: A third determination unit, configured to determine a conflict resolution strategy, where the conflict resolution strategy includes giving priority to the latest version, giving priority to node reliability, and conflict merging; The first determination unit is specifically configured to: Determine the priority of the node storing the conflict version according to the conflict resolution strategy.
[0017] Optionally, the third determination unit includes: A first determination module, configured to determine the system state of the distributed storage system; A second determination module, configured to determine the conflict type of the conflict version; A third determination module, configured to determine a conflict resolution strategy according to the system state and / or the conflict type.
[0018] Optionally, the second determination module is specifically configured to: Acquire the number of conflict nodes and the distribution of conflict nodes of the conflict version; Determine the conflict type according to the number of conflict nodes and the distribution of conflict nodes.
[0019] Optionally, the first determination unit is specifically configured to: When the conflict resolution strategy is to give priority to the latest version, determine the priority of the node storing the conflict version according to the time stamp; When the conflict resolution strategy is node reliability first, calculate the reliability of the node storing the conflict version according to network latency, processor load, and data access frequency, and determine the priority of the node storing the conflict version according to the reliability; When the conflict resolution strategy is conflict merging, determine that the priorities of the nodes storing the conflict version are the same.
[0020] The third aspect of the embodiments of the present application provides a distributed storage system, including: A processor, a memory, an input / output unit, and a bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method in the first aspect and any possible implementation manner of the first aspect.
[0021] The fourth aspect of the embodiments of the present application provides a computer-readable storage medium, on which a program is stored, and when the program is executed on a computer, the computer executes the method in the first aspect and any possible implementation manner of the first aspect.
[0022] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages: In the embodiments of the present application, first, a node generates a data version for the update of the target data, and through periodic conflict detection, it is found that multiple nodes may have conflicting updates to the same data. At this time, the system determines which node's version is more authoritative by assigning priorities to each node, ensuring that the system selects the most reliable version. Then, the version held by the node with the highest priority is determined as the target version. Finally, by broadcasting the target version to all relevant nodes, data consistency among the nodes is ensured. Therefore, even in the case of concurrent updates, conflicts can ultimately be resolved through priority management and version synchronization, thereby ensuring data consistency of the system. Description of the Drawings
[0023] Figure 1 It is a flowchart of an embodiment of the data consistency processing method in the embodiments of the present application; Figure 2 It is a flowchart of an embodiment of conflict detection in the embodiments of the present application; Figure 3 It is a flowchart of an embodiment of the data consistency processing method in the embodiments of the present application; Figure 4 It is a flowchart of an embodiment of determining the conflict resolution strategy in the embodiments of the present application; Figure 5Schematic flowchart of an embodiment of the conflict type for determining conflict versions in an embodiment of the present application; Figure 6 Schematic structural diagram of an embodiment of a distributed storage system in an embodiment of the present application; Figure 7 Schematic structural diagram of another embodiment of a distributed storage system in an embodiment of the present application. Detailed implementation manners
[0024] The embodiments of the present application provide a data consistency processing method, a distributed system, and a computer storage medium for ensuring data consistency in a distributed storage system.
[0025] Next, the embodiments in the present application will be described with reference to the accompanying drawings.
[0026] Please refer to Figure 1 , an embodiment of the data consistency processing method in an embodiment of the present application includes: 101. When a node triggers an update operation on target data, generate a data version; When a certain node in the distributed storage system initiates an update operation on target data, the system generates a data version for this update. This data version usually includes the updated data value and its associated metadata, such as version number, timestamp, node identifier of the source, etc. The data version is used to identify an independent modification behavior and is the basis for the system to support concurrent updates and version management.
[0027] 102. Periodically perform conflict detection on all data versions of the target data within a preset time window; To ensure that updates to the same target data by multiple nodes in the system do not cause inconsistency problems, the distributed storage system periodically performs conflict detection. Specifically, the system extracts all data versions generated by all nodes for the target data within a certain time window according to the preset detection period. This time window is the time range set by the system for collecting potential conflict versions. During the conflict detection process, the system compares the timestamps, write sources, version numbers, or other customized conflict detection rules of the versions to determine whether there are multiple versions that have made incompatible modifications to the same data content, that is, "conflict versions" appear.
[0028] 103. When conflict versions are detected, determine the priorities of the nodes storing the conflict versions and determine the highest priority; Once a conflicting version is detected, the distributed storage system needs to determine which nodes store these conflicting versions respectively and assign priorities to these nodes. The priorities of nodes can be set based on multiple dimensions, such as the role of the node (e.g., primary node or secondary node), network response latency, running stability, historical consistency performance, geographical location, or manually set priority weights, etc. The system will comprehensively consider the above factors to determine a node with the highest priority. A higher priority of this node means it has higher credibility or authority when resolving conflicts.
[0029] 104. Determine the target version through the node with the highest priority; After obtaining the node with the highest priority, the system uses the data version stored by this node as the final "target version". This target version is considered the authoritative version in this round of conflict resolution and is used to represent the final state of the target data after the conflict occurs.
[0030] 105. Broadcast the target version.
[0031] Subsequently, the distributed storage system will broadcast this target version to all relevant nodes within the system. The broadcast process can adopt a consistency protocol (such as the Gossip protocol or replication mechanisms like Paxos / Raft) to ensure that all replica nodes receive this target version and update their local storage, so that the entire system reaches a consistent state again on this data unit and maintains the overall data consistency.
[0032] In this embodiment, in the distributed storage system, first, the nodes generate data versions for the update of the target data, and through periodic conflict detection, it is found that multiple nodes may have conflicting updates to the same data. At this time, the system determines which node versions are more authoritative by assigning priorities to each node, ensuring that the system selects a most reliable version. Then, the version held by the node with the highest priority is determined as the target version. Finally, by broadcasting the target version to all relevant nodes, the data consistency of each node is ensured. Therefore, even if concurrent updates occur, conflicts can ultimately be resolved through priority management and version synchronization, thus ensuring the data consistency of the system.
[0033] Please refer to Figure 2 , in some embodiments of this application, step 102 in the above embodiment for periodically performing conflict detection on all data versions of the target data within a preset time window may include the following steps: 201. When the detection period arrives, respectively obtain the latest data versions from each node storing the target data; When the detection period arrives, the distributed storage system retrieves the latest data version on each node storing the target data separately. To ensure that the retrieved data represents the most representative data state within the current detection period, the system preferentially retrieves information containing complete version metadata, such as version numbers, timestamps, write order identifiers, and node sources. This information is not only used to determine the newness of versions but also provides auxiliary metrics for subsequent conflict judgment. Through this step, the system establishes a data set composed of the latest versions on each node.
[0034] 202. Perform conflict detection on the multiple latest data versions obtained.
[0035] After the system aggregates the latest data versions from each node, it performs conflict detection on these versions. Conflict detection is usually based on preset judgment criteria, such as comparing whether the data contents are consistent, whether the modification times overlap, or whether they are concurrent write operations from different nodes. When it is detected that there are multiple different versions and they cannot be directly merged into a consistent result, the system considers that there are conflicting versions of the current target data. This step not only identifies the existence of conflicts but also provides an accurate basis for judging the scope of conflicts and subsequent conflict resolution strategies.
[0036] In this embodiment, by regularly retrieving the latest data versions from each node and performing conflict detection, the system can timely detect version conflicts caused by concurrent updates. Conflict detection ensures that the data versions of all relevant nodes are effectively compared and reviewed, thus avoiding the situation where different versions are stored on multiple nodes and ensuring the accuracy of data consistency. Through conflict detection, the system can effectively identify and handle conflicts, providing a clear basis for subsequent selection of the target version by priority, thereby greatly reducing the risk of data inconsistency and ensuring the stability and reliability of the system.
[0037] Specifically, step 202 for performing conflict detection on the multiple latest data versions obtained may include the following steps: Calculate the hash values of the multiple latest data versions separately; Judge whether there are conflicting versions according to the calculated hash values.
[0038] After the system obtains multiple latest data versions, it calculates the hash value for the content of each data version separately. This hash value is used to accurately represent the uniqueness and integrity of the data content itself. By compressing complex data content into a hash result of a fixed length, the system can efficiently compare the differences between versions while avoiding the resource consumption caused by directly comparing data contents.
[0039] After calculating the hash values of all data versions, the system will compare these hash values. If there are two or more different hash values, it indicates that there are essential differences in the data content, and the system will determine these versions as conflicting versions accordingly. On the contrary, if all hash values are the same, it can be considered that these data versions are consistent in content and there are no conflicts.
[0040] By first calculating the hash values of multiple latest data versions respectively and then judging whether the data conflicts based on whether the hash values are consistent, the system can accurately identify and efficiently compare the content of data versions at extremely low computational cost. This mechanism avoids byte-by-byte comparison of the complete data content, greatly improving the performance while ensuring the accuracy of conflict detection. It is precisely this hash value-based judgment mechanism that enables the system to quickly detect potential conflicts in a multi-node environment, providing a reliable basis for subsequent conflict resolution and further enhancing the data consistency control ability.
[0041] Please refer to Figure 3 , another embodiment of the data consistency processing method in the embodiments of the present application includes: 301. When a node triggers an update operation on the target data, generate a data version; 302. Periodically perform conflict detection on all data versions of the target data within a preset time window; In this embodiment, steps 301-302 are similar to steps 101-102 in the foregoing embodiment, and will not be elaborated here.
[0042] 303. When detecting conflicting versions, determine a conflict resolution strategy, and the conflict resolution strategy includes giving priority to the latest version, giving priority to node reliability, and conflict merging; When the system discovers conflicting versions during conflict detection, it will automatically determine an applicable conflict resolution strategy. The system presets multiple strategy types, including the "give priority to the latest version" strategy (i.e., taking the version with the latest timestamp as the standard), the "give priority to node reliability" strategy (i.e., selecting the node version with high historical stability or credibility), and the "conflict merging" strategy (i.e., fusing the valid content of multiple versions to generate a new merged version), so as to flexibly handle different types of conflict situations.
[0043] 304. Determine the priorities of the nodes storing the conflicting versions according to the conflict resolution strategy, and determine the highest priority; The system evaluates the priorities of all nodes storing the conflicting versions according to the selected conflict resolution strategy. The evaluation dimensions may include the version generation time, the historical stability of the nodes, the network response situation, etc., and finally determine the node with the highest priority as the basis for subsequent target version selection.
[0044] Specifically, when the conflict resolution strategy is "latest version first", the priorities of the nodes storing the conflicting versions are determined according to the timestamps. When the conflict resolution strategy is "node reliability first", the reliability of the nodes storing the conflicting versions is calculated based on network latency, processor load, and data access frequency, and the priorities of the nodes storing the conflicting versions are determined according to the reliability. When the conflict resolution strategy is "conflict merging", it is determined that the priorities of the nodes storing the conflicting versions are the same.
[0045] After the system identifies multiple conflicting versions based on the conflict detection results, it will evaluate and determine the priorities of the nodes where each conflicting version is located according to a preset or dynamically selected conflict resolution strategy. If the "latest version first" strategy is selected, the system will compare the timestamps of each version, and the node where the newer version is located will be given a higher priority; if the "node reliability first" strategy is adopted, the system will comprehensively evaluate the nodes from multiple dimensions, including network latency (the lower the latency, the higher the reliability), processor load (the lower the load, the more sufficient the node's processing capacity), and data access frequency (the nodes with frequent access are more likely to maintain the latest and valid data). These metrics jointly determine the reliability score of the node and allocate priorities accordingly; if the "conflict merging" strategy is adopted, it is default that the priorities of all nodes storing the conflicting versions are the same, ensuring that their data are all adopted during subsequent merging processing.
[0046] 305. Determine the target version through the node with the highest priority. 306. Broadcast the target version.
[0047] In this embodiment, steps 305-306 are similar to steps 104-105 in the foregoing embodiment, and will not be elaborated here.
[0048] In this embodiment, when the system detects a data version conflict, it first selects a suitable conflict resolution method through a preset strategy, and then evaluates the node priorities according to the strategy logic, so as to determine the authoritative version through the node with the highest priority. It is precisely because the key step of "determining the conflict resolution strategy" introduces diverse and controllable judgment bases that the system can be highly flexible and accurate when facing different types of conflicts, ensuring that the selection of the target version takes into account the update timeliness, node credibility, and data integrity, thus significantly improving the system's conflict handling ability and data consistency level in complex concurrent scenarios.
[0049] Please refer to Figure 4 In some embodiments of the present application, step 303 of determining the conflict resolution strategy in the foregoing embodiment may include the following steps: 401. Determine the system state of the distributed storage system. When the distributed storage system detects multiple conflicting versions, the system first collects the current system running status information, including but not limited to metrics such as the overall network load, node online status, request distribution, and write pressure. These status information are used to reflect the current working environment of the system and help in the adaptive selection of subsequent strategies.
[0050] 402. Determine the conflict type of the conflicting versions; The system will further analyze the differences between the conflicting versions to identify their conflict types. For example, the conflict may be only concurrent writes in terms of time, may involve field-level modification conflicts, or even complex conflicts with inconsistent content structures. The determination of the conflict type is usually based on methods such as version difference comparison and field change trajectory analysis.
[0051] 403. Determine the conflict resolution strategy according to the system status and / or conflict type.
[0052] The system comprehensively considers the determined system status and conflict type and dynamically selects the most suitable conflict resolution strategy. If the system is in a high-load or node-frequently-fluctuating state, it may tend to choose "latest version first" to reduce the merging cost; if the conflict type is complex or the field coverage is strong, the "conflict merging" strategy may be given priority; while in the scenario where the node status is stable and the reliability difference is obvious, the system may choose the "node reliability first" strategy to ensure consistency and accuracy.
[0053] In this embodiment, since the system can simultaneously grasp the current running status and the type characteristics of the conflict itself before determining the conflict resolution strategy, it can dynamically match the most suitable strategy according to the actual situation. It is precisely the key step of "determining the conflict resolution strategy according to the system status and / or conflict type" that makes the strategy selection highly adaptable, so that the conflict handling can not only take into account performance in a high-pressure environment but also maintain accuracy in complex conflicts, significantly improving the flexibility and robustness of the distributed storage system in handling conflicts.
[0054] Please refer to Figure 5 , in some embodiments of the present application, step 402 in the above embodiment of determining the conflict type of the conflicting versions may include the following steps: 501. Obtain the number of conflicting nodes and the distribution of conflicting nodes of the conflicting versions; After the distributed storage system detects the existence of conflicting versions, it will count the number of nodes participating in the conflict and analyze the distribution of these nodes in the entire network topology. The number of nodes is used to measure the scale of the conflict, and the geographical or logical distribution of the nodes (such as whether they are concentrated in a certain network area or span multiple data centers) helps to judge the diffusion range and potential causes of the conflict.
[0055] 502. Determine the conflict type based on the number of conflict nodes and the conflict node distribution.
[0056] The system classifies the conflict type according to the obtained number of conflict nodes and the distribution pattern. For example, when the number of conflict nodes is small and concentrated, it can be judged as a conflict caused by local concurrent writing; while when the number of nodes is large and distributed across regions, it may belong to a complex conflict caused by system-level synchronization delay or wide spread. Based on this, the system marks the conflict type as "local concurrent conflict", "wide-area synchronization conflict", etc., to guide subsequent policy selection.
[0057] In this embodiment, since the system can judge the nature of the conflict by analyzing the number and distribution of conflict nodes, it can distinguish conflicts that are structurally similar but have different causes. It is precisely the key step of "determining the conflict type based on the number of conflict nodes and the conflict node distribution" that enables the system to clarify the complexity and cause pattern of the conflict, thus providing a clear basis for subsequent policy matching and improving the accuracy and efficiency of conflict handling.
[0058] Please refer to Figure 6 , an embodiment of the distributed storage system in the embodiment of the present application includes: A generation unit 601, configured to generate a data version when a node triggers an update operation on target data; A detection unit 602, configured to periodically perform conflict detection on all data versions of the target data within a preset time window; A first determination unit 603, configured to determine the priority of the node storing the conflict version and determine the highest priority when a conflict version is detected; A second determination unit 604, configured to determine the target version through the node with the highest priority; A broadcast unit 605, configured to broadcast the target version.
[0059] In this embodiment, first, the generation unit 601 generates a data version for the update of the target data, and the detection unit 602 discovers that multiple nodes may have conflicting updates to the same data through periodic conflict detection. At this time, the first determination unit 603 determines which node's version is more authoritative by assigning priorities to each node, ensuring that the system selects a most reliable version. Then, the second determination unit 604 determines the version held by the node with the highest priority as the target version. Finally, the broadcast unit 605 broadcasts the target version to all relevant nodes to ensure data consistency among the nodes. Therefore, even if concurrent updates occur, the distributed storage system can ultimately resolve conflicts through priority management and version synchronization, thus ensuring the data consistency of the system.
[0060] Optionally, the detection unit 602 includes: An acquisition module, configured to, when a detection period arrives, respectively acquire the latest data version from each node storing target data; A detection module, configured to perform conflict detection on the multiple latest data versions acquired.
[0061] Optionally, the detection module is specifically configured to: Calculate the hash values of the multiple latest data versions respectively; Judge whether there are conflict versions according to the calculated hash values.
[0062] Optionally, the distributed storage system further includes: A third determination unit, configured to determine a conflict resolution strategy, where the conflict resolution strategy includes latest version first, node reliability first, and conflict merging; The first determination unit 603 is specifically configured to: Determine the priorities of the nodes storing conflict versions according to the conflict resolution strategy.
[0063] Optionally, the third determination unit includes: A first determination module, configured to determine the system state of the distributed storage system; A second determination module, configured to determine the conflict type of the conflict versions; A third determination module, configured to determine the conflict resolution strategy according to the system state and / or the conflict type.
[0064] Optionally, the second determination module is specifically configured to: Acquire the number of conflict nodes and the distribution of conflict nodes of the conflict versions; Determine the conflict type according to the number of conflict nodes and the distribution of conflict nodes.
[0065] Optionally, the first determination unit 603 is specifically configured to: When the conflict resolution strategy is latest version first, determine the priorities of the nodes storing conflict versions according to the time stamp; When the conflict resolution strategy is node reliability first, calculate the reliability of the nodes storing conflict versions according to network latency, processor load, and data access frequency, and determine the priorities of the nodes storing conflict versions according to the reliability; When the conflict resolution strategy is conflict merging, determine that the priorities of the nodes storing conflict versions are the same.
[0066] In this embodiment, the functions of each unit and module correspond to the steps in the foregoing Figures 1 to 5 illustrated embodiment, and will not be elaborated herein.
[0067] Please refer to Figure 7 , another embodiment of the distributed storage system in the embodiment of the present application includes: A processor 701, a memory 702, an input / output unit 703, and a bus 704; The processor 701 is connected to the memory 702, the input / output unit 703, and the bus 704; A program is stored in the memory 702, and the processor 701 calls the program to execute Figures 1 to 5 The steps in the illustrated embodiment.
[0068] In this embodiment, the functions of the processor 701 correspond to the steps in the foregoing Figures 1 to 5 Illustrated embodiment, and will not be elaborated here.
[0069] This embodiment of the present application also provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, the computer is caused to execute the foregoing Figures 1 to 5 Method in any possible implementation manner.
[0070] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated here.
[0071] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the displayed or discussed mutual coupling or direct coupling or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0072] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0073] In addition, the functional units in each embodiment of the present application can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0074] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs.
Claims
1. A data consistency processing method, applied to a distributed storage system, characterized in that: include: When a node triggers an update operation on the target data, a data version is generated; Periodically performing conflict detection on all data versions of the target data within a preset time window; When a conflicting version is detected, determining the priority of the node storing the conflicting version, and determining the highest priority; Determine the target version through the node with the highest priority; The target version is broadcasted.
2. The method according to claim 1, characterized in that The periodically performing conflict detection on all data versions of the target data within a preset time window includes: When the detection cycle arrives, the latest data version is obtained from each node storing the target data; Perform conflict detection on the multiple latest data versions obtained.
3. The method according to claim 2, characterized in that The conflict detection of the acquired multiple latest data versions includes: Calculate the hash values of the multiple latest data versions respectively; Determine whether there is a conflicting version based on the calculated hash value.
4. The method according to any one of claims 1 to 3, characterized in that Before determining the priority of the node storing the conflicting version, the method further includes: Determine a conflict resolution strategy, wherein the conflict resolution strategy includes latest version priority, node reliability priority and conflict merging; Determining the priority of the node storing the conflicting version includes: The priority of the node storing the conflicting version is determined according to the conflict resolution strategy.
5. The method according to claim 4, characterized in that Determining the conflict resolution strategy includes: Determining a system status of the distributed storage system; determining a conflict type of the conflicting versions; A conflict resolution strategy is determined according to the system state and / or the conflict type.
6. The method according to claim 5, characterized in that Determining the conflict type of the conflicting version includes: Obtain the number of conflicting nodes and the distribution of conflicting nodes of the conflicting version; The conflict type is determined according to the number of conflicting nodes and the distribution of conflicting nodes.
7. The method according to claim 4, characterized in that Determining the priority of the node storing the conflicting version according to the conflict resolution strategy includes: When the conflict resolution strategy is that the latest version is given priority, the priority of the node storing the conflicting version is determined according to the timestamp; When the conflict resolution strategy is node reliability priority, the reliability of the node storing the conflicting version is calculated according to network delay, processor load and data access frequency, and the priority of the node storing the conflicting version is determined according to the reliability; When the conflict resolution strategy is conflict merging, it is determined that the priorities of the nodes storing the conflicting versions are consistent.
8. A distributed storage system, characterized in that: include: A generation unit, used for generating a data version when a node triggers an update operation on target data; A detection unit, used to periodically perform conflict detection on all data versions of the target data within a preset time window; A first determining unit, configured to determine the priority of the node storing the conflicting version when a conflicting version is detected, and determine the highest priority; A second determination unit, configured to determine a target version through a node with the highest priority; The broadcast unit is used to broadcast the target version.
9. A distributed storage system, characterized in that: include: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; A program is stored in the memory, and the processor calls the program to execute the method according to any one of claims 1 to 7. 10 . A computer-readable storage medium having a program stored thereon, wherein when the program is executed on a computer, the computer is caused to execute the method according to claim 1 .
Citation Information
Patent Citations
Data consistency method and device, equipment and storage medium
CN117271537A
Intelligent networking distributed storage interaction system and method based on multi-cabin cooperation
CN119292112A
Data conflict processing method, electronic equipment, readable medium and program product
CN119336524A