Merging conflict data processing method capable of realizing grouping sorting
By detecting the multi-dimensional relationship between conflicting data and sorting it, the problem of inefficient conflicting data positioning and processing in data merging is solved, and more efficient conflicting data repair and reduced invalid repair rate is achieved.
Patent Information
- Application Number
- CN202510013859.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-27
AI Technical Summary
During the data merging process, the location and processing of conflicting data are inefficient and high invalid repair rates, especially for data with complex dependencies.
By detecting the dependencies, source relationships and similarities between conflicting data, establishing multi-dimensional directed edge connections, performing weight sorting and prioritization sorting, and then grouping and sorting data blocks to determine the order of repairing conflicting data.
It improves the efficiency of conflict data repair processing, reduces the invalid repair rate, and ensures the smooth progress of the data merging process.
Smart Images

Figure CN120045157A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of data processing methods, and particularly relates to a method for processing merged conflict data that can achieve grouped sorting. Background Art
[0002] In the process of data merging or updating, due to various reasons, problems such as data conflicts often occur. When a conflict occurs, it is necessary to locate and process the conflict data one by one in a timely manner to ensure the smooth progress of processes such as data merging. During this process, for data that is dependent on each other, it is generally necessary to locate as much as possible the data in the conflict data that is dependent on other data, and then process the data that depends on this data to avoid invalidating the processing of the data that depends on this data during the subsequent processing of the dependent data. Currently, it is mainly monitored and screened by the system, lacking a relatively efficient processing method. Summary of the Invention
[0003] The purpose of the present invention is to provide, based on actual needs, a method for processing conflict data that can quickly locate the dependent data based on the dependency relationship between conflict data and determine its repair order, so as to improve the efficiency of conflict data repair processing and reduce the invalid repair rate, namely, a method for processing merged conflict data that can achieve grouped sorting.
[0004] To achieve the above purpose, the present invention adopts the following technical solutions.
[0005] A method for processing merged conflict data that can achieve grouped sorting includes the following steps:
[0006] Step 1: A step for completing the detection and extraction of conflict data
[0007] Specifically, when a conflict occurs in a data merging operation, obtain the conflict information that occurs in the database and combine it into a node data block;
[0008] Step 2: Multidimensional directed edge weight processing based on the data block dependency relationship
[0009] Specifically, define a set of node data blocks {dis}, define a unique identifier for each node data block in the set of node data blocks {dis} (in this embodiment, establish its unique identifier based on the database address of the node data block); analyze the dependency relationship, source attribution relationship, and similarity of each data block dis in the set of node data blocks {dis}, and establish a multidimensional directed edge connection based on the dependency relationship, source attribution relationship, and similarity.
[0010] The dependency-based relationship means that if there is a direct dependency relationship between two data blocks, a one-dimensional unidirectional edge with a weight value w=(0,1] is used to connect them; if there is an indirect dependency relationship between two data blocks, a one-dimensional unidirectional edge with a weight value w∈(0,1) is used to connect them according to the number of intermediate data blocks between the two data blocks.
[0011] The source-based relationship means that if two data blocks are derived from the same source data packet, a two-dimensional bidirectional edge with a weight value w′ = 1 is used to connect them.
[0012] The similarity-based relationship means that the similarity between two data blocks is detected, and a three-dimensional bidirectional edge with a weight value w″∈[0,1] is used to connect them according to the similarity.
[0013] Based on the foregoing multi-dimensional directed edge connections, a directed edge set T all ={T, T′, T″} is established, where T refers to a one-dimensional directed edge, T′ refers to a two-dimensional directed edge, and T″ refers to a three-dimensional directed edge.
[0014] Step 3: Priority sorting of directed edges based on multi-dimensional directed edge weights
[0015] Statistical directed edge set T all ={T, T′, T″}, and the weights of the edges are sorted in the order of one-dimensional directed edges, two-dimensional directed edges, and three-dimensional directed edges.
[0016] After sorting according to the weight values of each directed edge, a directed edge priority sequence Q T ;
[0017] Step 4: Grouping processing of data blocks based on data block dependencies
[0018] Initialize, and each node data block dis is used as a separate data block group.
[0019] Select directed edges from the directed edge priority sequence Q T in sequence from front to back. If the two node data blocks corresponding to the directed edge are already in the same data block group, then loop to select the next one; if the two node data blocks corresponding to the directed edge are not in the same data block group, then merge the data block groups where the two node data blocks are located.
[0020] After the loop ends, merge the remaining data block groups with only a single node data block into the same data block group, and there is no directed edge connection between the node data blocks in this data block group and any other node data blocks.
[0021] Step 5: Sorting of data groups based on data block dependencies
[0022] For each data block group, calculate the directed edge weights of all the data blocks within it respectively, and separately count the cumulative values of the one-dimensional directed edge weights, two-dimensional directed edge weights, and three-dimensional directed edge weights of each data block group. Then, sort the data block groups according to the magnitudes of the cumulative weights.
[0023] Determine the sorting of the data block groups based on the foregoing steps;
[0024] Step 6: Sort the data blocks within the data group based on the data block dependency
[0025] For the data blocks within each data block group, sort the data blocks according to their influence weights to obtain the first sorting sequence L of the data blocks within the group 1 ;
[0026] For the first sorting sequence L of the data block group 1 , traverse the data blocks in sequence from front to back, and analyze whether there is a one-dimensional directed edge pointing to another data block p' m located behind it within the group m ;
[0027] If there is, then move the corresponding data block p' m to the front of data block P m in descending order of the one-dimensional directed edge weights, and feed back the initial data block of the sequence for re-analysis;
[0028] If not, analyze the next data block until all data blocks are traversed;
[0029] Output the sorting result, and process the conflicting data in sequence according to the sorting order.
[0030] For the further improvement or specific implementation steps of the foregoing data analysis method for handling merging conflicts that can achieve grouped sorting,
[0031] In step 3, sorting the weights in the order of one-dimensional directed edges, two-dimensional directed edges, and three-dimensional directed edges specifically means: sorting according to the magnitudes of the one-dimensional directed edge weights, with the larger the one-dimensional directed edge weight, the more forward the sorting. If the one-dimensional directed edge weights are the same, then sort according to the magnitudes of the two-dimensional directed edge weights, with the larger the two-dimensional directed edge weight, the more forward the sorting. If the two-dimensional directed edge weights are the same, then sort according to the magnitudes of the three-dimensional directed edge weights, with the larger the three-dimensional directed edge weight, the more forward the sorting.
[0032] For the further improvement or specific implementation steps of the foregoing data analysis method for handling merging conflicts that can achieve grouped sorting, in step 5, sorting the data block groups according to the magnitudes of the cumulative weights specifically means:
[0033] For one-dimensional directed edges, the larger the accumulated weight value, the higher the sorting priority. If the accumulated weight values of one-dimensional directed edges are the same, then sort according to the size of the accumulated weight value of two-dimensional directed edges. The larger the accumulated weight value of two-dimensional directed edges, the higher the sorting priority. If the accumulated weight values of two-dimensional directed edges are the same, then sort according to the size of the accumulated weight value of three-dimensional directed edges. The larger the accumulated weight value of three-dimensional directed edges, the higher the sorting priority.
[0034] For further improvement or specific implementation steps of the above-mentioned data analysis method for pending merge conflicts that can achieve grouped sorting, step 3 further includes that if the sizes of multi-dimensional directed edge weights are all the same, then sort according to the size of the data block validity index. If the data block validity indexes are also the same, then sort according to the size of the data block importance index. The data block validity index and the data block importance index are given through statistical analysis or expert scoring.
[0035] For further improvement or specific implementation steps of the above-mentioned data analysis method for pending merge conflicts that can achieve grouped sorting, in step 1, the combination into node data blocks specifically means: cloning the database to the local, then tracing back from the last merge forward, traversing each merge history. If a certain merge record has two parent merges, it means that this merge is a merge record that has occurred. Then obtain the IDs of the two parent merges of this merge, re-execute the merge operation, analyze the merge result, determine whether this merge has conflicts, obtain the data blocks before and after the merge that have conflicts, including the data block ous before the merge and the data block ths after the merge, as well as the corresponding original data and conflict data block bas for the corresponding data. Combine the data block ous, the data block ths, and the data block bas to obtain the node data block dis.
[0036] For further improvement or specific implementation steps of the above-mentioned data analysis method for pending merge conflicts that can achieve grouped sorting, in step 6, the influence weight i of the i-th data block p
[0037]
[0038] where J refers to the total number of one-dimensional directed edges pointing to the data block, K refers to the total number of two-dimensional directed edges pointing to the data block, L refers to the total number of three-dimensional directed edges pointing to the data block, w j.i refers to the weight value of the j-th one-dimensional directed edge pointing to the data block p i , w' k.i refers to the weight value of the k-th two-dimensional directed edge pointing to the data block p i , w l ' . ' i refers to the weight value of the l-th three-dimensional directed edge pointing to the data block p i . Brief Description of the Drawings
[0039] Figure 1 It is a schematic diagram of the processing of the weighted value of the directed edge of the data block. Detailed Implementation Manner
[0040] The following describes the present invention in detail with reference to specific embodiments.
[0041] The data analysis method for processing merge conflicts with group sorting that can be achieved by the present invention is mainly used to provide a more efficient and effective method for sorting and optimizing conflict data, optimizing sorting according to the dependencies and similarities between conflict data, determining the conflict processing order, so as to more efficiently complete the conflict data processing during operations such as large-capacity processing and merging, reduce the secondary changes caused by the changes of the data on which the processed data depends, resulting in the invalidation of conflict processing and wasting system computing resources and other problems.
[0042] The data analysis method for processing merge conflicts with group sorting that can be achieved by the present invention mainly includes the following steps:
[0043] Step 1: The step for completing the detection and extraction of conflict data
[0044] This step mainly relies on the database self-checking system to locate and screen the conflict data in the database according to various identification information during the data merging process, and extract the data involved in the data conflict for subsequent independent repair, prevent interference with normal data, and at the same time reduce the data processing volume and the system computing power pressure. The specific content includes:
[0045] When a conflict occurs in the data merging operation, first obtain the conflict information that occurs in the database, clone the database to the local, and then trace back from the last merge forward, traverse each merge history. If a merge record has two parent merges, it means that this merge is a merge record that has occurred a merge. Then obtain the IDs of the two parent merges of this merge, re-execute the merge operation, analyze the merge result, determine whether this merge has a conflict, obtain the data blocks before and after the merge that have a conflict, including the data block ous before the merge and the data block ths after the merge, as well as the corresponding original data and the conflict data block bas of the corresponding data, and combine the data block ous, the data block ths, and the data block bas to obtain the node data block dis;
[0046] Step 2: Multidimensional directed edge weight processing based on data block dependency relationships
[0047] As Figure 1 shown, this step is mainly used to determine the dependency weights between data blocks according to the direct dependency relationships, homologous relationships, and similarity relationships between the screened conflict data blocks. The specific steps include:
[0048] Define a set of node data blocks {dis}, and define a unique identifier for each node data block in the set of node data blocks {dis} (in this embodiment, the unique identifier is established based on the database address of the node data block); analyze the dependency relationship, source attribution relationship, and similarity of each data block dis in the set of node data blocks {dis}, and establish multi-dimensional directed edge connections based on the dependency relationship, source attribution relationship, and similarity;
[0049] Based on the dependency relationship means that if there is a direct dependency relationship between two data blocks, a one-dimensional unidirectional edge connection with a weight value w=(0,1] is assigned to them; if there is an indirect dependency relationship between two data blocks, a one-dimensional unidirectional edge connection with a weight value w∈(0,1) is assigned according to the number of intermediate data blocks between the two data blocks;
[0050] Based on the source attribution relationship means that if two data blocks are derived from the same source data packet, a two-dimensional bidirectional edge connection with a weight value w′ = 1 is assigned to them;
[0051] Based on the similarity means that the similarity between two data blocks is detected, and a three-dimensional bidirectional edge connection with a weight value w″∈[0,1] is assigned according to their similarity;
[0052] Based on the foregoing multi-dimensional directed edge connections, establish a set of directed edges T all ={T, T′, T″}, where T refers to a one-dimensional directed edge, T′ refers to a two-dimensional directed edge, and T″ refers to a three-dimensional directed edge;
[0053] Step 3: Priority sorting of directed edges based on multi-dimensional directed edge weights
[0054] Statistical set of directed edges T all ={T, T′, T″} the weights of the edges, and perform weight sorting in the order of one-dimensional directed edges, two-dimensional directed edges, and three-dimensional directed edges. Specifically:
[0055] Sort according to the magnitude of the one-dimensional directed edge weight. The larger the one-dimensional directed edge weight, the higher the sorting. If the one-dimensional directed edge weights are the same, then sort according to the magnitude of the two-dimensional directed edge weight. The larger the two-dimensional directed edge weight, the higher the sorting. If the two-dimensional directed edge weights are the same, then sort according to the magnitude of the three-dimensional directed edge weight. The larger the three-dimensional directed edge weight, the higher the sorting;
[0056] After sorting according to the weight magnitude of each directed edge, according to the sorting result, establish a directed edge priority sequence Q T ;
[0057] If the multi-dimensional directed edge weights are all the same, sort according to the size of the data block validity index. If the data block validity indexes are also the same, sort according to the size of the data block importance index. The data block validity index and the data block importance index are given through statistical analysis or expert scoring;
[0058] Step 4: Data grouping processing based on data block dependencies
[0059] Initialize, and regard each node data block dis as a separate data block group.
[0060] Select directed edges from the directed edge priority sequence Q T in sequence from front to back. If the two node data blocks corresponding to the directed edge are already in the same data block group, then loop to select the next one; if the two node data blocks corresponding to the directed edge are not in the same data block group, then merge the data block groups where the two node data blocks are located;
[0061] After the loop ends, merge the remaining data block groups with only single node data blocks into the same data block group. There are no directed edge connections between the node data blocks in this data block group and any other node data blocks;
[0062] Step 5: Sorting of data groups based on data block dependencies
[0063] For each data block group, calculate the directed edge weights of all the data blocks inside it respectively, and respectively count the accumulated values of one-dimensional directed edge weights, two-dimensional directed edge weights, and three-dimensional directed edge weights of each data block group, and sort the data block groups according to the size of the accumulated weight values.
[0064] Specifically: The larger the accumulated value of one-dimensional directed edge weights, the higher the ranking. If the accumulated values of one-dimensional directed edge weights are the same, then sort according to the size of the accumulated value of two-dimensional directed edge weights. The larger the accumulated value of two-dimensional directed edge weights, the higher the ranking. If the accumulated values of two-dimensional directed edge weights are the same, then sort according to the size of the accumulated value of three-dimensional directed edge weights. The larger the accumulated value of three-dimensional directed edge weights, the higher the ranking;
[0065] Determine the sorting of the data block groups based on the foregoing steps;
[0066] Step 6: Sorting of data blocks within a data group based on data block dependencies
[0067] For the data blocks within each data block group, sort the data blocks according to their influence weights to obtain the first sorting sequence L of the data blocks within the group. 1 ;
[0068] Among them, the influence weight i of the i-th data block p
[0069] J refers to the total number of one-dimensional directed edges pointing to the data block, K refers to the total number of two-dimensional directed edges pointing to the data block, L refers to the total number of three-dimensional directed edges pointing to the data block, and w j.i refers to the weight of the j-th one-dimensional directed edge pointing to the data block p i ; w' k.i refers to the weight of the k-th two-dimensional directed edge pointing to the data block p i ; w l ' . ' i refers to the weight of the l-th three-dimensional directed edge pointing to the data block p i ;
[0070] For the first sorting sequence L of the data block group 1 , traverse the data blocks in sequence from front to back, and analyze whether there is a one-dimensional directed edge of the data block p m pointing to another data block p' m behind it in the group;
[0071] If so, move the corresponding data block p' m to the front of the data block P m in the order of decreasing one-dimensional directed edge weight, and feed back the initial data block of the sequence for re-analysis;
[0072] If not, analyze the next data block until all data blocks are traversed;
[0073] Output the sorting result, and process the conflicting data in sequence according to the sorting order.
[0074] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than limiting the protection scope of the present invention. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the essence and scope of the technical solutions of the present invention.
Claims
1. A method for processing merge conflict data that can achieve group sorting, characterized in that: The steps include: Step 1: Steps for completing conflict data detection and extraction Specifically, when a conflict occurs in a data merge operation, the conflict information in the database is obtained and combined into a node data block; Step 2: Multi-dimensional directed edge weight processing based on data block dependency Specifically, it means: defining a node data block set {dis}, defining a unique identifier for each node data block in the node data block set {dis}; analyzing the dependency, source-attribute relationship and similarity of each data block dis in the node data block set {dis}, and establishing a multi-dimensional directed edge connection based on the dependency, source-attribute relationship and similarity; Based on dependency, if there is a direct dependency between two data blocks, a one-dimensional unidirectional edge connection with a weight of w=(0,1] is assigned to them; if there is an indirect dependency between two data blocks, a one-dimensional unidirectional edge connection with a weight of w∈(0,1) is assigned to them according to the number of indirect data blocks between the two data blocks; Based on the source-belonging relationship, it means that if two data blocks originate from the same source data packet, they are connected by a two-dimensional bidirectional edge with a weight of w′=1; Similarity-based means: the similarity between two data blocks is detected, and a three-dimensional bidirectional edge connection is assigned with a weight w″∈[0,1] according to their similarity; Based on the aforementioned multi-dimensional directed edge connections, a directed edge set T is established. all ={T,T′,T″}, where T refers to a one-dimensional directed edge, T′ refers to a two-dimensional directed edge, and T″ refers to a three-dimensional directed edge; Step 3: Prioritize directed edges based on multi-dimensional directed edge weights Statistical directed edge set T all = The weights of the edges in {T, T′, T″} are sorted in the order of one-dimensional directed edges, two-dimensional directed edges, and three-dimensional directed edges; After sorting according to the weight of each directed edge, a directed edge priority sequence Q is established based on the sorting results. T ; Step 4: Data block grouping based on data block dependencies Initialize and treat each node data block dis as a separate data block group. From the directed edge priority sequence Q T Select the directed edges from front to back in sequence. If the two node data blocks corresponding to the directed edge are already in the same data block group, then select the next one in a loop; if the two node data blocks corresponding to the directed edge are not in the same data block group, then merge the data block groups where the two node data blocks are located; After the loop is finished, the remaining data block groups that only have a single node data block are merged into the same data block group. The node data blocks in the data block group do not have directed edge connections with any other node data blocks. Step 5: Sorting data groups based on data block dependencies For each data block group, the directed edge weights of all the data blocks inside it are calculated respectively, and the one-dimensional directed edge weight accumulation value, the two-dimensional directed edge weight accumulation value, and the three-dimensional directed edge weight accumulation value of each data block group are counted respectively, and the data block groups are sorted according to the size of the weight accumulation value. Determine the order of the data block groups based on the above steps; Step 6: Sort the data blocks within the data group based on data block dependencies For each data block group, sort the data blocks according to their influence weights to obtain a sorted sequence L1 of the data block group within the group; For the data block group, sort the sequence L1 once, traverse the data blocks in order from front to back, and analyze the data block p m Is there any other data block p' pointing to the back of the group? m One-dimensional directed edge of ; If it exists, the corresponding data block p′ is placed in the order of the one-dimensional directed edge weight from large to small m Move to data block P m The front side, and feed back the initial data block of the sequence for re-analysis; If it does not exist, analyze the next data block until all data blocks are traversed; Output the sorting results and complete the processing of conflicting data in sequence according to the sorting order.
2. The method for analyzing pending merge conflict data capable of realizing group sorting according to claim 1, characterized in that: The step 3, sorting the weights in the order of one-dimensional directed edges, two-dimensional directed edges, and three-dimensional directed edges, specifically means: sorting according to the size of the one-dimensional directed edge weights, the larger the one-dimensional directed edge weights, the higher the ranking; if the one-dimensional directed edge weights are the same, sorting according to the size of the two-dimensional directed edge weights, the larger the two-dimensional directed edge weights, the higher the ranking; if the two-dimensional directed edge weights are the same, sorting according to the size of the three-dimensional directed edge weights, the larger the three-dimensional directed edge weights, the higher the ranking.
3. The method for analyzing pending merge conflict data capable of realizing group sorting according to claim 1, characterized in that: The step 5, sorting the data block groups according to the size of the weight accumulation value, specifically refers to: The larger the one-dimensional directed edge weight cumulative value is, the higher the ranking is. If the one-dimensional directed edge weight cumulative values are the same, they are sorted according to the size of the two-dimensional directed edge weight cumulative value. The larger the two-dimensional directed edge weight cumulative value is, the higher the ranking is. If the two-dimensional directed edge weight cumulative values are the same, they are sorted according to the size of the three-dimensional directed edge weight cumulative value. The larger the three-dimensional directed edge weight cumulative value is, the higher the ranking is.
4. The method for analyzing pending merge conflict data capable of realizing group sorting according to claim 1, characterized in that: The step 3 also includes, if the weights of the multi-dimensional directed edges are the same, then sorting is based on the size of the data block validity index; if the data block validity indexes are also the same, then sorting is based on the size of the data block importance index, and the data block validity index and the data block importance index are given by statistical analysis or expert scoring.
5. The method for analyzing pending merge conflict data capable of realizing group sorting according to claim 1, characterized in that: The combination into node data blocks in step 1 specifically refers to: cloning the database locally, and then traversing each merge history from the last merge, if a merge record has two parent merges, it means that this merge is a merge record in which a merge has occurred, and then obtaining the IDs of the two parent merges of this merge, re-executing the merge operation, analyzing the merge result, determining whether this merge has a conflict, obtaining the data blocks before and after the merge in which the conflict occurs, including the data block ous before the merge and the data block ths after the merge, as well as the corresponding data corresponding to the original data and the conflicting data block bas, and combining the data block ous, the data block ths, and the data block bas to obtain the node data block dis.
6. The method for analyzing pending merge conflict data capable of realizing group sorting according to claim 1, characterized in that: In step 6, the i-th data block p i The influence weight Where J refers to the total number of one-dimensional directed edges pointing to the data block, K refers to the total number of two-dimensional directed edges pointing to the data block, L refers to the total number of three-dimensional directed edges pointing to the data block, and w j.i refers to the data block p i The weight of the j-th one-dimensional directed edge, w′ k.i refers to the data block p i The weight of the k-th two-dimensional directed edge, w l ' . ' i refers to the data block p i The weight of the lth three-dimensional directed edge of .