Detection method and system for preventing DDL repetition in data synchronization

By calculating semantic similarity, operation similarity value and operation data difference index during data synchronization, evaluating data synchronization quality, and optimizing duplicate DDL information, the problem of repeated execution of DDL statements during data synchronization in large enterprise databases is solved, and the efficiency and consistency of data synchronization are improved.

CN120030029AInactive Publication Date: 2025-05-23HANGZHOU KAIYUN JIZHI TECHNOLOGY CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510513755.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-05-23
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the database system of large enterprises, the repeated execution of DDL statements is prone to occur during data synchronization, resulting in waste of resources, inconsistency in data and degradation of performance.

Method used

By obtaining data change records during data synchronization, calculate semantic similarity, operation similarity values ​​and operation data difference index, evaluate the data synchronization quality, and compare it in the synchronization history library to determine and optimize duplicate DDL information.

Benefits of technology

It effectively avoids resource waste and performance degradation caused by repeated DDL execution, improves the accuracy and consistency of data synchronization, and promptly detects potential conflicts and operational semantic correlation problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120030029A_ABST
    Figure CN120030029A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data detection, and particularly discloses a detection method and system for preventing DDL duplication in data synchronization, and the method comprises the steps: obtaining a data change record in a data synchronization process, the data change record comprising data definition language DDL sample set information, an operation timestamp and operation content; acquiring multiple pieces of statement sample information according to the DDL sample set information, and acquiring semantic similarity of each piece of statement sample information; and obtaining an operation identifier corresponding to each piece of statement sample information according to the operation content. According to the method and the device, the problem of insufficient operation analysis in the prior art is comprehensively and deeply solved by acquiring the operation identifier, calculating the operation similarity value, the operation data difference index and the like. Operation key feature information can be accurately extracted, the actual significance and association of data operation are clarified in combination with context identification, and misjudgment caused by inaccurate information extraction and lack of context consideration is avoided.
Need to check novelty before this filing date? Find Prior Art

Claims

1. A detection method for preventing DDL duplication in data synchronization, characterized in that: include: Obtaining data change records during data synchronization, wherein the data change records include data definition language DDL sample set information, operation timestamp, and operation content; Acquire multiple statement sample information according to the data definition language DDL sample set information, and acquire the semantic similarity of each statement sample information; Obtaining an operation identifier corresponding to each statement sample information according to the operation content, and obtaining an operation similarity value according to the operation identifier; Acquire the operation data difference index corresponding to each statement sample information according to the operation timestamp; Obtaining a data synchronization quality assessment value according to the conflict index, semantic similarity, and operation similarity value; Acquire a synchronization history library, and compare it in a preset synchronization history library according to the data synchronization quality assessment value; If there are records with the same data synchronization quality assessment value, the statement sample information is determined to be duplicate DDL information, and synchronization optimization is performed on the statement sample information according to the data synchronization quality assessment value.

2. A detection method for data synchronization and DDL duplication prevention according to claim 1, characterized in that: The step of obtaining a plurality of statement sample information according to the data definition language DDL sample set information and obtaining the semantic similarity of each statement sample information includes: Acquire multiple key fields according to the data definition language DDL sample set information; Each of the key fields is concatenated into a character string according to a preset format, and the character string is standardized to obtain sentence sample information; Acquire matching keywords according to each of the sentence sample information; Obtaining a keyword distance value of a matching keyword corresponding to each of the sentence sample information; The semantic similarity is obtained according to the keyword distance value corresponding to each of the sentence sample information.

3. A detection method for data synchronization and DDL duplication prevention according to claim 1, characterized in that: The step of obtaining an operation identifier corresponding to each statement sample information according to the operation content, and obtaining an operation similarity value according to the operation identifier includes: Acquire key operation feature information according to each of the statement sample information, wherein the key operation feature information includes a plurality of data operation field information and data operation type information; Acquire context identification information according to the data operation field information; Acquire a data identifier according to the context identification information; Acquire the data identification association degree corresponding to the sentence sample information according to the data operation identifier; Acquire up and down operation information according to the data operation type information; Acquire an operation identifier according to the context operation information, Acquire the operation matching degree corresponding to the statement sample information according to the operation identifier; The operation similarity value is obtained according to the data identification association degree and the operation matching degree.

4. A detection method for data synchronization and DDL duplication prevention according to claim 1, characterized in that: The step of obtaining the operation data difference index corresponding to each statement sample information according to the operation timestamp includes: Obtaining the first operation timestamp corresponding to each statement sample information according to the operation timestamp; Obtain the most recent operation timestamp corresponding to each statement sample information according to the operation timestamp Acquire an operation time difference according to the first operation timestamp and the most recent operation timestamp; Obtain the data operation frequency within the operation time difference corresponding to each statement sample information; Get the total number of operations corresponding to each statement sample information; The operation data difference index is obtained according to the total number of operations and the data operation frequency.

5. A detection method for data synchronization and DDL duplication prevention according to claim 4, characterized in that: The step of obtaining the data synchronization quality evaluation value according to the conflict index, semantic similarity and operation similarity value comprises: Obtain the conflict index corresponding to each statement sample information; Obtaining a first evaluation coefficient according to the conflict index; Obtain the semantic similarity corresponding to each sentence sample information; Obtaining a second evaluation coefficient according to the conflict index; Obtain the operation similarity value corresponding to each statement sample information; Obtaining a third evaluation coefficient according to the conflict index; The data synchronization quality evaluation value is obtained according to the conflict index, the first evaluation coefficient, the semantic similarity, the second evaluation coefficient, the operation similarity value and the third evaluation coefficient corresponding to the sentence sample information.

6. A detection method for data synchronization and DDL duplication prevention according to claim 1, characterized in that: The step of performing synchronization optimization on the statement sample information according to the data synchronization quality evaluation value comprises: The statement sample information determined to be duplicate DDL information is collected to obtain a similar statement sample information set; Update the similar sentence sample information set to synchronize the history database, The similar statement sample information sets are removed from the data definition language DDL sample set information to obtain the optimized data definition language DDL sample set information.

7. A detection system for data synchronization and DDL duplication prevention, characterized in that: include: A first acquisition module is used to acquire data change records in the data synchronization process, wherein the data change records include data definition language DDL sample set information, operation timestamp and operation content; A second acquisition module is used to acquire multiple statement sample information according to the data definition language DDL sample set information, and acquire the semantic similarity of each statement sample information; A third acquisition module is used to acquire an operation identifier corresponding to each statement sample information according to the operation content, and acquire an operation similarity value according to the operation identifier; A fourth acquisition module, configured to acquire an operation data difference index corresponding to each statement sample information according to the operation timestamp; A fifth acquisition module, used to acquire a data synchronization quality evaluation value according to the conflict index, the semantic similarity and the operation similarity value; An optimization module, used for acquiring a synchronization history library and performing a comparison in a preset synchronization history library according to the data synchronization quality evaluation value; If there are records with the same data synchronization quality assessment value, the statement sample information is determined to be duplicate DDL information, and synchronization optimization is performed on the statement sample information according to the data synchronization quality assessment value.

8. A detection system for data synchronization and DDL duplication prevention according to claim 7, characterized in that: The second acquisition module includes: A first acquisition unit, configured to acquire a plurality of key fields according to the data definition language DDL sample set information; A processing unit, used for concatenating each of the key fields into a character string according to a preset format, and performing standardization processing on the character string to obtain sentence sample information; A second acquisition unit, configured to acquire matching keywords according to each of the sentence sample information; A first calculation unit, used to obtain a keyword distance value of a matching keyword corresponding to each of the sentence sample information; The third acquisition unit is used to acquire the semantic similarity according to the keyword distance value corresponding to each of the sentence sample information.

9. A detection system for data synchronization and DDL duplication prevention according to claim 7, characterized in that: The third acquisition module includes: A fourth acquisition unit, configured to acquire key operation feature information according to each of the statement sample information, wherein the key operation feature information includes a plurality of data operation field information and data operation type information; A fifth acquiring unit, configured to acquire context identification information according to the data operation field information; A sixth acquisition unit, configured to acquire a data identifier according to the context identification information; A seventh acquisition unit, configured to acquire a data identification association degree corresponding to the sentence sample information according to the data operation identifier; an eighth acquiring unit, configured to acquire up-down and down operation information according to the data operation type information; a ninth acquiring unit, configured to acquire an operation identifier according to the context operation information, A tenth obtaining unit, configured to obtain an operation matching degree corresponding to the sentence sample information according to the operation identifier; The second calculation unit is used to obtain the operation similarity value according to the data identification association degree and the operation matching degree.

10. A detection system for data synchronization and DDL duplication prevention according to claim 7, characterized in that: The fourth acquisition module includes: A first extraction unit, configured to obtain a first operation timestamp corresponding to each statement sample information according to the operation timestamp; The second extraction unit is used to obtain the most recent operation timestamp corresponding to each statement sample information according to the operation timestamp. A third extraction unit, configured to obtain an operation time difference according to the first operation timestamp and the most recent operation timestamp; A fourth extraction unit, used to obtain the data operation frequency within the operation time difference corresponding to each sentence sample information; A fifth extraction unit, used to obtain the total number of operations corresponding to each sentence sample information; The third calculation unit is used to obtain the operation data difference index according to the total number of operations and the data operation frequency.

Citation Information

Cited By

  • Data confusion relation evaluation method and device, equipment and storage medium

    CN120723746A

  • Performing commit operations with recovery on data system

    US12657206B1