Large-data-volume real-time increment capturing and high-speed processing method of comprehensive heterogeneous database architecture

By monitoring the binlog log of the database, real-time synchronization of data changes to the target database and search engines, the problem of data synchronization not real-time and cumbersome configuration under independent data management of the three operators is solved, efficient and accurate data synchronization and processing is achieved, and the stability and adaptability of the system are improved.

CN120234367APending Publication Date: 2025-07-01SPACE VISION (CHONGQING) TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311843628.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the authentication and authentication system platform, the independent data management of the three operators leads to large amount of data, time-consuming query, inability to synchronize in real time, and cumbersome configuration and management work.

Method used

By enabling the binlog function configuration of each database, using the data synchronization tool to listen to the binlog logs of related tables in real time, synchronize data changes to the target database and search engine, and use a distributed data synchronization architecture and machine learning algorithm for abnormal detection to achieve real-time data synchronization between multiple databases.

Benefits of technology

It improves the accuracy and timeliness of data synchronization, simplifies configuration and management work, improves the processing capabilities and stability of the software system, has adaptive adjustment and fault tolerance capabilities, and is adaptive to various database systems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
Patent Text Reader

Abstract

A large-data-volume real-time increment capturing and high-speed processing method of a comprehensive heterogeneous database architecture is applied to an authentication system platform and solves the problems that data of three operators are independently managed, the data volume is large, sources are many, a traditional query method is time-consuming, the data cannot be synchronized in real time, and the accuracy is not high. According to the method, the binlog function configuration of each database is started, the binlog of the related table is monitored in real time by using the data synchronization tool, and the data change is synchronized into the target database and the search engine, so that the accuracy and timeliness of data synchronization are improved, tedious configuration and management work are avoided, and the processing capacity of a software system is improved. Meanwhile, a secondary authentication and authentication interface of a set top box user is provided, and authentication and authentication data, user data, user account opening and user ordering data of three operators are summarized for data report query.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and more specifically, to a method for real-time incremental capture and high-speed processing of a large amount of data with a comprehensive heterogeneous database architecture. Background Art

[0002] In the current authentication and authorization system platform, the three major operators, namely China Mobile, China Telecom, and China Unicom, operate independently. The secondary authentication and authorization interfaces for set-top box users are handled by their respective systems. This independent system model results in the data management platform needing to aggregate all the authentication, authorization data, user data, user account opening, and user subscription data of these three operators for data report queries. Since the data sources of these three operators are numerous and the data volume is huge, with a monthly incremental data of 200 million, this poses a huge challenge to traditional data processing methods. Firstly, traditional data processing methods have the problem of long query time when querying these business data. Due to the huge data volume, traditional query methods require complex calculations and indexing operations, resulting in too long query time and unable to meet the real-time requirements. Secondly, traditional methods have the problem of inability to synchronize data in real time. Due to numerous data sources, traditional data processing methods cannot obtain and process changes in each database in real time, resulting in data asynchronization and affecting the accuracy and real-time nature of the data. Thirdly, traditional data processing methods also have cumbersome configuration and management work. In order to process a large amount of data, a large number of hardware and software resources need to be configured and maintained, increasing the management cost and workload. To solve these problems, we propose a method for real-time incremental capture and high-speed processing of a large amount of data with a comprehensive heterogeneous database architecture. This method can effectively improve the accuracy and timeliness of data synchronization by listening to the binlog logs of the databases and synchronizing the changes in the table data of each database to the target table in real time. This method simplifies the traditional cumbersome configuration and management work and improves the processing ability of the software system. Summary of the Invention

[0003] In the field of information technology and data processing, to achieve the above object, the present disclosure provides a method for real-time incremental capture and high-speed processing of a large amount of data with a comprehensive heterogeneous database architecture, which requires a server for running a data synchronization tool and a search engine database. The server should have sufficient computing power, storage space, and network bandwidth, and be configured with appropriate storage devices, such as disk arrays or cloud storage, to solve the data synchronization and data processing problems in the prior art. The technical solution includes the following steps: S1 Enable the binlog function configuration of each database source. In the database system, binlog is a log file that records all operations for changing database tables.

[0004] By enabling the binlog function configuration, S12 can record all data change operations on database tables, including operations such as insertion, update, and deletion. These operations will be recorded in the binlog, providing a basis for subsequent data synchronization.

[0005] S2 listens to the binlog logs of each database through a data synchronization tool.

[0006] S21 configures the connection information between the data synchronization tool and each database, including the database address, port, username, password, etc., for establishing a connection between the data synchronization tool and the database.

[0007] S22 The data synchronization tool starts a listening process, which is responsible for establishing a connection with the database and starting to listen to the binlog logs, enabling the binlog logging function in the database configuration.

[0008] S23 After establishing a connection with the database and starting to listen, the data synchronization tool reads the binlog log file from the database and parses its content. The binlog log file records all data change operations in the database, including insertion, update, and deletion, etc.

[0009] S24 By parsing the binlog logs, the data synchronization tool extracts the specific information of the data changes, such as the change timestamp, the type of change (insertion, update, or deletion), the changed data content, etc.

[0010] S25 According to business requirements and data processing rules, the data synchronization tool can filter and transform the extracted data change information. For example, it can filter out unnecessary data changes according to specific conditions, or perform format conversion on the data to meet the requirements of the target database.

[0011] S26 Based on the filtered and transformed data change information, the data synchronization tool performs corresponding data synchronization operations in the target database. If data is inserted or updated, the synchronization tool writes the changed data to the target database; if data is deleted, the synchronization tool deletes the corresponding data records from the target database.

[0012] S3 adds, updates, or deletes the corresponding data in the target database according to the changes in the binlog logs.

[0013] The binlog logs of the database record all modification operations on the database, including addition, update, and deletion, etc. The real-time data processing system can obtain real-time information on data changes in the database by listening to these logs.

[0014] When the S32 data synchronization tool detects a change in the binlog of the database, it immediately captures these changes and obtains relevant data change information. This information includes the type of change (insert, update, or delete), the content of the changed data, and the timestamp of the change, etc.

[0015] After receiving the data change information, the S33 data synchronization tool determines how to process these changes based on business requirements and data processing rules. For example, if the business requirement is to synchronize the changes to the target database, and the data processing rules require specific filtering or transformation of the data, then the data synchronization tool will perform corresponding operations on the data in the target database according to these rules.

[0016] The S34 data synchronization tool performs operations such as insert, update, or delete on the corresponding data in the target database according to the captured data change information. This process can effectively filter and transform data according to business requirements to ensure that the data in the target database is consistent with the data in the source database.

[0017] S4 enables the binlog function configuration in the target database.

[0018] S41 Log in to the management interface or command-line tool of the target database to ensure sufficient permissions to make configuration changes.

[0019] S42 In the configuration file of the target database (usually the my.cnf or my.ini file, depending on the database management system used), find the configuration item related to binlog. Ensure that the binlog logging function is enabled, that is, binlog_format is set to ROW or MIXE, and ensure that the binlog_expire_logs_auto_purge option is set to ON to automatically purge expired binlog log files.

[0020] S43 Set the storage path and file size limit of the binlog log file. This can ensure that the binlog log file does not grow indefinitely, thus avoiding storage space problems. Configure the binlog_expire_logs_days option as needed to set the retention days of the log file.

[0021] S44 After completing the configuration changes, restart the target database service to make the changes take effect. The specific restart command depends on the database management system used.

[0022] S45 verifies whether the binlog function of the target database is working properly by performing some database operations, such as inserting, updating, or deleting records. Tools or commands can be used to view the content of the generated binlog logs to ensure that the correct data change operations are recorded.

[0023] S5 The data synchronization tool plays a key role in the data synchronization process. It is responsible for reading the binlog logs of the target database and extracting the data change situations from them.

[0024] S51 The data synchronization tool establishes a connection with the target database and listens for the generation of binlog logs in real time. Once the data in the database changes (such as insert, update, or delete operations), the corresponding binlog events will be recorded.

[0025] S52 The data synchronization tool reads the binlog events and parses the detailed information of the data changes. This includes which data records have changed, the specific content of the changes, and the timestamps of the changes, etc.

[0026] S53 According to the business requirements and data processing rules, the data synchronization tool can transform and filter the parsed data. For example, some unnecessary data changes may be excluded, or the data format may need to be adjusted to meet the requirements of the search engine database.

[0027] S54 The processed data change information is synchronized to the search engine database. This usually involves inserting, updating, or deleting the corresponding index entries for the changed data. Since the search engine database has pre-built an efficient index structure, these changes can be quickly processed and reflected in the search query results.

[0028] S55 In this way, the search engine database is always synchronized with the target database to ensure the real-time and accuracy of the data. When users perform search queries, they can obtain accurate results from the latest data.

[0029] S56 The designs of the data synchronization tool and the search engine database both consider performance optimization and scalability. As the data volume grows and the query load increases, they can maintain efficient operation by expanding hardware resources or adding distributed nodes.

[0030] S57 During the data synchronization process, both the data synchronization tool and the search engine database have error handling and monitoring mechanisms. Any abnormal or inconsistent situations will be detected and trigger corresponding handling measures, such as retrying, logging, or notifying the administrator for intervention.

[0031] In addition, the present disclosure also has the following advantages: 1. It improves the accuracy and timeliness of data synchronization. By monitoring the binlog of the database in real time, it can obtain data changes in a timely manner and accurately synchronize and update the data. Moreover, it can customize and summarize data such as user information, user subscription information, program information, authentication record information, and authorization record information according to actual business needs. This method avoids problems such as missed synchronization and duplicate synchronization that may occur in traditional data processing methods. It also innovatively adopts a distributed data synchronization architecture, which can achieve real-time data synchronization between multiple databases, improving the reliability and scalability of data synchronization. At the same time, this method has an adaptive adjustment function, which can automatically adjust the data synchronization strategy according to the frequency and size of data changes. In addition, it innovatively uses machine learning algorithms to detect and warn of abnormal situations during the data synchronization process, enabling timely discovery and handling of problems in the data synchronization process, and improving the reliability and stability of data synchronization.

[0032] 2. It improves the processing ability of the software system. The present invention improves the processing ability of the software system by adopting automated data synchronization and update technology. This technology can automatically identify and synchronize data changes without manual intervention, simplifying the cumbersome configuration and management work. In addition, it adopts a distributed processing architecture, which can achieve multi-task parallel processing, improving the concurrent processing ability and response speed of the software system. It has a fault tolerance and recovery mechanism, which can ensure the rapid restoration of data synchronization and update in case of system failure or data anomaly, improving the stability and reliability of the software system.

[0033] 3. It has good versatility and scalability. The present invention can adapt to various different database systems, including MySQL, Oracle, PostgreSQL, etc., without custom development for specific databases by adopting cross-platform and cross-technology-stack data synchronization and update technology. At the same time, it adopts an abstract data synchronization interface, which can seamlessly connect with various different data sources and target databases. It also has an adaptive adaptation ability, which can automatically adjust the data synchronization rules and algorithms according to different database systems and data formats, further improving the versatility and adaptability of the system. In addition, this technology innovatively adopts a modular design concept, splitting the data synchronization and update functions into multiple independent modules, which can be flexibly combined and extended according to actual needs, improving the customizability and maintainability of the system.

[0034] We applied this method to a certain operator. A large amount of user behavior data is generated every day, and real-time data processing and analysis are required. The operator adopted a big data volume real-time incremental capture and high-speed processing method that integrates heterogeneous database architectures. After using this disclosure, the accuracy and timeliness of data synchronization have been significantly improved, and at the same time, problems such as missed synchronization and duplicate synchronization that may occur in traditional data processing methods have been avoided. Through real-time analysis of the data, the operator timely discovers changes and trends in user behavior, and then optimizes products and services. Before using this method, the company needed to spend several hours every day processing data, and there were often problems such as data out-of-sync and missed synchronization. After using this method, the data synchronization time has been shortened to a few minutes, and the accuracy rate of data synchronization has reached 99.9%. At the same time, through real-time data analysis, the operator found that the usage rate of a certain new function by users was relatively low, and then timely adjusted the product strategy, ultimately improving users' satisfaction and activity with the new function. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, without creative efforts, other related drawings can also be obtained based on these drawings.

[0036] Figure 1 It is a flowchart of the method of the present invention. EMBODIMENTS

[0037] Step 1: Enable the binlog function configuration at each database source. This usually requires corresponding configuration in the database management system.

[0038] S11 Modify the database configuration file: In the database configuration file, find the configuration items related to binlog, such as log_bin, binlog_format, etc., and make corresponding configurations according to actual needs.

[0039] S12 Restart the database service: After modifying the configuration file, the database service needs to be restarted to make the configuration take effect.

[0040] S13 Verify the binlog function: Verify whether the binlog function has been correctly enabled by viewing the database log file or executing relevant commands.

[0041] Step 2: Listen to the binlog logs of each database through a data synchronization tool. This process can be achieved by writing corresponding programs or scripts, and by using the binlog reading interface provided by the database, the data change situation can be obtained in real time.

[0042] S21 Select a suitable data synchronization tool: Select a suitable data synchronization tool according to actual needs, such as Canal, Maxwell, etc.

[0043] S22 Configure the data synchronization tool: Make corresponding configurations according to the documentation of the data synchronization tool and actual needs. The configuration content includes data source information, the path to monitor binlog logs, data filtering rules, etc.

[0044] S23 Start the data synchronization tool: Start the data synchronization tool to start monitoring binlog logs. When the data in the database changes, the data synchronization tool will read the binlog logs and extract the specific data change situations.

[0045] Step 3: When the data changes, the data synchronization tool will read the binlog logs and extract the specific data change situations. According to business requirements and data processing rules, perform operations such as adding, updating, or deleting the corresponding data in the target database. This process can include operations such as data filtering and transformation to meet business requirements and data consistency requirements.

[0046] S31 Extract data change information: Extract the specific data change information from the binlog logs, including change types (adding, updating, deleting, etc.), the specific content of the changed data, etc.

[0047] S32 Process data changes: According to business requirements and data processing rules, perform corresponding processing on the extracted data change information, including operations such as data filtering and transformation to meet business requirements and data consistency requirements. The processed data change information will be sent to the target database.

[0048] Step 4: Enable the binlog function configuration in the target database. This usually also requires corresponding configurations. For example, in MySQL, the binlog function can be enabled by modifying files or using specified commands.

[0049] S41 Configure the data synchronization tool: Make corresponding configurations according to the documentation of the data synchronization tool and actual needs. The configuration content includes data source information, the path to monitor binlog logs, data filtering rules, etc. For example, we can configure Canal to monitor the binlog logs of the MySQL database, filter out data changes that are not primary key changes, and only focus on data changes of primary keys.

[0050] Step 5: The data synchronization tool obtains the specific data changes by reading the binlog of the target database, and then synchronizes these changes to the search engine database. This process can be achieved by writing corresponding programs or scripts, using the binlog reading interface provided by the database to obtain data change situations in real time and synchronize them to the search engine database.

[0051] S51 Configure the search engine database: Configure the search engine database according to actual needs, such as Elasticsearch, Solr, etc. The configuration content may include data source information, index settings, etc.

[0052] S52 Write a synchronization program or script: Write the corresponding program or script according to actual needs to implement the function of synchronizing data changes in the target database to the search engine database. The synchronization process includes data conversion, index creation or update, etc.

[0053] S53 Start the synchronization program or script: Start the written synchronization program or script to start synchronizing data changes in the target database to the search engine database. When there are new data changes, the synchronization program or script will automatically synchronize these changes to the search engine database to ensure data real-time and consistency.

Claims

1. A method for real-time incremental capture and high-speed processing of a large amount of data in a comprehensive heterogeneous database architecture, characterized in that, By leveraging the binlog function configuration of the database and combining it with advanced data synchronization tools to monitor the binlog logs of relevant tables such as user information, user subscription information, program information, authentication record information, and authorization record information in real time, the real-time monitoring and synchronization of data changes in relevant tables are achieved. This not only improves the accuracy and timeliness of data synchronization, avoiding problems such as missed synchronization and duplicate synchronization that may occur in traditional data processing methods. Moreover, it simplifies the traditional cumbersome configuration and management work, significantly enhancing the processing capacity of the software system. This method of real-time incremental capture and high-speed processing of large amounts of data with a comprehensive heterogeneous database architecture will bring a revolutionary change to application scenarios that require real-time data processing, such as the authentication and authorization system platform. It solves the problems in the authentication and authorization system platform, where due to the independent management of data by the three major operators, China Mobile, China Telecom, and China Unicom, the data volume is huge and the sources are diverse, and traditional query methods are time-consuming and the data cannot be synchronized in real time with low accuracy.

2. A method for real-time incremental capture and high-speed processing of large amounts of data in a comprehensive heterogeneous database architecture, characterized in that, By integrating at least one database, one data synchronization tool, and at least one target database, the real-time monitoring and synchronization of the binlog logs of relevant tables such as user information, user subscription information, program information, authentication record information, and authorization record information for data changes are achieved, and the data changes are synchronized to the target database and the search engine in real time. Problems such as missed synchronization and duplicate synchronization that may occur in traditional data processing methods are avoided, the traditional cumbersome configuration and management work are simplified, and the processing capacity of the software system is significantly enhanced.

3. A method for real-time incremental capture and high-speed processing of a large amount of data in a comprehensive heterogeneous database architecture, characterized in that, A device is designed, which integrates a processor (using a high-performance central processing unit (CPU)), a memory (using a high-speed random access memory (RAM) and a long-life read-only memory (ROM)), and an input / output interface (using a high-speed data transmission interface such as USB, HDMI, or Thunderbolt, etc.). Adopting real-time data processing methods such as stream processing, batch processing, or hybrid processing, it can process a large amount of data and complete operations such as data cleaning, transformation, analysis, and storage in a short time. It not only has powerful data processing capabilities but also can communicate efficiently with external devices. By the processor executing the real-time data processing method, the memory stores the programs and data of the processor, and the input / output interface is used to communicate with external devices, realizing the rapid transmission and processing of data, further enhancing the real-time performance and response speed of data processing.