Data synchronization conversion method and device based on message queue and big data processing framework, equipment and medium

By introducing a big data processing framework and message queue into the data synchronization device, the problem that data synchronization devices in the prior art cannot adapt to the inconsistent selection and table structure of different technologies is achieved, and flexible and accurate data synchronization and conversion are achieved.

CN120011431APending Publication Date: 2025-05-16SHANDONG LANGCHAO YUNTOU INFORMATION TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202411832198.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-12
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

Existing message queue-based data synchronization devices usually only integrate one message middleware, which cannot adapt to the differences in technical selection of different companies, and synchronization errors occur when the table structure field names between the data source and the target source are inconsistent.

Method used

The data synchronization conversion method based on the message queue and big data processing framework is adopted. By configuring the data source and target source addresses, matching the table structure information, selecting the middleware type and configuring the message queue information, sending the data to the message queue, the big data processing framework processes the data and converts it into the format expected by the user, and finally storing or exporting the data when needed.

Benefits of technology

It realizes the flexibility and adaptability of data synchronization, can adapt to the technical selection of different companies, automatically match table structure fields and convert data formats to ensure real-time synchronization and accuracy of data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011431A_ABST
    Figure CN120011431A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data synchronous processing and storage, and particularly relates to a data synchronous conversion method, device and equipment based on a message queue and a big data processing framework and a medium. The method comprises the steps that a data source address is configured, and data needing to be synchronized is selected; configuring a target source address of the received data, and matching table structure information of the data source and the target source; selecting a middleware type, and configuring message queue information; sending the data needing to be synchronized to a message queue; the big data processing framework obtains the data in the message queue, processes and converts the data into a data format expected by a user according to the demand of the user, and then sends the data to the message queue; when the data are stored, the target source reads the processed data in the message queue and stores the processed data; and when the data is exported, reading the processed data in the message queue, reading the data as stream data, and then generating a file format selected by a user for exporting. Real-time synchronization and conversion of the data are ensured, and meanwhile the accuracy of the data is kept.
Need to check novelty before this filing date? Find Prior Art

Description

Background Art

[0002] Message Queue (MQ) is a container that stores messages during the transmission process. MQ is responsible for transmitting messages between two systems. These two systems can be heterogeneous, on different hardware, different operating systems, and written in different languages. They can communicate with each other with simple configuration and simple calls to a few MQ APIs, without having to consider the complexity of the underlying system and network. MQ can cope with a variety of abnormal situations, such as network congestion, temporary interruptions, etc. It mainly solves problems such as application decoupling, asynchronous messaging, and traffic clipping, and realizes high-performance, high-availability, scalable, and eventually consistent architecture. Big data processing framework is a collection of frameworks that can analyze and process big data, and can process and calculate data in big data systems.

[0003] There are many existing data synchronization devices, some of which are based on message queues for data forwarding or integrate a special big data processing framework within the system to process batch data. Devices based on message queues for data synchronization basically only integrate one type of message middleware for data forwarding, but there are many types of message middleware, such as ActiveMQ, RabbitMQ, RocketMQ, Kafka, ZeroMQ, etc. Different companies will use different message queues to meet the needs of business scenarios. At this time, if the synchronization device using the message middleware does not meet the current company's technology selection, it is necessary to continue to integrate related technologies. Some synchronization devices require the table structure field names between the data source and the target source to be consistent before they can synchronize data, and some special fields or fields cannot be mapped one by one for some reason, resulting in synchronization errors. Summary of the invention

[0004] Devices that perform data synchronization based on message queues will basically only integrate one type of message middleware for data forwarding. Different companies will use different message queues to adapt to business scenario needs. At this time, if the synchronization device using the message middleware does not conform to the current company's technical selection, it will need to continue to integrate related technologies. The present invention provides a data synchronization conversion method, device, equipment and medium based on a message queue and a big data processing framework.

[0005] In a first aspect, the technical solution of the present invention provides a data synchronization conversion method based on a message queue and a big data processing framework, comprising the following steps: Configure the data source address and select the data to be synchronized; Configure the target source address for receiving data and match the table structure information of the data source and target source; Select the middleware type and configure the message queue information; Send the data that needs to be synchronized to the message queue; The big data processing framework obtains the data in the message queue, processes the data and converts it into the data format expected by the user according to the user's needs, and then sends the data to the message queue; When data needs to be stored, the target source reads the processed data in the message queue for storage; When data needs to be exported, the processed data in the message queue is read, the data is read as stream data, and then the file format selected by the user is generated for export.

[0006] As a further limitation of the technical solution of the present invention, before the step of configuring the data source address and selecting the data to be synchronized includes: Configure the database connection environment to access and connect to related databases, obtain detailed data in the library, and integrate the dependency configuration of multiple message middleware; create the connection ends of each database, the connection information of the message queue, integrate the development environment of the big data processing framework Spark, and integrate Spark's dependencies in the development environment.

[0007] As a further limitation of the technical solution of the present invention, before the steps of configuring the data source address and selecting the data to be synchronized, the following steps are further included: Configure the database connection information. After the connection is successful, obtain the detailed database table information and structure and display it in the system.

[0008] As a further limitation of the technical solution of the present invention, the steps of configuring the data source address and selecting the data to be synchronized include: Select the data you want to synchronize; Configure the information to be synchronized to the target database. When configuring the target source, test whether the connection is available. If successful, obtain the library and table information of the target database.

[0009] As a further limitation of the technical solution of the present invention, the steps of configuring the target source address of the received data and matching the table structure information of the data source and the target source include: Select the table to be synchronized, automatically match the table structure information of the data source and target source by name, and establish a field mapping relationship. If the target source does not have a table structure, generate a table in the system and establish a field mapping relationship.

[0010] As a further limitation of the technical solution of the present invention, the steps of selecting the middleware type and configuring the message queue information include: Select the type of message middleware, generate the available default message queue information, or customize the corresponding message middleware connection information as needed, and send the configuration information to the backend. The backend reads the parameters in the information and distributes them to the corresponding configuration, tests the connection, and stores the message queue information in the database after the connection is successful.

[0011] As a further limitation of the technical solution of the present invention, the steps of the big data processing framework obtaining data in the message queue, processing the data and converting it into a data format expected by the user according to the needs of the user, and then sending the data to the message queue include: The message middleware obtains data by reading the message queue of Spark; Create a Spark Streaming application. The big data processing framework Spark analyzes the data synchronization format set by the user, filters, analyzes, and reorganizes the data according to the user's needs, converts it into the data format expected by the user, and then sends the data to the message queue.

[0012] As a further limitation of the technical solution of the present invention, the method further includes: Record data from the time configuration is completed, sent to the message queue, and received for storage, and monitor the synchronization of recorded data in real time.

[0013] In a second aspect, the technical solution of the present invention also provides a data synchronization conversion device based on a message queue and a big data processing framework, including a synchronization information configuration module, a data sending module, a data processing module, a storage data acquisition module and a file export processing module; The synchronization information configuration module is used to configure the data source address and select the data to be synchronized; configure the target source address for receiving data and match the table structure information of the data source and target source; select the middleware type and configure the message queue information; The data sending module is used to send the data that needs to be synchronized to the message queue; Data processing module: The local big data processing framework obtains the data in the message queue, processes the data according to the user's needs and converts it into the data format expected by the user, and then sends the data to the message queue; The storage data acquisition module is used to read the processed data in the message queue for storage when data needs to be stored; The file export processing module is used to read the processed data in the message queue when data needs to be exported, read the data as stream data, and then generate the file format selected by the user for export.

[0014] As a further limitation of the technical solution of the present invention, the device also includes an environment configuration module, which configures the connection environment of the database to access and connect to the relevant database, obtain detailed data in the library, and integrate the dependency configuration of multiple message middleware; creates the connection end of each database, the connection information of the message queue, integrates the development environment of the big data processing framework Spark, and integrates Spark's dependencies in the development environment; configures the database connection information, and after the connection is successful, obtains the detailed library table information and structure in the database and displays it in the system.

[0015] As a further limitation of the technical solution of the present invention, the synchronization information configuration module is specifically used to select data to be synchronized; configure information synchronized to the target database, test whether the connection is available when configuring the target source, and if successful, obtain the library table information of the target database.

[0016] As a further limitation of the technical solution of the present invention, the steps of configuring the target source address of the received data and matching the table structure information of the data source and the target source include: The synchronization information configuration module is used to select the table to be synchronized, automatically match the table structure information of the data source and the target source by name, and establish a field mapping relationship. If the target source does not have a table structure, a table is generated in the system and a field mapping relationship is established.

[0017] As a further limitation of the technical solution of the present invention, the synchronization information configuration module is also used to select the type of message middleware, generate available default message queue information, or customize the corresponding message middleware connection information as needed, and send the configuration information to the backend. The backend reads the parameters in the information and distributes them to the corresponding configuration, tests the connection, and stores the message queue information in the database when the connection is successful.

[0018] As a further limitation of the technical solution of the present invention, the data processing module is used for the message middleware to obtain data by reading the message queue of Spark; to create a Spark Streaming application, the big data processing framework Spark analyzes the data synchronization format set by the user, filters, analyzes, and reorganizes the data according to the user's needs, and converts it into the data format expected by the user, and then sends the data to the message queue.

[0019] As a further limitation of the technical solution of the present invention, the device also includes a log recording module for recording detailed information about data after configuration is completed, sent to the message queue, and received for storage, and for real-time monitoring of data synchronization.

[0020] In a third aspect, the technical solution of the present invention provides an electronic device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the data synchronization conversion method based on the message queue and the big data processing framework as described in the first aspect.

[0021] In a fourth aspect, the technical solution of the present invention also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable the computer to execute the data synchronization conversion method based on the message queue and big data processing framework as described in the first aspect.

[0022] It can be seen from the above technical solutions that the present invention has the following advantages: (1) On the basis of realizing the data synchronization function between different data sources, the types of message queues are enriched, and the applicable message middleware type can be selected by the user, and then the relevant information of the message queue can be customized. Different queues can be set according to different tasks, and when the message queue needs to be replaced, the user can flexibly configure it, and the adaptability is better; it can ensure the real-time synchronization and conversion of data while maintaining the accuracy of the data. (2) It can monitor and record the process of data synchronization in detail. Compared with the ordinary synchronization execution log record, it records the data source, target source information, message queue information, and the processing and execution of real-time data on this basis; (3) When synchronizing data between databases, the table structure of the data source and target source can automatically match the field mapping, and support customized manual field mapping relationships. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0024] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.

[0026] like Figure 1 As shown, an embodiment of the present invention provides a data synchronization conversion method based on a message queue and a big data processing framework, comprising the following steps: Step 1: Configure the data source address and select the data to be synchronized; Step 2: Configure the target source address for receiving data and match the table structure information of the data source and target source; Step 3: Select the middleware type and configure the message queue information; Step 4: Send the data to be synchronized to the message queue; Step 5: The big data processing framework obtains the data in the message queue, processes the data and converts it into the data format expected by the user according to the user's needs, and then sends the data to the message queue; Step 6: When data needs to be stored, the target source reads the processed data in the message queue for storage; Step 7: When data needs to be exported, read the processed data in the message queue, read the data as stream data, and then generate the file format selected by the user for export.

[0027] First, the user needs to configure the data source address in the system, which includes the database address, user name, password, and the data table or collection that needs to be synchronized. The system reads the data content specified by the user by connecting to the data source.

[0028] Next, the user configures the target source address for receiving data, which also includes the address of the target database, user name, password, etc. The system will automatically check and match the table structure information of the data source and target source based on the user configuration to ensure that the data can be correctly mapped and synchronized.

[0029] Users need to select appropriate message queue middleware in the system, such as Kafka, RabbitMQ, etc., and configure the corresponding message queue information, including queue name, connection address, port, etc. The system will establish a connection with the message queue based on the user configuration.

[0030] The system reads the data that needs to be synchronized from the data source and sends it to the configured message queue. This step ensures that the data can be sent to the message queue in real time or at a scheduled time for subsequent processing.

[0031] Big data processing frameworks (such as Spark, Flink, etc.) obtain data from the message queue and process the data according to the data processing rules configured by the user (such as data cleaning, conversion, aggregation, etc.). The processed data is sent back to the message queue for subsequent storage or export.

[0032] When data needs to be stored, the target source reads the processed data from the message queue and stores it in the specified database or data warehouse. This step ensures the persistence and queryability of the data.

[0033] When data needs to be exported, the system reads the processed data from the message queue and converts it into stream data. Then, according to the file format selected by the user (such as CSV, Excel, JSON, etc.), the corresponding file is generated and exported.

[0034] By configuring the data source and target source addresses, and selecting the middleware type and message queue, the technical solution of the present invention can flexibly adapt to different data synchronization and conversion requirements. At the same time, the introduction of the big data processing framework enables the system to process large-scale data and has good scalability. Through the real-time data transmission of the message queue and the efficient processing capability of the big data processing framework, the technical solution of the present invention can ensure the real-time synchronization and conversion of data while maintaining the accuracy of the data. The system can automatically check and match the table structure information of the data source and the target source, reducing the possibility of manual intervention and errors. At the same time, the big data processing framework can intelligently process the data according to the data processing rules configured by the user. Users can realize the automated process of data synchronization and conversion through simple configuration and selection. The system provides an intuitive user interface and a friendly operating experience, which reduces the threshold for use. The system supports data export in multiple file formats to meet the data needs of users in different scenarios. At the same time, the processing method of stream data makes data export more efficient and flexible.

[0035] In some embodiments, the steps of configuring the data source address and selecting the data to be synchronized include: Configure the database connection environment to access and connect to related databases, obtain detailed data in the library, and integrate the dependency configuration of multiple message middleware; create the connection ends of each database, the connection information of the message queue, integrate the development environment of the big data processing framework Spark, and integrate Spark's dependencies in the development environment; configure the database connection information, and after the connection is successful, obtain the detailed library table information and structure in the database and display it in the system.

[0036] In some embodiments, the steps of configuring the data source address and selecting the data to be synchronized include: Select the data you want to synchronize; Configure the information to be synchronized to the target database. When configuring the target source, test whether the connection is available. If successful, obtain the library and table information of the target database.

[0037] In some embodiments, the steps of configuring a target source address for receiving data and matching table structure information of a data source and a target source include: Select the table to be synchronized, automatically match the table structure information of the data source and target source by name, and establish a field mapping relationship. If the target source does not have a table structure, generate a table in the system and establish a field mapping relationship.

[0038] In some embodiments, the steps of selecting a middleware type and configuring message queue information include: Select the type of message middleware, generate the available default message queue information, or customize the corresponding message middleware connection information as needed, and send the configuration information to the backend. The backend reads the parameters in the information and distributes them to the corresponding configuration, tests the connection, and stores the message queue information in the database after the connection is successful.

[0039] In some embodiments, the steps of the big data processing framework obtaining data in the message queue, processing the data and converting it into a data format expected by the user according to the user's needs, and then sending the data to the message queue include: The message middleware obtains data by reading the message queue of Spark; Create a Spark Streaming application. The big data processing framework Spark analyzes the data synchronization format set by the user, filters, analyzes, and reorganizes the data according to the user's needs, converts it into the data format expected by the user, and then sends the data to the message queue.

[0040] In some embodiments, the method further comprises: Record data from the time configuration is completed, sent to the message queue, and received for storage, and monitor the synchronization of recorded data in real time.

[0041] The embodiment of the present invention also provides a data synchronization conversion device based on a message queue and a big data processing framework, including a synchronization information configuration module, a data sending module, a data processing module, a storage data acquisition module and a file export processing module; The synchronization information configuration module is used to configure the data source address and select the data to be synchronized; configure the target source address for receiving data and match the table structure information of the data source and target source; select the middleware type and configure the message queue information; The data sending module is used to send the data that needs to be synchronized to the message queue; Data processing module: The local big data processing framework obtains the data in the message queue, processes the data according to the user's needs and converts it into the data format expected by the user, and then sends the data to the message queue; The storage data acquisition module is used to read the processed data in the message queue for storage when data needs to be stored; The file export processing module is used to read the processed data in the message queue when data needs to be exported, read the data as stream data, and then generate the file format selected by the user for export.

[0042] In some embodiments, the device also includes an environment configuration module, which configures the connection environment of the database to access and connect to the relevant database, obtain detailed data in the library, and integrate the dependency configuration of multiple message middleware; creates the connection end of each database, the connection information of the message queue, integrates the development environment of the big data processing framework Spark, and integrates Spark's dependencies in the development environment; configures the database connection information, and after the connection is successful, obtains the detailed library table information and structure in the database and displays it in the system.

[0043] In some embodiments, the synchronization information configuration module is specifically used to select data to be synchronized; configure information to be synchronized to a target database, test whether the connection is available when configuring a target source, and if successful, obtain the library table information of the target database.

[0044] The synchronization information configuration module is used to select the table to be synchronized, automatically match the table structure information of the data source and the target source by name, and establish a field mapping relationship. If the target source does not have a table structure, a table is generated in the system and a field mapping relationship is established.

[0045] In some embodiments, the synchronization information configuration module is also used to select the type of message middleware, generate available default message queue information, or customize the corresponding message middleware connection information as needed, and send the configuration information to the backend. The backend reads the parameters in the information and distributes them to the corresponding configuration, tests the connection, and stores the message queue information in the database if the connection is successful.

[0046] In some embodiments, the data processing module is used for the message middleware to obtain data by reading the message queue of Spark; to create a Spark Streaming application, the big data processing framework Spark analyzes the data synchronization format set by the user, filters, analyzes, and reorganizes the data according to the user's needs, and converts it into the data format expected by the user, and then sends the data to the message queue.

[0047] In some embodiments, the device also includes a logging module for recording detailed information about data from the time the configuration is completed, the data is sent to the message queue, and the data is received and stored, and for real-time monitoring and recording of data synchronization.

[0048] In order to realize the functions of data flow, conversion and synchronization between the data source and the target source. The device is based on message queues and big data processing frameworks to perform data processing and synchronization functions. Message queues are used to transfer data between the data source and the target source, providing decoupling and asynchronous processing capabilities. The big data processing framework is used to convert and process the data that needs to be synchronized. The big data processing framework of this device uses Spark. It is mainly based on these two technologies to handle data synchronization between different databases.

[0049] This device provides more custom configurations, such as the configuration of the data source. If it is a database, configure the relevant connection information, configure the field mapping relationship between the data source and the target source, select the data to be synchronized, select the appropriate message middleware according to the business scenario, etc., configure the message queue, configure the specified data sending queue, and set the queue attributes according to the business scenario. The extracted data can also be processed and exported as JSON, CSV, Excel and other formats. Detailed log information will be recorded during data synchronization, which is convenient for users to find relevant records and troubleshoot problems. After the front end defines the data source, target library information, field mapping, and data range value, a series of information is transmitted to the back end. The back end receives the information, configures the data source and target library according to the information, and the specific logic of data synchronization, reads the data in the data source and sends it to the message queue. The big data processing device reads the data in the message queue, performs relevant processing, writes the data to the specified queue after processing, and finally reads the data in the queue after processing and stores it in the specified data target. Through this device, data synchronization between multiple databases can be realized, as well as the formatted export function of data can be realized. It can also select appropriate message middleware according to its current technology and customize the message queue for synchronized data.

[0050] The device provided by the present invention can be used as a tool to support different types of databases, such as relational databases: MySQL, Oracle, PostgreSQL, SQL Server, etc., and non-relational databases: MongoDB, Redis, Elasticsearch, HBase, etc.; configure database connection information, and obtain structural information of library table information; support custom creation of message queues, configuration of queue addresses, topics, producers, and consumer information; support flexible arrangement of data synchronization tasks in real-time scheduling, timed scheduling, event-driven scheduling, etc.; perform processing operations such as filtering, deduplication, combination, and aggregation on the collected data according to the data synchronization format set by the user, perform batch processing on historical data memory synchronization conversion, stream processing on real-time data analysis and forwarding, or timely combine and perform mixed processing; support conversion of different data formats, such as JSON, XML, CSV, etc., to facilitate adaptation to different types of data and user data export needs, and can remove invalid data and clean data during data conversion. Or add, delete and modify data according to customization to ensure data quality, support complex data mapping and conversion rules, intelligent mapping or custom mapping, and maximize the realization of mapping relationships; monitor the data synchronization process in real time to ensure the integrity, accuracy and consistency of data synchronization. When synchronization failures, data anomalies and other problems occur, it can trigger prompts and alarms, record the causes of related problems, and facilitate users to troubleshoot related problems. It also logs various aspects of synchronization to ensure subsequent tracking of problems and troubleshooting, and supports basic message fault tolerance and retry mechanisms.

[0051] The specific application process of the device provided by the embodiment of the present invention is as follows: (1) Configure the basic project environment, integrate the front-end and back-end separation development, related language development environment, configure the connection environment of MySQL, Oracle, PostgreSQL, SQL Server, MongoDB, Redis, Elasticsearch, HBase and other databases to access and connect to related databases and obtain detailed data in the library, integrate the dependency configuration of multiple message middleware, and facilitate users to select the message middleware they use, create the connection end of each database, the connection information of the message queue, the message topic, queue name, producer, consumer and other information, and the configuration information is obtained as a configurable item from the parameters received from the front end. The back end obtains the parameter information and distributes it to the corresponding configuration information. Integrate the development environment of the big data processing framework Spark, download related files, and integrate Spark dependencies in the development environment; (2) When using the tool, users should first test the connection to the database and configure the database connection information. After the connection is successful, the detailed database table information and structure will be obtained and displayed in the system. Different databases can be connected and the connection information will be retained, which is convenient for direct use after the subsequent test connection is successful without further configuration. After obtaining the relevant database information, you can click on the database table to view the detailed information. To operate the database table, you can query the data through SQL statements or some filtering conditions provided by the system. The front end will send the operations on the database back to the back end, and the back end will return the results to the system front end for display after processing; (3) After the user has selected the data to be synchronized, the information to be synchronized to the target database is configured. When configuring the target source, the connection is also tested to see if it is available. If successful, the database table information of the target database is obtained. Then the table to be synchronized is selected. The table structure information of the data source and the target source is automatically matched by name. In special cases such as special fields or large name differences, the user can manually establish the field mapping relationship after the automatic field matching is completed, or can directly perform manual field mapping without using the automatic field matching. If the target source does not have a table structure, the user can generate a table in the system, and the mapping relationship is completed at this point. (4) Then select the type of message middleware according to the situation. After selecting the middleware type, the available default message queue information is generated. You can also customize the corresponding message middleware connection information as needed, including the message middleware: queue address, authentication information, message subject, queue name, producer, consumer, etc. After the front-end configures the information, it will send the configuration information to the back-end. The back-end reads the parameters in the information and distributes them to the corresponding configuration. Test the connection. If the connection is successful, you can execute the next step. Because the data needs to be sent to the big data processing framework through the message queue first, this is the first message queue configuration. After the subsequent data processing is completed, it needs to be sent to the target source through the message middleware. The message queue information of this process must also be configured together. After the two message queues are configured, the message queue information will be stored in the database so that it can be used directly after testing without configuration next time. (5) Then select the synchronization time, support scheduled tasks, instant scheduling and other operations. After the front-end configures all operations, it will send the configuration information to the back-end, and the back-end will save the configuration information to the database. The user does not need to configure the configured information the next time, but only needs to test the connection successfully before using it. (6) After synchronization is turned on, the initial data will be sent to the message queue. Some message middleware can obtain data by reading the message queue through Spark. Unsupported message queues can read data through third-party connectors or custom code. Create a Spark Streaming application and configure various ways for Spark to process data. Later, the backend will select the corresponding data processing logic based on the synchronization method of the front-end user. The big data processing framework Spark is mainly responsible for analyzing the data synchronization format set by the user, filtering, analyzing, and reorganizing the data according to user needs, writing the corresponding task code and sending it to the framework, converting it into the data format expected by the user, and then sending the data to the message queue. The target source reads the processed data in the message queue and stores it to ensure the integrity, accuracy, and consistency of data synchronization; (7) If the user wants to export data instead of storing it in the database, he can choose to export it as a file when configuring the target source. The user selects the corresponding export file type, such as JSON, XML, or CSV. The execution is consistent with the above steps, except that the target database does not read the processed data in the message queue at the end. Instead, the system reads the processed data in the message queue, reads the data as stream data, and then generates the file format selected by the user for export; (8) Logging records detailed information about data from the time the configuration is completed, when it is sent to the message queue, and when it is received and stored, and monitors the data synchronization status in real time.

[0052] The embodiment of the present invention also provides an electronic device, which includes: a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus. The communication bus can be used for information transmission between the electronic device and the sensor. The processor can call the logic instructions in the memory to execute the following method: Step 1: configure the data source address and select the data to be synchronized; Step 2: configure the target source address of the received data and match the table structure information of the data source and the target source; Step 3: select the middleware type and configure the message queue information; Step 4: send the data to be synchronized to the message queue; Step 5: the big data processing framework obtains the data in the message queue, processes the data according to the user's needs and converts it into the data format expected by the user, and then sends the data to the message queue; Step 6: when the data needs to be stored, the target source reads the processed data in the message queue for storage; Step 7: when the data needs to be exported, read the processed data in the message queue, read the data as stream data, and then generate the file format selected by the user for export.

[0053] In addition, the logic instructions in the above-mentioned memory can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc., which can store program code.

[0054] An embodiment of the present invention provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions enable a computer to execute the method provided by the above method embodiment, for example, including: Step 1: configure the data source address and select the data to be synchronized; Step 2: configure the target source address for receiving data, and match the table structure information of the data source and the target source; Step 3: select the middleware type and configure the message queue information; Step 4: send the data to be synchronized to the message queue; Step 5: the big data processing framework obtains the data in the message queue, processes the data according to the user's needs and converts it into the data format expected by the user, and then sends the data to the message queue; Step 6: when data needs to be stored, the target source reads the processed data in the message queue for storage; Step 7: when data needs to be exported, read the processed data in the message queue, read the data as stream data, and then generate a file format selected by the user for export.

[0055] It should be understood that the order of execution of the steps in the above embodiment does not necessarily mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of the present invention.

[0056] An embodiment of the data synchronization conversion device based on the message queue and the big data processing framework is provided in an embodiment of the present invention. The device and the data synchronization conversion method based on the message queue and the big data processing framework in the above-mentioned embodiments belong to the same inventive concept. For details not described in detail in the embodiment of the data synchronization conversion device based on the message queue and the big data processing framework, reference can be made to the embodiment of the data synchronization conversion method based on the message queue and the big data processing framework.

[0057] The data synchronization conversion device based on the message queue and the big data processing framework is a unit and algorithm step of each example described in combination with the embodiments disclosed herein, which can be implemented by electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to the function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0058] Those skilled in the art will appreciate that various aspects of the data synchronization conversion method based on the message queue and the big data processing framework can be implemented as a system, method or program product. Therefore, various aspects of the present disclosure can be specifically implemented in the following forms, namely: a complete hardware implementation, a complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which can be collectively referred to as "circuit", "module" or "system" here.

[0059] Although the present invention has been described in detail with reference to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person of ordinary skill in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person of ordinary skill in the art may easily think of changes or substitutions within the technical scope disclosed by the present invention, and these shall be within the scope of protection of the present invention.

Claims

1. A data synchronization conversion method based on a message queue and a big data processing framework, characterized in that: The steps include: Configure the data source address and select the data to be synchronized; Configure the target source address for receiving data and match the table structure information of the data source and target source; Select the middleware type and configure the message queue information; Send the data that needs to be synchronized to the message queue; The big data processing framework obtains the data in the message queue, processes the data and converts it into the data format expected by the user according to the user's needs, and then sends the data to the message queue; When data needs to be stored, the target source reads the processed data in the message queue for storage; When data needs to be exported, the processed data in the message queue is read, the data is read as stream data, and then the file format selected by the user is generated for export.

2. The data synchronization conversion method based on message queue and big data processing framework according to claim 1 is characterized in that: The steps to configure the data source address and select the data to be synchronized include: Configure the database connection environment to access and connect to related databases, obtain detailed data in the library, and integrate the dependency configuration of multiple message middlewares; create the connection terminals of each database, the connection information of the message queue, integrate the development environment of the big data processing framework Spark, and integrate the dependencies of Spark in the development environment; Configure the database connection information. After the connection is successful, obtain the detailed database table information and structure and display it in the system.

3. The data synchronization conversion method based on message queue and big data processing framework according to claim 2 is characterized in that: The steps to configure the data source address and select the data to be synchronized include: Select the data you want to synchronize; Configure the information to be synchronized to the target database. When configuring the target source, test whether the connection is available. If successful, obtain the library and table information of the target database.

4. The data synchronization conversion method based on message queue and big data processing framework according to claim 3 is characterized in that: The steps of configuring the target source address for receiving data and matching the table structure information of the data source and the target source include: Select the table to be synchronized, automatically match the table structure information of the data source and target source by name, and establish a field mapping relationship. If the target source does not have a table structure, generate a table in the system and establish a field mapping relationship.

5. The data synchronization conversion method based on message queue and big data processing framework according to claim 4 is characterized in that: Select the middleware type and configure the message queue information as follows: Select the type of message middleware, generate the available default message queue information, or customize the corresponding message middleware connection information as needed, and send the configuration information to the backend. The backend reads the parameters in the information and distributes them to the corresponding configuration, tests the connection, and stores the message queue information in the database after the connection is successful.

6. The data synchronization conversion method based on message queue and big data processing framework according to claim 5 is characterized in that: The big data processing framework obtains data from the message queue, processes the data according to the user's needs and converts it into the data format expected by the user, and then sends the data to the message queue. The steps include: The message middleware obtains data by reading the message queue of Spark; Create a Spark Streaming application. The big data processing framework Spark analyzes the data synchronization format set by the user, filters, analyzes, and reorganizes the data according to the user's needs, converts it into the data format expected by the user, and then sends the data to the message queue.

7. The data synchronization conversion method based on message queue and big data processing framework according to claim 6 is characterized in that: The method further includes: Record data from the time configuration is completed, sent to the message queue, and received for storage, and monitor the synchronization of recorded data in real time.

8. A data synchronization conversion device based on a message queue and a big data processing framework, characterized in that: It includes a synchronization information configuration module, a data sending module, a data processing module, a storage data acquisition module and a file export processing module; The synchronization information configuration module is used to configure the data source address and select the data to be synchronized; configure the target source address for receiving data and match the table structure information of the data source and target source; select the middleware type and configure the message queue information; The data sending module is used to send the data that needs to be synchronized to the message queue; Data processing module: The local big data processing framework obtains the data in the message queue, processes the data according to the user's needs and converts it into the data format expected by the user, and then sends the data to the message queue; The storage data acquisition module is used to read the processed data in the message queue for storage when data needs to be stored; The file export processing module is used to read the processed data in the message queue when data needs to be exported, read the data as stream data, and then generate the file format selected by the user for export.

9. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; the memory stores computer program instructions executable by the at least one processor, and the computer program instructions are executed by the at least one processor so that the at least one processor can execute the data synchronization conversion method based on the message queue and big data processing framework as described in any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium stores computer instructions, which enable the computer to execute the data synchronization conversion method based on the message queue and big data processing framework as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Real-time data connection system, method, equipment and medium

    CN121255910A

  • Enterprise business data dynamic processing method and system based on open source data synchronization framework

    CN121255930A