Data-based report type data synchronization method and device

By initializing the paging control parameters in the Reader plug-in of datax, looping to obtain and parse the data of nested structures, generating intermediate structured data objects, and cache them to the target system according to batch policies, the data heterogeneity and scalability problems are solved, and synchronization accuracy and system stability are improved.

CN120336433AActive Publication Date: 2025-07-18BENXI IRON & STEEL (GROUP) INFORMATION AUTOMATION CO LTD

Patent Information

Application Number
CN202510820449.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-19
Publication Date
2025-07-18
Estimated Expiration
2045-06-19

AI Technical Summary

Technical Problem

The existing data synchronization technology has shortcomings in processing data heterogeneity and system scalability, resulting in low synchronization accuracy and high maintenance costs, especially in the synchronization synchronization of multi-source heterogeneous data synchronization, which is difficult to meet production requirements.

Method used

By initializing the paging control parameters in the Reader plug-in of datax, looping to obtain the paging data response and parsing the nested structure, generating intermediate structured data objects, pre-processing of data, and using the Writer plug-in to cache and submit to the target data system according to batch policies.

Benefits of technology

It improves the accuracy of multi-source heterogeneous data synchronization and the scalability of the system, reduces maintenance costs, and improves the stability and throughput of data synchronization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120336433A_ABST
    Figure CN120336433A_ABST
Patent Text Reader

Abstract

The invention provides a datax-based report type data synchronization method and device, and the method comprises the steps: initializing a paging control parameter based on a data structure of to-be-obtained data in a Reader plug-in, circularly transmitting a data obtaining request through an HTTP request mode, and circularly obtaining a paging data response of a paging size from a source end until a termination condition is reached, thereby achieving the synchronization of the to-be-obtained data. Analyzing the data fields of the nested structure in the paging data response to generate intermediate structured data objects corresponding to the data fields, and further performing data preprocessing on the intermediate structured data objects corresponding to the data fields of the nested structure and the data fields of the non-nested structure to form preprocessed data objects; transmitting the preprocessed data object to the Writer plug-in based on the Reader plug-in, caching the preprocessed data object according to a batch strategy based on the Writer plug-in, and submitting the preprocessed data object to the target data system, so that the synchronization accuracy and expandability of the multi-source heterogeneous data are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method and device for synchronizing report data based on DataX. Background Art

[0002] With the continuous development of big data technology, the integration demand of enterprises for heterogeneous data sources is increasing day by day. In traditional manufacturing industries such as iron and steel metallurgy, the acquisition, synchronization, and processing of massive data have become an important foundation for supporting digital operations. However, the current mainstream data synchronization technologies still face various challenges in practical applications, especially in the processing of data heterogeneity and system scalability.

[0003] First of all, the problem of data heterogeneity in the data synchronization scenario is becoming increasingly serious. Taking Ansteel Group as an example, the big data lake platform it constructs provides data services to external systems through API interfaces. However, the data structures returned by these interfaces are complex, usually organized in a multi-layer nested JSON format, with deep field paths, irregular structures, and a mixture of multi-level array nesting and Map structures in the returned JSON objects. Traditional ETL tools often have problems such as field misalignment and data omission when expanding into two-dimensional structured tables, resulting in the synchronization accuracy not meeting the production requirements.

[0004] Secondly, the scalability of existing systems is insufficient. In practical applications, business systems usually rely on customized Java programs to achieve the docking and synchronization of various data sources. This point-to-point development mode lacks a unified synchronization framework and intermediate data standard in the context of the rapid growth of data source types, resulting in the need to repeat development and testing for each newly added data source. Especially in the process of cross-system docking, the differences in field semantics, data precision, time format, etc. between different data sources are obvious. Manual docking not only takes time and effort but also has extremely high maintenance costs, bringing a heavy burden to the operation and maintenance team. Summary of the Invention

[0005] The present invention provides a method and device for synchronizing report data based on DataX to solve the deficiencies of insufficient synchronization accuracy and poor system scalability in the prior art.

[0006] The present invention provides a method for synchronizing report data based on DataX, including: In the Reader plugin of DataX, initialize the paging control parameters based on the data structure of the data to be obtained; the paging control parameters include the initial page number and the page size, and construct the URL address of the data acquisition request for the data to be obtained; the data to be obtained includes unstructured report data; Based on the Reader plugin, the data acquisition request is cyclically sent via HTTP requests, and the paged data responses of the paging size are cyclically obtained from the source end until a preset termination condition is reached. Then, the data fields with nested structures in the paged data responses are parsed to generate intermediate structured data objects corresponding to each data field; the source end includes various different types of data sources. Based on the Reader plugin, data preprocessing is performed on the intermediate structured data objects corresponding to the data fields with nested structures and the data fields with non-nested structures to form preprocessed data objects. Based on the Reader plugin, the preprocessed data objects are passed to the Writer plugin of datax, and the preprocessed data objects are cached according to the batch strategy and submitted to the target data system based on the Writer plugin.

[0007] According to a method for synchronizing report data based on datax provided by the present invention, the parsing of the data fields with nested structures in the paged data responses to generate intermediate structured data objects corresponding to each data field includes: Constructing a complete field path tree of the data fields with nested structures; For any data field with a nested structure, based on the field names, field types of each node in the complete field path tree of the any data field, and the order between each node, determining the intermediate data structure corresponding to the any data field; Based on the intermediate data structure corresponding to the any data field, generating an intermediate structured data object corresponding to the any data field.

[0008] According to a method for synchronizing report data based on datax provided by the present invention, the determining of the intermediate data structure corresponding to the any data field based on the field names, field types of each node in the complete field path tree of the any data field, and the order between each node includes: Based on the field names and field types of each node in the complete field path tree of the any data field, extracting the field semantic vector representations of the corresponding nodes; Based on the field semantic vector representations of each node in the complete field path tree of the any data field, the order between each node, and the field semantic vector representations of each field and the order between each field included in each standard nested data structure, performing matching to determine the standard nested data structure with the highest matching degree as the intermediate data structure corresponding to the any data field.

[0009] A report - type data synchronization method based on DataX according to the present invention. The method pre - processes the intermediate structured data objects corresponding to the data fields of the nested structure and the data fields of the non - nested structure based on the Reader plug - in to form pre - processed data objects, including: Determine the intermediate data type corresponding to the data field of the non - nested structure based on the field type of the data field of the non - nested structure; Convert the corresponding data field based on the intermediate data type corresponding to the data field of the non - nested structure to obtain an intermediate - type data object corresponding to the data field of the non - nested structure; Based on the type of the target data system, perform type conversion on the intermediate structured data object corresponding to the data field of the nested structure and the intermediate - type data object corresponding to the data field of the non - nested structure to obtain pre - processed data objects corresponding to each data field.

[0010] A report - type data synchronization method based on DataX according to the present invention. After performing type conversion on the intermediate structured data object corresponding to the data field of the nested structure and the intermediate - type data object corresponding to the data field of the non - nested structure to obtain pre - processed data objects corresponding to each data field, it further includes: Perform anomaly detection on the pre - processed data objects corresponding to each data field to obtain anomaly detection results for each data field; wherein, the anomaly detection result of any data field includes whether the any data field has an anomaly and the anomaly type when the any data field has an anomaly; Based on the anomaly handling strategies for each anomaly type in the preset strategy library, perform anomaly repair on the pre - processed data objects corresponding to the data fields with anomalies, so as to pass the pre - processed data objects after anomaly repair corresponding to the data fields with anomalies to the Writer plug - in of DataX.

[0011] A report - type data synchronization method based on DataX according to the present invention. The step of circularly obtaining the paged data response of the paging size from the source end until a preset termination condition is reached includes: Update the paging size based on the field type of the data fields in the paged data response returned by the current request, the data volume of the paged data response, and the request response time, and send the next request based on the updated paging size; Dynamically determine whether the preset termination condition is reached based on the identification field in the paged data response returned by the current request.

[0012] A report - type data synchronization method based on DataX according to the present invention. The step of caching the pre - processed data objects according to the batch strategy and submitting them to the target data system includes: Cache the pre - processed data objects corresponding to each data field in the memory buffer queue; Based on the set batch quantity or time window, batch - write the pre - processed data objects in the memory buffer queue to the target data system.

[0013] The present invention also provides a report - type data synchronization device based on DataX, including: An initialization unit, configured to initialize paging control parameters based on the data structure of the data to be acquired in the Reader plug - in of DataX; the paging control parameters include an initial page number and a page size, and construct a URL address for the data acquisition request for the data to be acquired; the data to be acquired includes unstructured report - type data; A data acquisition and parsing unit, configured to circularly send the data acquisition request in an HTTP request manner based on the Reader plug - in, circularly acquire paged data responses of the page size from the source end until a preset termination condition is reached, and parse the data fields with nested structures in the paged data responses to generate intermediate structured data objects corresponding to each data field; the source end includes various different types of data sources; A data pre - processing unit, configured to perform data pre - processing on the intermediate structured data objects corresponding to the data fields with nested structures and the data fields with non - nested structures based on the Reader plug - in to form pre - processed data objects; A data batch - writing unit, configured to transfer the pre - processed data objects to the Writer plug - in of DataX based on the Reader plug - in, and cache the pre - processed data objects according to a batch strategy and submit them to the target data system based on the Writer plug - in.

[0014] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the report - type data synchronization method based on DataX as described in any one of the above.

[0015] The present invention also provides a non - transitory computer - readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the report - type data synchronization method based on DataX as described in any one of the above.

[0016] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the report - type data synchronization method based on DataX as described in any one of the above.

[0017] A method and device for synchronizing report - type data based on DataX provided by the present invention initialize paging control parameters based on the data structure of the data to be obtained in the Reader plug - in of DataX. Then, based on the Reader plug - in, data acquisition requests are cyclically sent via HTTP requests to cyclically obtain paging data responses of the paging size from the source end until a preset termination condition is reached. The data fields with nested structures in the paging data responses are parsed to generate intermediate structured data objects corresponding to each data field. Furthermore, based on the Reader plug - in, data pre - processing is performed on the intermediate structured data objects corresponding to the data fields with nested structures and the data fields with non - nested structures to form pre - processed data objects. Then, based on the Reader plug - in, the pre - processed data objects are passed to the Writer plug - in of DataX, and based on the Writer plug - in, the pre - processed data objects are cached according to the batch strategy and submitted to the target data system, improving the synchronization accuracy and scalability for multi - source heterogeneous data. Brief Description of the Drawings

[0018] To more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 is a schematic flowchart of the method for synchronizing report - type data based on DataX provided by the present invention; Figure 2 is a schematic flowchart of the method for parsing data fields with nested structures provided by the present invention; Figure 3 is a schematic structural diagram of the device for synchronizing report - type data based on DataX provided by the present invention; Figure 4 is a schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments

[0020] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0021] Figure 1 is a schematic flowchart of the method for synchronizing report - type data based on DataX provided by the present invention, as Figure 1As shown, the method includes: Step 110: In the Reader plugin of DataX, initialize the paging control parameters based on the data structure of the data to be retrieved; the paging control parameters include the initial page number and the page size, and construct the URL address of the data retrieval request for the data to be retrieved; the data to be retrieved includes unstructured report data. Step 120: Based on the Reader plugin, send the data retrieval request in a loop via the HTTP request method, loop to retrieve the paging data response of the page size from the source end until a preset termination condition is reached, and parse the data fields with nested structures in the paging data response to generate intermediate structured data objects corresponding to each data field; the source end includes multiple different types of data sources. Step 130: Based on the Reader plugin, perform data preprocessing on the intermediate structured data objects corresponding to the data fields with nested structures and the data fields with non-nested structures to form preprocessed data objects. Step 140: Based on the Reader plugin, transfer the preprocessed data objects to the Writer plugin of DataX, and based on the Writer plugin, cache the preprocessed data objects according to the batch strategy and submit them to the target data system.

[0022] Here, by reconstructing the Reader plugin in the DataX framework, rewriting the core logic startRead method representing data retrieval and parsing in the Reader plugin, adding functions such as custom dynamic paging control, nested structure parsing, and type conversion, the parsed data is stored in the RecordSender object defined in the DataX framework, and the Writer plugin of DataX is used to implement data writing. That is, the reconstructed Reader plugin has the above functions such as dynamic paging control, nested structure parsing, and type conversion.

[0023] Specifically, when the reconstructed Reader plugin is initialized, the system will configure appropriate paging control parameters according to the data interface structure to be synchronized, that is, the data structure of the data to be obtained. These parameters can include the initial page number, the paging size, and the paging termination condition identifier, etc. The initial page number is used to define from which page to start obtaining data, and the paging size controls the amount of data requested each time, thus laying a foundation for subsequent batch writing operations. The paging termination condition identifier can be flexibly set according to different interface characteristics. For example, the paging termination condition identifier can be set to "has_next_page" or "next_page", so as to determine whether to continue to request the next page by judging whether "has_next_page" is false in the response or whether the "next_page" field is empty. Based on this, the Reader plugin constructs a complete URL address for the request of the data to be obtained, encapsulates the HTTP request information, and prepares for the subsequent loop paging acquisition process.

[0024] Next, the Reader plugin starts to send HTTP requests to the source interface in a loop, obtaining data page by page until the identifier field in the paged data response returned by the current request indicates that the preset termination condition has been reached. Among them, the source includes various types of data sources. Each request will return a paged data response in JSON format, and this paged data response may contain a large number of data fields with nested structures. In view of this feature, the embodiment of the present invention embeds a configurable path resolver and structure mapping module in the Reader plugin, so that the system can accurately extract data regions from complex nested JSON data. For example, in some typical interfaces, the main data may be located in the "list" array under the name "data", and each record nests multiple sub-structures such as price information, timestamp, and region fields. Therefore, by configuring the path resolver and structure mapping module of the Reader plugin, the system can accurately locate these nested data fields, parse them, convert the unstructured data into structured data, and thus generate corresponding intermediate structured data objects.

[0025] In some embodiments, in order to improve the data acquisition efficiency, after each HTTP request is sent, the paging size, which is a paging control parameter, can be updated based on the field type of the data fields in the paged data response returned by the current request, the amount of data in the paged data response, and the request response time, so as to send the next request based on the updated paging size.

[0026] Here, the parameter of the paging size is not only set in the initial stage, but also can be dynamically adjusted during the execution of the data acquisition loop according to the actual network environment and the characteristics of the returned data, so as to improve the stability and throughput efficiency of data synchronization. After the request is sent, the key characteristics in the paging data response returned can be analyzed. This analysis includes but is not limited to: the amount of data in the paging data response returned this time (such as the actual number of data records, the number of fields in the JSON data structure), the field types of the data fields (including the nesting levels of the nested data fields in the nested structure, the complexity of the field types, such as whether they contain arrays, objects, long texts, file encoding fields, etc.), and the time consumed from the request being sent to receiving the response (i.e., the request response time).

[0027] Specifically, when the system detects that the field structure in the current paging data response is relatively complex and the fields contain highly structured types such as nested objects and arrays, the system will determine that the batch of data requires a large amount of computing resources during the parsing and standardization process, and can reduce the paging size of the next request to avoid system processing bottlenecks. On the contrary, if the current response data structure is relatively simple, such as most fields are strings or numeric types and the number of data records has not reached the paging upper limit, the system can appropriately increase the paging size of the next page, thereby reducing the number of requests and improving the overall synchronization efficiency. In addition, the request response time is also an important input factor for dynamically adjusting the paging size. For example, if the system finds that the request response time for processing a certain page of data requests is significantly higher than the average level (such as higher than 3 seconds), it is inferred that there may be situations such as insufficient network bandwidth or high processing pressure on the interface server side. At this time, the paging size will also be reduced to reduce the data volume of the next request, thereby increasing the success rate of the request and reducing the system load. When the system shows a fast response speed for several consecutive requests (such as less than 1 second each time), the paging size can be appropriately increased to further improve the data pulling efficiency. Through the above self-adaptive control method of the paging size, the performance bottlenecks and the number of failed retries when facing data heterogeneity and network uncertainty can be reduced, and the stability, robustness and throughput of the entire data synchronization process are improved.

[0028] In some embodiments, as Figure 2 shown, the following method can be adopted to implement the parsing of nested structure data fields: Step 210, construct a complete field path tree for the data fields of the nested structure; Step 220, for any data field of the nested structure, based on the field names and field types of each node in the complete field path tree of the any data field and the order between each node, determine the intermediate data structure corresponding to the any data field; Step 230: Generate an intermediate structured data object corresponding to any of the data fields based on the intermediate data structure corresponding to the any of the data fields.

[0029] Among them, in order to accurately convert a complex JSON structure into an intermediate structured data object, the construction of a field path tree can be used as the basis to comprehensively identify the semantic and structural relationships between nested fields, and based on this, an intermediate data structure can be constructed to facilitate subsequent data standardization and writing operations. Since the paged data response returned by the source end often has a multi-layer nested structure, such as multi-level combination forms of nested arrays, nested objects, structures containing structures, etc. In order to effectively parse the data fields of these nested structures, a recursive scanning operation can be performed on the JSON structure in each paged data response to identify the data fields of the nested structure and construct a complete field path tree for each nested structure. In this path tree, each node represents a field in the data field of the nested structure, and the path of the node is determined by the hierarchical order of the fields. The path expression adopts the typical pattern of "parent field.child field.grandchild field". For example, a field "data.details.price" indicates that it is the "price" field in the "details" object under the "data" object. The field path tree not only records the name of each field, but also synchronously records its data type (such as string, array, object, number, etc.) and its position order in the hierarchical structure.

[0030] After the complete field path tree is constructed, according to the field name and field type of each node corresponding field in the complete field path tree, and the order of each node in the complete field path tree, select a standard nested data structure that matches the data field of the nested structure from the pre-constructed standard nested data structure library (such as Array <struct>or Map<String, Object>, etc.) as its corresponding intermediate data structure. In some other embodiments, based on the field names and field types of each node in the complete field path tree of the data fields of the nested structure, a pre-trained language model can be used to extract the field semantic vector representations of the corresponding nodes. Subsequently, based on the field semantic vector representations of each node in the complete field path tree of the data fields and the order between each node, as well as the field semantic vector representations of each field included in each standard nested data structure (which can be obtained based on a pre-trained language model) and the order between each field, semantic matching is performed to determine the standard nested data structure with the highest matching degree as the intermediate data structure corresponding to the data field.

[0031] After determining the intermediate data structure corresponding to the data field of any nested structure, an intermediate structured data object corresponding to the data field can be generated based on the definition of this intermediate data structure. This data object not only retains the hierarchical and type information of the original field but can also be further efficiently converted into a data structure supported by the target data system in subsequent processing flows, thereby helping to adapt to the differentiated structures of multi-source data interfaces.

[0032] After the parsing of the nested structure is completed, the Reader plugin further preprocesses the extracted data fields to generate a preprocessing data object that meets the format requirements of the target data system and has a clear structure. During this process, the data types of different data fields can be forcibly converted according to the preset field type mapping rules. It should be noted that before this, based on the field type of the data field of the non-nested structure, the intermediate data type corresponding to the data field of the non-nested structure can be determined, and this intermediate data type is a data type that can be conveniently converted into various target data system formats. Then, based on the intermediate data type corresponding to the data field of the non-nested structure, the corresponding data field is converted to obtain the intermediate type data object corresponding to the data field of the non-nested structure. Next, based on the type of the target data system (such as MySQL, Elasticsearch, HDFS, etc.), the intermediate structured data objects corresponding to the data fields of each nested structure and the intermediate type data objects corresponding to the data fields of each non-nested structure are type-converted to obtain the preprocessing data objects corresponding to each data field that meet the format requirements of the target data system.

[0033] In addition, anomaly detection can be performed on the preprocessed data objects corresponding to each data field to obtain the anomaly detection results for each data field. Among them, the anomaly detection result of any data field includes whether there is an anomaly in this data field and the type of anomaly when this data field has an anomaly. Subsequently, based on the anomaly handling policies for each anomaly type in the preset policy library, the preprocessed data objects corresponding to the data fields with anomalies are repaired for anomalies, further improving the availability and standardization of the overall data. For example, for missing or abnormally formatted field values, the default value policy can be automatically applied for filling, and for illegal field values, the discard policy can be applied, etc.

[0034] After the above preprocessing is completed, the Reader plug-in transfers the sorted preprocessed data objects (or the preprocessed data objects after anomaly repair) to the data pipeline of datax, and the Writer plug-in of datax receives them and is responsible for writing the preprocessed data objects corresponding to each data field into the target data system. In the embodiments of the present invention, the Writer plug-in can adopt a batch submission optimization strategy, which greatly improves the writing efficiency compared with the traditional one-by-one writing method. Specifically, the Writer plug-in maintains a memory buffer queue and temporarily stores the preprocessed data objects corresponding to each data field in the memory buffer queue after receiving them. Only when the cached data volume reaches a preset threshold (for example, every 100 pieces of data) or reaches a preset time window, a unified submission operation is triggered to batch-write the preprocessed data objects in the memory buffer queue into the target data system. This mechanism effectively reduces the impact of frequent IO operations on system performance and significantly improves the throughput. In addition, during the submission process, the Writer plug-in can select the optimal data writing method according to the type of the target data system, such as partition appending for Hive tables, batch insertion for MySQL, or high-speed writing interfaces for DWS, etc. Since the Reader plug-in has standardized the field types and formats, the Writer plug-in can efficiently complete the data writing action, improving the stability and compatibility of data synchronization.

[0035] In addition, in order to enhance the robustness and fault tolerance of the system, the embodiments of the present invention also design an adaptive anomaly handling and retry mechanism. During the process of the Reader plug-in initiating a paging request or performing field parsing, if abnormal situations such as network interruption, data format mismatch, and field missing occur, the system will automatically start the retry policy. This policy not only includes the limit of the maximum number of retries but also can dynamically adjust the retry interval to improve the system recovery ability. When some data fields still cannot be processed after multiple attempts, they can be marked as abnormal records and transferred to a dedicated log or intermediate storage area, and subsequent manual review or supplementary recording can be performed on them to avoid affecting the execution of the overall task.

[0036] In addition, the synchronization method of the present invention has good scalability. In the case of adding new data sources or interface changes, users only need to modify parameters such as URL templates, field path mappings, and data type definitions in the Reader plugin through configuration, without modifying the underlying logic or rewriting code, thus significantly reducing the maintenance cost. This method also supports a plug-in integrated monitoring mechanism, which can record metrics such as the success rate of each request response, the number of fields extracted, and the submission time-consuming, providing rich observability information for operation and maintenance personnel and helping to quickly locate potential problems in the data synchronization process.

[0037] In summary, the method provided by the embodiment of the present invention initializes paging control parameters based on the data structure of the data to be obtained in the Reader plugin of datax, and then, based on the Reader plugin, circularly sends data acquisition requests through HTTP requests to circularly obtain paging data responses of the paging size from the source end until a preset termination condition is reached, and parses the data fields with nested structures in the paging data responses to generate intermediate structured data objects corresponding to each data field. Furthermore, based on the Reader plugin, data preprocessing is performed on the intermediate structured data objects corresponding to the data fields with nested structures and the data fields with non-nested structures to form preprocessed data objects, and then the preprocessed data objects are passed to the Writer plugin of datax based on the Reader plugin, and the preprocessed data objects are cached according to the batch strategy and submitted to the target data system based on the Writer plugin, improving the synchronization accuracy and scalability for multi-source heterogeneous data.

[0038] The following describes the report class data synchronization device based on datax provided by the present invention. The report class data synchronization device based on datax described below can be correspondingly referred to the report class data synchronization method based on datax described above.

[0039] Based on any of the above embodiments, Figure 3 is a schematic structural diagram of the report class data synchronization device based on datax provided by the present invention. As Figure 3 shown, the device includes: An initialization unit 310, configured to initialize paging control parameters based on the data structure of the data to be obtained in the Reader plugin of datax; the paging control parameters include an initial page number and a paging size, and construct a URL address for the data acquisition request for the data to be obtained; the data to be obtained includes unstructured report class data; A data acquisition and parsing unit 320, configured to, based on a Reader plug-in, circularly send the data acquisition request in an HTTP request manner, circularly acquire paged data responses of the paging size from a source end until a preset termination condition is reached, and parse data fields with nested structures in the paged data responses to generate intermediate structured data objects corresponding to the respective data fields; the source end includes various different types of data sources; A data preprocessing unit 330, configured to perform data preprocessing on intermediate structured data objects corresponding to data fields with nested structures and data fields with non-nested structures based on a Reader plug-in to form preprocessed data objects; A data batch writing unit 340, configured to transfer the preprocessed data objects to a Writer plug-in of datax based on a Reader plug-in, and cache and submit the preprocessed data objects to a target data system based on the Writer plug-in according to a batch strategy.

[0040] The device provided by the embodiment of the present invention initializes paging control parameters based on the data structure of data to be acquired in a Reader plug-in of datax, thereby circularly sending a data acquisition request in an HTTP request manner based on the Reader plug-in, circularly acquiring paged data responses of the paging size from a source end until a preset termination condition is reached, parsing data fields with nested structures in the paged data responses to generate intermediate structured data objects corresponding to the respective data fields, then performing data preprocessing on intermediate structured data objects corresponding to data fields with nested structures and data fields with non-nested structures based on the Reader plug-in to form preprocessed data objects, and further transferring the preprocessed data objects to a Writer plug-in of datax based on the Reader plug-in, and caching and submitting the preprocessed data objects to a target data system based on the Writer plug-in according to a batch strategy, improving the synchronization accuracy and scalability for multi-source heterogeneous data.

[0041] Based on any of the above embodiments, the parsing of data fields with nested structures in the paged data responses to generate intermediate structured data objects corresponding to the respective data fields includes: Constructing a complete field path tree of the data fields with nested structures; For any data field with a nested structure, determining an intermediate data structure corresponding to the any data field based on the field names and field types of each node in the complete field path tree of the any data field and the order between the respective nodes; Generating an intermediate structured data object corresponding to the any data field based on the intermediate data structure corresponding to the any data field.

[0042] Based on any of the above embodiments, determining the intermediate data structure corresponding to any of the data fields based on the field names, field types of each node in the complete field path tree of any of the data fields, and the order between each node, includes: Based on the field names and field types of each node in the complete field path tree of any of the data fields, extracting the field semantic vector representation of the corresponding node; Based on the field semantic vector representation of each node and the order between each node in the complete field path tree of any of the data fields, and the field semantic vector representation of each field and the order between each field included in each standard nested data structure for matching, determining the standard nested data structure with the highest matching degree as the intermediate data structure corresponding to any of the data fields.

[0043] Based on any of the above embodiments, the data preprocessing of the intermediate structured data object corresponding to the nested structure data field and the non-nested structure data field by the Reader plugin to form a preprocessed data object, includes: Based on the field type of the non-nested structure data field, determining the intermediate data type corresponding to the non-nested structure data field; Based on the intermediate data type corresponding to the non-nested structure data field, converting the corresponding data field to obtain the intermediate type data object corresponding to the non-nested structure data field; Based on the type of the target data system, performing type conversion on the intermediate structured data object corresponding to the nested structure data field and the intermediate type data object corresponding to the non-nested structure data field to obtain the preprocessed data object corresponding to each data field.

[0044] Based on any of the above embodiments, after performing type conversion on the intermediate structured data object corresponding to the nested structure data field and the intermediate type data object corresponding to the non-nested structure data field to obtain the preprocessed data object corresponding to each data field, it further includes: Performing anomaly detection on the preprocessed data object corresponding to each data field to obtain the anomaly detection result of each data field; wherein, the anomaly detection result of any data field includes whether there is an anomaly in any data field, and the anomaly type when there is an anomaly in any data field; Based on the anomaly handling policies for each anomaly type in the preset policy library, performing anomaly repair on the preprocessed data object corresponding to the data field with an anomaly, so as to pass the preprocessed data object after anomaly repair corresponding to the data field with an anomaly to the Writer plugin of datax.

[0045] Based on any of the above embodiments, circularly obtaining paged data responses of the paging size from the source end until a preset termination condition is reached includes: Updating the paging size based on the field type of the data fields in the paged data response returned by the current request, the data volume of the paged data response, and the request response time, and sending a next request based on the updated paging size; Dynamically determining whether the preset termination condition is reached based on the identification field in the paged data response returned by the current request.

[0046] Based on any of the above embodiments, caching the preprocessed data objects according to a batch policy and submitting them to the target data system includes: Caching the preprocessed data objects corresponding to each data field in a memory buffer queue; Based on a set batch quantity or time window, batch-writing the preprocessed data objects in the memory buffer queue into the target data system.

[0047] Figure 4 FIG. is a schematic structural diagram of an electronic device provided by the present invention. As Figure 4 shown, the electronic device may include: a processor 410, a memory 420, a communication interface 430, and a communication bus 440. Among them, the processor 410, the memory 420, and the communication interface 430 communicate with each other through the communication bus 440. The processor 410 may call logical instructions in the memory 420 to execute a method for synchronizing report class data based on datax. The method includes: in the Reader plug-in of datax, initializing paging control parameters based on the data structure of the data to be obtained; the paging control parameters include an initial page number and a paging size, and constructing a URL address for a data acquisition request for the data to be obtained; the data to be obtained includes unstructured report class data; based on the Reader plug-in, circularly sending the data acquisition request through an HTTP request method, circularly obtaining paged data responses of the paging size from the source end until a preset termination condition is reached, and parsing the data fields with nested structures in the paged data response to generate intermediate structured data objects corresponding to each data field; the source end includes various different types of data sources; based on the Reader plug-in, performing data preprocessing on the intermediate structured data objects corresponding to the data fields with nested structures and the data fields without nested structures to form preprocessed data objects; based on the Reader plug-in, passing the preprocessed data objects to the Writer plug-in of datax, and caching the preprocessed data objects according to a batch policy based on the Writer plug-in and submitting them to the target data system.

[0048] In addition, when the logical instructions in the above-mentioned memory 420 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0049] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-transitory computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer can execute the data synchronization method for report data based on DataX provided by the above-mentioned various methods. The method includes: in the Reader plug-in of DataX, initializing paging control parameters based on the data structure of the data to be acquired; the paging control parameters include an initial page number and a paging size, and constructing a URL address for the data acquisition request for the data to be acquired; the data to be acquired includes unstructured report data; based on the Reader plug-in, circularly sending the data acquisition request in an HTTP request manner, circularly acquiring the paging data response of the paging size from the source end until a preset termination condition is reached, and parsing the data fields with nested structures in the paging data response to generate intermediate structured data objects corresponding to each data field; the source end includes various different types of data sources; based on the Reader plug-in, performing data preprocessing on the intermediate structured data objects corresponding to the data fields with nested structures and the data fields with non-nested structures to form preprocessed data objects; based on the Reader plug-in, transferring the preprocessed data objects to the Writer plug-in of DataX, and caching the preprocessed data objects according to a batch strategy based on the Writer plug-in and submitting them to the target data system.

[0050] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the data synchronization method for report data based on DataX provided above. The method includes: in the Reader plug-in of DataX, initializing paging control parameters based on the data structure of the data to be obtained; the paging control parameters include an initial page number and a paging size, and constructing a URL address for the data acquisition request for the data to be obtained; the data to be obtained includes unstructured report data; based on the Reader plug-in, circularly sending the data acquisition request in an HTTP request manner, circularly obtaining paging data responses of the paging size from the source end until a preset termination condition is reached, and parsing data fields with nested structures in the paging data responses to generate intermediate structured data objects corresponding to each data field; the source end includes multiple different types of data sources; based on the Reader plug-in, performing data preprocessing on the intermediate structured data objects corresponding to the data fields with nested structures and the data fields with non-nested structures to form preprocessed data objects; based on the Reader plug-in, passing the preprocessed data objects to the Writer plug-in of DataX, and caching the preprocessed data objects according to a batch strategy based on the Writer plug-in and submitting them to the target data system.

[0051] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0052] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disc, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0053] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.< / struct>

Claims

1. A method for synchronizing report data based on DataX, characterized in that, Including: In the Reader plugin of DataX, initialize paging control parameters based on the data structure of the data to be obtained; The paging control parameters include an initial page number and a paging size, and construct a URL address for the data acquisition request for the data to be obtained; the data to be obtained includes unstructured report data; Based on the Reader plugin, circularly send the data acquisition request in the form of an HTTP request, circularly obtain paged data responses of the paging size from the source end until a preset termination condition is reached, and parse the data fields in the nested structure in the paged data response to generate intermediate structured data objects corresponding to each data field; the source end includes various different types of data sources; Based on the Reader plugin, perform data preprocessing on the intermediate structured data objects corresponding to the data fields in the nested structure and the data fields in the non-nested structure to form preprocessed data objects; Based on the Reader plugin, pass the preprocessed data objects to the Writer plugin of DataX, and based on the Writer plugin, cache the preprocessed data objects according to a batch strategy and submit them to the target data system.

2. The method for synchronizing report data based on DataX according to claim 1, wherein The parsing of the data fields in the nested structure in the paged data response to generate intermediate structured data objects corresponding to each data field includes: Construct a complete field path tree for the data fields in the nested structure; For any data field in the nested structure, based on the field names, field types, and the order between each node in the complete field path tree of the any data field, determine the intermediate data structure corresponding to the any data field; Based on the intermediate data structure corresponding to the any data field, generate an intermediate structured data object corresponding to the any data field.

3. The method for synchronizing report data based on DataX according to claim 2, wherein The determining of the intermediate data structure corresponding to the any data field based on the field names, field types, and the order between each node in the complete field path tree of the any data field includes: Based on the field names and field types of each node in the complete field path tree of the any data field, extract the field semantic vector representations of the corresponding nodes; Based on the field semantic vector representations of each node and the order between each node in the complete field path tree of the any data field, and the field semantic vector representations of each field and the order between each field included in each standard nested data structure, perform matching to determine the standard nested data structure with the highest matching degree as the intermediate data structure corresponding to the any data field.

4. The method for synchronizing report data based on DataX according to claim 1, wherein The performing of data preprocessing on the intermediate structured data objects corresponding to the data fields in the nested structure and the data fields in the non-nested structure based on the Reader plugin to form preprocessed data objects includes: Based on the field type of the data field in the non-nested structure, determine the intermediate data type corresponding to the data field in the non-nested structure; Based on the intermediate data type corresponding to the data field in the non-nested structure, perform conversion on the corresponding data field to obtain an intermediate type data object corresponding to the data field in the non-nested structure; Based on the type of the target data system, perform type conversion on the intermediate structured data objects corresponding to the data fields of the nested structure and the intermediate type data objects corresponding to the data fields of the non-nested structure to obtain the preprocessed data objects corresponding to each data field.

5. The method for synchronizing report data based on DataX according to claim 4, wherein After performing the type conversion on the intermediate structured data objects corresponding to the data fields of the nested structure and the intermediate type data objects corresponding to the data fields of the non-nested structure to obtain the preprocessed data objects corresponding to each data field, the following steps are further included: Perform anomaly detection on the preprocessed data objects corresponding to each data field to obtain the anomaly detection results for each data field; wherein, the anomaly detection result of any data field includes whether the any data field has an anomaly, and the type of anomaly when the any data field has an anomaly; Based on the anomaly handling policies for each type of anomaly in the preset policy library, perform anomaly repair on the preprocessed data objects corresponding to the data fields with anomalies, so as to pass the preprocessed data objects after anomaly repair corresponding to the data fields with anomalies to the Writer plugin of datax.

6. The method for synchronizing report data based on DataX according to claim 1, wherein The step of cyclically obtaining the paged data responses of the paging size from the source end until a preset termination condition is reached includes: Update the paging size based on the field type of the data fields in the paged data response returned by the current request, the data volume of the paged data response, and the request response time, so as to send the next request based on the updated paging size; Dynamically determine whether the preset termination condition is reached based on the identification field in the paged data response returned by the current request.

7. The method for synchronizing report data based on DataX according to claim 1, characterized in that, The step of caching the preprocessed data objects according to the batch policy and submitting them to the target data system includes: Cache the preprocessed data objects corresponding to each data field in the memory buffer queue; Based on the set batch quantity or time window, batch-write the preprocessed data objects in the memory buffer queue into the target data system.

8. A report data synchronization device based on DataX, characterized in that, It includes: An initialization unit, configured to initialize the paging control parameters in the Reader plugin of datax based on the data structure of the data to be obtained; The paging control parameters include the initial page number and the paging size, and construct the URL address of the data acquisition request for the data to be obtained; the data to be obtained includes unstructured report data; A data acquisition and parsing unit, configured to cyclically send the data acquisition request in an HTTP request manner based on the Reader plugin, cyclically obtain the paged data responses of the paging size from the source end until a preset termination condition is reached, and parse the data fields of the nested structure in the paged data response to generate the intermediate structured data objects corresponding to each data field; the source end includes various different types of data sources; A data preprocessing unit, configured to perform data preprocessing on the intermediate structured data objects corresponding to the data fields of the nested structure and the data fields of the non-nested structure based on the Reader plugin to form preprocessed data objects; The data batch writing unit is used to transfer the preprocessed data object to the Writer plug-in of DataX based on the Reader plug-in, and cache the preprocessed data object according to the batch strategy based on the Writer plug-in and submit it to the target data system.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein, When the processor executes the program, it implements the DataX-based report data synchronization method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the DataX-based report data synchronization method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Data processing method and device

    CN108011687A

  • Adaptive recruitment decision-making system and method based on recruitment behavior data

    CN115063119A

  • Data analysis method and device, electronic equipment and readable medium

    CN115114890A

  • Data storage and processing method based on heterogeneous technology

    CN117251414A

  • Data capturing method and device and related equipment

    CN117370464A

Cited By

  • Data synchronization method and system for intelligent dirty data detection and restoration based on DataX

    CN121301477A

  • API (Application Program Interface) paging request method and device based on DataX

    CN121614196A