Data quality detection method and device for DataX and medium
By reading the data source in parallel shards and loading the detection rule library in DataX, generating exception logs and visual reports, the integration of DataX data quality detection is solved, and efficient and accurate data quality detection is achieved.
Patent Information
- Application Number
- CN202510335294.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-20
- Publication Date
- 2025-07-04
AI Technical Summary
DataX in the prior art does not have complete data quality detection functions. When used in combination with data synchronization tools such as DataX, there are problems such as high integration difficulty and low detection efficiency.
Through DataX's reader plug-in, data sources are read in parallel, and the quality detection engine is used to load the predefined detection rule library. The detection field is empty, the format is deviated from the preset template or the required fields are missing. Exception logs are generated, and they are classified into the distributed file system by timestamp, a visual summary report is generated, and a repair script library is called for repair.
It realizes efficient and accurate data quality detection during the data input and output process of DataX, improves detection efficiency, and ensures the reliability and effectiveness of data processing.
Smart Images

Figure CN120256262A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and in particular, to a method, device, and medium for data quality detection for DataX. Background Art
[0002] DataX is a data synchronization tool open-sourced by Alibaba. It provides a rich plug-in system and can achieve data transmission between various data sources. Data quality detection is to evaluate and verify aspects such as the accuracy, integrity, consistency, and timeliness of data to ensure that the data can meet business requirements. In the big data era, the amount of data has increased sharply, the data sources are diverse, and data quality problems have become more prominent. The demand for efficient data quality detection methods and systems is also becoming increasingly urgent.
[0003] With the rapid development of information technology, enterprises and organizations have accumulated a large amount of data. However, there are often quality problems in this data, such as data missing, errors, duplicates, etc. Low-quality data will seriously affect the accuracy of data analysis and the reliability of decision-making. Although DataX, as a widely used data synchronization tool, can efficiently achieve data transmission between different data sources, the native DataX does not have a perfect data quality detection function. Currently, data quality detection tools on the market often need to be deployed and configured separately. When combined with data synchronization tools such as DataX, there are problems such as high integration difficulty and low detection efficiency.
[0004] Through the above analysis, the problems and defects existing in the prior art are as follows:
[0005] DataX in the prior art does not have a perfect data quality detection function. When combined with data synchronization tools such as DataX, there are problems such as high integration difficulty and low detection efficiency. Summary of the Invention
[0006] Embodiments of this application provide a method, device, and medium for data quality detection for DataX, which can solve the problems that DataX in the prior art does not have a perfect data quality detection function and there are problems such as high integration difficulty and low detection efficiency when combined with data synchronization tools such as DataX.
[0007] In a first aspect, an embodiment of the present application provides a method for data quality detection for DataX. The method includes: parallelly sharding and reading a data source through a reader plugin of DataX to obtain data to be processed; when the data to be processed is transmitted to a memory buffer, triggering a quality detection engine, and loading a predefined detection rule library through the quality detection engine. The detection rule library includes null judgment rules, compliance templates, and integrity constraints; if it is detected that a field is empty, the format deviates from a preset template, or a required field is missing, an exception log including the problem type, field identifier, and original data is generated; the exception log is classified and stored in a distributed file system according to timestamps, a visual summary report is generated, and the problem type, field identifier, and original data are parsed, and a predefined repair script library is called for repair.
[0008] In an implementation manner of the present application, parallelly sharding and reading a data source through a reader plugin of DataX to obtain data to be processed specifically includes: parsing the distribution characteristics of connection parameters based on the connection parameters of the data source; calculating an optimal sharding strategy according to the distribution characteristics, allocating independent threads for each shard, and concurrently reading through the reader plugin to obtain shard data; caching the shard data in a memory buffer in preset batches; during the reading process, monitoring the throughput rate of the shards in real time, and if it is detected that the shard delay exceeds a threshold, triggering a resharding operation.
[0009] In an implementation manner of the present application, during the reading process, monitoring the throughput rate of the shards in real time, and if it is detected that the shard delay exceeds a threshold, triggering a resharding operation specifically includes: monitoring the CPU occupancy rate and network delay of the shards, and constructing a shard health index; if the health index is lower than a preset threshold, splitting the current shard into multiple sub-shards; allocating new threads for the sub-shards, migrating them to uncompleted data reading tasks, and updating the shard metadata to a ZooKeeper cluster.
[0010] In an implementation manner of the present application, classifying and storing the exception log in a distributed file system according to timestamps and generating a visual summary report specifically includes: extracting the exception log in the distributed file system and performing aggregation statistics according to the problem type; calculating a data quality score based on the aggregation result and generating an interactive HTML report; pushing the interactive HTML report to a message queue for a downstream system to trigger an automatic repair process.
[0011] In an implementation manner of the present application, parsing the problem type, field identifier, and original data, and calling a predefined repair script library for repair specifically includes: selecting an optimal repair strategy according to the script matching degree, and if the score is lower than the threshold, transferring it to manual processing; re-injecting the repaired data into the DataX data transmission stream and triggering a secondary quality detection; if the secondary quality detection passes, updating the data quality score and closing the exception work order.
[0012] In one implementation manner of the present application, calculating an optimal sharding strategy according to distribution characteristics specifically includes: collecting metadata of a data source, where the metadata includes table size, index distribution, and storage engine type; constructing a sharding weight matrix based on the metadata; combining network bandwidth and the number of CPU cores, and using a greedy algorithm to select sharding boundaries, and during the sharding execution process, correcting the weight matrix in real time.
[0013] In one implementation manner of the present application, after calling a predefined repair script library for repair, the method further includes: extracting abnormal logs with a problem type higher than a preset frequency, and extracting rule optimization feature vectors; inputting the feature vectors into a rule model to generate a candidate set of new rules; performing conflict detection on the candidate set, and if the candidate set passes the verification of the conflict detection, automatically merging the candidate set into the detection rule library; triggering incremental data quality recheck based on the candidate set of new rules, and updating the visualization summary report.
[0014] In one implementation manner of the present application, inputting the feature vectors into a rule model to generate a candidate set of new rules specifically includes: using a Transformer model to learn the relevance of the feature vectors to maximize the rule coverage rate to obtain a rule model; deploying the rule model as a microservice to enable online incremental learning and real-time rule push.
[0015] In a second aspect, an embodiment of the present application further provides a device for data quality detection for DataX. The device includes at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: read a data source in parallel shards through a reader plugin of DataX to obtain data to be processed; when the data to be processed is transmitted to a memory buffer, trigger a quality detection engine, and load a predefined detection rule library through the quality detection engine, where the detection rule library includes null judgment rules, compliance templates, and integrity constraints; if it is detected that a field is empty, the format deviates from a preset template, or a required field is missing, generate an abnormal log including the problem type, field identifier, and original data; classify and store the abnormal log in a distributed file system according to timestamps, generate a visualization summary report, and parse the problem type, field identifier, and original data, and call a predefined repair script library for repair.
[0016] In a third aspect, an embodiment of the present application further provides a non-volatile computer storage medium for data quality detection of DataX, storing computer-executable instructions, and the computer-executable instructions are configured to: read data sources in parallel and in slices through the reader plug-in of DataX to obtain data to be processed; when the data to be processed is transmitted to the memory buffer, trigger the quality detection engine, and load a predefined detection rule library through the quality detection engine, where the detection rule library includes null judgment rules, compliance templates, and integrity constraints; if it is detected that a field is empty, the format deviates from the preset template, or a required field is missing, generate an exception log including the problem type, field identifier, and original data; classify and store the exception log in the distributed file system according to the timestamp, generate a visual summary report, parse the problem type, field identifier, and original data, and call a predefined repair script library for repair.
[0017] A method, device, and medium for data quality detection of DataX provided by an embodiment of the present application utilize the writer and reader mechanisms of DataX to achieve fast input and output of data streams. By embedding self-developed data quality detection scripts, data quality detection is performed during the data input and output process of DataX, supporting common detection rules such as null judgment, data compliance, and data integrity, and outputting the detection results to a text file, which can efficiently and accurately perform data quality detection and improve the reliability and effectiveness of data processing. Combining data quality detection closely with the data transmission process of DataX, without additional data transmission and processing steps, greatly improves the detection efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0019] Figure 1 is a flowchart of a method for data quality detection of DataX provided by an embodiment of the present application;
[0020] Figure 2 is a schematic internal structure diagram of a device for data quality detection of DataX provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments of this application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts belong to the scope of protection of this application.
[0022] The embodiments of this application provide a method, device, and medium for data quality detection for DataX, which solve the problems in the prior art that DataX does not have a perfect data quality detection function, and when combined with data synchronization tools such as DataX, there are problems of high integration difficulty and low detection efficiency.
[0023] The technical solutions proposed in the embodiments of this application will be described in detail below with reference to the drawings.
[0024] Figure 1 It is a flowchart of a method for data quality detection for DataX provided by the embodiments of this application. As Figure 1 shown, a method for data quality detection for DataX provided by the embodiments of this application specifically includes the following steps:
[0025] Step 10: Read the data source in parallel slices through the reader plugin of DataX to obtain the data to be processed.
[0026] In this step, the writer plugin of DataX is responsible for writing data to the target data source, and the reader plugin is responsible for reading data from the source data source. By expanding and customizing the writer and reader plugins of DataX, it can adapt to the data formats and transmission requirements of different data sources, and achieve fast and stable transmission of the data stream. When reading data from a relational database, the optimized algorithm of the reader plugin can efficiently read a large amount of data and quickly transmit it to the subsequent data quality detection link.
[0027] As an optional embodiment, reading the data source in parallel slices through the reader plugin of DataX to obtain the data to be processed may specifically include: Step 101: Analyze the distribution characteristics of the connection parameters based on the connection parameters of the data source; Step 102: Calculate the optimal slicing strategy according to the distribution characteristics, allocate independent threads to each slice, and concurrently read the sliced data through the reader plugin.
[0028] In this step, obtain the connection parameters of the data source: database URL, username, password, table name, query conditions, analyze the connection parameters, and determine the distribution characteristics of the data source.
[0029] As an alternative embodiment, calculating an optimal sharding strategy according to distribution characteristics may specifically include:
[0030] Step 1021: Collect metadata of the data source. The metadata includes table size, index distribution, and storage engine type; Step 1022: Construct a sharding weight matrix based on the metadata; Step 1023: Combine the network bandwidth and the number of CPU cores, and use the greedy algorithm to select sharding boundaries. During the sharding execution process, the weight matrix is corrected in real time.
[0031] In this step, the structure and dimension of the sharding weight matrix can be defined according to the type and importance of the metadata. For example, a two-dimensional matrix can be defined, where the rows represent different tables or data blocks, and the columns represent different sharding factors, including table size, index complexity, and storage engine performance. According to the metadata, calculate the weight values of each table or data block on each sharding factor. The weight values can be calculated through normalization, standardization, or a custom weight function. Fill the calculated weight values into the weight matrix to form a complete sharding weight matrix. The greedy algorithm is an algorithm that makes the best or optimal choice at each step in the current state. In this embodiment, the greedy algorithm can gradually adjust the sharding boundaries according to the weight matrix and resource limitations, so that the data volume of each sharding is as equal as possible and meets the resource limitations. During the sharding execution process, monitor the changes in data distribution and resource usage in real time. According to the monitoring results, correct the weight values in the weight matrix in real time to reflect the dynamic changes in data distribution.
[0032] Step 103: Cache the sharded data into the memory buffer in preset batches; Step 104: During the reading process, monitor the throughput rate of the shards in real time. If it is detected that the shard delay exceeds the threshold, trigger a resharding operation.
[0033] In this step, caching into the memory buffer in preset batches is to ensure the decoupling of data transmission and quality detection.
[0034] As an alternative embodiment, during the reading process, monitor the throughput rate of the shards in real time. If it is detected that the shard delay exceeds the threshold, trigger a resharding operation, which may specifically include: Step 1041: Monitor the CPU occupancy rate and network delay of the shards, and construct a shard health index; Step 1042: If the health index is lower than the preset threshold, split the current shard into multiple sub-shards; Step 1043: Allocate new threads for the sub-shards, and migrate them to the unfinished data reading tasks, and update the shard metadata to the ZooKeeper cluster.
[0035] In this step, a health metric can be defined, which comprehensively considers the CPU occupancy rate and network latency. For example, the health metric = α×(1 - CPU occupancy rate) + β×(1 - network latency / maximum tolerable latency). According to the system performance and business requirements, set the threshold of the health metric and update the metadata of the sub-shards to the ZooKeeper cluster or other distributed coordination services.
[0036] Step 20: When the data to be processed is transferred to the memory buffer, trigger the quality detection engine, and load the predefined detection rule library through the quality detection engine. The detection rule library includes null-checking rules, compliance templates, and integrity constraints.
[0037] In this step, during the data input and output process of DataX, embed the self-developed data quality detection script to monitor the data flow in real time. When the data is read from the reader plugin and before it is transferred to the writer plugin, the data quality detection script is automatically started to analyze each data record and determine whether there are quality problems according to the preset detection rules. For the data read from the CSV file, the detection script will check in real time whether the data format conforms to the CSV specification during the data transfer process.
[0038] Step 30: If it is detected that a field is empty, the format deviates from the preset template, or a required field is missing, generate an exception log containing the problem type, field identifier, and original data.
[0039] In this step, the data quality detection supports a variety of common detection rules, including null-checking, data compliance, data integrity, etc. The null-checking detection rule is used to check whether a data field is a null value, and if it is null, it is marked as a quality problem; the data compliance detection rule checks whether the data meets specific format, range, etc. requirements according to business needs, such as whether the phone number field conforms to the phone number format specification; the data integrity detection rule ensures that all necessary fields in the data record exist and there is no missing situation. When detecting user information data, the null-checking detection will check whether keywords such as name and ID number are empty, the data compliance detection will check whether the ID number conforms to the 18-digit number format requirement, and the data integrity detection will ensure that all required fields have values.
[0040] Step 40: Classify and store the exception logs in the distributed file system according to the timestamp, generate a visual summary report, parse the problem type, field identifier, and original data, and call the predefined repair script library for repair.
[0041] As an alternative embodiment, the exception logs are classified and stored in a distributed file system according to timestamps, and a visual summary report is generated, which may specifically include: Step 401: Extract the exception logs in the distributed file system and perform aggregation statistics according to problem types; Step 402: Calculate the data quality score based on the aggregation result and generate an interactive HTML report; Step 403: Push the interactive HTML report to a message queue for the downstream system to trigger an automatic repair process.
[0042] In this step, for example, the generated HTML report is pushed to a message queue such as Kafka or RabbitMQ. A message queue is an asynchronous communication mechanism that can decouple the upstream log processing system from the downstream repair system. The downstream system includes an operation and maintenance system and a self-healing system for faults. The downstream system obtains the report from the message queue and, based on the information in the report, that is, the problem type and the data quality score, can trigger the corresponding automatic repair process. For example, it can restart the faulty service, adjust the system configuration, and send an alarm notification to the operation and maintenance personnel.
[0043] As an alternative embodiment, the problem type, field identifier, and original data are parsed, and a predefined repair script library is called for repair, which may specifically include: Step 404: Select the optimal repair strategy according to the script matching degree. If the score is lower than the threshold, transfer it to manual processing; Step 405: Re-inject the repaired data into the DataX data transmission stream and trigger a secondary quality inspection; Step 406: If the secondary quality inspection passes, update the data quality score and close the exception work order.
[0044] In this step, the predefined repair script library is traversed, the matching degree of each script is calculated according to the problem type and the field identifier, and a matching degree threshold is set. If the matching degree of the optimal strategy is lower than this threshold, it is considered that there is no suitable automatic repair script and it needs to be transferred to manual processing; the repaired data is re-injected into the DataX data transmission stream to ensure that the data can continue to be used by the subsequent processing flow. After the data is re-injected, a secondary quality inspection process is triggered.
[0045] As an alternative embodiment, after calling the predefined repair script library for repair, the method may further include: Extracting the exception logs with a problem type higher than a preset frequency, extracting rule optimization feature vectors; Inputting the feature vectors into a rule model to generate a new rule candidate set; Performing conflict detection on the candidate set. If the check of the conflict detection passes, automatically merge the candidate set into the detection rule library; Triggering an incremental data quality re-inspection based on the new rule candidate set and updating the visual summary report.
[0046] In this step, the filtered abnormal logs are parsed to extract feature information related to rule optimization. According to the input feature vector, the extracted feature vector is input into the rule model to generate a candidate set of new rules and generate new detection rules. Conflict detection is performed on the generated candidate set of new rules to ensure that the new rules do not conflict with the rules in the existing rule library, which may include overlapping rule conditions, contradictory trigger thresholds, and conflicting response actions.
[0047] As an alternative embodiment, inputting the feature vector into the rule model to generate a candidate set of new rules may specifically include: using a Transformer model to learn the relevance of the feature vector to maximize the rule coverage rate to obtain the rule model; deploying the rule model as a microservice to enable online incremental learning and real-time rule push.
[0048] In this step, the Transformer model is trained using the training dataset to optimize the model parameters so that the model can accurately learn the relevance between feature vectors and make the generated rules cover as many abnormal scenarios as possible. After the model training is completed, the model is used to predict new feature vectors to generate a candidate set of new rules, and the generated rules include clear detection conditions, trigger thresholds, and response actions for subsequent data quality detection and processing.
[0049] The above is the method embodiment proposed in this application. Based on the same inventive concept, the embodiment of this application also provides a device for data quality detection of DataX, and its structure is as Figure 2 shown.
[0050] Figure 2 This is a schematic internal structure diagram of a device for data quality detection of DataX provided by the embodiment of this application. As Figure 2 shown, the device includes:
[0051] At least one processor 201;
[0052] And a memory 202 communicatively connected to at least one processor;
[0053] Among them, the memory 202 stores instructions executable by at least one processor. The instructions are executed by at least one processor 201, enabling at least one processor 201 to: read the data source in parallel slices through the reader plug-in of DataX to obtain the data to be processed; when the data to be processed is transmitted to the memory buffer, trigger the quality detection engine, and load the predefined detection rule library through the quality detection engine. The detection rule library includes null judgment rules, compliance templates, and integrity constraints; if it is detected that a field is empty, the format deviates from the preset template, or a required field is missing, generate an exception log including the problem type, field identifier, and original data; classify and store the exception log in the distributed file system according to the timestamp, generate a visual summary report, parse the problem type, field identifier, and original data, and call the predefined repair script library for repair.
[0054] Some embodiments of the present application provide a non-volatile computer storage medium for data quality detection of DataX corresponding to Figure 1 which stores computer-executable instructions. The computer-executable instructions are set to: read the data source in parallel slices through the reader plug-in of DataX to obtain the data to be processed; when the data to be processed is transmitted to the memory buffer, trigger the quality detection engine, and load the predefined detection rule library through the quality detection engine. The detection rule library includes null judgment rules, compliance templates, and integrity constraints; if it is detected that a field is empty, the format deviates from the preset template, or a required field is missing, generate an exception log including the problem type, field identifier, and original data; classify and store the exception log in the distributed file system according to the timestamp, generate a visual summary report, parse the problem type, field identifier, and original data, and call the predefined repair script library for repair.
[0055] The embodiments in the present application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the embodiments of the Internet of Things devices and media, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiments.
[0056] The systems and media provided by the embodiments of the present application correspond one-to-one with the methods. Therefore, the systems and media also have beneficial technical effects similar to those of the corresponding methods. Since the beneficial technical effects of the methods have been described in detail above, the beneficial technical effects of the systems and media will not be elaborated here.
[0057] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0058] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0059] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0060] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0061] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0062] The memory may include non-permanent memory in the computer-readable medium, random access memory (RAM), and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0063] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0064] It should also be noted that the term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0065] The above description is only for the embodiments of the present application and is not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the scope of the claims of the present application.
Claims
1. A method for data quality detection of DataX, characterized in that The method includes: Parallelly sharding and reading the data source through the reader plugin of DataX to obtain the data to be processed; When the data to be processed is transmitted to the memory buffer, trigger the quality detection engine, and load the predefined detection rule library through the quality detection engine. The detection rule library includes null judgment rules, compliance templates, and integrity constraints; If it is detected that a field is empty, the format deviates from the preset template, or a required field is missing, generate an exception log including the problem type, field identifier, and original data; Classify and store the exception log in the distributed file system according to the timestamp, generate a visual summary report, and parse the problem type, field identifier, and original data, and call the predefined repair script library for repair.
2. The method for data quality detection for DataX according to claim 1, wherein The step of parallelly sharding and reading the data source through the reader plugin of DataX to obtain the data to be processed specifically includes: Based on the connection parameters of the data source, analyze the distribution characteristics of the connection parameters; Calculate the optimal sharding strategy according to the distribution characteristics, allocate independent threads for each shard, and concurrently read the shard data through the reader plugin; Cache the shard data in the memory buffer in preset batches; During the reading process, monitor the throughput rate of the shards in real time. If it is detected that the shard delay exceeds the threshold, trigger the resharding operation.
3. The method for data quality detection for DataX according to claim 2, wherein The step of monitoring the throughput rate of the shards in real time during the reading process and triggering the resharding operation if it is detected that the shard delay exceeds the threshold specifically includes: Monitor the CPU occupancy rate and network delay of the shard, and construct a shard health index; If the health index is lower than the preset threshold, split the current shard into multiple sub-shards; Allocate new threads for the sub-shards, migrate them to the unfinished data reading tasks, and update the shard metadata to the ZooKeeper cluster.
4. A method for data quality detection of DataX according to claim 1, characterized in that, The step of classifying and storing the exception log in the distributed file system according to the timestamp and generating a visual summary report specifically includes: Extract the exception log in the distributed file system and perform aggregation statistics according to the problem type; Calculate the data quality score based on the aggregation result and generate an interactive HTML report; Push the interactive HTML report to the message queue for the downstream system to trigger the automatic repair process.
5. A method for data quality detection of DataX according to claim 4, characterized in that The step of parsing the problem type, field identifier, and original data and calling the predefined repair script library for repair specifically includes: Select the optimal repair strategy according to the script matching degree. If the score is lower than the threshold, transfer it to manual processing; Re-inject the repaired data into the DataX data transmission stream and trigger the secondary quality detection; If the secondary quality detection passes, update the data quality score and close the exception work order.
6. A method for data quality detection of DataX according to claim 1, characterized in that, The step of calculating the optimal sharding strategy according to the distribution characteristics specifically includes: Collect the metadata of the data source, and the metadata includes table size, index distribution, and storage engine type; Construct a sharding weight matrix based on the metadata; Combined with the network bandwidth and the number of CPU cores, use the greedy algorithm to select the sharding boundary, and in the process of sharding execution, correct the weight matrix in real time.
7. A method for data quality detection of DataX according to claim 1, characterized in that After the step of calling the predefined repair script library for repair, the method further includes: Extract the abnormal logs of the problem types higher than the preset frequency, and optimize the feature vectors of the extraction rules; Input the feature vectors into the rule model to generate a candidate set of new rules; Perform conflict detection on the candidate set. If the verification of the conflict detection is passed, automatically merge the candidate set into the detection rule library; Trigger incremental data quality recheck based on the candidate set of new rules and update the visual summary report.
8. A method for data quality detection of DataX according to claim 7, characterized in that, Input the feature vectors into the rule model to generate a candidate set of new rules, specifically including: Use the Transformer model to learn the relevance of the feature vectors to maximize the rule coverage rate and obtain the rule model; Deploy the rule model as a microservice to enable online incremental learning and real-time rule push.
9. An apparatus for data quality detection of DataX, characterized in that, The device includes: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can: Read the data source in parallel slices through the reader plugin of DataX to obtain the data to be processed; When the data to be processed is transmitted to the memory buffer, trigger the quality detection engine, and load the predefined detection rule library through the quality detection engine. The detection rule library includes null judgment rules, compliance templates and integrity constraints; If it is detected that a field is empty, the format deviates from the preset template, or a required field is missing, generate an abnormal log including the problem type, field identifier and original data; Classify and store the abnormal logs in the distributed file system according to the timestamp, generate a visual summary report, parse the problem type, field identifier and original data, and call the predefined repair script library for repair.
10. A non-volatile computer storage medium for data quality detection of DataX, storing computer-executable instructions, characterized in that, The computer-executable instructions are set to: Read the data source in parallel slices through the reader plugin of DataX to obtain the data to be processed; When the data to be processed is transmitted to the memory buffer, trigger the quality detection engine, and load the predefined detection rule library through the quality detection engine. The detection rule library includes null judgment rules, compliance templates and integrity constraints; If it is detected that a field is empty, the format deviates from the preset template, or a required field is missing, generate an abnormal log including the problem type, field identifier and original data; Classify and store the abnormal logs in the distributed file system according to the timestamp, generate a visual summary report, parse the problem type, field identifier and original data, and call the predefined repair script library for repair.
Citation Information
Cited By
HTTP (Hyper Text Transport Protocol) intelligent parameter change sensing and directional matching method and device
CN120929652A
An HTTP intelligent parameter variable sensing and directional matching method and device
CN120929652B
Integration method of judicial library system and financial core system, equipment and medium
CN121117087A
A method, device and medium for integrating a treasurer system and a financial core system
CN121117087B
Data synchronization method and system for intelligent dirty data detection and restoration based on DataX
CN121301477A