Unified data quality auditing method, device and equipment based on Apache Arrow
Through the unified data quality audit method based on Apache Arrow, the problems of high intrusion in data quality auditing of data processing frameworks in the existing technology are solved, and comprehensive quality inspection and automatic data error correction of multiple data processing components are achieved, and the consistency and reusability of data quality management are improved.
Patent Information
- Application Number
- CN202510089467.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-06
AI Technical Summary
The existing data processing frameworks have high invasiveness in data quality audits, high customization development costs, and difficulty in migrating between different frameworks, resulting in the challenge of consistency and reusability of data quality management.
The unified data quality audit method based on Apache Arrow is adopted, and the data of various data processing frameworks is converted into Apache Arrow format through the data access module, data auditing is used to perform data auditing using the responsibility chain mode, and diversion processing is performed through the data output module, supporting automatic data error correction and repair.
It realizes comprehensive quality inspection of process data of multiple data processing components, supports rapid iteration and customization, reduces development and maintenance costs, and improves the consistency and reusability of data quality management.
Smart Images

Figure CN120104600A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present invention generally relate to the field of data processing technology, and more particularly to a unified data quality audit method and device based on Apache Arrow. Background Art
[0002] During data processing, monitoring of abnormal data and automatic error correction often cause great intrusion to business data processing programs. This requirement requires data processing frameworks such as DataX, Kettle, and Spark to be adapted and customized to achieve data quality audits. This intrusion not only increases development difficulty, but also increases maintenance costs.
[0003] In addition, data quality audit procedures in data processing pipelines often need to be highly customized to adapt to specific data processing frameworks. Such customized programs have high development costs and are difficult to migrate between different frameworks, making consistency and reusability of data quality management between different platforms a challenge. Summary of the invention
[0004] To solve the above problems, the present invention provides a set of configurable and extensible data quality control methods and devices, which can perform comprehensive quality checks on process data in various data processing components; can automatically trigger data error correction and repair mechanisms, support rapid iteration and customization, and enhance the extensibility of the system through abstract classes and template methods, making it easy for developers to quickly adapt new data quality audit components.
[0005] According to an embodiment of the present invention, a method, apparatus and device for unified data quality audit based on Apache Arrow are provided.
[0006] In a first aspect of the present invention, a method for unified data quality audit based on Apache Arrow is provided. The method comprises:
[0007] S01: The data access module accesses process data from various data processing frameworks, converts all accessed data into data in Apache Arrow format using the adapter mode, constructs the Arrow Flight gRPC client, and sends the data in Apache Arrow format to the data quality control module;
[0008] S02: The data quality control module, as the Arrow Flight gRPC server, receives the data in Apache Arrow format converted by the data access module, processes the data according to one or more audit rules using the chain of responsibility model, and outputs the data in Apache Arrow format, and sends the audited data to the data output module;
[0009] S03: The data output module receives the Apache Arrow format data in response from the Arrow Flight gRPC server of the data quality control module, which includes three types of data: data source data processed by the quality audit control module, exception detail data, and audit indicator data. These three types of data are diverted for processing: the data source data processed by the quality audit control module is converted into data of various data processing components; the exception detail data is stored in external storage; and the audit indicator data is stored in the database for analysis of audit quality indicators.
[0010] Furthermore, the implementation steps of converting all the accessed data into data in Apache Arrow format using the adapter pattern in S01 are: defining a basic data adapter interface; implementing an adapter of a data processing component; wherein the method of defining the basic data adapter interface includes: a method of inferring the Arrow Schema of data, and a method of converting data into Apache Arrow format.
[0011] Furthermore, the Arrow Schema method for inferring data is used to accurately infer the structure and type of data and generate a Schema definition that conforms to the Apache Arrow format. The implementation steps are:
[0012] Dataset sampling: sampling large-scale datasets;
[0013] Data type inference: traverse each column of the data to perform type analysis, remove null values and analyze the characteristics of non-null values, and infer the Apache Arrow type according to the following rules: Integer type: select the appropriate bit width according to the value range; Floating point type: use float64 type by default; Boolean type: map to bool type; Time and date type: convert to timestamp type; String type: select string or large_string according to length; List type: map to Apache Arrow array; Dictionary type: convert to structure; Nested data: map to nested structure; Other types: convert to string type by default;
[0014] Schema construction: Create an Apache Arrow field list. Each field contains: field name, data type, whether it can be null, and related metadata. Merge field information to build a complete Schema object.
[0015] Furthermore, the method of converting data into the Apache Arrow format is used to achieve efficient cross-language and cross-platform data exchange and processing through a unified data representation method, and the implementation steps are:
[0016] The method to obtain the inferred data Schema information of the Arrow Schema of the inferred data;
[0017] Initialize Apache Arrow memory management and buffers;
[0018] Data type conversion: According to the mapping relationship between the actual data and the inferred Schema field, the data type is converted: Integer type: select the appropriate bit width according to the value range; floating point number: convert to 32-bit or 64-bit floating point number; string: convert to fixed-length or variable-length string; Boolean value: convert to Apache Arrow Boolean type; date and time: convert to Apache Arrow timestamp type; list is converted to Apache Arrow array; dictionary is converted to structure; nested data is converted to nested structure;
[0019] Memory layout organization: Organize data by column, allocate memory buffer for each column, use parallel processing to accelerate conversion, reuse memory buffer, and achieve zero-copy operation;
[0020] Exception handling: Null value handling: mark the null value position, maintain the null value mask, and handle different types of null value representations; Exception value handling: data out-of-bounds check, type mismatch handling, and format error handling;
[0021] Result output: Generate Apache Arrow data table.
[0022] Furthermore, the processing flow of the data quality control module described in S02 is:
[0023] Input layer: receives source data in the unified Apache Arrow format
[0024] Rule chain construction: dynamically assemble audit rule chains based on configuration
[0025] Rule execution: A single rule is processed and the result is output in Apache Arrow format. The result is used as the input of the next rule for further processing. Parallel processing and distributed expansion are supported.
[0026] Result output: Generate unified quality reports and processed data.
[0027] Furthermore, the functions implemented by the audit rules described in S02 include: repairing abnormal data; filtering abnormal data; storing abnormal detailed data; and audit verification.
[0028] Furthermore, the steps of converting the data source data processed by the quality audit control module into the data of various data processing components in S03 are: defining an output adapter base class; implementing adapters of various data processing components; wherein the method of defining the output adapter base class includes: converting the Apache Arrow Schema into a schema in a target format; converting the Apache Arrow data into a target component format.
[0029] Furthermore, the schema for converting the Apache Arrow Schema to the target format is used to reassemble the data from the Apache Arrow format into the original format suitable for the target platform, ensuring that the data can flow smoothly to the next stage. The implementation steps are:
[0030] Analyze the structure of Apache Arrow Schema, including field names, data types, and field order;
[0031] Map the data type of each field in the Apache Arrow Schema to the data type of the target component; use the mapped data type to build the Schema of the target component.
[0032] Furthermore, the steps for converting Apache Arrow data into the target component format are:
[0033] Map Apache Arrow Schema to target component Schema: Map the fields in Apache Arrow Schema to the Schema of the target component format, including mapping of field names, data types, and field order;
[0034] Data type conversion: According to the mapping relationship, the data type in Apache Arrow is converted into the data type supported by the target format;
[0035] Data conversion: Convert Apache Arrow data to the target format. For Spark, use the pyApacheArrow library to directly convert Apache ArrowTable to SparkDataFrame; for DataX, convert Apache Arrow data to DataX's data format.
[0036] In a second aspect of the present invention, a device for unified data quality audit based on Apache Arrow is provided. The device comprises:
[0037] Data access module: used to access process data of various data processing frameworks, convert all accessed data into Apache Arrow format data using the adapter mode, construct the Arrow Flight gRPC client, and send the Apache Arrow format data to the data quality control module;
[0038] Data quality control module: used to receive the data in Apache Arrow format converted by the data access module, process the data according to one or more audit rules using the chain of responsibility model, and then output the data in Apache Arrow format, and send the audited data to the data output module;
[0039] Data output module: used to receive Apache Arrow format data in response from the Arrow Flight gRPC server of the data quality control module, which includes three types of data: data source data processed by the quality audit control module, exception detail data, and audit indicator data. These three types of data are diverted for processing: the data source data processed by the quality audit control module is converted into data of various data processing components; the exception detail data is stored in external storage; and the audit indicator data is stored in the database for analysis of audit quality indicators.
[0040] In a third aspect of the present invention, an electronic device is provided, comprising: a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the program, the method according to the first aspect of the present invention is implemented.
[0041] In a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the method according to the first aspect of the present invention is implemented.
[0042] The meanings of the above mentioned English abbreviations:
[0043] Apache Arrow: A cross-language columnar memory format designed to accelerate data analysis and distributed computing tasks through zero-copy data sharing
[0044] Arrow Flight gRPC: Arrow data transmission protocol framework based on gRPC
[0045] Schema: metadata description that defines the data structure
[0046] float64: double-precision floating point number, occupies 64 bits of storage space
[0047] bool: Boolean data type, value is true or false
[0048] timestamp: timestamp data type, accurate to nanoseconds
[0049] string: string data type, UTF-8 encoding
[0050] large_string: Large string data type
[0051] pyApache Arrow: A Python library providing a columnar in-memory format
[0052] Apache ArrowTable: Tabular data structure in Arrow format
[0053] Spark: Distributed Computing Platform
[0054] DataFrame: A distributed dataset based on RDD
[0055] DataX: A tool for offline synchronization of heterogeneous data sources
[0056] The present invention provides a configurable and extensible data quality control method and device, which can perform comprehensive quality checks on process data in various data processing components; can automatically trigger data error correction and repair mechanisms, support rapid iteration and customization, and enhance the extensibility of the system through abstract classes and template methods, making it easy for developers to quickly adapt new data quality audit components.
[0057] It should be understood that the contents described in the summary of the invention are not intended to limit the key or important features of the embodiments of the present invention, nor are they intended to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The above and other features, advantages and aspects of the embodiments of the present invention will become more apparent with reference to the following detailed description in conjunction with the accompanying drawings.
[0059] Figure 1 A flow chart of a method for unified data quality audit based on Apache Arrow according to an embodiment of the present invention is shown;
[0060] Figure 2 A block diagram of a device for unified data quality audit based on Apache Arrow according to an embodiment of the present invention is shown;
[0061] Figure 3 A schematic diagram of a device for unified data quality audit based on Apache Arrow according to an embodiment of the present invention is shown;
[0062] Figure 4 It shows a schematic diagram of the overall process according to an embodiment of the present invention;
[0063] Figure 5 A schematic diagram of a data access module process according to an embodiment of the present invention is shown;
[0064] Figure 6 A schematic diagram of a data quality control module process according to an embodiment of the present invention is shown;
[0065] Figure 7 A schematic diagram of a data output module flow according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0066] In order to make the purpose, technical solution and advantages of the embodiments of the present invention clearer, the technical solution in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0067] According to an embodiment of the present invention, a method, apparatus and device for unified data quality auditing based on Apache Arrow are proposed, which can perform comprehensive quality checks on process data in various data processing components; can automatically trigger data error correction and repair mechanisms, support rapid iteration and customization, and enhance the scalability of the system through abstract classes and template methods, so as to facilitate developers to quickly adapt new data quality auditing components.
[0068] The principle and spirit of the present invention are explained in detail below with reference to several representative embodiments of the present invention.
[0069] Figure 1 1 is a flow chart of a method for unified data quality audit based on Apache Arrow according to an embodiment of the present invention. The method includes:
[0070] S01: The data access module accesses process data from various data processing frameworks, converts all accessed data into data in Apache Arrow format using the adapter mode, constructs the Arrow Flight gRPC client, and sends the data in Apache Arrow format to the data quality control module;
[0071] S02: The data quality control module, as the Arrow Flight gRPC server, receives the data in Apache Arrow format converted by the data access module, processes the data according to one or more audit rules using the chain of responsibility model, and outputs the data in Apache Arrow format, and sends the audited data to the data output module;
[0072] S03: The data output module receives the Apache Arrow format data in response from the Arrow Flight gRPC server of the data quality control module, which includes three types of data: data source data processed by the quality audit control module, exception detail data, and audit indicator data. These three types of data are diverted for processing: the data source data processed by the quality audit control module is converted into data of various data processing components; the exception detail data is stored in external storage; and the audit indicator data is stored in the database for analysis of audit quality indicators.
[0073] It should be noted that, although the operations of the method of the present invention are described in a specific order in the above embodiments and the accompanying drawings, this does not require or imply that the operations must be performed in the specific order, or that all the operations shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps.
[0074] In order to explain the above-mentioned unified data quality audit method based on Apache Arrow more clearly, a specific embodiment is used for illustration below. However, it should be noted that this embodiment is only for better illustrating the present invention and does not constitute an improper limitation to the present invention.
[0075] The following is a specific example to further illustrate a unified data quality audit method based on Apache Arrow:
[0076] like Figure 4 As shown, the data quality audit device includes three modules: a data access module, a data quality control module and a data output module.
[0077] 1. The data access module is responsible for accessing process data from various data processing frameworks and converting it into data in the Apache Arrow format with efficient columnar memory storage and zero-copy features. Metadata information such as the field name of the data is stored in the Apache Arrow Schema for subsequent modules to automatically identify the structure and type of the data.
[0078] 2. The data quality control module, as the Arrow Flight gRPC server, receives the data in Apache Arrow format converted by the data access module. After being processed by one or more audit rules, the output of each audit rule is data in Apache Arrow format, which serves as the input of the next audit rule, thereby realizing the chain connection of multiple audit rules and improving scalability.
[0079] 3. The data output module receives Apache Arrow format data from the data quality audit control module and performs diversion processing: business data is output to the write link of the data processing framework for subsequent processing; abnormal detail data is stored in external storage; and audit indicator data is stored in the database to support the analysis of audit quality indicators.
[0080] The specific contents are as follows:
[0081] The data access module is responsible for receiving data from various data processing components, but there are the following pain points:
[0082] 1. Complexity of data format conversion: The data grids output by different data processing components (DataX, Kettle, Spark) use different data formats and serialization methods. The metadata definition and type system of each component are different, and multiple links such as field mapping, type conversion, and special value processing need to be processed.
[0083] 2. Performance bottleneck: Format conversion of large-scale data will bring additional CPU and memory overhead, and frequent data serialization and deserialization will affect the overall throughput.
[0084] 3. Scalability issues: How to automatically adapt when adding new data processing components to ensure fast data access.
[0085] In order to solve the above pain points, Figure 5 As shown in the figure, a unified data conversion interface is defined to unify the data of various data processing components into the Apache Arrow format. Then the Arrow Flight gRPC client is constructed to send the connected Apache Arrow data to the data quality control module.
[0086] The Apache Arrow format is suitable for serializing data in memory without storing it on disk. The data in the Apache Arrow format includes the data and its schema. The schema defines the field name and data type of the data, which unifies the data format of the data connected to various data processing components and facilitates subsequent processing. In order to achieve the scalability of data format conversion, the adapter mode is used to implement data source conversion, provide a standardized conversion interface, and convert data to the Apache Arrow format in memory.
[0087] The implementation steps include:
[0088] 1. Define the basic data adapter interface.
[0089] The basic data adapter interface mainly includes two methods: the Arrow Schema method for inferring data and the method for converting data into Apache Arrow format. Each data processing component inherits the basic data adapter interface and implements its own conversion logic.
[0090] Specifically, the main goal of the Arrow Schema method for inferring data is to accurately infer the structure and type of the data and generate a Schema definition that conforms to the Arrow format, providing a basis for subsequent data conversion. The implementation needs to consider the diversity and complexity of the data to ensure the accuracy and efficiency of type inference. Implementation steps:
[0091] 1) Dataset sampling: Sampling large-scale datasets, usually taking the first 1,000 rows or a specified number of samples;
[0092] 2) Data type inference: Traverse each column of the data to perform type analysis, remove null values and analyze the characteristics of non-null values. Infer the Arrow type according to the following rules:
[0093] Integer type: select the appropriate bit width (int8 / int16 / int32 / int64) according to the value range
[0094] Floating point type: float64 type is used by default
[0095] Boolean type: mapped to bool type
[0096] Date and time type: Convert to timestamp type
[0097] String type: select string or large_string according to the length
[0098] List type: mapped to Arrow array
[0099] Dictionary type: Convert to structure
[0100] Nested data: mapped as nested structures
[0101] Other types: converted to string type by default;
[0102] 3) Schema construction: Create an Arrow field list, each field contains: field name, data type, whether it can be empty, and related metadata. Merge field information to build a complete Schema object.
[0103]
[0104] Method for converting data into Apache Arrow format: Through a unified data representation method, efficient data exchange and processing across languages and platforms can be achieved. Data access and computing performance can be improved through columnar storage and optimized memory layout. Zero-copy operations are supported to reduce memory consumption and provide an efficient memory management mechanism. At the same time, the standardization and consistency of the data format are ensured, complex data structures are supported, and finally high performance, high availability and good interoperability of the data processing system are achieved. Implementation steps:
[0105] 1) Obtain the data schema information inferred by method 1;
[0106] 2) Initialize Arrow memory management and buffers;
[0107] 3) Data type conversion. According to the mapping relationship between the actual data and the inferred Schema field, the data type is converted:
[0108] Integer type: choose the appropriate bit width according to the value range
[0109] Floating point: Convert to 32-bit or 64-bit floating point
[0110] String: Convert to fixed-length or variable-length string
[0111] Boolean value: Convert to Arrow Boolean type
[0112] Datetime: Convert to Arrow timestamp type
[0113] Convert list to Arrow array
[0114] Convert dictionary to structure
[0115] Nested data is converted into nested structures;
[0116] 4) Memory layout organization: Organize data by column, allocate memory buffer for each column, use parallel processing to accelerate conversion, reuse memory buffer, and achieve zero-copy operation;
[0117] 5) Exception handling;
[0118] Null value processing: mark null value positions, maintain null value masks, and handle different types of null value representations.
[0119] Abnormal value processing: data out-of-bounds check, type mismatch processing, format error processing;
[0120] 6) Result output: Generate Arrow data table.
[0121]
[0122] 2. Adapter implementation of data processing components
[0123] Each data processing component, such as Spark and Datax, inherits the basic data adapter interface and implements the corresponding methods. The following code example is a Spark adapter implementation, which converts Spark DataFrame to ArrowSchema and DataFrame to ArrowTable.
[0124]
[0125] The following is an example of DataX's adapter implementation:
[0126]
[0127]
[0128]
[0129] like Figure 6 As shown in the figure, the data quality control module uses the responsibility chain model to implement the data quality audit process, and uses the unified Arrow format to realize the data flow between rules. Each audit rule is an independent processing unit with standardized input and output interfaces. It supports flexible rule arrangement and can dynamically adjust the rule execution order.
[0130] The unified Arrow format ensures data structure consistency, supports dynamic addition of new audit rules, has low coupling between rules, and is easy to maintain. The processing flow is as follows:
[0131] 1. Input layer: receives source data in a unified Arrow format;
[0132] 2. Rule chain construction: dynamically assemble the audit rule chain according to the configuration;
[0133] 3. Rule execution: A single rule is processed and the result is output in Arrow format. The result is used as the input of the next rule for further processing. Parallel processing and distributed expansion are supported.
[0134] 4. Result output: Generate unified quality reports and processed data.
[0135] The functions implemented by each data quality audit rule are as follows:
[0136] 1. Fix abnormal data. For example, set the null value data to the default value, and set the number that exceeds the threshold to a specific value (such as mean, maximum value, minimum value, etc.).
[0137] 2. Filter abnormal data: Determine whether to filter abnormal data based on user configuration to prevent it from flowing to the next link.
[0138] 3. Abnormal detail data storage: Identify abnormal data based on the configured audit rules. For example, for null value audit, extract null value records as abnormal details and store them for subsequent analysis.
[0139] 4. Audit verification: Generate corresponding audit indicators according to the configured audit rules. For example, the indicators for null value audit include total number of records and number of null value records. Based on these indicators, configure the audit verification formula to determine whether the audit result is passed. For example, for null value audit, you can set the verification formula: number of null value records / total number of records <0.1, that is, when the null value ratio is less than 10%, the audit is passed.
[0140] like Figure 7 As shown, the data output module receives the Apache Arrow format data in response from the Arrow Flight gRPC server of the data quality control module, which mainly includes three types of data: data source data processed by the quality audit control module, exception detail data, and audit indicator data.
[0141] Data is diverted and processed: data source data processed by the quality audit control module is converted into data of various data processing components for subsequent processing; abnormal detail data is stored in external storage; audit indicator data is stored in the database to support the analysis of audit quality indicators.
[0142] In order to ensure the universality and scalability of data writing to various data processing components, an adapter module is also used to convert data in the unified Apache Arrow format into data of various data processing components. The steps include:
[0143] 1. Define the output adapter base class and implement two methods: 1) convert the Arrow Schema to the schema of the target format; 2) convert the Arrow data to the target component format.
[0144] 1) The goal of converting the Arrow Schema to the target format schema is to reassemble the data from the Arrow format into the original format suitable for the target platform, ensuring that the data can flow smoothly to the next stage.
[0145] The implementation logic is as follows:
[0146] (1) Analyze the structure of Apache Arrow Schema, including field names, data types, field order, etc.
[0147] (2) Map the data type of each field in the Arrow Schema to the data type of the target component. For example, map the int32 type of Arrow to the IntegerType of Spark or the corresponding integer type of DataX.
[0148] (3) Use the mapped data types to build the schema of the target component.
[0149] 2) The implementation logic for converting Arrow data to the target component format is as follows:
[0150] (1) Map Arrow Schema to target component schema: Map the fields in Arrow Schema to the schema of the target component format. This includes mapping of field names, data types, and field order.
[0151] (2) Data type conversion: According to the mapping relationship, the data type in Arrow is converted to the data type supported by the target format. For example, convert Arrow's int32 to Spark's IntegerType.
[0152] (3) Data conversion: Convert Arrow data to the target format. For Spark, use the pyarrow library to directly convert ArrowTable to Spark DataFrame; for DataX, convert Arrow data to DataX data format.
[0153]
[0154] 2. Adapter implementation of various data processing components. Inherit the adapter base class and implement the above methods.
[0155]
[0156]
[0157] Based on the same inventive concept, the present invention also proposes a unified data quality audit device based on Apache Arrow. The implementation of the device can refer to the implementation of the above method, and the repeated parts will not be repeated. Figure 2 As shown, the device 100 includes:
[0158] Data access module 101: used to access process data of various data processing frameworks, convert all accessed data into data in Apache Arrow format using the adapter mode, construct Arrow Flight gRPC client, and send data in Apache Arrow format to the data quality control module;
[0159] Data quality control module 102: used to receive the data in Apache Arrow format converted by the data access module, process the data according to one or more audit rules using the chain of responsibility model, and then output the data in Apache Arrow format, and send the audited data to the data output module;
[0160] Data output module 103: used to receive Apache Arrow format data in response from the Arrow Flight gRPC server of the data quality control module, including three types of data: data source data processed by the quality audit control module, exception detail data and audit indicator data, and these three types of data are diverted for processing: the data source data processed by the quality audit control module is converted into data of various data processing components; the exception detail data is stored in external storage; and the audit indicator data is stored in the database for analysis of audit quality indicators.
[0161] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the described module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.
[0162] like Figure 3 As shown, the device includes a central processing unit (CPU), which can perform various appropriate actions and processes according to computer program instructions stored in a read-only memory (ROM) or loaded from a storage unit into a random access memory (RAM). In the RAM, various programs and data required for the operation of the device can also be stored. The CPU, ROM and RAM are connected to each other through a bus. An input / output (I / O) interface is also connected to the bus.
[0163] Multiple components in the device are connected to the I / O interface, including: input units, such as keyboards, mice, etc.; output units, such as various types of displays, speakers, etc.; storage units, such as disks, optical disks, etc.; and communication units, such as network cards, modems, wireless communication transceivers, etc. The communication unit allows the device to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunication networks.
[0164] The processing unit performs the various methods and processes described above, such as methods S01 to S03. For example, in some embodiments, methods S01 to S03 may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit. In some embodiments, part or all of the computer program may be loaded and / or installed on the device via a ROM and / or a communication unit. When the computer program is loaded into the RAM and executed by the CPU, one or more steps of methods S01 to S03 described above may be performed. Alternatively, in other embodiments, the CPU may be configured to execute methods S01 to S03 in any other appropriate manner (e.g., by means of firmware).
[0165] The functions described above herein may be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip systems (SOCs), load programmable logic devices (CPLDs), and the like.
[0166] The program code for implementing the method of the present invention can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0167] In the context of the present invention, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0168] In addition, although each operation is described in a specific order, this should be understood as requiring such operation to be performed in the specific order shown or in a sequential order, or requiring that all illustrated operations should be performed to obtain desired results. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details are included in the above discussion, these should not be interpreted as limiting the scope of the present invention. Some features described in the context of a separate embodiment can also be implemented in a single implementation in combination. On the contrary, the various features described in the context of a single implementation can also be implemented in multiple implementations individually or in any suitable sub-combination mode.
[0169] Although the subject matter has been described in language specific to structural features and / or methodological logical actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. On the contrary, the specific features and actions described above are merely example forms of implementing the claims.
Claims
1. A unified data quality audit method based on Apache Arrow, characterized in that: The method includes: S01: The data access module accesses process data from various data processing frameworks, converts all accessed data into Apache Arrow format data using the adapter mode, constructs the Arrow Flight gRPC client, and sends the Apache Arrow format data to the data quality control module; S02: The data quality control module, as the Arrow Flight gRPC server, receives the data in Apache Arrow format converted by the data access module, processes the data according to one or more audit rules using the chain of responsibility model, and outputs the data in Apache Arrow format, and sends the audited data to the data output module; S03: The data output module receives the Apache Arrow format data in response from the Arrow Flight gRPC server of the data quality control module, which includes three types of data: data source data processed by the quality audit control module, exception detail data, and audit indicator data. These three types of data are diverted for processing: the data source data processed by the quality audit control module is converted into data of various data processing components; the exception detail data is stored in external storage; and the audit indicator data is stored in the database for analysis of audit quality indicators.
2. According to claim 1, a unified data quality audit method based on Apache Arrow is characterized in that: The implementation steps of converting all the accessed data into the data in Apache Arrow format using the adapter pattern described in S01 are: defining a basic data adapter interface; implementing an adapter of a data processing component; wherein the methods included in defining the basic data adapter interface are: a method for inferring the Arrow Schema of data, and a method for converting data into the Apache Arrow format.
3. According to claim 2, a unified data quality audit method based on Apache Arrow is characterized in that: The Arrow Schema method for inferring data is used to accurately infer the structure and type of data and generate a Schema definition that conforms to the Apache Arrow format. The implementation steps are: Dataset sampling: sampling large-scale datasets; Data type inference: traverse each column of the data to perform type analysis, remove null values and analyze the characteristics of non-null values, and infer the Apache Arrow type according to the following rules: Integer type: select the appropriate bit width according to the value range; Floating point type: use float64 type by default; Boolean type: map to bool type; Time and date type: convert to timestamp type; String type: select string or large_string according to length; List type: map to Apache Arrow array; Dictionary type: convert to structure; Nested data: map to nested structure; Other types: convert to string type by default; Schema construction: Create an Apache Arrow field list. Each field contains: field name, data type, whether it can be null, and related metadata. Merge field information to build a complete Schema object.
4. According to claim 3, a unified data quality audit method based on Apache Arrow is characterized in that: The method of converting data into the Apache Arrow format is used to achieve efficient cross-language and cross-platform data exchange and processing through a unified data representation method, and the implementation steps are: The method to obtain the inferred data Schema information of the Arrow Schema of the inferred data; Initialize Apache Arrow memory management and buffers; Data type conversion: According to the mapping relationship between the actual data and the inferred Schema field, the data type is converted: Integer type: select the appropriate bit width according to the value range; Floating point number: convert to 32-bit or 64-bit floating point number; String: converted to fixed-length or variable-length string; Boolean value: converted to Apache Arrow Boolean type; Date and time: converted to Apache Arrow timestamp type; List converted to Apache Arrow array; Dictionary converted to structure; Nested data converted to nested structure; Memory layout organization: Organize data by column, allocate memory buffer for each column, use parallel processing to accelerate conversion, reuse memory buffer, and achieve zero-copy operation; Exception handling: Null value handling: mark the null value position, maintain the null value mask, and handle different types of null value representations; Exception value handling: data out-of-bounds check, type mismatch handling, and format error handling; Result output: Generate Apache Arrow data table.
5. According to claim 1, a unified data quality audit method based on Apache Arrow is characterized in that: The processing flow of the data quality control module described in S02 is: Input layer: receives source data in the unified Apache Arrow format Rule chain construction: dynamically assemble audit rule chains based on configuration Rule execution: A single rule is processed and the result is output in Apache Arrow format. The result is used as the input of the next rule for further processing. Parallel processing and distributed expansion are supported. Result output: Generate unified quality reports and processed data.
6. The method for unified data quality audit based on Apache Arrow according to claim 1, characterized in that: The functions implemented by the audit rules described in S02 include: repairing abnormal data; filtering abnormal data; storing abnormal detailed data; and audit verification.
7. The method for unified data quality audit based on Apache Arrow according to claim 1, characterized in that: The steps of converting the data source data processed by the quality audit control module into the data of various data processing components described in S03 are: defining an output adapter base class; implementing adapters of various data processing components; wherein the method of defining the output adapter base class includes: converting the Apache Arrow Schema into a schema in a target format; converting the Apache Arrow data into a target component format.
8. The method for unified data quality audit based on Apache Arrow according to claim 7, characterized in that: The schema for converting the Apache Arrow Schema to the target format is used to reassemble the data from the Apache Arrow format into the original format suitable for the target platform, ensuring that the data can flow smoothly to the next stage. The implementation steps are: Analyze the structure of Apache Arrow Schema, including field names, data types, and field order; Map the data type of each field in the Apache Arrow Schema to the data type of the target component; use the mapped data type to build the Schema of the target component.
9. The method for unified data quality audit based on Apache Arrow according to claim 7, characterized in that: The steps to convert Apache Arrow data to the target component format are: Map Apache Arrow Schema to target component Schema: Map the fields in Apache Arrow Schema to the Schema of the target component format, including mapping of field names, data types, and field order; Data type conversion: According to the mapping relationship, the data type in Apache Arrow is converted into the data type supported by the target format; Data conversion: Convert Apache Arrow data to the target format. For Spark, use the pyApache Arrow library to directly convert Apache ArrowTable to Spark DataFrame; for DataX, convert Apache Arrow data to DataX data format.
10. A unified data quality audit device based on Apache Arrow, characterized in that: The device includes: Data access module: used to access process data of various data processing frameworks, convert all accessed data into Apache Arrow format data using the adapter mode, construct the Arrow Flight gRPC client, and send the Apache Arrow format data to the data quality control module; Data quality control module: used to receive the data in Apache Arrow format converted by the data access module, process the data according to one or more audit rules using the chain of responsibility model, and then output the data in Apache Arrow format, and send the audited data to the data output module; Data output module: used to receive Apache Arrow format data in response from the Arrow Flight gRPC server of the data quality control module, which includes three types of data: data source data processed by the quality audit control module, exception detail data, and audit indicator data. These three types of data are diverted for processing: the data source data processed by the quality audit control module is converted into data of various data processing components; the exception detail data is stored in external storage; and the audit indicator data is stored in the database for analysis of audit quality indicators.
11. An electronic device comprising a memory and a processor, wherein a computer program is stored in the memory, wherein: When the processor executes the program, the method according to any one of claims 1 to 9 is implemented.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 9 is implemented.