Data quality assessment and optimization method, low-code platform and computer device
By automating the multi-dimensional evaluation and optimization operators through a low-code platform, the problems of high difficulty and low efficiency in data governance are solved, a flexible and convenient data governance solution is achieved, and the efficiency of data quality evaluation and optimization is improved.
Patent Information
- Application Number
- CN202511114382.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-11
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-08-11
AI Technical Summary
Existing technologies for data governance are difficult and inefficient, heavily reliant on programming, which makes data quality assessment and optimization complex and time-consuming.
It provides a low-code platform that automates the diagnosis of data quality through multi-dimensional evaluation indicators and dynamically schedules and optimizes operators for data processing, thereby reducing the technical threshold and improving governance efficiency.
It enables flexible and convenient configuration of data governance solutions in a low-code development environment, significantly reducing the difficulty of data governance and improving governance efficiency, as well as enhancing the efficiency of data quality assessment and optimization.
Smart Images

Figure CN120596875B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to data quality assessment and optimization methods, low-code platforms, and computer equipment. Background Technology
[0002] Data governance is crucial for ensuring data quality. However, current data governance solutions heavily rely on programming for data quality assessment and optimization, requiring users to be proficient in multiple programming languages and familiar with complex algorithm libraries, making data governance quite challenging. Furthermore, these solutions necessitate repeatedly writing large amounts of code to meet the processing needs of different datasets, resulting in inefficient data governance.
[0003] There is currently no effective solution to the problem of high difficulty and low efficiency in data governance in related technologies. Summary of the Invention
[0004] This embodiment provides a data quality assessment and optimization method, a low-code platform, and a computer device to address the problems of high difficulty and low efficiency in data governance in related technologies.
[0005] Firstly, this embodiment provides a data quality assessment and optimization method applied to a low-code platform; the method includes:
[0006] Obtain the raw dataset input to the low-code platform and the metadata of the raw dataset;
[0007] Based on multi-dimensional evaluation metrics that match the original dataset, the original dataset is analyzed according to the metadata to obtain the quality evaluation results of the original dataset;
[0008] Identify multiple preset optimization operators in the low-code platform that match the quality assessment results;
[0009] The original dataset is optimized using multiple preset optimization operators that match the quality assessment results to obtain the target dataset.
[0010] In some embodiments, after obtaining the input raw dataset and the metadata of the raw dataset, the method further includes:
[0011] In response to a first user instruction input through the evaluation interface of the low-code platform, the multi-dimensional evaluation metrics of the original dataset selected by the first user instruction are determined; the multi-dimensional evaluation metrics include pixel-level evaluation metrics, semantic-level evaluation metrics, and structural-level evaluation metrics.
[0012] In some embodiments, determining a plurality of preset optimization operators in the low-code platform that match the quality assessment results includes:
[0013] Based on the quality assessment results, determine the optimization strategy for the original dataset;
[0014] Determine the combination of optimization operators in the low-code platform that matches the optimization strategy; the combination of optimization operators includes multiple preset optimization operators.
[0015] In some embodiments, optimizing the original dataset using multiple preset optimization operators that match the quality assessment results to obtain the target dataset includes:
[0016] The preset optimization operators that match the quality assessment results are arranged in a process flow to obtain a corresponding optimization flow; the optimization flow is used to indicate the execution order of each preset optimization operator;
[0017] Based on the optimization process, the original dataset is optimized by calling each of the preset optimization operators to obtain the target dataset.
[0018] In some embodiments, the method further includes:
[0019] Based on the second user command input through the parameter configuration interface of the low-code platform, the algorithm parameters of each preset optimization operator are dynamically adjusted.
[0020] In some embodiments, the method further includes:
[0021] The optimization processing task of the original dataset is divided into multiple optimization task fragments.
[0022] Each of the optimized task shards is allocated to multiple computing nodes of the low-code platform so that the corresponding optimized task shard is executed through each computing node.
[0023] In some of these embodiments, each of the preset optimization operators includes a data repair operator and a generative enhancement operator.
[0024] Secondly, this embodiment provides a low-code platform, including:
[0025] The data access module is used to acquire the raw dataset input to the low-code platform and the metadata of the raw dataset;
[0026] The data evaluation module is used to analyze the original dataset based on the metadata, using multi-dimensional evaluation indicators that match the original dataset, to obtain the quality evaluation result of the original dataset.
[0027] A data optimization module is used to determine multiple preset optimization operators in the low-code platform that match the quality assessment results;
[0028] The data optimization module is further configured to optimize the original dataset using multiple preset optimization operators that match the quality assessment results, thereby obtaining the target dataset.
[0029] Thirdly, this embodiment provides a computer device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the data quality assessment and optimization method described in the first aspect above.
[0030] Fourthly, this embodiment provides a storage medium storing a computer program that, when executed by a processor, implements the data quality assessment and optimization method described in the first aspect above.
[0031] Compared with related technologies, the data quality assessment and optimization method, low-code platform, and computer device provided in this embodiment obtain the original dataset and its metadata input to the low-code platform; analyze the original dataset based on the metadata according to multi-dimensional evaluation indicators matching the original dataset to obtain the quality assessment result of the original dataset; determine multiple preset optimization operators in the low-code platform that match the quality assessment result; and optimize the original dataset using the multiple preset optimization operators that match the quality assessment result to obtain the target dataset. This solves the problem of high data governance difficulty and low efficiency, and enables flexible and convenient configuration of data governance solutions with the help of the low-code development environment, significantly improving governance efficiency while reducing the difficulty of data governance.
[0032] Details of one or more embodiments of this application are set forth in the following drawings and description to make other features, objects and advantages of this application more readily apparent. Attached Figure Description
[0033] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0034] Figure 1 This is a schematic diagram of the structure of a low-code platform provided in an embodiment of this application;
[0035] Figure 2 This is a flowchart of a data quality assessment and optimization method provided in an embodiment of this application;
[0036] Figure 3 This is a flowchart of an embodiment of the optimized operator matching method provided in this application;
[0037] Figure 4 This is a flowchart of a data optimization method provided in an embodiment of this application;
[0038] Figure 5 This is a structural block diagram of a data quality assessment and optimization apparatus provided in an embodiment of this application.
[0039] In the diagram: 10, Low-code platform; 100, Data access module; 200, Data evaluation module; 300, Data optimization module; 400, Distributed computing module; 500, Visual interaction module; 600, Acquisition module; 700, Evaluation module; 800, Optimization module. Detailed Implementation
[0040] To better understand the purpose, technical solution, and advantages of this application, the application is described and illustrated below in conjunction with the accompanying drawings and embodiments.
[0041] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.
[0042] To clearly describe the data quality assessment and optimization method applied to the low-code platform 10 provided in the embodiments of this application, the low-code platform 10 will be described in detail first with reference to the accompanying drawings. Figure 1This is a schematic diagram of the structure of a low-code platform 10 provided in this embodiment, such as... Figure 1 As shown, the low-code platform 10 includes a data access module 100, a data evaluation module 200, and a data optimization module 300.
[0043] The data access module 100 is used to acquire the raw dataset and metadata of the raw dataset from the input low-code platform 10.
[0044] The data evaluation module 200 is used to analyze the original dataset based on metadata and multi-dimensional evaluation indicators that match the original dataset, and obtain the quality evaluation results of the original dataset.
[0045] The data optimization module 300 is used to determine multiple preset optimization operators in the low-code platform 10 that match the quality assessment results;
[0046] The data optimization module 300 is also used to optimize the original dataset using multiple preset optimization operators that match the quality assessment results, so as to obtain the target dataset.
[0047] In this embodiment, the low-code platform 10 is a system platform with a graphical interface (such as drag-and-drop components and a visual process designer) as its core interaction mechanism. Core functions can be developed through configuration-based operations without the need to write complex code. In data governance application scenarios, the low-code platform 10 provides functions such as flexibly selecting data optimization operators and dynamically configuring relevant parameters through a visual operation interface, so as to quickly realize the orchestration and construction of data governance processes, significantly reducing the development threshold and technical difficulty.
[0048] The data access module 100 employs a distributed storage architecture (such as the Hadoop Distributed File System or Google Distributed File System) to ensure strong scalability and stability. It supports petabyte (PB) level image and video data fragmentation and storage. Through data fragmentation technology, it distributes large-scale datasets across multiple storage nodes, effectively improving data storage reliability and read / write performance. The data access module 100 acquires the raw dataset from the low-code platform 10. Through its built-in format parsing engine, it automatically identifies various common data formats (such as JPEG, PNG, BMP, and MP4) and extracts multiple metadata elements from the raw dataset, such as resolution, shooting time, and tag information. This metadata will be used for subsequent data quality assessment and processing, providing crucial foundational information for data governance.
[0049] The Data Evaluation Module 200 is dedicated to achieving multi-dimensional quantitative analysis of datasets such as images and videos. Utilizing a multi-dimensional indicator system (such as pixel-level, semantic-level, and structural-level indicators), covering data repetition rate, brightness uniformity, sharpness, resolution consistency, and label distribution balance, it conducts systematic quantitative analysis of the dataset to accurately identify data defects and potential problems. Through the reasonable setting of evaluation standards and algorithm models, it achieves comprehensive diagnosis and evaluation of data quality. Specifically, it pre-determines multi-dimensional evaluation indicators that match the original dataset. Using the evaluation operators corresponding to these indicators, it analyzes the original dataset based on metadata to obtain the quality evaluation results.
[0050] Furthermore, the data optimization module 300 provides an intelligent optimization operator library that integrates diverse optimization operators. These optimization operators encapsulate data enhancement, repair, and generation operations in data governance into standardized, reusable functional components, while also supporting parameterized dynamic configuration for flexible customization, providing a rich selection of technologies for data optimization. After completing the data quality assessment, the assessment results are correlated with optimization operators of different functions in the intelligent optimization operator library to identify multiple preset optimization operators in the low-code platform 10 that match the quality assessment results. These preset optimization operators then optimize the original dataset to obtain the target dataset, thus forming an assessment-optimization closed loop.
[0051] It's worth noting that in Low-Code Platform 10, the model invocation and integration module adopts a modular design and standardized interfaces, supporting the rapid integration and dynamic replacement of various algorithm models. The orchestration engine of Low-Code Platform 10 allows users to design data governance processes through a visual interface using drag-and-drop functionality, ensuring smooth interaction between the user and the platform. Specifically, the low-code orchestration engine uses a visual process canvas to enable drag-and-drop combination of evaluation modules and optimization operators, generating custom data governance pipelines. It also features conditional branching, version management, process cloning, and sharing capabilities.
[0052] The visual interface supports drag-and-drop operations to connect data quality assessment and optimization operators, enabling the rapid construction of personalized data governance workflows. It also provides a graphical parameter configuration interface, allowing for flexible configuration of operator parameters via sliders and drop-down menus. These parameters include adjusting the super-resolution factor, denoising intensity, and the number of sampling steps in the generation model, achieving precise control over the data processing. Furthermore, the Low-Code Platform 10 includes built-in templates for commonly used data governance workflows. These templates contain operator combinations corresponding to the data governance objectives, such as image denoising and enhancement, and label equalization. Users can also customize templates according to their specific application needs and perform import / export operations, improving the reusability and efficiency of the data governance workflow.
[0053] The low-code platform provided in this application comprises a data access module, a data evaluation module, and a data optimization module. The data access module acquires the original dataset and its metadata input to the low-code platform. The data evaluation module analyzes the original dataset based on multi-dimensional evaluation metrics matching the original dataset and the metadata, obtaining a quality evaluation result. The data optimization module identifies multiple preset optimization operators in the low-code platform that match the quality evaluation result and optimizes the original dataset using these operators to obtain the target dataset. Based on this, the low-code development environment automatically diagnoses dataset quality through multi-dimensional evaluation metrics and dynamically schedules appropriate optimization operators to perform targeted data optimization. This solution transforms the traditional data governance process into a visual orchestration, enabling users to complete complex algorithm configurations, logic definitions, and data mining tasks, lowering the technical threshold while improving data governance efficiency. It solves the problems of high difficulty and low efficiency in data governance, achieving flexible and convenient configuration of data governance solutions using a low-code development environment, significantly improving governance efficiency while reducing the difficulty of data governance.
[0054] In some other embodiments, the low-code platform 10 also includes a distributed computing module 400.
[0055] Specifically, the distributed computing module 400 builds a task scheduling system based on a distributed computing framework (such as the open-source computing framework Apache Spark). This system intelligently shards data optimization tasks and distributes them across multiple computing nodes, allowing each node to execute the corresponding optimization task shard, thus improving data optimization efficiency. Simultaneously, it monitors the processing progress and resource utilization of each computing node in real time, including key indicators such as CPU utilization, GPU utilization, and memory usage. This data is then visualized to help users intuitively understand task execution and promptly identify and resolve potential performance bottlenecks.
[0056] It should be noted that the above task scheduling system supports dynamic expansion of the graphics processor acceleration cluster. It automatically adjusts the allocation of computing resources according to the task's computing requirements and the node's load to ensure efficient task execution. This enables real-time data processing under high concurrency conditions, meets the needs of multiple signal source inputs, and is suitable for scenarios requiring real-time monitoring and rapid response.
[0057] In some other embodiments, the low-code platform 10 also includes a visual interaction module 500.
[0058] Specifically, after the data evaluation module completes the data quality assessment of the original dataset, it generates a corresponding data quality analysis report and visualizes it. The data quality analysis report can be in formats including but not limited to PDF and JPG, and its content includes a list of defective data, indicator radar charts, and comparison charts of typical cases. By visualizing the data evaluation results, users can gain a comprehensive understanding of the dataset's quality status, providing strong support for data governance decisions.
[0059] Furthermore, the data processing module supports real-time preview of single samples. For example, when a user clicks on any image in the dataset, the comparison before and after image restoration is visually displayed, allowing for an intuitive understanding of the data optimization effect. Simultaneously, the visualization interaction module 500 provides interactive operations such as image zooming and rotation for detailed observation.
[0060] The data quality assessment and optimization method for low-code platforms provided in the above embodiments of this application will be explained and described in detail below with reference to the accompanying drawings. Figure 2 This is a flowchart of a data quality assessment and optimization method provided in this embodiment, such as... Figure 2 As shown, the process includes the following steps:
[0061] Step S210: Obtain the original dataset of the input low-code platform and the metadata of the original dataset;
[0062] Specifically, the raw dataset is obtained from the input low-code platform. This raw dataset can be one or more combinations of text, image, and video datasets. Data extraction is then performed on the raw dataset to obtain its metadata. This metadata includes information such as resolution, capture time, and labeling information. This metadata will be used for subsequent data quality assessment and processing, providing crucial foundational information for data governance.
[0063] Step S220: Based on multi-dimensional evaluation metrics that match the original dataset, analyze the original dataset according to metadata to obtain the quality evaluation results of the original dataset;
[0064] Specifically, multi-dimensional evaluation metrics matching the original dataset are determined. These multi-dimensional evaluation metrics include pixel-level, semantic-level, and structural-level evaluation metrics. For example, multiple suitable evaluation metrics are selected based on the data type of the original dataset. In other embodiments, the multi-dimensional evaluation metrics of the original dataset selected by the first user instruction can also be determined based on the first user instruction input through the evaluation interface of the low-code platform.
[0065] Among them, pixel-level evaluation metrics are used to comprehensively evaluate the pixel-level quality of images. Image-related metrics include the uniformity of image brightness, the distribution of image contrast (which can be reflected by the contrast entropy value), and the consistency of image resolution. Video-related metrics include the consistency of motion of target objects in the video and whether the audio and video are synchronized. Text-related metrics include typo detection and grammatical error detection. Semantic-level evaluation aims to provide semantic basis for data quality assessment. By statistically analyzing the frequency of label categories and calculating the Gini coefficient, it judges whether the label distribution is reasonable, so as to deeply analyze the balance of label distribution, or identify the occlusion of targets in the image and statistically analyze the target occlusion rate, or accurately find the missing key information in the dataset by locating missing values. Structural-level evaluation metrics include data duplication rate and data missingness. By accurately calculating the data duplication rate, duplicate samples in the dataset can be effectively identified, while the accurate location of missing values can provide an important reference for the optimization of data structure.
[0066] Furthermore, based on the selected evaluation indicators, the corresponding evaluation operators are invoked to evaluate and analyze the original dataset to obtain the quality evaluation results of the original dataset, and a corresponding data quality analysis report is generated for visualization. The data quality analysis report includes, but is not limited to, a list of defective data, indicator radar charts, and typical case comparison charts. In this way, by visualizing the data evaluation results, users can fully understand the quality status of the dataset and provide strong support for data governance decisions.
[0067] For example, when the original dataset is an image dataset, duplicate data detection, brightness uniformity analysis, and label distribution evaluation are performed on the image dataset. The evaluation steps are described in detail below.
[0068] 1) Duplicate Data Detection Process: Images are pre-normalized to 32×32 pixels, then grayscale processed and subjected to Discrete Cosine Transform (DCT). Low-frequency regions are extracted based on the DCT results, generating an 8×8 feature matrix. The feature matrix is averaged and binarized to obtain a 64-bit perceptual hash value. Finally, a perceptual hash algorithm is used to filter similar images. Similarity retrieval is achieved through Hamming distance calculation, which supports batch vectorization operations, enabling second-level similarity retrieval for millions of images in a cluster environment. The system also provides dynamically configurable similarity thresholds, allowing for the setting of preset threshold templates for different business scenarios such as security images and medical images. The threshold adjustment step can be accurate to 0.1, ensuring a recall rate ≥95% and an accuracy ≥98% for duplicate data identification.
[0069] 2) Brightness Uniformity Analysis Process: Each image is divided into a grid, with dynamically configurable granularity (e.g., 4×4, 8×8, and 16×16). Brightness analysis is performed on the gridded images, based on the L channel values of the Lab color space established by the Commission on Illumination (CIE) (which better reflects human visual characteristics), avoiding errors caused by the non-linear response of the RGB channels. A spatial weighting factor (center weight 1.2, edge weight 0.8) is introduced to optimize the standard deviation calculation, enhancing the perception of differences in visually sensitive areas. Finally, a brightness distribution heatmap is provided through a visualization module to intuitively display the brightness analysis results. Interactive highlighting of grid areas is supported; clicking on any grid displays the corresponding area's average brightness, extreme values, and deviation rate from the global mean, assisting users in locating overexposed or underexposed local image areas.
[0070] 3) Label Distribution Evaluation Process: The label distribution evaluation uses the Gini coefficient calculated based on the Lorenz curve, supporting label weight configurations such as positive and negative sample weighting and business priority labeling. The specific formula for calculating the Gini coefficient G is as follows:
[0071] (1)
[0072] In equation (1), n represents the total number of samples; This represents the weight of the i-th sample. When the Gini coefficient exceeds the set threshold (default is 0.7, configurable range is 0.5 to 0.9), the system automatically triggers the equalization recommendation algorithm, generates the corresponding data augmentation scheme (such as SMOTE oversampling, clustering undersampling), and matches it with the generative augmentation function of the intelligent optimization operator library to optimize the label distribution structure and form an evaluation-optimization closed loop.
[0073] Step S230: Identify multiple preset optimization operators in the low-code platform that match the quality assessment results;
[0074] It should be noted that the low-code platform's intelligent optimization operator library includes different categories of preset optimization operators, including but not limited to data repair operators, generative enhancement operators, and text correction and enhancement operators. It supports graphical configuration of algorithm parameters, and each operator uses standardized encapsulation with unified input / output interfaces. Examples include super-resolution operators, denoising operators, text-to-image repair operators, and image-to-image enhancement operators.
[0075] Specifically, after completing the data quality assessment, based on the quality assessment results of the original dataset, the data to be optimized and its data quality issues are identified in the original dataset. Then, multiple suitable preset optimization operators are selected from the intelligent optimization operator library. In other embodiments, multiple optimization operators selected by the second user command can also be determined based on the current data quality assessment results input through a visual operation interface. That is, suitable optimization operators are selected from the optimization operator library through the visual operation interface provided by the low-code platform, and the process is arranged by dragging, connecting, etc. For example, for low-resolution images, a combination of "super-resolution operator → texturing image enhancement operator" can be added. Simultaneously, users can flexibly set the parameters of each operator in the parameter configuration interface to generate personalized execution plans.
[0076] Step S240: The original dataset is optimized using multiple preset optimization operators that match the quality assessment results to obtain the target dataset.
[0077] Specifically, multiple preset optimization operators matching the quality assessment results are retrieved to optimize the data to be optimized in the original dataset, resulting in the target dataset and forming an evaluation-optimization closed loop. During processing, real-time preview of single samples is supported. For example, the processing progress and intermediate results can be visualized in real time. When a user clicks on any image in the dataset through the system interface, a comparison view before and after image restoration is displayed, allowing for an intuitive understanding of the data optimization effect.
[0078] The low-code platform provides a graphical parameter configuration interface, which allows users to flexibly configure operator parameters through interactive methods such as sliders and drop-down menus. These parameters include adjusting the super-resolution factor, denoising intensity, and the number of sampling steps in the generated model, enabling precise control over the data processing process.
[0079] It should be noted that the low-code platform also supports setting data governance process templates. The data governance process templates contain combinations of operators corresponding to the data governance objectives, such as image noise reduction and enhancement, and label equalization. At the same time, users can customize the templates according to their actual application needs and perform import and export operations, which helps to improve the reusability and work efficiency of the data governance process.
[0080] Furthermore, after the optimization process is completed, the target dataset is exported and a secondary quality assessment is performed to verify the effectiveness of the initial optimization strategy. Based on the assessment results, the data governance process is dynamically adjusted to improve data quality. For example, if the data quality indicators do not meet the preset requirements, the root cause is traced back to the specific step (such as improper interpolation algorithm selection or incorrect super-resolution operator parameter configuration) for root cause analysis, and then the data governance process is adjusted (such as changing the optimization operator). Conversely, if the secondary assessment results meet the preset requirements, the process can be solidified as a governance template.
[0081] Data governance is crucial for ensuring data quality. However, current data governance solutions heavily rely on programming for data quality assessment and optimization, requiring users to be proficient in multiple programming languages and familiar with complex algorithm libraries, making data governance quite challenging. Furthermore, these solutions necessitate repeatedly writing large amounts of code to meet the processing needs of different datasets, resulting in inefficient data governance.
[0082] Compared to existing technologies, this application acquires the original dataset and its metadata from the input low-code platform; analyzes the original dataset based on the metadata using multi-dimensional evaluation metrics that match the original dataset, obtaining a quality assessment result; identifies multiple preset optimization operators in the low-code platform that match the quality assessment result; and optimizes the original dataset using these preset optimization operators to obtain the target dataset. Based on this, it utilizes a low-code development environment to automatically diagnose dataset quality using multi-dimensional evaluation metrics and dynamically schedules appropriate optimization operators to perform targeted data optimization. This solution transforms the traditional data governance process into a visual orchestration, lowering the technical barrier while improving efficiency. It solves the problems of high difficulty and low efficiency in data governance, enabling flexible and convenient configuration of data governance solutions using a low-code development environment, reducing the difficulty of data governance, improving governance efficiency, and effectively shortening the governance cycle.
[0083] In some embodiments, after obtaining the input raw dataset and the metadata of the raw dataset, the following steps are also included:
[0084] In response to a first user instruction input through the evaluation interface of a low-code platform, determine the multi-dimensional evaluation metrics for the original dataset selected by the first user instruction; the multi-dimensional evaluation metrics include pixel-level evaluation metrics, semantic-level evaluation metrics, and structural-level evaluation metrics.
[0085] Specifically, the low-code platform provides an evaluation interface where users can flexibly select the required evaluation dimensions, such as resolution detection and label distribution analysis. Based on the first user command input through the evaluation interface, multi-dimensional evaluation metrics for the original dataset selected by the first user command are determined. These multi-dimensional evaluation metrics mainly include pixel-level evaluation metrics, semantic-level evaluation metrics, and structural-level evaluation metrics, among other metric types.
[0086] Among them, pixel-level evaluation metrics are used to comprehensively evaluate the pixel-level quality of images. Image-related metrics include the uniformity of image brightness, the distribution of image contrast, and the consistency of image resolution. Video-related metrics include the consistency of motion of target objects in the video and whether the audio and video are synchronized. Text-related metrics include typo detection and grammatical error detection. Semantic-level evaluation aims to provide semantic basis for data quality assessment. By statistically analyzing the frequency of label categories and calculating the Gini coefficient, it judges whether the label distribution is reasonable, so as to deeply analyze the balance of label distribution, or identify the occlusion of targets in the image and statistically analyze the target occlusion rate, or accurately find the missing key information in the dataset by locating missing values. Structural-level evaluation metrics include data duplication rate and data missingness. By accurately calculating the data duplication rate, duplicate samples in the dataset can be effectively identified, and the accurate location of missing values can provide an important reference for the optimization of data structure.
[0087] In this embodiment, in response to a first user instruction input through the evaluation interface of a low-code platform, a multi-dimensional evaluation metric for the original dataset selected by the first user instruction is determined. The multi-dimensional evaluation metric includes pixel-level evaluation metrics, semantic-level evaluation metrics, and structural-level evaluation metrics. This enables the multi-dimensional data evaluation system to systematically ensure data quality, providing a high-confidence data source for downstream model training, thereby improving the model's generalization ability and prediction accuracy.
[0088] In some of these embodiments, such as Figure 3 As shown, step S230, which involves determining multiple preset optimization operators in the low-code platform that match the quality assessment results, includes the following steps:
[0089] Step S231: Determine the optimization strategy for the original dataset based on the quality assessment results;
[0090] Step S232: Determine the combination of optimization operators in the low-code platform that matches the optimization strategy; the combination of optimization operators includes multiple preset optimization operators.
[0091] Specifically, after completing the data quality assessment, based on the assessment results of the original dataset, the data to be optimized and its data quality problems are identified, and then the optimization strategy for the original dataset is determined. For example, if the assessment results indicate that the original dataset has a large number of missing values, the optimization strategy includes finding and filling in the missing values for key fields by tracing the original records, and applying interpolation algorithms to fill in numerical fields; if the assessment results indicate that the image resolution in the original dataset does not meet the preset requirements, the optimization strategy includes increasing the image resolution to ensure image quality.
[0092] Furthermore, based on the optimization strategy of the original dataset, intelligent algorithms determine the matching combination of optimization operators in the low-code platform. The combination of optimization operators typically includes multiple preset optimization operators. For example, when the optimization strategy indicates that image resolution needs to be increased, a combination of super-resolution operators and generative enhancement operators is automatically triggered.
[0093] In this embodiment, based on the quality assessment results, an optimization strategy for the original dataset is determined, and a combination of optimization operators matching the optimization strategy is identified in the low-code platform. The combination of optimization operators includes multiple preset optimization operators, thereby providing differentiated solutions for different types of data defects, such as low resolution, noise pollution, and label skew, forming an evaluation-optimization closed loop. This helps to achieve accurate data optimization, providing a higher quality data foundation for downstream machine learning model training, thereby improving the training effect and performance of the model.
[0094] In some of these embodiments, such as Figure 4 As shown, step S240 optimizes the original dataset using multiple preset optimization operators that match the quality assessment results to obtain the target dataset, including the following steps:
[0095] Step S241: Arrange the preset optimization operators that match the quality assessment results into a process flow to obtain the corresponding optimization flow; the optimization flow is used to indicate the execution order of each preset optimization operator.
[0096] Step S242: Based on the optimization process, call each preset optimization operator to optimize the original dataset to obtain the target dataset.
[0097] Specifically, the adapted preset optimization operators are orchestrated to obtain a corresponding optimization process, which indicates the execution order of each preset optimization operator. Based on the optimization process, each preset optimization operator is invoked to optimize the data to be optimized, resulting in the target dataset.
[0098] For example, for low-resolution images, a combination of "super-resolution operator → text-based image enhancement operator" can be added. The super-resolution operator is used to upgrade the image resolution to the target scale and reconstruct the basic texture structure, while the text-based image enhancement operator repairs blurred areas of the image based on the input text description. This supports fine-grained repair at the pixel level to deep enhancement at the semantic level, achieving multi-level data optimization and comprehensively ensuring data quality.
[0099] In this embodiment, the preset optimization operators that match the quality assessment results are arranged into a process to obtain a corresponding optimization process. The optimization process is used to indicate the execution order of each preset optimization operator. Then, based on the optimization process, each preset optimization operator is called to optimize the original dataset to obtain the target dataset, thereby achieving accurate data optimization and improving data quality.
[0100] In some embodiments, the above-described data quality assessment and optimization method further includes the following steps:
[0101] Based on second-user commands input through the parameter configuration interface of the low-code platform, the algorithm parameters of each preset optimization operator are dynamically adjusted.
[0102] Specifically, the low-code platform provides a graphical parameter configuration interface, which dynamically adjusts the algorithm parameters of each preset optimization operator based on second-user commands input through the interface. Algorithm parameters, such as super-resolution factor, denoising intensity, and number of sampling steps in the generated model, can be flexibly set via interactive methods such as sliders and drop-down menus.
[0103] It's worth noting that, to ensure real-time responsiveness and interpretability of parameter adjustments, the platform incorporates a parameter linkage engine. When a user modifies algorithm parameters, the system automatically calculates and displays the affected related parameters. Furthermore, all adjustments trigger an instant preview function. For example, a before-and-after comparison example is generated in the sidebar, ensuring users can intuitively perceive the impact of the parameters.
[0104] In this embodiment, based on the second user command input through the parameter configuration interface of the low-code platform, the algorithm parameters of each preset optimization operator are dynamically adjusted, thereby achieving precise control over the data processing process and helping to improve the data optimization effect.
[0105] In some embodiments, the above-described data quality assessment and optimization method further includes the following steps:
[0106] The optimization task of the original dataset is split into multiple optimization task splits.
[0107] Each optimization task shard is distributed to multiple compute nodes on the low-code platform so that the corresponding optimization task shard can be executed on each compute node.
[0108] Specifically, on the low-code platform, a task scheduling system is built based on a distributed computing framework. This system can intelligently shard data optimization tasks and distribute them to multiple computing nodes, so that each computing node can execute the corresponding optimization task shard, thereby improving data optimization efficiency. Simultaneously, it monitors the processing progress and resource utilization of each computing node in real time, including key indicators such as CPU utilization, GPU utilization, and memory usage, and visualizes the processing progress and resource utilization to help users intuitively understand the task execution status, enabling them to promptly identify and resolve potential performance bottlenecks.
[0109] It should be noted that the above task scheduling system supports dynamic expansion of the graphics processor acceleration cluster. It automatically adjusts the allocation of computing resources according to the task's computing requirements and the node's load, thereby ensuring efficient task execution.
[0110] In this embodiment, the optimization processing task of the original dataset is sharded to obtain multiple optimization task shards, and each optimization task shard is allocated to multiple computing nodes of the low-code platform so that the corresponding optimization task shard can be executed by each computing node, thereby achieving efficient execution of optimization tasks and improving data governance efficiency.
[0111] In some of these embodiments, the preset optimization operators include data repair operators and generative enhancement operators.
[0112] Specifically, the preset optimization operators include data restoration operators and generative enhancement operators. Examples include super-resolution operators, denoising operators, text-to-image restoration operators, and image-to-image enhancement operators. The following examples illustrate the different optimization operators.
[0113] 1) Super-resolution operators: Supports bicubic interpolation and the Enhanced Deep Residual Network (EDSR) model. Users can configure magnification factors such as 2x, 4x, or 8x as needed. The bicubic interpolation algorithm improves image resolution by using polynomial fitting while maintaining image smoothness; the EDSR model relies on the deep residual network to learn image texture priors, enabling the generation of higher-quality super-resolution images and effectively improving image details.
[0114] 2) Denoising Operator: Image noise is detected through the collaboration of Local Binary Pattern (LBP) and Gaussian Mixture Model, accurately identifying salt-and-pepper noise and Gaussian noise and quantifying noise density (0~100%). Median filtering with adaptive window (dynamically adjustable within the range of 3×3 to 11×11) is used, or an optimized 3D block matching strategy (BM3D) is adopted, which introduces spatial location constraints and color histogram similarity measure. This ensures that when the noise density is less than or equal to 30%, the Structural Similarity Index Measure (SSIM) is greater than or equal to 0.92, so as to better preserve the structure of image edges, textures and other structures, while improving the processing speed by 40% compared with the traditional method.
[0115] 3) Text-generated image restoration operator: Based on the text-generated image model (such as the Stable Diffusion Model), text-driven local generation is realized. When the user inputs a text description, such as "repair the blurry cat face", the generation model can generate a high-definition completed image, thereby automatically filling in the missing parts of the image according to the user's description, effectively improving the clarity and integrity of the image.
[0116] 4) Graph-to-Graph Enhancement Operator: By using the ControlNet model, the image is enhanced while preserving the structural features of the input image, in order to generate a higher resolution, richer colors, and clearer details, thus meeting the user's demand for improved image quality.
[0117] This embodiment organically combines visual algorithms with large model generation technology to achieve a rapid improvement in data quality.
[0118] The following description and illustration of this embodiment uses medical image dataset governance as an example, specifically including the following steps:
[0119] S1. Data Acquisition and Input.
[0120] Five thousand medical images in Digital Imaging and Communications in Medicine (DICOM) format were uploaded to a low-code platform via File Transfer Protocol (FTP). Upon receiving the medical image dataset, a format parsing engine converted each image from DICOM to PNG format and stored it in a distributed file storage cluster (Hadoop HDFS). During storage, metadata such as image resolution and capture time was extracted to prepare for subsequent data processing.
[0121] On the system's evaluation interface, users can select multiple evaluation dimensions such as resolution detection, noise level analysis, and label integrity check based on their input instructions, and then click the "Execute Evaluation" button to start a comprehensive quality evaluation of the dataset.
[0122] S2. Analyze the data quality assessment results.
[0123] The output data quality assessment results include: 1200 images in the dataset have a resolution of 256×256, significantly lower than the clinically required resolution of 512×512; low-resolution images will affect diagnostic accuracy; 300 images in the dataset have a noise standard deviation greater than the preset value, indicating significant salt-and-pepper noise, which will interfere with the effective information in the images, reducing image quality and diagnostic accuracy; and 50 images in the dataset have missing lesion region annotations, with a label integrity of 99%. Missing labels will affect data usability and model training performance.
[0124] S3, Optimize data arrangement process.
[0125] In the low-code platform's visual interface, add an input node to import the medical image data to be processed into the workflow. Then, connect the resolution detection operator to perform resolution detection on each medical image. Next, set a conditional branch: when an image resolution less than 512×512 is detected, connect the EDSR super-resolution operator (setting the magnification factor to 2x) and the denoising operator (selecting the BM3D algorithm and setting the noise intensity to 15) in sequence; if the image resolution meets the resolution standard, output it directly.
[0126] For images processed by super-resolution and denoising operators, a text enhancement operator is connected, and a prompt word (such as "enhance medical image contrast and preserve lesion details") is input to further improve image quality.
[0127] It should be noted that specific parameters can be set for each operator in the parameter configuration interface. For example, the EDSR model uses pre-trained medical image weights to better adapt to the characteristics of medical images; the block size of the BM3D algorithm is set to 8×8 to ensure that image details are preserved while removing noise; and the sampling step number of the generation model is set to 30 to balance image generation quality and processing efficiency.
[0128] S4, Distributed Computing and Visualization.
[0129] The Apache Spark cluster automatically starts the task scheduling system, dividing the optimization task of 5,000 medical images into chunks and distributing them across 10 computing nodes for parallel processing. The graphics processing unit (GPU) nodes are responsible for super-resolution and generative model calculations, fully leveraging the powerful computing capabilities of the GPU to improve processing speed; the central processing unit (CPU) nodes handle tasks such as format conversion, ensuring the efficient operation of the entire processing flow.
[0130] During processing, a visual interface displays the processing progress in real time, helping users quickly understand the data processing status and showing detailed information such as the average processing time for each image. Users can also click on any image to preview the processing results in real time. Additionally, the visual interface displays the resource utilization of each node, facilitating task monitoring.
[0131] S5. Result Output and Iteration.
[0132] After optimizing all images, the resulting high-quality medical image dataset is exported for user convenience, such as applications in medical image analysis and machine learning model training. Before exporting, the system performs a secondary verification of the optimized dataset to ensure data integrity and accuracy. Users can also evaluate the optimized dataset on the platform, comparing the evaluation results with the quality metrics of the original dataset to intuitively understand the improvement in data quality. If the optimization effect does not meet expectations, the data governance process can be adjusted and optimized based on the evaluation results, and the data processing task can be re-executed to achieve continuous improvement and iteration of data quality.
[0133] This embodiment also provides a data quality assessment and optimization apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below refer to combinations of software and / or hardware that perform a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0134] Figure 5 This is a structural block diagram of the data quality assessment and optimization device in this embodiment, as shown below. Figure 5 As shown, the device includes:
[0135] Module 600 is used to obtain the raw dataset of the input low-code platform and the metadata of the raw dataset;
[0136] Evaluation module 700 analyzes the original dataset based on metadata, using multi-dimensional evaluation metrics that match the original dataset, to obtain the quality evaluation results of the original dataset;
[0137] Optimization module 800 identifies multiple preset optimization operators in the low-code platform that match the quality assessment results;
[0138] The optimization module 800 optimizes the original dataset using multiple preset optimization operators that match the quality assessment results, thereby obtaining the target dataset.
[0139] The apparatus provided in this embodiment acquires the original dataset and its metadata from the input low-code platform; analyzes the original dataset based on the metadata using multi-dimensional evaluation metrics that match the original dataset to obtain a quality assessment result; determines multiple preset optimization operators in the low-code platform that match the quality assessment result; and optimizes the original dataset using these preset optimization operators to obtain the target dataset. This solves the problems of high difficulty and low efficiency in data governance, enabling flexible and convenient configuration of data governance solutions using a low-code development environment, significantly improving governance efficiency while reducing the difficulty of data governance.
[0140] In some of these embodiments, in Figure 5 Based on this, the device also includes a selection module for determining the multi-dimensional evaluation metrics of the original dataset selected by the first user instruction in response to a first user instruction input through the evaluation interface of the low-code platform; the multi-dimensional evaluation metrics include pixel-level evaluation metrics, semantic-level evaluation metrics, and structural-level evaluation metrics.
[0141] In some embodiments, the optimization module 800 is further configured to determine an optimization strategy for the original dataset based on the quality assessment results; determine a combination of optimization operators in the low-code platform that matches the optimization strategy; the combination of optimization operators includes multiple preset optimization operators.
[0142] In some embodiments, the optimization module 800 is further configured to orchestrate the preset optimization operators that match the quality assessment results to obtain a corresponding optimization process; the optimization process is used to indicate the execution order of each preset optimization operator; based on the optimization process, each preset optimization operator is called to optimize the original dataset to obtain the target dataset.
[0143] In some embodiments, the optimization module 800 is also used to dynamically adjust the algorithm parameters of each preset optimization operator based on a second user instruction input through the parameter configuration interface of the low-code platform.
[0144] In some of these embodiments, in Figure 5 Based on this, the device also includes a scheduling module, which is used to divide the optimization processing task of the original dataset into multiple optimization task fragments; and to allocate each optimization task fragment to multiple computing nodes of the low-code platform so that the corresponding optimization task fragment can be executed through each computing node.
[0145] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.
[0146] This embodiment also provides a computer device, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the steps in any of the above method embodiments.
[0147] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0148] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0149] S1, obtain the raw dataset of the input low-code platform and the metadata of the raw dataset;
[0150] S2, based on multi-dimensional evaluation metrics that match the original dataset, analyzes the original dataset according to metadata to obtain the quality evaluation results of the original dataset;
[0151] S3 identifies multiple pre-defined optimization operators in the low-code platform that match the quality assessment results;
[0152] S4 optimizes the original dataset using multiple preset optimization operators that match the quality assessment results, thus obtaining the target dataset.
[0153] It should be noted that the specific examples in this embodiment can refer to the examples described in the above embodiments and optional implementations, and will not be repeated in this embodiment.
[0154] Furthermore, in conjunction with the data quality assessment and optimization methods provided in the above embodiments, this embodiment can also provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the data quality assessment and optimization methods described in the above embodiments.
[0155] It should be understood that the specific embodiments described herein are merely illustrative of the application and not intended to limit it. All other embodiments derived by those skilled in the art based on the embodiments provided in this application without inventive effort are within the scope of protection of this application.
[0156] Obviously, the accompanying drawings are merely some examples or embodiments of this application. Those skilled in the art can apply this application to other similar situations based on these drawings without any creative effort. Furthermore, it is understood that although the work done in this development process may be complex and lengthy, for those skilled in the art, certain design, manufacturing, or production modifications made based on the technical content disclosed in this application are merely conventional technical means and should not be considered as insufficient disclosure of this application.
[0157] The term "embodiment" in this application refers to a specific feature, structure, or characteristic described in connection with an embodiment that may be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily imply the same embodiment, nor does it imply that it is mutually exclusive with or independent of other embodiments. It will be clearly or implicitly understood by those skilled in the art that the embodiments described in this application may be combined with other embodiments without conflict.
[0158] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the appended claims.
Claims
1. A method of data quality assessment and optimization, characterized by, The method is applied to a low-code platform, and comprises the following steps: An original data set input into the low-code platform and metadata of the original data set are acquired; In response to a first user instruction input through an evaluation interface of the low-code platform, a multi-dimensional evaluation index of the original data set selected by the first user instruction is determined; the multi-dimensional evaluation index comprises a pixel-level evaluation index, a semantic-level evaluation index, and a structure-level evaluation index; Based on the multi-dimensional evaluation index matched with the original data set, the original data set is analyzed according to the metadata to obtain a quality evaluation result of the original data set; A plurality of preset optimization operators matched with the quality evaluation result in the low-code platform are determined; The original data set is optimized by using the plurality of preset optimization operators matched with the quality evaluation result to obtain a target data set; The optimization of the original data set by using the plurality of preset optimization operators matched with the quality evaluation result to obtain the target data set comprises the following steps: based on an orchestration engine of the low-code platform, the data evaluation module of the low-code platform and each preset optimization operator are combined through a visual operation interface to obtain a corresponding optimization process; the optimization process is used to indicate an execution order of each preset optimization operator; based on the optimization process, each preset optimization operator is called to optimize the original data set to obtain the target data set.
2. The data quality assessment and optimization method of claim 1, wherein, The determination of the plurality of preset optimization operators matched with the quality evaluation result in the low-code platform comprises the following steps: Based on the quality evaluation result, an optimization strategy of the original data set is determined; A combination of optimization operators matched with the optimization strategy in the low-code platform is determined; the combination of optimization operators comprises a plurality of preset optimization operators.
3. The data quality assessment and optimization method of claim 1, wherein, The method further comprises the following steps: Based on a second user instruction input through a parameter configuration interface of the low-code platform, algorithm parameters of each preset optimization operator are dynamically adjusted.
4. The data quality assessment and optimization method of claim 1, wherein, The method further comprises the following steps: The optimization processing task of the original data set is divided into a plurality of optimization task shards; Each optimization task shard is distributed to a plurality of computing nodes of the low-code platform, so that the corresponding optimization task shard is executed by each computing node.
5. The data quality assessment and optimization method of any one of claims 1 to 4, wherein, Each preset optimization operator comprises a data repair operator and a generative enhancement operator.
6. A low-code platform characterized in that, The method comprises the following steps: A data access module is configured to acquire an original data set input into the low-code platform and metadata of the original data set; The data access module is further configured to determine, in response to a first user instruction input through an evaluation interface of the low-code platform, a multi-dimensional evaluation index of the original data set selected by the first user instruction; the multi-dimensional evaluation index comprises a pixel-level evaluation index, a semantic-level evaluation index, and a structure-level evaluation index; A data evaluation module is configured to analyze, based on a multi-dimensional evaluation index matched with the original data set, the original data set according to the metadata to obtain a quality evaluation result of the original data set; a data optimization module configured to determine a plurality of preset optimization operators in the low-code platform that match the quality evaluation result; the data optimization module is further configured to perform optimization processing on the original data set by using the plurality of preset optimization operators that match the quality evaluation result, to obtain a target data set; the data optimization module is further configured to, based on an orchestration engine of the low-code platform, combine the data evaluation module of the low-code platform and each preset optimization operator through a visual operation interface to obtain a corresponding optimization process; the optimization process is configured to indicate an execution order of each preset optimization operator; based on the optimization process, each preset optimization operator is called to perform optimization processing on the original data set, to obtain the target data set. 7.A computer device, comprising a memory and a processor, and characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to execute the steps of the data quality evaluation and optimization method in any one of claims 1 to 5.
8. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to implement the steps of the data quality evaluation and optimization method in any one of claims 1 to 5.
Citation Information
Patent Citations
Data management method and device
CN113722302A
Multi-modal large model construction method based on low codes
CN119148997A