Data quality evaluation and optimization method, low-code platform and computer equipment

Through the multi-dimensional evaluation and optimization operator automation processing of the low-code platform, the problem of difficult and inefficient data governance has been solved, a flexible and convenient data governance solution has been implemented, and data governance efficiency has been improved.

CN120596875AActive Publication Date: 2025-09-05E SURFING VISION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202511114382.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-11
Publication Date
2025-09-05
Estimated Expiration
2045-08-11

AI Technical Summary

Technical Problem

Existing data governance solutions rely on programming implementation, which makes data quality assessment and optimization difficult and inefficient. It requires proficiency in multiple programming languages ​​and repetitive code writing.

Method used

Provides a low-code platform that automatically diagnoses data quality through multi-dimensional evaluation indicators and dynamically schedules optimization operators for data processing, lowering technical barriers and improving governance efficiency.

Benefits of technology

It enables flexible and convenient configuration of data governance solutions in a low-code development environment, reduces the difficulty of data governance and significantly improves governance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120596875A_ABST
    Figure CN120596875A_ABST
Patent Text Reader

Abstract

The invention relates to a data quality evaluation and optimization method, a low-code platform and computer equipment, and the data quality evaluation and optimization method comprises the steps: obtaining an original data set inputted into the low-code platform and metadata of the original data set; based on the multi-dimensional evaluation index matched with the original data set, analyzing the original data set according to the metadata to obtain a quality evaluation result of the original data set; determining a plurality of preset optimization operators matched with the quality evaluation result in the low-code platform; and performing optimization processing on the original data set through a plurality of preset optimization operators matched with the quality evaluation result to obtain a target data set. Through the data processing method and device, the problems of high data processing difficulty and low efficiency are solved, the data processing scheme is flexibly and conveniently configured by means of a low-code development environment, and the processing efficiency is remarkably improved while the data processing difficulty is reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, and in particular to data quality assessment and optimization methods, low-code platforms and computer equipment. Background Art

[0002] Data governance is key to ensuring data quality. However, current data governance solutions rely heavily on programming for data quality assessment and optimization, requiring users to master multiple programming languages ​​and familiarize themselves with complex algorithm libraries, making data governance difficult. Furthermore, these solutions require repetitive coding to meet the processing requirements of different data sets, resulting in inefficient data governance.

[0003] There is currently no effective solution to the problem of difficulty and low efficiency in data governance in related technologies. Summary of the Invention

[0004] In this embodiment, a data quality assessment and optimization method, a low-code platform, and a computer device are provided to solve the problem of difficult and inefficient data governance in related technologies.

[0005] First, in this embodiment, a data quality assessment and optimization method is provided, which is applied to a low-code platform; the method includes:

[0006] Obtaining an original data set input into the low-code platform and metadata of the original data set;

[0007] Analyzing the original dataset according to the metadata based on a multi-dimensional evaluation index that matches the original dataset to obtain a quality evaluation result of the original dataset;

[0008] Determining a plurality of preset optimization operators in the low-code platform that match the quality assessment result;

[0009] The original data set is optimized by using a plurality of preset optimization operators that match the quality assessment result to obtain a target data set.

[0010] In some embodiments, after obtaining the input original data set and metadata of the original data set, the method further includes:

[0011] In response to a first user instruction input through the evaluation interface of the low-code platform, the multidimensional evaluation indicators of the original data set selected by the first user instruction are determined; the multidimensional evaluation indicators include pixel-level evaluation indicators, semantic-level evaluation indicators and structural-level evaluation indicators.

[0012] In some embodiments, determining a plurality of preset optimization operators in the low-code platform that match the quality assessment result includes:

[0013] Determining an optimization strategy for the original data set based on the quality assessment result;

[0014] Determine an optimization operator combination in the low-code platform that matches the optimization strategy; the optimization operator combination includes multiple preset optimization operators.

[0015] In some embodiments, the optimizing the original data set by using a plurality of the preset optimization operators that match the quality assessment result to obtain the target data set includes:

[0016] Performing process arrangement on each of the preset optimization operators that matches the quality assessment result to obtain a corresponding optimization process; the optimization process is used to indicate the execution order of each of the preset optimization operators;

[0017] Based on the optimization process, each of the preset optimization operators is called to optimize the original data set to obtain the target data set.

[0018] In some embodiments, the method further comprises:

[0019] Based on a second user instruction input through the parameter configuration interface of the low-code platform, the algorithm parameters of each of the preset optimization operators are dynamically adjusted.

[0020] In some embodiments, the method further comprises:

[0021] Slicing the optimization processing task of the original data set to obtain multiple optimization task slices;

[0022] Each of the optimization task slices is distributed to multiple computing nodes of the low-code platform so that the corresponding optimization task slice is executed by each of the computing nodes.

[0023] In some embodiments, each of the preset optimization operators includes a data repair operator and a generative enhancement operator.

[0024] Secondly, in this embodiment, a low-code platform is provided, including:

[0025] A data access module, used to obtain the original data set input into the low-code platform and the metadata of the original data set;

[0026] A data evaluation module is configured to analyze the original data set according to the metadata based on a multi-dimensional evaluation index that matches the original data set, and obtain a quality evaluation result of the original data set;

[0027] A data optimization module, configured to determine a plurality of preset optimization operators in the low-code platform that match the quality assessment result;

[0028] The data optimization module is further configured to optimize the original data set by using a plurality of preset optimization operators that match the quality assessment result to obtain a target data set.

[0029] In a third aspect, a computer device is provided in this embodiment, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the data quality assessment and optimization method described in the first aspect is implemented.

[0030] In a fourth aspect, a storage medium is provided in this embodiment, on which a computer program is stored. When the program is executed by a processor, the data quality assessment and optimization method described in the first aspect is implemented.

[0031] Compared with related technologies, the data quality assessment and optimization method, low-code platform and computer equipment provided in this embodiment obtain the original data set input into the low-code platform and the metadata of the original data set; based on the multi-dimensional evaluation indicators matching the original data set, the original data set is analyzed according to the metadata to obtain the quality assessment result of the original data set; multiple preset optimization operators matching the quality assessment results in the low-code platform are determined; the original data set is optimized through multiple preset optimization operators matching the quality assessment results to obtain the target data set, which solves the problem of difficult and inefficient data governance, and realizes the flexible and convenient configuration of data governance solutions with the help of a low-code development environment, significantly improving governance efficiency while reducing the difficulty of data governance.

[0032] The details of one or more embodiments of the present application are set forth in the following drawings and description to make other features, objects, and advantages of the present application more readily apparent. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0034] Figure 1 This is a schematic diagram of the structure of the low-code platform provided in one embodiment of the present application;

[0035] Figure 2 This is a flow chart of a data quality assessment and optimization method provided by an embodiment of the present application;

[0036] Figure 3 This is a flowchart of an optimization operator matching method provided in one embodiment of the present application;

[0037] Figure 4 This is a flow chart of a data optimization method provided by an embodiment of the present application;

[0038] Figure 5 This is a structural block diagram of a data quality assessment and optimization device provided in one embodiment of the present application.

[0039] In the figure: 10, low-code platform; 100, data access module; 200, data evaluation module; 300, data optimization module; 400, distributed computing module; 500, visual interaction module; 600, acquisition module; 700, evaluation module; 800, optimization module. DETAILED DESCRIPTION

[0040] In order to more clearly understand the purpose, technical solutions and advantages of the present application, the present application is described and illustrated below in conjunction with the accompanying drawings and embodiments.

[0041] Unless otherwise defined, technical or scientific terms used in this application shall have the ordinary meanings as understood by persons of ordinary skill in the art to which this application belongs. The terms "a," "an," "the," "these," and similar expressions in this application do not denote limitations on quantity and may be singular or plural. The terms "comprise," "include," "have," and any variations thereof, as used in this application, are intended to cover non-exclusive inclusions. For example, a process, method, system, product, or device comprising a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include unlisted steps or modules (units) or other steps or modules (units) inherent to the process, method, product, or device. The terms "connected," "connected," "coupled," and similar expressions used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. As used in this application, "plurality" means two or more. "And / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone; A and B exist simultaneously; or B exists alone. Generally, the character " / " indicates that the objects in the preceding and following relationship are in an "or" relationship. The terms "first", "second", "third", etc. involved in this application are only used to distinguish similar objects and do not represent a specific ordering of the objects.

[0042] In order to clearly describe the data quality assessment and optimization method applied to the low-code platform 10 provided in the embodiment of the present application, the low-code platform 10 is first described in detail with reference to the accompanying drawings. Figure 11 is a schematic diagram of the structure of a low-code platform 10 provided in this embodiment, as shown in FIG. Figure 1 As shown, the low-code platform 10 includes a data access module 100, a data evaluation module 200 and a data optimization module 300.

[0043] The data access module 100 is used to obtain the original data set input into the low-code platform 10 and the metadata of the original data set;

[0044] A data evaluation module 200 is configured to analyze the original data set according to the metadata based on multi-dimensional evaluation indicators that match the original data set to obtain a quality evaluation result of the original data set;

[0045] The data optimization module 300 is used to determine a plurality of preset optimization operators in the low-code platform 10 that match the quality assessment results;

[0046] The data optimization module 300 is further configured to optimize the original data set by using a plurality of preset optimization operators that match the quality assessment results to obtain a target data set.

[0047] In this embodiment, the low-code platform 10 is a system platform with a graphical interface (such as drag-and-drop components and a visual process designer) as its core interaction platform. Core function development can be completed through configuration operations, without the need for writing complex code. In the application scenario of data governance, the low-code platform 10 provides functions such as flexible selection of data optimization operators and dynamic configuration of relevant parameters through a visual operation interface, which can quickly realize the orchestration and construction of data governance processes, significantly reducing the development threshold and technical difficulty.

[0048] The data access module 100 uses a distributed storage architecture (such as the Hadoop Distributed File System and the Google Distributed File System) to ensure strong scalability and stability, and can support sharded storage of petabyte-scale image and video data. Through data sharding technology, it disperses and stores large-scale data sets across multiple storage nodes, effectively improving the reliability and read and write performance of data storage. The data access module 100 is used to obtain the original data set input into the low-code platform 10. Through the built-in format parsing engine, it automatically identifies multiple common data formats (such as JPEG, PNG, BMP, MP4) and extracts multiple metadata of the original data set, such as resolution, shooting time, and tag information. The metadata of the original data set will be used for subsequent data quality assessment and processing, providing important basic information for data governance.

[0049] The data evaluation module 200 is dedicated to achieving multi-dimensional quantitative analysis of datasets such as images and videos. Leveraging a multi-dimensional indicator system (such as pixel-level, semantic-level, and structural-level evaluation indicators), covering data repetition rate, brightness uniformity, clarity, resolution consistency, and label distribution balance, it conducts systematic quantitative analysis of the dataset to accurately identify data defects and potential problems. By rationally setting evaluation standards and algorithmic models, it achieves comprehensive diagnosis and assessment of data quality. Specifically, it predetermines multi-dimensional evaluation indicators that match the original dataset. Using the evaluation operators corresponding to the multi-dimensional evaluation indicators, the original dataset is analyzed based on metadata to obtain a quality assessment result for the original dataset.

[0050] Furthermore, the intelligent optimization operator library provided by the data optimization module 300 integrates a variety of optimization operators. The optimization operator encapsulates operations such as data enhancement, repair, and generation in data governance into standardized, reusable functional components, while supporting parameterized dynamic configuration for flexible customization, providing a rich range of technical options for data optimization. After completing the data quality assessment, the quality assessment results are associated with the optimization operators of different functions in the intelligent optimization operator library to determine multiple preset optimization operators in the low-code platform 10 that match the quality assessment results. The original data set is optimized through multiple preset optimization operators that match the quality assessment results to obtain the target data set, thereby forming an evaluation-optimization closed loop.

[0051] It should be noted that in the low-code platform 10, the model call and integration module adopts a modular design and standardized interface to support the rapid integration and dynamic replacement of multiple algorithm models. The orchestration engine of the low-code platform 10 supports users to design data governance processes in a drag-and-drop manner through a visual operation interface to ensure interaction between users and the platform. It should be noted that the low-code orchestration engine uses a visual process canvas to realize the drag-and-drop combination of evaluation modules and optimization operators to generate custom data governance pipelines, and also has conditional branch settings, version management, process cloning and sharing functions.

[0052] Among them, the visual operation interface supports operations such as dragging the mouse to connect data quality assessment operators and optimization operators to quickly build personalized data governance processes. At the same time, it provides a graphical parameter configuration interface, which can flexibly configure operator parameters through interactive methods such as sliders and drop-down menus. For example, it can adjust the super-resolution multiple, denoising intensity, and the number of sampling steps of the generated model to achieve precise control of the data processing process. Not only that, the low-code platform 10 has built-in commonly used data governance process templates. The templates contain operator combinations corresponding to data governance purposes, such as image noise reduction and enhancement, label equalization, etc. At the same time, users can customize templates according to actual application needs and perform import and export operations, which improves the reusability and work efficiency of the data governance process.

[0053] The low-code platform provided in the embodiment of the present application is composed of a data access module, a data evaluation module, and a data optimization module. The data access module is used to obtain the original data set input into the low-code platform and the metadata of the original data set; the data evaluation module is used to analyze the original data set according to the metadata based on the multi-dimensional evaluation indicators that match the original data set to obtain the quality evaluation results of the original data set; the data optimization module is used to determine multiple preset optimization operators in the low-code platform that match the quality evaluation results, and optimize the original data set through the multiple preset optimization operators that match the quality evaluation results to obtain the target data set. Based on this, the low-code development environment is used to automatically diagnose the quality of the data set through multi-dimensional evaluation indicators, and dynamically schedule the adapted optimization operators to perform targeted data optimization. This solution transforms the traditional data governance process into visual orchestration, enabling users to complete complex algorithm configuration, logic definition, and data mining tasks, reducing the technical threshold while improving data governance efficiency. It solves the problem of difficult and inefficient data governance and realizes the flexible and convenient configuration of data governance solutions with the help of a low-code development environment, significantly improving governance efficiency while reducing the difficulty of data governance.

[0054] In some other embodiments, the low-code platform 10 also includes a distributed computing module 400.

[0055] Specifically, distributed computing module 400 builds a task scheduling system based on a distributed computing framework (such as the open-source computing framework Apache Spark). This system intelligently partitions data optimization tasks and distributes them to multiple computing nodes, allowing each computing node to execute its corresponding optimization task fragment, thereby improving data optimization efficiency. Simultaneously, it monitors the processing progress and resource utilization of each computing node in real time, including key metrics such as CPU utilization, GPU utilization, and memory usage. It also provides a visual display of processing progress and resource utilization, helping users intuitively understand task execution and promptly identify and resolve potential performance bottlenecks.

[0056] It should be noted that the above-mentioned task scheduling system supports the dynamic expansion of graphics processor acceleration clusters, and automatically adjusts the allocation of computing resources according to the computing requirements of the task and the load of the node to ensure efficient execution of tasks, thereby enabling real-time data processing under high concurrency conditions and meeting the needs of multiple signal source inputs. It is suitable for real-time monitoring and rapid response scenarios.

[0057] In some other embodiments, the low-code platform 10 also includes a visualization interaction module 500.

[0058] Specifically, after the data assessment module completes the data quality assessment of the original dataset, it generates a corresponding data quality analysis report and displays it visually. The report formats include, but are not limited to, PDF and JPG. The report includes a list of defective data, a radar chart of indicators, and a comparison chart of typical cases. By visualizing the data assessment results, users can gain a comprehensive understanding of the quality of the dataset, providing strong support for data governance decisions.

[0059] Furthermore, during data processing, a single-sample real-time preview function is supported. For example, when a user clicks on any image in the dataset, a visual comparison of the image before and after restoration is displayed, allowing the user to intuitively experience the data optimization effect. Furthermore, the visualization interaction module 500 provides interactive operations such as image zooming and rotation, allowing users to observe the image in detail.

[0060] The following continues to explain and illustrate in detail the data quality assessment and optimization method applied to the low-code platform provided in the above embodiment of this application in conjunction with the accompanying drawings. Figure 2 This is a flow chart of a data quality assessment and optimization method provided in this embodiment. Figure 2 As shown, the process includes the following steps:

[0061] Step S210, obtaining the original data set input into the low-code platform and the metadata of the original data set;

[0062] Specifically, the low-code platform acquires the original dataset, which can be one or more combinations of text, image, and video datasets, and extracts data from the original dataset to obtain metadata. This metadata includes information such as resolution, capture time, and tags. This metadata is used for subsequent data quality assessment and processing, providing important foundational information for data governance.

[0063] Step S220 , analyzing the original dataset according to the metadata based on the multi-dimensional evaluation indicators that match the original dataset to obtain a quality evaluation result of the original dataset;

[0064] Specifically, multi-dimensional evaluation indicators that match the original data set are determined, and the multi-dimensional evaluation indicators include pixel-level evaluation indicators, semantic-level evaluation indicators, and structural-level evaluation indicators. For example, multiple evaluation indicators that are adapted are selected based on the data type of the original data set. In other embodiments, the multi-dimensional evaluation indicators of the original data set selected by the first user instruction can also be determined based on the first user instruction input through the evaluation interface of the low-code platform.

[0065] Among them, pixel-level evaluation indicators are used to comprehensively evaluate the pixel-level quality of images. Image-related indicators include the uniformity of image brightness, image contrast distribution (which can be reflected by contrast entropy), consistency of image resolution, etc. Video-related indicators include the motion consistency of target objects in the video, whether the audio and picture are synchronized, etc. Text-related indicators include typo detection and grammatical error detection, etc.; semantic-level evaluation aims to provide a semantic basis for data quality assessment. By counting the frequency of label categories and calculating the Gini coefficient, it can determine whether the label distribution is reasonable, so as to deeply analyze the balance of label distribution, or identify the occlusion of targets in the image, calculate the target occlusion rate, or accurately find the missing key information in the dataset through missing value location, etc.; structural-level evaluation indicators include data repetition rate, data missingness, etc. By accurately calculating the data repetition rate, repeated samples in the dataset can be effectively identified, and the precise location of missing values ​​can provide an important reference for the optimization of data structure.

[0066] Furthermore, according to the selected evaluation indicators, the corresponding evaluation operators are called to evaluate and analyze the original data set to obtain the quality evaluation results of the original data set, and generate the corresponding data quality analysis report for visual display. The content of the data quality analysis report includes but is not limited to a list of defective data, an indicator radar chart, a typical case comparison chart, etc., so as to help users fully understand the quality status of the data set through the visualization of data evaluation results, and provide strong support for data governance decisions.

[0067] Exemplarily, when the original data set is an image data set, duplicate data detection, brightness uniformity analysis, and label distribution evaluation are performed on the image data set. Each evaluation step is described in detail below.

[0068] 1) Duplicate Data Detection Process: Each image is pre-normalized to 32×32 pixels, then grayscaled and subjected to a discrete cosine transform (DCT). Low-frequency regions are extracted based on the DCT results, generating an 8×8 feature matrix. A 64-bit perceptual hash value is obtained by averaging and binarizing the feature matrix. Finally, similar images are screened using the perceptual hashing algorithm. Similarity retrieval is performed using a Hamming distance calculation that supports batch vectorization. In a cluster environment, similarity retrieval can be performed for millions of images in seconds. The system also provides dynamically configurable similarity thresholds, with preset threshold templates tailored to different business scenarios, such as security imaging and medical imaging. The threshold adjustment step size can be adjusted accurately to 0.1, ensuring a recall rate of ≥95% and an accuracy rate of ≥98% for duplicate data identification.

[0069] 2) Brightness Uniformity Analysis Process: Each image is divided into a grid with dynamically configurable granularity, such as 4×4, 8×8, and 16×16. Luminance analysis is then performed on the gridded image. Image luminance analysis is calculated based on the L channel values ​​of the Lab color space established by the Commission International Eclairage (CIE), which better reflects human visual characteristics and avoids errors caused by nonlinear responses of the RGB channels. Spatial weighting factors (center weight of 1.2 and edge weight of 0.8) are introduced to optimize standard deviation calculation and enhance the perception of differences in visually sensitive areas. Finally, a visualization module provides a luminance distribution heatmap to intuitively display the luminance analysis results. Interactive highlighting of grid areas is supported. Clicking on any grid cell displays the corresponding area's luminance mean, extreme values, and deviation rate from the global mean, helping users locate overexposed or underexposed image regions.

[0070] 3) Label distribution evaluation process: Label distribution evaluation uses the Gini coefficient calculation based on the Lorenz curve definition, and supports label weight configuration such as positive and negative sample weighting and business priority marking. The specific calculation formula of the Gini coefficient G is as follows:

[0071] (1)

[0072] In formula (1), n ​​represents the total number of samples; Represents the weight of the i-th sample. When the Gini coefficient exceeds the set threshold (default 0.7, configurable range 0.5 to 0.9), the system automatically triggers the balanced recommendation algorithm, generates the corresponding data enhancement scheme (such as SMOTE oversampling and clustering undersampling), and matches it with the generative enhancement function associated with the intelligent optimization operator library to optimize the label distribution structure, forming an evaluation-optimization closed loop.

[0073] Step S230, determining multiple preset optimization operators in the low-code platform that match the quality assessment results;

[0074] It should be noted that the low-code platform's intelligent optimization operator library includes different categories of preset optimization operators, including but not limited to data repair operators, generative enhancement operators, and text error correction and enhancement operators. It supports graphical configuration of algorithm parameters, and each operator uses standardized packaging and has a unified input and output interface. For example, super-resolution operators, denoising operators, text-based image repair operators, and image-based image enhancement operators.

[0075] Specifically, after completing the data quality assessment, the data to be optimized and its data quality problems in the original data set are determined based on the quality assessment results of the original data set, and then a plurality of preset optimization operators that are suitable are selected from the intelligent optimization operator library. In other embodiments, the plurality of optimization operators selected by the second user instruction can also be determined based on the second user instruction input through the visual operation interface for the current data quality assessment result, that is, the appropriate optimization operator is selected from the optimization operator library through the visual operation interface provided by the low-code platform, and the process is orchestrated by dragging, connecting, etc. For example, for low-resolution images, a combination of "super-resolution operator → Vincent image enhancement operator" is added. At the same time, users can flexibly set the parameters of each operator in the parameter configuration interface to generate a personalized execution plan.

[0076] Step S240 , optimizing the original data set by using a plurality of preset optimization operators that match the quality assessment results to obtain a target data set.

[0077] Specifically, multiple preset optimization operators matching the quality assessment results are invoked to optimize the data to be optimized in the original dataset, obtaining the target dataset and forming a closed evaluation-optimization loop. During the processing, a single-sample real-time preview function is supported. For example, real-time visualization of processing progress and intermediate results is provided. When users click on any image in the dataset through the system interface, a visual comparison of the image before and after restoration is displayed, allowing them to intuitively experience the data optimization effect.

[0078] Among them, the low-code platform provides a graphical parameter configuration interface, which can flexibly configure operator parameters through interactive methods such as sliders and drop-down menus, such as adjusting the super-resolution multiple, denoising intensity, and the number of sampling steps of the generated model, to achieve precise control of the data processing process.

[0079] It should be noted that the low-code platform also supports the setting of data governance process templates. The data governance process templates contain operator combinations corresponding to the data governance purposes, such as image noise reduction and enhancement, label equalization, etc. At the same time, users can customize templates according to actual application needs and perform import and export operations, which helps to improve the reusability and work efficiency of the data governance process.

[0080] Furthermore, after the optimization process is complete, the target dataset is exported and subjected to a secondary quality assessment to verify the effectiveness of the initial optimization strategy. Based on the assessment results, the data governance process is dynamically adjusted to drive data quality improvement. For example, if data quality indicators fail to meet preset requirements, root cause analysis is performed to identify specific issues (such as improper interpolation algorithm selection or incorrect super-resolution operator parameter configuration), leading to adjustments to the data governance process (such as replacing optimization operators). Conversely, if the secondary assessment results meet the preset requirements, the process can be solidified as a governance template.

[0081] Data governance is key to ensuring data quality. However, current data governance solutions rely heavily on programming for data quality assessment and optimization, requiring users to master multiple programming languages ​​and familiarize themselves with complex algorithm libraries, making data governance difficult. Furthermore, these solutions require repetitive coding to meet the processing requirements of different data sets, resulting in inefficient data governance.

[0082] Compared with the existing technology, the present application obtains the original data set input into the low-code platform and the metadata of the original data set; based on the multi-dimensional evaluation indicators that match the original data set, the original data set is analyzed according to the metadata to obtain the quality evaluation results of the original data set; multiple preset optimization operators that match the quality evaluation results in the low-code platform are determined; the original data set is optimized through multiple preset optimization operators that match the quality evaluation results to obtain the target data set. Based on this, the low-code development environment is utilized to automatically diagnose the quality of the data set through multi-dimensional evaluation indicators, and dynamically schedule the adapted optimization operators to perform targeted data optimization. This solution transforms the traditional data governance process into visual orchestration, which improves efficiency while lowering the technical threshold, solves the problem of difficult and inefficient data governance, and realizes the flexible and convenient configuration of data governance solutions with the help of a low-code development environment, reduces the difficulty of data governance, improves governance efficiency, and effectively shortens the governance cycle.

[0083] In some embodiments, after obtaining the input original data set and metadata of the original data set, the following steps are further included:

[0084] In response to a first user instruction input through the evaluation interface of the low-code platform, multidimensional evaluation indicators of the original data set selected by the first user instruction are determined; the multidimensional evaluation indicators include pixel-level evaluation indicators, semantic-level evaluation indicators, and structural-level evaluation indicators.

[0085] Specifically, the low-code platform provides an evaluation interface where users can flexibly select the desired evaluation dimensions, such as resolution detection and label distribution analysis. Based on the first user instruction entered through the evaluation interface, the platform determines the multi-dimensional evaluation indicators for the original dataset selected by the first user instruction. These multi-dimensional evaluation indicators primarily include pixel-level evaluation indicators, semantic-level evaluation indicators, and structural-level evaluation indicators.

[0086] Among them, pixel-level evaluation indicators are used to comprehensively evaluate the pixel-level quality of images. Image-related indicators include the uniformity of image brightness, image contrast distribution, consistency of image resolution, etc. Video-related indicators include the motion consistency of target objects in the video, whether the audio and picture are synchronized, etc. Text-related indicators include typo detection and grammatical error detection, etc.; semantic-level evaluation aims to provide a semantic basis for data quality assessment. By counting the frequency of label categories, calculating the Gini coefficient, etc., it determines whether the label distribution is reasonable, so as to deeply analyze the balance of label distribution, or identify the occlusion of targets in the image, calculate the target occlusion rate, or accurately find the missing key information in the dataset through missing value location, etc.; structural-level evaluation indicators include data repetition rate, data missingness, etc. By accurately calculating the data repetition rate, repeated samples in the dataset can be effectively identified, and the precise location of missing values ​​can provide an important reference for the optimization of data structure.

[0087] Through this embodiment, in response to a first user instruction input through the evaluation interface of the low-code platform, multi-dimensional evaluation indicators of the original data set selected by the first user instruction are determined. The multi-dimensional evaluation indicators include pixel-level evaluation indicators, semantic-level evaluation indicators, and structural-level evaluation indicators. This enables systematic data quality to be guaranteed through multi-dimensional data evaluation, providing a high-confidence data source for downstream model training, thereby improving model generalization capabilities and prediction accuracy.

[0088] In some of these embodiments, Figure 3 As shown, determining multiple preset optimization operators in the low-code platform that match the quality assessment results in step S230 includes the following steps:

[0089] Step S231, determining an optimization strategy for the original data set based on the quality assessment result;

[0090] Step S232, determine the optimization operator combination that matches the optimization strategy in the low-code platform; the optimization operator combination includes multiple preset optimization operators.

[0091] Specifically, after completing the data quality assessment, the data to be optimized and its data quality issues are determined based on the quality assessment results of the original dataset, and then the optimization strategy for the original dataset is determined. For example, if the assessment results indicate that the original dataset contains a large number of missing values, the optimization strategy may include tracing the data source to find the original records to complete the missing values ​​for key fields, or applying interpolation algorithms to fill in the missing values ​​for numeric fields. If the assessment results indicate that the image resolution in the original dataset does not meet the preset requirements, the optimization strategy may include increasing the image resolution to ensure image quality.

[0092] Furthermore, based on the optimization strategy of the original dataset, an intelligent algorithm determines the matching optimization operator combination in the low-code platform. This optimization operator combination typically includes multiple preset optimization operators. For example, if the optimization strategy indicates that image resolution needs to be increased, the combination of the super-resolution operator and the generative enhancement operator is automatically triggered.

[0093] Through this embodiment, the optimization strategy of the original data set is determined based on the quality assessment results, and the optimization operator combination that matches the optimization strategy in the low-code platform is determined. The optimization operator combination includes multiple preset optimization operators, thereby providing differentiated solutions for different types of data defects, such as low resolution, noise pollution, label skew, etc., forming an evaluation-optimization closed loop, which helps to achieve accurate data optimization and provide a higher-quality data foundation for downstream machine learning model training, so as to improve the training effect and performance of the model.

[0094] In some of these embodiments, Figure 4 As shown, step S240 optimizes the original data set by using multiple preset optimization operators that match the quality assessment results to obtain the target data set, including the following steps:

[0095] Step S241: Arrange the processes of the preset optimization operators that match the quality assessment results to obtain corresponding optimization processes; the optimization processes are used to indicate the execution order of the preset optimization operators;

[0096] Step S242: Based on the optimization process, various preset optimization operators are called to optimize the original data set to obtain a target data set.

[0097] Specifically, the adapted preset optimization operators are orchestrated to obtain a corresponding optimization process, which is used to indicate the execution order of the preset optimization operators. Based on the optimization process, the preset optimization operators are called to optimize the data to be optimized to obtain the target data set.

[0098] For example, for low-resolution images, a combination of "super-resolution operator → Vincent image enhancement operator" is added. The super-resolution operator is used to increase the image resolution to the target scale and reconstruct the basic texture structure. The Vincent image enhancement operator repairs the blurred areas of the image based on the input text description, supporting fine restoration from pixel level to deep enhancement at semantic level, realizing multi-level data optimization and comprehensively ensuring data quality.

[0099] Through this embodiment, the preset optimization operators that match the quality assessment results are orchestrated to obtain the corresponding optimization process. The optimization process is used to indicate the execution order of each preset optimization operator. Then, based on the optimization process, each preset optimization operator is called to optimize the original data set to obtain the target data set, thereby achieving accurate data optimization and improving data quality.

[0100] In some embodiments, the data quality assessment and optimization method further includes the following steps:

[0101] Based on the second user instruction entered through the parameter configuration interface of the low-code platform, the algorithm parameters of each preset optimization operator are dynamically adjusted.

[0102] Specifically, the low-code platform provides a graphical parameter configuration interface that dynamically adjusts the algorithm parameters of each preset optimization operator based on second-party user instructions entered through the parameter configuration interface. Algorithm parameters such as super-resolution multiple, denoising intensity, and the number of sampling steps for the generated model can be flexibly set through interactive methods such as sliders and drop-down menus.

[0103] It's important to note that to ensure real-time response and interpretability of parameter adjustments, the platform features a built-in parameter linkage engine. When users modify algorithm parameters, the system automatically calculates and displays the affected associated parameters. Furthermore, all adjustments trigger an instant preview. For example, before-and-after comparison examples are generated in the sidebar, ensuring users can intuitively perceive the impact of parameters.

[0104] Through this embodiment, based on the second user instruction input through the parameter configuration interface of the low-code platform, the algorithm parameters of each preset optimization operator are dynamically adjusted to achieve precise control of the data processing process, which helps to improve the data optimization effect.

[0105] In some embodiments, the data quality assessment and optimization method further includes the following steps:

[0106] The optimization processing task of the original data set is divided into slices to obtain multiple optimization task slices;

[0107] Allocate each optimization task slice to multiple computing nodes of the low-code platform so that the corresponding optimization task slice can be executed through each computing node.

[0108] Specifically, on the low-code platform, a task scheduling system is built based on a distributed computing framework. This system can intelligently slice and slice data optimization tasks and distribute them to multiple computing nodes, so that each computing node can execute the corresponding optimization task slice, improving data optimization efficiency. At the same time, the processing progress and resource utilization of each computing node are monitored in real time, including key indicators such as CPU utilization, GPU utilization, and memory utilization. The processing progress and resource utilization are visually displayed to help users intuitively understand the task execution status and promptly identify and resolve potential performance bottlenecks.

[0109] It should be noted that the above-mentioned task scheduling system supports dynamic expansion of the graphics processor acceleration cluster, and automatically adjusts the allocation of computing resources according to the computing requirements of the task and the load of the node, thereby ensuring efficient execution of the task.

[0110] Through this embodiment, the optimization processing tasks of the original data set are sharded to obtain multiple optimization task shards, and each optimization task shard is distributed to multiple computing nodes of the low-code platform, so that the corresponding optimization task shard is executed by each computing node, thereby achieving efficient execution of optimization tasks and improving data governance efficiency.

[0111] In some of the embodiments, each preset optimization operator includes a data repair operator and a generative enhancement operator.

[0112] Specifically, each preset optimization operator includes a data restoration operator and a generative enhancement operator. For example, a super-resolution operator, a denoising operator, a text-based image restoration operator, and a graphic-based image enhancement operator. The following examples illustrate different optimization operators.

[0113] 1) Super-resolution Operator: Supports the bicubic interpolation algorithm and the Enhanced Deep Residual Network (EDSR) model, with users able to configure magnification factors such as 2x, 4x, or 8x as needed. The bicubic interpolation algorithm uses polynomial fitting to increase image resolution while maintaining image smoothness. The EDSR model leverages a deep residual network to learn image texture priors, generating higher-quality super-resolution images and effectively enhancing image detail.

[0114] 2) Denoising Operator: This operator uses the Local Binary Pattern (LBP) algorithm and the Gaussian Mixture Model to collaboratively detect image noise, accurately identifying salt and pepper noise and Gaussian noise, and quantifying the noise density (0–100%). It also employs a median filter with adaptive windows (dynamically adjusted from 3×3 to 11×11), or an optimized three-dimensional block matching strategy (BM3D), which introduces spatial position constraints and a color histogram similarity metric. This ensures that the Structural Similarity Index Measure (SSIM) is greater than or equal to 0.92 when the noise density is less than or equal to 30%, better preserving image edges, textures, and other structures. This also improves processing speed by 40% compared to traditional methods.

[0115] 3) Text-based image restoration operator: This operator implements text-driven local generation based on text-based image models (such as the Stable Diffusion Model). When a user enters a text description, such as "repair a blurred cat face," the generative model generates a high-definition complete image, automatically filling in the missing parts of the image based on the user's description, effectively improving image clarity and completeness.

[0116] 4) Image generation and enhancement operator: Through the ControlNet model, the image is enhanced while preserving the structural features of the input image to generate an image version with higher resolution, richer colors, and clearer details, meeting users' needs for improved image quality.

[0117] Through this embodiment, the visual algorithm is organically combined with the large model generation technology to achieve a rapid improvement in data quality.

[0118] The following describes and illustrates this embodiment using medical image dataset management as an example, which specifically includes the following steps:

[0119] S1. Data collection and input.

[0120] 5,000 medical images in Digital Imaging and Communications in Medicine (DICOM) format were uploaded to the low-code platform via the File Transfer Protocol (FTP). After receiving the medical image dataset, a format parsing engine was used to convert each image from DICOM to PNG format and store it in a distributed file storage cluster (Hadoop HDFS). During storage, metadata such as image resolution and capture time were extracted to prepare for subsequent data processing.

[0121] In the system's evaluation interface, select multiple evaluation dimensions such as resolution detection, noise level analysis, and label integrity check according to user input instructions, and click the Execute Evaluation button to start a comprehensive quality assessment of the dataset.

[0122] S2. Analyze data quality assessment results.

[0123] The data quality assessment results are output, including that 1,200 images in the dataset have a resolution of 256×256, which is significantly lower than the 512×512 resolution standard required for clinical diagnosis. Low-resolution images will affect diagnostic accuracy; 300 images in the dataset have a noise standard deviation greater than the preset value, that is, there is obvious salt and pepper noise, which will interfere with the effective information in the image and reduce image quality and diagnostic accuracy; 50 images in the dataset lack lesion area annotations, and the label completeness is 99%. Missing labels will affect data availability and model training effects.

[0124] S3. Orchestrate data optimization processes.

[0125] In the low-code platform's visual interface, add an input node to introduce the medical image data to be processed. Next, connect the resolution detection operator to perform resolution detection on each medical image. Next, set a conditional branch. If the image resolution is less than 512×512, connect the EDSR super-resolution operator (with a magnification of 2x) and the denoising operator (with the BM3D algorithm and a noise intensity of 15). If the image resolution meets the required resolution, output it directly.

[0126] Among them, for images processed by the super-resolution operator and the denoising operator, the Wensheng image enhancement operator is connected and a prompt word is input (such as the prompt word is "enhance the contrast of medical images and retain lesion details") to further improve the image quality.

[0127] It's important to note that the parameter configuration interface allows you to set specific parameters for each operator. For example, the EDSR model uses pre-trained medical imaging weights to better adapt to the characteristics of medical imaging; the BM3D algorithm uses an 8×8 block size to ensure that image details are preserved while removing noise; and the generative model uses a sampling step count of 30 to balance image quality and processing efficiency.

[0128] S4. Distributed computing and visualization.

[0129] The Apache Spark cluster automatically activated its task scheduling system, dividing the optimization task for 5,000 medical images into 10 compute nodes for parallel processing. The GPU nodes were responsible for super-resolution and model generation, leveraging the GPU's powerful computing power to increase processing speed. The CPU nodes handled tasks such as format conversion, ensuring the efficiency of the entire processing flow.

[0130] During processing, a visual interface displays real-time progress, helping users quickly understand data processing status and displaying detailed information such as the average processing time for each image. Clicking on any image also allows for a real-time preview of the image processing results. Furthermore, the visual interface displays resource utilization for each node, making it easy for users to monitor task execution.

[0131] S5. Result output and iteration.

[0132] After optimizing all images, the generated high-quality medical image dataset will be exported for subsequent use, such as in medical imaging analysis and machine learning model training. Prior to export, the system will perform a secondary verification of the optimized dataset to ensure data integrity and accuracy. Users can also evaluate the optimized dataset on the platform and, by comparing the evaluation results with the quality indicators of the original dataset, gain an intuitive understanding of the improvement in data quality. If the optimization effect does not meet expectations, the data governance process can be adjusted and optimized based on the evaluation results, and the data processing task can be re-executed to achieve continuous improvement and iteration of data quality.

[0133] This embodiment also provides a data quality assessment and optimization device, which is used to implement the above-mentioned embodiments and preferred embodiments. Details already described will not be repeated here. The terms "module," "unit," "subunit," etc. used below refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0134] Figure 5 This is a structural block diagram of the data quality assessment and optimization device of this embodiment. Figure 5 As shown, the device includes:

[0135] An acquisition module 600 is used to acquire an original data set input into the low-code platform and metadata of the original data set;

[0136] Evaluation module 700, which analyzes the original dataset according to metadata based on multi-dimensional evaluation indicators that match the original dataset to obtain a quality evaluation result of the original dataset;

[0137] The optimization module 800 determines a plurality of preset optimization operators in the low-code platform that match the quality assessment results;

[0138] The optimization module 800 optimizes the original data set by using a plurality of preset optimization operators that match the quality assessment results to obtain a target data set.

[0139] Through the device provided by this embodiment, the original data set and the metadata of the original data set input into the low-code platform are obtained; based on the multi-dimensional evaluation indicators matching the original data set, the original data set is analyzed according to the metadata to obtain the quality evaluation result of the original data set; multiple preset optimization operators matching the quality evaluation result in the low-code platform are determined; the original data set is optimized by multiple preset optimization operators matching the quality evaluation result to obtain the target data set, which solves the problem of difficult and inefficient data governance, and realizes the flexible and convenient configuration of data governance solutions with the help of the low-code development environment, which reduces the difficulty of data governance while significantly improving governance efficiency.

[0140] In some of these embodiments, Figure 5 On the basis of this, the device also includes a selection module for determining the multidimensional evaluation indicators of the original data set selected by the first user instruction in response to a first user instruction input through the evaluation interface of the low-code platform; the multidimensional evaluation indicators include pixel-level evaluation indicators, semantic-level evaluation indicators and structural-level evaluation indicators.

[0141] In some of the embodiments, the optimization module 800 is also used to determine the optimization strategy of the original data set based on the quality assessment results; determine the optimization operator combination in the low-code platform that matches the optimization strategy; the optimization operator combination includes multiple preset optimization operators.

[0142] In some of the embodiments, the optimization module 800 is also used to orchestrate the processes of the preset optimization operators that match the quality assessment results to obtain the corresponding optimization processes; the optimization processes are used to indicate the execution order of the preset optimization operators; based on the optimization processes, the preset optimization operators are called to optimize the original data set to obtain the target data set.

[0143] In some of the embodiments, the optimization module 800 is also used to dynamically adjust the algorithm parameters of each preset optimization operator based on a second user instruction input through the parameter configuration interface of the low-code platform.

[0144] In some of these embodiments, Figure 5 On the basis of, the device also includes a scheduling module for sharding the optimization processing tasks of the original data set to obtain multiple optimization task shards; and allocating each optimization task shard to multiple computing nodes of the low-code platform to execute the corresponding optimization task shard through each computing node.

[0145] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can be located in the same processor; or the above modules can be located in different processors in any combination.

[0146] This embodiment further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0147] Optionally, the computer device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor, and the input / output device is connected to the processor.

[0148] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0149] S1, obtains the original data set input into the low-code platform and the metadata of the original data set;

[0150] S2, based on the multi-dimensional evaluation indicators matching the original dataset, analyzes the original dataset according to the metadata to obtain the quality evaluation results of the original dataset;

[0151] S3, determining multiple preset optimization operators in the low-code platform that match the quality assessment results;

[0152] S4, optimizes the original data set through multiple preset optimization operators that match the quality assessment results to obtain the target data set.

[0153] It should be noted that, for specific examples in this embodiment, reference may be made to the examples described in the above embodiments and optional implementation modes, and will not be repeated in this embodiment.

[0154] In addition, in conjunction with the data quality assessment and optimization methods provided in the above embodiments, a storage medium may also be provided in this embodiment to implement the data quality assessment and optimization methods. The storage medium stores a computer program that, when executed by a processor, implements any of the data quality assessment and optimization methods in the above embodiments.

[0155] It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit it. Based on the embodiments provided in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0156] Obviously, the accompanying drawings are merely examples or embodiments of the present application. A person skilled in the art can also apply the present application to other similar situations based on these drawings without inventive effort. Furthermore, it is understandable that, although the work involved in this development process may be complex and lengthy, certain design, manufacturing, or production changes based on the technical content disclosed in this application are merely routine technical means for a person skilled in the art and should not be considered to constitute a deficiency in the disclosure of the present application.

[0157] The term "embodiment" as used in this application refers to specific features, structures, or characteristics described in conjunction with the embodiment that can be included in at least one embodiment of this application. The appearance of this phrase in various places in the specification does not necessarily mean that the embodiment is the same, nor does it mean that it is mutually exclusive with other embodiments and is independent or optional. It is understood, either explicitly or implicitly, by those skilled in the art that the embodiments described in this application can be combined with other embodiments when there is no conflict.

[0158] The above-described embodiments merely represent several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of patent protection. It should be noted that a person of ordinary skill in the art may make various modifications and improvements without departing from the spirit of the present application, and these modifications and improvements fall within the scope of protection of the present application. Therefore, the scope of protection of the present application shall be determined by the appended claims.

Claims

1. A data quality assessment and optimization method, characterized in that: Applied to low-code platforms; The method comprises: Obtaining an original data set input into the low-code platform and metadata of the original data set; Analyzing the original dataset according to the metadata based on a multi-dimensional evaluation index that matches the original dataset to obtain a quality evaluation result of the original dataset; Determining a plurality of preset optimization operators in the low-code platform that match the quality assessment result; The original data set is optimized by using a plurality of preset optimization operators that match the quality assessment result to obtain a target data set.

2. The data quality assessment and optimization method according to claim 1, characterized in that: After obtaining the input original data set and metadata of the original data set, the method further includes: In response to a first user instruction input through the evaluation interface of the low-code platform, the multidimensional evaluation indicators of the original data set selected by the first user instruction are determined; the multidimensional evaluation indicators include pixel-level evaluation indicators, semantic-level evaluation indicators and structural-level evaluation indicators.

3. The data quality assessment and optimization method according to claim 1, characterized in that: The determining of a plurality of preset optimization operators in the low-code platform that match the quality assessment result includes: Determining an optimization strategy for the original data set based on the quality assessment result; Determine an optimization operator combination in the low-code platform that matches the optimization strategy; the optimization operator combination includes multiple preset optimization operators.

4. The data quality assessment and optimization method according to claim 1, characterized in that: The optimizing process of the original data set by using the plurality of preset optimization operators matching the quality assessment results to obtain the target data set includes: Performing process arrangement on each of the preset optimization operators that matches the quality assessment result to obtain a corresponding optimization process; the optimization process is used to indicate the execution order of each of the preset optimization operators; Based on the optimization process, each of the preset optimization operators is called to optimize the original data set to obtain the target data set.

5. The data quality assessment and optimization method according to claim 4, characterized in that: The method further comprises: Based on a second user instruction input through the parameter configuration interface of the low-code platform, the algorithm parameters of each of the preset optimization operators are dynamically adjusted.

6. The data quality assessment and optimization method according to claim 1, characterized in that: The method further comprises: Slicing the optimization processing task of the original data set to obtain multiple optimization task slices; Each of the optimization task slices is distributed to multiple computing nodes of the low-code platform so that the corresponding optimization task slice is executed by each of the computing nodes.

7. The data quality assessment and optimization method according to any one of claims 1 to 6, characterized in that: Each of the preset optimization operators includes a data repair operator and a generative enhancement operator.

8. A low-code platform, characterized in that include: A data access module, used to obtain the original data set input into the low-code platform and the metadata of the original data set; A data evaluation module is configured to analyze the original data set according to the metadata based on a multi-dimensional evaluation index that matches the original data set, and obtain a quality evaluation result of the original data set; A data optimization module, configured to determine a plurality of preset optimization operators in the low-code platform that match the quality assessment result; The data optimization module is further configured to optimize the original data set by using a plurality of preset optimization operators that match the quality assessment result to obtain a target data set.

9. A computer device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to run the computer program to perform the steps of the data quality assessment and optimization method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the data quality assessment and optimization method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Data management method and device

    CN113722302A

  • Dragging editing method based on webpage end

    CN118585186A

  • Multi-modal large model construction method based on low codes

    CN119148997A

  • Intelligent operation risk analysis method and system based on deep learning

    CN120125006A

  • System and method for sandboxing

    WO2025114678A1

Cited By

  • Data quality and preparation cost evaluation device for domain large model

    CN121256299A