Data quality evaluation system based on large model
By constructing a data quality evaluation system based on a large model and utilizing transfer learning and blockchain technology, the problems of strong subjectivity, low efficiency, and insufficient credibility in industrial data evaluation have been solved, achieving high-precision, transparent, and reliable data quality evaluation and asset valuation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGSU JINYICHENG TECHNOLOGY DEVELOPMENT CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-05-12
AI Technical Summary
Existing technologies for data quality assessment in the industrial sector suffer from several problems, including difficulty in implementing evaluation standards, strong reliance on subjective factors, low efficiency, high costs, and insufficient credibility and public trust in the evaluation results. In particular, when faced with complex and heterogeneous industrial data resources, it is difficult to achieve efficient and reliable quality assessment.
We construct a data quality evaluation system based on a large model. By introducing transfer learning and knowledge distillation techniques for adaptive training and combining blockchain technology for full-process coding and on-chain, we achieve collaboration between large and small models, ensuring the transparency of the evaluation process and the credibility of the results.
It has enabled high-precision automated quality assessment of data from different fields and with different structures, improving the accuracy, transparency and credibility of the assessment results, reducing human interference, improving work efficiency and system reliability, and opening up the value chain from data quality assessment to data asset valuation.
Smart Images

Figure CN122019520A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of data governance and artificial intelligence technology, and in particular to a data quality evaluation system based on a large model. Background Technology
[0002] With the deepening development of the big data era, data has become a key production factor. In the industrial sector, massive amounts of data hold immense value in driving production optimization, service innovation, and model transformation. To standardize and promote the realization of data resource value, the national government has successively issued important policies such as the "Interim Provisions on Accounting Treatment Related to Enterprise Data Resources" and the "Guiding Opinions on Strengthening Data Asset Management," providing clear policy guidance and implementation paths for industrial enterprises to explore the assetization of data resources and promote the valuation, inclusion, and capitalization of data assets. Against this backdrop, conducting scientific and reliable assessments of industrial data assets becomes a prerequisite for unlocking their value. Regardless of the valuation model used, data quality remains the core foundation and key input for measuring and determining the value of data assets.
[0003] Currently, data quality evaluation mainly relies on the national standard "Information Technology Data Quality Evaluation Indicators" (GB / T36344-2018), which establishes an evaluation framework from multiple dimensions, including standardization, completeness, accuracy, consistency, timeliness, and accessibility. However, in the practical application of this standard for asset-based assessment of complex, massive, and heterogeneous data resources in the industrial sector, a series of prominent challenges arise: 1. Difficulty in Implementing Evaluation Standards and High Subjectivity: The industrial sector is characterized by numerous sub-sectors, and the quality connotations and importance weights of various data resources (such as equipment operation data, process parameters, and supply chain information) vary greatly. Selecting a suitable subset of indicators from the standards and scientifically determining the weights of each indicator for a specific evaluation object heavily relies on expert experience. This makes the evaluation process highly subjective and inconsistent, hindering the formation of objective and reusable evaluation models, a common problem for both evaluation agencies and industrial enterprises.
[0004] 2. Low evaluation efficiency and high cost: Industry is a data-rich sector, and the datasets to be evaluated are often large in scale and complex in structure. Traditional quality evaluation methods based on manual sampling or rule calculation are time-consuming and cannot meet the timeliness requirements of assetization assessment. Conducting evaluations of full or large-scale samples will generate prohibitive human and computing resource costs, restricting the large-scale implementation of data assetization.
[0005] 3. Insufficient credibility and public trust in the evaluation process and results: Currently, the data element market has diverse participants, but there is a lack of data quality evaluation institutions uniformly recognized by authoritative regulatory authorities. Evaluation results generated by centralized systems or single institutions lack transparency in the process, are difficult to verify in terms of conclusions, and lack a fair audit and traceability mechanism. This makes it difficult to gain widespread recognition from transaction parties and regulatory authorities, thereby affecting the market recognition and liquidity of data assets based on the results.
[0006] Therefore, the industry urgently needs an innovative solution to overcome the aforementioned shortcomings. Although artificial intelligence technology has brought hope, simply applying large models faces problems such as high deployment costs and difficulties in transferring domain knowledge; while relying solely on small models has limited evaluation capabilities. At the same time, ensuring the transparency of the automated evaluation process and the immutability of the results remains a technological blind spot.
[0007] In view of this, there is an urgent need to build a new industrial data quality evaluation system and method that can integrate the powerful cognition of large models with the high efficiency and accuracy of small models, and use blockchain technology to solidify the evaluation traces, thereby achieving efficient, intelligent and reliable quality evaluation of industrial data resources, and providing solid and reliable technical support for the accurate valuation and compliant entry of data assets. Summary of the Invention
[0008] The purpose of this invention is to provide a data quality assessment system based on a large model. This system introduces a large model and utilizes transfer learning or knowledge distillation techniques for adaptive training, constructing an intelligent assessment system that coordinates large and small models. This enables high-precision automated quality assessment of data from different domains and with different structures. Simultaneously, by using blockchain technology to encode and record the entire assessment process and results on the blockchain, the immutability of the assessment process and the traceability of the results are ensured, significantly improving the accuracy, transparency, and credibility of the assessment results.
[0009] The technical solution to achieve the purpose of this invention is: the data quality evaluation system based on a large model in this invention includes: The hardware layer consists of a virtualized pool of communication, computing, storage, and security resources, formed by building network devices, computing devices, storage devices, and security devices. The service layer, deployed above the hardware layer, includes network communication services, intelligent computing services, data storage services, information security services, digital wallet services, and basic general services. The application layer, deployed above the service layer, includes a quality assessment module and an on-chain coding module; The business layer, deployed above the application layer, includes a system management module, a configuration management module, and an evaluation management module; The interface layer, deployed above the business layer, is configured to connect to external platforms based on the network communication service.
[0010] Furthermore, in the aforementioned service layer, the intelligent computing service is configured to adaptively train the data quality evaluation model trained on the large model platform using transfer learning or knowledge distillation; the digital wallet service is configured to connect with smart contracts.
[0011] Furthermore, in the application layer described above, the quality assessment module is configured to call the intelligent computing service, use the selected assessment model to evaluate multiple quality indicators of the data to be assessed, and generate a data quality coefficient. The encoding and on-chain module is configured to call the digital wallet service to encode and encrypt the relevant indicators of the quality assessment process and then upload them to the smart contract.
[0012] Furthermore, the aforementioned quality indicators include data standardization, completeness, accuracy, consistency, timeliness, and accessibility.
[0013] Furthermore, in the aforementioned business layer, the system management module is configured to manage the external platform information connected to the system and the system's own configuration information; The configuration management module is configured to configure the selection of the evaluation model, the information of the evaluated data, and the extraction rules of the evaluated data; The evaluation management module is configured to select an evaluation model and a data sample to be evaluated according to the configuration of the configuration management module, and call the application layer to perform quality evaluation and quality coefficient calculation.
[0014] Furthermore, the aforementioned interface layer is configured to connect to one or more of the following through a preset communication protocol: a high-quality dataset construction platform, a blockchain platform, an evaluated data system, a data asset management platform, and a system operation and maintenance platform.
[0015] Furthermore, the aforementioned preset communication protocols include HTTPS, API, gRPC, and SDK.
[0016] Furthermore, the aforementioned application layer is configured to operate according to the following logic: S1. After receiving the business request, configure the parameters; S2. Determine if a corresponding model exists: If it exists, then determine the model; if it does not exist, then obtain the data quality evaluation model from the high-quality dataset construction platform, optimize it, and then determine the model. S3. Extract sample data from the data system being evaluated; S4. Calculate the data quality coefficient of the sample data based on the model driven in step S2 to obtain the data quality coefficient; S5. Perform quality evaluation coding on the associated blockchain platform on the data quality coefficient; S6. Feedback business requests based on the coding results and complete the data quality evaluation.
[0017] Furthermore, the calculation process for the aforementioned data quality coefficient is as follows: A. Parameter Configuration: System parameters are set using CLI or Web models, including the industry attributes of the data being evaluated, the total amount of data being evaluated (C), and the data structure (S). B. Model Selection: Analyze the parameter attributes of the data being evaluated. The system uses a labeling method to select a suitable evaluation model. If the corresponding model does not exist, a suitable data quality evaluation model is obtained from a high-quality dataset platform, and the evaluation model is obtained after optimization of a small model. C. Sample extraction model calculation: Based on the total amount of data, data type, and data structure, the sample extraction model is selected according to the following formula:
[0018] Where R: Sample extraction model C: Total amount of data CNT: Total Data Resource Boundary Point SUM(S): Types of data structures. N: The dividing point between different types of data structures. Ri: The i-th type of stochastic model; D. Extracting sample data: Extracting the data resources to be evaluated according to the selected random sampling model, obtaining sample data resources with a sample size of N, and sending them to the intelligent computing service. E. Selection of data quality indicators for intelligent computing services: Based on the industry to which the data resources belong, sample data, and other information, use the optimized calculation model to configure the data quality coefficient calculation indicator system and the weight value Xi of each indicator; F. Calculation of coefficients for each indicator system by intelligent computing services: The intelligent computing service relies on the evaluation model to statistically analyze the total amount of data to be evaluated and the data resources that meet the indicator requirements, and calculates the coefficients for each indicator system according to the following formula. : ; G. Intelligent computing service calculates data quality coefficient: The data quality coefficient α is calculated according to the following formula:
[0019] Where α: data quality coefficient in data resources. n: The number of data quality indicators selected i: The i-th data quality evaluation index Xi: The weight value of the i-th data quality evaluation indicator. : The coefficient of the i-th data quality evaluation indicator.
[0020] Furthermore, the steps for the above quality assessment coding are as follows: a. Data Coding: Encode information such as the industry of the data, data sample parameters, evaluation model parameters, evaluation time, and evaluation results in the data quality calculation process according to specified rules; b. Construct a secure channel: Connect the system to the smart contract address and construct a trusted channel between the system and the smart contract based on network security services and network communication services; c. Data upload to the blockchain: The encoded data quality evaluation information is sent to the blockchain platform via a secure channel, and the upload result response is received and recorded.
[0021] Furthermore, the data quality coefficient calculated by the application layer Upon request from the data asset assessment system, the data quality coefficient is fed back to the data asset assessment system. The data asset assessment system then uses the following calculation model to obtain the data utility of the data asset, and thus obtains the final data asset valuation: U= (1+l)(1-r); Where U: data asset valuation, α: Data quality coefficient β: Data flow coefficient l: Data monopoly coefficient r: Risk coefficient for realizing data value.
[0022] The present invention has the following positive effects: (1) Through the layered decoupling design, the present invention makes the responsibilities of each module of the system clear, easy to maintain, expand and upgrade. For example, the algorithm of the service layer can be replaced independently or the business logic of the application layer can be updated. At the same time, hardware virtualization and resource pooling provide elastic and stable basic support, ensuring that the system can efficiently and reliably handle large-scale data evaluation tasks. Moreover, this architecture lays a solid system foundation for realizing automated and intelligent evaluation processes and transparent and reliable result traceability, fundamentally improving work efficiency and system reliability.
[0023] (2) This invention utilizes intelligent computing services to adaptively train large models using transfer learning or knowledge distillation, achieving collaboration between large and small models. This inherits the powerful generalization and feature extraction capabilities of large models while gaining the advantages of rapid inference and low-cost deployment of small models, thus enabling "better, better, and faster" high-precision quality evaluation of datasets of different domains and sizes. Simultaneously, the connection between digital wallet services and smart contracts provides a technical channel for subsequent on-chain traceability of key evaluation information, which is a core element in building a trustworthy evaluation mechanism.
[0024] (3) In this invention, the quality assessment module performs automated evaluation by calling intelligent computing services, which reduces human involvement and the influence of subjective factors, making the evaluation results more accurate. The coding and on-chain module encodes the evaluation process and results on the blockchain, which is the first time that blockchain evidence storage technology has been systematically introduced into the field of data quality evaluation. This makes the evaluation process and conclusions traceable by everyone, ensuring the transparency of the process and the credibility of the results, and providing an immutable chain of evidence for the value recognition of data assets.
[0025] (4) This invention lists multi-dimensional evaluation indicators for data quality and constructs a comprehensive and systematic evaluation indicator system covering standardization, completeness, accuracy, consistency, timeliness and accessibility. It comprehensively measures from multiple key dimensions of the data life cycle, ensuring the comprehensiveness and scientific nature of the quality evaluation, so that the obtained "data quality coefficient" can more realistically and from multiple perspectives reflect the actual usability of the data.
[0026] (5) Through the modular design of system management, configuration management and evaluation management, this invention allows users to flexibly configure evaluation tasks, select models and manage data sources through a user-friendly front-end (such as Vue.js), which greatly reduces the technical threshold for using the system. Moreover, it realizes the full-process automated management from task configuration and resource scheduling to computation invocation. Users only need to initiate a request, and the subsequent steps are all completed automatically by the system, resulting in extremely high work efficiency.
[0027] (6) This invention clarifies the linkage capability between the system and the high-quality dataset construction platform, blockchain platform, evaluated data system, data asset management platform, and system operation and maintenance platform. This design makes the system of this invention not an isolated tool, but can be integrated into the existing data technology ecosystem to achieve integrated collaboration of model acquisition, data access, asset valuation, and operation and maintenance monitoring, thereby improving the practicality and application value of the system.
[0028] (7) By supporting multiple mainstream and secure communication protocols such as HTTPS, API, GRPC, and SDK, this invention ensures that the system can perform efficient, stable, and secure data interaction and service calls with various heterogeneous external platforms, thereby enhancing the system's compatibility and integration capabilities.
[0029] (8) The automated workflow of the application layer in this invention describes a complete, closed-loop automated workflow from receiving a request to receiving a feedback result. It covers all key steps such as model judgment, data extraction, calculation, and on-chain processing, ensuring end-to-end automatic execution of the evaluation task. Furthermore, it explicitly proposes a mechanism for automatically acquiring and optimizing high-quality datasets from a platform when there is a lack of readily available models, demonstrating the system's self-learning and adaptive capabilities and ensuring the feasibility of the system's evaluation when facing new domain data.
[0030] (9) This invention describes in detail the automated calculation model and process of data quality coefficients, and proposes a formulaic method for intelligently selecting sample extraction models based on total data volume and data structure, making the sampling process scientific and reasonable, and taking into account both efficiency and representativeness. Furthermore, it defines a complete mathematical model from indicator weight allocation to the calculation of each indicator coefficient, and finally to the synthesis of the quality coefficients. This model quantifies complex quality evaluation into a calculable coefficient, minimizing the interference of human judgment, making the calculation results objective, consistent, and repeatable, and significantly improving the credibility and practicality of the data quality coefficients.
[0031] (10) This invention designs encoding, chain building, chain uploading, and response recording as a standardized and verifiable evidence storage process. 2) By constructing a trusted channel based on network security services for data transmission, the security of the chain uploading process itself is ensured, preventing information from being stolen or tampered with during transmission. This process translates the abstract concept of "blockchain evidence storage" into concrete and operable technical steps, which is a key technical guarantee for achieving transparency and trustworthiness in the evaluation process.
[0032] (11) This invention also describes the collaboration between the system and the data asset valuation system, clarifying that the "data quality coefficient" generated by this system can be used as a key input parameter for the downstream data asset valuation system, thus opening up the value chain from "data quality evaluation" to "data asset valuation". By providing standardized data quality coefficient output, the asset valuation model can obtain a more objective and reliable underlying data quality basis, thereby calculating a more reasonable and more market-recognized data asset utility and valuation, promoting the value discovery and circulation of high-quality data elements. Attached Figure Description
[0033] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings, wherein... Figure 1 This is a framework diagram of the data quality evaluation system based on a large model in this invention; Figure 2 This is a logical relationship diagram of the data flow coefficient calculation system in this invention; Figure 3 This is a diagram illustrating the application layer's working mode in this invention. Figure 4 This is a flowchart of the application layer workflow in this invention. Detailed Implementation
[0034] To make the objectives, technical solutions, and positive effects of this invention clearer, the following will provide a detailed and complete description of a data quality evaluation system based on a large model, in conjunction with specific embodiments. Those skilled in the art should understand that the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.
[0035] This invention provides a data quality evaluation system based on a large model, aiming to solve the problems of strong subjectivity, low efficiency, and insufficient reliability of results in existing technologies. The core idea of this system is to construct a five-layer decoupled elastic architecture, integrate intelligent computing with collaborative large and small models with blockchain notarization technology, and achieve fully automated, high-precision, transparent and reliable industrial data quality evaluation.
[0036] refer to Figure 1 The data quality evaluation system based on a large model in this invention comprises, from bottom to top, a hardware layer, a service layer, an application layer, a business layer, and an interface layer. This layered and decoupled design ensures clear responsibilities and well-defined interfaces for each module, facilitating independent maintenance, expansion, and upgrades. For example, the algorithm in the service layer can be optimized or the business logic in the application layer can be replaced without altering other layers, significantly improving the system's maintainability and scalability. Furthermore, this architecture lays a solid foundation for subsequent implementation of automated and intelligent evaluation processes and transparent and reliable result traceability.
[0037] As the physical foundation of the entire system, the hardware layer uses virtualization technologies (such as KVM, VMware, or Docker container technology) to abstract, pool, and uniformly manage heterogeneous hardware resources such as network devices (switches, routers), computing devices (CPU / GPU server clusters), storage devices (disk arrays, distributed storage), and security devices (firewalls, WAFs). This constructs four logical resource pools: communication, computing, storage, and security. This virtualization and resource pooling solution provides a flexible, on-demand allocation foundation, enabling dynamic resource scheduling based on the load of upper-layer applications. This ensures the system can efficiently and reliably handle large-scale, high-concurrency data quality evaluation tasks in industrial applications, fundamentally improving resource utilization and system reliability.
[0038] The service layer is deployed above the hardware layer. By encapsulating and scheduling the underlying virtualized resources, it provides stable and reusable basic service capabilities for upper-layer applications. This layer mainly includes: network communication services, intelligent computing services, data storage services, information security services, digital wallet services, and basic general services such as log monitoring.
[0039] Intelligent computing service is one of the core services of this system. It is configured to utilize the underlying computing resource pool to specifically handle model-related computational tasks. Its key function is to adaptively train a pre-trained general-purpose data quality evaluation model obtained from an external large model platform (such as a high-quality dataset construction platform) using model compression and adaptation techniques such as transfer learning or knowledge distillation. For example, a large model trained on general text data can be quickly fine-tuned (through transfer learning) or refined (through knowledge distillation) into a lightweight, high-precision domain-specific small model using a small amount of labeled data from a specific industrial field (such as equipment manufacturing or process industries). This process achieves synergy between large and small models, inheriting the powerful generalization and feature extraction capabilities of the large model while gaining the advantages of rapid inference and low-cost deployment of the small model. This enables "better, better, and faster" high-precision quality evaluation of datasets of different domains and sizes.
[0040] Digital wallet services are a core component in building a trusted evaluation mechanism. They are configured to manage digital identities, key pairs, and other data required for interaction with the blockchain platform and are interconnected with smart contracts deployed on the blockchain. This service provides a secure and standardized technical channel for the application layer to trace and upload critical evaluation information on the blockchain, ensuring the feasibility of subsequent operations.
[0041] The application layer, deployed above the service layer, is the core layer for executing business logic and is directly responsible for the main functions of data quality evaluation and evidence storage. It includes a quality assessment module and an encoding-on-chain module. The quality assessment module reduces human intervention and the influence of subjective factors. Configured to invoke the intelligent computing service of the service layer based on task instructions issued by the business layer, the quality assessment module uses an adaptively trained evaluation model to automatically and intelligently evaluate multiple quality indicators of the data to be evaluated, ultimately generating a quantified data quality coefficient. The encoding-on-chain module systematically introduces blockchain evidence storage technology. It is configured to invoke the digital wallet service of the service layer to encode and encrypt key information in the quality assessment process (such as input parameters, model version, sampling method, intermediate results, final coefficients, etc.) according to predetermined rules, and then upload it to the smart contract. This allows the evaluation process and conclusions to be traceable by all relevant parties, ensuring transparency and credibility of the results, and providing an immutable chain of evidence for the valuation of data assets.
[0042] The business layer is deployed above the application layer and is primarily responsible for user interaction and business process management. It uses front-end frameworks such as Vue.js to build the user interface and includes system management, configuration management, and evaluation management modules. The system management module is used to uniformly manage the connection information and configuration parameters of the system itself and all external platforms it connects to.
[0043] The configuration management module provides a user-friendly graphical interface, allowing users to flexibly configure evaluation tasks, including: the selection of the evaluation model, information about the data source being evaluated (such as database connection strings), and data extraction rules. This significantly lowers the technical barrier to entry for using the system.
[0044] The evaluation management module receives business requests from users and automatically selects models and data samples to be evaluated based on the settings of the configuration management module. It then calls the lower-level application layer to perform the specific quality evaluation and coefficient calculations. This achieves fully automated management of the entire process, from task configuration and resource scheduling to computation invocation. Users only need to initiate a request; all subsequent steps are completed automatically by the system, resulting in extremely high work efficiency.
[0045] The interface layer is deployed at the top of the system, serving as a bridge for interaction between the system and the external ecosystem. Based on the network communication services provided by the service layer, it connects with various external platforms through multiple pre-defined communication protocols. This invention's system is designed to interact with multiple external systems, such as high-quality dataset construction platforms, blockchain platforms, evaluated data systems, data asset management platforms, and system operation and maintenance platforms. This design ensures that the system is not an isolated tool but can be integrated into the existing data technology ecosystem, achieving integrated collaboration in model acquisition, data access, asset valuation, and operation and maintenance monitoring, thus enhancing the system's practicality and application value. The pre-defined communication protocols include, but are not limited to, HTTPS, API, gRPC, and SDK, ensuring efficient, stable, and secure data interaction and service calls between the system and various heterogeneous external platforms, enhancing the system's compatibility and integration capabilities.
[0046] The application layer is configured to operate according to the following logic: S1. After receiving the business request, configure the parameters; S2. Determine if a corresponding model exists: If it exists, then determine the model; if it does not exist, then obtain the data quality evaluation model from the high-quality dataset construction platform, optimize it, and then determine the model. S3. Extract sample data from the data system being evaluated; S4. Calculate the data quality coefficient of the sample data based on the model driven in step S2 to obtain the data quality coefficient; S5. Perform quality evaluation coding on the associated blockchain platform on the data quality coefficient; S6. Feedback business requests based on the coding results and complete the data quality evaluation.
[0047] The calculation process for the data quality coefficient is as follows: A. Parameter Configuration: System parameters are set using CLI or Web models, including the industry attributes of the data being evaluated, the total amount of data being evaluated (C), and the data structure (S). B. Model Selection: Analyze the parameter attributes of the data being evaluated. The system uses a labeling method to select a suitable evaluation model. If the corresponding model does not exist, a suitable data quality evaluation model is obtained from a high-quality dataset platform, and the evaluation model is obtained after optimization of a small model. C. Sample extraction model calculation: Based on the total amount of data, data type, and data structure, the sample extraction model is selected according to the following formula:
[0048] Where R: Sample extraction model C: Total amount of data CNT: Total Data Resource Boundary Point SUM(S): Types of data structures. N: The dividing point between different types of data structures. Ri: The i-th type of stochastic model; The parameters N and CNT are empirical parameters, related to the data industry classification. These two parameters vary greatly across different industries and are set by industry experts during system initialization. During system operation, they can be dynamically adjusted based on feedback from evaluation results. D. Extracting sample data: Extracting the data resources to be evaluated according to the selected random sampling model, obtaining sample data resources with a sample size of N, and sending them to the intelligent computing service. E. Selection of data quality indicators for intelligent computing services: Based on the industry to which the data resources belong, sample data, and other information, use the optimized calculation model to configure the data quality coefficient calculation indicator system and the weight value Xi of each indicator; F. Calculation of coefficients for each indicator system by intelligent computing services: The intelligent computing service relies on the evaluation model to statistically analyze the total amount of data to be evaluated and the data resources that meet the indicator requirements, and calculates the coefficients for each indicator system according to the following formula. : ; G. Intelligent computing service calculates data quality coefficient: The data quality coefficient α is calculated according to the following formula:
[0049] Where α: data quality coefficient in data resources. n: The number of data quality indicators selected i: The i-th data quality evaluation index Xi: The weight value of the i-th data quality evaluation indicator. : The coefficient of the i-th data quality evaluation indicator.
[0050] The steps for quality assessment coding are as follows: a. Data Encoding: Information such as the industry of the data, data sample parameters, evaluation model parameters, evaluation time, and evaluation results during the data quality calculation process are encoded according to specified rules. The specified rules are related to smart contracts and blockchain platforms. To adapt to different blockchain platforms, a mode in which multiple encoding rules coexist is adopted. b. Construct a secure channel: Connect the system to the smart contract address and construct a trusted channel between the system and the smart contract based on network security services and network communication services; that is, construct a trusted channel with the blockchain node based on the TLS1.3 protocol, and ensure the security of the communication link by carrying out localization of the TLS1.3 protocol and using national cryptographic algorithms. c. Data upload to the blockchain: The encoded data quality evaluation information is sent to the blockchain platform via a secure channel, and the upload result response is received and recorded.
[0051] Data quality coefficients calculated at the application layer Upon request from the data asset assessment system, the data quality coefficient is fed back to the data asset assessment system. The data asset assessment system then uses the following calculation model to obtain the data utility of the data asset, and thus obtains the final data asset valuation: U= (1+l)(1-r); Where U: data asset valuation, α: Data quality coefficient β: Data flow coefficient l: Data monopoly coefficient r: Risk coefficient for realizing data value.
[0052] The following describes how the system of the present invention works collaboratively, using a specific industrial data quality evaluation scenario, to demonstrate its complete, closed-loop automated workflow and self-learning and adaptive capabilities.
[0053] For example, using carrier communication data as an example, briefly illustrate the workflow of this system: 1. The administrator inputs the industry classification, total data volume, data application scope, data structure, data update frequency, data source and other tag information of the communication data to be evaluated, as well as the address information of the data system to be evaluated and the address information of the data asset evaluation system; 2. Based on the tag information, the system calculates the matching degree between the data to be evaluated and the evaluation model. If there is an evaluation model with a matching degree greater than δ, the model M is selected as the data quality evaluation model for the data to be evaluated, and the process proceeds to step 5. If there is no evaluation model with a matching degree greater than δ, the next step is performed. The matching threshold δ can be configured according to the business requirements for evaluation accuracy and efficiency. For example, in scenarios requiring high accuracy, a higher δ value (e.g., above 0.75) can be set; in scenarios requiring rapid response, a lower δ value (e.g., around 0.6) can be set. Its specific value can be optimized and calibrated based on the feedback effects of historical evaluation tasks. 3. Connect to a high-quality dataset platform, send the label information of the communication data to be evaluated to the platform, and obtain the appropriate data quality evaluation model M1; 4. Perform knowledge distillation on the evaluation model M1 to train a local evaluation model M with a simpler structure that is more suitable for evaluating the quality of communication data, and save it. 5. After determining the evaluation model, the system selects different CNT and N values according to the data types of the data to be evaluated, and reasonably selects the random data extraction model R according to the data structure and data volume. 6. Connect the system to be evaluated using the address information and random data extraction model R, and extract sample data D. 7. Send the sample data D to the local intelligent evaluation model M, and calculate the data quality index α according to the following procedure: 7.1 Based on the evaluation model, determine the evaluation indicators and indicator weights for the communication data to be evaluated; 7.2 Calculate the index of each indicator for the sample data D respectively; 7.3 Calculate the communication data quality index α based on the weights of each index within each index level; 8. Repeat steps 6-7 three times to calculate the variance Δ of the three obtained quality indices; 9. If the Δ value is greater than 0.05, fine-tune the CNT and N values, reselect the random data extraction model R, and then repeat steps 6-8 to obtain a data quality index α with a delta value less than 0.05. 10. Disconnect from the communication data system being evaluated; 11. Encode information such as data tagging, data evaluation model, data evaluation time, data evaluation structure, and data system address according to blockchain rules; 12. Use the TLS protocol modified with Chinese national cryptography to link with the smart contract address and build a secure and trusted channel; 13. Send the encoded data quality evaluation information to the smart contract, and disconnect from the smart contract upon successful transmission; 14. Using the same TLS protocol modified by the national cryptographic standard, connect to the data asset assessment system to build a secure and trusted channel. After sending the data quality index to the data asset assessment system, disconnect the connection between them.
[0054] When a user initiates a quality evaluation request for a database of a smart manufacturing production line through the business layer front end, the system operates according to the following logic: S1. Parameter Configuration: Users or the system initialize task parameters through the CLI command line or Web interface, including specifying the industry attributes of the data being evaluated (such as "discrete manufacturing - assembly line"), the total amount of data C, and the data structure S (such as tables, time series data, images, etc.).
[0055] S2. Model Decision and Acquisition: The system first determines whether a pre-built evaluation model matching the current industry and data characteristics exists in the local model library. If it exists, the model is loaded directly. If not, the system's self-learning and adaptive capabilities are demonstrated: the system automatically retrieves a basic large-scale data quality evaluation model from the high-quality dataset construction platform through the interface layer, and immediately utilizes the intelligent computing services of the service layer, combined with the limited configuration information or sample data of the current task, to perform rapid transfer learning optimization, thereby obtaining a lightweight "small model" adapted to the current task. This ensures the feasibility of the system's evaluation when facing new fields and new types of data.
[0056] S3. Intelligent Sampling: Based on the total data volume C and data structure complexity SUM(S) configured in step S1, the system automatically selects a sample extraction model using a scientific and formulaic method to balance evaluation efficiency and result representativeness. Specifically, the selection follows the logical formula below:
[0057] Where R represents the final selected sample extraction model; CNT and N are preset cutoff points (e.g., CNT = 100 million samples, N = 10 types); R1, R2, and R3 represent different random sampling algorithms (e.g., R1 is stratified random sampling, R2 is systematic sampling, and R3 is simple random sampling). After selecting the model, the system automatically extracts N samples from the data system being evaluated.
[0058] S4. Automated Quality Assessment and Coefficient Calculation: This is the core calculation step of this invention, minimizing the interference of human judgment. Sample data and the selected model are submitted to the intelligent computing service. The specific calculation process is as follows: Indicator and Weight Configuration (Step E): Based on the industry and sample characteristics of the data, the system automatically configures a set of data quality evaluation indicators and the weight values Xi for each indicator using an optimized calculation model. These indicators construct a comprehensive and systematic evaluation indicator system covering standardization, completeness, accuracy, consistency, timeliness, and accessibility, ensuring a comprehensive evaluation.
[0059] Indicator coefficient calculation (step F): The intelligent computing service runs the evaluation model, counts the amount of data in the sample that meets the requirements of each indicator, and calculates the coefficient of each indicator according to the following formula. : ; Composite coefficient synthesis (step G): Finally, the data quality coefficient α is calculated according to the following formula: ; Where α: data quality coefficient in data resources. n: The number of data quality indicators selected i: The i-th data quality evaluation index Xi: The weight value of the i-th data quality evaluation indicator. : The coefficient of the i-th data quality evaluation indicator.
[0060] This model quantifies complex quality assessments into an objective, consistent, and repeatable coefficient, significantly improving the credibility and practicality of the assessment results.
[0061] S5. Trusted Evidence Storage (Encoding on the Chain): To solidify the evaluation process and give it credibility, the system immediately initiates the encoding and chain-based process, which is a standardized and verifiable evidence storage process. Data Encoding: The on-chain encoding module serializes the key metadata of this evaluation (such as industry, sample parameters, model hash, evaluation time, final coefficient α, etc.) in JSON or Protobuf format and calculates its hash value.
[0062] b. Construct a secure channel: The module calls the digital wallet service to obtain identity credentials and constructs a trusted channel between the smart contract address and the network security service (such as the TLS protocol) to prevent theft or tampering during transmission.
[0063] c. Data On-Chain: The encoded and hashed data packet is sent to the blockchain network through a secure channel, triggering a pre-set smart contract to permanently record the data hash on the chain. The system receives and records the transaction receipt (TxID) returned by the blockchain. This process translates the abstract concept of "blockchain notarization" into concrete operations, providing a key technological guarantee for achieving transparency and trustworthiness in the evaluation process.
[0064] S6. Result Feedback and Value Transmission: The application layer returns the calculated data quality coefficient α and the blockchain transaction receipt to the business layer. The business layer then displays the results to the user, completing an end-to-end automated quality assessment.
[0065] Furthermore, the "data quality coefficient" α generated by this system can serve as a key input parameter for downstream data asset valuation systems. When a data asset valuation system initiates a request through the interface layer, this system can feed back the coefficient α. The data asset valuation system can then incorporate it into its own valuation model (e.g., as the basis for calculating the data asset valuation U: U = ...). (1+l)(1-r), thus calculating a more reasonable and market-recognized value of data assets. This opens up the value chain from "data quality evaluation" to "data asset valuation", and promotes the value discovery and circulation of high-quality data elements by providing objective and credible underlying quality basis.
[0066] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely specific embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A data quality evaluation system based on a large model, characterized in that, include: The hardware layer consists of a virtualized pool of communication, computing, storage, and security resources, formed by building network devices, computing devices, storage devices, and security devices. The service layer, deployed above the hardware layer, includes network communication services, intelligent computing services, data storage services, information security services, digital wallet services, and basic general services. The application layer, deployed above the service layer, includes a quality assessment module and an on-chain coding module; The business layer, deployed above the application layer, includes a system management module, a configuration management module, and an evaluation management module; The interface layer, deployed above the business layer, is configured to connect to external platforms based on the network communication service.
2. The data quality assessment system based on a large model according to claim 1, characterized in that: In the service layer, the intelligent computing service is configured to adaptively train the data quality evaluation model trained on the large model platform using transfer learning or knowledge distillation; the digital wallet service is configured to connect with smart contracts.
3. The data quality assessment system based on a large model according to claim 1, characterized in that: In the application layer, the quality assessment module is configured to call the intelligent computing service, use the selected assessment model to evaluate multiple quality indicators of the data to be assessed, and generate a data quality coefficient. The encoding and on-chain module is configured to call the digital wallet service to encode and encrypt the relevant indicators of the quality assessment process and then upload them to the smart contract.
4. The data quality evaluation system based on a large model according to claim 3, characterized in that: The aforementioned quality indicators include data standardization, completeness, accuracy, consistency, timeliness, and accessibility.
5. The data quality evaluation system based on a large model according to claim 1, characterized in that: In the business layer, the system management module is configured to manage the external platform information connected to the system and the system's own configuration information; The configuration management module is configured to configure the selection of the evaluation model, the information of the evaluated data, and the extraction rules of the evaluated data; The evaluation management module is configured to select an evaluation model and a data sample to be evaluated according to the configuration of the configuration management module, and call the application layer to perform quality evaluation and quality coefficient calculation.
6. The data quality evaluation system based on a large model according to claim 1, characterized in that: The interface layer is configured to connect to one or more of the following via a preset communication protocol: a high-quality dataset construction platform, a blockchain platform, an evaluated data system, a data asset management platform, and a system operation and maintenance platform.
7. The data quality assessment system based on a large model according to claim 6, characterized in that: The preset communication protocols include HTTPS, API, gRPC, and SDK.
8. The data quality evaluation system based on a large model according to claim 3, characterized in that: The application layer is configured to operate according to the following logic: S1. After receiving the business request, configure the parameters; S2. Determine if a corresponding model exists: If it exists, then determine the model; if it does not exist, then obtain the data quality evaluation model from the high-quality dataset construction platform, optimize it, and then determine the model. S3. Extract sample data from the data system being evaluated; S4. Calculate the data quality coefficient of the sample data based on the model driven in step S2 to obtain the data quality coefficient; S5. Perform quality evaluation coding on the associated blockchain platform on the data quality coefficient; S6. Feedback business requests based on the coding results and complete the data quality evaluation.
9. The data quality evaluation system based on a large model according to claim 8, characterized in that: The calculation process for the data quality coefficient is as follows: A. Parameter Configuration: System parameters are set using CLI or Web models, including the industry attributes of the data being evaluated, the total amount of data being evaluated (C), and the data structure (S). B. Model Selection: Analyze the parameter attributes of the data being evaluated. The system uses a labeling method to select a suitable evaluation model. If the corresponding model does not exist, a suitable data quality evaluation model is obtained from a high-quality dataset platform, and the evaluation model is obtained after optimization of a small model. C. Sample extraction model calculation: Based on the total amount of data, data type, and data structure, the sample extraction model is selected according to the following formula: Where R: sample extraction model, C: Total amount of data CNT: Total Data Resource Boundary Point SUM(S): Types of data structures. N: The dividing point between different types of data structures. Ri: The i-th type of stochastic model; D. Extracting sample data: Extracting the data resources to be evaluated according to the selected random sampling model, obtaining sample data resources with a sample size of N, and sending them to the intelligent computing service. E. Selection of data quality indicators for intelligent computing services: Based on the industry to which the data resources belong, sample data, and other information, use the optimized calculation model to configure the data quality coefficient calculation indicator system and the weight value Xi of each indicator; F. Calculation of coefficients for each indicator system by intelligent computing services: The intelligent computing service relies on the evaluation model to statistically analyze the total amount of data to be evaluated and the data resources that meet the indicator requirements, and calculates the coefficients for each indicator system according to the following formula. : ; G. Intelligent computing service calculates data quality coefficient: The data quality coefficient α is calculated according to the following formula: ; Where α: data quality coefficient in data resources. n: The number of data quality indicators selected i: The i-th data quality evaluation index Xi: The weight value of the i-th data quality evaluation indicator. : The coefficient of the i-th data quality evaluation indicator.
10. The data quality assessment system based on a large model according to claim 8, characterized in that: The steps for the quality evaluation coding are as follows: a. Data Coding: Encode the industry to which the data belongs, data sample parameters, evaluation model parameters, evaluation time, and evaluation results in the data quality calculation process according to specified rules; b. Construct a secure channel: Connect the system to the smart contract address and construct a trusted channel between the system and the smart contract based on network security services and network communication services; c. Data upload to the blockchain: The encoded data quality evaluation information is sent to the blockchain platform via a secure channel, and the upload result response is received and recorded.
11. The data quality assessment system based on a large model according to claim 9, characterized in that: Data quality coefficients calculated at the application layer Upon request from the data asset assessment system, the data quality coefficient is fed back to the data asset assessment system. The data asset assessment system then uses the following calculation model to obtain the data utility of the data asset, and thus obtains the final data asset valuation: U= (1+l)(1-r); Where U: data asset valuation, α: Data quality coefficient β: Data flow coefficient l: Data monopoly coefficient r: Risk coefficient for realizing data value.