Intelligent data quality evaluation system and method based on multi-modal analysis and dynamic baseline prediction

Through intelligent data quality evaluation methods of multimodal analysis and dynamic baseline prediction, multiple limitations of data quality evaluation in the prior art are solved, efficient evaluation of unstructured data and time-series data and cross-system consistency verification are realized, and evaluation efficiency and accuracy are significantly improved.

CN120196916AInactive Publication Date: 2025-06-24WUHAN DAJIANG DATA CO LTD

Patent Information

Application Number
CN202510258295.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-24
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art has problems such as single-dimensional evaluation limitations, missing unstructured data evaluation, low detection accuracy of time-series data fluctuations, and low cross-system consistency verification efficiency in data quality evaluation.

Method used

The intelligent data quality evaluation method of multimodal analysis and dynamic baseline prediction is adopted. Different types of data are received through the multimodal data access module, and the dimensional routing allocation engine selects the appropriate evaluation path, uses the large model analysis core to perform semantic verification, feature fusion analysis and dynamic baseline prediction, and cross-system comparison is realized through the distributed consistency verification protocol, and finally generates a quality report of trustworthiness traceability.

Benefits of technology

A comprehensive evaluation of unstructured data is achieved, the accuracy of time series data detection is improved, the efficiency of cross-system consistency verification is significantly improved, and the maintenance cost of data quality management is reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120196916A_ABST
    Figure CN120196916A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent data quality evaluation system and method based on multi-modal analysis and dynamic baseline prediction, and belongs to the technical field of data quality management. According to the invention, a system architecture comprising a multi-modal data access module, a dimension routing distribution engine, a large model analysis core and a visual report generator is constructed. Structured, unstructured and time series data can be accessed into the system, after metadata acquisition, evaluation paths are distributed according to data types, quality analysis is executed by a large model analysis core, and finally, a visualization report generator outputs a multi-dimensional quality scoring matrix and a root cause analysis report, so that the comprehensiveness, accuracy and efficiency of data quality evaluation are effectively improved, and the data quality evaluation efficiency is improved. And the method is suitable for heterogeneous data management scenes of multiple industries such as finance, medical treatment, e-commerce and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data quality management, and particularly relates to an intelligent data quality evaluation method and system integrating multi-modal analysis, large language models, and distributed computing technologies. Background Art

[0002] In the field of data quality management, with the explosive growth of data volume and the increasing richness of data types, traditional data quality evaluation technologies are facing numerous challenges. Currently, there are various data quality evaluation solutions on the market, but they all have limitations to varying degrees.

[0003] Single-dimension evaluation limitation: Some existing technologies only support the evaluation of structured data, such as CN114003710A, and cannot handle unstructured data. In practical applications, unstructured data widely exists, such as image reports in the medical field, text contracts in the financial field, etc. These data contain important information, but due to the limitations of evaluation technologies, their quality cannot be effectively evaluated.

[0004] Missing evaluation of unstructured data: For example, US20220365925A1 does not cover the quality detection of text / images, which leads to evaluation blind spots in scenarios such as medical image reports. The text descriptions and image information in medical image reports are crucial for diagnosis. If their quality cannot be evaluated, it may affect the accuracy of diagnosis.

[0005] Low accuracy in detecting fluctuations of time-series data: WO2022156680A1 uses a fixed threshold method to detect fluctuations of time-series data, and the false alarm rate is >35% in the scenario of financial transaction records. The data of financial transaction records fluctuates frequently, and the fixed threshold method is difficult to adapt to complex changes, easily generating a large number of false alarms and affecting the judgment of real abnormal situations.

[0006] Low efficiency in cross-system consistency verification: Traditional cross-system consistency verification methods require full-scale data comparison, and it takes more than 12 hours to verify in a 100-node cluster. In the case of a large amount of data, full-scale comparison not only consumes a large amount of resources but also has extremely low efficiency, unable to meet business scenarios with high real-time requirements. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to overcome the defects of the prior art and provide an intelligent data quality evaluation system and method based on multi-modal analysis and dynamic baseline prediction.

[0008] To solve the above technical problems, the present invention provides the following technical solutions:

[0009] An intelligent data quality evaluation method based on multi-modal analysis and dynamic baseline prediction of the present invention includes the following steps:

[0010] a. The multi-modal data access module receives structured data, unstructured data, and time-series data;

[0011] b. In the dimension routing and allocation engine, select the corresponding evaluation path according to the data type:

[0012] c. Structured data triggers semantic verification of the large model;

[0013] d. Unstructured data starts multi-modal feature fusion analysis;

[0014] e. Time-series data loads the dynamic baseline prediction model;

[0015] f. Implement cross-system data comparison through the distributed consistency verification protocol;

[0016] g. Generate a visual quality report containing the credibility traceability score.

[0017] The present invention also discloses an intelligent data quality evaluation system for the method described in claim 1, including:

[0018] The multi-modal data access unit is configured to be compatible with database table, text, image, and log stream data input;

[0019] The dimension routing and allocation unit has a built-in meta-data analysis algorithm to dynamically allocate evaluation paths;

[0020] The large model analysis core unit integrates the following sub-modules:

[0021] The unstructured data semantic integrity evaluation module;

[0022] The time-series dynamic baseline prediction module based on the Temporal Convolutional Network (TCN);

[0023] The distributed consistency verification engine.

[0024] As a preferred technical solution of the present invention, the multi-modal feature fusion analysis includes:

[0025] Perform entity recognition (NER) and semantic coherence verification on text data;

[0026] Extract key region features from image data and perform integrity detection;

[0027] Generate a comprehensive score through a weighted fusion algorithm:

[0028] Comprehensive score = 0.6 text score + 0.4 image score.

[0029] As a preferred technical solution of the present invention, the dynamic baseline prediction model adopts the following technology:

[0030] Extract temporal features using a Temporal Convolutional Network (TCN);

[0031] Predict the fluctuation baseline for future time windows based on the attention mechanism;

[0032] Determine abnormal periods through the residual 3σ principle, and the abnormal determination formula is:

[0033] residual = |actual value - predicted value| > 3 * standard deviation.

[0034] As a preferred technical solution of the present invention, the distributed consistency verification protocol includes:

[0035] Generate a data digest through the feature fingerprint hashing algorithm;

[0036] Adopt differential comparison technology to locate inconsistent fields;

[0037] Verification efficiency optimization formula: the time consumption T = O(log N), where N is the number of nodes.

[0038] As a preferred technical solution of the present invention, the credibility traceability scoring model includes:

[0039] Data lineage integrity verification sub-module;

[0040] Change record auditing sub-module;

[0041] Access log analysis sub-module;

[0042] Authentication information verification sub-module;

[0043] Scoring calculation formula:

[0044] Credibility = 0.4 lineage score + 0.3 change score + 0.2 access score + 0.1 authentication score.

[0045] As a preferred technical solution of the present invention, the distributed consistency verification engine supports the following alternative solutions:

[0046] Use blockchain technology to achieve audit tracing;

[0047] Adopt a lightweight consensus algorithm to optimize throughput;

[0048] Deploy local verification agents at edge computing nodes.

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0050] 1. Breakthrough in the evaluation ability of unstructured data of the present invention: the evaluation coverage rate of medical image reports has increased from 0% to 88%, which can more comprehensively evaluate the quality of unstructured data and provide more accurate data support for related industries.

[0051] 2. The timing detection accuracy of the present invention is significantly improved: the false alarm rate in the financial transaction flow scenario is reduced by 26.3 percentage points, effectively reducing false alarms and improving the detection accuracy of abnormal timing data.

[0052] 3. The verification efficiency of the present invention is exponentially improved: the time-consuming for cross-system consistency verification is shortened by 15 times, greatly improving the efficiency of cross-system data consistency verification and meeting the business requirements with high real-time requirements.

[0053] 4. The comprehensive maintenance cost of the present invention is reduced: the manpower requirement for rule maintenance is reduced by 70%, reducing the manpower cost of the enterprise in data quality management. Description of the Drawings

[0054] The drawings are used to provide a further understanding of the present invention, and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0055] Figure 1 : System architecture diagram;

[0056] Figure 2 : Dimension routing allocation flowchart;

[0057] Figure 3 : Unstructured data evaluation model structure diagram;

[0058] Figure 4 : Timing dynamic baseline prediction algorithm flowchart;

[0059] Figure 5 : Distributed consistency verification protocol timing diagram. Detailed Embodiments

[0060] The following describes the preferred embodiments of the present invention with reference to the drawings. It should be understood that the preferred embodiments described herein are only used to explain and illustrate the present invention and are not used to limit the present invention.

[0061] Embodiment 1

[0062] As Figures 1-5 shown, the present invention provides an intelligent data quality evaluation system and method based on multimodal analysis and dynamic baseline prediction, including:

[0063] System Architecture

[0064] Multimodal data access module: supports the unified access of structured data (database tables), unstructured data (text / images), and timing data (log streams), and realizes the fusion processing of various types of data.

[0065] Dimension Routing Allocation Engine: Automatically allocate evaluation paths based on metadata analysis. By analyzing the metadata of data, determine the data type, and then select the most suitable evaluation path to improve the pertinence and efficiency of evaluation.

[0066] Large Model Analysis Core: Integrate multiple key technologies, including unstructured data semantic integrity evaluation, time-series data dynamic baseline prediction, and cross-system consistency distributed verification. Utilize the powerful capabilities of large models to deeply analyze different types of data and improve the accuracy of evaluation.

[0067] Visualization Report Generator: Generate a multi-dimensional quality scoring matrix and root cause analysis report. Present the data quality evaluation results in an intuitive visual way, facilitating users to quickly understand the data quality status and deeply analyze the root causes of problems.

[0068] Workflow

[0069] Metadata Collection Phase: Extract data lineage, field constraints, etc. through the API gateway to provide basic information for subsequent evaluation.

[0070] Dimension Routing Selection Phase: The dimension routing engine selects the evaluation path according to the data type. Structured data triggers large model semantic verification, unstructured data starts multi-modal feature extraction, and time-series data loads the dynamic baseline prediction model.

[0071] Quality Analysis Phase: The large model core performs quality analysis and outputs scores for each dimension. For different types of data, corresponding technologies are used for analysis to obtain accurate quality scores.

[0072] Report Generation Phase: The visualization module aggregates the results and generates an interactive report for easy viewing and analysis by users.

[0073] When specifically used:

[0074] (I) Implementation of the Multi-modal Data Access Module

[0075] By developing a unified data access interface, support the input of multiple data formats. For structured data, use database connection technology to read data; for unstructured data, use file reading and image recognition technologies to collect data; for time-series data, use log collection tools to obtain it in real time.

[0076] (II) Working Principle of the Dimension Routing Allocation Engine

[0077] After the metadata collection is completed, parse the metadata to extract key information such as data type and data structure. According to the pre-set rules, judge the data type, such as structured data, unstructured data, or time-series data, and allocate it to the corresponding evaluation path.

[0078] (3) Operating mechanism of the core of large model analysis

[0079] Semantic integrity evaluation of unstructured data: Use text entity recognition (NER) technology to identify entities in text data, and verify the logical coherence of the text through semantic coherence analysis algorithms; for image data, use image key region detection algorithms to extract key region features and perform integrity detection. Finally, generate a comprehensive score through a weighted fusion algorithm.

[0080] Dynamic baseline prediction of time series data: Use a temporal convolutional network (TCN) to extract features from time series data, predict the fluctuation baseline of future time windows based on the attention mechanism, and determine abnormal periods through the residual 3σ principle.

[0081] Cross-system consistency distributed verification: Generate a data digest through a feature fingerprint hashing algorithm, and use differential comparison technology to locate inconsistent fields to improve verification efficiency.

[0082] (4) Implementation of the visualization report generator

[0083] Integrate the scores of each dimension output by the core of the large model analysis to generate a multi-dimensional quality score matrix. Use data visualization technology to present the score matrix and the root cause analysis results to users in the form of charts, reports, etc., to achieve interactive operations.

[0084] (5) Application of the data credibility traceability scoring model

[0085] Obtain corresponding data through the data lineage integrity verification sub-module, change record audit sub-module, access log analysis sub-module, and authentication information verification sub-module, and calculate the credibility score according to the formula: credibility = 0.4 * lineage integrity + 0.3 * change record + 0.2 * access log + 0.1 * authentication information.

[0086] Finally, it should be noted that the above are only preferred embodiments of the present invention and are not used to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An intelligent data quality assessment method based on multimodal analysis and dynamic baseline prediction, characterized in that: The following steps are involved: a. The multi-modal data access module receives structured data, unstructured data and time series data; b. In the dimension routing distribution engine, select the corresponding evaluation path according to the data type: 1) Structured data triggers semantic verification of large models; 2) Unstructured data initiates multimodal feature fusion analysis; 3) Time series data loading dynamic baseline prediction model; c. Realize cross-system data comparison through distributed consistency verification protocol; d. Generate a visual quality report including a credibility traceability score.

2. An intelligent data quality assessment system for the method of claim 1, characterized in that: include: A multimodal data access unit configured to be compatible with database table, text, image and log stream data input; Dimensional routing allocation unit, with built-in metadata analysis algorithm, dynamically allocates evaluation paths; The core unit of large model analysis integrates the following submodules: Unstructured data semantic integrity assessment module; A temporal dynamic baseline prediction module based on a temporal convolutional network (TCN); Distributed consistency verification engine.

3. The method according to claim 1, characterized in that The multimodal feature fusion analysis includes: Perform entity recognition (NER) and semantic coherence verification on text data; Extract key area features from image data and perform integrity detection; Generate a comprehensive score through a weighted fusion algorithm: Comprehensive score = 0.6 text score + 0.4 image score.

4. The method according to claim 1, characterized in that The dynamic baseline prediction model adopts the following technologies: Use the temporal convolutional network (TCN) to extract temporal features; Predict the fluctuation baseline of future time windows based on the attention mechanism; The abnormal period is determined by the residual 3σ principle, and the abnormal determination formula is: residual=|actual value-predicted value|>3*standard deviation.

5. The method according to claim 1, characterized in that The distributed consistency verification protocol includes: Generate data summary through feature fingerprint hash algorithm; Use differential comparison technology to locate inconsistent fields; Verification efficiency optimization formula: time consumption T = O (log N), N is the number of nodes.

6. The system according to claim 2, characterized in that The credibility tracing scoring model includes: Data lineage integrity verification submodule; Change record audit submodule; Access log analysis submodule; Authentication information verification submodule; Rating calculation formula: Credibility = 0.4 bloodline score + 0.3 change score + 0.2 access score + 0.1 certification score.

7. The system according to claim 2, characterized in that The distributed consistency verification engine supports the following alternatives: Use blockchain technology to enable audit trails; Use lightweight consensus algorithm to optimize throughput; Deploy local verification agents on edge computing nodes.

Citation Information

Patent Citations

  • Intelligent fusion and circulation method based on fraud early warning data

    CN114003710A

  • Control method, analysis device, and recording medium

    US20220365925A1

  • Circuit system and method

    WO2022156680A1

Cited By

  • Multi-modal data quality evaluation method based on deep learning

    CN120804084A

  • Data quality evaluation method and system based on rule engine and machine learning

    CN121030265A

  • Agricultural data quality evaluation method and system based on multi-source heterogeneous data fusion

    CN122221181A