Material data quality evaluation method based on large model Agent

Through the material data quality evaluation method based on the big model Agent, combined with the rule set, RAG technology, machine learning algorithms and intelligent agents, the problem of uneven material data quality is solved, efficient and intelligent data quality evaluation is achieved, and the modern material science needs for high-quality data is met.

CN119993341AActive Publication Date: 2025-05-13SHANGHAI UNIV

Patent Information

Application Number
CN202510072136.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-13
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

In modern materials science, the quality of material data is uneven, the data is insufficient in accuracy and reliability, and traditional data quality management methods are inefficient, making it difficult to meet the requirements of efficiency and accuracy.

Method used

The material data quality evaluation method based on the big model Agent is adopted, and comprehensive quality evaluation results are generated through the multi-dimensional evaluation of the cyclic detection mechanism based on rulesets, the RAG technology and LLMs model, and the fusion analysis of the intelligent Agent.

Benefits of technology

It significantly improves data processing efficiency, analysis depth and interpretability of evaluation results, provides intelligent and automated data quality evaluation solutions, ensuring the reliability and scientificity of material data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993341A_ABST
    Figure CN119993341A_ABST
Patent Text Reader

Abstract

The invention discloses a material data quality evaluation method based on a large model Agent, and relates to the technical field of material science and intelligent data management. According to the method, firstly, based on a rule set and through a real-time feedback and cyclic detection mechanism, original material data is corrected, and it is ensured that data entering follow-up evaluation has high quality; and secondly, based on the knowledge base, deep analysis is performed on the speciality, logic consistency and rationality of the to-be-evaluated material data by utilizing an RAG technology and an LLMs model, and a data quality evaluation result based on the knowledge base is generated. And meanwhile, various machine learning algorithms are used for carrying out multi-aspect detailed detection on the to-be-evaluated material data to obtain a plurality of data quality evaluation results based on the machine learning algorithms, and visual and quantitative evaluation results are provided through a visual method. And then, multi-dimensional evaluation information is integrated through a large model Agent, and a high-interpretation and structured comprehensive quality evaluation result is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of material science and intelligent data management, and in particular to a material data quality evaluation method based on a large model agent. Background Art

[0002] With the rapid development of materials science, data has become a key factor in driving innovation and breakthroughs. The amount of data generated by modern materials research is growing exponentially, and the data types are becoming increasingly diverse, covering a wide range of information such as the composition, structure, performance, and processing technology of materials. These data come from a wide range of sources, including experimental measurements, theoretical calculations, numerical simulations, and literature reports. Against the backdrop of this rapid growth in data, the problem of uneven data quality has become particularly prominent, as shown in the following aspects:

[0003] Insufficient data accuracy: Material data comes from a wide range of sources, including experimental measurements, simulation calculations, and literature extraction. However, the accuracy of these data is often affected by measurement errors, modeling biases, and inconsistencies in manual annotations, and cannot fully meet the needs of high-precision material research.

[0004] Low data reliability: Due to the diversity of data collection standards and methods, data from different sources vary significantly and there is a lack of a unified quality assurance mechanism, making it difficult for researchers to verify the credibility of data in practical applications.

[0005] Inefficient data quality assessment: Traditional data quality management methods mainly rely on manual inspection and simple rule constraints, which can be applied to small-scale data sets, but are inefficient when faced with large and complex data sets, and it is difficult to meet the requirements of modern materials science for efficiency and accuracy.

[0006] Traditional data quality management methods mainly rely on manual inspection and simple rule-based constraints. This method is effective in dealing with small-scale data sets, but it is powerless when faced with complex large-scale data sets in the field of modern materials science. Manual inspection is time-consuming and labor-intensive, and is easily interfered by subjective factors, while simple rule constraints are difficult to adapt to the diversity and dynamic characteristics of material data. In addition, due to the high degree of specialization of material data, non-professionals often make misjudgments when processing data due to lack of domain knowledge, and there are significant deficiencies in cross-domain data integration and adaptive data evaluation. These problems not only limit the application value of data, but also hinder the development of intelligent research in the field of materials science.

[0007] In view of this, there is an urgent need for an intelligent and automated data quality assessment technology that can adapt to the diversity and complexity of material data. This technology has efficient rule-based judgment capabilities, dynamic knowledge calibration capabilities, and intelligent analysis capabilities, thereby providing comprehensive guarantees for the reliability and scientificity of material data and meeting the urgent needs of modern material science research for high-quality data. Summary of the invention

[0008] The purpose of this application is to provide a material data quality evaluation method based on a large model agent to achieve automated analysis of the quality of material data with diversity, dynamic characteristics, and cross-domain integration.

[0009] To achieve the above objectives, this application provides the following solutions.

[0010] In a first aspect, the present application provides a material data quality evaluation method based on a large model agent, comprising:

[0011] Based on the rule set, the original material data is corrected using a cyclic detection mechanism to obtain the material data to be evaluated;

[0012] Based on the knowledge base, RAG technology and LLMs model are used to conduct quality assessment on the material data to be assessed, and obtain the data quality assessment results based on the knowledge base;

[0013] Combine multiple machine learning algorithms to conduct multi-dimensional evaluation of the material data to be evaluated, obtain multi-dimensional test result visualization images and multiple data quality evaluation results based on machine learning algorithms;

[0014] Based on the large model agent, the data quality assessment results based on the knowledge base, the visualization images of the multi-dimensional detection results and the data quality assessment results based on multiple machine learning algorithms are integrated and analyzed to obtain a comprehensive quality assessment result.

[0015] According to the specific embodiments provided in this application, this application has the following technical effects.

[0016] The present application provides a material data quality evaluation method based on a large model agent. First, the present application corrects the original material data based on a rule set and through a real-time feedback and loop detection mechanism to ensure that the data entering the subsequent evaluation has a high quality. Secondly, based on the knowledge base, the RAG technology and the LLMs model are used to conduct an in-depth analysis of the professionalism, logical consistency and rationality of the material data to be evaluated, and generate data quality evaluation results based on the knowledge base. At the same time, a variety of machine learning algorithms perform a variety of detailed inspections on the material data to be evaluated, and obtain multiple data quality evaluation results based on machine learning algorithms, and provide intuitive and quantitative evaluation results through visualization methods. Then, the multi-dimensional evaluation information is integrated through the large model agent to generate a highly interpretable and structured comprehensive quality evaluation result. Compared with traditional data quality management methods, this application significantly improves data processing efficiency, analysis depth and interpretability of evaluation results, and provides intelligent and automated innovative solutions for modern materials science research. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 A schematic diagram of a flow chart of a material data quality evaluation method based on a large model agent provided in one embodiment of the present application;

[0019] Figure 2 A schematic diagram of a material data quality evaluation method based on a large model agent provided in one embodiment of the present application;

[0020] Figure 3 A flowchart of rule definition and screening evaluation provided for an embodiment of the present application;

[0021] Figure 4 A flow chart of data point quality assessment based on a knowledge base provided in one embodiment of the present application;

[0022] Figure 5 A flowchart of data point quality assessment based on a machine learning algorithm provided in one embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, the present application is further described in detail below with reference to the accompanying drawings and specific implementation methods.

[0025] In an exemplary embodiment, Figure 1 As shown, a material data quality evaluation method based on a large model agent is provided, including the following steps 101 to 104.

[0026] Step 101 : Based on a rule set, a cyclic detection mechanism is used to modify the original material data to obtain the material data to be evaluated.

[0027] Step 102, based on the knowledge base, using RAG technology and LLMs model, quality assessment is performed on the material data to be assessed, and a data quality assessment result based on the knowledge base is obtained.

[0028] Step 103, combining multiple machine learning algorithms to perform multi-dimensional evaluation on the material data to be evaluated, and obtaining a multi-dimensional detection result visualization image and multiple data quality evaluation results based on machine learning algorithms.

[0029] Step 104, based on the large model Agent, a fusion analysis is performed on the data quality assessment results based on the knowledge base, the multi-dimensional detection result visualization images, and the data quality assessment results based on multiple machine learning algorithms to obtain a comprehensive quality assessment result.

[0030] In the above steps 101 to 104, firstly, by using the data preprocessing and rule screening module, users can flexibly define evaluation rules. The system automatically converts the rules into program codes through LLMs and realizes fully automated data screening and detection. Real-time feedback and status marking help users intuitively identify data quality problems and make corrections to ensure the accuracy of subsequent evaluations. Secondly, based on the knowledge base reasoning ability of LLMs, the system deeply analyzes the matching degree, logical consistency and rationality of data and domain knowledge, extracts key knowledge points, and provides support for evaluation. Thirdly, machine learning algorithms are applied to perform outlier detection, outlier identification, dimensionality reduction analysis and performance prediction, extract multidimensional features and generate preliminary quality evaluation results. Then, with the help of the multimodal understanding ability of intelligent agents, visualization results such as dimensionality reduction distribution maps and abnormal marking maps are analyzed to generate intuitive explanatory content. Finally, algorithm analysis, knowledge base reasoning and visualization analysis results are integrated to generate a comprehensive quality evaluation report covering problem location, quantitative indicators and optimization suggestions. This method effectively combines machine learning, knowledge base reasoning and multimodal analysis technology to provide an efficient and intelligent data quality evaluation solution, which greatly improves data optimization and decision-making efficiency.

[0031] The embodiment of the present application proposes a material data quality evaluation method based on a large model agent, aiming to break through the limitations of traditional methods in terms of evaluation singularity, domain knowledge fusion and result interpretation. This method innovatively integrates rule detection, knowledge base reasoning, machine learning analysis and multimodal visual analysis, and achieves unified fusion of multi-algorithm results through intelligent agents to generate comprehensive and intuitive comprehensive evaluation reports. In particular, the system accurately identifies data problems and provides optimization suggestions through knowledge reasoning and multimodal analysis capabilities driven by LLMs, significantly improving the intelligence, accuracy and decision support efficiency of data quality evaluation, and providing an efficient and innovative solution for quality assessment and optimization in complex data scenarios.

[0032] In another exemplary embodiment of the present application, Figure 2 As shown, the above step 101 can be replaced by the following step A, the above step 102 can be replaced by the following step B, the above step 103 can be replaced by the following step C, and the above step 104 can be replaced by the following step D.

[0033] Step A: In the data quality assessment application of this application, the raw data is preprocessed through the data processing layer, including feature engineering, missing value detection, and preset rule screening. In addition, the combination of feature engineering and loop detection mechanism realizes flexible and efficient rule management and comprehensive data quality detection. In this process, the preliminary screening results of the data will be fed back to the user in real time, and the user can judge whether to enter the complete data quality assessment process based on the feedback. This step ensures that the data entering the subsequent evaluation stage has a high initial quality, thereby effectively avoiding the subsequent waste of computing resources caused by low-quality data. Reference Figure 3 , a specific embodiment of step A specifically includes the following steps 201-204.

[0034] Step 201: Define and manage preset rule sets. In the preset rule evaluation module, the system allows users to define a variety of rule sets R preset Through large language models (LLMs), the rules R of the user’s natural language input preset is automatically converted into a programming language format to form a preset rule function f detect , and store it in the rule base, thus achieving efficient and flexible rule management. This process significantly improves the efficiency and flexibility of rule definition and supports dynamic rule management.

[0035] f detect =f LLM (R preset ,P rule )

[0036] Step 202: Rule-based automated detection. The system preprocesses the raw data S, including feature engineering operations to enrich data features, and automatically detects missing values ​​and outliers. Then, the system uses the rule set R preset Automatically detect quality issues in data, including missing values, range checks and other rule detection, and generate detection rule detection results R rule :

[0037] R rule =f detect (S)

[0038] Among them, f detect It is a preset rule detection function, which is used to detect the preset Perform quality inspection on the original data S. Specifically, f detect The data will be tested in multiple dimensions. Several common detection rules are shown in Table 1.

[0039] Table 1 Detection rules

[0040]

[0041]

[0042] Step 203: Real-time feedback and status marking. After the detection is completed, the system will result The result box displayed in the main operation interface. The status marking signals (such as green, yellow and red marks) provide intuitive feedback on the quality status of the current data, helping users to quickly identify and deal with problems. The marking signal rules are as follows: Green - data quality meets the preset rules; Yellow - minor problems that require user review; Red - serious problems that need to be corrected immediately.

[0043] Step 204: Data correction and loop detection mechanism. The user annotates and corrects the data S based on the feedback. If the user re-enters the corrected data S′, the corrected data is resubmitted to the preset rule evaluation module for a new round of detection:

[0044] R rule =f detect (S′)

[0045] This cyclic detection mechanism ensures that the corrected data complies with the requirements of the rules and provides a solid foundation for subsequent comprehensive quality assessment.

[0046] Step B: In the data quality assessment application of this application, the knowledge base assessment module based on large-scale language models (LLMs) combines the built-in knowledge base with the external material knowledge base, and uses RAG (Retrieval-Augmented Generation) technology to achieve efficient knowledge retrieval and reasoning, providing intelligent support for data quality assessment. By dynamically calling professional information in the domain knowledge base, the system can effectively detect unreasonable parts in the data and automatically report possible quality problems in the data, providing semantic support and accurate knowledge background for the interpretability of subsequent data quality assessment.

[0047] Reference Figure 4 , a specific implementation of step B specifically includes the following steps 301-304.

[0048] Step 301: Build and manage knowledge base. The module integrates the built-in LLMs knowledge base and the external material knowledge base uploaded by users, supporting the storage and management of multi-source heterogeneous data. en , for the text content T in the knowledge base know Encode and convert complex material knowledge into efficient retrieval structure V vec , to support subsequent calls and analysis.

[0049] V vec =f en (T know )

[0050] Step 302: Knowledge retrieval and data matching analysis. Based on RAG technology RAG , the system retrieves content C related to the target data S from the knowledge base re , and match and analyze the data content. know Under the guidance of LLM , verify the rationality of the data, identify potential quality issues, and obtain data quality assessment results based on the knowledge base R know .

[0051]

[0052] C re =Top k (sim(f en (S),V vec ))

[0053] R know =f LLM (S,C re ,P know )

[0054] Among them, f en (S) is the vector representation of the target data; V vec is the vector representation of each content in the knowledge base; sim(f en (S),V vec ) is the similarity score between the target data and the knowledge base content; Top k It means selecting the k pieces of content with the highest similarity to the target data S.

[0055] Step 303: Scoring and quantitative quality assessment. The system scores each piece of data S according to the analysis results of the knowledge base. score , generating quantitative quality reference indicators. The system uses a large language model to analyze data S and queries to obtain relevant domain knowledge C re Perform a comprehensive analysis and score the data according to the set scoring rules. During the scoring process, the system will comprehensively consider multiple factors, including the degree of match between the data and the knowledge base content, logical consistency, and the rationality of domain knowledge. At the same time, the scoring method uses the large language model based on the prompt words of different scoring dimensions, combined with the domain background and related content, to help the large language model (LLM) model infer and judge the score of the data point S under the scoring rule. This process provides users with intuitive and accurate quality assessment results, ensuring that the assessment is highly operational and of reference value.

[0056] Step 304: The system determines the score based on the score S. scoreGenerate corresponding data quality improvement suggestions. The proposal of data quality improvement suggestions depends on the scoring results, and the improvement suggestions are based on the matching degree of data and knowledge base content, logical consistency and rationality of domain knowledge. Specifically, the scoring method in the embodiment of the present application uses the prompt of the large language model based on different scoring dimensions, refers to the relevant content, and allows LLMs to judge the score of this data point S under each scoring rule. The system will generate specific improvement measures for data quality through the reasoning process of the large language model (LLM) and combined with the scoring dimensions. These suggestions may include adjusting the data collection process, supplementing missing data, optimizing data annotation, correcting data logic errors, etc., to improve the overall quality of the data. In the reasoning process, LLM, in addition to considering the defects of the data itself, will also conduct in-depth analysis of the data based on domain knowledge to identify potential optimization directions. Through the improvement suggestions generated by this reasoning, the system can ensure that the data meets the needs of subsequent applications, while improving the availability and reliability of the data, and helping the subsequent analysis and decision-making process.

[0057] Step C: In the data quality assessment application of this application, the data point quality assessment module based on machine learning algorithms relies on a variety of mature traditional algorithms, combined with the explanatory analysis function of LLMs, to conduct a comprehensive test and multi-dimensional analysis of data quality. In this process, a variety of machine learning algorithms are combined to conduct a multi-dimensional evaluation of data from aspects such as outlier detection, outlier identification, dimensionality reduction visualization, and regression classification prediction model analysis to quantify data quality and ensure the reliability of the evaluation. It assists users in discovering potential problems, identifying low-quality data, and providing an algorithmic quantitative detection basis for subsequent data quality assessments. Reference Figure 5 , a specific implementation of step C specifically includes the following steps 401-404.

[0058] Step 401: Algorithm selection and parameter configuration. The user selects the required evaluation algorithm in the algorithm selection area of ​​the system interface, including outlier detection A out , Outlier DetectionA ano , Dimensionality reduction visualization A dim , and the machine learning model predicts A ML According to the specific scenario requirements, users can configure the relevant algorithm parameters C out , C ano , C dim and C ML , such as threshold, distance measurement method, dimensionality reduction, etc., to optimize the algorithm performance.

[0059] Step 402: Data analysis and calculation processing. The system performs algorithm analysis and calculation on the input data according to the algorithm selected by the user and the corresponding configuration, and calculates the input data according to the visualization function A. draw Output the corresponding visualization image O fig :

[0060] O out =A out (S,C out )

[0061] O ano =A ano (S,C ano )

[0062] O dim =A dim (S,C dim )

[0063] O ML =A ML (S,C ML )

[0064] O fig =A draw (O out ,O ano ,O dim ,O ML )

[0065] Among them, outlier detection O out :Use outlier algorithms such as box plots, Z scores, or DBSCAN to identify and remove outliers in the data to ensure the rationality of data distribution.

[0066] Outlier Detection ano : Use outlier detection algorithms such as isolation forest and local outlier factor (LOF) to detect outliers in the data and reduce the interference of noise on the analysis results.

[0067] Dimensionality reduction visualization dim :Use dimensionality reduction algorithms such as PCA and t-SNE to map high-dimensional data to low-dimensional space, helping users to intuitively observe data distribution and identify potential anomalies.

[0068] Machine learning model predicts O ML :By training different machine learning prediction models, such as random forests, neural networks, etc., the performance indicators of the data are evaluated, and the difference between the predicted values ​​and the actual values ​​is compared to judge the data quality.

[0069] Visualization results fig :O fig To pass A draw Multiple visualizations are drawn, covering the identification of outliers and abnormal points, the distribution of data points after dimensionality reduction, and the comparison between the prediction results of the machine learning model and the actual values. This visualization can effectively show the data distribution characteristics and model performance, which is convenient for further analysis and understanding.

[0070] Step 403: Algorithm result display and interpretability analysis. After the system completes the calculation, the evaluation results are intuitively displayed on the operation interface. The display content includes: algorithm detection indicators, such as the number of outliers, the proportion of abnormal points, machine learning prediction results, and the visualization results of each algorithm. Then, the system calls f LLM , in different prompt words P out , P ano , P dim and P ML Under the guidance of T out , T ano , T dim and T ML Perform interpretive analysis on the results and generate a detailed evaluation report R out , R ano , R dim and R ML (The report includes quality problem location, possible cause analysis and optimization suggestions):

[0071] R out =f LLM (S,O out ,T out ,P out )

[0072] R ano =f LLM (S,O ano ,T ano ,P ano )

[0073] R dim =f LLM (S,O dim ,T dim ,P dim )

[0074] R ML =f LLM (S,O ML ,T ML ,P ML )

[0075] Step 404: Data quality optimization suggestions based on analysis results. Based on the analysis results of the machine learning algorithm, the system generates personalized data quality optimization suggestions for the user. The user is prompted to supplement the experimental data to cover the key data distribution area, or to correct the data collection process to reduce the impact of noise. In addition, the system can also suggest that the user adjust the sampling range, improve the data annotation, or further verify the authenticity of the abnormal data, providing specific action directions for subsequent data quality improvement.

[0076] Step D: In the data quality assessment application of the present invention, the quality assessment report generation module based on intelligent agent technology integrates machine learning algorithms, knowledge base reasoning and multimodal visual analysis to automatically generate a quality assessment report with high interpretability and scientificity. In this process, the evaluation report is generated by the intelligent agent based on LLM, key indicators and analysis results are integrated, and the interpretability analysis of the evaluation results returned by various algorithms is performed in combination with domain knowledge, providing users with intuitive and efficient decision support. The specific implementation of step D specifically includes the following steps 501-504:

[0077] Step 501: Multi-dimensional quality assessment and intelligent analysis based on machine learning algorithms. The system conducts a comprehensive quality assessment by integrating the multi-dimensional algorithm results generated in the data quality assessment, such as outlier detection, abnormal point analysis, dimensionality reduction distribution, and performance prediction. In this process, the intelligent agent combines the technical background and application value of each indicator and generates a comprehensive quality assessment based on the prompt word P. ML Under the guidance of the AI, the AI ​​generates analytical content of each algorithm result by reasoning, deeply reveals the source and potential impact of data quality issues, and provides technical reference support for users. The AI ​​Agent uses the reasoning ability of the large language model to comprehensively analyze the evaluation results and generate a quantitative analytical report:

[0078] R1=f Agent (R out ,R ano ,R dim ,R ML ,P ML )

[0079] Step 502: Knowledge base-driven domain feature evaluation and high-value information extraction. Relying on the domain reasoning ability of large language models (LLMs), the system comprehensively analyzes the matching degree, logical consistency and rationality of the data and the knowledge base content. know Under the guidance of , the intelligent agent uses domain knowledge to conduct in-depth reasoning and analysis, further refines high-value information, and generates in-depth evaluation results with domain characteristics. These results not only help users accurately locate data problems, but also provide a scientific basis for subsequent improvements. By combining domain knowledge and content generated by reasoning, the system ensures the comprehensiveness and scientificity of the evaluation process:

[0080] R2=f Agent (R know ,P know )

[0081] Step 503: Interpretation of the visualization results of multi-modal intelligent analysis. fig Under the prompt of figThe intelligent agent automatically analyzes and intelligently interprets the data (such as dimensionality reduction distribution diagrams, outlier detection diagrams, etc.). It extracts key information from the charts and generates accurate analysis based on the context, helping users intuitively understand the potential problems in the data and their impact on the overall results. The intelligent agent generates detailed chart interpretations through comprehensive analysis and converts complex chart information into easy-to-understand natural language descriptions, thereby further improving the readability and application value of the report:

[0082] R3=f Agent (O fig ,O fig )

[0083] Step 504: Data quality assessment report integration and summary. By integrating the machine learning algorithm analysis results, knowledge base reasoning conclusions and multi-modal visual analysis content, the system sum Under the guidance of , a structured comprehensive quality assessment report is automatically generated. The report comprehensively covers data problem location, quality quantitative indicators, optimization suggestions and intuitive display. Through the in-depth analysis and prompt-driven summary capabilities of the intelligent agent, the report helps users fully understand the current status of data, accurately identify data quality issues, and clarify the direction of improvement with a clear structure and rich content, providing solid and reliable support for the application of high-quality data:

[0084] R = f Agent (R1, R2, R3, P sum )

[0085] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0086] This article uses specific examples to illustrate the principles and implementation methods of this application. The description of the above embodiments is only used to help understand the method and core ideas of this application. At the same time, for those skilled in the art, according to the ideas of this application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as limiting this application.

Claims

1. A material data quality evaluation method based on a large model agent, characterized in that: include: Based on the rule set, the original material data is corrected using a cyclic detection mechanism to obtain the material data to be evaluated; Based on the knowledge base, RAG technology and LLMs model are used to conduct quality assessment on the material data to be assessed, and the data quality assessment results based on the knowledge base are obtained; Combine multiple machine learning algorithms to conduct multi-dimensional evaluation of the material data to be evaluated, obtain multi-dimensional test result visualization images and multiple data quality evaluation results based on machine learning algorithms; Based on the large model agent, the data quality assessment results based on the knowledge base, the visualization images of the multi-dimensional detection results and the data quality assessment results based on multiple machine learning algorithms are integrated and analyzed to obtain a comprehensive quality assessment result.

2. The material data quality evaluation method based on large model agent according to claim 1 is characterized in that: Based on the rule set, a cyclic detection mechanism is used to modify the original material data to obtain the material data to be evaluated, including: Based on the large language model, the rules in the natural language form input by the user are converted into programming language form to build a rule set; Initialize the value of n to 0; Initialize the original material data to the material data after the nth correction; Using a preset rule function, based on the rule set, the material data after the nth correction is tested to obtain a test result, and the test result is fed back to the user; Obtaining the material data after the n+1th correction; the material data after the n+1th correction is obtained by the user manually correcting the material data after the nth correction according to the test data; Increment the value of n by 1, and return to the step of "using a preset rule function to detect the material data after the nth correction based on the rule set, obtaining a detection result, and feeding back the detection result to the user", until the amount of material data that needs to be corrected by the user in the detection result is less than the correction threshold, or the value of n is not less than the correction number threshold, and output the material data after the nth correction as the material data to be evaluated.

3. The material data quality evaluation method based on large model Agent according to claim 1 is characterized in that: Based on the knowledge base, RAG technology and LLMs model are used to conduct quality assessment on the material data to be assessed, and the data quality assessment results based on the knowledge base are obtained, including: Using RAG technology to detect material knowledge related to the material data to be evaluated in the knowledge base as target material knowledge; Based on the target material knowledge, the LLMs model is used to perform quality analysis on the material data to be evaluated, and a data quality evaluation result based on the knowledge base is obtained.

4. The material data quality evaluation method based on large model Agent according to claim 3 is characterized in that: The material knowledge related to the material data to be evaluated in the knowledge base is detected by using RAG technology as target material knowledge, which previously also includes: Each material knowledge in the knowledge base and the material data to be evaluated are encoded to obtain a retrieval structure of each material knowledge and a retrieval structure of the material data to be evaluated.

5. The material data quality evaluation method based on large model Agent according to claim 1 is characterized in that: Combine multiple machine learning algorithms to conduct multi-dimensional evaluation of the material data to be evaluated, obtain multi-dimensional test result visualization pictures and multiple data quality evaluation results based on machine learning algorithms, including: Different machine learning algorithms are used to perform outlier detection, abnormal point detection, dimension reduction visualization and machine learning model prediction on the material data to be evaluated, and the number of outliers, abnormal point ratio, dimension reduction visualization results and machine learning prediction results are obtained; Use visualization functions to visualize the number of outliers, the proportion of abnormal points, the dimensionality reduction visualization results, and the machine learning prediction results to obtain a visualization image of the multi-dimensional detection results; Based on the characteristic knowledge of different machine learning algorithms, the LLMs model was used to perform explanatory analysis on the number of outliers, the proportion of abnormal points, the dimensionality reduction visualization results and the machine learning prediction results, and four data quality assessment results based on machine learning algorithms were obtained.

6. The material data quality evaluation method based on large model Agent according to claim 1 is characterized in that: The comprehensive quality assessment results are: R=f Agent (R1,R2,R3,P sum ); R1=f Agent (R out ,R ano ,R dim ,R ML ,P ML ); R2=f Agent (R know ,P know ); R3=f Agent (O fig ,P fig ); Among them, R is the comprehensive quality assessment result, f Agent () is the large model Agent, R1, R2 and R3 are the first quality assessment result, the second quality assessment result and the third quality assessment result respectively, R out , R ano , R dim and R ML are the data quality assessment results based on four machine learning algorithms, P ML , P know , P fig , P sum They are the prompt words for multidimensional quality assessment, visual result interpretation, knowledge base-driven domain characteristic assessment, and comprehensive quality assessment. out is the data quality assessment result about the number of outliers, R ano is the data quality assessment result on the proportion of outliers, R dim is the data quality assessment result of the dimension reduction visualization result, R ML For data quality assessment results on machine learning prediction results, R know is the data quality assessment result based on the knowledge base, O fig Visualize images of multi-dimensional detection results.

7. The material data quality evaluation method based on large model Agent according to claim 1 is characterized in that: Based on the knowledge base, RAG technology and LLMs model are used to evaluate the quality of the material data to be evaluated, and the data quality evaluation results based on the knowledge base are obtained, which also includes: Scoring the data quality assessment results based on the knowledge base to obtain scoring results; A data quality improvement suggestion is generated based on the scoring result.

Citation Information

Patent Citations

  • Data compliance detection method and system based on large language model and AI-Agent

    CN118228250A

  • Enterprise annual report analysis method based on LLM and RAG

    CN118520867A

  • BOM data multi-dimensional analysis system and analysis method based on large language model

    CN119004372A

  • Method and apparatus for analyzing medical data using large language model

    KR102747558B1

  • Dynamic evaluation of language model prompts for model selection and output validation and methods and systems of the same

    US12147513B1

Cited By

  • Regenerated material recovery method, device and equipment based on multi-modal large language model

    CN120804344A

  • Carbon footprint data quality detection method and device and storage medium

    CN121071761A

  • Data quality detection method, device and storage medium for carbon footprint

    CN121071761B

  • Risk discovery and early warning method and system based on large model

    CN121436642A