Model analysis method and device based on knowledge graph, equipment and medium
By building a professional data set and a multi-dimensional analysis index system based on the knowledge graph, the shortcomings of the existing evaluation system in specific fields are solved, the accurate evaluation and adaptive evaluation of the model are achieved, and the application effect of the model in the medical and financial fields is improved.
Patent Information
- Application Number
- CN202510275123.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-07-18
AI Technical Summary
The existing model evaluation system lacks flexibility in specific fields such as medical care and finance, and it is difficult to accurately measure the professional knowledge mastery of the model, resulting in incomplete and inaccurate evaluation results.
Using a knowledge graph-based method, we collect data from the target field, build a professional data set, extract keywords and label them, generate a knowledge graph, establish a professional corpus, design a multi-dimensional analysis index system, input a model and conduct knowledge verification, and generate analysis results.
It improves the accuracy and adaptability of model evaluation, can more comprehensively measure the semantic understanding ability and knowledge matching of the model, enhances the credibility of the evaluation results, and optimizes the application ability of the large model in specific fields.
Smart Images

Figure CN120338062A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data analysis technology, and in particular to a knowledge graph-based model analysis method, device, equipment and storage medium. Background Art
[0002] In the field of natural language processing (NLP), with the rapid development of large-scale pre-trained models, the evaluation system of general large models has become mature. However, in the application of specific fields (such as medical and financial), the existing evaluation system still has serious deficiencies, making it difficult to accurately evaluate the actual capabilities of the model in the professional field. These deficiencies are specifically reflected in the following aspects:
[0003] In the medical field, the model may have a high level of language fluency in generating medical texts, but it is difficult to accurately judge whether the medical guidelines it interprets (such as the NCCN Guidelines) meet the standards. Existing evaluation systems usually ignore the differences between different medical standards, and it is difficult to accurately judge whether the medical advice generated by the model is in line with the latest medical practices. In addition, the training data of many models comes from public medical papers or encyclopedia-type medical knowledge bases, but lacks data support such as actual clinical records, case analysis, and surgical guidelines, so that the evaluation results cannot fully reflect the true capabilities of the model.
[0004] In the financial field, the existing evaluation system mainly focuses on the grammatical correctness of the model-generated financial report analysis or market forecast, but it is difficult to measure its ability to understand the relationship between complex financial logic, regulatory terms and financial data. Financial regulatory rules vary in different markets and legal environments. The existing evaluation system cannot effectively determine whether the model can correctly parse and adapt to different regulatory systems, resulting in a decrease in the accuracy of financial compliance analysis. The training data of the model is usually based on public texts such as financial news and financial reports, but lacks high-value data such as financial regulations, securities analyst reports, and corporate internal audit reports, resulting in the lack of effectiveness of the model evaluation in scenarios such as financial risk control and legal compliance review.
[0005] There are significant differences in the evaluation needs of different fields, but the existing evaluation system lacks flexibility and is difficult to adapt to the special requirements of various industries. In the medical field, the evaluation of diagnostic reasoning tasks requires contextual reasoning in combination with clinical case data, but existing evaluation methods usually rely only on text similarity matching, which makes it difficult to measure the clinical decision-making ability of the model. In the financial field, market trend forecasting involves multidimensional time series data, and existing evaluation methods usually focus on the accuracy of text generation and fail to include the performance of the model in data-driven tasks in the evaluation scope. The lack of adaptability of this evaluation system limits the evaluation ability of the model in complex field tasks, making it difficult to fully measure the practical application value of the model. Summary of the invention
[0006] The main object of the present invention is to provide a model analysis method, device, equipment and storage medium based on a knowledge graph, aiming to solve the technical problem that the evaluation system in the prior art lacks the integration of professional knowledge in the target field and the in-depth evaluation of semantic understanding, resulting in the inability to accurately measure the knowledge mastery ability of the large model in a specific field.
[0007] To achieve the above object, the present invention provides a model analysis method based on a knowledge graph, including:
[0008] Collecting target field data, and screening and classifying the target field data to generate a professional data set;
[0009] Performing keyword extraction and annotation processing on the core concepts and professional terms in the professional data set to generate an annotation result;
[0010] Constructing a target field knowledge graph including a domain entity set, an entity relationship network and an attribute feature library based on the annotation result;
[0011] Establishing a professional corpus based on the professional data set, and generating an analysis task set based on the professional corpus;
[0012] Constructing a multi-dimensional analysis index system, and each analysis task in the analysis task set is associated with at least one analysis dimension in the multi-dimensional analysis index system;
[0013] Inputting the analysis task into the model to be analyzed, and obtaining the task output result of the model to be analyzed;
[0014] Performing knowledge verification on the task output result according to the target field knowledge graph to generate a knowledge verification result;
[0015] Analyzing the knowledge verification result based on the multi-dimensional analysis index system to generate an analysis result.
[0016] Furthermore, to achieve the above object, the present invention provides a model analysis device based on a knowledge graph, including:
[0017] A data processing module, configured to collect target field data, and screen and classify the target field data to generate a professional data set;
[0018] A keyword extraction and annotation module, configured to perform keyword extraction and annotation processing on the core concepts and professional terms in the professional data set to generate an annotation result;
[0019] A knowledge graph construction module, configured to construct a target field knowledge graph including a domain entity set, an entity relationship network and an attribute feature library based on the annotation result;
[0020] An analysis task generation module, configured to establish a professional corpus based on the professional data set, and generate an analysis task set based on the professional corpus;
[0021] An evaluation index management module, configured to construct a multi-dimensional analysis index system, and each analysis task in the analysis task set is at least associated with one analysis dimension in the multi-dimensional analysis index system;
[0022] A task execution module, configured to input the analysis task into the model to be analyzed, and obtain the task output result of the model to be analyzed;
[0023] A knowledge verification module, configured to perform knowledge verification on the task output result according to the knowledge graph of the target domain, and generate a knowledge verification result;
[0024] A result analysis module, configured to analyze the knowledge verification result based on the multi-dimensional analysis index system, and generate an analysis result.
[0025] Further, to achieve the above object, the present invention further provides a computer device, the computer device includes a memory, a processor, and a knowledge graph-based model analysis program stored on the memory and executable on the processor, and when the knowledge graph-based model analysis program is executed by the processor, the steps of the knowledge graph-based model analysis method as described above are implemented.
[0026] Further, to achieve the above object, the present invention further provides a computer-readable storage medium, on which a knowledge graph-based model analysis program is stored, and when the knowledge graph-based model analysis program is executed by a processor, the steps of the knowledge graph-based model analysis method as described above are implemented.
[0027] Beneficial effects: The present invention relates to the technical field of data analysis and can be applied to business scenarios such as medical health, fintech, and cultural research. It discloses a model analysis method based on a knowledge graph, including: collecting target domain data, screening and classifying it to generate a professional data set; extracting keywords and performing annotation processing on core concepts and professional terms in the professional data set to generate an annotation result; constructing a target domain knowledge graph containing a domain entity set, an entity relationship network, and an attribute feature library based on the annotation result; establishing a professional corpus based on the professional data set and generating an analysis task set; constructing a multi-dimensional analysis index system and associating the analysis tasks with analysis dimensions; inputting the analysis tasks into the model to be analyzed to obtain task output results; verifying the knowledge of the task output results based on the target domain knowledge graph to generate a knowledge verification result; analyzing the knowledge verification result based on the multi-dimensional analysis index system to generate an analysis result. By introducing the target domain knowledge graph, the present invention realizes a structured comparison of the model output content, improving the accuracy of evaluation; through the multi-dimensional analysis index system, comprehensively measures the semantic understanding ability, knowledge matching degree, and reasoning ability of the model, enhancing the adaptability of evaluation; through the professional corpus and analysis task set, makes the evaluation task more suitable for specific domain application scenarios, improving the credibility of the evaluation result, thereby optimizing the knowledge mastery effect and application ability of the large model in a specific domain. Brief Description of the Drawings
[0028] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings:
[0029] Figure 1 is a schematic diagram of an application environment of the model analysis method based on a knowledge graph in an embodiment of the present invention;
[0030] Figure 2 is a schematic flowchart of an embodiment of the model analysis method based on a knowledge graph of the present invention;
[0031] Figure 3 is a schematic diagram of functional modules of a preferred embodiment of the model analysis device based on a knowledge graph of the present invention;
[0032] Figure 4 is a schematic diagram of the structure of a computer device in an embodiment of the present invention;
[0033] Figure 5 is another schematic diagram of the structure of a computer device in an embodiment of the present invention. Detailed Embodiments
[0034] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0035] The model analysis method based on a knowledge graph provided by the embodiments of the present invention can be applied in such asFigure 1 In the application environment, the client communicates with the server through the network. The server can collect target domain data through the client, screen and classify it to generate a professional data set; extract and label keywords for the core concepts and professional terms in the professional data set to generate a labeling result; build a target domain knowledge graph including a domain entity set, an entity relationship network, and an attribute feature library based on the labeling result; establish a professional corpus based on the professional data set and generate an analysis task set; build a multi-dimensional analysis index system and associate the analysis tasks with the analysis dimensions; input the analysis tasks into the model to be analyzed to obtain task output results; verify the knowledge of the task output results based on the target domain knowledge graph to generate a knowledge verification result; analyze the knowledge verification result based on the multi-dimensional analysis index system to generate an analysis result. By introducing the target domain knowledge graph, the present invention realizes the structured comparison of the model output content, improves the accuracy of evaluation; through the multi-dimensional analysis index system, comprehensively measures the semantic understanding ability, knowledge matching degree and reasoning ability of the model, and enhances the adaptability of evaluation; through the professional corpus and the analysis task set, the evaluation task is more in line with the specific domain application scenario, improves the credibility of the evaluation result, and thus optimizes the knowledge mastery effect and application ability of the large model in a specific domain. Among them, the client can be, but is not limited to, various personal computers, laptop computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0036] Please refer to Figure 2 , Figure 2 which is a schematic flowchart of an embodiment of the model analysis method based on a knowledge graph provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than here.
[0037] As Figure 2 shown, the model analysis method based on a knowledge graph proposed by the present invention includes the following steps:
[0038] S10, collect target domain data, and screen and classify the target domain data to generate a professional data set;
[0039] In this embodiment, collecting target domain data and screening and classifying it to generate a professional data set is the basic link of the whole method. This step involves multiple technical features, including the selection of data sources, the structured arrangement of data, data cleaning and quality control, the setting of data classification rules, and the storage and management of professional data sets.
[0040] The selection of data sources directly affects the quality and coverage of professional datasets. Data can be collected from open datasets, industry databases, data released by research institutions, enterprise internal knowledge bases, and other channels. In the medical field, data sources such as hospital electronic medical record systems, clinical guidelines, and medical papers can be selected; in the financial field, financial statements, market data, regulatory regulations, etc. can be collected. To improve the integrity and professionalism of the data, methods such as web scraping, API interface acquisition, and manual entry can be combined to ensure that the data sources are extensive and authoritative.
[0041] Data structuring is an important step to ensure that data can be efficiently processed by large models. The original data may be unstructured text, tabular data, images, or audio records. To facilitate subsequent analysis, the data needs to be standardized and format-converted. For example, in the medical field, electronic medical record data from different hospitals can be converted into a unified medical data format (such as HL7, FHIR standards), and in the financial field, financial data from different companies can be converted into a unified XBRL format.
[0042] The main purpose of data cleaning and quality control is to remove redundant information, correct incorrect data, and fill in missing information to improve the accuracy and reliability of the data. Methods such as rule matching, statistical analysis, and anomaly detection can be used for data cleaning. For example, in medical data processing, abnormal values caused by input errors can be removed, such as data with blood pressure and blood sugar values beyond the physiological range. In financial data processing, duplicate data caused by inconsistent formats can be discovered and removed by cross-comparing financial report information.
[0043] The setting of data classification rules determines the hierarchical structure of professional datasets. Different classification criteria can be set according to the knowledge system and application requirements of the field. For example, in the medical field, data can be divided into subcategories such as disease diagnosis and treatment records, clinical trial data, and medical research literature; in the financial field, data can be divided into categories such as company financial data, market data, and regulatory policies. Machine learning classification algorithms such as text clustering and topic modeling can be combined during the classification process to improve the automation and accuracy of data classification.
[0044] The storage management of professional datasets needs to ensure the accessibility, security, and efficient retrieval of the data. Relational databases (such as MySQL, PostgreSQL) can be used to store structured data, NoSQL databases (such as MongoDB, Elasticsearch) can be used to store unstructured text data, and distributed file systems (such as HDFS) can also be used to store large-scale data. In the medical field, data storage needs to comply with data privacy protection standards such as HIPAA; in the financial field, data storage needs to meet regulatory requirements such as SOX.
[0045] Different implementation methods can be adjusted according to specific application scenarios. For example, in the data collection stage, data can be obtained through automated crawling or manual entry. In the data classification stage, rule-based classification methods or machine learning algorithms can be used for automatic classification. In terms of data storage management, local databases or cloud storage solutions such as AWS S3 and Google Cloud Storage can be used to meet different computing and storage requirements.
[0046] Data screening can be combined with manual review and automated rule matching. For example, in medical data processing, rules for physiological rationality checks can be set, and in financial data processing, historical data comparison methods can be used to detect data anomalies.
[0047] By constructing a professional dataset, the integrity, accuracy, and domain coverage of the data are improved, providing a reliable data foundation for subsequent keyword extraction, knowledge graph construction, and model evaluation.
[0048] S20, perform keyword extraction and annotation processing on the core concepts and professional terms in the professional dataset to generate an annotation result;
[0049] In this embodiment, performing keyword extraction and annotation processing on the core concepts and professional terms in the professional dataset to generate an annotation result is a key step to ensure that the model can accurately understand domain-specific terms and concepts. This step involves multiple technical features, including keyword extraction, semantic analysis, annotation rule setting, core concept annotation, professional term annotation, and storage management of the annotation result, etc.
[0050] Keyword extraction is to identify terms of great significance from the professional dataset for subsequent semantic analysis and knowledge structure construction. It can be extracted based on statistical methods (such as TF-IDF), deep learning models (such as BERT), or rule matching methods. For example, in the medical field, frequently occurring disease names, drug names, and treatment methods are extracted; in the financial field, industry terms, financial indicators, and market-related terms are extracted.
[0051] Semantic analysis is used to identify the context relationship between the extracted keywords and the original text to ensure the accuracy of the keywords. Word vector models (such as Word2Vec, FastText) can be used to calculate the semantic similarity between the keywords and the context, filtering out words irrelevant to the domain. For example, in medical data processing, the difference between "hypertension" (disease) and "high blood pressure" (symptom) can be distinguished; in financial data processing, "yield rate" can be identified as a financial indicator rather than ordinary "income".
[0052] The annotation rule setting determines how to attach structured information to keywords, making them suitable for subsequent knowledge graph construction and model training. Different types of annotation rules can be set, such as entity category annotation, part-of-speech annotation, semantic category annotation, etc. For example, in the medical field, different annotation labels can be set for diseases, symptoms, and treatment methods; in the financial field, categories such as company names, financial indicators, and market events can be distinguished.
[0053] Core concept annotation refers to the detailed annotation of the core concepts identified in keyword extraction, including concept definitions, semantic categories, context relationships, etc. For example, in the medical field, "diabetes" is annotated as "chronic disease category" and relevant symptoms are attached; in the financial field, "price-earnings ratio" is annotated as "financial indicator".
[0054] Professional term annotation is used to identify terms in a specific field and provide synonyms, usage scenarios, and semantic restrictions. For example, in the medical field, "hyperglycemia" can be annotated with "high blood sugar" as a synonym; in the financial field, "stock repurchase" can be annotated with "company repurchase" as an alternative term.
[0055] The storage management of annotation results is an important link to ensure that the annotation information can be used efficiently in subsequent steps. The annotation results can be stored in JSON, XML, or database structures, making them easy to retrieve and update. For example, in medical data processing, the annotation results can be stored in a clinical terminology database; in financial data processing, a professional term index library can be constructed.
[0056] By constructing a complete professional dataset, it is ensured that the data used in the evaluation process has high quality, strong domain adaptability, and structured characteristics, thereby improving the knowledge coverage breadth and accuracy of the large model. Through keyword extraction and annotation processing, core concepts and professional terms are extracted, enabling the evaluation task of the large model to be verified based on precise semantic analysis and structured knowledge. Through the semantic annotation of professional terms and core concepts, the construction of the knowledge graph becomes more systematic and comprehensive, ensuring that the model can be comprehensively evaluated based on multi-dimensional indicators (such as semantic accuracy, norm compliance, and concept system depth) during evaluation.
[0057] S30, construct a target domain knowledge graph containing a domain entity set, an entity relationship network, and an attribute feature library based on the annotation results;
[0058] In this embodiment, a target domain knowledge graph is constructed to achieve the structured organization of domain entities, the automatic reasoning of semantic relationships, and the precise management of attribute features. This process covers entity set construction, relationship network establishment, and attribute feature induction. By constructing structured information, the model's understanding and reasoning abilities in a specific domain are improved.
[0059] In the process of constructing a knowledge graph, first, domain-related core concepts and professional terms are extracted based on the annotation results and summarized into an entity set. The named entity recognition technology is used to automatically identify important concepts in text data, and combined with the existing domain term library, the entity names are standardized. For example, in medical text data, the expressions of entities related to "hyperglycemia" and "diabetes" are unified, and in financial data, entity mappings are established between "market adjustment" and "financial crisis".
[0060] The semantic relationships between entities are constructed through relation extraction technology, involving hierarchical relationships, causal relationships, time series relationships, etc. Based on syntactic analysis and knowledge embedding methods, the logical connections between concepts in the text are extracted and a relation network is constructed. For example, in medical data, the causal relationship that "hypertension" causes "heart disease" can be defined, and in financial data, the impact of "interest rate hikes" on "stock market fluctuations" can be described.
[0061] The organization of entity attributes includes information such as time, space, and logical features, enabling the knowledge graph to carry richer context information. By supplementing attribute features, the ability of knowledge reasoning is improved. For example, in a medical knowledge base, information such as indications and contraindications is added to drugs, and in a financial knowledge base, the time attribute of industry dynamics is marked.
[0062] When constructing the entity set, a rule-based method can be selected to match standard terms using a predefined dictionary, or a deep learning method (such as BERT-CRF) can be used to automatically identify domain terms. In medical text processing, an entity dictionary is constructed based on the ICD-10 disease classification system. In financial literature analysis, financial indicators are identified based on a financial term library.
[0063] When establishing a relation network, explicit relationships can be extracted based on the dependency syntactic analysis method, or implicit associations can be mined based on knowledge graph embedding technology (such as TransE). In medical texts, the logical relationships between diseases, symptoms, and treatment methods can be identified; in financial analysis, an association network of market policies, corporate decisions, and investment returns can be constructed.
[0064] When generalizing attribute features, a structured database (such as RDF) can be used to store triple knowledge, or a graph database (such as Neo4j) can be used to build an efficient retrieval system. In the medical field, information such as diagnostic criteria, clinical manifestations, and treatment methods can be associated with each disease; in the financial field, dynamic attributes of corporate financial report indicators over time can be constructed.
[0065] By constructing a knowledge graph for the target domain, knowledge enhancement for large model analysis tasks is achieved, enabling the model to more accurately match professional knowledge in semantic parsing, concept reasoning, and evaluation tasks, and improving the accuracy and interpretability of analysis. Through the construction of an entity set, the complete expression of core concepts is ensured; through the construction of a relationship network, knowledge associations become clearer; through the induction of attribute features, domain knowledge can support complex reasoning and evaluation tasks, thereby enhancing the adaptability of large models in different domains and the quality of evaluation.
[0066] S40, establishing a professional corpus based on the professional dataset, and generating an analysis task set based on the professional corpus;
[0067] In this embodiment, by constructing a professional corpus and generating an analysis task set based on the corpus, targeted evaluation of the large model is ensured, making the evaluation task professional, targeted, and scalable. This process involves multiple aspects such as data construction, content classification, structured storage, and task generation mechanism of the corpus, enabling the analysis task to cover the performance of the model at different levels, thereby enhancing the scientificity and applicability of the evaluation.
[0068] In the process of constructing the professional corpus, high-quality texts are first screened from the professional dataset to ensure the authority and accuracy of the data source. For example, in the medical field, the corpus can include "Internal Medicine", "Pharmacopoeia", and various clinical guidelines; in the financial field, the corpus can cover national fiscal policies, financial market regulations, and industry analysis reports.
[0069] Secondly, the selected texts are classified into three categories: core literature, annotated literature, and derivative materials. Core literature refers to authoritative basic theory literature, such as clinical guidelines in the medical field and regulatory regulations in the financial field. Annotated literature refers to the explanatory and extended content based on core literature, such as expert comments, research papers, and interpretive articles. Derivative materials include case analyses, research reports, and practical experience sharing in actual applications, such as financial market event analysis, medical record records, etc.
[0070] To enhance the usability of the corpus, the text data needs to be structured, extracting core concepts, key terms, and context information, and establishing an index. Natural language processing technologies (such as TF-IDF, BERT, etc.) can be used to extract keywords, identify entities, and model topics from the corpus to improve information retrieval efficiency. For example, in medical texts, diseases, drugs, and treatment plans are extracted, and in financial texts, financial indicators, policy events, and market trends are extracted.
[0071] Generate an analysis task set based on a professional corpus to ensure that the evaluation tasks cover the core capabilities of large models. The analysis tasks mainly focus on key dimensions such as semantic understanding, compliance with norms, and depth of the concept system, and design diverse evaluation methods.
[0072] The task types can include:
[0073] Text parsing tasks: Provide texts in specific domains and require the model to perform structured interpretation. For example, in the medical field, provide patient medical records and require the model to identify key symptoms, diagnosis conclusions, and treatment suggestions; in the financial field, provide market analysis reports and require the model to summarize market trends and investment strategies.
[0074] Knowledge reasoning tasks: Design question chains to examine the reasoning ability of the model. For example, in the medical field, input the cause of the disease and require the model to deduce possible complications; in the financial field, input policy adjustments and require the model to infer market reactions.
[0075] Case comparison tasks: Provide multiple instances and require the model to analyze the differences. For example, in the medical field, compare the applicability of two different treatment plans; in the financial field, compare the financial report data of two enterprises and analyze their operating conditions.
[0076] The task set should consider difficulty grading, from basic recognition tasks to high-order reasoning tasks, to construct a hierarchical evaluation system. For example, in medical text analysis, the basic task is to identify medical terms in the text, the intermediate task is to understand clinical decisions, and the advanced task is to reason about complex cases. In the financial field, the basic task is to extract financial report data, the intermediate task is to analyze market trends, and the advanced task is to predict the economic cycle.
[0077] To adapt to the ability levels of different models, the analysis task set should have a dynamic adjustment mechanism, which can adjust the task complexity according to the initial performance of the model. For example, if the model performs well in basic tasks, the task difficulty will be automatically increased, and more reasoning and analysis levels will be added to ensure that the evaluation covers the ability boundary of the model.
[0078] By constructing a professional corpus, it is ensured that the evaluation tasks have high-quality data support, improving the domain adaptability of large models. Through the design of the analysis task set, the evaluation tasks cover the semantic understanding, reasoning ability, and knowledge matching ability of the model, making the evaluation results more targeted and interpretable. The multi-level task difficulty regulation mechanism ensures that the evaluation tasks can be dynamically adapted to large models with different ability levels, providing precise guidance for model optimization.
[0079] S50, construct a multi-dimensional analysis index system, and each analysis task in the analysis task set is at least associated with one analysis dimension in the multi-dimensional analysis index system;
[0080] In this embodiment, by constructing a multi-dimensional analysis index system, the evaluation process of the analysis task set has structured and multi-level evaluation capabilities. This analysis index system is mainly used to measure the performance of large models under different task types, and to evaluate the evaluation object in a fine-grained manner through different analysis dimensions, ensuring the comprehensiveness and accuracy of the evaluation results.
[0081] The construction of the multi-dimensional analysis index system first requires determining the categories of analysis dimensions. In different application scenarios, the settings of analysis dimensions should be able to accurately measure the core capabilities of the model. For example:
[0082] Semantic accuracy analysis dimension: Measures whether the model accurately identifies and uses professional terms, concepts, and entities during text understanding and content generation. For example, in medical texts, determine whether the model can correctly distinguish between similar disease names (such as hypertension and hypotension); in the financial field, judge whether its use of economic terms is accurate.
[0083] Specification compliance analysis dimension: Evaluates whether the output content of the model conforms to established specifications, industry standards, or domain rules. For example, in medical applications, determine whether the diagnostic suggestions generated by the model comply with clinical practice guidelines; in financial analysis, evaluate its compliance with financial report analysis.
[0084] Concept system depth analysis dimension: Measures the model's reasoning ability and understanding of the relationships between concepts in domain knowledge. For example, in a medical scenario, determine whether the model can infer the causes and possible complications of diseases; in financial analysis, analyze its understanding of the relationships between economic indicators.
[0085] After determining the analysis dimensions, it is necessary to establish a hierarchical structure of the analysis dimensions. Each analysis dimension can be further broken down into multiple fine-grained sub-indicators to more accurately measure the capabilities of the model. For example:
[0086] The semantic accuracy analysis dimension can be broken down into: professional term recognition rate, entity resolution accuracy rate, context matching degree, etc.
[0087] The specification compliance analysis dimension can be broken down into: industry standard compliance rate, regulatory consistency, expert review consistency, etc.
[0088] The concept system depth analysis dimension can be broken down into: cross-document knowledge transfer ability, logical reasoning accuracy rate, complex problem-solving completeness, etc.
[0089] To ensure that the tasks in the analysis task set can reflect the evaluation objectives, it is necessary to establish a mapping relationship between the analysis tasks and the analysis dimensions. That is, each analysis task is associated with at least one analysis dimension, so that the task evaluation results can be mapped to specific ability indicators. For example:
[0090] Task Example 1 (Medical field): Analyze the diagnosis conclusion in the medical record text. This task is associated with the semantic accuracy analysis dimension (evaluating the accuracy of the use of diagnostic terms) and the compliance analysis dimension (evaluating compliance with clinical guidelines).
[0091] Task Example 2 (Financial field): Analyze the content of the enterprise financial report. This task is associated with the semantic accuracy analysis dimension (evaluating the accuracy of the extraction of financial indicators) and the depth of the conceptual system analysis dimension (evaluating its reasoning ability for the financial health status).
[0092] In terms of setting up the analysis dimensions, the method of expert definition + data-driven modeling can be adopted.
[0093] Domain experts sort out the core competency indicators and, in combination with existing industry standards, initially set up the categories of analysis dimensions.
[0094] Use data analysis methods (such as cluster analysis, principal component analysis) to conduct pattern mining on the existing evaluation data, discover potential evaluation dimensions, and optimize the analysis index system.
[0095] In terms of the mapping between the analysis tasks and the analysis dimensions, the method of manual setting + automatic learning can be adopted.
[0096] In the initial stage, experts manually associate the existing analysis tasks with the analysis dimensions to ensure the rationality of the tasks.
[0097] Use machine learning models (such as Bayesian networks, decision trees) to conduct pattern learning on a large amount of historical evaluation data, automatically recommend appropriate analysis dimensions, and improve the mapping efficiency.
[0098] In terms of the calculation of the analysis indicators, the method combining rule matching + data statistics + deep learning can be adopted.
[0099] Rule matching: For structured data (such as medical diagnosis codes, legal provisions), use knowledge graph or pattern matching-based methods to directly judge the compliance.
[0100] Data statistics: For numerical indicators (such as term recognition rate, text matching degree), use statistical analysis methods to calculate evaluation indicators such as accuracy rate, recall rate, and F1 value.
[0101] Deep learning: For complex evaluation tasks (such as concept reasoning, semantic consistency evaluation), use pre-trained language models such as BERT and GPT for feature extraction, and combine with classification models to calculate scores.
[0102] In terms of optimizing the evaluation process, an adaptive evaluation framework can be adopted to dynamically adjust the evaluation tasks according to the initial performance of the model.
[0103] If the model performs well on basic tasks (such as term recognition), high-difficulty tasks (such as logical reasoning) will be automatically added.
[0104] If the model performs abnormally on a specific task (such as performing poorly on specific domain knowledge), the test samples in this field will be focused on increasing to improve the pertinence of the evaluation.
[0105] By constructing a multi-dimensional analysis index system, the evaluation process of the analysis task set can be made more scientific, detailed, and interpretable. Each analysis task is associated with specific analysis dimensions, thus ensuring that the evaluation task can accurately measure the ability boundary of the large model. Through hierarchical analysis index design, the evaluation results can intuitively show the advantages and disadvantages of the model in different capabilities, providing targeted guidance for model optimization.
[0106] S60, input the analysis task into the model to be analyzed, and obtain the task output result of the model to be analyzed;
[0107] In this embodiment, inputting the analysis task into the model to be analyzed to obtain the task output result is the core link of large model evaluation. This process involves formatted input of tasks, model processing, and result acquisition. The analysis tasks are screened according to the previously generated task set and matched with the multi-dimensional analysis index system to ensure that the input tasks cover different analysis dimensions. The input method can adopt natural language instructions, structured data, tabular data, or multi-modal input according to the model type to meet the evaluation requirements of different fields. After the task is input, the model generates corresponding task output results based on internal training data, knowledge bases, and reasoning capabilities. The forms of output results include natural language texts, structured data, classification labels, or scoring values. To ensure that the output results meet the evaluation criteria, result preprocessing is required, including formatting, denoising, screening out irrelevant content, and converting to a standardized format for subsequent analysis.
[0108] In different implementation methods, the evaluation efficiency can be improved by optimizing the input strategy and output processing method. In terms of input optimization, the task templating method can be adopted to preset the task description format to reduce model understanding deviation. For example, in the medical field, a standardized diagnosis description template can be used, and in the financial field, the financial report analysis task can be input in a fixed format. In terms of output acquisition, the multiple sampling technique can be adopted to enable the model to generate multiple candidate results for the same task, and the optimal solution can be selected by comparing the consistency. In addition, a confidence scoring mechanism can be used to provide a confidence evaluation for each output result to screen out high-confidence results and improve the reliability of the evaluation results.
[0109] By standardizing the task input and optimizing the task output, the evaluation process becomes more stable and reliable, effectively reducing the model generation error and improving the accuracy of the evaluation. The flexibility of the task input method enhances the applicability of the evaluation system, enabling it to adapt to different types of models and ensuring the integrity and clarity of the input content, thereby enhancing the pertinence of the evaluation task. The formatting of the output results improves the readability and parsability of the data, providing accurate basic data for subsequent knowledge verification and analysis.
[0110] S70, perform knowledge verification on the task output result according to the target domain knowledge graph to generate a knowledge verification result;
[0111] In this embodiment, knowledge verification is a crucial step to ensure that the task output result conforms to the domain knowledge specification. The task output result is verified through the target domain knowledge graph to detect entity matching, relationship consistency, and attribute feature compliance, ensuring the accuracy and rationality of the content generated by the model. The basic principle of knowledge verification is to parse the task output result into structured data and layer-by-layer compare the existing entities, relationships, and attributes in the knowledge graph to identify possible errors, omissions, or logical conflicts.
[0112] First, parse the task output result, extract the involved domain entities, including proper nouns, specific terms, concept labels, etc., and match them with the entity set in the target domain knowledge graph to determine whether there are unrecognized entities or entities with incorrect mappings. Entity matching adopts methods based on semantic similarity calculation, such as cosine similarity, Word2Vec, or BERT embedding vector calculation, to improve the matching accuracy. For entities that fail to match, further analyze whether they belong to spelling mistakes, synonym variations, or model reasoning errors, and record the reasons in the knowledge verification result.
[0113] Secondly, parse the entity relationships involved in the task output result, map them to the task output relationship network, and compare it with the entity relationship network of the knowledge graph. This process includes three layers of verification: hierarchical consistency detection, checking whether the entities follow the established hierarchical structure, such as "disease" should belong to "medical classification" rather than "financial concept"; causal logic detection, confirming whether the causal relationships in the output result conform to the rules in the knowledge graph, such as "hypertension" causing "heart disease" rather than the opposite; interpretation relationship comparison, ensuring that the interpretations derived by the model conform to the existing knowledge system, such as whether the interpretation of "financial market decline" as "lack of market confidence" is consistent with the economic model.
[0114] Then, extract the attribute features in the task output results, including numerical indicators, time series information, text descriptions, etc., and compare them with the attribute feature library in the knowledge graph to detect whether there are attribute deviations or situations beyond the standard range. For example, in the medical field, if the model gives "the normal blood pressure value is 200 / 150 mmHg", then this value exceeds the medical standard and should be marked as abnormal. In the financial field, if the model predicts that the profit growth rate of an enterprise is "500%", while industry data shows that it generally does not exceed "30%", then it is necessary to further analyze the data source and calculation logic.
[0115] Finally, integrate the results of entity matching, relationship comparison, and attribute detection to generate the knowledge verification results. The knowledge verification results are stored in the form of a structured report, including items with successful matches, items with failed matches, logical conflict items, and abnormal attribute items, and are classified and marked as "high risk", "requiring manual review", or "low risk and acceptable" levels. The knowledge verification results are not only used to evaluate the credibility of the current task output results, but also can be used to guide model adjustment, such as optimizing training data, adjusting inference logic, and enhancing domain-specific knowledge, so as to improve the overall performance of the model.
[0116] In different fields, knowledge verification can be optimized by different technical means. In the medical field, the medical knowledge graph can be combined, disease matching can be performed through the ICD-10 standard classification system, term normalization can be carried out using the SNOMED CT terminology set, and the clinical decision support system (CDSS) technology can be used to verify the medical rationality of the model inference results. For example, when parsing the task output results, the BERT pre-trained model can be used for medical concept parsing, and the standardized expressions of diagnostic terms can be compared based on the UMLS (Unified Medical Language System) database to ensure that the generated medical advice complies with medical norms.
[0117] In the financial field, the economics knowledge graph can be combined to conduct a structured comparison of information such as financial statements, economic indicators, and market sentiment. For example, use XBRL (eXtensible Business Reporting Language) to parse enterprise financial data, extract key financial indicators, and conduct compliance detection based on the Basel Accord or International Financial Reporting Standards (IFRS). For market analysis tasks, time series modeling techniques can be used to backtest historical data to evaluate the rationality of model predictions, such as detecting whether it conforms to financial principles such as mean reversion and market momentum.
[0118] By introducing a knowledge graph for verifying the task output results, the output content of the large model becomes more credible, effectively detecting semantic errors, logical conflicts, and domain mismatch problems. Knowledge verification ensures the standardized expression of terms at the entity matching level, improving the model's understanding ability of domain terms; at the relationship comparison level, it guarantees the correctness of the reasoning logic and avoids unreasonable inferences in the content generated by the model; at the attribute detection level, it restricts the compliance of numerical values and texts, reducing the risk of the spread of incorrect information.
[0119] S80, analyze the knowledge verification results based on the multi-dimensional analysis index system to generate an analysis result.
[0120] In this embodiment, analyze the knowledge verification results based on the multi-dimensional analysis index system to evaluate the semantic accuracy, logical consistency, and domain fit of the task output results, ensuring that the reasoning ability of the model and the generated content meet the expected standards. The knowledge verification results usually include information such as entity matching situations, relationship comparison consistency, and attribute feature compliance. The analysis process needs to combine predefined analysis dimensions, including but not limited to the semantic accuracy analysis dimension, the specification fit analysis dimension, and the concept system depth analysis dimension, to systematically evaluate the quality of the task output results.
[0121] First, extract the entity matching situation in the knowledge verification results and calculate based on the semantic accuracy analysis dimension. This analysis dimension mainly measures the matching rate of the terms in the task output results with the standard terms in the knowledge graph, including the proportion of entities with successful matches, the number of entities with failed matches, and the matching deviation items. The matching success rate can be determined by calculating the ratio of the matched entity items to the total entity items in the task output results, and the deviation items are calculated in combination with the context semantic similarity to form an entity matching report.
[0122] Secondly, extract the relationship comparison anomaly items in the knowledge verification results and analyze them based on the specification fit analysis dimension. This analysis dimension measures whether the relationship descriptions in the task output results conform to the standard relationship definitions in the knowledge graph, including hierarchical consistency, causal logic rationality, and explanatory relationship accuracy. For the relationship comparison anomaly items, the proportion of abnormal relationship items can be calculated and classified and analyzed in combination with the error types (such as hierarchical conflicts, causal errors, explanatory deviations) to form a relationship comparison report.
[0123] Then, extract the attribute feature detection results in the knowledge verification results and evaluate them based on the concept system depth analysis dimension. This analysis dimension measures whether the attribute descriptions in the task output results conform to the attribute standards in the knowledge graph, including the attribute matching success rate, the degree of attribute deviation, and the abnormal attribute items. For the attribute consistency analysis, the proportion of successfully matched attribute items can be calculated, and the number of attribute items exceeding the standard range can be counted to form an attribute consistency report.
[0124] Finally, based on the results of the semantic accuracy analysis dimension, the specification compliance analysis dimension, and the concept system depth analysis dimension, various scoring indicators are calculated, and the analysis results are generated. The analysis results are stored in the form of structured data, including the scores of each analysis dimension, the statistics of abnormal items, and the comprehensive evaluation conclusion.
[0125] In different application scenarios, different methods can be used to optimize the analysis process. In the medical field, clinical data standards (such as SNOMED CT, ICD-10) can be combined to calculate the matching rate of diagnostic terms and evaluate the rationality of diagnostic suggestions based on a medical reasoning model. For example, if the task output result contains "Chronic kidney disease can be improved by a high-sodium diet", the specification compliance analysis dimension will detect that this conclusion does not conform to the medical guidelines and mark it as an abnormal item.
[0126] In the financial field, financial analysis standards can be combined to evaluate whether the market predictions in the task output results conform to the economic model. For example, if the model predicts that "The appreciation of the US dollar will cause a significant increase in the price of gold", the concept system depth analysis dimension will detect that this conclusion deviates from the historical market trend and adjust the scoring result.
[0127] By analyzing the knowledge verification results based on the multi-dimensional analysis index system, the evaluation of the task output results becomes more comprehensive and accurate. The semantic accuracy analysis dimension ensures the standardization of term matching, reducing problems such as term misuse or semantic drift; the specification compliance analysis dimension improves the logical consistency of the task output results, preventing reasoning errors or causal reversals; the concept system depth analysis dimension enhances the accuracy of attribute features, avoiding numerical or text descriptions that exceed the reasonable range.
[0128] The present invention relates to the technical field of data analysis and can be applied to business scenarios such as medical health, fintech, and cultural research. It discloses a model analysis method based on a knowledge graph, including: collecting target domain data, screening and classifying it to generate a professional dataset; performing keyword extraction and annotation processing on core concepts and professional terms in the professional dataset to generate an annotation result; constructing a target domain knowledge graph containing a domain entity set, an entity relationship network, and an attribute feature library based on the annotation result; establishing a professional corpus based on the professional dataset and generating an analysis task set; constructing a multi-dimensional analysis index system and associating the analysis tasks with analysis dimensions; inputting the analysis tasks into the model to be analyzed to obtain task output results; verifying the knowledge of the task output results based on the target domain knowledge graph to generate a knowledge verification result; and analyzing the knowledge verification result based on the multi-dimensional analysis index system to generate an analysis result. By introducing the target domain knowledge graph, the present invention realizes the structured comparison of the model output content, improves the accuracy of evaluation; through the multi-dimensional analysis index system, comprehensively measures the semantic understanding ability, knowledge matching degree, and reasoning ability of the model, enhances the adaptability of evaluation; through the professional corpus and the analysis task set, makes the evaluation task more suitable for specific domain application scenarios, improves the credibility of the evaluation result, and thus optimizes the knowledge mastery effect and application ability of the large model in a specific domain.
[0129] In one embodiment, the above S20 includes:
[0130] S201, performing keyword extraction on the professional dataset to extract core concepts and professional terms;
[0131] S202, performing semantic role annotation on the core concepts and generating a unique identifier for each core concept using a uniform resource identifier to generate a core concept annotation table;
[0132] S203, performing context relevance annotation on the professional terms to generate a professional term annotation table;
[0133] S204, generating a structured annotation result according to the core concept annotation table and the professional term annotation table.
[0134] In this embodiment, keyword extraction and annotation processing are key links to ensure that the core concepts and professional terms in the professional dataset can be effectively identified, classified, and structurally stored. Through this process, the understanding ability of the large model for professional terms can be improved, providing basic data support for subsequent knowledge graph construction and semantic analysis.
[0135] First, extract keywords from the professional dataset to extract core concepts and professional terms. Keyword extraction uses natural language processing (NLP) techniques, and methods such as term frequency-inverse document frequency (TF-IDF), latent semantic analysis (LSA), or BERT embedding vectors based on deep learning can be used. TF-IDF is suitable for identifying high-frequency but non-general domain-specific terms, while LSA identifies potential core concepts in the text through topic modeling. BERT performs dynamic word vector calculation in combination with context and can more accurately identify the actual meaning of terms. For example, in medical texts, concepts such as "diabetes" and "insulin resistance" can be extracted, and in financial data, terms such as "capital market" and "liquidity risk" can be extracted.
[0136] Then, perform semantic role annotation on the extracted core concepts and generate a unique identifier for each core concept. Semantic role annotation is based on named entity recognition (NER) technology, and the BIO (Beginning-Inside-Outside) annotation method can be used to identify entity categories, such as disease names, company names, legal terms, etc. Uniform Resource Identifiers (URIs) are used to ensure the uniqueness of concepts and can be stored in a standard format (such as DOI, UUID, or RDF-based URI schemes). For example, if the extracted concept is "inflation", a unique identifier URI: finance:concept / inflation can be generated. All annotation results are stored in the core concept annotation table, which contains concept names, definition descriptions, URIs, and associated attribute fields.
[0137] Next, perform context relevance annotation on professional terms to ensure the matching degree of terms with the context. Context relevance annotation mainly uses similarity calculation based on word vectors or co-occurrence analysis based on rules to judge whether a term has a stable semantic association with its context. For example, in medical texts, "hypertension" is usually related to "cardiovascular diseases" rather than "nervous system diseases". At the same time, during the annotation process, synonym relationships, domain-specific interpretations, and recommended usages can be established between terms. For example, a synonym relationship can be established between "myocardial infarction" and "myocardial infarction", and in the financial field, "liquidity" can correspond to different interpretations in different contexts. The annotation results are stored in the professional term annotation table, which contains term names, definition descriptions, synonym lists, and context applicability scores.
[0138] Finally, based on the core concept annotation table and the professional term annotation table, a structured annotation result is generated. The structured annotation result is stored in a standard format (such as JSON, XML, or RDF triples) to ensure data scalability and interoperability. This structured data can be used for subsequent knowledge graph construction, machine learning model training, and large model inference optimization. For example, in medical text analysis tasks, the structured annotation result can be used to train medical language models and improve their recognition and reasoning abilities for professional terms.
[0139] In this embodiment, through keyword extraction and annotation processing, the terms in the professional dataset are standardized and structured, improving the large model's recognition ability for professional terms. Semantic role annotation enhances the logical relevance between concepts and helps improve the accuracy of text parsing. Context relevance annotation ensures the rationality of term usage and avoids understanding errors caused by context deviation. The finally generated structured annotation result can be used for knowledge graph construction, domain model training, and intelligent question answering system optimization, enhancing the large model's reasoning ability and applicability in different fields.
[0140] In one embodiment, the above S30 includes:
[0141] S301, extracting core concept entities from the annotation result and generating a domain entity set of the target domain knowledge graph based on the core concept entities;
[0142] S302, constructing an entity relationship classification system for the target domain knowledge graph that includes hierarchical relationships, causal relationships, and explanatory relationships;
[0143] S303, constructing an attribute feature description framework for the target domain knowledge graph that includes time attributes, space attributes, and logical attributes;
[0144] S304, generating an entity relationship network and an attribute feature library of the target domain knowledge graph based on the entity relationship classification system and the attribute feature description framework.
[0145] In this embodiment, the target domain knowledge graph serves as an evaluation benchmark, enabling the evaluation of the large model to conduct systematic analysis in combination with domain knowledge and improving the accuracy and professionalism of the evaluation. The role of the knowledge graph is to provide entity matching, relationship comparison, and attribute consistency verification, thereby assisting the large model evaluation system in evaluating in dimensions such as semantic accuracy, specification compliance, and concept system depth.
[0146] First, extract the core concept entities from the annotation results as the basic data source for the knowledge graph. These core concepts come from keyword extraction and annotation processing in professional datasets, covering the professional terms and basic knowledge of the target field. For example, in the medical field, core concepts may include "hypertension", "insulin resistance", etc.; in the financial field, they may include "interest rate adjustment", "market liquidity", etc.; in the legal field, they may include "contract establishment", "tort liability", etc. This process can utilize named entity recognition (NER) and natural language processing technology (NLP) to ensure that the entity library of the knowledge graph can cover the domain knowledge required for large model evaluation.
[0147] Next, construct an entity relationship classification system as the relationship matching criterion in the evaluation task. When generating text, the large model may involve hierarchical relationships, causal relationships, or explanatory relationships between domain knowledge. For example, in a medical evaluation task, the model may answer "the causes of diabetes", and its output needs to match causal relationships such as "high-sugar diet" and "insulin resistance", rather than being wrongly associated with "fracture" or "Alzheimer's disease". In a financial evaluation task, the model may answer "the impact of the central bank's interest rate hike", and its output should match "decline in inflation" rather than "long-term stability of the stock market". This kind of matching requires the relationship network provided by the knowledge graph as a benchmark, enabling the evaluation to accurately judge whether the large model correctly understands the logical relationships between concepts.
[0148] Then, establish an attribute feature description framework to support the evaluation of the attribute consistency of the model output content. For example, in a legal evaluation task, the model may need to answer "the conditions for contract revocation", and its output should conform to the time attribute (contract signing time), spatial attribute (contract applicable area), and logical attribute (meeting the revocation conditions stipulated by the Contract Law) stipulated by the legal provisions. If the model output is inconsistent with the legal provisions stored in the knowledge graph, it can be determined that the normative compliance of its answer is relatively low.
[0149] Finally, based on the entity relationship classification system and the attribute feature description framework, construct the entity relationship network and attribute feature library of the knowledge graph in the target field, and use them in the large model evaluation task to provide a structured knowledge benchmark to support the semantic accuracy analysis, relationship matching analysis, and normative compliance scoring of the large model evaluation system.
[0150] In the evaluation task of medical large models, the target domain knowledge graph is used to determine whether the model accurately understands medical knowledge such as diseases, drugs, and treatment plans. For example, the evaluation task may require the model to answer "Can diabetic patients take aspirin?" The answer of the model needs to match the medical guideline information in the knowledge graph to ensure that it meets clinical standards. If the model wrongly suggests that "diabetic patients can take aspirin for a long time", while the knowledge graph clearly indicates that "long-term use of aspirin may increase the risk of bleeding", the evaluation system can determine that the model has a low degree of compliance with medical knowledge norms.
[0151] In the evaluation task of financial large models, the target domain knowledge graph is used to verify whether the model correctly understands economic concepts and the logic of market operation. For example, the evaluation task may require the model to answer "the impact of interest rate adjustment on the stock market". If the model outputs that "raising interest rates usually drives the stock market up", while the economic relationship marked in the knowledge graph is that "raising interest rates usually leads to a decline in the stock market", the evaluation system can determine that the model has an incorrect relationship match in economic knowledge.
[0152] In the evaluation task of legal large models, the target domain knowledge graph is used to detect whether the model's answers comply with legal provisions. For example, the evaluation task may require the model to answer "the conditions for contract revocation". If the model answers that "a contract can be revoked at any time", while the legal provisions marked in the knowledge graph are that "contract revocation must meet specific legal conditions, such as fraud, material misunderstanding, or gross inequality of bargaining power", the evaluation system can determine that the model has a low degree of compliance with legal knowledge norms.
[0153] In this embodiment, by constructing the target domain knowledge graph and using it in the evaluation task of large models, the professionalism and accuracy of the large model evaluation system can be effectively improved. The entity matching ability of the knowledge graph enables the evaluation task to accurately judge whether the model output content contains correct domain concepts; the relationship comparison ability enables the evaluation task to detect whether the model correctly understands the logical relationship between concepts; the attribute consistency verification enables the evaluation system to ensure that the model output content meets the domain specifications and standard requirements.
[0154] In one embodiment, the above S40 includes:
[0155] S401, classifying the professional data set into a domain core text set, an annotated literature set, and a derivative material set to form the professional corpus;
[0156] S402, performing semantic annotation on the core concepts in the domain core text set to generate core concept labels;
[0157] S403, performing tagging processing on the specification relevance content in the annotated literature set to generate specification association labels;
[0158] S404. Perform case annotation on the actual application scenario content in the derivative data set to generate application scenario tags.
[0159] S405. Extract knowledge points from the professional corpus based on the core concept tags, specification association tags, and application scenario tags.
[0160] S406. Generate the analysis task set according to the extracted knowledge points.
[0161] In this embodiment, by constructing a professional corpus, a high-quality data foundation is provided for subsequent analysis tasks. First, classify the professional data set so that different types of data can play their respective roles in the analysis tasks. The professional data set is divided into a domain core text set, an annotated literature set, and a derivative data set. Among them, the domain core text set contains the most representative literature in the field, such as the "Diagnosis and Treatment Guidelines" in the medical field, etc.; the annotated literature set covers the interpretations of experts and scholars, research papers, paraphrases, etc., and is used to expand and explain in detail the concepts of the core text; the derivative data set includes the application scenarios, case studies, practical experiences, etc. in the field, and provides practical application support for theoretical knowledge.
[0162] After completing the data classification, perform annotation on different data sets to enhance the usability of the data. Core concept annotation is mainly used for the domain core text set, performing semantic annotation on key concepts to form core concept tags. For example, in the medical field, professional terms such as "cancer staging" and "immune response" are annotated. Specification relevance annotation is used for the annotated literature set to mark the relevance of the text to domain standards or rules, such as the citation of the "Securities Law" in financial literature. Application scenario annotation is applied to the derivative data set to mark specific practical scenarios, such as the impact of financial market fluctuations on investment strategies, etc., to generate application scenario tags.
[0163] After completing the annotation, use the core concept tags, specification association tags, and application scenario tags to extract knowledge points from the professional corpus. These knowledge points can include medical diagnosis standards, key factors of financial models, etc., and provide a basis for the generation of analysis tasks. Finally, based on the extracted knowledge points, generate an analysis task set, and these tasks are used for subsequent model evaluation. For example, in medical evaluation, tasks such as "check whether the model correctly explains the causes of 'diabetes complications'" can be generated.
[0164] In the medical large model evaluation task, medical literature is first classified. For example, the "NCCN Oncology Guidelines" is used as the core text, doctors' interpretations of the guidelines are used as annotated literature, and clinical cases are used as derivative materials. Then, medical core terms such as "tumor staging" and "immunotherapy" are labeled to generate core concept tags. For the literature interpreting the guidelines, the medical standards involved are marked, such as "According to the WHO classification, lung cancer is divided into small cell lung cancer and non-small cell lung cancer", to form standardized association tags. For clinical cases, such as the differences in the tolerance of different patients to chemotherapy, application scenario tags are formed. Finally, the generated analysis tasks include "evaluating whether the model correctly interprets the treatment plan for 'non-small cell lung cancer'" and "checking whether the model's application of 'PD-L1 inhibitors' complies with the guidelines".
[0165] In this embodiment, by constructing a professional corpus based on a professional data set and extracting knowledge points from it to generate analysis tasks, the professionalism and pertinence of the large model evaluation task are improved. First, by classifying the data set, different types of data can be effectively utilized in the evaluation task, ensuring that the evaluation task can cover multiple levels of theoretical knowledge, professional interpretation, and practical application. Second, through core concept annotation, standardized association annotation, and application scenario annotation, the resolution of the evaluation system for the large model's understanding ability is improved, enabling the evaluation task to more finely check the semantic accuracy, standard compliance, and practical application ability of the model. Finally, the analysis tasks generated based on knowledge points make the evaluation task more targeted, capable of accurately measuring the performance of the model in the field of professional knowledge, and providing a reliable basis for the optimization of the large model.
[0166] In one embodiment, the above S50 includes:
[0167] S501, according to the characteristics of the professional data set and the content in the professional corpus, determine the analysis dimension categories of the multi-dimensional analysis index system, and the analysis dimension categories include but are not limited to the semantic accuracy analysis dimension, the standard compliance analysis dimension, and the concept system depth analysis dimension;
[0168] S502, based on the analysis dimension categories, construct the hierarchical structure of the multi-dimensional analysis index system;
[0169] S503, for each analysis task in the analysis task set, assign at least one corresponding analysis dimension according to the task objective and content characteristics, and establish the corresponding relationship between the analysis task and the analysis dimension based on the multi-dimensional analysis index system.
[0170] In this embodiment, a multi-dimensional analysis index system is constructed to enable the analysis task to be evaluated based on scientific and reasonable evaluation criteria, ensuring the objectivity and professionalism of the evaluation results. First, the categories of analysis dimensions need to be determined. Based on the characteristics of the professional dataset and the content of the professional corpus, the evaluation dimensions applicable to the analysis task are summarized. The design of the analysis dimensions covers three core aspects: semantic accuracy, compliance with norms, and depth of the concept system. Among them, the semantic accuracy analysis dimension is used to measure the precision of word meanings, context consistency, and grammatical rationality of the text; the compliance with norms analysis dimension is used to evaluate whether the text content conforms to industry standards, policies and regulations, or the domain knowledge system; the depth of the concept system analysis dimension is used to judge the hierarchy, abstract progression, and logical complexity of the text content in terms of concept understanding.
[0171] Secondly, when constructing the hierarchical structure of the analysis dimensions, the above analysis dimensions need to be further refined to make them quantifiable and applicable to different tasks. For example, the semantic accuracy dimension can be refined into term accuracy and context consistency. Term accuracy focuses on whether the professional terms in the text conform to the standard definitions, while context consistency measures the coherence degree of the text in the context. The compliance with norms dimension can be refined into the matching degree with the standard answer and logical compliance. The former examines the consistency between the text and the authoritative answer, and the latter is used to check whether the text content violates logic or standards. The depth of the concept system dimension is refined into the complexity of logical association and the progression of abstract levels. The former measures the complexity of the logical relationship between concepts in the text, and the latter evaluates whether the text reflects the logical progression of concepts from shallow to deep.
[0172] After completing the construction of the dimension hierarchy, it is necessary to set quantification criteria and define evaluation criteria for each sub-index. For example, the term accuracy sub-index can be evaluated based on the definition consistency threshold. If the matching degree of the terms in the text with the standard definition is lower than the set threshold, it is determined that the terms are misused; the logical compliance sub-index can be detected using the relationship network matching degree threshold. If the text content conflicts with the relationship network stored in the knowledge graph during the logical reasoning process, it is determined that the logic is non-compliant.
[0173] Establishing the mapping relationship between tasks and dimensions is a key link in the analysis task design. During this process, at least one corresponding evaluation dimension needs to be assigned according to the goals and content characteristics of the analysis task. Finally, by generating a task-dimension mapping table, the association relationship between each analysis task and semantic accuracy, compliance with norms, or depth of the concept system is clarified, making the evaluation process of the analysis task traceable and consistent.
[0174] Finally, a comprehensive evaluation standard is set to ensure that the scoring method of the analysis task is reasonable and comparable. Specifically, it can include:
[0175] Allocate the weight ratios of evaluation dimensions. For example, set the weight ratio of the semantic accuracy dimension to 50%, the specification compliance degree to 30%, and the depth of the concept system to 20%.
[0176] Determine the scoring method. For example, for the sub-index of term accuracy, a weighted calculation method can be used to calculate the term matching degree; for the sub-index of logical compliance, a deduction rule for conflicting items can be adopted. If there are relationship items in the text that violate logic, the corresponding score will be deducted from the total score.
[0177] Generate a comprehensive evaluation output standard, clarify the scoring rules and weight distribution of each sub-index, so that the final analysis result has a high degree of credibility.
[0178] In the evaluation task of medical large models, the analysis task may involve "judging whether the model correctly interprets the pathological mechanism of 'type 2 diabetes'". For this task, it is necessary to associate the semantic accuracy dimension (judging whether the concept of 'insulin resistance' is correct), the standard answer matching degree (whether the model's answer conforms to medical guidelines), and the logical association complexity (evaluating whether the model correctly describes the multi-factor influence of 'glucose metabolism imbalance'). After constructing the evaluation dimensions, set the scoring criteria. For example, if the term matching degree is lower than 75%, points will be deducted; if the consistency of comparison with medical guidelines is lower than 85%, it will be determined as content deviation. Finally, a comprehensive evaluation will be carried out according to the term matching ratio of 40%, the standard answer matching degree ratio of 40%, and the logical complexity ratio of 20%.
[0179] In the evaluation task of financial large models, the analysis task may involve "judging whether the model correctly analyzes the impact of interest rate hikes on the stock market". The evaluation dimensions of this task include the term accuracy dimension (judging whether the concept of 'interest rate hike' is correct), the specification compliance degree dimension (judging whether the model quotes authoritative economic theories), and the abstract level progression dimension (judging whether the model analyzes layer by layer starting from macroeconomic principles). After constructing the dimension hierarchical structure, set the scoring rules. For example, if the term matching degree is lower than 85%, 1 point will be deducted; if authoritative economic literature is not cited, 2 points will be deducted. Finally, a weighted calculation will be carried out according to the term matching ratio of 30%, the specification compliance degree ratio of 50%, and the abstract level progression ratio of 20% to output the evaluation result.
[0180] In this embodiment, by constructing a multi-dimensional analysis index system, the evaluation results of the large model are made more comprehensive, scientific, and quantifiable. First, by refining the evaluation dimension categories, the evaluation system can accurately evaluate different task types, improving the pertinence of the evaluation. Second, by establishing a hierarchical structure, the analysis dimensions are broken down into multiple sub-indicators, making the evaluation process have higher resolution and traceability. Further, by setting quantitative criteria, the objectivity of the evaluation system is achieved, making the evaluation results of the large model consistent and comparable. Finally, by combining task dimension mapping and scoring criterion setting, the analysis tasks and evaluation indicators form a systematic association, providing an effective evaluation basis for the optimization of the large model.
[0181] In one embodiment, the above S70 includes:
[0182] S701, extracting domain entities in the task output result, matching the domain entities with the domain entity set in the target domain knowledge graph, and recording the successfully matched entity items and the failed-to-match entity items;
[0183] S702, extracting the entity relationships in the task output result, constructing a task output relationship network according to the entity relationships, and comparing the task output relationship network with the entity relationship network in the target domain knowledge graph, and recording the relationship comparison anomaly items;
[0184] S703, extracting the attribute features of the domain entities in the task output result, comparing the attribute features with the attribute feature library in the target domain knowledge graph, and recording the attribute items with successfully matched attribute features, the attribute items with failed-to-match attribute features, and the attribute items beyond the attribute standard range in the target domain knowledge graph;
[0185] S704, generating a knowledge verification result according to the successfully matched entity items and the failed-to-match entity items, the relationship comparison anomaly items, the failed-to-match attribute items, and the attribute items beyond the attribute standard range in the target domain knowledge graph.
[0186] In this embodiment, knowledge verification aims to evaluate the accuracy, consistency, and rationality of the task output result based on the target domain knowledge graph to generate a knowledge verification result. First, extract the domain entities in the task output result, identify the key entities in the task output through text parsing technology, and match them with the domain entity set in the target domain knowledge graph. The successfully matched entity items indicate that the content output by the model conforms to the standard knowledge in the knowledge graph, while the failed-to-match entity items mean that the model may have used incorrect terms or failed to correctly understand related concepts.
[0187] Secondly, extract the entity relationships in the task output results, construct a task output relationship network, and compare it with the entity relationship network in the target domain knowledge graph. Entity relationship refers to the association between entities in different domains, such as causal relationship, hierarchical relationship, or parallel relationship.
[0188] Furthermore, extract the attribute features of the domain entities in the task output results and compare them with the attribute feature library in the knowledge graph. At the same time, if the description of a certain concept by the model conforms to the attribute settings in the knowledge graph, this item should be recorded as an attribute item with successful attribute feature matching; if the attribute information described by the model does not exist in the knowledge graph, it should be marked as an attribute item with failed matching.
[0189] Finally, combine all the matching information in the above verification process to comprehensively generate a knowledge verification result, including entity items with successful and failed matches, relationship comparison abnormal items, attribute items with failed matches, and attribute items exceeding the standard range. This result can be used in subsequent analysis processes to evaluate the model's knowledge mastery in specific tasks.
[0190] In this embodiment, by constructing a knowledge verification system based on the knowledge graph, the reliability and accuracy of the large model analysis task are improved. First, through domain entity matching, ensure that the terms output by the model are consistent with the standard definitions in the knowledge base to avoid semantic misuse. Secondly, through entity relationship comparison, verify whether the logical relationships in the task output conform to the domain knowledge system to ensure the rigor of the reasoning process. Furthermore, through attribute feature comparison, check whether the numerical values, time, and description information output by the model are within the standard range to avoid factual errors. Finally, combining the knowledge verification results, the knowledge mastery of the model can be systematically evaluated, providing accurate feedback for the optimization and application of the large model.
[0191] In one embodiment, the above S80 includes:
[0192] S801, extract the entity matching report, relationship comparison report, and attribute consistency report in the knowledge verification result;
[0193] S802, based on the entity matching report, analyze the entity matching success rate according to the semantic accuracy analysis dimension in the multi-dimensional analysis index system;
[0194] S803, based on the relationship comparison report, analyze the proportion of relationship comparison abnormal items according to the specification fit analysis dimension in the multi-dimensional analysis index system;
[0195] S804, based on the attribute consistency report, analyze the attribute consistency matching rate according to the concept system depth analysis dimension in the multi-dimensional analysis index system;
[0196] S805. Generate the analysis result according to the entity matching success rate of the semantic accuracy analysis dimension, the proportion of relationship comparison anomaly items in the specification compliance analysis dimension, and the attribute consistency matching rate of the concept system depth analysis dimension.
[0197] In this embodiment, based on the multi-dimensional analysis index system, the knowledge verification result is analyzed to generate the final analysis result of the evaluation task. First, extract the entity matching report, relationship comparison report, and attribute consistency report from the knowledge verification result. These reports are derived from the entity matching, relationship comparison, and attribute feature consistency detection of the task output result during the previous knowledge verification process. The entity matching report records whether the domain entities involved in the model output content can be matched with the knowledge graph. The relationship comparison report stores the comparison of the relationship descriptions in the task output result with the relationship network of the knowledge graph. The attribute consistency report shows whether the attribute features involved in the task output conform to the standards in the knowledge graph.
[0198] Second, based on the entity matching report, calculate the entity matching success rate according to the semantic accuracy analysis dimension in the multi-dimensional analysis index system. The semantic accuracy analysis dimension is used to measure the degree of understanding of domain concepts by the model in the task output. The entity matching success rate evaluates the model's ability to correctly identify and use terms and concepts by calculating the proportion of successfully matched entity items in the total number of task output entities. For example, if the task output result involves 10 domain entities and 8 of them can be matched with the knowledge graph, the entity matching success rate of the semantic accuracy analysis dimension is 80%.
[0199] Then, based on the relationship comparison report, analyze the proportion of relationship comparison anomaly items according to the specification compliance analysis dimension in the multi-dimensional analysis index system. The specification compliance analysis dimension mainly measures whether the logical structure of the task output result conforms to the domain knowledge system. The calculation method of the proportion of relationship comparison anomaly items is: proportion of abnormal relationship items = (number of relationship items that do not conform to the known relationships in the knowledge graph) / (total number of relationship items in the task output). For example, if the task output result involves 5 entity relationships and 2 of them are inconsistent with the standard relationships in the knowledge graph, the proportion of relationship comparison anomaly items is 40%. This indicator is used to evaluate whether the model can correctly organize and express the logical relationships between domain knowledge.
[0200] Next, based on the attribute consistency report, calculate the attribute consistency matching rate according to the concept system depth analysis dimension in the multi-dimensional analysis index system. The concept system depth analysis dimension is used to measure the model's mastery of domain attribute features in the task output. The attribute consistency matching rate evaluates the model's ability to control the details of domain information by calculating the proportion of items that conform to the standard attributes in the knowledge graph in the task output. For example, if the task output result contains 15 attribute descriptions, and 12 of them conform to the standard descriptions in the knowledge graph, the attribute consistency matching rate is 80%.
[0201] Finally, combine the entity matching success rate in the semantic accuracy analysis dimension, the proportion of abnormal items in the relationship comparison in the specification compliance analysis dimension, and the attribute consistency matching rate in the concept system depth analysis dimension to generate the final analysis result. This analysis result is used to comprehensively evaluate the model's knowledge understanding and expression ability in different tasks and can be used as an important basis for subsequent optimization of the large model.
[0202] In this embodiment, by constructing a multi-dimensional analysis index system and systematically analyzing the knowledge verification results, the evaluation of the large model can cover multiple aspects such as term matching, relationship logic, and concept system depth, ensuring the accuracy and reliability of the evaluation results. Through the entity matching success rate based on the semantic accuracy analysis dimension, the accuracy of the model in using domain terms can be evaluated; through the proportion of abnormal items in the relationship comparison based on the specification compliance analysis dimension, the deviation of the model in knowledge logical reasoning can be identified; through the attribute consistency matching rate based on the concept system depth analysis dimension, the depth of the model's mastery of domain information can be measured. Combining these analysis results can be used to optimize the training data of the large model, adjust the model parameters, and improve the evaluation method, thereby enhancing the application ability of the large model in the professional field.
[0203] In one embodiment, a model analysis device based on a knowledge graph is provided. The model analysis device based on the knowledge graph corresponds one-to-one with the model analysis method based on the knowledge graph in the above embodiment. Refer to Figure 3 , Figure 3 is a schematic diagram of the functional modules of a preferred embodiment of the model analysis device based on the knowledge graph of the present invention. Data processing module 10, keyword extraction and annotation module 20, knowledge graph construction module 30, analysis task generation module 40, evaluation index management module 50, task execution module 60, knowledge verification module 70, and result analysis module 80. The detailed description of each functional module is as follows:
[0204] The data processing module 10 is used to collect target domain data, screen and classify the target domain data, and generate a professional data set;
[0205] The keyword extraction and annotation module 20 is used to perform keyword extraction and annotation processing on the core concepts and professional terms in the professional dataset, and generate an annotation result;
[0206] The knowledge graph construction module 30 is used to construct a target domain knowledge graph including a domain entity set, an entity relationship network, and an attribute feature library based on the annotation result;
[0207] The analysis task generation module 40 is used to establish a professional corpus based on the professional dataset, and generate an analysis task set based on the professional corpus;
[0208] The evaluation index management module 50 is used to construct a multi-dimensional analysis index system, and each analysis task in the analysis task set is at least associated with one analysis dimension in the multi-dimensional analysis index system;
[0209] The task execution module 60 is used to input the analysis task into the model to be analyzed, and obtain the task output result of the model to be analyzed;
[0210] The knowledge verification module 70 is used to perform knowledge verification on the task output result according to the target domain knowledge graph, and generate a knowledge verification result;
[0211] The result analysis module 80 is used to analyze the knowledge verification result based on the multi-dimensional analysis index system, and generate an analysis result.
[0212] In one embodiment, the keyword extraction and annotation module 20 is specifically used for:
[0213] Perform keyword extraction on the professional dataset to extract core concepts and professional terms;
[0214] Perform semantic role annotation on the core concepts, and generate a unique identifier for each core concept using a uniform resource identifier, generating a core concept annotation table;
[0215] Perform context relevance annotation on the professional terms, generating a professional term annotation table;
[0216] Generate a structured annotation result according to the core concept annotation table and the professional term annotation table.
[0217] In one embodiment, the knowledge graph construction module 30 is specifically used for:
[0218] Extract core concept entities from the annotation result, and generate a domain entity set of the target domain knowledge graph based on the core concept entities;
[0219] Construct an entity relationship classification system for the target domain knowledge graph including hierarchical relationships, causal relationships, and explanatory relationships;
[0220] Construct an attribute feature description framework including time attributes, space attributes, and logical attributes for the target domain knowledge graph;
[0221] Generate an entity relationship network and an attribute feature library of the target domain knowledge graph based on the entity relationship classification system and the attribute feature description framework.
[0222] In one embodiment, the analysis task generation module 40 is specifically configured to:
[0223] Classify the professional data set into a domain core text set, an annotated literature set, and a derivative data set to form the professional corpus;
[0224] Perform semantic annotation on the core concepts in the domain core text set to generate core concept labels;
[0225] Perform tagging on the canonical relevance content in the annotated literature set to generate canonical association labels;
[0226] Perform case annotation on the actual application scenario content in the derivative data set to generate application scenario labels;
[0227] Extract knowledge points from the professional corpus based on the core concept labels, canonical association labels, and application scenario labels;
[0228] Generate the analysis task set according to the extracted knowledge points.
[0229] In one embodiment, the evaluation index management module 50 is specifically configured to:
[0230] Determine the analysis dimension categories of the multi-dimensional analysis index system according to the characteristics of the professional data set and the content in the professional corpus. The analysis dimension categories include, but are not limited to, the semantic accuracy analysis dimension, the canonical compliance analysis dimension, and the concept system depth analysis dimension;
[0231] Construct a hierarchical structure of the multi-dimensional analysis index system based on the analysis dimension categories;
[0232] For each analysis task in the analysis task set, assign at least one corresponding analysis dimension according to the task objective and content characteristics, and establish a correspondence between the analysis task and the analysis dimension based on the multi-dimensional analysis index system.
[0233] In one embodiment, the knowledge verification module 70 is specifically configured to:
[0234] Extract the domain entities in the task output result, match the domain entities with the set of domain entities in the target domain knowledge graph, and record the entity items with successful matches and the entity items with failed matches;
[0235] Extract the entity relationships in the task output result, construct a task output relationship network based on the entity relationships, and compare the task output relationship network with the entity relationship network in the target domain knowledge graph, and record the items with abnormal relationship comparisons;
[0236] Extract the attribute features of the domain entities in the task output result, compare the attribute features with the attribute feature library in the target domain knowledge graph, and record the attribute items with successful attribute feature matches, the attribute items with failed matches, and the attribute items that exceed the attribute standard range in the target domain knowledge graph;
[0237] Generate a knowledge verification result based on the entity items with successful matches and the entity items with failed matches, the items with abnormal relationship comparisons, the attribute items with failed matches, and the attribute items that exceed the attribute standard range in the target domain knowledge graph.
[0238] In one embodiment, the result analysis module 80 is specifically configured to:
[0239] Extract the entity matching report, relationship comparison report, and attribute consistency report in the knowledge verification result;
[0240] Based on the entity matching report, analyze the entity matching success rate according to the semantic accuracy analysis dimension in the multi-dimensional analysis index system;
[0241] Based on the relationship comparison report, analyze the proportion of items with abnormal relationship comparisons according to the specification fit analysis dimension in the multi-dimensional analysis index system;
[0242] Based on the attribute consistency report, analyze the attribute consistency matching rate according to the concept system depth analysis dimension in the multi-dimensional analysis index system;
[0243] Generate the analysis result according to the entity matching success rate in the semantic accuracy analysis dimension, the proportion of items with abnormal relationship comparisons in the specification fit analysis dimension, and the attribute consistency matching rate in the concept system depth analysis dimension.
[0244] In one embodiment, a computer device is provided. The computer device can be a server, and its internal structure diagram can be as Figure 4As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the server side of a model analysis method based on a knowledge graph.
[0245] In one embodiment, a computer device is provided. The computer device can be a client, and its internal structure diagram can be as Figure 5 shown in the figure. The computer device includes a processor, a memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server through a network connection. When the computer program is executed by the processor, it realizes the functions or steps on the client side of a model analysis method based on a knowledge graph
[0246] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented:
[0247] Collect target domain data, and perform screening and classification processing on the target domain data to generate a professional data set;
[0248] Perform keyword extraction and annotation processing on the core concepts and professional terms in the professional data set to generate an annotation result;
[0249] Construct a target domain knowledge graph including a domain entity set, an entity relationship network, and an attribute feature library based on the annotation result;
[0250] Establish a professional corpus based on the professional data set, and generate an analysis task set based on the professional corpus;
[0251] Construct a multi-dimensional analysis index system, and each analysis task in the analysis task set is at least associated with one analysis dimension in the multi-dimensional analysis index system;
[0252] Input the analysis task into the model to be analyzed, and obtain the task output result of the model to be analyzed;
[0253] Perform knowledge verification on the task output result according to the target domain knowledge graph, and generate a knowledge verification result;
[0254] Analyze the knowledge verification result based on the multi-dimensional analysis index system, and generate an analysis result.
[0255] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0256] Collect target domain data, and perform screening and classification processing on the target domain data to generate a professional data set;
[0257] Perform keyword extraction and annotation processing on the core concepts and professional terms in the professional data set to generate an annotation result;
[0258] Construct a target domain knowledge graph including a domain entity set, an entity relationship network, and an attribute feature library based on the annotation result;
[0259] Establish a professional corpus based on the professional data set, and generate an analysis task set based on the professional corpus;
[0260] Construct a multi-dimensional analysis index system, and each analysis task in the analysis task set is associated with at least one analysis dimension in the multi-dimensional analysis index system;
[0261] Input the analysis task into the model to be analyzed, and obtain the task output result of the model to be analyzed;
[0262] Perform knowledge verification on the task output result according to the target domain knowledge graph, and generate a knowledge verification result;
[0263] Analyze the knowledge verification result based on the multi-dimensional analysis index system, and generate an analysis result.
[0264] It should be noted that for the functions or steps that can be achieved by the above computer-readable storage medium or computer device, reference can be made to the relevant descriptions on the server side and the user side in the foregoing method embodiments. To avoid repetition, they will not be described in detail here.
[0265] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in this application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.
[0266] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0267] It should be noted that if non-company software tools or components appear in the embodiments of this application, they are only used for illustrative introduction and do not represent actual use. The above embodiments are only used to illustrate the technical solutions of the present invention, not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for analyzing a model based on a knowledge graph, characterized in that, Including the following steps: Collect data in the target domain, and screen and classify the target domain data to generate a professional dataset; Perform keyword extraction and annotation processing on the core concepts and professional terms in the professional dataset to generate an annotation result; Construct a target domain knowledge graph including a domain entity set, an entity relationship network, and an attribute feature library based on the annotation result; Establish a professional corpus based on the professional dataset, and generate an analysis task set based on the professional corpus; Construct a multi-dimensional analysis index system, and each analysis task in the analysis task set is associated with at least one analysis dimension in the multi-dimensional analysis index system; Input the analysis task into the model to be analyzed, and obtain the task output result of the model to be analyzed; Perform knowledge verification on the task output result according to the target domain knowledge graph to generate a knowledge verification result; Analyze the knowledge verification result based on the multi-dimensional analysis index system to generate an analysis result.
2. The method for model analysis based on a knowledge graph according to claim 1, wherein Perform keyword extraction and annotation processing on the core concepts and professional terms in the professional dataset to generate an annotation result, including: Extract keywords from the professional dataset to extract core concepts and professional terms; Perform semantic role annotation on the core concepts, and generate a unique identifier for each core concept using a uniform resource identifier to generate a core concept annotation table; Perform context relevance annotation on the professional terms to generate a professional term annotation table; Generate a structured annotation result according to the core concept annotation table and the professional term annotation table.
3. The model analysis method based on a knowledge graph according to claim 1, wherein Construct a target domain knowledge graph including a domain entity set, an entity relationship network, and an attribute feature library based on the annotation result, including: Extract core concept entities from the annotation result, and generate a domain entity set of the target domain knowledge graph based on the core concept entities; Construct an entity relationship classification system for the target domain knowledge graph including hierarchical relationships, causal relationships, and explanatory relationships; Construct an attribute feature description framework for the target domain knowledge graph including time attributes, space attributes, and logical attributes; Generate the entity relationship network and attribute feature library of the target domain knowledge graph based on the entity relationship classification system and the attribute feature description framework.
4. The method for model analysis based on a knowledge graph according to claim 1, wherein Establish a professional corpus based on the professional dataset, and generate an analysis task set based on the professional corpus, including: Classify the professional dataset into a domain core text set, an annotation literature set, and a derivative data set to form the professional corpus; Perform semantic annotation on the core concepts in the domain core text set to generate core concept labels; Perform tagging on the normative relevance content in the annotation literature set to generate normative association labels; Perform case annotation on the actual application scenario content in the derivative data set to generate application scenario labels; Extract knowledge points from the professional corpus based on the core concept labels, normative association labels, and application scenario labels; Generate the analysis task set according to the extracted knowledge points.
5. The model analysis method based on a knowledge graph according to claim 1, wherein Construct a multi-dimensional analysis index system, where each analysis task in the set of analysis tasks is associated with at least one analysis dimension in the multi-dimensional analysis index system, including: Determine the analysis dimension categories of the multi-dimensional analysis index system according to the characteristics of the professional data set and the content in the professional corpus. The analysis dimension categories include, but are not limited to, the semantic accuracy analysis dimension, the specification compliance analysis dimension, and the concept system depth analysis dimension; Based on the analysis dimension categories, construct a hierarchical structure of the multi-dimensional analysis index system; For each analysis task in the set of analysis tasks, allocate at least one corresponding analysis dimension according to the task objective and content characteristics, and establish a correspondence between the analysis task and the analysis dimension based on the multi-dimensional analysis index system.
6. The method for model analysis based on a knowledge graph according to claim 1, wherein Perform knowledge verification on the task output result according to the target domain knowledge graph, and generate a knowledge verification result, including: Extract the domain entities in the task output result, and match the domain entities with the set of domain entities in the target domain knowledge graph, and record the entity items with successful matches and the entity items with failed matches; Extract the entity relationships in the task output result, construct a task output relationship network according to the entity relationships, and compare the task output relationship network with the entity relationship network in the target domain knowledge graph, and record the relationship comparison abnormal items; Extract the attribute characteristics of the domain entities in the task output result, and compare the attribute characteristics with the attribute feature library in the target domain knowledge graph, and record the attribute items with successful matches, the attribute items with failed matches, and the attribute items that exceed the attribute standard range in the target domain knowledge graph; Generate a knowledge verification result according to the entity items with successful matches and the entity items with failed matches, the relationship comparison abnormal items, the attribute items with failed matches, and the attribute items that exceed the attribute standard range in the target domain knowledge graph.
7. The method for model analysis based on a knowledge graph according to claim 1, wherein Analyze the knowledge verification result based on the multi-dimensional analysis index system to generate an analysis result, including: Extract the entity matching report, relationship comparison report, and attribute consistency report in the knowledge verification result; Based on the entity matching report, analyze the entity matching success rate according to the semantic accuracy analysis dimension in the multi-dimensional analysis index system; Based on the relationship comparison report, analyze the proportion of relationship comparison abnormal items according to the specification compliance analysis dimension in the multi-dimensional analysis index system; Based on the attribute consistency report, analyze the attribute consistency matching rate according to the concept system depth analysis dimension in the multi-dimensional analysis index system; Generate the analysis result according to the entity matching success rate in the semantic accuracy analysis dimension, the proportion of relationship comparison abnormal items in the specification compliance analysis dimension, and the attribute consistency matching rate in the concept system depth analysis dimension.
8. A model analysis device based on a knowledge graph, characterized in that, The model analysis device based on the knowledge graph includes: A data processing module for collecting target domain data, screening and classifying the target domain data, and generating a professional data set; A keyword extraction and annotation module for performing keyword extraction and annotation processing on the core concepts and professional terms in the professional data set to generate an annotation result; A knowledge graph construction module, configured to construct a target domain knowledge graph including a domain entity set, an entity relationship network, and an attribute feature library based on the annotation result; An analysis task generation module, configured to establish a professional corpus based on the professional data set, and generate an analysis task set based on the professional corpus; An evaluation index management module, configured to construct a multi-dimensional analysis index system, and each analysis task in the analysis task set is associated with at least one analysis dimension in the multi-dimensional analysis index system; A task execution module, configured to input the analysis task into a model to be analyzed, and obtain a task output result of the model to be analyzed; A knowledge verification module, configured to perform knowledge verification on the task output result according to the target domain knowledge graph, and generate a knowledge verification result; A result analysis module, configured to analyze the knowledge verification result based on the multi-dimensional analysis index system, and generate an analysis result.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a knowledge graph-based model analysis program stored on the memory and executable on the processor. When the knowledge graph-based model analysis program is executed by the processor, the steps of the knowledge graph-based model analysis method according to any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, A knowledge graph-based model analysis program is stored on the storage medium. When the knowledge graph-based model analysis program is executed by a processor, the steps of the knowledge graph-based model analysis method according to any one of claims 1-7 are implemented.
Citation Information
Cited By
Oil and gas field service management and control method and device
CN120723756A
Oil and gas field service management method and device
CN120723756B
Task disassembly and federation allocation method and system based on multi-level knowledge graph
CN120822800A
Government affair knowledge graph ontology construction and optimization method and device, equipment and medium
CN120930760A
Legal information analysis method and system, computer device, medium and product
CN120973884A