Data assessment method, device and system based on large model and storage medium

Through large models, collect and analyze multi-source heterogeneous data, generate knowledge graphs, and dynamically adjust assessment indicators, the lag and inaccuracy of traditional performance appraisal methods are solved, effective processing of unstructured data and dynamic adjustment of scores are achieved, and the accuracy and efficiency of assessment results are improved.

CN120355308AActive Publication Date: 2025-07-22CETC BIGDATA RES INST CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510846474.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-07-22
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

Traditional performance appraisal methods are difficult to adapt to changes in the external environment, lack the ability to process unstructured data, and lack a dynamic adjustment mechanism to score rules, resulting in lag and inaccurate assessment results.

Method used

A large model is used to collect multi-source heterogeneous data by deploying the target Agent, and generate a knowledge graph of policy subjects-indicator items-quantitative standards. Combined with graph neural networks and multimodal scoring engines, the assessment indicators are dynamically adjusted, and the semantic and spatiotemporal features are extracted using RoBERTa-wwm and 3D-ResNet, a confidence compensation mechanism is introduced, and the dynamic routing controller adjusts the scoring path.

Benefits of technology

It has achieved a timely response to changes in policies and regulations, comprehensively handled a variety of assessment materials, improved the accuracy and stability of assessment results, reduced manual intervention, and improved the unity of assessment efficiency and standards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355308A_ABST
    Figure CN120355308A_ABST
Patent Text Reader

Abstract

The invention discloses a data assessment method, device and system based on a large model and a storage medium. The method comprises the steps of collecting multi-source heterogeneous data; analyzing the multi-source heterogeneous data, and generating an index knowledge graph containing a policy subject-index item-quantitative standard triple; calculating to obtain a weight distribution matrix Wp of the priorities; in a multi-modal scoring engine, a RoBERTa-wwm model is adopted to extract a deep semantic feature vector Vt; extracting a spatio-temporal feature vector Vv through 3D-ResNet; when the cosine similarity delta of the Vt and the Vv is lower than a threshold gamma, generating a confidence coefficient compensation parameter beta through a generator; generating a path weight vector according to the confidence coefficient compensation parameter beta; calculating early fusion path output, middle fusion path output or late fusion path output; calculating a final score vector Sf; and outputting a final assessment score value based on the final score vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of big data technology, and particularly to a data assessment method, device, system, and storage medium based on a large model. Background Art

[0002] In the management process of various organizations and institutions, performance assessment is an important means to measure the operation effect of personnel, departments, or systems. Traditional performance assessment methods mainly rely on manually set assessment indicators and scoring criteria, usually based on quantitative data (such as KPI indicators) and qualitative evaluations (such as subjective scores) for assessment. However, these methods have certain limitations in practical applications.

[0003] First of all, the setting of assessment indicators often depends on the experience of managers and is difficult to adapt to the dynamic changes in the external environment. For example, the adjustment of policies and regulations may affect the assessment criteria, and traditional methods lack the ability to respond to external changes in a timely manner, resulting in the lag of assessment indicators behind actual needs.

[0004] Secondly, the forms of assessment materials are becoming increasingly diverse. In addition to traditional text data, they also include unstructured data such as videos and images. However, existing assessment methods mainly focus on structured data, and the processing ability for unstructured data is limited, resulting in insufficient utilization of information in the assessment process and affecting the comprehensiveness and accuracy of the evaluation.

[0005] In addition, the assessment process usually adopts fixed scoring rules and lacks a dynamic adjustment mechanism for different data qualities. When the assessment data has uncertainties or low confidence levels, traditional scoring methods are difficult to effectively identify and make corresponding adjustments, which may affect the objectivity and reliability of the assessment results.

[0006] With the development of information technology, people have begun to explore data-driven methods to optimize the performance assessment process in order to improve the adaptability and accuracy of the assessment system. In this context, how to more efficiently collect, analyze, and utilize assessment data and conduct reasonable evaluations in combination with different types of assessment materials has become the focus of current research and application. Summary of the Invention

[0007] To solve the above technical problems, this application provides a data assessment method, device, system, and storage medium based on a large model.

[0008] The technical solutions provided in this application are described below: In the first aspect of this application, a target Agent deployed on the data bus is used to collect multi-source heterogeneous data from policy websites, industry databases, and institutional systems; Regularly detect the multi-source heterogeneous data. When the similarity change ΔS of the policy text in the multi-source heterogeneous data exceeds the preset threshold α, parse the multi-source heterogeneous data to generate an index knowledge graph containing a triple of policy subject-index item-quantification standard; Use a graph neural network to decompose the department KPI into a tree-shaped requirement graph and calculate the weight distribution matrix W_p of the priority; Update the index knowledge graph based on the distribution matrix W_p and input the updated index knowledge graph into a pre-configured multi-modal scoring engine; In the multi-modal scoring engine, use the RoBERTa-wwm model to extract the deep semantic feature vector V_t for the text-based assessment materials; Extract the spatio-temporal feature vector V_v from the video stream data corresponding to the text-based assessment materials through 3D-ResNet; When the cosine similarity Δ between V_t and V_v is lower than the threshold γ, generate a confidence compensation parameter β based on the cosine similarity Δ through a generator; Generate a path weight vector through a dynamic routing controller according to the confidence compensation parameter β; Based on the path weight vector, calculate the early fusion path output, the mid-term fusion path output or the late fusion path output through the weighted aggregation layer in the multi-modal scoring engine; Calculate the final scoring vector S_f based on the early fusion path output, the mid-term fusion path output or the late fusion path output; Output the final assessment score value based on the final scoring vector.

[0009] Optionally, the use of a graph neural network to decompose the department KPI into a tree-shaped requirement graph and calculate the weight distribution matrix W_p of the priority includes: Obtain the KPI hierarchical relationship, assessment historical data and business context information of the department; Map the KPI hierarchical relationship to a directed graph, where the nodes in the directed graph represent each KPI target, the edges represent the causal relationship between KPI targets, and the weights represent the influence degree between KPI targets; Enhance the weights in the directed graph by combining the assessment historical data and the business context information to generate a tree-shaped knowledge graph; Calculate the importance scores of each KPI target in the directed graph through GNN; Construct the weight distribution matrix W_p based on the importance scores.

[0010] Optionally, parsing the multi-source heterogeneous data to generate an index knowledge graph containing a triple of policy subject-index item-quantification standard includes: Identify the policy subjects in the policy texts of the multi-source heterogeneous data through BERT-NER; Extract the keywords in the policy texts through the TF-IDF algorithm; Extract the index items in the policy texts through dependency parsing; Extract the quantization criteria in the policy texts through regular matching and rule parsing methods; Construct a triple of the policy subject-index item-quantization criteria; Store the triple into the knowledge graph to obtain an index knowledge graph. In the index knowledge graph, the nodes include policy subject nodes, index item nodes, and quantization criteria nodes, and the edges include the edges connecting each node.

[0011] Optionally, when the cosine similarity Δ between the V_t and the V_v is lower than the threshold γ, based on the cosine similarity Δ, generating a confidence compensation parameter β by a generator includes: Calculate the cosine similarity Δ between the text feature vector V_t and the video feature vector V_v; If Δ is lower than the threshold γ, input the cosine similarity into a pre-configured variational autoencoder VAE and output the confidence compensation parameter β.

[0012] Optionally, calculating the final score vector S_f based on the output of the early fusion path, the output of the middle fusion path, or the output of the late fusion path includes: When the confidence compensation parameter β is approximately equal to the first preset value, adopt the early fusion path, and the early fusion path includes concatenating or weighted summing the text feature vector and the video feature vector to obtain the final score vector; When the confidence compensation parameter β is approximately equal to the second preset value, adopt the middle fusion path, and the middle fusion path includes calculating the attention scores of the text feature vector and the video feature vector respectively and fusing them in the hidden layer to obtain the final score vector; When the confidence compensation parameter β is approximately equal to the third preset value, adopt the late fusion path, and the late fusion path includes calculating the score vectors of the text feature vector and the video feature vector respectively and performing weighted averaging on the score vectors to obtain the final score vector.

[0013] Optionally, outputting the final assessment score value based on the final score vector includes: Convert the final score vector into a low-dimensional score feature through a fully connected layer; Process the low-dimensional score feature through a score prediction layer to obtain a normalized preliminary score; Analyze the assessment historical data and obtain the demand discrete level characteristics or demand range characteristics; Construct a mapping function according to the demand discrete level characteristics or the demand range characteristics; Input the preliminary score into the mapping function to obtain the discrete level or the assessment score value.

[0014] The second aspect of this application provides a data assessment device based on a large model, and the device includes: A data acquisition unit, configured to collect multi-source heterogeneous data from a policy website, an industry database, and an institutional system through a target Agent deployed on a data bus; An index knowledge graph construction unit, configured to periodically detect the multi-source heterogeneous data, and when it detects that the similarity change ΔS of the policy text in the multi-source heterogeneous data exceeds a preset threshold α, analyze the multi-source heterogeneous data to generate an index knowledge graph including a policy subject-index item-quantification standard triple; A matrix construction unit, configured to decompose the department KPI into a tree-shaped demand graph by using a graph neural network and calculate a weight distribution matrix W_p of priorities; An input unit, configured to update the index knowledge graph based on the distribution matrix W_p and input the updated index knowledge graph into a pre-configured multi-modal scoring engine; A semantic feature extraction unit, configured to extract a deep semantic feature vector V_t from text-based assessment materials by using a RoBERTa-wwm model in the multi-modal scoring engine; A spatio-temporal feature extraction unit, configured to extract a spatio-temporal feature vector V_v from the video stream data corresponding to the text-based assessment materials through a 3D-ResNet; A parameter compensation unit, configured to generate a confidence compensation parameter β through a generator based on the cosine similarity Δ when the cosine similarity Δ between the V_t and the V_v is lower than a threshold γ; A path weight vector generation unit, configured to generate a path weight vector by a dynamic routing controller according to the confidence compensation parameter β; A path fusion unit, configured to calculate an early fusion path output, a mid-term fusion path output, or a late fusion path output through a weighted aggregation layer in the multi-modal scoring engine based on the path weight vector; A scoring vector generation unit, configured to calculate a final scoring vector S_f based on the early fusion path output, the mid-term fusion path output, or the late fusion path output; A score value generation unit, configured to output a final assessment score value based on the final scoring vector.

[0015] The matrix construction unit is specifically configured to: Obtain the KPI hierarchical relationship, assessment historical data, and business context information of the department; Map the KPI hierarchical relationship into a directed graph, where the nodes in the directed graph represent each KPI target, the edges represent the causal relationship between KPI targets, and the weights represent the influence degree between KPI targets; Enhance the weights in the directed graph by combining the assessment historical data and the business context information to generate a tree-shaped knowledge graph; Calculate the importance scores of each KPI target in the directed graph through GNN; Construct a weight distribution matrix W_p based on the importance scores.

[0016] The third aspect of this application provides a data assessment system based on a large model, including: A processor, a memory, an input / output unit, and a bus; The processor is connected to the memory, the input / output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method of the first aspect and any optional method in the first aspect.

[0017] The fourth aspect of this application provides a computer-readable storage medium, on which a program is stored, and when the program is executed on a computer, it executes the method of the first aspect and any optional method in the first aspect.

[0018] From the above technical solutions, it can be seen that this application has the following advantages: 1. By adopting the method based on policy text similarity detection, it can timely perceive the changes in policies, regulations, and industry standards, and parse out the mapping relationship of "policy subject - indicator item - quantification standard" through the knowledge graph, so as to dynamically adjust the assessment indicators, ensure that the assessment system always meets the latest requirements, and avoid the lag of traditional manual adjustment.

[0019] 2. Through the multi-modal scoring engine, it can simultaneously process various types of assessment materials such as text and video, realize cross-modal information fusion, avoid the one-sidedness of assessment caused by relying only on a single data source, and thus improve the comprehensiveness and objectivity of the assessment results.

[0020] 3. Adopt the RoBERTa-wwm model to extract the deep semantic features of the text, and combine the 3D-ResNet to extract the spatio-temporal features of the video. When there are significant differences between the two, introduce a confidence compensation mechanism to dynamically adjust the scoring strategy, ensure that a reasonable assessment result can still be given when the data confidence is low, and improve the accuracy and stability of scoring.

[0021] 4. Through the dynamic routing controller, automatically adjust the fusion path according to the confidence compensation parameter β, and calculate different scoring results of early fusion, mid-term fusion, and late fusion in the weighted aggregation layer, so as to adapt to different assessment scenarios and make the scoring results more accurate and in line with the actual situation.

[0022] 5. This method is based on large models and graph neural networks, and can complete the whole process of data collection, index update, assessment scoring, etc. without manual intervention, reducing the subjectivity and workload of manual scoring, improving the assessment efficiency, and at the same time enhancing the unity and traceability of assessment criteria. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] In order to more clearly illustrate the technical solutions in the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0024] Figure 1 It is a schematic flowchart of an embodiment of the data assessment method based on a large model provided in the present application; Figure 2 It is a schematic flowchart of an implementation manner of step S102 in the data assessment method based on a large model provided in the present application; Figure 3 It is a schematic structural diagram of an embodiment of the data assessment device based on a large model provided in the present application; Figure 4 It is a schematic structural diagram of an embodiment of the data assessment system based on a large model provided in the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] Please refer to Figure 1 , the present application first provides an embodiment of a data assessment method based on a large model, and this embodiment includes: S101. Collect multi-source heterogeneous data from policy websites, industry databases, and institutional systems through the target Agent deployed on the data bus; Deploy the target Agent in the data bus to be responsible for obtaining multi-source heterogeneous data from different data sources. These data sources include but are not limited to: Policy websites: containing content such as government organization policies, industry standards, and regulatory updates; Industry databases: collecting information such as market trends, industry development reports, and technical standards; Institutional systems: involving performance data, operation data, employee assessment materials, etc. within an enterprise or organization.

[0026] The target Agent regularly obtains this data through API interfaces, web scraping technologies, or data subscription mechanisms, and performs preliminary structuring on it, storing it in a data buffer for subsequent analysis.

[0027] In this embodiment, the target Agent, as a data collection and preprocessing module, can be implemented using large models.

[0028] In this solution, the target Agent is not only a traditional data collection module but also combines large models (such as LLM, Transformer, etc.) for tasks such as intelligent parsing, denoising, and structured data conversion, thereby improving data quality and accuracy.

[0029] Traditional web crawlers face problems such as complex HTML structures and text noise. The target Agent combines LLM (such as GPT-4, Claude) to understand web page content and automatically extract key information.

[0030] The specific steps are as follows: Access the policy website of the industry association and scrape the HTML page.

[0031] Use LLM to parse the web page DOM structure and extract key information such as policy text, release date, and keywords.

[0032] Combine BERT and GPT for text summarization and denoising to remove irrelevant information such as advertisements and navigation.

[0033] Use TF-IDF + BERT to calculate text similarity and remove duplicate policy documents.

[0034] For example, if the implementation details are missing in some policy texts, the large model can be used to complete the details to enhance data integrity. Use T5 and GPT for policy classification and automatically classify them into categories such as "fiscal policy, industrial policy, environmental protection regulations", etc.

[0035] The target Agent can perform data fusion based on API access to open databases and in combination with LLM. For example, it can connect to data sources such as the enterprise credit information publicity system and the WIND economic database through the target Agent to obtain macroeconomic and market research data. Use the large model to analyze industry trends and extract core data points, such as: "The new energy vehicle market is expected to grow by 15% in 2025" "The semiconductor industry is affected by policies and its growth rate slows down in the second half of 2024" Another example: Use GPT-4 + LangChain to parse industry reports and extract key indicators (market size, growth rate, policy impact, etc.) from unstructured PDFs.

[0036] The following provides an example of code implementation: from langchain.document_loaders import PyPDFLoader from langchain.chains.question_answering import load_qa_chain from langchain.chat_models import ChatOpenAI loader = PyPDFLoader("industry_report.pdf") pages = loader.load() # Let GPT-4 analyze the industry report chain = load_qa_chain(ChatOpenAI(model="gpt-4"), chain_type="stuff") result = chain.run({"question": "Please extract the core trends of this industry", "documents": pages}) print(result).

[0037] The target Agent can also directly connect to the enterprise ERP / CRM system through the API to extract department assessment data, such as: sales target achievement rate, customer satisfaction, and operation cost optimization rate.

[0038] This embodiment improves the degree of intelligence in data collection, ensures the accuracy, timeliness, and structurality of data, and lays a solid foundation for subsequent KPI evaluation and assessment.

[0039] S102. Regularly detect the multi-source heterogeneous data. When the similarity change ΔS of the policy text in the multi-source heterogeneous data exceeds the preset threshold α, parse the multi-source heterogeneous data to generate an index knowledge graph containing triples of policy subject-index item-quantification standard; The system regularly detects the collected multi-source heterogeneous data, and uses text vectorization methods such as BERT and TF-IDF to calculate the similarity between the newly collected policy text and the text in the existing policy library; When the similarity change ΔS between the newly detected policy text and the existing policy exceeds the preset threshold α, the system determines that the policy may contain new assessment criteria or indicator changes. Analyze the changed policy text, extract the key entities and their relationships; generate triples of "policy subject - indicator item - quantification standard", and add them to the indicator knowledge graph to form a structured assessment indicator system.

[0040] In a specific implementation, in order to extract effective policy indicator information from multi-source heterogeneous data in step S102, natural language processing (NLP) techniques can be used. By methods such as named entity recognition (NER), keyword extraction, dependency syntactic analysis, regular matching, and rule parsing, parse the policy subject, indicator item, and quantification standard from the policy text, and finally form structured triples of "policy subject - indicator item - quantification standard" and construct an indicator knowledge graph. Refer to Figure 2 , one implementation of this step includes: S1021. Identify the policy subject in the policy text of the multi-source heterogeneous data through BERT-NER; The policy subject usually refers to institutions, industry organizations, enterprises, or other decision-making entities. In this step, a named entity recognition (NER) model based on BERT is used to identify the policy subject in the policy text.

[0041] Select the BERT-BiLSTM-CRF architecture, which combines the context understanding ability of BERT and the sequence modeling ability of BiLSTM, and uses the CRF layer to improve the accuracy of named entity recognition.

[0042] S1022. Extract the keywords in the policy text through the TF-IDF algorithm; In order to extract the core content of the policy document, the TF-IDF (term frequency - inverse document frequency) method is used to extract key concepts from the policy text, providing a basis for subsequent indicator item recognition. Set a policy corpus, count the frequency of each word in the policy text, and calculate the inverse document frequency (IDF).

[0043] S1023. Extract the indicator items in the policy text through dependency analysis; In the policy text, the indicator item usually describes the specific assessment requirements of the policy for a certain field, such as "penetration rate of intelligent manufacturing enterprises". The dependency parsing method is used to find the assessment indicator items in the policy text. Specifically, StanfordCoreNLP, spaCy, or LTP can be used for dependency analysis to identify relationships such as subject-predicate-object, attributive-middle, and verb-object, and extract the core indicator items.

[0044] S1024. Extract the quantitative criteria in the policy text through regular matching and rule parsing methods; Quantitative criteria refer to specific numerical values or quantitative indicators in policy requirements, such as "reach 40%", "not less than 10 billion yuan", "reduce 30%", etc.

[0045] Regular rules can be set: (\d+(\.\d+)?\s*(%|billion yuan|10,000 square meters|kilometer|piece)); An example sentence is as follows: "By 2025, the penetration rate of intelligent manufacturing enterprises will reach 40%." Then, the regular matching result is as follows: Quantitative criteria: 40%.

[0046] If the sentence is more complex, it can be parsed in combination with large model rules, such as using GPT-4 / T5 for parsing: Input: "What is the target penetration rate of intelligent manufacturing enterprises?" Output: "40%".

[0047] S1025. Construct a triple of the policy subject - indicator item - quantitative criteria; Based on the above analysis, a triple of "policy subject - indicator item - quantitative criteria" can be constructed.

[0048] For example: Input policy text: "A certain organization proposed that by 2025, the penetration rate of intelligent manufacturing enterprises will reach 40%." Extraction result: (A certain organization, Penetration rate of intelligent manufacturing enterprises, 40%).

[0049] S1026. Store the triple into the knowledge graph to obtain an indicator knowledge graph. In the indicator knowledge graph, the nodes include policy subject nodes, indicator item nodes, and quantitative criteria nodes, and the edges include the edges connecting each node.

[0050] Finally, store the triple into the knowledge graph to achieve structured storage of policy data and support subsequent querying and reasoning.

[0051] Node types: Policy subject (such as "A certain organization"); Indicator item (such as "Penetration rate of intelligent manufacturing enterprises"); Quantitative criteria (such as "40%"); Edge types: Implementation (from policy subject to indicator item); Target (from indicator item to quantitative criteria).

[0052] Finally, Neo4j can be used to build a knowledge graph. An example is as follows: “CREATE(p: PolicySubject{name: "An Organization"}) CREATE(i: IndicatorItem{name: "Penetration Rate of Intelligent Manufacturing Enterprises"}) CREATE(q: QuantificationStandard{value: "40%"}) CREATE(p)-[: Implemented]->(i) CREATE(i)-[: AimedAt]->(q)”.

[0053] This implementation method can automatically and efficiently parse policy documents, form a structured policy indicator knowledge graph, and provide support for subsequent KPI evaluation and intelligent decision-making.

[0054] S103. Use a graph neural network to decompose the department KPI into a tree-shaped requirement graph, and calculate the priority weight distribution matrix W_p; Use a graph neural network (GNN) to analyze the department KPI, decompose it into a tree-shaped requirement graph. First, obtain the KPI data of different departments and convert it into a standardized representation form, such as a hierarchical indicator structure.

[0055] Construct a graph structure with the KPI and its sub-items. The nodes represent indicators, and the edges represent the association relationships; use GAT (Graph Attention Network) or GCN (Graph Convolutional Network) to calculate the influence weights between KPIs. Combine historical assessment data and use a graph neural network to calculate the weight distribution matrix W_p of each KPI node; this matrix is used to measure the assessment priorities of different indicators and provide a basis for subsequent indicator system updates.

[0056] Specifically, a specific implementation method of this step includes: Obtain the KPI hierarchical relationship, assessment historical data, and business context information of the department; map the KPI hierarchical relationship into a directed graph. In the directed graph, the nodes represent each KPI target, the edges represent the causal relationships between KPI targets, and the weights represent the influence degrees between KPI targets; enhance the weights in the directed graph by combining the assessment historical data and the business context information to generate a tree-shaped knowledge graph; calculate the importance scores of each KPI target in the directed graph through GNN; construct the weight distribution matrix W_p based on the importance scores.

[0057] In this step, a graph neural network (GNN) is used to hierarchically decompose the key performance indicators (KPIs) of each department, construct a tree-like requirement graph, and calculate the influence weights between KPIs based on historical assessment data to generate a priority weight distribution matrix \(W_p\). This matrix can be used to measure the importance of different assessment indicators and provide a decision-making basis for the subsequent update of the indicator knowledge graph.

[0058] The KPIs of different departments may vary in data format, hierarchical relationship, and assessment criteria. Therefore, it is necessary to standardize the KPI data first.

[0059] The KPI data mainly comes from: Enterprise internal management systems (such as ERP, performance management systems); Industry assessment criteria (such as ISO quality management system, institutional performance assessment system); Historical assessment data (past KPI scores, assessment results, assessment rules).

[0060] During the standardization process, the KPI data is converted into a hierarchical indicator structure to make it suitable for subsequent GNN processing. To calculate the influence weights between KPIs using GNN, the KPIs and their sub-items need to be constructed into a graph structure, where: Node: Represents KPI indicators at all levels; Edge: Represents the dependency relationship or influence relationship between KPIs (such as causal relationship, weight transfer).

[0061] Let the KPI system be \(G=(V, E)\): V (nodes): Represents KPIs and their sub-items; E (edges): Represents the hierarchical relationship and influence relationship between KPIs.

[0062] The edges of the graph can be divided into: Hierarchical dependency edges (e.g., "Net profit" is a sub-item of "Operating efficiency"); Influence relationship edges (e.g., "Customer satisfaction" may affect "Operating revenue").

[0063] An example of the specific representation is as follows: G = { Nodes: [Operating revenue, Net profit, Return on assets, Service response time, Complaint rate, Production efficiency, Quality pass rate], Edges: (From Operating efficiency to Operating revenue), (From Operating efficiency to Net profit), (From Operating efficiency to Return on assets), (Customer satisfaction to service response time) (Customer satisfaction to complaint rate) }

[0064] In this embodiment, past assessment results (such as KPI scores, enterprise performance data) can be used for regression analysis, and the weights between two nodes can be enhanced.

[0065] In the assessment system, the importance score is used to measure the weight of each KPI indicator in the overall assessment system. To quantify the relative importance between different KPIs, this step constructs a weight distribution matrix W_p based on the importance score to guide indicator fusion and subsequent assessment score calculation.

[0066] The importance score can be calculated in various ways. For example, it can be calculated based on the allocation of the hierarchical structure, for example: In the hierarchical KPI system, the weights should follow the principle of top-down allocation: Let the weight of the parent KPI be w_p, and the weight w_c of its sub-KPI needs to satisfy: ; For example, assume that the weight of "operational efficiency" is 1.0, and it has three sub-items: operating income, net profit, and return on assets.

[0067] If allocated according to experience: Operating income: 0.5; Net profit: 0.3; Return on assets: 0.2; Then the sum of the importance scores of these sub-KPIs is 1.0.

[0068] Once the importance scores S = [s1, s2,..., sn] of all KPIs are obtained, the weight distribution matrix W_p can be constructed.

[0069] After ensuring that the sum of all scores is 1, construct the weight distribution matrix: ; Among them, the diagonal element W_p[i, i] represents the self-weight of this KPI; the non-diagonal element W_p[i, j] represents the relative influence of KPI i on KPI j.

[0070] S104. Update the indicator knowledge graph based on the distribution matrix W_p, and input the updated indicator knowledge graph into a pre-configured multi-modal scoring engine; ​Based on the calculation result of \(W_p\), adjust the weights of each assessment indicator in the indicator knowledge graph. If the weight values of certain KPIs change significantly, the system adjusts the associated assessment indicators; determine whether to add, modify, or delete some assessment indicators according to the latest knowledge graph; form an updated indicator system and input it into the multi-modal scoring engine.

[0071] In the assessment system, the indicator knowledge graph \(G\) adopts a triple structure of "policy subject - indicator item - quantification standard", that is: \(G=(V, E, A)\); Among them: \(V\): The set of nodes, including policy subject nodes, indicator item nodes, and quantification standard nodes; \(E\): The set of edges, representing the association relationships between different nodes, such as subordination, dependence, influence, etc.; \(A\): Node attributes, including descriptions of assessment indicators, current weights, calculation methods, etc.

[0072] In the assessment system, each indicator item has a corresponding weight value, which is used to measure the importance of the indicator in the overall assessment. After the new weight distribution matrix \(W_p\) is calculated, it is necessary to update the weights of each indicator in the indicator knowledge graph to ensure that the assessment system can accurately reflect the latest business requirements.

[0073] The system first compares the weight value of the current assessment indicator with the newly calculated weight value and checks the magnitude of their changes. If the weight change of an indicator exceeds a preset threshold (e.g., 10%), it is considered that the importance of the indicator has changed significantly and needs to be adjusted.

[0074] Since the assessment system usually requires a certain degree of stability, the newly calculated weight value is not directly used, but a smooth update method is adopted for adjustment. For example, the new weight and the old weight can be weighted averaged to make the change of the weight more stable and avoid the assessment criteria from changing too frequently due to short-term fluctuations.

[0075] If an indicator is not only affected by its own business data but also has a certain association relationship with other assessment items (such as an assessment item depends on the performance of another indicator), the system will consider these factors comprehensively for adjustment. For example, the performance assessment of a certain department may be affected by multiple KPIs, then the final weight of this indicator will be adjusted according to the changes of relevant indicators to make it more in line with the overall assessment goal.

[0076] After the weight update of all metrics is completed, the system checks the overall weight distribution to ensure that the sum of the weights of all metrics remains within a reasonable range (e.g., 100%). If it is found that the sum of the weights exceeds or is lower than the expected value, the system normalizes the weights of all metrics to keep them within the reasonable assessment criteria.

[0077] Finally, according to the dynamic changes in the importance of KPIs, corresponding operations are performed; if the importance weight of some assessment items has increased significantly and this item is not included in the original assessment system, this metric item is automatically added, and a connection with the relevant KPIs is established. If the weights of some assessment items are consistently lower than the threshold and the assessment data indicates that their influence is weak, this metric item can be removed to simplify the assessment system. If the weight changes of some metrics lead to adjustments in their priorities, such as a significant increase in the weight of a sub-item, it may be necessary to re-divide the hierarchical structure and adjust the metric attribution relationship.

[0078] The updated knowledge graph needs to be passed as input to the Multimodal Scoring Engine (MSE). Specifically, first, the knowledge graph is converted into a tensor or embedding representation for easy model calculation. The GNN is used to extract node features, such as the weights of metric items, historical assessment data, assessment text content, etc. Based on the latest weight distribution W_p, the multimodal scoring engine combines text, video, and metric data to calculate the assessment score.

[0079] The following is a further illustration through an example: Suppose there are the following KPIs in the original assessment system: production efficiency (weight 0.4), quality pass rate (weight 0.3), work safety accident rate (weight 0.2), energy consumption level (weight 0.1) In the new round of assessment: The importance of the quality pass rate is increased to 0.4; The influence of the energy consumption level drops to 0.05; The carbon emission compliance rate (0.05) is newly added.

[0080] Then the updated metric knowledge graph is adjusted as follows: Update the weight values of each metric; Adjust the priority of the quality pass rate in the assessment system; Add a carbon emission compliance rate node to the knowledge graph and establish relationships with other KPIs.

[0081] In this step, the accuracy of the assessment system is enhanced to better meet the actual business needs; the scoring mechanism is optimized to ensure that the scoring engine can accurately reflect the contribution degree of each assessment metric.

[0082] S105. In the multi-modal scoring engine, use the RoBERTa-wwm model to extract the deep semantic feature vector V_t from the text-based assessment materials; In this step, use the RoBERTa-wwm (Whole Word Masking) deep learning model to extract features from the text-based assessment materials; this model can capture context semantic relationships and extract the deep semantic feature vector V_t. Preprocess the text data (remove stop words, tokenize, vectorize); calculate the text semantic embedding through the RoBERTa-wwm model to generate the feature vector V_t.

[0083] During the assessment process, text-based assessment materials usually include policy documents, work reports, business data descriptions, etc. These texts often have complex semantic relationships. To accurately understand and analyze the text content, this step uses the RoBERTa-wwm (Whole Word Masking) pre-trained language model to extract deep semantic features from the text data.

[0084] First, remove irrelevant characters, HTML tags, special symbols, etc. to make the text content more standardized. Use Chinese tokenization tools (such as Jieba, PKUSEG) to tokenize the text and retain the complete word structure. Filter out common words with no practical meaning such as "de", "le", "shi" to reduce noise interference to the model. Finally, convert the processed text into a word vector representation that the model can recognize and prepare to input it into the RoBERTa-wwm model.

[0085] In the RoBERTa-wwm model, input the preprocessed text into the RoBERTa-wwm model through the input embedding layer, and use its specific word embedding layer (Embedding Layer) to convert the text into a high-dimensional vector representation. RoBERTa-wwm adopts the Transformer structure and uses the multi-head self-attention mechanism (Multi-Head Attention) to analyze the semantic associations within the sentence and extract important information in the sentence. Compared with the sub-word level masking training of BERT, RoBERTa-wwm adopts the whole-word level masking strategy, enabling the model to better understand the semantics of complete words and improving the understanding ability of texts such as professional terms and policy documents. After multiple layers of Transformer encoding, the model finally outputs a high-dimensional feature vector V_t to represent the overall semantic information of the text. The finally obtained V_t is the deep semantic feature vector of the text and can be used for subsequent multi-modal fusion scoring to ensure that the text information can accurately affect the final assessment result.

[0086] S106. Extract the spatio-temporal feature vector V_v from the video stream data corresponding to the text-based assessment materials through 3D-ResNet; 3D-ResNet is a spatio-temporal feature extraction network that can extract both the spatial information (frame images) and temporal information (changes in frame sequences) of a video. First, frame sampling is performed on the video stream and converted into 3D input data; 3D-ResNet is used for feature extraction, and the spatio-temporal feature vector V_v is output.

[0087] The assessment materials are not limited to text, but may also include video records, such as work scene monitoring, task execution process videos, etc. To extract key information from the video, 3D-ResNet (3D residual network) is used in this step for spatio-temporal feature extraction.

[0088] Since videos usually contain a large number of frames, directly processing all frames will lead to excessive computational complexity. Therefore, an equal-interval or key-frame sampling method is adopted to extract a representative frame sequence from the video. The video frames are normalized, such as adjusting brightness, contrast, and normalizing the size, to improve the generalization ability of the model. The frame sequence is converted into a 3D tensor to adapt to the input format of 3D-ResNet for processing video data.

[0089] In the 3D-ResNet model, 3D-ResNet uses 3D convolution, which can process both spatial information (intra-frame image features) and temporal information (changes in frame sequences) simultaneously, capturing the dynamic features of the video.

[0090] Through skip connections, part of the original information is allowed to be directly transmitted in the deep network, solving the problem of gradient disappearance in traditional deep networks and improving the stability of the model.

[0091] The spatio-temporal features extracted are dimensionally reduced through pooling layers, enabling the model to more effectively aggregate video information and reduce computational costs.

[0092] After being processed by 3D-ResNet, the output video feature vector V_v contains the dynamic spatio-temporal information of the video, such as behavior patterns, scene changes, video content, etc., providing support for subsequent multi-modal fusion scoring.

[0093] Through the above-mentioned feature extraction methods of RoBERTa-wwm and 3D-ResNet, the system can extract key features from text data and video data respectively, ensuring that the assessment system can accurately evaluate different types of assessment materials and providing comprehensive support for the final scoring.

[0094] S107. When the cosine similarity Δ between the V_t and the V_v is lower than the threshold γ, based on the cosine similarity Δ, a confidence compensation parameter β is generated by the generator; Calculate the cosine similarity Δ between V_t and V_v to measure the matching degree between the text and the video. When Δ is lower than the threshold γ (indicating a large deviation between the text and the video information), the system uses a generative adversarial network (GAN) or Bayesian optimization method to generate a confidence compensation parameter β to correct the scoring deviation.

[0095] In an alternative implementation, this step includes: calculating the cosine similarity Δ between the text feature vector V_t and the video feature vector V_v; if Δ is lower than the threshold γ, input the cosine similarity into a pre-configured variational autoencoder VAE and output the confidence compensation parameter β.

[0096] In this alternative implementation, the matching degree between the two is evaluated by calculating the cosine similarity Δ between the text feature vector V_t and the video feature vector V_v. If the similarity is low (i.e., Δ is lower than the set threshold γ), it may mean that the relevance between the text and the video is weak or there is uncertainty. Therefore, it is necessary to introduce a variational autoencoder (VAE) for compensation calculation to generate the confidence compensation parameter β.

[0097] A variational autoencoder (VAE) is a probabilistic generative model used to learn the latent distribution of data and can compensate for missing information in the presence of uncertainty.

[0098] In this step, the VAE is used to learn the joint distribution of the text feature V_t and the video feature V_v, and generate the compensation parameter β in the case of low similarity to improve the credibility of the fused features.

[0099] Specifically, in this embodiment, the VAE processing flow is as follows: Concatenate the text feature vector V_t and the video feature vector V_v to form the input feature X = [V_t, V_v], and send it into the VAE for encoding.

[0100] Extract the latent representation Z of the input feature through a neural network (such as a multi-layer perceptron MLP), and learn its mean μ and standard deviation σ to form Gaussian distribution parameters: Z ~ N(μ, σ); Here, Z is a low-dimensional latent variable representation of the input feature X, which can capture the latent patterns of the data.

[0101] Since Z follows a normal distribution, in order to achieve differentiable gradient optimization, the reparameterization trick is introduced: Z = μ + σ * ϵ; where ε is a random noise following the standard normal distribution N(0, 1), ensuring that the gradient can be propagated smoothly.

[0102] Through a decoder, the latent variable Z is remapped back to the original feature space to generate a compensation parameter β: β = Decoder(Z); β reflects the confidence compensation value between the text and video features. If the text and video have a low matching degree, the VAE will predict an appropriate compensation value to make the finally fused features more valuable for reference.

[0103] In this embodiment, the range of β can be taken between [0, 1], indicating the degree of compensation: If β approaches 0, it means the compensation effect is low, and the influence of the original cosine similarity is maintained.

[0104] If β approaches 1, it means the compensation effect is strong, and the credibility of low-similarity data is improved.

[0105] Subsequently, the dynamic routing controller will adjust the fusion path of the multi-modal scoring engine based on β to ensure that the matching information between the text and video can be reasonably fused, ultimately improving the accuracy of the assessment score.

[0106] This optional implementation method calculates the cosine similarity Δ between the text feature vector V_t and the video feature vector V_v, and when Δ is lower than the threshold γ, uses a variational autoencoder (VAE) for compensation calculation to generate a confidence compensation parameter β. The VAE reasonably compensates for low-similarity situations by learning the joint distribution of text-video features, making the final multi-modal scoring more stable and accurate.

[0107] S108. Generate a path weight vector through the dynamic routing controller according to the confidence compensation parameter β; The confidence compensation parameter β is used as a dynamic adjustment factor and input to the dynamic routing controller. Using the attention mechanism, the weights of each fusion path (early fusion, middle fusion, late fusion) are calculated to generate a path weight vector.

[0108] The dynamic routing controller is responsible for calculating the weights of different fusion paths according to β and generating the final path weight vector. Its core components include: Input layer: Receive the confidence compensation parameter β and multi-modal features (V_t, V_v) as inputs.

[0109] Attention mechanism: Use the attention mechanism to calculate the weights of different fusion paths (early fusion, middle fusion, late fusion).

[0110] When calculating, set different path strategies, such as W_e (early fusion weight), W_m (middle fusion weight), W_l (late fusion weight).

[0111] Taking β as the dynamic weight coefficient, the fusion weights of different paths are adjusted. Since the weights of the three fusion paths need to satisfy the probability distribution property (i.e., the sum of the weights of all paths is 1), the Softmax function is used for normalization.

[0112] After normalization, the final generated path weight vector W = [W_e', W_m', W_l'] is used to guide the subsequent fusion calculation.

[0113] S109. Based on the path weight vector, calculate the early fusion path output, the mid-term fusion path output, or the late fusion path output through the weighted aggregation layer in the multi-modal scoring engine; Early fusion: Information fusion is performed at the feature extraction stage; it is applicable to high-confidence data.

[0114] Mid-term fusion: Cross-modal fusion is performed at the feature transformation stage (attention mechanism); it is applicable to the situation where the data is partially matched but there are certain errors.

[0115] Late fusion: Fusion is performed at the final scoring calculation stage; it is applicable to low-confidence data, and the score is adjusted by the compensation parameter β.

[0116] In a specific implementation manner, this step includes: When the confidence compensation parameter β is approximately equal to the first preset value, the early fusion path is adopted, and the early fusion path includes concatenating or weighted summing the text feature vector and the video feature vector to obtain the final score vector; When the confidence compensation parameter β is approximately equal to the second preset value, the mid-term fusion path is adopted, and the mid-term fusion path includes calculating the attention scores of the text feature vector and the video feature vector respectively, and fusing them in the hidden layer to obtain the final score vector; When the confidence compensation parameter β is approximately equal to the third preset value, the late fusion path is adopted, and the late fusion path includes calculating the score vectors of the text feature vector and the video feature vector respectively, and performing weighted averaging on the score vectors to obtain the final score vector.

[0117] In this embodiment, according to different confidence compensation parameters, a fusion strategy is selected. When β is relatively high (the text and video have a high degree of matching), the early fusion path dominates (W_e' is relatively large), and vector concatenation or interactive attention is directly performed between the text feature vector V_t and the video feature vector V_v to generate the fusion feature. It is applicable to scenarios where the text and video are highly matched, such as a speech script and the corresponding video explanation.

[0118] When β is moderate (the text-video matching degree is average), the mid-term fusion path dominates (W_m' is larger). First, the high-level features of the text and video are extracted separately (such as the semantic features encoded by Transformer and the depth features extracted by 3D-ResNet), and then they are fused in the feature space. This is applicable to scenarios where the content is partially matched but there are some modal differences, such as news reports and corresponding video materials.

[0119] When β is low (the text-video matching degree is low), the late fusion path dominates (W_l' is larger). The text score and video score are calculated separately, and finally, a weighted combination is performed at the score layer. This is applicable to cases where the text and video are weakly correlated, such as when the text describes the product function and the video shows the product advertisement.

[0120] The calculated path weight vector W = [W_e', W_m', W_l'] is used as the input for the subsequent weighted aggregation layer to guide the calculation method of the scoring model.

[0121] In this step, through the dynamic routing controller, using the confidence compensation parameter β as an adjustment factor, combined with the attention mechanism, the weights of different fusion paths are calculated, and the Softmax normalization is used to generate the path weight vector W.

[0122] S110. Calculate the final score vector S_f based on the output of the early fusion path, the output of the mid-term fusion path, or the output of the late fusion path; Based on the path weight vector, calculate the final score vector S_f for the output of the early fusion path, the output of the mid-term fusion path, or the output of the late fusion path.

[0123] In this step, using the path weight vector W = [W_e', W_m', W_l'], the output results of different fusion paths are weighted and calculated to generate the final score vector S_f.

[0124] In S108, the dynamic routing controller calculated the weight allocation methods of different fusion paths (early, mid-term, late) according to the confidence compensation parameter β. Based on this weight information, in this step, the weighted aggregation strategy is used to calculate the final score vector S_f to ensure that the scoring result can integrate different modal information and maintain adaptability to the data matching degree.

[0125] In the multi-modal scoring system, the output data forms of different fusion paths are different, mainly including: Output of the early fusion path: Fusion is directly performed at the feature level (for example, the text feature vector V_t and the video feature vector V_v are directly concatenated and then sent into the scoring network).

[0126] Its output is a fused feature vector, representing the comprehensive feature of this data.

[0127] Output of the mid-term fusion path: First, calculate the high-dimensional semantic representations of the text and the video separately, and then perform fusion in the latent variable space (for example, obtain the fused features through information interaction using Transformer or LSTM).

[0128] Its output is a scored vector S_m after information fusion, which can capture deeper semantic relationships.

[0129] Output of the late-term fusion path: After the text and the video calculate their respective scores separately, perform fusion at the score level (for example, calculate the text score and the video score respectively, and then perform weighted combination at the final stage). Its output is a final decision scored vector, which can be used in cases where the matching degree between the text and the video is relatively low.

[0130] In another implementation, the system can directly select the output of a certain fusion path as the final scoring result instead of using the weighted fusion method to calculate the final scored vector S_f. This selection process is based on the applicability of different fusion strategies in the current scenario and the influence of the confidence compensation parameter β on the information fusion strategy.

[0131] When the system detects significant differences in the reliability of different fusion paths, it can select the most suitable fusion strategy for the current data matching situation. For example: If the text and video features match highly (cosine similarity Δ is high), the output of the late-term fusion path can be directly adopted because the text and the video can calculate scores separately and then fuse directly, reducing information loss.

[0132] If the matching degree between the text and video features is relatively low (cosine similarity Δ is low), the output of the early-term fusion path can be selected because early-term fusion can perform data alignment at the feature level and alleviate the problem of information mismatch.

[0133] If the data volume of the text and the video is large and deep feature interaction is required for fusion, the output of the mid-term fusion path can be selected because this path can utilize deep learning networks (such as Transformer and LSTM) for feature interaction and improve the fusion effect.

[0134] S111. Output the final assessment score value based on the final scored vector.

[0135] In this step, the system needs to convert the final scored vector into a specific assessment score value for subsequent processing and decision-making by the assessment system. This process can be achieved through linear transformation or deep neural network mapping to ensure that the scoring result meets the established assessment criteria and can accurately reflect the actual performance of the object being assessed.

[0136] To obtain a clear assessment score from the high-dimensional final score vector S_f, the system can adopt the following two methods: When the assessment scoring system is relatively fixed and the influence weights of each feature on the final score are relatively stable, a linear mapping method can be used to directly convert S_f into a numerical assessment score.

[0137] Use a weight vector W_{out} to perform weighted summation on the final score vector S_f: Score = W_{out} * S_f + b; Among them, W_{out} is a pre-trained weight matrix used to convert the high-dimensional score vector into an assessment score value. b is a bias term used to normalize or offset-adjust the score. The calculated score may need to be adjusted to a specified score range (for example, 0 - 100 points).

[0138] When the scoring system is relatively complex and the non-linear relationships of multiple features need to be considered, a neural network mapping method can be used to predict the assessment score value from the final score vector S_f.

[0139] Adopt a multi-layer perceptron (MLP), use the final score vector S_f as the input, and output the final assessment score.

[0140] Structure example: Input layer: Receive the final score vector S_f (such as 256 dimensions).

[0141] Hidden layer: The first layer: 128 neurons, ReLU activation function.

[0142] The second layer: 64 neurons, ReLU activation function.

[0143] The third layer: 32 neurons, ReLU activation function.

[0144] Output layer: A single neuron, Sigmoid or Tanh activation function, output a score value between 0 and 1.

[0145] In one implementation, an embodiment of this step includes: Convert the final score vector into low-dimensional score features through a fully connected layer; process the low-dimensional score features through a score prediction layer to obtain a normalized preliminary score; analyze the assessment historical data, and obtain demand discrete level features or demand range features; construct a mapping function according to the demand discrete level features or the demand range features; input the preliminary score into the mapping function to obtain discrete levels or assessment score values.

[0146] In this embodiment, this step uses methods such as a fully connected layer, a scoring prediction layer, and a mapping function to gradually convert the final scoring vector into an assessment score value.

[0147] Specifically, the final scoring vector S_f may be a high-dimensional representation (such as 256-dimensional or 512-dimensional), and there may be redundant information when directly used for calculating the score. Therefore, it is first necessary to reduce the dimension and retain the most important information. This vector comes from the scoring calculation of the fusion path (early fusion / mid-term fusion / late fusion). It may contain features of different modality data (such as scoring information of text and video). One or more layers of fully connected neural networks (FC layers) are used to map the high-dimensional scoring vector to a low-dimensional space: Fscore = ReLU(W1S_f + b1); Among them, W1 is the weight matrix and b_1 is the bias term. After passing through ReLU (or other non-linear activation functions), the feature expression ability can be ensured.

[0148] This vector retains the core information of the final scoring vector and removes redundant information (such as some low-correlation dimensions). The dimension can be set to 16, 32, or 64, depending on the complexity of the scoring task.

[0149] Based on the dimension-reduced scoring feature Fscore, a normalized score is calculated to ensure that the score value is within the standard range (such as between 0 and 1). An additional fully connected layer is used to convert the low-dimensional scoring feature Fscore into a single score value: Scorenorm = σ(W2Fscore + b2), where W2 is the weight matrix of the scoring prediction layer and b2 is the bias term. The activation function uses Sigmoid (σ) to ensure that the output range is between 0 and 1. The obtained Scorenorm is a normalized preliminary score value, which will be used later to map to specific assessment levels or score values.

[0150] To make the assessment score meet the business requirements, it is necessary to refer to historical data and determine the scoring criteria (discrete levels or score ranges). Read the past assessment score data and analyze the scoring distributions of different indicators and assessment objectives. Statistically calculate the mean and standard deviation of the historical scores and observe the trend of the division of assessment levels.

[0151] If the assessment score is a graded score (such as four levels of A, B, C, and D), then it is necessary to analyze the boundary values of different levels from the historical data. If the assessment score is a continuous value range (such as 60 - 100 points), then it is necessary to analyze the distribution of the score range.

[0152] According to the historical data, construct a suitable scoring mapping function to ensure that the calculated preliminary score can be correctly converted into the final assessment score value that meets the business requirements.

[0153] One implementation approach is that a classification mapping function (suitable for discrete scoring) can be constructed based on the discrete grading characteristics of requirements. If the scoring system adopts fixed grades (such as A / B / C / D), interval division rules are constructed, which can ensure that the assessment scores can fall within the predefined grade range and meet the business requirements.

[0154] If the assessment scores are in a continuous interval (such as 60 - 100 points), a linear transformation mapping can be constructed, which can convert the normalized scores between 0 and 1 into the target assessment score interval. Finally, the mapping function is used to convert the normalized scores into actual assessment scores or grades.

[0155] The above embodiments have described the method in the present application. Next, embodiments of the device, system, and storage medium in the present application will be described.

[0156] Refer to Figure 3 , an embodiment of a data assessment device based on a large model is provided in the present application, and this embodiment includes: A data acquisition unit 301, configured to collect multi-source heterogeneous data from a policy website, an industry database, and an institutional system through a target Agent deployed on a data bus; An index knowledge graph construction unit 302, configured to periodically detect the multi-source heterogeneous data, and when the similarity change ΔS of the policy text in the multi-source heterogeneous data exceeds a preset threshold α, parse the multi-source heterogeneous data to generate an index knowledge graph containing a triple of policy subject - index item - quantization standard; A matrix construction unit 303, configured to decompose the department KPI into a tree-shaped requirement graph by using a graph neural network, and calculate a weight distribution matrix W_p of priorities; An input unit 304, configured to update the index knowledge graph based on the distribution matrix W_p, and input the updated index knowledge graph into a pre-configured multi-modal scoring engine; A semantic feature extraction unit 305, configured to extract a deep semantic feature vector V_t from text-based assessment materials by using a RoBERTa-wwm model in the multi-modal scoring engine; A spatio-temporal feature extraction unit 306, configured to extract a spatio-temporal feature vector V_v from the video stream data corresponding to the text-based assessment materials by using a 3D-ResNet; A parameter compensation unit 307, configured to generate a confidence compensation parameter β by a generator based on the cosine similarity Δ when the cosine similarity Δ between the V_t and the V_v is lower than a threshold γ; A path weight vector generation unit 308, configured to generate a path weight vector by a dynamic routing controller according to the confidence compensation parameter β; The path fusion unit 309 is configured to calculate the early fusion path output, the mid-term fusion path output, or the late fusion path output based on the path weight vector through the weighted aggregation layer in the multi-modal scoring engine; The scoring vector generation unit 310 is configured to calculate the final scoring vector S_f based on the early fusion path output, the mid-term fusion path output, or the late fusion path output; The scoring value generation unit 311 is configured to output the final assessment scoring value based on the final scoring vector.

[0157] Optionally, the matrix construction unit 303 is specifically configured to: Obtain the KPI hierarchical relationship, the assessment historical data, and the business context information of the department; Map the KPI hierarchical relationship to a directed graph, where the nodes in the directed graph represent each KPI target, the edges represent the causal relationship between KPI targets, and the weights represent the influence degree between KPI targets; Enhance the weights in the directed graph by combining the assessment historical data and the business context information to generate a tree-like knowledge graph; Calculate the importance scores of each KPI target in the directed graph through GNN; Construct a weight distribution matrix W_p based on the importance scores.

[0158] Optionally, the indicator knowledge graph construction unit 302 is specifically configured to: Identify the policy subjects in the policy text of the multi-source heterogeneous data through BERT-NER; Extract the keywords in the policy text through the TF-IDF algorithm; Extract the indicator items in the policy text through dependency analysis; Extract the quantization criteria in the policy text through regular matching and rule parsing methods; Construct a triple of the policy subject - indicator item - quantization criteria; Store the triple in the knowledge graph to obtain an indicator knowledge graph, where the nodes in the indicator knowledge graph include policy subject nodes, indicator item nodes, and quantization criteria nodes, and the edges include the edges connecting each node.

[0159] Optionally, the parameter compensation unit 307 is specifically configured to: Calculate the cosine similarity Δ between the text feature vector V_t and the video feature vector V_v; If Δ is lower than the threshold γ, input the cosine similarity into a pre-configured variational autoencoder VAE and output a confidence compensation parameter β.

[0160] Optionally, the scoring value generation unit 311 is specifically configured to: Convert the final scoring vector into low-dimensional scoring features through a fully connected layer; Process the low-dimensional scoring features through a scoring prediction layer to obtain a normalized preliminary score; Analyze the assessment historical data and obtain a demand discrete level feature or a demand range feature; Construct a mapping function according to the demand discrete level feature or the demand range feature; Input the preliminary score into the mapping function to obtain a discrete level or an assessment scoring value.

[0161] Please refer to Figure 4 , this application also provides a data assessment device based on a large model, including: A processor 401, a memory 402, an input / output unit 403, and a bus 404; The processor 401 is connected to the memory 402, the input / output unit 403, and the bus 404; The memory 402 stores a program, and the processor 401 calls the program to execute any of the above methods.

[0162] This application also relates to a computer-readable storage medium, on which a program is stored. When the program runs on a computer, the computer is enabled to execute any of the above methods.

[0163] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be described herein again.

[0164] In several embodiments provided in this application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0165] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0166] In addition, each functional unit in various embodiments of the present application may be integrated into one processing unit, may exist physically alone for each unit, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0167] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.

Claims

1. A data assessment method based on large models, characterized in that The method includes: Collecting multi-source heterogeneous data from policy websites, industry databases, and institutional systems through a target Agent deployed on the data bus; Regularly detecting the multi-source heterogeneous data. When the similarity change ΔS of the policy text in the multi-source heterogeneous data exceeds a preset threshold α, parsing the multi-source heterogeneous data to generate an index knowledge graph containing a triple of policy subject-index item-quantification standard; Using a graph neural network to decompose the department KPI into a tree-shaped requirement graph and calculating the weight distribution matrix W_p of priorities; Updating the index knowledge graph based on the distribution matrix W_p and inputting the updated index knowledge graph into a pre-configured multi-modal scoring engine; In the multi-modal scoring engine, using the RoBERTa-wwm model to extract the deep semantic feature vector V_t for text-based assessment materials; Extracting the spatio-temporal feature vector V_v from the video stream data corresponding to the text-based assessment materials through 3D-ResNet; When the cosine similarity Δ between V_t and V_v is lower than the threshold γ, generating a confidence compensation parameter β based on the cosine similarity Δ through a generator; Generating a path weight vector by a dynamic routing controller according to the confidence compensation parameter β; Based on the path weight vector, calculating the early fusion path output, the mid-term fusion path output, or the late fusion path output through the weighted aggregation layer in the multi-modal scoring engine; Calculating the final score vector S_f based on the early fusion path output, the mid-term fusion path output, or the late fusion path output; Outputting the final assessment score value based on the final score vector.

2. The data assessment method based on the large model according to claim 1, wherein The step of using a graph neural network to decompose the department KPI into a tree-shaped requirement graph and calculating the weight distribution matrix W_p of priorities includes: Obtaining the KPI hierarchical relationship, assessment historical data, and business context information of the department; Mapping the KPI hierarchical relationship to a directed graph, where the nodes in the directed graph represent each KPI target, the edges represent the causal relationship between KPI targets, and the weights represent the influence degree between KPI targets; Enhancing the weights in the directed graph by combining the assessment historical data and the business context information to generate a tree-shaped knowledge graph; Calculating the importance scores of each KPI target in the directed graph through GNN; Constructing the weight distribution matrix W_p based on the importance scores.

3. The data assessment method based on a large model according to claim 1, wherein Parsing the multi-source heterogeneous data to generate an index knowledge graph containing a triple of policy subject-index item-quantification standard, including: Identifying the policy subject in the policy text of the multi-source heterogeneous data through BERT-NER; Extracting the keywords in the policy text through the TF-IDF algorithm; Extracting the index items in the policy text through dependency analysis; Extracting the quantification standard in the policy text through regular matching and rule parsing methods; Constructing the triple of policy subject-index item-quantification standard; The triple is stored in the knowledge graph to obtain the index knowledge graph. In the index knowledge graph, the nodes include policy subject nodes, index item nodes, and quantization standard nodes, and the edges include the edges connecting each node.

4. The data assessment method based on a large model according to claim 1, wherein When the cosine similarity Δ between the V_t and the V_v is lower than the threshold γ, based on the cosine similarity Δ, generating a confidence compensation parameter β by a generator includes: Calculating the cosine similarity Δ between the text feature vector V_t and the video feature vector V_v; If Δ is lower than the threshold γ, inputting the cosine similarity into a pre-configured variational autoencoder VAE and outputting the confidence compensation parameter β.

5. The data assessment method based on the large model according to claim 4, wherein Calculating the final score vector S_f based on the output of the early fusion path, the output of the mid-term fusion path, or the output of the late fusion path includes: When the confidence compensation parameter β is approximately equal to the first preset value, an early fusion path is adopted. The early fusion path includes concatenating or weighted summing the text feature vector and the video feature vector to obtain the final score vector; When the confidence compensation parameter β is approximately equal to the second preset value, a mid-term fusion path is adopted. The mid-term fusion path includes respectively calculating the attention scores of the text feature vector and the video feature vector and fusing them in the hidden layer to obtain the final score vector; When the confidence compensation parameter β is approximately equal to the third preset value, a late fusion path is adopted. The late fusion path includes respectively calculating the score vectors of the text feature vector and the video feature vector and performing weighted averaging on the score vectors to obtain the final score vector.

6. The data assessment method based on the large model according to claim 2, wherein Outputting the final assessment score value based on the final score vector includes: Converting the final score vector into a low-dimensional score feature through a fully connected layer; Processing the low-dimensional score feature through a score prediction layer to obtain a normalized preliminary score; Parsing the assessment historical data and obtaining a demand discrete level feature or a demand range feature; Constructing a mapping function according to the demand discrete level feature or the demand range feature; Inputting the preliminary score into the mapping function to obtain a discrete level or an assessment score value.

7. A data assessment device based on a large model, characterized in that The device includes: A data acquisition unit for collecting multi-source heterogeneous data of policy websites, industry databases, and institutional systems through a target Agent deployed on the data bus; An index knowledge graph construction unit for periodically detecting the multi-source heterogeneous data. When the similarity change ΔS of the policy text in the multi-source heterogeneous data exceeds a preset threshold α, parsing the multi-source heterogeneous data and generating an index knowledge graph including a policy subject-index item-quantization standard triple; A matrix construction unit for decomposing the department KPI into a tree-shaped demand graph by using a graph neural network and calculating the weight distribution matrix W_p of the priority; An input unit for updating the index knowledge graph based on the distribution matrix W_p and inputting the updated index knowledge graph into a pre-configured multi-modal scoring engine; A semantic feature extraction unit, which is used to extract a deep semantic feature vector V_t for text-based assessment materials in the multi-modal scoring engine by using the RoBERTa-wwm model; A spatio-temporal feature extraction unit, which is used to extract a spatio-temporal feature vector V_v for the video stream data corresponding to the text-based assessment materials by using 3D-ResNet; A parameter compensation unit, which is used to generate a confidence compensation parameter β by a generator based on the cosine similarity Δ when the cosine similarity Δ between the V_t and the V_v is lower than a threshold γ; A path weight vector generation unit, which is used to generate a path weight vector by a dynamic routing controller according to the confidence compensation parameter β; A path fusion unit, which is used to calculate an early fusion path output, a mid-term fusion path output or a late fusion path output by a weighted aggregation layer in the multi-modal scoring engine based on the path weight vector; A scoring vector generation unit, which is used to calculate a final scoring vector S_f based on the early fusion path output, the mid-term fusion path output or the late fusion path output; A scoring value generation unit, which is used to output a final assessment scoring value based on the final scoring vector; 8. The data assessment device based on the large model according to claim 7, wherein The matrix construction unit is specifically used for: Obtaining the KPI hierarchical relationship, assessment historical data and business context information of the department; Mapping the KPI hierarchical relationship into a directed graph, where nodes in the directed graph represent each KPI target, edges represent the causal relationship between KPI targets, and weights represent the influence degree between KPI targets; Enhancing the weights in the directed graph by combining the assessment historical data and the business context information to generate a tree-like knowledge graph; Calculating the importance scores of each KPI target in the directed graph by using GNN; Constructing a weight distribution matrix W_p based on the importance scores; 9. A data assessment system based on a large model, characterized in that, The system includes: A processor, a memory, an input-output unit and a bus; The processor is connected to the memory, the input-output unit and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 6; 10. A computer-readable storage medium, characterized in that, A program is stored on the computer-readable storage medium, and when the program is executed on a computer, it executes the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • ESG index evaluation system and method based on multi-source data fusion, and storage medium

    CN117370876A

  • Intelligent multimedia interactive teaching and checking system and method

    CN119379506A

  • Multi-source heterogeneous data fusion and processing method based on big data

    CN119783037A

  • Relevance evaluation system and method, program, and recording medium

    WO2017072822A1