An evaluation method and apparatus for knowledge-based question answering

By using an evaluation method oriented towards professional knowledge question answering, dynamically adjusting retrieval parameters and optimizing question answer quality, the system addresses the issues of insufficient timeliness of knowledge updates and the coherence of multi-turn dialogues, thus achieving an efficient and personalized professional question answering system.

CN122088674APending Publication Date: 2026-05-26INSPUR SOFTWARE CO LTD

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSPUR SOFTWARE CO LTD
Filing Date
2026-01-20
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing professional question-answering systems suffer from insufficient timeliness of knowledge updates, weak coherence in multi-turn dialogues, insufficient accuracy of professional word segmentation, poor model interpretability, and high security and compliance risks, resulting in high error rates and difficulty in meeting personalized needs.

Method used

An evaluation method oriented towards professional knowledge question answering is adopted. Through question answer quality evaluation module, knowledge freshness evaluation module, and cross-modal alignment evaluation module, combined with a large language model and knowledge base, the retrieval parameters are dynamically adjusted to optimize question answer quality. Multi-factor weighted model and RLHF technology are introduced to achieve knowledge self-optimization.

Benefits of technology

It improves the accuracy of knowledge retrieval and the quality of personalized answers, reduces the cost of manual review, enhances multimodal semantic alignment capabilities, and meets the needs of time-sensitive fields.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122088674A_ABST
    Figure CN122088674A_ABST
Patent Text Reader

Abstract

This invention relates to the field of data processing and evaluation technology, specifically providing an evaluation method and apparatus for professional knowledge question-and-answer systems. Regarding quality scoring, user evaluations and answer data are directly input into a question-and-answer quality evaluation module to generate quality scoring indicators with confidence intervals, quantifying the accuracy, relevance, and user satisfaction of the answers. Regarding freshness, user question and answer data synchronously drive a knowledge freshness evaluation module to measure the timeliness of the knowledge contained in the answers. Regarding cross-modal alignment, the same question-and-answer data also drives a cross-modal alignment evaluation module, outputting cross-modal alignment indicators to assess the semantic consistency between multimodal content. The quality score, freshness, and cross-modal alignment are aggregated into a knowledge self-optimization strategy engine module. This module uses a decision machine to comprehensively analyze the indicators, dynamically adjust retrieval parameters, and optimize the large model suggestion engineering. Compared with existing technologies, this invention can improve the overall accuracy of retrieval results while meeting personalized response needs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing and evaluation technology, specifically providing an evaluation method and apparatus for professional knowledge question answering. Background Technology

[0002] Current question-answering systems targeting specific professional domains primarily rely on deep learning and natural language processing (NLP) technologies to construct their core architecture. These systems achieve semantic understanding and representation through pre-trained language models (such as BERT and Transformer), capture key semantic features using attention mechanisms, and integrate structured relationships from knowledge graphs (such as "disease-symptom-drug" triples) through multi-turn dialogue state tracking (such as LSTM-based context encoding) to address the complexity of professional questions. Generative models (such as GPT and T5) optimize the professionalism of answers by fine-tuning domain-specific corpora (such as legal provisions and medical guidelines). Some systems also introduce multimodal technologies to integrate text, images, and other data to enhance the comprehensiveness of the answers. However, existing technologies are still limited by insufficient precision in professional word segmentation (such as Chinese medical terminology) and weak coherence in multi-turn dialogues, leading to difficulties in disambiguating complex semantics.

[0003] At the knowledge-driven and data analysis level, professional question-answering systems rely on the construction and application of domain knowledge graphs. For example, in the medical and legal fields, authoritative data (medical journals, court judgments) is crawled to build an "entity-relationship" knowledge base, and embedding technologies (such as TransE) are used to calculate semantic similarity, replacing traditional keyword matching. Meanwhile, automated modeling technologies (such as AutoML and neural architecture search) optimize feature engineering and model selection, lowering the barrier to domain modeling, while context-aware technologies improve the real-time nature of personalized answers by dynamically collecting user environment data (such as geographical location and historical interactions). However, the static nature of knowledge bases leads to insufficient timeliness, especially in rapidly changing fields such as technology and medicine, where delayed information updates can easily cause incorrect answers. Furthermore, the high cost of professional annotation and the scarcity of data further restrict the model's generalization ability.

[0004] The quality assessment of existing question-answering systems faces challenges such as a lack of multi-dimensional indicators and security and compliance risks. While academia has proposed a four-dimensional evaluation system of "accuracy, completeness, readability, and professionalism," practical applications still rely on single similarity calculations (such as F1 scores and BLEU scores), and readability is generally insufficient (e.g., only 50% of answers on Zhihu meet the standard). At the security level, professional data (such as patient medical records) requires dynamic anonymization and role-based access control to ensure compliance; some systems incorporate blockchain technology for data ownership registration. However, poor model interpretability (black-box decision-making reduces credibility) and real-time bottlenecks (delayed updates to streaming data) remain core challenges, requiring hybrid architectures such as incremental learning, edge computing, and attention weight visualization to mitigate these issues.

[0005] How to further integrate dynamic knowledge base updates with unsupervised domain adaptation technologies to promote the evolution of professional question-answering systems towards credibility and efficiency is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0006] This invention addresses the shortcomings of the prior art by providing a highly practical evaluation method for professional knowledge question answering.

[0007] A further technical objective of this invention is to provide a reasonably designed, safe, and applicable evaluation device for professional knowledge question answering.

[0008] The technical solution adopted by this invention to solve its technical problem is: An evaluation method for professional knowledge question answering is proposed. In this method, user evaluation and answer data are directly input into a question answering quality evaluation module based on quality scores. The question answering quality evaluation module generates a quality score index with a confidence interval by analyzing user scores, feedback text and corresponding question and answer content, thereby quantitatively evaluating the accuracy, relevance and user satisfaction of the answers. Regarding the freshness of knowledge, user question and answer data synchronously drive the knowledge freshness evaluation module. The knowledge freshness evaluation module calls the large language model and knowledge base, and outputs the knowledge freshness index through the timeliness analysis algorithm to measure the timeliness of the knowledge contained in the answer. For cross-modal alignment, the same question-and-answer data also drives the cross-modal alignment evaluation module, which integrates cross-modal large model, semantic model and knowledge base, adopts multimodal semantic alignment algorithm, outputs cross-modal alignment index, and evaluates the semantic consistency between multimodal content; The system integrates quality scores, freshness, and cross-modal alignment into the knowledge self-optimization strategy engine module. This module uses a decision machine to comprehensively analyze the indicators, dynamically adjust retrieval parameters, and optimize the large model suggestion process.

[0009] Furthermore, in the question-and-answer quality evaluation module, commonly used indicators such as accuracy, timeliness, compliance, and readability are selected as the question-and-answer quality evaluation indicators, and a dynamic weight adjustment strategy is implemented. The evaluation model expression is as follows: ; coefficient This represents the scores for each dimension. This indicates the frequency of each dimension in historical Q&A. This indicates that among all evaluation dimensions, the first... The frequency of each dimension in historical Q&A; Indicates the first The dynamic weighting coefficients corresponding to the quality evaluation indicators of each question and answer; Simultaneously, large model bias detection and optimization are also required, and the model expression is as follows: ; parameter For factual error count, parameter The expression difference score is calculated using semantic similarity, with coefficients... The domain weight index is represented by N, which represents the total number of question and answer samples participating in this large model bias detection.

[0010] Furthermore, the knowledge freshness evaluation module introduces the following expression: ; formula It is a time decay term, representing the freshness weight of information at time t, simulating the natural aging of information over time; This is the initial freshness weight, representing the proportion of the initial state of object k that contributes to its freshness. ; It is the decay rate, which controls the speed at which information ages. ; It is an exponential decay model, where newly generated information has high freshness but its value decreases over time, which is consistent with the natural aging process. formula This is a version update item, reflecting the effect of version updates on improving freshness; This is the version update weight, representing the proportion of time a version update contributes to the freshness factor. ; It is a version weight function, which is a dynamic function that depends on the version update strategy and can be expressed using a discrete or continuous model.

[0011] Furthermore, if Using a discrete model, during version v update Otherwise, it is 0; if Using a continuous model ,in The most recent update time is considered a constant. This is the version aging factor. This represents the current time point in the calculation of the metric. At this point, the freshness model expression is as follows: ; at the same time, and Satisfy constraints Ensure that the sum of the weights of the two parts is 100%; Or, depending on the actual situation, Use piecewise functions.

[0012] Furthermore, in order to maximize the similarity of matched "image-text" pairs and minimize the similarity of non-matching "image-text" pairs, the cross-modal alignment evaluation module reconstructs the cross-modal alignment loss function, as shown in the following expression: ; in, The text feature vectors are encoded using the BERT model; The image feature vector is encoded using the CLIP model; Cosine similarity is a commonly used function for calculating the similarity between two sets of feature vectors. This is a temperature parameter, with a default value of 0.05. It controls the weight of difficult negative samples. Difficult negative samples are those negative samples that are highly similar in features to positive samples, making them difficult for the model to distinguish. Indicates the index of the text sample. The index of the image sample that matches the text sample; To ensure training stability, a gradient clipping constraint is added, expressed as follows: ; Prevent multimodal feature collapse.

[0013] Furthermore, in the knowledge self-optimization strategy engine module, the incentive signal is introduced into a multi-factor weighted model, expressed as follows: ; in, Weighting user ratings User satisfaction or user ratings can be standardized from 1 to 5 stars to 0 to 1, or calculated using a question-and-answer quality evaluation model; For factual error weighting, The level of factual error is set as a fixed value based on the field or industry. Assigning weights based on urgency, The urgency level is determined by a specific value based on the scenario. The KL divergence coefficient is... Two probability distributions The difference between them is measured. This defines stability constraints, and the confidence effect restricts new strategies. Compared to the old strategy The update and rereading.

[0014] Furthermore, the stability constraints are calculated as follows: ; The continuous form is an integral, which means that if and Completely identical, KL=0; KL>0 indicates that there are differences in the distribution, and the larger the value, the more significant the difference. The sum of all weight coefficients is 1, that is... .

[0015] An evaluation device for knowledge-based question answering includes: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to invoke the machine-readable program to execute an evaluation method oriented towards knowledge-based question answering.

[0016] Compared with existing technologies, the evaluation method and apparatus for knowledge-based question answering of the present invention have the following outstanding advantages: Compared to traditional static timeliness assessment, this invention improves the accuracy of knowledge retrieval and large-model responses, making it particularly suitable for time-sensitive fields such as healthcare, science and technology, and government affairs. In open-domain image and text retrieval tasks, it enhances multimodal semantic alignment capabilities. It enables dynamic quantitative assessment of question-and-answer quality, improving the relevance of quality scores to human evaluation and reducing the cost of manual review. It supports multi-indicator fusion decision optimization, improving the overall accuracy of retrieval results in knowledge base scenarios while meeting personalized response needs. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0018] Figure 1 This is a flowchart illustrating an evaluation method oriented towards professional knowledge question answering; Figure 2 This is a flowchart illustrating the question-and-answer quality assessment module in an evaluation method oriented towards professional knowledge question-and-answer. Figure 3 This is a flowchart illustrating the knowledge freshness evaluation module in an evaluation method oriented towards professional knowledge question answering. Figure 4 This is a flowchart illustrating the cross-modal alignment evaluation module in an evaluation method oriented towards professional knowledge question answering. Figure 5 This is a flowchart illustrating the knowledge self-optimization strategy engine module in an evaluation method oriented towards professional knowledge question answering. Figure 6 This is a flowchart of an implementation example of an evaluation method oriented towards professional knowledge question answering; Figure 7This is a system architecture diagram of an evaluation method oriented towards professional knowledge question answering. Detailed Implementation

[0019] To enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to specific embodiments. Obviously, the described embodiments are merely some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] The following is a preferred embodiment: Example 1:

[0021] In this embodiment, an evaluation method for professional knowledge-based question-and-answer is proposed. Based on the quality score, user evaluation and answer data are directly input into the question-and-answer quality evaluation module. The question-and-answer quality evaluation module generates a quality score index with a confidence interval by analyzing user scores, feedback text and corresponding question-and-answer content, and quantitatively evaluates the accuracy, relevance and user satisfaction of the answer. Regarding freshness, user question and answer data synchronously drive the knowledge freshness evaluation module. The knowledge freshness evaluation module calls the Large Language Model (LLM) and knowledge base, and outputs the knowledge freshness index through timeliness analysis algorithms (such as time decay model and knowledge update comparison) to measure the timeliness level of the knowledge contained in the answer. For cross-modal alignment, the same question-and-answer data also drives the cross-modal alignment evaluation module, which integrates cross-modal large models (such as CLIP / ViLT), semantic models (such as BERT) and knowledge bases, and adopts multimodal semantic alignment algorithms (such as text-image similarity calculation and entity consistency verification) to output cross-modal alignment index (Alignment) to evaluate the semantic consistency between multimodal content; The system aggregates quality scores, freshness, and cross-modal alignment into a knowledge self-optimization strategy engine module. This module uses a decision machine (such as a threshold-based rule engine or reinforcement learning agent) to comprehensively analyze the metrics, dynamically adjust retrieval parameters (such as vector retrieval weights and metadata filters), and optimize the large model suggestion engineering. This process forms a closed-loop iteration, continuously improving the accuracy of knowledge retrieval and the quality of personalized output with question-and-answer interactions, while also providing potential basis for updating the retrieval model weights.

[0022] In the question-and-answer quality evaluation module, commonly used indicators such as accuracy, timeliness, compliance, and readability are selected as evaluation metrics. A dynamic weight adjustment strategy is implemented, and the evaluation model expression is as follows: ; coefficient This represents the scores for each dimension. This indicates the frequency of each dimension in historical Q&A. This indicates that among all evaluation dimensions, the first... The frequency of each dimension in historical Q&A; Indicates the first The quality evaluation indicators for each question and answer (accuracy, timeliness, compliance, readability, etc.) The dynamic weight coefficient corresponding to each dimension (this coefficient is calculated based on the historical call frequency of the corresponding dimension).

[0023] Simultaneously, large model bias detection and optimization are also required, and the model expression is as follows: ; parameter For factual error count, parameter The expression difference score is calculated using semantic similarity, with coefficients... The domain weight index is represented by N, which represents the total number of question and answer samples participating in this large model bias detection.

[0024] In the knowledge freshness evaluation module, the following expression is introduced: ; formula It is a time decay term, representing the freshness weight of information at time t, simulating the natural aging of information over time; This is the initial freshness weight, representing the proportion of the initial state of object k that contributes to its freshness. ; It refers to the decay rate, which controls the speed at which information ages. ; It is an exponential decay model, where newly generated information has high freshness but its value decreases over time, which is consistent with the natural aging process. formula This is a version update item, reflecting the effect of version updates on improving freshness; This is the version update weight, representing the proportion of time a version update contributes to the freshness factor. ; It is a version weight function, which is a dynamic function that depends on the version update strategy and can be expressed using a discrete or continuous model.

[0025] if Using a discrete model, during version v update Otherwise, it is 0; if Using a continuous model ,in The most recent update time is considered a constant. This is the version aging factor. This indicates the current time point in time when the freshness indicator is calculated. At this point, the freshness model expression is as follows: ; at the same time, and Satisfy constraints Ensure that the sum of the weights of the two parts is 100%; Or, depending on the actual situation, Use piecewise functions.

[0026] In the cross-modal alignment evaluation module, to maximize the similarity of matched "image-text" pairs and minimize the similarity of non-matching "image-text" pairs, the cross-modal alignment loss function is reconstructed, as shown in the following expression: ; in, The text feature vector (where the superscript is...) (Text tags), encoded by the BERT model; Image feature vector (where superscript) (Image tagging), encoded using the CLIP model; The cosine similarity is a commonly used function for calculating the similarity between two sets of feature vectors. This is a temperature parameter, with a default value of 0.05. It controls the weight of difficult negative samples. Difficult negative samples are those negative samples that are highly similar in features to positive samples, making them difficult for the model to distinguish. , This is the index identifier for the sample, where Indicates the index of the text sample. This indicates the index of the image sample that matches the text sample.

[0027] To ensure training stability, a gradient clipping constraint is added, expressed as follows: ; Prevent multimodal feature collapse.

[0028] In the knowledge self-optimization strategy engine module, RLHF (Reinforcement Learning from Human Feedback) is a technique that combines reinforcement learning with human preferences, aiming to make the output of AI models (especially large language models) more aligned with human values ​​and intentions. Its core idea is to transform subjective human evaluations into quantifiable reward signals to guide model optimization. Here, the incentive signal is introduced into a multi-factor weighted model, expressed as follows: ; in, Weighting user ratings User satisfaction or user ratings can be standardized from 1 to 5 stars to 0 to 1, or calculated using a question-and-answer quality evaluation model; For factual error weighting, The level of factual error is set as a fixed value based on the field or industry. Assigning weights based on urgency, The urgency level is determined by a specific value based on the scenario. The KL divergence coefficient is... Two probability distributions The difference between them is measured. This defines stability constraints, and the confidence effect restricts new strategies. Compared to the old strategy Repeated updates are performed to avoid performance crashes caused by policy mutations.

[0029] Furthermore, the stability constraints are calculated as follows: ; The continuous form is an integral, which means that if and Completely identical, KL=0; KL>0 indicates that there are differences in the distribution, and the larger the value, the more significant the difference. The sum of all weight coefficients is 1, that is... .

[0030] like Figure 1As shown, this paper describes three parallel evaluation strategies driven by user ratings and feedback, and user questions and answers, as input sources. User ratings and feedback, and user questions and answers serve as direct input sources, driving the question-and-answer quality evaluation module to obtain question-and-answer quality index scores and biases. This module primarily evaluates the quality of answers based on user ratings and feedback, combined with user questions and the obtained answers, and outputs the scores and biases. User questions and answers drive the knowledge freshness evaluation module, which requires the use of large-scale models, knowledge bases, and other tools, and evaluates according to a specific algorithm to obtain knowledge freshness indices. User questions and answers also drive the cross-modal alignment evaluation module, which requires the use of cross-modal large-scale models, knowledge bases, semantic large-scale models, and other tools, and evaluates according to a specific algorithm to obtain cross-modal alignment indices.

[0031] Next, the quality evaluation indicators (scores and biases), knowledge freshness indicators, and cross-modal alignment indicators are summarized to drive a knowledge self-optimization strategy, adjusting retrieval parameters based on the decision-making machine's judgment. This process gradually optimizes the retrieval strategy and the semantic output of the large model as the number of question-and-answer sessions increases, ultimately yielding highly accurate and personalized answers.

[0032] like Figure 2 As shown, when the system detects low user ratings (ratings < 3 stars / 5 stars), it triggers an analysis process, decomposing the rating bias into four independent problem domains using a feedback classification model: The first category, for cases of factual errors, involves activating the knowledge gap assessment module, which can then connect to the expert-assisted review module to obtain information on knowledge gaps, such as knowledge errors or knowledge deficiencies. The second category, targeting situations with low satisfaction, involves activating the issue element analysis module to examine the completeness of the user's issue elements, analyze the relationship between this completeness and the responses already given, and generate model parameter tuning, such as large model temperature parameters, prompt words, and supplementary engineering points. The third category addresses situations where timeliness is insufficient or information is urgent. It involves calling the external knowledge joint evaluation module to address issues such as insufficient timeliness or urgent responses caused by failure to update the knowledge base in a timely manner, thereby introducing external knowledge and promptly supplementing it with the latest information. The fourth category involves discrete responses, including dispersion analysis. This involves analyzing the degree of dispersion between user questions and responses to obtain dispersion parameters. The results of these four types of analysis are then aggregated into a question-and-answer quality evaluation parameter system to construct a multi-dimensional evaluation matrix. The question-and-answer quality evaluation (weight) coefficients are calculated based on the multi-factor weighted model and used to control the updating of the RLHF strategy.

[0033] like Figure 3As shown, the process begins with knowledge retrieval, which retrieves target information from the knowledge base. Based on the retrieved knowledge, timestamps are extracted to parse the time-related metadata (such as publication, creation / modification, and validity periods) associated with the knowledge entries. Simultaneously, the knowledge version weight update on the right executes a dynamic weight adjustment algorithm, which may be based on factors such as version iteration frequency, source authority, or user feedback. The timestamp extraction and knowledge version weight update outputs are then fed into the terminal's knowledge freshness calculation module as joint inputs.

[0034] This core component integrates time-series features (timestamps) and version weight parameters, and uses a predefined knowledge freshness model (metric function) to generate quantifiable knowledge freshness evaluation indicators. The entire process adopts a unidirectional workflow design, clearly presenting a standardized processing path for calculating knowledge timeliness.

[0035] like Figure 4 As shown in the flowchart, this system illustrates a closed-loop knowledge processing system based on machine learning. The system receives raw data starting with "knowledge set input"; then, through "knowledge type determination," it performs structured classification and representation analysis of the knowledge; in the "knowledge vectorization processing" stage, embedding techniques (such as word embedding or neural network encoding) are used to transform the knowledge into dense vector representations; in the "similarity calculation" stage, distance metrics (such as cosine similarity) are used to evaluate the correlation between knowledge entities; the "loss function calculation" stage quantifies and evaluates model performance based on specific learning objectives (such as contrastive learning or classification tasks); finally, through "gradient pruning constraint determination," the system implements numerical stability control of the optimization algorithm, performing norm threshold truncation on the backpropagation gradient to prevent gradient explosion, thus completing the iterative optimization of the training process.

[0036] like Figure 5 As shown, the flowchart describes the core control logic of the question-answering system's quality evaluation and optimization mechanism. Its information input sources include three parallel evaluation dimensions: 1) "Question-answering quality evaluation parameters and deviations" (quantifying basic question-answering accuracy indicators and errors), 2) "Knowledge freshness indicators" (assessing the timeliness of the knowledge base), and 3) "Cross-modal alignment indicators" (detecting the consistency of multimodal data such as text and images).

[0037] The outputs of these three dimensions converge on the "threshold control" module via arrows. This module undertakes the core regulatory function. Its output is divided into two paths: the main path points to "multi-factor weighted evaluation" (which calculates and integrates the weights of different dimension indicators), and the parallel path points to "evaluation threshold control" (which dynamically sets the evaluation standard thresholds for each dimension).

[0038] The entire process culminates in a closed loop, indicated by a dashed arrow pointing to "weight update / parameter adjustment," signifying that the system dynamically optimizes indicator weights and algorithm parameters based on evaluation results, achieving adaptive iterative upgrades. This design forms a complete closed-loop control chain from multi-dimensional evaluation input → threshold control → comprehensive evaluation → dynamic optimization.

[0039] like Figure 6 As shown, this flowchart uses a vertical layout to describe the evaluation processing mechanism based on user feedback and its subsequent decision-making process. The process starts at the "User Feedback" node, which is connected to the "Evaluation System" node via directed arrows, representing the initial processing layer of feedback data. Subsequently, the evaluation results are input into the core decision-making unit—the diamond-shaped decision box "Decentralization Trigger".

[0040] The output of the freewheeling trigger contains two mutually exclusive paths: When the combined condition of "low score + high frequency of questions and answers" is met, the process leads to the "weight update / parameter adjustment" node. This node represents the dynamic optimization process of model parameters or system weights, designed to respond to the potential problems indicated by the combination of low scores and high interaction frequency.

[0041] When feedback is identified as an "evaluation that does not require a response," the process directs to the "Record Evaluation Parameters" node. This node represents the operation of persistently storing evaluation-related metadata (such as scores, features, etc.) without triggering an immediate intervention mechanism.

[0042] All process elements (rectangles represent processing steps or data nodes, and diamonds represent decision nodes) and their names strictly follow the original diagram. This diagram clearly defines a control flow that triggers different processing strategies (active tuning or passive archiving) based on evaluation results.

[0043] like Figure 7 As shown in the diagram, this diagram illustrates a simplified system architecture, which is based on a professional layered structure and divided into three layers from bottom to top: the platform layer, the application layer, and the presentation layer.

[0044] First, the platform layer serves as the foundational support, including a large model cluster (covering inference models, vectorized models, speech models, and image recognition models) that provides core AI capabilities; database systems such as MySQL and MongoDB for structured data storage, vector libraries and graph vector libraries to support high-dimensional data processing; and component libraries (such as Redis for cache management, Log4j for logging, Nacos for service discovery, and Docker for containerized deployment), forming the reliable infrastructure of the entire system.

[0045] Secondly, the upper layer is the application layer, responsible for implementing business logic. It includes modules such as question-and-answer quality evaluation, freshness evaluation, similarity calculation, image recognition, and policy control. These modules handle real-time interaction, content analysis, and decision-making processes by calling platform layer resources. Finally, the presentation layer, located at the top, serves as the user interface layer, including front-end interfaces such as mini-programs, apps, web applications, and edge devices, visualizing and presenting the functions of the application layer to the end user. All elements within each layer collaborate closely, with modules clearly focused on core functions and excluding software development auxiliary components, ensuring efficient and scalable system operation. Example 2:

[0046] Implementation of a solution for a municipal government Q&A platform: The expression for knowledge freshness is as follows: ; Initial freshness weight , representing the proportion of the initial state of object k that contributes to freshness ( ), and differentiated configuration based on local policy types, such as , Attenuation rate It can control the rate of information aging ( Establish a policy lifecycle prediction model and dynamically calculate the λ value based on the policy revision cycle, such as... , .

[0047] For policy version parameters affected by time t, due to... To calculate, if , ,but , .

[0048] Version weight function It is a dynamic function, where t = current time - policy release date, therefore it can also be written as... In addition, in this embodiment Using a piecewise function, as follows: ; When a new policy is released less than one month ago (or a specified threshold) ; Therefore, based on the above calculation rules, the knowledge freshness of livelihood policies that have been published for 0, 6, 20, and 40 months is calculated as follows: ; ; ; ; That is, the knowledge freshness of livelihood policies that have been published for 0, 6, 20, and 40 months are 1.0000, 0.8867, 0.6494, and 0.3995 respectively. The significance of this is to illustrate the characteristics of policy knowledge changing over time. When there are multiple documents of similar types, the evaluation method based on knowledge freshness can provide more accurate answers to user inquiries.

[0049] First, input source 1: Citizens' answer rating (3 stars) and text feedback on "Application conditions for talent rental subsidies in the High-tech Zone in 2025": "The answer did not mention the new regulations on overseas academic qualification certification".

[0050] Input source 2: Original question records and policy clause answer texts provided by the system.

[0051] API calls to government knowledge bases (including policy bases and service guide bases) and cross-modal large models (integrating text / image / table recognition).

[0052] According to the methodology in this case, the evaluation module is divided into question-and-answer quality evaluation, knowledge freshness evaluation, and cross-modal alignment evaluation. Its specific execution process and output metrics are as follows: (1) Q&A quality evaluation: The keyword "overseas academic qualification certification" in user feedback was analyzed and the coverage of the answer text was compared; the timeliness of the policy clauses was tested, and the final Q&A quality score was 72 points.

[0053] (2) Knowledge freshness evaluation: Scanning the "Talent Policy of High-tech Zone" entry in the knowledge base, it was found that the PDF file updated by the Human Resources and Social Security Bureau 3 days ago was not included in the answer system. The freshness index of this knowledge was calculated to be 0.63 based on the recorded parameters.

[0054] Cross-modal alignment assessment: By comparing the response text with scanned copies of policy documents, it was found that the academic qualification verification process lacked textual descriptions, and the modal alignment index was 54% (< the standard threshold of 65%).

[0055] Based on this method, the pseudocode is designed as follows: # ====================================================== # Pseudocode for a knowledge self-optimization strategy decision engine; # Input source: Quantitative output results from the three major evaluation modules; # Triggering condition: Joint diagnosis using multi-dimensional evaluation indicators; # ====================================================== def knowledge_self_optimization(quality_score, # From [Question and Answer Quality Evaluation] module → (Indicator Scoring and Bias) freshness_index, # From [Knowledge Freshness Evaluation] module → (Knowledge Freshness Index) modality_alignment, # From [Cross-Modality Alignment Evaluation] module → (Cross-Modality Alignment Index) THRESHOLD_QUALITY = 80, # Preset quality pass rate (percentage); THRESHOLD_FRESHNESS = 0.8, # Freshness baseline value [0-1]; THRESHOLD_MODALITY = 0.65 # Modal alignment reference value [0-1]; ): Knowledge self-optimization strategy trigger logic: When both conditions are met: 1) The quality score of the Q&A is below the passing mark. 2) Knowledge freshness is insufficient compared to the benchmark. 3) When cross-modal alignment fails to meet the standard Execute emergency optimization protocol # --- Core Judgments of Decision Tree --- if (quality_score <THRESHOLD_QUALITY and freshness_index <THRESHOLD_FRESHNESS and modality_alignment <THRESHOLD_MODALITY): # Initiate URGENT level optimization (corresponding flowchart [Knowledge Self-Optimization Strategy]) optimization_level = URGENT # Perform dynamic adjustment of search parameters (corresponding flowchart [Search Parameter Adjustment]) adjust_retrieval_parameters(policy_doc_weight_increment = 0.4, # policy document retrieval weight +40% cross_modal_weight_increment = 0.3 # cross-modal content extraction weight +30%) # Perform real-time updates to the knowledge base (corresponding flowchart [Retrieval Weight Update]) update_knowledge_base(new_regulation = "2025 Human Resources and Social Security Regulation

[15] No.", linked_resource = "Overseas Academic Credential Verification Process.pdf") return optimization_level # Returns the policy execution status, etc. # --- Regular optimization branch (implied logic in the expanded flowchart) --- else: # Implement an incremental optimization strategy return execute_gradual_optimization(...) Enhanced technical features: (1) Threshold parameterization: The THRESHOLD_ series of constants enables configurable threshold management to adapt to different scenario requirements (such as setting stricter standards in government scenarios).

[0056] (2) Incremental weight adjustment: The *_weight_increment parameter is used to achieve non-overlapping weight superposition, while retaining the historical optimization trajectory.

[0057] (3) Dual-path optimization mechanism: the else branch implies an incremental optimization path that is not explicitly represented in the flowchart, thus improving the completeness of the decision-making logic.

[0058] (4) Resource association storage: The linked_resource parameter enables dynamic association between text terms and cross-modal resources such as PDFs / images, solving the problem of text and image separation.

[0059] The pseudocode design fully follows the tree-like decision structure of the flowchart, taking the three major evaluation indicators as decision inputs, and triggering the collaborative optimization of the retrieval system and knowledge base through clear threshold conditions, forming a closed-loop self-evolving system.

[0060] After three rounds of question-and-answer loop optimization: (1) Dynamic adjustment of search parameters: The priority weight of policy document search was increased from 0.5 to 0.82, and the cross-modal content relevance coefficient β was optimized from 0.3 to 0.68.

[0061] (2) Real-time updates of the knowledge base: The system automatically captures new regulations from the official website of the Human Resources and Social Security Bureau, reducing the knowledge update delay from 72 hours to 4 hours.

[0062] (3) Improved service efficiency: (a) The accuracy rate of answering the same question "overseas academic qualification certification materials" increased from 63% to 92%, the average citizen satisfaction rating reached 4.5 stars (from 3.2 stars), and the transfer rate of manual customer service decreased by 67%. Example

[0063] An evaluation device for knowledge-based question answering includes: at least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to invoke the machine-readable program to execute an evaluation method oriented towards knowledge-based question answering.

[0064] The processor can be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor can be a microprocessor or any conventional processor.

[0065] Memory is used to store computer programs and / or modules. The processor implements various functions of the electronic device by running or executing the computer programs and / or modules stored in the memory, and by accessing data stored in the memory. Memory can mainly include a program storage area and a data storage area. The program storage area can store the operating system, at least one application program required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, memory can also include high-speed random access memory, and can also include non-volatile memory, such as hard disks, RAM, plug-in hard disks, smart memory cards (SMC), secure digital cards (SD cards), flash memory cards, at least one disk storage device, flash memory devices, or other volatile solid-state storage devices.

[0066] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. An evaluation method for professional knowledge question answering, characterized in that, Regarding the quality rating, user reviews and answer data are directly input into the question-and-answer quality evaluation module. The question-and-answer quality evaluation module analyzes user ratings, feedback texts, and corresponding question-and-answer content to generate quality rating indicators with confidence intervals, which quantitatively evaluate the accuracy, relevance, and user satisfaction of the answers. Regarding the freshness of knowledge, user question and answer data synchronously drive the knowledge freshness evaluation module. The knowledge freshness evaluation module calls the large language model and knowledge base, and outputs the knowledge freshness index through the timeliness analysis algorithm to measure the timeliness of the knowledge contained in the answer. For cross-modal alignment, the same question-and-answer data also drives the cross-modal alignment evaluation module, which integrates cross-modal large model, semantic model and knowledge base, adopts multimodal semantic alignment algorithm, outputs cross-modal alignment index, and evaluates the semantic consistency between multimodal content; The system integrates quality scores, freshness, and cross-modal alignment into the knowledge self-optimization strategy engine module. This module uses a decision machine to comprehensively analyze the indicators, dynamically adjust retrieval parameters, and optimize the large model suggestion process.

2. The evaluation method for professional knowledge-based question answering according to claim 1, characterized in that, In the question-and-answer quality evaluation module, commonly used indicators such as accuracy, timeliness, compliance, and readability are selected as evaluation indicators. A dynamic weight adjustment strategy is implemented, and the evaluation model expression is as follows: ; coefficient This represents the scores for each dimension. This indicates the frequency of each dimension in historical Q&A. This indicates that among all evaluation dimensions, the first... The frequency of each dimension in historical Q&A; Indicates the first The dynamic weighting coefficients corresponding to the quality evaluation indicators of each question and answer; Simultaneously, large model bias detection and optimization are also required, and the model expression is as follows: ; parameter For factual error count, parameter The expression difference score is calculated using semantic similarity, with coefficients... The domain weight index is represented by N, which represents the total number of question and answer samples participating in this large model bias detection.

3. The evaluation method for professional knowledge-based question answering according to claim 2, characterized in that, The knowledge freshness evaluation module introduces the following expression: ; formula It is a time decay term, representing the freshness weight of information at time t, simulating the natural aging of information over time; This is the initial freshness weight, representing the proportion of the initial state of object k that contributes to its freshness. ; It is the decay rate, which controls the speed at which information ages. ; It is an exponential decay model, where newly generated information has high freshness but its value decreases over time, which is consistent with the natural aging process. formula This is a version update item, reflecting the effect of version updates on improving freshness; This is the version update weight, representing the proportion of time a version update contributes to the freshness factor. ; It is a version weight function, which is a dynamic function that depends on the version update strategy and can be expressed using a discrete or continuous model.

4. The evaluation method for professional knowledge question answering according to claim 3, characterized in that, if Using a discrete model, during version v update Otherwise, it is 0; if Using a continuous model ,in The most recent update time is considered a constant. This is the version aging factor. This represents the current time point in the calculation of the metric. At this point, the freshness model expression is as follows: ; at the same time, and Satisfy constraints Ensure that the sum of the weights of the two parts is 100%; Or, depending on the actual situation, Use piecewise functions.

5. The evaluation method for professional knowledge question answering according to claim 4, characterized in that, In the cross-modal alignment evaluation module, to maximize the similarity of matched "image-text" pairs and minimize the similarity of non-matched "image-text" pairs, the cross-modal alignment loss function is reconstructed, as shown in the following expression: ; in, The text feature vectors are encoded using the BERT model; The image feature vector is encoded using the CLIP model; Cosine similarity is a commonly used function for calculating the similarity between two sets of feature vectors. This is a temperature parameter, with a default value of 0.

05. It controls the weight of difficult negative samples. Difficult negative samples are those negative samples that are highly similar in features to positive samples, making them difficult for the model to distinguish. Indicates the index of the text sample. The index of the image sample that matches the text sample; To ensure training stability, a gradient clipping constraint is added, expressed as follows: ; Prevent multimodal feature collapse.

6. The evaluation method for professional knowledge question answering according to claim 5, characterized in that, In the knowledge self-optimization strategy engine module, the incentive signal is introduced into a multi-factor weighted model, as shown in the following expression: ; in, Weighting user ratings User satisfaction or user ratings can be standardized from 1 to 5 stars to 0 to 1, or calculated using a question-and-answer quality evaluation model; For factual error weighting, The level of factual error is set as a fixed value based on the field or industry. Assigning weights based on urgency, The urgency level is determined by a specific value based on the scenario. The KL divergence coefficient is... Two probability distributions The difference between them is measured. This defines stability constraints, and the confidence effect restricts new strategies. Compared to the old strategy The update and rereading.

7. The evaluation method for professional knowledge question answering according to claim 6, characterized in that, The stability constraints are calculated as follows: ; The continuous form is an integral, which means that if and Completely identical, KL=0; KL>0 indicates a difference in distribution, the larger the value, the more significant the difference, and the sum of all weight coefficients is 1, i.e. .

8. An evaluation device for professional knowledge question answering, characterized in that, include: At least one memory and at least one processor; The at least one memory is used to store a machine-readable program; The at least one processor is configured to invoke the machine-readable program to perform the method according to any one of claims 1 to 7.