Knowledge base question answering system question recall optimization method, computer device and medium

CN122817440APending Publication Date: 2026-09-25BEIJING JIZHI DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611088701.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-21
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0003]在进行问题召回时,通常可以将用户问题进行编码以进行向量检索,或者基于用户问题进行关键词匹配,但由于各问题在编码过程中相互独立,难以捕捉问题间的深层语义交互,在区分表述相近但含义不同的问题时容易出现误召回的情况

Benefits of technology

[0016]本说明书提供的多个实施方式中,首先,构建问题对训练数据集,问题对包括正样本和负样本,正样本为知识库中的主问题与其对应的相似问构成的问题对,负样本为知识库中的主问题与其非相似问构成的问题对;接着,使用交叉编码器教师模型对问题对训练数据集中的各问题对进行语义相似度评分,生成各问题对所对应的教师评分;然后,对教师评分进行软标签处理,得到软标签训练数据;再接着,基于软标签训练数据,利用余弦相似度损失函数对双编码器学生模型进行蒸馏训练,以使双编码器学生模型输出的两个问题向量之间的余弦相似度逼近软标签训练数据中的教师评分;再接着,将蒸馏训练后的双编码器学生模型部署至知识库问答系统,用于将用户输入的问题编码为向量,并在知识库中预先生成的主问题向量库中进行向量检索,以召回与用户输入的问题最相似的主问题并返回其对应的答案,如此,可以提高知识库问答系统的问题召回精度和召回效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122817440A_ABST
    Figure CN122817440A_ABST
Patent Text Reader

Abstract

The embodiment of the specification provides a question recall optimization method of a knowledge base question answering system, a computer device and a medium, and relates to the technical field of artificial intelligence. The method comprises the following steps: constructing a question pair training data set, the question pair comprising positive samples and negative samples, the positive sample being a question pair formed by a main question and a similar question thereof, and the negative sample being a question pair formed by the main question and a non-similar question thereof; using a cross-encoder teacher model to score the semantic similarity of each question pair, to generate a teacher score corresponding to each question pair; based on the teacher score, using a cosine similarity loss function to perform distillation training on a double-encoder student model; deploying the model after the distillation training to the knowledge base question answering system, to encode a question input by a user into a vector, and performing vector retrieval in a main question vector library, to recall a main question most similar to the question input by the user and return a corresponding answer. In this way, the question recall accuracy and recall efficiency of the knowledge base question answering system can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments described in this specification relate to the field of artificial intelligence technology, specifically to a question retrieval optimization method, computer equipment, and medium for a knowledge base question answering system. Background Technology

[0002] In knowledge-based question-answering systems, a common question retrieval architecture is the FAQ (Feature Questions) framework: the knowledge base pre-stores several main questions (standard question formats) and their corresponding answers. Each main question is associated with multiple manually labeled similar questions (variant question formats), forming a knowledge structure of "main question + set of similar questions + answer". When a user enters a new question, the system retrieves the semantically most similar main question from the knowledge base and returns the corresponding answer to the user.

[0003] When conducting issue recall, user issues can usually be encoded for vector retrieval or keyword matching can be performed based on user issues. However, since each issue is independent during the encoding process, it is difficult to capture the deep semantic interaction between issues. This can easily lead to false recall when distinguishing issues with similar expressions but different meanings.

[0004] Therefore, there is an urgent need to provide an optimization method for question recall in knowledge base question answering systems to improve the accuracy and efficiency of question recall. Summary of the Invention

[0005] In view of this, this specification provides several embodiments of a question retrieval optimization method, computer device, and medium for a knowledge base question answering system, in order to improve the accuracy and efficiency of question retrieval in the knowledge base question answering system.

[0006] This specification provides a question retrieval optimization method for a knowledge base question answering system. The method includes: constructing a question pair training dataset, wherein the question pairs include positive samples and negative samples, the positive samples being question pairs consisting of a main question in the knowledge base and its corresponding similar questions, and the negative samples being question pairs consisting of a main question in the knowledge base and its dissimilar questions; using a cross-encoder teacher model to perform semantic similarity scoring on each question pair in the question pair training dataset, generating teacher scores corresponding to each question pair; performing soft labeling on the teacher scores to obtain soft-label training data; based on the soft-label training data, using a cosine similarity loss function to distill and train a dual-encoder student model, so that the cosine similarity between the two question vectors output by the dual-encoder student model approximates the teacher scores in the soft-label training data; deploying the distilled dual-encoder student model to the knowledge base question answering system, for encoding the user-input question into a vector, and performing vector retrieval in a pre-generated main question vector library in the knowledge base to recall the main question most similar to the user-input question and return its corresponding answer.

[0007] In some implementations, the soft labeling process includes a rating normalization mapping; the normalization mapping includes: mapping the teacher rating to the [0, 1] interval using min-max normalization; or, mapping the teacher rating to the interval near [0, 1] using a sigmoid function.

[0008] In some implementations, the soft labeling process includes low-quality score filtering; the low-quality score filtering includes: identifying samples whose normalized teacher scores fall within a preset neutral range as low-quality samples and removing them.

[0009] In some implementations, the soft labeling process includes score distribution calibration; the score distribution calibration includes: statistically analyzing the score distribution characteristics of positive and negative samples in the training dataset for the question; when the number of samples in the middle score interval is less than a preset requirement, binning the teacher scores and balancing the sample sampling in each score interval to make the gradient distribution of similarity scores in the training data reasonable.

[0010] In some implementations, the negative samples are constructed as follows: for each main question in the knowledge base, other main questions that do not belong to the set of similar questions of the main question are randomly selected from the knowledge base, and the current main question and the randomly selected other main questions are combined to form a negative sample; wherein the ratio of the number of positive samples to negative samples is 1:1.

[0011] In some implementations, the method further includes: collecting newly added main questions and their similar question annotation data in the knowledge base in real time, and / or collecting user questions and main question pairs verified as similar by the cross-encoder teacher model during the operation of the knowledge base question answering system as incremental data; performing quality screening on the incremental data to obtain high-quality incremental data; when the accumulated high-quality incremental data reaches a preset threshold, triggering incremental training of the currently deployed dual-encoder student model and performing incremental updates of the model, and performing vector encoding based on the newly added main questions to update the main question vector library in the knowledge base.

[0012] In some implementations, the method further includes: after each incremental training, evaluating the issue recall accuracy of the updated dual-encoder student model on a preset validation set; if the issue recall accuracy is lower than the model accuracy before the incremental update, then rolling back to the dual-encoder student model version before the incremental update and abandoning this incremental update.

[0013] In some implementations, the dual-encoder student model is a model based on a sentence-transformer architecture.

[0014] This specification provides a computer device including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the question recall optimization method of the knowledge base question answering system described in any of the above embodiments.

[0015] This specification provides a computer-readable storage medium storing a computer program thereon, which, when executed, implements the question recall optimization method of the knowledge base question-answering system described in any of the above embodiments.

[0016] In the various implementation methods provided in this specification, firstly, a question pair training dataset is constructed, comprising positive and negative samples. Positive samples are question pairs consisting of the main question in the knowledge base and its corresponding similar questions, while negative samples are question pairs consisting of the main question in the knowledge base and its dissimilar questions. Next, a cross-encoder teacher model is used to perform semantic similarity scoring on each question pair in the training dataset, generating teacher scores for each question pair. Then, soft-labeling is applied to the teacher scores to obtain soft-label training data. Next, based on the soft-label training data, a cosine similarity loss function is used to distill the dual-encoder student model, so that the cosine similarity between the two question vectors output by the dual-encoder student model approximates the teacher scores in the soft-label training data. Finally, the distilled dual-encoder student model is deployed to the knowledge base question answering system to encode the user-input question into a vector and perform vector retrieval in a pre-generated main question vector library in the knowledge base to recall the main question most similar to the user-input question and return its corresponding answer. This improves the question retrieval accuracy and efficiency of the knowledge base question answering system. Attached Figure Description

[0017] Figure 1 This is a flowchart illustrating the question retrieval optimization method for the knowledge base question-answering system provided in the embodiments of this specification; Figure 2 A schematic diagram of a question recall optimization device for a knowledge base question-answering system provided in the embodiments of this specification; Figure 3 A schematic diagram of a computer device provided for an embodiment of this specification. Detailed Implementation

[0018] To enable those skilled in the art to better understand the solutions described in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this specification, and not all of them. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.

[0019] This specification provides an example application scenario of a question retrieval optimization method for a knowledge base question answering system. The application scenario can be a knowledge base question answering system that can perform question retrieval based on a deployed distillation model.

[0020] In this embodiment, the knowledge base question-answering system can be a system that automatically answers natural language questions posed by users based on pre-stored knowledge entries. For example, in an enterprise customer service scenario, the knowledge base question-answering system can store common questions and their corresponding answers. When a user enters "How to process returns and exchanges," the system retrieves the most matching similar question from the knowledge base and returns the corresponding answer.

[0021] This specification provides a method for optimizing question recall in a knowledge base question-answering system. Please refer to [link / reference]. Figure 1 , Figure 1 This is a flowchart illustrating the optimization of question retrieval in a knowledge base question-answering system according to an embodiment of this specification. This embodiment provides the method and operation steps shown in the flowchart, but based on conventional or non-creative work, it may include more or fewer operation steps. The order of steps listed in the embodiment is merely one possible execution order among many, and does not represent the only possible execution order. In actual system or server product execution, the method shown in the embodiment can be executed sequentially or in parallel (e.g., in a parallel processor or multi-threaded processing environment). Specifically, as follows... Figure 1 As shown, the question recall optimization method of this knowledge base question answering system may include the following steps.

[0022] Step S110: Construct a training dataset of question pairs. Question pairs include positive samples and negative samples. Positive samples are question pairs consisting of the main question in the knowledge base and its corresponding similar questions. Negative samples are question pairs consisting of the main question in the knowledge base and its dissimilar questions.

[0023] In some cases, in knowledge-based question-answering systems, question retrieval is the process of selecting one or more main questions from a pool of main questions stored in the knowledge base that are semantically most similar to the user's input question. For knowledge-based question-answering systems, it is necessary to improve the accuracy and speed of question retrieval to enhance the overall question-answering performance.

[0024] A question pair is a pairing structure consisting of two question texts, used to characterize the similarity relationship between semantic units to be compared.

[0025] Specifically, positive samples refer to semantically similar question pairs, that is, question pairs consisting of a main question and its corresponding similar questions in the knowledge base. The main question is a predefined and stored standard question in the knowledge base, representing a specific semantic intent. For example, for the intent "how to apply for a refund," the main question could be "how to apply for a refund." Similar questions are variant questions that are semantically identical or highly similar to the main question but expressed differently. For example, for the main question "how to apply for a refund," similar questions could include "how to get a refund," "how to process a refund," "what are the steps to apply for a refund," etc. A main question and its corresponding similar questions constitute positive samples, meaning that the question pair is semantically labeled as similar.

[0026] Specifically, negative samples refer to semantically dissimilar question pairs, that is, question pairs consisting of the main question and its dissimilar questions in the knowledge base. Dissimilar questions refer to other questions that are semantically different from the main question and do not belong to the set of similar questions to the main question. For example, for the main question "How to apply for a refund", dissimilar questions could be "How to change the password" or "How to check the order status". By constructing positive and negative samples, the training dataset can provide clear criteria for judging similarity and dissimilarity for subsequent models, thereby supporting supervised distillation training.

[0027] Step S120: Use the cross-encoder teacher model to perform semantic similarity scoring on each question pair in the question pair training dataset, and generate the teacher score corresponding to each question pair.

[0028] The teacher model is a high-precision model that serves as a knowledge source, providing training targets for subsequent student models. The cross-encoder teacher model refers to a model that uses a cross-encoder as the teacher. It is a pre-trained language model employing a cross-encoder architecture, which can concatenate two input texts and input them into the model. Through the model's internal attention mechanism, it achieves deep semantic interaction and joint encoding between the two input texts.

[0029] In this embodiment, a cross-encoder teacher model can be invoked to jointly encode the two question texts (i.e., the main question and the similar question) in each question pair in the training dataset, outputting a semantic similarity score that reflects the degree of semantic similarity between the two. This semantic similarity score is used as the original teacher score, which can be used as the soft label target value in subsequent distillation training. In this way, the high-precision semantic understanding capability of the cross-encoder teacher model can be used to provide accurate similarity judgments for each question pair, laying the foundation for subsequent knowledge transfer.

[0030] Specifically, the cross-encoder teacher model is a cross-encoder model trained on a question repetition detection dataset. For example, the cross-encoder teacher model can be built on either a cross-encoder model or a quora-distilroberta-base model.

[0031] Step S130: Perform soft labeling on the teacher ratings to obtain soft label training data.

[0032] In some cases, to improve the quality and applicability of raw teacher ratings as training labels, corresponding preprocessing operations can be performed, such as soft label processing. Soft labels refer to continuous value labels relative to hard labels (such as binary labels of 0 or 1). They can retain the confidence information and fuzzy boundary information of the teacher model's judgment of samples, and can provide richer training signals for the student model.

[0033] In this embodiment, the soft-labeled training data is a set of triplet datasets that has been processed with soft labels and can be directly used for model training or distillation training. Each triplet contains the main question, candidate questions, and processed teacher ratings. Soft labeling addresses issues such as the inconsistent range of teacher ratings in the original teacher model output, skewed rating distribution, and interference from low-confidence samples, making the subsequent distillation training data more stable and reliable, thereby improving the distillation effect.

[0034] Step S140: Based on the soft-label training data, use the cosine similarity loss function to perform distillation training on the dual encoder student model so that the cosine similarity between the two question vectors output by the dual encoder student model approximates the teacher rating in the soft-label training data.

[0035] In this embodiment, the parameters of the dual-encoder student model can be iteratively updated by using teacher ratings in the soft-label training data as the regression target and cosine similarity loss function as the optimization target.

[0036] A student model refers to the model to be trained. A dual-encoder student model means that a dual encoder is used as a student model. In other words, a dual-encoder student model is a neural network model that uses a dual-encoder architecture. It encodes two input questions separately through independent encoders, generating their own question vector representations, and then determines the semantic similarity between the two questions by calculating the similarity between the vectors.

[0037] The cosine similarity loss function is used to measure the difference between the cosine similarity between two vectors and the target value. Specifically, the cosine similarity loss function calculates the cosine similarity between the main question vector and the candidate question vectors output by the student model, and uses the mean squared error between this cosine similarity and the teacher ratings in the soft-labeled training data as the loss value. This loss value is then used to update the parameters of the student model through backpropagation. The cosine similarity measures the degree of directional alignment between the two vectors by calculating the cosine of the angle between them.

[0038] For example, the cosine similarity value ranges from -1 to 1. The closer the value is to 1, the more consistent the directions of the two vectors are and the more semantically similar they are. As an example, for a positive sample consisting of the main question and its corresponding similar questions, that is, the question pair is semantically labeled as similar, its teacher rating should be close to 1.

[0039] Distillation training uses the output of the cross-encoder teacher model as the learning objective. It optimizes the parameters of the dual-encoder student model to make its predictions approximate the output of the cross-encoder teacher model. Through distillation training, the dual-encoder student model learns to approximate the question similarity judgment ability of teacher ratings in the soft-label training data. In other words, it makes the cosine similarity between the two question vectors output by the dual-encoder student model approximate the teacher ratings in the soft-label training data. Thus, the high-precision semantic similarity judgment ability of the cross-encoder teacher model can be effectively transferred to the dual-encoder student model, allowing it to maintain the high-speed retrieval advantage of the dual-encoder architecture while achieving similarity judgment accuracy close to that of the cross-encoder teacher model.

[0040] Specifically, the dual-encoder student model can be a model based on a sentence-transformer architecture. For example, the dual-encoder student model can be built based on either the all-MiniLM-L6-v2 model or the all-mpnet-base-v2 model.

[0041] For example, the parameters for distillation training may include, but are not limited to, the number of training epochs, batch size, learning rate, and warmup proportion.

[0042] As an example, the training epochs are 4, the batch size is 16, and the learning rate is 2e. -5 The preheating step ratio is 0.1. Of course, it's understandable that the parameters for distillation training can be determined based on the actual parameters.

[0043] The number of training rounds refers to the total number of times the entire training dataset is input into the model for forward and backward propagation optimization.

[0044] Batch size refers to the number of question pairs that the model processes in parallel at one time when updating its internal weight parameters. Specifically, it involves taking 16 triples (main question, candidate questions, teacher scores) from the training set each time to allow the model to calculate the loss and gradient.

[0045] The learning rate is set to 2e -5 That is, 0.00002, which is the step size coefficient that the optimizer (such as AdamW) uses to update the model weights during backpropagation.

[0046] Setting the warm-up step ratio to 0.1 means that at the very beginning of training, the learning rate is not directly set to 2e. -5 Instead, it starts with 0 or a very small value, and gradually increases linearly to 2e in the first 10% of the total training steps. -5 .

[0047] Step S150: Deploy the dual-encoder student model trained by distillation to the knowledge base question answering system to encode the user input question into a vector and perform vector retrieval in the pre-generated main question vector library in the knowledge base to recall the main question most similar to the user input question and return its corresponding answer.

[0048] The questions entered by users refer to the unanswered questions posed by end users of the knowledge base question-answering system in natural language.

[0049] In this embodiment, the dual-encoder student model, having completed distillation training, can be loaded and applied to the online operating environment of a knowledge-based question-answering system to process questions input by actual users. Specifically, user questions in text form can be input into the dual-encoder student model, and the model's forward computation can transform and encode them into a dense vector representation of fixed dimensions.

[0050] The main question vector library is a set of vector indexes pre-encoded into vectors by a dual-encoder student model from all main questions in the knowledge base. It is used to support efficient vector similarity retrieval, enabling the use of the user's input question vector as the query vector. The most similar vector is found in the main question vector library by calculating the similarity between vectors. In other words, one or more main questions in the main question vector library with the highest cosine similarity to the user's input question vector are used as the recall results, and the pre-associated answers of the recall results are obtained from the knowledge base and provided to the user.

[0051] In the above embodiments, the dual-encoder student model trained by distillation can utilize its high-precision question coding capability to achieve end-to-end fast question-answering processing from user input question to final answer, which can improve the response speed of question recall while ensuring recall accuracy.

[0052] In some implementations, step S130, soft label processing may include score normalization mapping. That is, the teacher scores output by the cross-encoder teacher model are subjected to numerical range transformation processing to map data with different numerical ranges to a unified target numerical interval through a specific transformation.

[0053] In this embodiment, the normalization mapping may include: using minimum-maximum normalization to map teacher ratings to the interval [0, 1]; or using the sigmoid function to map teacher ratings to the interval near [0, 1].

[0054] Specifically, when mapping teacher ratings to the [0, 1] interval using min-maximum normalization, each teacher rating can be transformed to the [0, 1] interval through a linear transformation based on the minimum and maximum values ​​of the teacher ratings. For example, for any teacher rating, the minimum rating can be subtracted from the teacher rating and then divided by the difference between the maximum and minimum ratings.

[0055] Specifically, the sigmoid function is a sigmoid activation function, and sigmoid(x) equals 1 divided by 1 plus e raised to the power of negative x. When using the sigmoid function to map teacher ratings to the interval [0, 1], the teacher ratings can be input into the sigmoid function, and a nonlinear transformation can be used to map them to the interval (0, 1).

[0056] As an example, the numerical distribution range of any teacher rating output by the cross-encoder teacher model may be unknown or may contain extreme values. When the range of teacher ratings output by the cross-encoder teacher model is known, min-max normalization can be used to map the teacher ratings to the [0, 1] interval; when the range of teacher ratings output by the cross-encoder teacher model is unknown or contains extreme values, the sigmoid function can be used to map the teacher ratings to the interval near [0, 1]. Thus, the teacher ratings processed through normalization mapping ensure that they fall within the [0, 1] label range required by the cosine similarity loss function, allowing distillation training to proceed normally. The closer the value is to 1, the more consistent the directions and the more semantically similar the two vectors are.

[0057] In some implementations, step S130 may include low-quality score filtering. That is, after normalization mapping, samples whose normalized teacher scores fall within a preset neutral range may be identified as low-quality samples and removed. This can identify and remove samples that may lead to a decline in training effectiveness due to high uncertainty in the teacher model's judgment.

[0058] The preset neutral interval is a pre-defined scoring range. Scores within this range represent a low level of confidence in the teacher model's judgment regarding the similarity of corresponding question pairs. When a teacher's score falls within the preset neutral interval, it indicates that the teacher model lacks a clear tendency to judge the similarity of question pairs. If it continues to be used as a training label, it will introduce noise rather than a valid signal, and therefore should be removed. Since the teacher scores output by the cross-encoder teacher model are in the [0, 1] interval after normalization mapping, the preset neutral interval is the middle region of this [0, 1] interval.

[0059] As an example, the preset neutral interval can be the range [0.45, 0.55]. Alternatively, it can be within the range [0.45, 0.55], such as [0.48, 0.52]. Of course, the preset neutral interval also includes the range [0.45, 0.55], for example, it can be [0.4, 0.6].

[0060] In the above implementation, filtering by low-quality scores can improve the confidence of labels in distillation training data, reduce the interference of noisy samples on training, and thus improve the accuracy of the model after distillation.

[0061] In some implementations, step S130 may include rating distribution calibration. That is, the sample distribution of each rating interval of the teacher ratings is adjusted, or in other words, the distribution of each teacher rating used as a sample in the soft-label training data is adjusted.

[0062] Specifically, rating distribution calibration may include: statistically analyzing the rating distribution characteristics of positive and negative samples in the training dataset; when the number of samples in the middle rating interval is less than a preset requirement, binning the teacher ratings and balancing the sample sampling within each rating interval to ensure a reasonable gradient distribution of similarity ratings in the training data.

[0063] For example, the distribution of teacher ratings for positive and negative samples in the training dataset can be analyzed to obtain the sample quantity distribution for each rating interval, and intervention and adjustment can be made for imbalanced sample distribution.

[0064] The intermediate rating interval is the middle region between the two ends of the rating range [0, 1]. For example, the intermediate rating interval could be the range [0.2, 0.8]. When the number of samples in the intermediate rating interval is less than the preset requirement, it indicates that the training dataset lacks progressive similarity rating samples, which may cause the dual-encoder student model to have difficulty learning the ability to make continuous transition judgments from similarity to dissimilarity.

[0065] The preset requirement is a minimum standard for the number of samples in the intermediate scoring range. For example, the proportion of samples in the intermediate scoring range to the total number of samples can be set to be no less than a certain threshold.

[0066] Binning is a process that divides a continuous scoring interval into multiple discrete sub-intervals and assigns each sample to the corresponding sub-interval based on its scoring value.

[0067] Sampling balancing involves supplementing sampling for scoring intervals with insufficient sample size and undersampling for scoring intervals with excessive sample size after binning, so as to make the sample size in each scoring interval roughly balanced.

[0068] In the above implementation, by calibrating the score distribution, it can be ensured that the gradient distribution of similarity scores in the training dataset is reasonable, so that the student model can fully learn the ability to judge full gradient similarity from completely dissimilar to highly similar during the distillation training process, thereby improving the problem recall accuracy of the model under various similarity boundary conditions.

[0069] In some implementations, negative samples can be constructed as follows: for each main question in the knowledge base, other main questions that do not belong to the set of similar questions of the main question are randomly selected from the knowledge base, and the current main question and the randomly selected other main questions are combined to form a negative sample.

[0070] In this model, the ratio of positive to negative samples is 1:1. Setting this ratio to 1:1 ensures that the model encounters a balanced number of positive and negative examples during training, avoiding training bias caused by class imbalance. Of course, depending on the specific circumstances, the ratio of positive to negative samples can be other than 2:1 or 3:2, etc., and is not limited here.

[0071] For any main question, its set of similar questions refers to the set of all similar questions in the knowledge base that are associated with the main question, and each similar question in the set of similar questions is semantically the same as or highly similar to the main question.

[0072] By constructing the above method, we can ensure that the two problems in the negative sample are semantically dissimilar, thus providing a clear basis for judging negative examples for model training.

[0073] In some implementations, the problem recall optimization method further includes the following steps S210-S230.

[0074] Step S210: Collect newly added main questions and their similar question annotation data in the knowledge base in real time, and / or collect user questions and main question pairs that are verified as similar by the cross-encoder teacher model during the operation of the knowledge base question answering system as incremental data.

[0075] In some cases, it is necessary to perform online incremental updates so that the parameters of the already trained model can be fine-tuned using new data without a full retraining, thereby improving the model's performance on new data.

[0076] Specifically, when a new main question is added to the knowledge base or similar question annotations are added to an existing main question, the newly added question pairing data is collected. When the online knowledge base question answering system processes user questions, if a pairing between a user question and a main question is determined to be similar by the cross-encoder teacher model (the semantic similarity score of the cross-encoder teacher model for the output of this question pair exceeds a preset threshold), then the pairing is recorded as potentially valid training data. Incremental data refers to the new sample data collected in the above process that can be used to further train the model. In this way, through real-time data collection, the training set can be continuously expanded using the real matching data generated during system operation, enabling the model to adapt to the dynamic changes of the knowledge base.

[0077] Step S220: Perform quality screening on the incremental data to obtain high-quality incremental data.

[0078] Specifically, quality checks can be performed on the collected incremental data to exclude low-quality samples that do not meet training requirements, thus achieving quality screening. Quality screening includes, but is not limited to: removing question pairs that are duplicates of existing training data, removing samples with teacher rating confidence levels below a preset threshold, and removing samples with abnormal formats or incomplete content. High-quality incremental data refers to incremental samples that meet training quality requirements and are retained after quality screening. Through quality screening, the reliability of the data used for incremental training can be ensured, preventing low-quality data from contaminating the accuracy of the model after distillation training.

[0079] Step S230: When the accumulated high-quality incremental data reaches the preset threshold, the current deployed dual encoder student model is incrementally trained and the model is incrementally updated. Vector encoding is performed based on the newly added main question to update the main question vector library in the knowledge base.

[0080] The preset threshold is a pre-defined minimum data volume standard for triggering incremental training. This avoids system instability caused by frequent training due to insufficient data.

[0081] Specifically, when the amount of accumulated high-quality incremental data reaches a preset threshold, the incremental training process can be initiated. Incremental training refers to using the existing weights of the currently deployed dual-encoder student model as the initial state, further optimizing the model's parameters using newly added high-quality incremental data, and then replacing the currently deployed old model with the new model after incremental training, so that it can be effective in subsequent question-answering services, thus achieving incremental updates.

[0082] In addition, in this embodiment, for newly added main questions in the knowledge base, the incrementally updated dual-encoder student model can be used to encode them into vectors, and the vectors can be added to the main question vector library, so that the newly added main questions can be recalled by subsequent user queries.

[0083] Through the incremental training and vector library update mechanisms described above, the distillation model can continuously adapt to changes in the knowledge base and maintain the long-term timeliness of recall accuracy.

[0084] In some implementations, the preset threshold or trigger threshold for incremental training can be set to accumulate 200 high-quality incremental data points. The learning rate used for incremental training can be 1 / 5 of the learning rate used in distillation training, and the number of training rounds is 1 to 2.

[0085] In some implementations, it is necessary to monitor and roll back incrementally after training in order to improve the stability and reliability of the online knowledge base question answering system.

[0086] Therefore, after incremental training in step S230, the problem recall optimization method may also include the following steps S310-S320.

[0087] Step S310: After each incremental training, evaluate the question recall accuracy of the updated dual-encoder student model on a predefined validation set.

[0088] The validation set is a pre-defined set of question pairs from the knowledge base that are not involved in the training process, and is used to independently evaluate the generalization performance of the model.

[0089] For example, the issue recall precision in this step can refer to the proportion of the model that correctly recalls the main issue on the validation set.

[0090] Step S320: If the problem recall accuracy is lower than the model accuracy before the incremental update, roll back to the dual encoder student model version before the incremental update and abandon this incremental update.

[0091] Rollback means restoring the model state to a previous version before incremental training.

[0092] Through the aforementioned monitoring and rollback mechanisms, it can be ensured that each incremental training is carried out on the premise of positively improving model performance, avoiding model performance degradation due to fluctuations in incremental data quality or distribution shifts, thereby ensuring the stability and reliability of the online system.

[0093] This specification provides a question retrieval optimization device for a knowledge base question-answering system. Please refer to [link / reference]. Figure 2 The problem recall optimization device may include a training data construction module 410, a teacher model scoring module 420, a scoring preprocessing module 430, a distillation training module 440, and a model deployment optimization module 450.

[0094] The training data construction module 410 is used to construct a training dataset of question pairs. The question pairs include positive samples and negative samples. Positive samples are question pairs consisting of the main question in the knowledge base and its corresponding similar questions, while negative samples are question pairs consisting of the main question in the knowledge base and its dissimilar questions. The teacher model scoring module 420 is used to perform semantic similarity scoring on each question pair in the question pair training dataset using the cross-encoder teacher model, and generate the teacher score corresponding to each question pair. The scoring preprocessing module 430 is used to perform soft labeling on teacher scores to obtain soft label training data. Distillation training module 440 is used to distill the dual encoder student model based on soft-label training data and use the cosine similarity loss function to make the cosine similarity between the two question vectors output by the dual encoder student model approximate the teacher rating in the soft-label training data. The model deployment optimization module 450 is used to deploy the distillation-trained dual-encoder student model to the knowledge base question answering system. It encodes the user-input question into a vector and performs vector retrieval in the pre-generated main question vector library in the knowledge base to recall the main question most similar to the user-input question and return its corresponding answer.

[0095] The specific functions and effects of the issue recall optimization device can be explained by referring to other embodiments in this specification, and will not be repeated here. Each module in the issue recall optimization device can be implemented entirely or partially through software, hardware, or a combination thereof. Each module can be embedded in the processor of the computer device in hardware form or independent of it, or it can be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0096] This specification also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a computer, implements the problem recall and optimization method described in any of the above embodiments.

[0097] This specification also provides a computer program product containing instructions that, when executed by a computer, cause the computer to implement the problem recall optimization method described in any of the above embodiments.

[0098] This specification also provides a computer device, including a memory and a processor, wherein the memory stores a computer program that, when executed by the processor, implements the problem recall optimization method described in any of the above embodiments.

[0099] In some implementations, please refer to Figure 3 The computer device can be a terminal, and its internal structure diagram can be as follows: Figure 3 As shown, the computer device includes a processor, memory, and a communication interface connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. When the computer program is executed by the processor, it implements a problem recall optimization method.

[0100] It is understood that the specific examples in this document are only intended to help those skilled in the art better understand the embodiments described herein, and are not intended to limit the scope of the invention.

[0101] It is understood that in the various embodiments described in this specification, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments described in this specification.

[0102] It is understood that the various implementation methods described in this specification can be implemented individually or in combination, and the implementation methods in this specification are not limited in this respect.

[0103] Unless otherwise stated, all technical and scientific terms used in the embodiments of this specification have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this specification. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items. The singular forms "a," "the," and "the" as used in the embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0104] It is understood that the processor in the embodiments of this specification can be an integrated circuit chip with signal processing capabilities. In implementation, each step of the above method embodiments can be completed by integrated logic circuits in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this specification. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this specification can be directly implemented by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above methods.

[0105] It is understood that the memory in the embodiments of this specification may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. Non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. Volatile memory may be random access memory (RAM). It should be noted that the memory in the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.

[0106] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this specification.

[0107] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the aforementioned method implementations, and will not be repeated here.

[0108] The above description is merely a specific embodiment of this specification, but the scope of protection of this invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in this specification should be included within the scope of protection of this specification. Therefore, the scope of protection of this invention should be determined by the scope of the claims.

Claims

1. A question recall optimization method for a knowledge base question answering system, characterized in that, The method includes: Construct a training dataset of question pairs, wherein the question pairs include positive samples and negative samples. The positive samples are question pairs consisting of the main question in the knowledge base and its corresponding similar questions, and the negative samples are question pairs consisting of the main question in the knowledge base and its dissimilar questions. The cross-encoder teacher model is used to perform semantic similarity scoring on each question pair in the training dataset, generating teacher scores for each question pair. The teacher ratings are processed with soft labels to obtain soft-label training data. Based on the soft-label training data, the dual-encoder student model is distilled using the cosine similarity loss function so that the cosine similarity between the two question vectors output by the dual-encoder student model approximates the teacher rating in the soft-label training data. The dual-encoder student model trained by distillation is deployed to the knowledge base question answering system to encode the user-input question into a vector and perform vector retrieval in the pre-generated main question vector library in the knowledge base to recall the main question most similar to the user-input question and return its corresponding answer.

2. The problem recall optimization method according to claim 1, characterized in that, The soft label processing includes a rating normalization mapping; the normalization mapping includes: The teacher ratings are mapped to the [0, 1] interval using minimum-maximum normalization. Alternatively, the sigmoid function can be used to map the teacher ratings to the interval [0, 1].

3. The problem recall optimization method according to claim 2, characterized in that, The soft label processing includes low-quality score filtering; the low-quality score filtering includes: Samples whose normalized teacher ratings fall within the preset neutral range are identified as low-quality samples and removed.

4. The problem recall optimization method according to any one of claims 1-3, characterized in that, The soft label processing includes score distribution calibration; The scoring distribution calibration includes: The scoring distribution characteristics of positive and negative samples in the training dataset are statistically analyzed. When the number of samples in the middle scoring interval is less than the preset requirement, the teacher scores are binned and the samples are balanced within each scoring interval to make the gradient distribution of similarity scores in the training data reasonable.

5. The problem recall optimization method according to claim 1, characterized in that, The negative samples are constructed in the following way: For each main question in the knowledge base, other main questions that do not belong to the set of similar questions of the main question are randomly selected from the knowledge base, and the current main question and the randomly selected other main questions are used to form a negative sample; wherein, the ratio of the number of positive samples to negative samples is 1:

1.

6. The problem recall optimization method according to claim 1, characterized in that, The method further includes: Real-time collection of newly added main questions and their similar question annotation data in the knowledge base, and / or collection of user questions and main question pairs verified as similar by the cross-encoder teacher model during the operation of the knowledge base question answering system, as incremental data; The incremental data is subjected to quality screening to obtain high-quality incremental data; Once the accumulated high-quality incremental data reaches a preset threshold, it triggers incremental training of the currently deployed dual-encoder student model, incremental updates of the model, and vector encoding based on the newly added main question to update the main question vector library in the knowledge base.

7. The problem recall optimization method according to claim 6, characterized in that, The method further includes: After each incremental training, the problem recall accuracy of the updated dual-encoder student model is evaluated on a predefined validation set. If the recall accuracy of the problem is lower than the model accuracy before the incremental update, then roll back to the dual-encoder student model version before the incremental update and abandon this incremental update.

8. The problem recall optimization method according to claim 1, characterized in that, The dual-encoder student model is based on a sentence-transformer architecture.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the problem recall optimization method according to any one of claims 1 to 8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the problem recall optimization method according to any one of claims 1 to 8.