A multi-level document retrieval method

By using the problem state classification model and lag impact evaluation model of the SSM convolution module in a multi-round dialogue system, the problem of problem state and dynamically optimized the knowledge retrieval strategy is solved, and the problems of knowledge retrieval error and computing resource consumption in traditional systems are improved, and user interaction experience and knowledge retrieval efficiency are improved.

CN119415482BActive Publication Date: 2025-05-27无锡锡商银行股份有限公司
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510035545.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-09
Publication Date
2025-05-27
Estimated Expiration
2045-01-09

AI Technical Summary

Technical Problem

Traditional multi-round dialogue systems rely too much on exponential decay strategies based on time dimensions in knowledge retrieval, resulting in reduced effectiveness of key information in the far rounds, resulting in knowledge retrieval errors, and increased computing resource consumption, affecting user interaction experience.

Method used

The problem state classification model based on the SSM convolution module is used to identify the problem state of the user's single-round questions. Different knowledge retrieval strategies are adopted according to the recognized status, and a single-round search knowledge set is obtained, and a lag impact assessment model is constructed to identify potential lag potential risks. When there is a potential risk of lag, perform an adaptive correction preset exponential decay strategy.

Benefits of technology

By accurately identifying the problem status and combining historical dialogue information, dynamic optimization and update of knowledge retrieval sets can be achieved, the accuracy and efficiency of knowledge retrieval in multiple rounds of dialogue scenarios can be improved, the effectiveness loss of historical information can be reduced, and resource usage can be optimized to reduce response time.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119415482B_ABST
    Figure CN119415482B_ABST
Patent Text Reader

Abstract

The present invention discloses a multi-level document retrieval method, which specifically relates to the technical field of multi-level document retrieval. In a single-round knowledge retrieval, the state of the question is first discriminated, and then the corresponding retrieval strategy is executed according to the state to obtain a retrieval knowledge set based on the current-round question, and multiple pieces of information in the process of question state recognition by the question state classification model are obtained. A lag impact evaluation model is constructed to timely identify potential lag hazards, optimize resource usage, reduce response time, and improve the user interaction experience. When there are no lag hazards in the question state classification model, the single-round retrieval knowledge set is updated through multi-round knowledge rolling. The exponential decay strategy is used to obtain the knowledge influence degree coefficient and combined with the in-round knowledge importance ranking coefficient to comprehensively sort the knowledge, and a knowledge ranking table is obtained for display to the user, improving the accuracy and efficiency of knowledge retrieval in a multi-round dialogue scenario.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multi-level document retrieval, and more specifically, the present invention relates to a multi-level document retrieval method. Background Art

[0002] With the continuous development of artificial intelligence technology, especially in the field of natural language processing (NLP), multi-turn dialogue systems have gradually become an important part of human-computer interaction. Such systems can understand and process continuous user inputs and provide an efficient communication experience through appropriate responses. However, traditional dialogue systems often face challenges in knowledge retrieval. Especially in scenarios where multi-turn conversations require combining context information, an exponential decay strategy based on the time dimension is often used to update the multi-turn retrieval knowledge set. Over-reliance on the exponential decay strategy based on the time dimension may not accurately capture key information in distant turns, resulting in a gradual decrease in the effectiveness of key information in distant turns and ultimately errors in knowledge retrieval. For example, a user has a conversation with a customer service robot and starts by asking about the malfunction of a certain mobile phone of the company. In the first turn of the conversation, the user mentions the specific model of the mobile phone and asks about related malfunction information. By the 6th turn, the user suddenly mentions the malfunction-related problem again, but this time the question is about the accessory warranty policy. Due to the exponential decay strategy based on the time dimension, the system significantly reduces the weight of the first-turn conversation. Therefore, in the 6th turn, the system cannot effectively retrieve the specific model information from the first turn, resulting in a mismatch between the retrieved warranty policy and the mobile phone model, and ultimately the reply to the user cannot solve the user's problem. Moreover, in real-time conversation scenarios and when performing convolutional operations and retrievals on large-scale historical information, the exponential decay strategy based on the time dimension will cause a significant increase in the system's computational resource consumption, which may lead to delays in the retrieval response time or even lags. Since the system cannot provide more accurate retrieval-related knowledge for the user's questions, it seriously affects the user's interaction experience. As a result, the user may leave the platform, leading to a decrease in the platform's credibility. Summary of the Invention

[0003] In order to overcome the above-mentioned defects of the prior art, an embodiment of the present invention provides a multi-level document retrieval method to solve the problems raised in the above background art.

[0004] To achieve the above object, the present invention provides the following technical solutions:

[0005] A multi-level document retrieval method includes the following steps:

[0006] Step S1, identify the problem state of the user's single-turn question according to the problem state classification model based on the SSM convolutional module, obtain the problem state, and adopt different knowledge retrieval strategies according to the problem state to obtain the single-turn retrieval knowledge set corresponding to the problem state;

[0007] Step S2, obtain multiple pieces of information in the problem status recognition process of the problem status classification model, construct a stuttering impact assessment model, generate a stuttering impact assessment index, and determine whether there is a stuttering hidden danger in the problem status classification model;

[0008] Step S3, when there is no stuttering hidden danger in the problem status classification model, perform multi-round knowledge rolling update on the single-round retrieval knowledge set, obtain the knowledge influence degree coefficient by using the preset exponential decay strategy, and comprehensively sort the knowledge in combination with the in-round knowledge importance ranking coefficient to obtain a knowledge ranking table for display to the user;

[0009] Step S4, when there is a stuttering hidden danger in the problem status classification model, perform adaptive correction on the preset exponential decay strategy.

[0010] In a preferred embodiment, in step S1, first preprocess the text of the user's current-round question, perform word segmentation on the question text, convert each word into a corresponding word vector to obtain the semantic vector of the current-round question; at the same time, preprocess and perform word segmentation and word vector conversion operations on the text of the user's historical multi-round questions and the corresponding-round retrieval knowledge set to obtain the semantic vectors of the historical multi-round questions and the semantic vector matrix of the retrieval knowledge set respectively;

[0011] Construct a question text semantic vector matrix based on the semantic vector of the current-round question and the semantic vectors of the historical multi-round questions, input the question text semantic vector matrix into the question SSM convolution module, and perform convolution on the question text semantic vector matrix according to the convolution kernel of the question SSM convolution module;

[0012] Input the semantic vector matrix of the retrieval knowledge set into the retrieval knowledge SSM convolution module, and perform convolution on the question text semantic vector matrix according to the convolution kernel of the retrieval knowledge SSM convolution module;

[0013] Horizontally splice the convolution output feature vectors extracted from the question SSM convolution module and the retrieval knowledge SSM convolution module to obtain the final output feature vector, and input the horizontally spliced final output feature vector into the fully connected layer to obtain the output of the fully connected layer;

[0014] Input the output of the fully connected layer into the Softmax activation function to obtain the probability value of the problem status, and the obtained expression is as follows, , where represents the probability value that the question of the user's current-round question belongs to the problem status , , represents the output of the fully connected layer;

[0015] Problem status Indicates an independent new problem;

[0016] The problem status M2 indicates that it is not independent and is the detail of the problem in the previous round;

[0017] The problem status M3 indicates that it is not independent and has an intersection with the problem in the previous round;

[0018] Classify the problem status of the user's current round of question according to the probability value output by the Softmax activation function, as follows:

[0019] If > And > , then mark the problem status of the user's current round of question as M1, obtain the semantic vector of the user's current round of question, perform semantic similarity matching with the semantic vectors of the existing knowledge in the knowledge base, sort the existing knowledge in the knowledge base from large to small according to the semantic similarity, and select the top top_k pieces of knowledge as the single-round retrieval knowledge set;

[0020] If > And > , then mark the problem status of the user's current round of question as M2, and directly use the retrieval knowledge set of the previous round of question as the single-round retrieval knowledge set of the current round;

[0021] If > And > , then mark the problem status of the user's current round of question as M3, splice the previous round of question and the current round of question, obtain the semantic vector of the spliced text, then perform semantic similarity matching with the semantic vectors of the existing knowledge in the knowledge base, sort the existing knowledge in the knowledge base from large to small according to the semantic similarity, and select the top top_k pieces of knowledge as the single-round retrieval knowledge set.

[0022] In a preferred embodiment, in step S2, the multiple pieces of information in the problem status recognition process include the model execution time anomaly coefficient, the resource consumption status coefficient, the convolutional layer number anomaly fluctuation coefficient, and the user churn rate.

[0023] In a preferred embodiment, the acquisition logic of the model execution time anomaly coefficient is as follows:

[0024] Obtain the execution inference time for classifying the problem status by the problem status classification model each time the user asks a question , where g represents the number of the user's question, g = {1, 2,..., G}, and G is a positive integer;

[0025] Calculate the average value of the execution inference time , the expression is as follows , calculate the standard deviation of the execution inference time based on the average value of the execution inference time , the expression is as follows ;

[0026] Compare the execution inference time of each time with the standard deviation of the execution inference time, and mark the execution inference time as abnormal. When the execution inference time is greater than or equal to the standard deviation of the execution inference time, the execution inference time greater than or equal to the standard deviation of the execution inference time is marked as an abnormal execution inference time;

[0027] Calculate the abnormal coefficient of the model execution time , the expression is as follows , where represents the h-th abnormal execution inference time, h = {1, 2,..., H}, and H is a positive integer.

[0028] In a preferred embodiment, the acquisition logic of the resource consumption status coefficient is as follows:

[0029] Obtain the resource consumption data during the operation of the problem status classification model, including but not limited to CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate;

[0030] It is necessary to normalize the obtained CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate, and calculate the resource consumption status coefficient based on the normalized CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate , the expression is as follows , where respectively represent the normalized CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate, , , , respectively represent the maximum acceptable expected values of the CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate, respectively represent of the preset proportionality coefficient, and are all greater than 0.

[0031] In a preferred embodiment, the acquisition logic of the abnormal fluctuation coefficient of the number of convolutional layers is as follows:

[0032] Obtain the dataset of the number of convolutional layers when the problem status classification model performs convolutional operations, and mark the dataset of the number of convolutional layers as , Denote the convolutional layer value when the problem status classification model performs the r-th convolutional operation, where r = {1, 2, ..., R} and R is a positive integer; extract the unique convolutional layer values from the convolutional layer number dataset to obtain the unique value set , Denote the unique value of the k-th convolutional layer value in the convolutional layer number dataset, where k = {1, 2, ..., K} and K is a positive integer; traverse the convolutional layer values in the convolutional layer number dataset and count each occurrence times ; calculate the probability of the unique value of each convolutional layer value , and the expression is as follows ;

[0033] Calculate the entropy value of the convolutional layer number , and the expression is as follows ;

[0034] Calculate the abnormal fluctuation coefficient of the convolutional layer number , and the expression is as follows , where Denote the standard deviation of the convolutional layer values in the convolutional layer number dataset, and the expression is as follows , Denote the average value of the convolutional layer values in the convolutional layer number dataset, and the expression is as follows .

[0035] In a preferred embodiment, the acquisition logic of the user churn rate is as follows:

[0036] Record the total number of active users within the preset fixed time period ZT and the number of users who stop using the problem status classification model , and calculate the user churn rate , and the expression is as follows .

[0037] In a preferred embodiment, normalize the obtained model execution time anomaly coefficient, resource consumption status coefficient, convolutional layer number abnormal fluctuation coefficient, and user churn rate, and construct a stuttering impact evaluation model based on the normalized model execution time anomaly coefficient, resource consumption status coefficient, convolutional layer number abnormal fluctuation coefficient, and user churn rate to generate a stuttering impact evaluation index , and the formula on which the stuttering impact evaluation model is based is as follows , where respectively denote the preset proportionality coefficients of the model execution time anomaly coefficient, resource consumption status coefficient, convolutional layer number abnormal fluctuation coefficient, and user churn rate, and are all greater than 0;

[0038] Compare the carding impact evaluation index with the preset carding impact evaluation index threshold to determine whether there is a carding hidden danger in the problem status classification model, as follows:

[0039] If the carding impact evaluation index is greater than the carding impact evaluation index threshold, generate a carding hidden danger signal; if the carding impact evaluation index is less than or equal to the carding impact evaluation index threshold, generate a model running stable signal.

[0040] In a preferred embodiment, in step S3, a preset exponential decay strategy is used to obtain the knowledge influence degree coefficient, as follows:

[0041] Obtain the retrieval knowledge set of the user's historical multi-round questions, sort the retrieval knowledge set of the historical multi-round questions from far to near according to the time sequence, and record the earliest round as and the current round as , , L is a positive integer; for the round number , the knowledge influence degree coefficient of the retrieval knowledge set of its participating in multi-round knowledge rolling update is calculated as follows ; ;

[0042] For the knowledge in each round of the retrieval knowledge set, no longer simply rely on the semantic similarity score, and use the in-round knowledge importance ranking coefficient based on the semantic similarity ranking for evaluation. For the knowledge , the calculation expression of its in-round knowledge importance ranking coefficient in the l-round retrieval knowledge set is as follows , where represents the round of the retrieval knowledge set, represents the number of knowledge in the round of the retrieval knowledge set. If the knowledge is in the round of the retrieval knowledge set, then represents the ranking of the knowledge in ;

[0043] Normalize the knowledge influence degree coefficient and the in-round knowledge importance ranking coefficient, and obtain the comprehensive ranking value of the knowledge according to the normalized knowledge influence degree coefficient and the in-round knowledge importance ranking coefficient. The acquisition expression is as follows , where represents the comprehensive ranking value of the knowledge, and according to the comprehensive ranking value of the knowledge, sort the knowledge from large to small to obtain a knowledge ranking table for display to the user.

[0044] In a preferred embodiment, in step S4, when there is a potential risk of lag in the problem status classification model, the preset exponential decay strategy is adaptively corrected as follows:

[0045] Obtain the frequency of knowledge zs being mentioned in the retrieval knowledge set in the th round; ;

[0046] Calculate the semantic similarity between the question in the th round and the current round of question according to the similarity measurement method; ;

[0047] Use the frequency of knowledge zs being mentioned in the retrieval knowledge set in the th round and the semantic similarity between the question in the th round and the current round of question as correction factors to correct the knowledge influence degree coefficient , and its correction function expression is as follows , where in the formula , represents the corrected knowledge influence degree coefficient, respectively represent the preset proportionality coefficients of the frequency and the semantic similarity , and are both greater than 0.

[0048] Technical effects and advantages of the present invention:

[0049] In the single-round knowledge retrieval of the present invention, the status of the question is first discriminated, and then the corresponding retrieval strategy is executed according to the status to obtain the retrieval knowledge set based on the current round of question, and multiple pieces of information in the process of question status recognition by the question status classification model are obtained to construct a lag impact evaluation model, generate a lag impact evaluation index, determine whether there is a potential risk of lag in the question status classification model, timely identify potential lag risks, optimize resource usage, reduce response time, and improve the user interaction experience. When there is no potential risk of lag in the question status classification model, the single-round retrieval knowledge set is updated by multi-round knowledge rolling, and the preset exponential decay strategy is used to obtain the knowledge influence degree coefficient and combine it with the in-round knowledge importance ranking coefficient to comprehensively sort the knowledge, obtain a knowledge ranking table for display to the user, improve the accuracy and efficiency of knowledge retrieval in the multi-round dialogue scenario, realize the dynamic optimization and update of the knowledge retrieval set by accurately identifying the question status and combining historical dialogue information. When there is a potential risk of lag in the question status classification model, the preset exponential decay strategy is adaptively corrected to ensure that key information can still be accurately extracted in the real-time dialogue scenario and reduce the loss of the effectiveness of historical information. Brief Description of the Drawings

[0050] For the convenience of those skilled in the art to understand, the present invention will be further described below in conjunction with the accompanying drawings;

[0051] Figure 1 It is a schematic structural diagram of the method of the embodiment of the present invention. Specific embodiments

[0052] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the protection scope of the present invention.

[0053] Embodiment: Figure 1 A multi-level document retrieval method of the present invention is given, including the following steps:

[0054] Step S1, identify the problem state of the user's single-round question according to the problem state classification model based on the SSM convolution module, obtain the problem state, and adopt different knowledge retrieval strategies according to the problem state to obtain the single-round retrieval knowledge set corresponding to the problem state;

[0055] Step S2, obtain multiple pieces of information in the process of the problem state classification model for problem state identification, construct a carding impact assessment model, generate a carding impact assessment index, and determine whether there is a carding hidden danger in the problem state classification model;

[0056] Step S3, when there is no carding hidden danger in the problem state classification model, perform multi-round knowledge rolling update on the single-round retrieval knowledge set, adopt a preset exponential decay strategy to obtain the knowledge influence degree coefficient and combine the in-round knowledge importance ranking coefficient to comprehensively sort the knowledge, and obtain a knowledge ranking table for display to the user;

[0057] Step S4, when there is a carding hidden danger in the problem state classification model, perform adaptive correction on the preset exponential decay strategy;

[0058] Step S1, identify the problem state of the user's single-round question according to the problem state classification model based on the SSM convolution module, obtain the problem state, and adopt different knowledge retrieval strategies according to the problem state to obtain the single-round retrieval knowledge set corresponding to the problem state;

[0059] First, preprocess the text of the user's current-round question, including operations such as removing stop words and punctuation marks to standardize the question text. Then, tokenize the question text and convert each word into its corresponding word vector to obtain the semantic vector of the current-round question. Commonly used word vector generation methods include the context-based word embedding method BERT and the static word embedding methods Word2Vec or GloVe.

[0060] At the same time, preprocess and perform tokenization and word vector conversion operations on the text of the user's historical multi-round questions and the corresponding round's retrieval knowledge set, respectively obtaining the semantic vector of the historical multi-round questions and the semantic vector matrix of the retrieval knowledge set.

[0061] Construct a question text semantic vector matrix based on the semantic vector of the current-round question and the semantic vectors of the historical multi-round questions , where represents the semantic vector of the current-round question, represents the semantic vectors of the historical multi-round questions, where ={1, 2,..., };

[0062] Input the question text semantic vector matrix into the question SSM convolution module, and perform convolution on the question text semantic vector matrix according to the convolution kernel of the question SSM convolution module as follows , where represents the convolutional output feature vector of the question, which is used to represent the association between the current-round question and the historical multi-round questions. The convolution kernel consists of multiple state transition matrices ( , , , ). is the state transition matrix, which controls how the historical state is related to the current state. is the input matrix, which represents how the current question affects the state of the module. and are the output matrices, which control how the convolution kernel combines the semantic features of each question;

[0063] Input the semantic vector matrix of the retrieval knowledge set into the retrieval knowledge SSM convolution module, and perform convolution on the question text semantic vector matrix according to the convolution kernel of the retrieval knowledge SSM convolution module as follows , where represents the convolutional output feature vector of the retrieval knowledge;

[0064] Input the convolutional output feature vectors and Perform horizontal splicing to obtain the final output feature vector , and input the final output feature vector after horizontal splicing into the fully connected layer to obtain the output of the fully connected layer. The role of this layer is to perform further linear combination and feature extraction on the spliced comprehensive features. Through linear transformation, the fully connected layer converts the high-dimensional feature vector into a low-dimensional feature representation: , where represents the weight matrix of the fully connected layer, is the bias term, is the output of the fully connected layer;

[0065] Input the output of the fully connected layer into the Softmax activation function to obtain the probability value of the problem state. The obtained expression is as follows, , where represents the probability value that the question asked by the user in the current round belongs to the problem state , , represents the output of the fully connected layer;

[0066] Convert the output of the fully connected layer into a probability distribution. The role of Softmax is to convert the output into the probabilities of three states (M1, M2, M3);

[0067] The problem state represents an independent new question;

[0068] The problem state M2 represents a non-independent question that is a detail of the previous round's question;

[0069] The problem state M3 represents a non-independent question that has an intersection with the previous round's question;

[0070] Classify the problem state of the question asked by the user in the current round according to the probability value output by the Softmax activation function, as follows:

[0071] If > and > , then mark the problem state of the question asked by the user in the current round as M1, obtain the semantic vector of the question asked by the user in the current round, perform semantic similarity matching with the semantic vectors of the existing knowledge in the knowledge base, sort the existing knowledge in the knowledge base from large to small according to the semantic similarity, and select the top top_k pieces of knowledge as the single-round retrieval knowledge set;

[0072] It should be noted that for semantic similarity matching with the semantic vectors of existing knowledge in the knowledge base, common similarity measurement methods, including cosine similarity, Euclidean distance, etc., can be used to calculate the similarity between the semantic vector of the user's question and the semantic vectors of each piece of knowledge in the knowledge base, which will not be elaborated here;

[0073] If > and > , then mark the question status of the user's current round of question as M2, and directly use the retrieval knowledge set of the previous round of question as the single-round retrieval knowledge set of the current round;

[0074] If > and > , then mark the question status of the user's current round of question as M3, splice the previous round of question and the current round of question, obtain the semantic vector of the spliced text, and then perform semantic similarity matching with the semantic vectors of existing knowledge in the knowledge base. Sort the existing knowledge in the knowledge base from large to small according to semantic similarity, and select the top top_k pieces of knowledge as the single-round retrieval knowledge set;

[0075] Step S2, obtain multiple pieces of information in the process of problem status recognition by the problem status classification model, construct a lag impact evaluation model, generate a lag impact evaluation index, and determine whether there is a lag risk in the problem status classification model;

[0076] The multiple pieces of information in the problem status recognition process include the model execution time anomaly coefficient, the resource consumption status coefficient, the convolutional layer number anomaly fluctuation coefficient, and the user churn rate;

[0077] The model execution time anomaly coefficient is used to measure the deviation degree between the time required for the problem status classification model to classify the user's question each time and the expected execution time. By evaluating the execution inference time of the model in different situations, it can be judged whether its execution efficiency is stable and whether there is a potential lag risk. It mainly reflects the difference between the actual execution inference time and the expected execution inference time. When the model execution inference time significantly exceeds or is lower than the expected value, the system may experience response delays or efficiency anomalies due to reasons such as tight computing resources, too high model complexity, and sudden increase in data volume. By obtaining the model execution time anomaly coefficient, early warnings can be issued when the model execution inference time is abnormal, and further resource optimization or model adjustment measures can be taken; the higher the model execution time anomaly coefficient, the higher the probability that there is a lag risk in the problem status classification model;

[0078] The acquisition logic of the model execution time anomaly coefficient is as follows:

[0079] Obtain the execution inference time for the question status classification model to classify the question status each time a user asks a question , where g represents the number of the user's question, g = {1, 2,..., G}, and G is a positive integer;

[0080] Calculate the average value of the execution inference time , and the expression is as follows , and calculate the standard deviation of the execution inference time based on the average value of the execution inference time , and the expression is as follows ;

[0081] Compare the execution inference time of each time with the standard deviation of the execution inference time, and mark the execution inference time as abnormal. When the execution inference time is greater than or equal to the standard deviation of the execution inference time, the execution inference time greater than or equal to the standard deviation of the execution inference time is marked as an abnormal execution inference time;

[0082] Calculate the abnormal coefficient of the model execution time , and the expression is as follows , where in the formula represents the h-th abnormal execution inference time, h = {1, 2,..., H}, and H is a positive integer;

[0083] The resource consumption status coefficient is used to measure the difference and pressure between the computing resources consumed by the question status classification model during operation and the available resources. It mainly reflects whether the resources such as CPU / GPU usage rate and memory occupancy consumed by the model during inference and processing of user questions reach the tolerance limit of the system, and evaluates whether there is a risk of lag caused by resource overload in the model. When the resource consumption status coefficient is higher, it indicates that the difference between the computing resources consumed by the question status classification model during operation and the available resources is smaller and the pressure is greater, and the probability of the question status classification model having a lag risk is higher;

[0084] The acquisition logic of the resource consumption status coefficient is as follows:

[0085] Obtain the resource consumption data of the question status classification model during operation, including but not limited to CPU usage rate, GPU usage rate, memory occupancy ratio, disk I / O usage rate;

[0086] It should be noted that the resource consumption data can be obtained in real time through performance monitoring tools such as top, Prometheus, Grafana, etc., and these tools can provide detailed information about resource consumption usage;

[0087] Normalize the obtained CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate, and calculate the resource consumption status coefficient based on the normalized CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate , and the expression is as follows , where respectively represent the normalized CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate, , , , respectively represent the maximum acceptable expected values of the CPU usage rate, GPU usage rate, memory occupancy ratio, and disk I / O usage rate, respectively represent 's preset proportionality coefficients, and are all greater than 0;

[0088] It should be noted that is set according to the actual situation. For example, the expert weight assignment method is adopted, that is, experts in related fields are invited to determine the preset proportionality coefficients of each index through professional opinion surveys and comprehensive evaluations;

[0089] The convolutional layer abnormal fluctuation coefficient is used to measure the abnormal fluctuation degree of the convolutional layer when the problem state classification model performs convolutional operations. By obtaining the convolutional layer abnormal fluctuation coefficient, the impact of the abnormal fluctuation of the convolutional layer on the running performance of the problem state classification model can be identified in a timely manner, so as to provide constructive opinions for optimizing the problem state classification model. The larger the convolutional layer abnormal fluctuation coefficient, the greater the abnormal fluctuation degree of the convolutional layer when the problem state classification model performs convolutional operations, the deeper the impact on the running performance of the problem state classification model, and the higher the probability that there is a lagging hidden danger in the problem state classification model;

[0090] The acquisition logic of the convolutional layer abnormal fluctuation coefficient is as follows:

[0091] Obtain the convolutional layer dataset when the problem state classification model performs convolutional operations, and mark the convolutional layer dataset as , represents the convolutional layer value when the problem state classification model performs the r-th convolutional operation, r = {1, 2,..., R}, and R is a positive integer; extract the unique convolutional layer values from the convolutional layer dataset to obtain the unique value set , represents the unique value of the k-th convolutional layer value in the convolutional layer dataset, k = {1, 2,..., K}, and K is a positive integer; traverse the convolutional layer values in the convolutional layer dataset and count the occurrence times of each ​; Calculate the probability of unique values of the numerical values of each convolutional layer , The expression is as follows ;

[0092] Calculate the entropy value of the number of convolutional layers , The expression is as follows ;

[0093] The entropy value of the number of convolutional layers is used to measure the complexity and uncertainty of changes in the problem state classification model, reflect the fluctuation degree of the number of convolutional layers, and is used to evaluate the impact of these fluctuations on the performance of the problem state classification model. If the entropy value is very large, it indicates that there are significant abnormal fluctuations in the number of convolutional layers, which may affect the stability and performance of the model, and the probability of potential lags in the problem state classification model is higher;

[0094] Calculate the abnormal fluctuation coefficient of the number of convolutional layers , The expression is as follows , where represents the standard deviation of the numerical values of the convolutional layers in the convolutional layer number dataset, and the expression is as follows , represents the average value of the numerical values of the convolutional layers in the convolutional layer number dataset, and the expression is as follows ;

[0095] The user churn rate is used to measure the proportion of users who stop using the problem state classification model within a certain period of time. This indicator reflects the satisfaction of users with the performance and response speed of the problem state classification model, as well as the attractiveness and reliability of the problem state classification model during the user interaction process. A higher user churn rate may indicate potential lags of different degrees in the problem state classification model;

[0096] The acquisition logic of the user churn rate is as follows:

[0097] Record the total number of active users within the preset fixed time period ZT and the number of users who stop using the problem state classification model , Calculate the user churn rate , The expression is as follows ;

[0098] It should be noted that the preset fixed time period can be set by those skilled in the art according to the usage habits and interaction frequencies of users, as well as referring to historical data and the performance benchmarks of the problem state classification model, to set an expected time period;

[0099] Normalize the obtained model execution time anomaly coefficient, resource consumption status coefficient, convolutional layer number anomaly fluctuation coefficient, and user churn rate. Construct a lag impact assessment model based on the normalized model execution time anomaly coefficient, resource consumption status coefficient, convolutional layer number anomaly fluctuation coefficient, and user churn rate, and generate a lag impact assessment index. The formula on which the lag impact assessment model is based is as follows In the formula respectively represent the preset proportionality coefficients of the model execution time anomaly coefficient, resource consumption status coefficient, convolutional layer number anomaly fluctuation coefficient, and user churn rate, and are all greater than 0;

[0100] It should be noted that is set according to the actual situation. For example, the expert weighting method is adopted, that is, experts in related fields are invited to determine the preset proportionality coefficients of each index through professional opinion surveys and comprehensive evaluations;

[0101] It can be seen from the above calculation expressions that the larger the model execution time anomaly coefficient, the larger the resource consumption status coefficient, the larger the convolutional layer number anomaly fluctuation coefficient, and the larger the user churn rate, the larger the lag impact assessment index, indicating that the probability of the problem status classification model having a lag hidden danger is higher. On the contrary, the smaller the model execution time anomaly coefficient, the smaller the resource consumption status coefficient, the smaller the convolutional layer number anomaly fluctuation coefficient, and the smaller the user churn rate, the smaller the lag impact assessment index, indicating that the probability of the problem status classification model having a lag hidden danger is lower;

[0102] Compare the lag impact assessment index with the preset lag impact assessment index threshold to determine whether the problem status classification model has a lag hidden danger, specifically as follows:

[0103] If the lag impact assessment index is greater than the lag impact assessment index threshold, it indicates that the problem status classification model has a lag hidden danger and a lag hidden danger signal is generated; if the lag impact assessment index is less than or equal to the lag impact assessment index threshold, it indicates that the problem status classification model does not have an obvious lag hidden danger and a model operation stable signal is generated;

[0104] Step S3, when the problem status classification model does not have a lag hidden danger, perform multi-round knowledge rolling update on the single-round retrieval knowledge set, adopt the preset exponential decay strategy to obtain the knowledge influence degree coefficient, and combine the in-round knowledge importance ranking coefficient to comprehensively sort the knowledge to obtain a knowledge ranking table for display to the user;

[0105] Adopt the preset exponential decay strategy to obtain the knowledge influence degree coefficient, specifically as follows:

[0106] Obtain the retrieval knowledge set of the user's historical multi-round questions, sort the retrieval knowledge set of the historical multi-round questions in chronological order from far to near, and record the earliest round as , and record the current round as , , where L is a positive integer; for the round number , the knowledge influence degree coefficient of the round retrieval knowledge set participating in the multi-round knowledge rolling update The calculation formula is as follows ; ;

[0107] For the knowledge in each round of the retrieval knowledge set, no longer simply rely on the semantic similarity score, and use the in-round knowledge importance ranking coefficient based on the semantic similarity ranking for evaluation. For the knowledge , the calculation expression of its in-round knowledge importance ranking coefficient in the l-round retrieval knowledge set is as follows , where represents the round of retrieval knowledge set, represents the number of knowledge in the round of retrieval knowledge set. If the knowledge is in the round of retrieval knowledge set, then represents the ranking of the knowledge in ;

[0108] Normalize the knowledge influence degree coefficient and the in-round knowledge importance ranking coefficient, and obtain the comprehensive ranking value of the knowledge according to the normalized knowledge influence degree coefficient and the in-round knowledge importance ranking coefficient. The acquisition expression is as follows , where represents the comprehensive ranking value of the knowledge, and according to the comprehensive ranking value of the knowledge, sort the knowledge from large to small to obtain the knowledge ranking list for display to the user;

[0109] Step S4, when there is a potential lag in the question status classification model, adaptively correct the preset exponential decay strategy, specifically as follows:

[0110] Obtain the frequency of the knowledge zs being mentioned in the round of retrieval knowledge set ;

[0111] Calculate the semantic similarity between the round of question and the current round of question according to the similarity measurement method ;

[0112] It should be noted that for common similarity measurement methods, including cosine similarity, Euclidean distance, etc., in this application, the cosine similarity calculation method is used as an implementation to calculate the semantic similarity between the questions in the current round of questions , and its calculation expression is as follows , where in the formula represents the semantic vector representation of the questions in the round of questions, represents the semantic vector representation of the questions in the current round of questions represents the dot product operation of and represents the Euclidean norm of the semantic vector of the questions in the round of questions, represents the Euclidean norm of the semantic vector of the questions in the current round of questions;

[0113] The frequency of the knowledge zs mentioned in the retrieval knowledge set in the round of questions and the semantic similarity between the questions in the current round of questions are used as correction factors to correct the knowledge influence degree coefficient , and its correction function expression is as follows , where in the formula represents the corrected knowledge influence degree coefficient, respectively represent the preset proportional coefficients of the frequency , semantic similarity , and are both greater than 0;

[0114] It should be noted that before correcting according to the frequency , semantic similarity , it is necessary to normalize the frequency , semantic similarity . Common normalization methods include Min-Max normalization, Z-Score standardization, etc., which are set according to the actual situation and will not be elaborated here;

[0115] In single-round knowledge retrieval, the present invention first discriminates the status of the question, and then executes the corresponding retrieval strategy according to the status to obtain a retrieval knowledge set based on the current-round question, and obtains a plurality of pieces of information in the process of identifying the question status by the question status classification model, constructs a carding impact evaluation model, generates a carding impact evaluation index, determines whether there is a carding hidden danger in the question status classification model, timely identifies potential carding hidden dangers, optimizes resource usage, reduces the response time, and improves the user interaction experience. When there is no carding hidden danger in the question status classification model, multi-round knowledge rolling update is performed on the single-round retrieval knowledge set, and a preset exponential decay strategy is used to obtain the knowledge influence degree coefficient and combine it with the in-round knowledge importance ranking coefficient to comprehensively sort the knowledge, obtain a knowledge ranking table for display to the user, improve the accuracy and efficiency of knowledge retrieval in multi-round dialogue scenarios, dynamically optimize and update the knowledge retrieval set by accurately identifying the question status and combining historical dialogue information. When there is a carding hidden danger in the question status classification model, the preset exponential decay strategy is adaptively corrected to ensure that key information can still be accurately extracted in real-time dialogue scenarios and reduce the loss of the effectiveness of historical information;

[0116] The above formulas are all dimensionless and take their numerical values for calculation. The formula is obtained by collecting a large amount of data for software simulation to get a formula closest to the actual situation. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0117] It should be understood that in various embodiments of the present application, the magnitudes of the serial numbers of the above processes do not mean the order of execution, and the order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0118] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A multi-level document retrieval method, characterized in that: The steps include: Step S1, identifying the question state of a single-round question of a user according to the question state classification model based on the SSM convolution module, obtaining the question state, and adopting different knowledge retrieval strategies according to the question state to obtain the single-round retrieval knowledge set corresponding to the question state; Step S2, obtaining multiple information of the problem state classification model for the problem state identification process, constructing a jamming impact assessment model, generating a jamming impact assessment index, and determining whether the problem state classification model has a jamming risk; Step S3: when the problem status classification model does not have the risk of freezing, multiple rounds of knowledge rolling updates are performed on the single-round retrieval knowledge set, a preset exponential decay strategy is used to obtain the knowledge influence coefficient, and the knowledge is comprehensively ranked in combination with the knowledge importance ranking coefficient within the round to obtain a knowledge ranking table for display to the user; Step S4, when the problem state classification model has a hidden danger of freezing, adaptively correct the preset exponential decay strategy; In step S3, a preset exponential decay strategy is used to obtain the knowledge influence coefficient, as follows: Get the retrieval knowledge set of the user's historical multi-round questions, sort the retrieval knowledge set of historical multi-round questions from far to near according to the chronological order, the earliest round is l = 1, the current round is l = L, l = {1, 2, ..., L}, L is a positive integer; for the round number l, the knowledge influence coefficient YX of the L-round retrieval knowledge set participating in the multi-round knowledge rolling update l The calculation formula is as follows The multiple pieces of information in the problem state identification process include model execution time abnormality coefficient, resource consumption state coefficient, convolution layer abnormal fluctuation coefficient, and user churn rate; The logic for obtaining the model execution time anomaly coefficient is as follows: Get the execution reasoning time t of the question status classification model for question status classification each time a user asks a question g , g represents the number of times the user asks a question, g = {1, 2, ..., G}, G is a positive integer; Calculate the average execution inference time tp, the expression is as follows The standard deviation of the execution reasoning time tb is calculated based on the average execution reasoning time. The expression is as follows Compare each execution reasoning time with the execution reasoning time standard deviation, and mark the execution reasoning time as abnormal. When the execution reasoning time is greater than or equal to the execution reasoning time standard deviation, the execution reasoning time greater than or equal to the execution reasoning time standard deviation is marked as abnormal execution reasoning time. Calculate the model execution time anomaly coefficient mxz, the expression is as follows In the formula, ty h represents the execution reasoning time of the hth exception, h={1,2,...,H}, H is a positive integer; The logic for obtaining the resource consumption status coefficient is as follows: Obtain resource consumption data of the problem status classification model during operation, including but not limited to CPU usage, GPU usage, memory usage, and disk I / O usage; The obtained CPU usage, GPU usage, memory usage, and disk I / O usage need to be normalized, and the resource consumption state coefficient zyx is calculated based on the normalized CPU usage, GPU usage, memory usage, and disk I / O usage. The expression is as follows: Where cp, gp, nc, and io represent the normalized CPU usage, GPU usage, memory usage, and disk I / O usage, respectively. z 、gp z 、nc z ,io z They represent the maximum acceptable expected values ​​of CPU usage, GPU usage, memory usage, and disk I / O usage, respectively. a1, a2, a3, and a4 represent The preset proportional coefficient, and a1, a2, a3, a4 are all greater than 0; The logic for obtaining the abnormal fluctuation coefficient of the number of convolution layers is as follows: Get the convolution layer data set of the problem state classification model when performing convolution operation, and mark the convolution layer data set as JC = {cs r },cs r Represents the convolution layer value when the problem state classification model performs the convolution operation for the rth time, r = {1, 2, ..., R}, R is a positive integer; extract the unique convolution layer value from the convolution layer data set to obtain the unique value set WY = {wy k },wy k Represents the unique value of the kth convolutional layer value in the convolutional layer data set, k = {1, 2, ..., K}, K is a positive integer; traverse the convolutional layer values ​​in the convolutional layer data set, and count each wy k The number of occurrences f(wy k );Calculate the probability P(wy) of a unique value for each convolutional layer value k ), the expression is as follows Calculate the entropy value HX of the convolution layer, the expression is as follows Calculate the abnormal fluctuation coefficient jiv of the number of convolution layers, expressed as follows Where jcb represents the standard deviation of the convolutional layer values ​​in the convolutional layer data set, and the expression is as follows jcj represents the average value of the convolution layer in the convolution layer data set, and the expression is as follows The logic for obtaining the user churn rate is as follows: Record the total number of active users yzs and the number of users who stop using the problem status classification model lsy within the preset fixed time period ZT, and calculate the user churn rate yhl, which is expressed as follows The obtained model execution time abnormality coefficient, resource consumption state coefficient, convolution layer abnormal fluctuation coefficient, and user churn rate are normalized, and a jam impact assessment model is constructed based on the normalized model execution time abnormality coefficient, resource consumption state coefficient, convolution layer abnormal fluctuation coefficient, and user churn rate to generate a jam impact assessment index kdzs. The jam impact assessment model is based on the following formula: Where mxz represents the model execution time abnormality coefficient, zyx represents the resource consumption state coefficient, jiv represents the abnormal fluctuation coefficient of the number of convolution layers, yhl represents the user churn rate, α, β, γ, δ represent the preset proportional coefficients of the model execution time abnormality coefficient, resource consumption state coefficient, abnormal fluctuation coefficient of the number of convolution layers, and user churn rate, respectively, and α, β, γ, δ are all greater than 0; Compare the jam impact assessment index with the preset jam impact assessment index threshold to determine whether the problem status classification model has jam risks, as follows: If the jamming impact assessment index is greater than the jamming impact assessment index threshold, a jamming potential risk signal is generated; if the jamming impact assessment index is less than or equal to the jamming impact assessment index threshold, a model operation stability signal is generated.

2. A multi-level document retrieval method according to claim 1, characterized in that: In step S1, the text of the user's current round of questions is first preprocessed, the question text is segmented, and each word is converted into a corresponding word vector to obtain the semantic vector of the current round of questions; at the same time, the text of the user's historical multiple rounds of questions and the corresponding round of retrieval knowledge set are preprocessed and segmented to convert into word vectors, respectively obtaining the semantic vectors of the historical multiple rounds of questions and the semantic vector matrix of the retrieval knowledge set; Based on the semantic vector of the current round of questions and the semantic vectors of multiple rounds of questions in the past, a question text semantic vector matrix is ​​constructed, and the question text semantic vector matrix is ​​input into the question SSM convolution module, and the question text semantic vector matrix is ​​convolved according to the convolution kernel of the question SSM convolution module; The semantic vector matrix of the retrieval knowledge set is input into the retrieval knowledge SSM convolution module, and the semantic vector matrix of the question text is convolved according to the convolution kernel of the retrieval knowledge SSM convolution module; The convolution output feature vectors extracted from the question SSM convolution module and the retrieval knowledge SSM convolution module are horizontally spliced ​​to obtain the final output feature vector, and the final output feature vector after horizontal splicing is passed to the fully connected layer to obtain the fully connected layer output; The output of the fully connected layer is passed into the Softmax activation function to obtain the probability value of the problem state. The expression is as follows: Where P_M i Indicates that the question asked by the user in the current round belongs to question state M i The probability value, M i ∈{M1, M2, M3}, Z represents the output of the fully connected layer; Problem state M1 represents an independent new problem; Problem state M2 means it is not independent and is the details of the previous round of problems; Problem state M3 means it is not independent and has intersection with the previous round of problems; According to the probability value output by the Softmax activation function, the status of the question asked by the user in the current round is classified as follows: If P_M1>P_M2 and P_M1>P_M3, then mark the question status of the user's current question as M1, obtain the semantic vector of the user's current question, and match it with the semantic vector of the existing knowledge in the knowledge base for semantic similarity. Sort the existing knowledge in the knowledge base from large to small according to the semantic similarity, and select the top_k pieces of knowledge as the single-round retrieval knowledge set; If P_M2>P_M1 and P_M2>P_M3, then the status of the question asked by the user in the current round is marked as M2, and the retrieval knowledge set of the previous round of questions is directly used as the single-round retrieval knowledge set of the current round; If P_M3>P_M1 and P_M3>P_M2, the question status of the user's current round of questions is marked as M3, the previous round of questions are concatenated with the current round of questions, the semantic vector of the concatenated text is obtained, and then the semantic similarity is matched with the semantic vector of the existing knowledge in the knowledge base. The existing knowledge in the knowledge base is sorted from large to small according to the semantic similarity, and the top_k pieces of knowledge are selected as the single-round retrieval knowledge set.

3. A multi-level document retrieval method according to claim 1, characterized in that: For the knowledge in each round of retrieval knowledge set, we no longer rely solely on the semantic similarity score, but use the intra-round knowledge importance ranking coefficient based on semantic similarity ranking for evaluation. For knowledge zs, its intra-round knowledge importance ranking coefficient in the retrieval knowledge set of round l is calculated as follows: Epoch l Represents the retrieved knowledge set of round l, Count(Epoch l ) represents the number of knowledge in the retrieval knowledge set in round l. If knowledge zs is in the retrieval knowledge set in round l, then Rank(zs,l) represents the rank of knowledge zs in Epoch l Medium ranking; The knowledge influence coefficient and the knowledge importance ranking coefficient within the round are normalized, and the comprehensive ranking value of knowledge is obtained according to the normalized knowledge influence coefficient and the knowledge importance ranking coefficient within the round. The expression for obtaining the value is as follows: Scor represents the comprehensive ranking value of knowledge, and according to the comprehensive ranking value of knowledge, the knowledge is sorted from large to small to obtain a knowledge ranking table for display to users.

4. A multi-level document retrieval method according to claim 1, characterized in that: In step S4, when the problem state classification model has a hidden danger of freezing, the preset exponential decay strategy is adaptively modified as follows: The frequency PV of the knowledge zs being mentioned in the retrieval knowledge set in round l l ; Calculate the semantic similarity sim(l,L) between the question in the first round and the question in the current round L according to the similarity measurement method; The frequency PV of knowledge zs being mentioned in the retrieval knowledge set in round l l And the semantic similarity sim(l,L) between the first round of questions and the current round of questions L is used as a correction factor to the knowledge influence coefficient YX l The correction function expression is as follows In the formula represents the modified knowledge influence coefficient, ω represents the frequency PV l , the preset proportionality coefficient of the semantic similarity sim(l,L), and ω is greater than 0.

Citation Information

Patent Citations

  • Machine reading understanding method for guiding attention based on knowledge

    CN111241807A

  • Service knowledge base management system and method based on knowledge big model

    CN118917390A