Intelligent question and answer method and device, equipment and medium
By obtaining multimodal data, calculating semantic similarity, conducting context analysis and generating new problem sequences in an intelligent question-and-answer system, the problem of inaccurate understanding deviations and inaccurate responses of existing systems in complex semantic and diversified problems is solved, and higher question-and-answer accuracy and user experience are achieved.
Patent Information
- Application Number
- CN202510496016.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-21
AI Technical Summary
Existing intelligent question-answer systems are difficult to accurately understand user intentions and generate reasonable answers when facing complex semantics and diversified problems, especially when user inputs have ambiguity, ambiguity or multi-level semantics.
By obtaining multimodal data, the semantic similarity between them and candidate answers is calculated, the initial candidate answer is filtered out, and context analysis is used to generate intent prediction results. Then, a semantic generation model is used to generate a new question sequence similar to the semantics of the user's questions, an extended question and answer library is built, and closed-loop optimization is performed based on user feedback data to generate the optimized question and answer results.
It significantly improves the accuracy of question-and-answer matching, enhances the diversity and comprehensiveness of the knowledge base, improves the system's ability to cover user questions, and provides more accurate and in line with user needs.
Smart Images

Figure CN120011484A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of intelligent question answering technology, and in particular to an intelligent question answering method, device, equipment and medium. Background Art
[0002] With the development of artificial intelligence technology, intelligent question-answering systems have been widely used in many fields, such as customer service, education, and medical care. However, existing intelligent question-answering systems often face problems of understanding bias and inaccurate responses when faced with complex semantics and diverse questions. Especially when user input is ambiguous, ambiguous, or has multi-level semantics, it is difficult for the system to accurately predict user intent and generate reasonable answers.
[0003] Currently, the main methods rely on rule-based matching methods or simple keyword search methods. These methods are prone to problems such as insufficient semantic matching and misunderstanding of intent when dealing with semantically rich and context-dependent conversations, which affects the accuracy of the system's response and the user's interactive experience. Therefore, how to improve the semantic understanding and answer generation effect of the intelligent question-answering system through more accurate semantic similarity matching and intent prediction mechanisms, and then optimize the overall performance of the system and user experience, has become an important topic in intelligent question-answering technology. Summary of the invention
[0004] In view of this, the purpose of the present invention is to provide an intelligent question-answering method, device, equipment and medium, which can greatly improve the accuracy of question-answer matching and optimize the user query experience. The specific scheme is as follows:
[0005] In a first aspect, the present application discloses an intelligent question-answering method, which is applied to a question-answering system, comprising:
[0006] Obtaining a user question containing multimodal data and a candidate answer set corresponding to the user question;
[0007] Determining a first semantic similarity between the user question and the candidate answer set, and determining an initial candidate answer from the candidate answer set according to the first semantic similarity;
[0008] Based on the initial candidate answer, historical conversation information in the process of interaction with the user is obtained, and the historical conversation information is contextually analyzed using a sliding window to generate an intention prediction result for the user's question;
[0009] According to the intention prediction result, a new question sequence with similar semantics to the user's question is generated using a semantic generation model, and an extended question and answer library is constructed based on the new question sequence;
[0010] Obtain user feedback data, and based on the user feedback data and the data in the extended question and answer library, use a closed-loop optimization model to generate a question and answer result that optimizes the initial candidate answer, and push the question and answer result to the user; the user feedback data is the scoring data of the user's generated answer feedback for different questions.
[0011] Optionally, the multimodal data includes any one or a combination of text data, image data, voice data, expression data, and human posture data; wherein, after obtaining the user question containing the multimodal data and the candidate answer set corresponding to the user question, the method further includes:
[0012] The user question and the candidate answer set are respectively format converted and feature extracted to generate corresponding cross-modal semantic embeddings.
[0013] Optionally, format conversion and feature extraction are performed on the user question to generate corresponding cross-modal semantic embedding, including:
[0014] Performing word vector conversion on the text data to obtain a first feature vector, and inputting the first feature vector into a pre-trained word embedding model to obtain a first feature representation of the text data;
[0015] Extracting a second feature vector of the image data using a preset convolutional neural network, and determining a second feature representation of the image data based on the second feature vector;
[0016] Processing the speech data using a preset acoustic feature extraction method to obtain a third feature vector, and determining a third feature representation of the speech data based on the third feature vector;
[0017] Extracting a fourth eigenvector of the expression data using a preset facial expression analysis tool, and determining a fourth feature representation of the expression data based on the fourth eigenvector;
[0018] Extracting a fifth eigenvector of the human body posture data using a preset human body posture estimation model, and determining a fifth feature representation of the human body posture data based on the fifth eigenvector;
[0019] The first feature representation, the second feature representation, the third feature representation, the fourth feature representation and the fifth feature representation are aligned, and the aligned first feature representation, the second feature representation, the third feature representation, the fourth feature representation and the fifth feature representation are respectively mapped to a shared semantic space using a preset multimodal embedding model to obtain a unified cross-modal semantic embedding.
[0020] Optionally, obtaining a candidate answer set corresponding to the user's question includes:
[0021] Acquire a first candidate answer corresponding to the user's question from the original question-and-answer database, and construct a candidate answer set based on the first candidate answer;
[0022] Correspondingly, after constructing the extended question-answer library based on the new question sequence, the method further includes:
[0023] A second candidate answer corresponding to the new question sequence is obtained from the extended question and answer library, and the second candidate answer is added to the candidate answer set.
[0024] Optionally, after constructing the extended question-answer library based on the new question sequence, the method further includes:
[0025] Checking the second semantic similarity between the new question sequence generated in the extended question and answer library and the existing question sequence in the original question and answer library, and reviewing the answers to the questions in the new question sequence;
[0026] When the second semantic similarity is greater than or equal to a preset similarity threshold, and / or the review result of the answer to the question does not meet the preset answer standard, a target question-answer pair is determined, and quality control is performed on the target question-answer pair.
[0027] Optionally, the step of generating a question-and-answer result after optimizing the initial candidate answer by using a closed-loop optimization model based on the user feedback data and the data in the extended question-and-answer library includes:
[0028] Determine a quality score for each question-answer pair in the extended question-answer library based on the user feedback data and the data in the extended question-answer library;
[0029] Determining optimization parameters of the expanded question and answer library using the quality score, and inputting the optimization parameters and the user feedback data into a closed-loop optimization model to generate a question and answer result after optimizing the initial candidate answer;
[0030] The calculation formula of the quality score is: ; The calculation formula of the optimization parameter is: ; is the quality score of the jth question-answer pair, n is the total number of question-answer pairs in the extended question-answer database, Score the user's feedback on the i-th question, is the user's i-th question and the jth question in the new question sequence The semantic similarity between is the adjustment coefficient, and m is the number of questions in the question sequence under the current user intention category.
[0031] Optionally, after using the closed-loop optimization model to generate the question-answering result after optimizing the initial candidate answer, the method further includes:
[0032] Generate a corresponding voice broadcast according to the question and answer result, and use the voice-driven face model to drive the virtual interactive agent to perform lip synchronization and facial movements according to the voice broadcast;
[0033] According to the intention prediction result, the posture synthesis model is used to drive the virtual interaction agent to generate a coherent body movement sequence.
[0034] In a second aspect, the present application discloses an intelligent question-answering device, which is applied to a question-answering system, comprising:
[0035] A data acquisition module, used to acquire user questions containing multimodal data and a candidate answer set corresponding to the user questions;
[0036] A semantic matching module, configured to determine a first semantic similarity between the user question and the candidate answer set, and determine an initial candidate answer from the candidate answers according to the first semantic similarity;
[0037] An intention prediction module, configured to obtain historical conversation information during the interaction with the user based on the initial candidate answer, and perform context analysis on the historical conversation information using a sliding window to generate an intention prediction result for the user's question;
[0038] An extended question-answer library construction module is used to generate a new question sequence with similar semantics to the user's question based on the intention prediction result using a semantic generation model, and to construct an extended question-answer library based on the new question sequence;
[0039] The question and answer result generation module is used to obtain user feedback data, and based on the user feedback data and the data in the extended question and answer library, use a closed-loop optimization model to generate a question and answer result that optimizes the initial candidate answer, and push the question and answer result to the user; the user feedback data is the scoring data of the user's generated answer feedback for different questions.
[0040] In a third aspect, the present application discloses an electronic device, comprising a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the intelligent question and answer method as described above.
[0041] In a fourth aspect, the present application discloses a computer-readable storage medium for storing a computer program; wherein the computer program, when executed by a processor, implements the intelligent question-answering method as described above.
[0042] The present application provides an intelligent question-answering method, which is applied to a question-answering system, including: obtaining a user question containing multimodal data and a candidate answer set corresponding to the user question; determining a first semantic similarity between the user question and the candidate answer set, and determining an initial candidate answer from the candidate answer set according to the first semantic similarity; based on the initial candidate answer, obtaining historical dialogue information in the process of interaction with the user, and using a sliding window to perform context analysis on the historical dialogue information to generate an intention prediction result for the user question; based on the intention prediction result, using a semantic generation model to generate a new question sequence that is semantically similar to the user question, and building an extended question-answering library based on the new question sequence; obtaining user feedback data, and based on the user feedback data and the data in the extended question-answering library, using a closed-loop optimization model to generate a question-answering result optimized for the initial candidate answer, and pushing the question-answering result to the user; the user feedback data is scoring data of the user's feedback on generated answers to different questions.
[0043] The beneficial technical effects of the present application are as follows: for user questions in multimodal data, the initial candidate answers are first screened out by calculating the semantic similarity between them and the candidate answers, thereby improving the accuracy of question-answer matching; secondly, a sliding window prediction mechanism is introduced to generate intention prediction results for user questions by predicting user needs, and the knowledge base of the question-answering system is enriched through semantically similar question generation technology to obtain an expanded question-answering base, which significantly enhances the diversity and comprehensiveness of the knowledge base and improves the system's coverage of user questions; finally, based on user feedback and the expanded question-answering base, more accurate answers that meet user needs are provided through closed-loop optimization.
[0044] In addition, the present application provides an intelligent question-and-answer device, equipment, and storage medium, which correspond to the above-mentioned intelligent question-and-answer method and have the same effect as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying creative work.
[0046] Figure 1 A flow chart of an intelligent question-answering method disclosed in this application;
[0047] Figure 2 A schematic diagram of the structure of a question-answering system disclosed in this application;
[0048] Figure 3 This is a schematic diagram of the structure of an intelligent question-answering device disclosed in this application;
[0049] Figure 4 This is a structural diagram of an electronic device disclosed in this application. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0051] Currently, intelligent question-answering systems often face problems of understanding bias and inaccurate responses when faced with complex semantics and diverse questions. Especially when user input is vague, ambiguous, or has multi-level semantics, it is difficult for the system to accurately predict user intentions and generate reasonable answers.
[0052] To this end, this application provides an intelligent question-and-answer solution that can accurately and effectively predict user intentions, greatly improve the accuracy of question-and-answer matching, and enhance the level of interactive intelligence and user experience.
[0053] The embodiment of the present invention discloses an intelligent question-answering method, see Figure 1 As shown, applied to a question answering system, the method includes:
[0054] Step S11: obtaining a user question containing multimodal data and a candidate answer set corresponding to the user question.
[0055] In order to meet the needs of real-time interaction, the intelligent system generates and feeds back user requests based on the streaming output mechanism and the input of feedback parameters. In the embodiment of the present application, the multimodal data input by the user when asking questions is first obtained. After receiving the user request, the system will immediately perform an initial judgment of intent and semantic analysis without waiting for all processing to be completed. At this time, the system will give a set of candidate answers corresponding to the user's question.
[0056] In a specific implementation, the multimodal data of the user's question may include any one or a combination of text data, image data, voice data, expression data, and human posture data input by the user when asking the question. By integrating information of multiple modalities such as text, image, voice, expression, and posture, the synchronous processing of multimodal data is achieved, and the accuracy of question-answer matching is improved.
[0057] In another specific implementation, the candidate answer set is a set including at least one candidate answer generated for a user's question. The sources of the candidate answer set include the following two categories: one is a first candidate answer corresponding to the user's question that is initially screened from the original question and answer library of the intelligent system; the other is a second candidate answer corresponding to the generated new question sequence that is obtained from the expanded question and answer library based on a streaming output mechanism.
[0058] It should be pointed out that the extended question and answer library is a question and answer knowledge base constructed based on user questions and used to store a series of different questions with semantics similar to the user's questions. The question is generated based on the inferred user intention. For a user question, a new question sequence corresponding to it and containing at least one question can be determined in the extended question and answer library. The grammatical structure of these questions may be different, but the semantics should be consistent with the user's original question. Exemplarily, assuming that the user asks "How to program?" Then, the new question sequence can include a series of questions such as "How to learn programming?" and "Where should programming start?". When the extended question and answer library has been constructed, and there is a new question sequence corresponding to the user's question in the extended question and answer library, the system can match one or more candidate answers for each generated question in the new question sequence, obtain a second candidate answer, and add it to the original candidate answer set constructed by the first candidate answer to form the current candidate answer set.
[0059] Step S12: determining a first semantic similarity between the user question and the candidate answer set, and determining an initial candidate answer from the candidate answer set according to the first semantic similarity.
[0060] In an embodiment of the present application, the semantic similarity between the user's question and the candidate answer set is calculated, and the most relevant answer is selected from the candidate answer set based on the semantic similarity as the output of this round of question and answer.
[0061] Specifically, by using cosine similarity Calculate the similarity between the user's question and the candidate answer set. The formula is: ; where Q is an embedding vector in M that represents the user's question, , is a candidate answer set, where each The semantic embedding vector representing the candidate answer, , d is the vector dimension; represents the dot product of the question vector and the candidate answer vector, q represents the semantic embedding vector of the user question Q obtained after cross-modal embedding, ; and Represents the modulus of the question vector and the candidate answer vector respectively. The similarity value reflects the semantic similarity between the user's question and each candidate answer.
[0062] In this embodiment of the present application, the semantic similarity between all candidate answers in the candidate answer set and the user's question is calculated. , select the candidate answer with the maximum similarity ,Right now As the initial candidate answer determined in this round of question and answer.
[0063] Step S13: Based on the initial candidate answer, historical dialogue information in the process of interaction with the user is obtained, and a sliding window is used to perform context analysis on the historical dialogue information to generate an intention prediction result for the user's question.
[0064] In order to better understand the user's intention, in the embodiment of the present application, the system adopts a three-round conversation sliding window mechanism, combining the historical context information of the user and the system during the interaction process to construct a semantically enhanced representation. Specifically, when the user asks questions in each round, the system automatically captures the content of the last three rounds of question-and-answer interactions between the user and the system (including user questions and system replies) to form a sliding window context sequence.
[0065] Among them, the context It is the historical conversation information during the interaction between the user and the system. The sliding window size is w, which is used to analyze the user's context at the current moment. Each context unit Corresponding to an embedding vector .
[0066] When using a sliding window to process the context, assume that the context window at the current time t is , which contains the user's conversation history in the past w time steps. The window size w=3 means that the most recent three rounds of conversation are used as the current context. For each sliding window Compute its semantic representation , that is, merging the vector representations of all context units in the current window into a new semantic embedding representation: ; This formula represents the window The average vector representation of , utilizes the context information within the window to generate the global semantic embedding of the window.
[0067] Furthermore, the semantic vector q of the user's question and the current sliding window The average semantic vector representation of , calculate the prediction result of user intention. Among them, For sliding window The average value of each context vector in is used to enhance the context information of the question semantics.
[0068] In the embodiment of the present application, when generating the intention prediction result for the user's question, an intention prediction model is used. , the model combines the question and context information to predict the user's intention. The input of the intention prediction model is the semantic representation of the question and the semantic representation of the context window, and the output is the predicted user intention label I. The prediction process can be described by the following formula: ;in, It is an intention prediction model, which calculates the user's intention based on the semantic vector of the question and context; the intention prediction result I is a classification result, which represents the most likely intent category determined by the system based on the user's question and context. It plays a "guiding control role" in the subsequent semantic generation model, used to guide the generation model to generate a sequence of questions related to the semantics of the intent.
[0069] Step S14: according to the intention prediction result, a new question sequence with similar semantics to the user's question is generated using a semantic generation model, and an extended question and answer library is constructed based on the new question sequence.
[0070] In the embodiment of the present application, based on the intention prediction result I obtained in the above steps, the category of the user's intention can be inferred. For example, suppose the user asks "How to program?" and the intention prediction result is (such as "programming tutorial"), we can then determine that we need to generate a series of variant questions with similar semantics for this intent.
[0071] Specifically, the intention prediction result I is used as the input of the process of generating a new question sequence, and the input also includes the user's original question text On this basis, using the semantic generation model Generate a set of semantically similar question sequences The mathematical model of the generation process can be expressed as: ;in, is the original question entered by the user when asking a question, That is, the intention prediction result obtained from the intention prediction model. Generate models for semantics.
[0072] Furthermore, the generated new question sequence contains multiple questions After generation, it will be stored in the extended question and answer library That is, the extended question-answering database is constructed based on the new question sequences generated in response to user questions. It should be noted that the extended question-answering database not only stores these question sequences, but also includes the corresponding standard answers. , that is, the standard answer to each generated question, ensuring that the new question can be effectively matched with the data in the original question and answer library.
[0073] This step can be completed as follows: ;in, Indicates the expansion of the question-answer database. Represents the generated new questions and the corresponding standard answers In this way, the generated new question sequences and their answer pairs can be inserted into the extended question-answer library to further enrich the content of the database.
[0074] Step S15: Obtain user feedback data, and based on the user feedback data and the data in the extended question and answer library, use a closed-loop optimization model to generate a question and answer result that optimizes the initial candidate answer, and push the question and answer result to the user; the user feedback data is the scoring data of the user's generated answer feedback for different questions.
[0075] In the embodiment of the present application, the extended question and answer database and the user feedback data are analyzed. Provides scoring data for user-generated answer feedback for different questions, where: Ask a question for the i-th user, The user's feedback score for the i-th user's question. The user feedback data provides a quality evaluation of the questions and answers generated during the interaction with the user, while the extended question-answer database provides relevant semantic matching and answer information. The data in the extended question-answer database is a question-answer pair consisting of a new question sequence and its corresponding standard answer. ,in, To generate new questions, is the standard answer corresponding to the new question.
[0076] It should be pointed out that by analyzing the extended Q&A database and user feedback data, the quality score of each question-answer pair in the extended Q&A database can be obtained, and then the optimization parameters of the extended Q&A database can be determined using the quality score. The optimization parameter is an optimization adjustment factor generated based on the user feedback data and the low-scoring questions in the extended Q&A database. Furthermore, by analyzing the user feedback data and the optimization parameters of the extended Q&A database through a closed-loop optimization model, the various parts of the Q&A system are comprehensively adjusted to generate optimized Q&A results and improve the overall Q&A matching effect. Ultimately, the optimized Q&A system can better match user needs and generate more accurate answers.
[0077] The present application provides an intelligent question-answering method, which is applied to a question-answering system, including: obtaining a user question containing multimodal data and a candidate answer set corresponding to the user question; determining a first semantic similarity between the user question and the candidate answer set, and determining an initial candidate answer from the candidate answer set according to the first semantic similarity; based on the initial candidate answer, obtaining historical dialogue information in the process of interaction with the user, and using a sliding window to perform context analysis on the historical dialogue information to generate an intention prediction result for the user question; based on the intention prediction result, using a semantic generation model to generate a new question sequence that is semantically similar to the user question, and building an extended question-answering library based on the new question sequence; obtaining user feedback data, and based on the user feedback data and the data in the extended question-answering library, using a closed-loop optimization model to generate a question-answering result optimized for the initial candidate answer, and pushing the question-answering result to the user; the user feedback data is scoring data of the user's feedback on generated answers to different questions.
[0078] The beneficial technical effects of the present application are as follows: for user questions in multimodal data, the initial candidate answers are first screened out by calculating the semantic similarity between them and the candidate answers, thereby improving the accuracy of question-answer matching; secondly, a sliding window prediction mechanism is introduced to generate intention prediction results for user questions by predicting user needs, and the knowledge base of the question-answering system is enriched through semantically similar question generation technology to obtain an expanded question-answering base, which significantly enhances the diversity and comprehensiveness of the knowledge base and improves the system's coverage of user questions; finally, based on user feedback and the expanded question-answering base, more accurate answers that meet user needs are provided through closed-loop optimization.
[0079] Based on the above embodiments, in a feasible implementation, in order to achieve effective fusion of multimodal data, after obtaining the user question and the candidate answer set, the embodiment of the present application performs format conversion and feature extraction on the user question and the candidate answer set respectively to generate their corresponding cross-modal semantic embeddings. The two are two parallel processes that are processed separately but have the same structure. Specifically, the format conversion and feature extraction of the user question to generate the corresponding cross-modal semantic embedding include the following steps:
[0080] Step 1: Perform word vector conversion on the text data to obtain a first feature vector, and input the first feature vector into a pre-trained word embedding model to obtain a first feature representation of the text data.
[0081] In the embodiment of the present application, the input text data is assumed to be T. When format conversion and feature extraction are performed on the text data T, each text unit is firstly converted into a word vector to obtain a feature vector of the text. ; where d represents the dimension of the word vector. Further, the feature representation of the text data is obtained through the trained word embedding model , where each element represents the feature representation corresponding to a single text unit.
[0082] Step 2: Use a preset convolutional neural network to extract a second feature vector of the image data, and determine a second feature representation of the image data based on the second feature vector.
[0083] In the embodiment of the present application, the input image data is assumed to be I. When performing format conversion and feature extraction on the image data I, a feature vector of the image is first extracted using a convolutional neural network (CNN) for each image. , and further obtain the feature representation of image data , where each element represents the feature representation corresponding to a single image.
[0084] Step 3: Process the speech data using a preset acoustic feature extraction method to obtain a third feature vector, and determine a third feature representation of the speech data based on the third feature vector.
[0085] In the embodiment of the present application, the input voice data is assumed to be S. When the voice data S is format converted and feature extracted, an acoustic feature extraction method, such as MFCC (Mel-frequency cepstral coefficients), is used to process each voice signal to obtain a feature vector of the voice. , and further obtain the feature representation of speech data , where each element represents the feature representation corresponding to a single speech signal.
[0086] Step 4: extracting a fourth eigenvector of the expression data using a preset facial expression analysis tool, and determining a fourth feature representation of the expression data based on the fourth eigenvector.
[0087] In the embodiment of the present application, the input expression data is assumed to be E. When performing format conversion and feature extraction on the expression data E, firstly, a facial expression analysis tool such as a FACS (Facial Action Coding System) method is used to extract a feature vector of the expression for each expression frame. , and further obtain the feature representation of expression data , where each element represents the feature representation corresponding to a facial expression frame.
[0088] Step 5: Use a preset human posture estimation model to extract the fifth eigenvector of the human posture data, and determine the fifth feature representation of the human posture data based on the fifth eigenvector.
[0089] In the embodiment of the present application, the input human posture data is assumed to be P. When performing format conversion and feature extraction on the human posture data P, firstly, a human posture estimation model such as OpenPose (open source human posture estimation model) is used to extract the characteristic vector of the posture for each human posture input to obtain the characteristic representation of the posture data. , and further obtain the feature representation of human posture data , where each element represents the feature representation corresponding to a pose skeleton key point graph.
[0090] Step 6: Align the first feature representation, the second feature representation, the third feature representation, the fourth feature representation and the fifth feature representation, and use a preset multimodal embedding model to map the aligned first feature representation, the second feature representation, the third feature representation, the fourth feature representation and the fifth feature representation to a shared semantic space to obtain a unified cross-modal semantic embedding.
[0091] In the embodiment of the present application, after obtaining the feature representations corresponding to each of the multimodal data, the corresponding feature representations are aligned, and then a unified cross-modal semantic embedding is generated through a cross-modal learning method. ;in Embedding vector representing each modality data.
[0092] Specifically, when the aligned feature representations are mapped to the same vector space through the cross-modal learning method, the multimodal embedding model is used to map T, I, S, E, and P, calculate the representation of each modality data in the shared semantic space, and obtain a unified cross-modal embedding matrix M. The multimodal embedding model is a mapping function Completed, Ability to map data of different modalities into a shared semantic space, .
[0093] It is worth noting that the candidate answer set also needs to be feature extracted based on its corresponding multimodal data and embedded through the same embedding function The semantic vector representation is obtained. In this way, it can be ensured that the user's question and the answer in the candidate answer set are in the same semantic space.
[0094] Based on the above embodiment, this embodiment will specifically explain steps S14 and S15 in the above embodiment. In order to ensure the quality of the extended question and answer library and ensure that the user can obtain high-quality and accurate answers, after the extended question and answer library is constructed based on the new question sequence, the following steps may also be included:
[0095] Step 1: Check the second semantic similarity between the new question sequence generated in the extended question and answer library and the existing question sequence in the original question and answer library, and review the answers to the questions in the new question sequence;
[0096] Step 2: When the second semantic similarity is greater than or equal to a preset similarity threshold, and / or the review result of the answer to the question does not meet the preset answer standard, determine the target question-answer pair and perform quality control on the target question-answer pair.
[0097] In the embodiment of the present application, in order to ensure the data quality in the extended question and answer library, a quality control module is set in the intelligent system, and the quality control module is used to verify the question and answer pairs corresponding to the generated new question sequence. This step is performed in the following manner:
[0098] Check new question sequences using semantic similarity metrics such as cosine similarity or similarity calculations from deep learning models The original question sequence in the question-answering database The second semantic similarity between to ensure that the generated questions do not have too much duplication.
[0099] The answer to the generated The audit is conducted to ensure that the answers in the expanded Q&A database are highly accurate and semantically consistent. It is understandable that although the standard answer is automatically generated by the system or automatically matched from the original Q&A database, there are still potential problems such as semantic drift and insufficient ambiguous coverage, so it is necessary to verify its accuracy through a manual assisted audit mechanism. The audit process may include a dual strategy of confidence screening of the natural language processing model + manual audit to ensure that the answers finally included in the database will not mislead users or reduce the reliability of the Q&A system.
[0100] If the similarity is too high or there is a problem with the answer, the system will go back and regenerate or adjust the question and answer pair. The mathematical representation of this process is: ;in, is the set similarity threshold. If the similarity exceeds this threshold, the new question is considered to be too similar to the existing question and needs to be adjusted.
[0101] Based on the content of the above embodiment, it can be seen that analyzing user feedback data Compare the answers to questions in the extended Q&A database , which is ultimately used to determine the optimization parameters in the input closed-loop optimization model. The following is an explanation of the specific process of determining the optimization parameters for the extended question-answer database:
[0102] Specifically, the step of generating the question-and-answer result after optimizing the initial candidate answer by using a closed-loop optimization model based on the user feedback data and the data in the extended question-and-answer library comprises the following steps:
[0103] Step 1: Determine the quality score of each question-answer pair in the extended question-answer database based on the user feedback data and the data in the extended question-answer database.
[0104] First, we need to rate the user's feedback on the i-th user's question Perform weighted processing and then combine it with the corresponding user questions With the answer Analyze to get the quality score of each question-answer pair : ;in, is the quality score of the jth question-answer pair, n is the total number of question-answer pairs in the extended question-answer database, Score the user's feedback on the i-th question, is the user's i-th question and the jth question in the new question sequence The semantic similarity between
[0105] Step 2: Determine optimization parameters of the expanded question and answer library using the quality score, and input the optimization parameters and the user feedback data into a closed-loop optimization model to generate a question and answer result after optimizing the initial candidate answer.
[0106] In the present application example, according to the quality score obtained , the system can further evaluate the quality of questions in the extended Q&A database, filter out question-answer pairs that do not meet the standards, that is, filter out question-answer pairs with low quality scores, and delete or replace them to provide a basis for subsequent optimization steps. Through this analysis, the system can identify questions that do not match user feedback or have low quality, and then optimize or replace these questions.
[0107] For example, if the quality score of a question-answer pair is below a certain threshold , the question and answer pair will be marked as pending update: .
[0108] In the present application example, according to the quality score obtained , the system generates the optimization parameters for expanding the question-answer database , used to adjust and optimize the matching and generation strategies of the question answering system. Specifically, the optimization parameters are generated The process can be expressed by the following formula: ;in, is the adjustment coefficient, m is the number of questions in the question sequence under the current user intention category; is the user's i-th question and the jth question in the new question sequence The semantic similarity between .
[0109] It can be seen that user feedback data provides a quality evaluation of questions and answers generated during the interaction with users, while the extended question and answer database provides relevant semantic matching and answer information. By determining the optimization parameters of the extended question and answer database, it can be reflected which question and answer pairs are more suitable in quality and which ones need to be optimized. In this way, the optimization parameters will be used to adjust the questions and answers in the question and answer database to improve the performance of the overall question and answer system.
[0110] Furthermore, when generating optimization parameters After that, the system needs to update and expand the Q&A library based on the optimization results to ensure that the Q&A pairs in the library better meet user needs and improve the accuracy of semantic matching and intent prediction. , which can filter out low-quality question-answer pairs. For the replaced low-quality question-answer pairs, the generated optimization parameters can be used and user feedback data To generate new question-answer pairs, or to improve the quality of existing question-answer pairs by adjusting them. The updated question-answer pairs will be added back to the extended question-answer library and continue to be used by the question-answering system.
[0111] In a feasible implementation, by adjusting the optimization parameters, the system can also optimize the data structure in the extended question and answer library to improve retrieval efficiency and matching accuracy. For example, the index structure of the database is optimized or the storage format is adjusted to increase query speed. Ultimately, the updated extended question and answer library can contain question and answer pairs that better meet user needs, improving the accuracy of semantic matching and intent prediction.
[0112] Furthermore, in an embodiment of the present application, user feedback data and optimization parameters of the extended knowledge base are input into a closed-loop optimization model, and the semantic matching process, intent prediction process, and final answer generation process in the process of generating question and answer results are jointly optimized to improve the overall performance of the system.
[0113] The core of the closed-loop optimization model is to achieve dynamic performance improvement and feedback adjustment. The input of the model can be expressed as: In the process of combining the semantic matching process, the intent prediction process and the final answer generation process, the following steps are specifically included:
[0114] First, input the data Finally, the semantic matching process matches questions by comparing the semantic similarity between user input and existing questions. The optimized semantic matching algorithm will combine the semantic similarity scores in user feedback data to update the matching strategy and enhance the accuracy of matching. The update is performed using the following formula: ;in, is the semantic similarity, which indicates the matching degree between the user's question and the generated question; The rating of user feedback indicates the user's satisfaction with the problem; and is the weight coefficient, which is obtained through optimization calculation and is used to balance the impact of the original similarity and the user feedback score.
[0115] Secondly, the intent prediction process is optimized. The intent prediction process predicts the user's query intent by analyzing the text entered by the user. In the closed-loop optimization process, user feedback data helps optimize the accuracy of the model in identifying user intent. Specifically, the system optimizes the intent prediction model through the following adjustment formula: ;in, Indicates user intent The predicted probability of is the optimized semantic similarity, and m represents the current user intention category. The optimized intent prediction model can better identify and match the actual needs of users, thereby improving the system's accuracy in predicting intent.
[0116] Finally, the answer generation process is optimized. The answer generation process is responsible for generating specific answers from the intent prediction results. In the closed-loop optimization process, the generation module needs to be adjusted based on the optimized intent prediction results and semantic matching to generate more accurate and appropriate answers. Adjustment is performed using the following formula: ;in For the generated answer; is a generative model that predicts outcomes based on user intent , semantic similarity and optimization parameters Generate answers. Through closed-loop optimization, the generative model can provide more accurate answers that meet user needs based on user feedback and optimized parameters.
[0117] After completing the joint optimization of the above process, the closed-loop optimization model outputs the optimized question-answering results. This result not only combines user feedback data and optimization parameters, but also improves the overall performance of the system by optimizing semantic matching, intent prediction, and answer generation. Ultimately, the optimized question-answering system can better match user needs and generate more accurate answers. The specific output is:
[0118] ;in, represents the optimized generated answer set; The probability predicted for the intention; is the minimum acceptance probability threshold, and answers below this threshold will be discarded. Through this output, the system can push optimized question-answering results to users to improve user experience.
[0119] Based on the above embodiment, in a feasible implementation, in order to enhance the expressiveness and immersiveness of the system interaction, the system introduces an audio-driven dynamic expression and gesture synthesis mechanism after generating the question and answer results. This mechanism can drive the virtual interactive agent, such as a digital human, to perform expressions and body movements in real time according to the generated voice or answer content. Specifically, the following steps may also be included:
[0120] Generate a corresponding voice broadcast according to the question and answer result, and use the voice-driven face model to drive the virtual interactive agent to perform lip synchronization and facial movements according to the voice broadcast;
[0121] According to the intention prediction result, the posture synthesis model is used to drive the virtual interaction agent to generate a coherent body movement sequence.
[0122] In a specific implementation, the voice-driven face model can be a Wav2Lip model, which uses the Wav2Lip model to accurately drive lip synchronization and facial movements according to the voice; the posture synthesis model can be an E-NeRF model, which uses the E-NeRF model to generate a natural and coherent body movement sequence according to semantic intent and emotional tags. In this way, multi-modal consistency of expression, movement, and voice is achieved, and the realism and naturalness of human-computer interaction are enhanced.
[0123] It is worth noting that the present invention adopts a streaming output mechanism and feedback parameter input to meet the real-time interaction requirements. The system generates and feedbacks user requests through an in-stream output mechanism. This mechanism achieves a smoother and low-latency response experience by optimizing the output strategy of the system's generation module, specifically including: after the system receives the user request, it immediately performs an initial judgment of the intention and semantic analysis without waiting for all processing to be completed; the generation module adopts a step-by-step generation strategy to output part of the results in real time during the content generation process; the output content can be synchronously transmitted to the voice broadcast module and the digital human rendering module, so that the voice, expression and gesture are updated in coordination, and the natural interaction effect of "speaking and moving" and "answering and displaying" is achieved. It can be seen that this mechanism significantly reduces the system response delay and is suitable for application scenarios with high real-time requirements such as virtual human dialogue and voice question and answer. Furthermore, after the streaming output is completed, the system will input the user feedback data and optimization parameters into the closed-loop optimization model.
[0124] Based on the above embodiments, this embodiment exemplarily provides modules included in a question-answering system. The question-answering system may include: a data collection module, a semantic matching module, an intention prediction module, an answer generation module, a question-answering quality assessment and optimization module, a knowledge base update module, and a user interaction and visualization module.
[0125] The data collection module 10 is used to obtain data related to user questions and system answers from various data sources. Specifically, the module collects the questions input by the user in real time from the interaction between the user and the system. 、The answer returned by the system , and user feedback data on the answers , including evaluation scores, modification suggestions and additional contextual information. The collected data will provide the basis for subsequent semantic analysis and question-answering quality assessment.
[0126] The semantic matching module 20 is mainly used to match the questions input by the user and questions in the extended Q&A library Perform semantic similarity matching. By calculating the similarity between questions, the system uses natural language processing (NLP) technology and deep learning models to evaluate the relevance of user questions and system answers. The system uses pre-trained word embedding or semantic vector models to obtain high-quality semantic matching results. The calculated semantic similarity results will serve as the basis for subsequent intent prediction and answer generation.
[0127] The intention prediction module 30 is used to receive the similarity data provided by the semantic matching module, and predict the user's real needs by analyzing the user's intention. The system uses a deep learning algorithm to train the intention prediction model, identifies the core intention of the user's question based on a large amount of historical data and user feedback, and outputs possible answer categories. The accuracy of intention prediction directly affects the answer generation effect of the system, so this module adopts a dynamic adjustment mechanism to ensure that the model is continuously optimized to cope with different types of user needs.
[0128] The answer generation module 40 is based on the user needs output by the intention prediction module and combines the existing answers in the knowledge base The system generates personalized answers that meet user needs through natural language generation (NLG) technology, and the content generated by the system. The system generates the best answer based on multiple dimensions such as question type, user preferences, contextual information, etc. to ensure the relevance and readability of the answer. The generated answers will be evaluated for quality, and the generation process will be further optimized through feedback data.
[0129] The question-answer quality evaluation and optimization module 50 is responsible for evaluating the quality of the answers generated by the system. and quality score , the system continuously optimizes the question-answer pairs in the question-answer database. Specifically, the system calculates the quality score of each question-answer pair and uses the score results to update the knowledge items in the question-answer database to ensure that the system always provides high-quality answers. At the same time, the system adjusts the optimization parameters based on user feedback to continuously improve the accuracy of semantic matching and answer generation.
[0130] The knowledge base update module 60 is used to dynamically adjust the expanded question and answer base according to the quality evaluation results and optimization parameters. Update the questions and answers in the Q&A database, delete low-quality entries, and add new relevant content. This module uses a closed-loop feedback mechanism to enable the knowledge base to continuously adapt to new user needs and question types, improving the flexibility and accuracy of the overall Q&A system.
[0131] The user interaction and visualization module 70 is used to provide an intuitive interface for users to input questions and view the answers generated by the system. Users can use this interface to provide feedback on the quality and accuracy of the answers, and the system will further optimize the question and answer library based on these feedbacks. In addition, this module also supports multi-scenario simulation and interactive operations to help users understand the system's operating process and question and answer logic, thereby improving user experience.
[0132] It can be seen that through modular design, it is possible to efficiently process user input, accurately predict user intent, and generate high-quality personalized answers. Through continuous optimization and updating of the extended question and answer library, the system can self-learn and adjust in multiple rounds of interactions to improve overall performance. In particular, through the closed-loop feedback mechanism of the question and answer quality evaluation and optimization module, the system can continuously improve the knowledge base based on user feedback to ensure that the most relevant and accurate answers are provided for each interaction. The intelligent and adaptive characteristics of the system enable it to be applied in multiple fields, including customer service, technical support, and smart assistants, to provide users with more efficient and accurate services.
[0133] Correspondingly, the present application embodiment also discloses an intelligent question-answering device, see Figure 3 As shown, the device is applied to a question-answering system and includes:
[0134] A data acquisition module 11 is used to acquire a user question containing multimodal data and a candidate answer set corresponding to the user question;
[0135] A semantic matching module 12, configured to determine a first semantic similarity between the user question and the candidate answer set, and determine an initial candidate answer from the candidate answers according to the first semantic similarity;
[0136] The intention prediction module 13 is used to obtain historical dialogue information in the process of interaction with the user based on the initial candidate answer, and use a sliding window to perform context analysis on the historical dialogue information to generate an intention prediction result for the user's question;
[0137] An extended question and answer library construction module 14 is used to generate a new question sequence with similar semantics to the user's question based on the intention prediction result using a semantic generation model, and construct an extended question and answer library based on the new question sequence;
[0138] The question and answer result generation module 15 is used to obtain user feedback data, and based on the user feedback data and the data in the extended question and answer library, use a closed-loop optimization model to generate a question and answer result that optimizes the initial candidate answer, and push the question and answer result to the user; the user feedback data is the scoring data of the user's generated answer feedback for different questions.
[0139] Among them, for more specific working processes of the above-mentioned modules, please refer to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0140] It can be seen that the above scheme of this embodiment is applied to a question-and-answer system, including: obtaining a user question containing multimodal data and a candidate answer set corresponding to the user question; determining a first semantic similarity between the user question and the candidate answer set, and determining an initial candidate answer from the candidate answer set according to the first semantic similarity; based on the initial candidate answer, obtaining historical conversation information during the interaction with the user, and using a sliding window to perform context analysis on the historical conversation information to generate an intention prediction result for the user question; based on the intention prediction result, using a semantic generation model to generate a new question sequence that is semantically similar to the user question, and building an extended question-and-answer library based on the new question sequence; obtaining user feedback data, and based on the user feedback data and the data in the extended question-and-answer library, using a closed-loop optimization model to generate a question-and-answer result optimized for the initial candidate answer, and pushing the question-and-answer result to the user; the user feedback data is the scoring data of the user's generated answer feedback for different questions.
[0141] The beneficial technical effects of the present application are as follows: for user questions in multimodal data, the initial candidate answers are first screened out by calculating the semantic similarity between them and the candidate answers, thereby improving the accuracy of question-answer matching; secondly, a sliding window prediction mechanism is introduced to generate intention prediction results for user questions by predicting user needs, and the knowledge base of the question-answering system is enriched through semantically similar question generation technology to obtain an expanded question-answering base, which significantly enhances the diversity and comprehensiveness of the knowledge base and improves the system's coverage of user questions; finally, based on user feedback and the expanded question-answering base, more accurate answers that meet user needs are provided through closed-loop optimization.
[0142] Furthermore, the present application also discloses an electronic device. Figure 4 This is a structural diagram of an electronic device according to an exemplary embodiment. The content in the diagram cannot be regarded as any limitation on the scope of use of the present application.
[0143] Figure 4 A schematic diagram of the structure of an electronic device provided in an embodiment of the present application. The electronic device may specifically include: at least one processor 21, at least one memory 22, a power supply 23, a communication interface 24, an input / output interface 25, and a communication bus 26. The memory 22 is used to store a computer program, which is loaded and executed by the processor 21 to implement the relevant steps in the intelligent question-answering method disclosed in any of the aforementioned embodiments. In addition, the electronic device in this embodiment may specifically be a computer.
[0144] In this embodiment, the power supply 23 is used to provide working voltage for each hardware device on the electronic device; the communication interface 24 can create a data transmission channel between the electronic device and the external device, and the communication protocol it follows is any communication protocol that can be applied to the technical solution of the present application, and is not specifically limited here; the input and output interface 25 is used to obtain external input data or output data to the outside world, and its specific interface type can be selected according to specific application needs and is not specifically limited here.
[0145] In addition, the memory 22 as a carrier for resource storage may be a read-only memory, a random access memory, a disk or an optical disk, etc. The resources stored thereon may include an operating system 221, a computer program 222 and data 223, etc. The data 223 may include various data. The storage method may be temporary storage or permanent storage.
[0146] The operating system 221 is used to manage and control various hardware devices on the electronic device and the computer program 222, which can be Windows Server, Netware, Unix, Linux, etc. In addition to including a computer program that can be used to complete the intelligent question-answering method performed by the electronic device disclosed in any of the aforementioned embodiments, the computer program 222 can further include a computer program that can be used to complete other specific tasks.
[0147] Furthermore, the embodiment of the present application also discloses a computer-readable storage medium, where the computer-readable storage medium includes a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a magnetic disk or an optical disk, or any other form of storage medium known in the technical field. Wherein, the computer program implements the aforementioned intelligent question-and-answer method when executed by the processor. For the specific steps of the method, reference can be made to the corresponding contents disclosed in the aforementioned embodiments, which will not be repeated here.
[0148] In this specification, each embodiment is described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part.
[0149] The steps of the intelligent question-answering method or algorithm described in conjunction with the embodiments disclosed herein can be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module can be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the technical field.
[0150] Finally, it should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the sentence "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.
[0151] The above is a detailed introduction to the intelligent question-answering method, device, equipment and medium provided by the present invention. Specific examples are used in this article to illustrate the principles and implementation methods of the present invention. The description of the above embodiments is only used to help understand the method of the present invention and its core idea; at the same time, for general technical personnel in this field, according to the idea of the present invention, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as a limitation on the present invention.
Claims
1. An intelligent question-answering method, characterized in that: Applied to question answering systems, including: Obtaining a user question containing multimodal data and a candidate answer set corresponding to the user question; Determining a first semantic similarity between the user question and the candidate answer set, and determining an initial candidate answer from the candidate answer set according to the first semantic similarity; Based on the initial candidate answer, historical conversation information in the process of interaction with the user is obtained, and the historical conversation information is contextually analyzed using a sliding window to generate an intention prediction result for the user's question; According to the intention prediction result, a new question sequence with similar semantics to the user's question is generated using a semantic generation model, and an extended question and answer library is constructed based on the new question sequence; Obtain user feedback data, and based on the user feedback data and the data in the extended question and answer library, use a closed-loop optimization model to generate a question and answer result that optimizes the initial candidate answer, and push the question and answer result to the user; the user feedback data is the scoring data of the user's generated answer feedback for different questions.
2. The intelligent question-answering method according to claim 1, characterized in that: The multimodal data includes any one or a combination of text data, image data, voice data, expression data, and human posture data; wherein, after obtaining the user question containing the multimodal data and the candidate answer set corresponding to the user question, the method further includes: The user question and the candidate answer set are respectively format converted and feature extracted to generate corresponding cross-modal semantic embeddings.
3. The intelligent question-answering method according to claim 2, characterized in that: The user question is formatted and features are extracted to generate corresponding cross-modal semantic embedding, including: Performing word vector conversion on the text data to obtain a first feature vector, and inputting the first feature vector into a pre-trained word embedding model to obtain a first feature representation of the text data; Extracting a second feature vector of the image data using a preset convolutional neural network, and determining a second feature representation of the image data based on the second feature vector; Processing the speech data using a preset acoustic feature extraction method to obtain a third feature vector, and determining a third feature representation of the speech data based on the third feature vector; Extracting a fourth eigenvector of the expression data using a preset facial expression analysis tool, and determining a fourth feature representation of the expression data based on the fourth eigenvector; Extracting a fifth eigenvector of the human body posture data using a preset human body posture estimation model, and determining a fifth feature representation of the human body posture data based on the fifth eigenvector; The first feature representation, the second feature representation, the third feature representation, the fourth feature representation and the fifth feature representation are aligned, and the aligned first feature representation, the second feature representation, the third feature representation, the fourth feature representation and the fifth feature representation are respectively mapped to a shared semantic space using a preset multimodal embedding model to obtain a unified cross-modal semantic embedding.
4. The intelligent question-answering method according to claim 1, characterized in that: Obtaining a candidate answer set corresponding to the user's question, including: Acquire a first candidate answer corresponding to the user's question from the original question-and-answer database, and construct a candidate answer set based on the first candidate answer; Correspondingly, after constructing the extended question-answer library based on the new question sequence, the method further includes: A second candidate answer corresponding to the new question sequence is obtained from the extended question and answer library, and the second candidate answer is added to the candidate answer set.
5. The intelligent question-answering method according to claim 4, characterized in that: After constructing the extended question-answer library based on the new question sequence, the method further includes: Checking the second semantic similarity between the new question sequence generated in the extended question and answer library and the existing question sequence in the original question and answer library, and reviewing the answers to the questions in the new question sequence; When the second semantic similarity is greater than or equal to a preset similarity threshold, and / or the review result of the answer to the question does not meet the preset answer standard, a target question-answer pair is determined, and quality control is performed on the target question-answer pair.
6. The intelligent question-answering method according to claim 1, characterized in that: The step of generating a question-and-answer result after optimizing the initial candidate answer by using a closed-loop optimization model based on the user feedback data and the data in the extended question-and-answer database includes: Determine a quality score for each question-answer pair in the extended question-answer library based on the user feedback data and the data in the extended question-answer library; Determining optimization parameters of the expanded question and answer library using the quality score, and inputting the optimization parameters and the user feedback data into a closed-loop optimization model to generate a question and answer result after optimizing the initial candidate answer; The calculation formula of the quality score is: ; The calculation formula of the optimization parameter is: ; is the quality score of the jth question-answer pair, n is the total number of question-answer pairs in the extended question-answer database, Score the user's feedback on the i-th question, is the user's i-th question and the jth question in the new question sequence The semantic similarity between is the adjustment coefficient, and m is the number of questions in the question sequence under the current user intention category.
7. The intelligent question-answering method according to any one of claims 1 to 6, characterized in that: After the closed-loop optimization model is used to generate the question-answering result after optimizing the initial candidate answer, the method further includes: Generate a corresponding voice broadcast according to the question and answer result, and use the voice-driven face model to drive the virtual interactive agent to perform lip synchronization and facial movements according to the voice broadcast; According to the intention prediction result, the posture synthesis model is used to drive the virtual interaction agent to generate a coherent body movement sequence.
8. An intelligent question-answering device, characterized in that: Applied to question answering systems, including: A data acquisition module, used to acquire user questions containing multimodal data and a candidate answer set corresponding to the user questions; A semantic matching module, configured to determine a first semantic similarity between the user question and the candidate answer set, and determine an initial candidate answer from the candidate answers according to the first semantic similarity; An intention prediction module, configured to obtain historical conversation information during the interaction with the user based on the initial candidate answer, and perform context analysis on the historical conversation information using a sliding window to generate an intention prediction result for the user's question; An extended question-answer library construction module is used to generate a new question sequence with similar semantics to the user's question based on the intention prediction result using a semantic generation model, and to construct an extended question-answer library based on the new question sequence; The question and answer result generation module is used to obtain user feedback data, and based on the user feedback data and the data in the extended question and answer library, use a closed-loop optimization model to generate a question and answer result that optimizes the initial candidate answer, and push the question and answer result to the user; the user feedback data is the scoring data of the user's generated answer feedback for different questions.
9. An electronic device, characterized in that: The electronic device includes a processor and a memory; wherein the memory is used to store a computer program, and the computer program is loaded and executed by the processor to implement the intelligent question-answering method as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: Used to store computer programs; wherein the computer program, when executed by a processor, implements the intelligent question-answering method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Telephone customer service processing method and system based on personalized robot
CN118433311A
Text robot application system based on large model
CN119474323A
Dynamic question and answer matching method and system based on context awareness
CN119691133A
Method, apparatus and device for recognizing abnormal behavior on the basis of voice and image features
WO2021169209A1
Multi-modal knowledge-based question answering method and system for 5g message
WO2025060773A1
Cited By
Response method and electronic equipment
CN120633848A
Device and method for realizing telephone question and answer customer service agent based on large model RAG
CN121071099A