Intelligent question answering method and device, storage medium and computer device

CN122547925APending Publication Date: 2026-08-11KANG JIAN INFORMATION TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-02
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

然而,这种方式依赖信息单一,无法掌控用户完整语义,这就导致召回的信息与用户实际需求偏差较大

Benefits of technology

[0015] According to the present invention, an intelligent question-answering method, apparatus, storage medium, and computer device, compared with the current method of generating answers solely based on the user's current question, the present invention receives user question information, determines the user's intent based on the user question information, and obtains multi-round historical question-answering information consistent with the user's intent; then, it determines the user's emotion based on the user question information, determines the response style for the user question information based on the user's intent and the user's emotion, and determines the response content for the user question information based on the user question information and the multi-round historical question-answering information; finally, it responds to the user question information based on the response style and the response content. Therefore, by comprehensively analyzing the user question information and the multi-round historical question-answering information consistent with the user question information, the present invention can comprehensively and accurately understand the user's intent, making the answers more targeted and practical; by analyzing the user's emotion and intent to determine the response style, it can adapt to the needs of different user emotional states, thereby enhancing user emotional identification and satisfaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122547925A_ABST
    Figure CN122547925A_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent question-answering method, apparatus, storage medium, and computer device, relating to the fields of digital healthcare and fintech. It enables a more comprehensive and accurate understanding of user intent, resulting in more targeted and practical answers. The method includes: receiving user question information; determining user intent based on the user question information and acquiring multi-round historical question-answering information consistent with the user intent; determining user emotion based on the user question information; determining a response style for the user question information based on the user intent and the user emotion; determining response content for the user question information based on the user question information and the multi-round historical question-answering information; and responding to the user question information based on the response style and the response content. This invention is applicable to intelligent question-answering scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of digital healthcare technology and financial technology technology, and in particular to an intelligent question-answering method, device, storage medium and computer equipment. Background Technology

[0002] Users face numerous challenges in their daily lives and work, such as inquiries about medication plans, financial insurance products, and car rental information. As a result, users are increasingly eager to obtain the answers they need quickly and accurately.

[0003] Currently, answers are typically generated solely based on the question posed by the user. However, this approach relies on limited information and fails to capture the user's complete semantic understanding, leading to a significant discrepancy between the retrieved information and the user's actual needs. Summary of the Invention

[0004] This invention provides an intelligent question-answering method, device, storage medium, and computer equipment, which are mainly able to more comprehensively and accurately understand user intent, making the answers more targeted and practical.

[0005] According to a first aspect of the present invention, an intelligent question-answering method is provided, comprising: Receive user question information, determine user intent based on the user question information, and obtain multi-round historical question and answer information with the agreement graph of the user intent; Based on the user's question information, determine the user's emotion; based on the user's intent and the user's emotion, determine the response style for the user's question information; based on the user's question information and multi-round historical question and answer information, determine the response content for the user's question information. The system responds to the user's question based on the response style and the response content.

[0006] Optionally, determining the response content based on the user's question information and multi-round historical question-and-answer information includes: When the user's question information is multimodal information, determine the multimodal feature vector corresponding to the multimodal information, wherein the multimodal information includes at least two types of information selected from text information, video information, image information, and voice information; The vector magnitude of the multimodal feature vector is determined, the vector weight of the multimodal feature vector is determined based on the vector magnitude, and the multimodal feature vector is fused based on the vector weight to obtain a multimodal fused feature vector. Based on the multimodal fusion feature vector and multi-round historical question-and-answer information, the response content for the user's question is determined.

[0007] Optionally, determining the response content for the user's question based on the multimodal fusion feature vector and multi-round historical question-and-answer information includes: Determine the historical question-and-answer feature vector of multi-round historical question-and-answer information; The time interval between the question time of the user's question and the question and answer time of each historical question and answer is determined, and the time weight of each historical question and answer feature vector is determined based on the time interval. The multimodal feature vector and each historical question-and-answer feature vector are divided into multiple subspaces respectively. The subspace similarity of the multimodal feature vector and each historical question-and-answer feature vector in each subspace is determined respectively. Based on the subspace similarity, the spatial weight of each historical question-and-answer feature vector is determined respectively. The multiple subspaces include at least two of the following: entity subspace, intent subspace, and sentiment subspace. Based on the time weight and the spatial weight, the vector weight of the corresponding historical question and answer feature vector is determined, and based on the vector weight, each historical question and answer feature vector is weighted and fused to obtain a historical question and answer fused feature vector. The multimodal fusion feature vector and the historical question-and-answer fusion feature vector are interactively processed, and the interaction processing result is input into a preset response model to predict the response content, thereby obtaining the response content for the user's question information.

[0008] Optionally, obtaining multi-turn historical question-and-answer information related to the user intent agreement graph includes: Determine the intent complexity of the user intent and the average information complexity of all candidate historical question-and-answer information that agree with the user intent graph, respectively. Based on the intent complexity and the average information complexity, the information acquisition rounds of candidate historical question and answer information are determined, and the candidate historical question and answer information of the corresponding information acquisition rounds is stored in a cache pool. Multi-round historical question and answer information that agrees with the user intent graph is obtained from the candidate historical question and answer information in the cache pool.

[0009] Optionally, determining the response style for the user's question based on the user's intent and the user's emotion includes: Determine the intensity of the user's intent and the intensity of the user's emotion; Based on the user's intent and its corresponding intent intensity, and the user's emotion and its corresponding emotion intensity, determine the response style for the user's question.

[0010] Optionally, the method further includes: Determine the user characteristic data of multiple historical users who have engaged in historical intelligent question-and-answer interactions, and classify each historical user into different user categories based on the user characteristic data; In response to the input operation of the target user for the target question information, the target user's target feature data is obtained. Based on the target feature data and the user feature data, the target user category to which the target user belongs is determined in different user categories. And the target historical question information that is similar to the target question information is matched in the historical question information of the historical users under the target user category. The target question information is responded to based on the historical response strategy of the target historical question information, wherein the historical response strategy includes historical response content and historical response style.

[0011] Optionally, responding to the target question information based on the historical response strategy of the target historical question information includes: Determine historical user feedback information that responded to the target historical question information, and extract style evaluation information for the historical response style and content evaluation information for the historical response content from the historical user feedback information; Based on the style evaluation information, the historical response style is adjusted; based on the content evaluation information, the historical response content is adjusted; based on the style adjustment results and content adjustment results, the current response strategy for the target question is determined; and based on the current response strategy, the target question is answered.

[0012] According to a second aspect of the present invention, an intelligent question-answering device is provided, comprising: The acquisition unit is used to receive user question information, determine user intent based on the user question information, and acquire multi-round historical question and answer information that agrees with the user intent. The determining unit is configured to determine the user's emotion based on the user's question information, determine the response style for the user's question information based on the user's intent and the user's emotion, and determine the response content for the user's question information based on the user's question information and multi-round historical question and answer information. The response unit is used to respond to the user's question information based on the response style and the response content.

[0013] According to a third aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described intelligent question-answering method.

[0014] According to a fourth aspect of the present invention, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the above-described intelligent question-answering method.

[0015] According to the present invention, an intelligent question-answering method, apparatus, storage medium, and computer device, compared with the current method of generating answers solely based on the user's current question, the present invention receives user question information, determines the user's intent based on the user question information, and obtains multi-round historical question-answering information consistent with the user's intent; then, it determines the user's emotion based on the user question information, determines the response style for the user question information based on the user's intent and the user's emotion, and determines the response content for the user question information based on the user question information and the multi-round historical question-answering information; finally, it responds to the user question information based on the response style and the response content. Therefore, by comprehensively analyzing the user question information and the multi-round historical question-answering information consistent with the user question information, the present invention can comprehensively and accurately understand the user's intent, making the answers more targeted and practical; by analyzing the user's emotion and intent to determine the response style, it can adapt to the needs of different user emotional states, thereby enhancing user emotional identification and satisfaction. Attached Figure Description

[0016] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings: Figure 1 A flowchart of an intelligent question-answering method provided by an embodiment of the present invention is shown; Figure 2 This invention provides a flowchart of another intelligent question-answering method according to an embodiment of the invention. Figure 3 This diagram illustrates the structure of an intelligent question-answering device according to an embodiment of the present invention. Figure 4 This invention provides a schematic diagram of the structure of another intelligent question-answering device according to an embodiment of the invention. Figure 5 A schematic diagram of the physical structure of a computer device provided in an embodiment of the present invention is shown. Detailed Implementation

[0017] The present invention will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the present application can be combined with each other.

[0018] Currently, the method of generating answers based solely on the question asked by the user cannot grasp the user's complete semantics, which leads to a significant discrepancy between the retrieved information and the user's actual needs.

[0019] To address the aforementioned problems, embodiments of the present invention provide an intelligent question-answering method, such as... Figure 1 As shown, the method includes: 101. Receive user questions, determine user intent based on user questions, and obtain multi-round historical question and answer information that aligns with the user intent.

[0020] In this embodiment of the invention, user questions are received in real time via API interfaces or message queues, supporting multimodal input such as text and voice. For example, in an e-commerce customer service scenario, a user enters "Does this phone come in black?" through a web chat window. The system encapsulates this text string into a JSON format message, including the user ID, question timestamp, and original text content, and pushes it to the subsequent processing module. The user's intent is determined based on the question information, such as "product inquiry," "size consultation," or "discount inquiry." For example, if a user asks "I want to return an item," combined with the previous statement "The clothes I just received are damaged," the intent is identified as "after-sales return." A cache pool stores all candidate historical question-and-answer information for the user. Multiple rounds of historical question-and-answer information are matched against a consensus graph in the cache pool based on the user's intent. For example, if a user currently asks "Does it come in red?", the intent is identified as "color consultation," which also falls under product consultation. Therefore, question-and-answer information with the same "product consultation" intent is filtered from the cache pool, such as the historical question-and-answer information "Does this phone come in black?".

[0021] 102. Determine user sentiment based on user question information; determine response style based on user intent and user sentiment; determine response content based on user question information and multi-round historical question and answer information.

[0022] In this embodiment of the invention, lexical features, punctuation and tone, and contextual information in user questions can be identified to determine the user's emotions. For example, lexical feature analysis: analyzing the words in user questions reveals positive emotion words such as "happy," "satisfied," and "fantastic," and negative emotion words such as "angry," "disappointed," and "terrible." The presence of these clearly emotionally charged words in the question allows for a preliminary assessment of the user's emotional state. For instance, if a user asks, "This product is so disappointing," the word "disappointed" directly indicates a negative emotion. Furthermore, adverbs of degree of attention, such as "very" and "extremely," intensify the emotional intensity. For example, "I am very angry" expresses a stronger negative emotion than simply "I am angry." Punctuation and Tone Analysis: Identify punctuation and tone in user questions. Exclamation marks often express strong emotions, such as "This service is terrible!", conveying anger. Multiple consecutive question marks indicate escalating doubt into dissatisfaction, such as "What's going on? Why isn't it resolved? What are you doing?". Ellipses suggest helplessness or unspoken feelings. Furthermore, the overall tone of the text, such as questioning, complaining, or praising, helps determine the user's emotions. For example, a user asking "How could you handle this?" suggests strong agitation and dissatisfaction. Contextual Analysis: Utilize the context of previous question-and-answer sessions to further determine user emotions. If a user has expressed dissatisfaction in previous rounds, and the current question, while not using explicit emotional language, continues to revolve around the same issue, it suggests the user remains negatively affected. For example, if a user complained about slow delivery in previous rounds, and the current question, "When will it arrive?", indicates anxiety and dissatisfaction stemming from waiting.

[0023] Furthermore, based on the identified user emotions and intentions, a mapping relationship between emotion and intention-based response styles is established. When the user's intention is product inquiry and their emotion is positive, the response style should be professional and detailed, but also warm and friendly, such as using phrases like "We are very pleased that you are interested in our product. Let me introduce it to you in detail...". If the user's intention is after-sales complaint and their emotion is negative and angry, the response style should be patient and sincere, emphasizing reassurance, such as "We understand your anger at this moment. We will definitely resolve the issue for you as soon as possible...". If the user's intention is feedback and their emotion is neutral, the response style should focus on respect and seriousness, such as "Thank you for your valuable suggestion. We will consider it carefully...". The initially determined response style can also be fine-tuned based on the user's historical interaction habits and preferences. If the user prefers concise and direct responses in past interactions, redundant expressions should be appropriately reduced in the current response style; if the user prefers a more formal language style, more standardized vocabulary and sentence structures should be used in the response.

[0024] Furthermore, key content is extracted from user questions, including core elements such as the product, problem, and need involved. For example, if a user asks, "I bought [product name] and encountered [specific problem], how can I solve it?", "[product name]" and "[specific problem]" are extracted as key information. Simultaneously, content related to the current question is searched in multiple rounds of historical Q&A information. This includes checking if the user has asked similar questions before, the system's previous responses, and subsequent user feedback. For example, if a user previously inquired about the usage of a certain medicine and is now asking a related question, supplementary answers can be provided based on parts not covered in previous responses or new questions raised by the user. At the same time, attention is paid to whether historical Q&A includes information on problem-solving progress and related commitments, ensuring that the current response is consistent and coherent with historical information. Based on the extracted key information and related historical Q&A, matching answers are searched from a pre-built knowledge base. The knowledge base contains multiple question-answer pairs, such as product information, frequently asked questions, and business process descriptions. For example, when a user asks about product malfunctions, common causes and solutions are retrieved from the knowledge base; for business processing questions, accurate processing procedures and required materials are provided; for insurance recommendation questions, insurance products that meet the user's needs are offered with a brief introduction. Information from the knowledge base is integrated with relevant content from historical Q&A sessions to form a complete response. The response is ensured to be logically clear, accurate, and consistent with a predetermined response style. For example, if the response style is warm and friendly, more approachable language is used; if the style is professional and formal, accurate industry terminology and standardized sentence structures are used. Simultaneously, the response is checked to ensure it fully answers the user's question, avoiding omissions of key information. This method accurately identifies user emotions, determines an appropriate response style based on user intent and emotions, and generates accurate, coherent, and user-relevant responses by combining information from multiple rounds of historical Q&A. This effectively improves the efficiency and satisfaction of user-system communication in intelligent interaction systems, enhances user trust and reliance on the system, and is applicable to various intelligent interaction scenarios, such as medical and fintech scenarios.

[0025] 103. Respond to user questions based on response style and content.

[0026] In this embodiment of the invention, considering the user's historical interaction records and preferences, the response style and content are personalized and fine-tuned. If the user prefers concise replies in past interactions, the language in the current response is appropriately simplified, and redundant information is removed; if the user prefers detailed explanations, relevant background knowledge and extended information are added to the response content. The response statements are organized according to the selected response style and adjusted response content. The statements are ensured to be fluent, clear, and conform to natural language expression habits. When organizing the statements, conjunctions and punctuation are used appropriately to enhance the logic and coherence of the statements. Furthermore, the format of the response is standardized, including font, font size, color, paragraph spacing, etc. Appropriate formats are selected based on different interaction scenarios and user devices to ensure that users can read the response content clearly and comfortably. For example, in mobile applications, a larger font size and concise layout are used to facilitate viewing on small screens; on web pages, an aesthetically pleasing and easy-to-read format can be designed according to the page layout. The organized and standardized response content is then output to the user through appropriate channels. In text-based interaction scenarios, such as online chat windows and text messages, the response text is sent directly to the user. In voice-based interaction scenarios, such as intelligent voice assistants, the response text is converted into speech and played. When outputting responses, attention should be paid to the timing and frequency to avoid prolonged user waiting or information overload. For example, for responses to complex questions, key information can be output in stages to keep users informed of the progress. This invention can accurately match appropriate response styles and content based on different user questions and scenarios, generating natural, accurate, and personalized responses, effectively improving the user's interaction experience with the system.

[0027] According to the intelligent question-answering method provided by the present invention, compared with the current method of generating answers solely based on the user's current question, the present invention receives user question information, determines the user's intent based on the user question information, and obtains multi-round historical question-and-answer information with the user's intent agreement graph; then, it determines the user's emotion based on the user question information, determines the response style for the user question information based on the user's intent and the user's emotion, and determines the response content for the user question information based on the user question information and the multi-round historical question-and-answer information; finally, it responds to the user question information based on the response style and the response content. Therefore, by comprehensively analyzing the user question information and the multi-round historical question-and-answer information with the user question information agreement graph, the present invention can comprehensively and accurately understand the user's intent, making the answers more targeted and practical; by analyzing the user's emotion and intent to determine the response style, it can adapt to the needs of different user emotional states, thereby enhancing user emotional identification and satisfaction.

[0028] Furthermore, to better illustrate the above-described intelligent question-answering process, as a refinement and extension of the above embodiments, this invention provides another intelligent question-answering method, such as... Figure 2 As shown, the method includes: 201. Receive user questions, determine user intent based on user questions, and obtain multi-round historical question and answer information that aligns with the user intent agreement graph.

[0029] In this embodiment of the invention, on a webpage, user-inputted questions are received through preset online chat windows, consultation forms, etc.; on a mobile application, user questions are obtained using built-in chat interfaces, problem feedback entry points, etc.; for voice interaction scenarios, voice recognition technology is used to convert user voice questions into text information. For example, on an e-commerce platform's webpage, if a user enters "How is the battery life of this phone?" in the consultation box on the product details page, the system instantly captures this text information as the user's question. After receiving the user's question, a preliminary check is performed on the completeness of the question, checking for obvious garbled characters, missing key content, etc. If the information is found to be incomplete, a prompt is proactively sent to the user, guiding them to complete the information. For example, if the user only enters "How…", the system replies, "Your question seems incomplete; please provide more details so that I can better answer your question." Furthermore, keywords are extracted from user questions, and user intent is determined based on these keywords. For example, if the extracted keywords mainly relate to the coverage of insurance products, the user intent is likely to be inquiring about the coverage of insurance products; if the keywords revolve around after-sales issues, the user intent is determined to be after-sales consultation or complaint; if the keywords revolve around medication dosage, the user intent is determined to be medication plan consultation. Further, to accurately answer user questions, it is also necessary to determine multi-round historical question-and-answer information consistent with the user intent agreement graph. Based on this, step 202 specifically includes: determining the intent complexity of the user intent and the average information complexity of all candidate historical question-and-answer information consistent with the user intent agreement graph; determining the information acquisition round of the candidate historical question-and-answer information based on the intent complexity and the average information complexity; storing the candidate historical question-and-answer information of the corresponding information acquisition round in a cache pool; and retrieving multi-round historical question-and-answer information consistent with the user intent agreement graph from the candidate historical question-and-answer information in the cache pool.

[0030] Specifically, the process involves determining multi-dimensional intent attributes such as the scope of the knowledge domain involved, the level of abstraction of the intent, and the difficulty of resolving the intent. Based on these multi-dimensional intent attributes, the complexity of the user's intent is determined. For example, a scope score is assigned based on the scope of the knowledge domain, an abstraction score is assigned based on the level of abstraction, and a difficulty score is assigned based on the difficulty of resolving the intent. Weighting coefficients are then determined for each of these factors. The scope score, abstraction score, and difficulty score are then weighted and summed to obtain a comprehensive score. Finally, the complexity of the intent is determined based on this comprehensive score. A higher comprehensive score indicates greater intent complexity. At the same time, the complexity of an intent can also be assessed based on any of the following attributes: the scope of the knowledge domain, the level of abstraction of the intent, and the difficulty of resolving the intent. For example, if an intent involves multiple different and weakly related knowledge domains, such as simultaneously involving medical and physical knowledge, then the intent complexity is high. Highly abstract intents, such as "exploring the future development direction of mankind," have higher intent complexity than specific and clear intents, such as "checking the price of a certain book." Intents that require multi-step reasoning and the integration of multiple factors to resolve, i.e., those with greater difficulty in resolving the intent, are judged to have high intent complexity.

[0031] Furthermore, after determining the user's intent, all candidate historical question-and-answer messages with the same intent are retrieved from the historical question-and-answer database. For each candidate historical question-and-answer message, the information complexity is determined based on the intent complexity evaluation method described above. Specifically, for each candidate historical question-and-answer message, multi-dimensional attribute information such as information richness, logical structure complexity, and information professionalism is obtained. Then, each dimension of attribute information is scored, and each score result is weighted and summed. The information complexity of the corresponding candidate historical question-and-answer message is determined based on the weighted sum. Alternatively, information complexity can be determined based on any one of the following: information richness, logical structure complexity, or information professionalism. For example, when assessing complexity based on information richness, candidate historical question-and-answer information containing a large amount of detailed data, cases, explanations, etc., has high content richness and correspondingly high complexity. When assessing complexity based on the complexity of information logical structure, candidate historical question-and-answer information with a clear and hierarchical logical structure has low complexity, while candidate historical question-and-answer information with complex and difficult-to-organize logical relationships has high complexity. When assessing complexity based on the level of information specialization, candidate historical question-and-answer information involving many professional terms and in-depth professional knowledge has high complexity. After determining the information complexity of each candidate historical question-and-answer information, the average information complexity of all candidate historical question-and-answer information is determined.

[0032] Furthermore, after calculating the average of the intent complexity and information complexity, the information acquisition round N is calculated according to the following formula:

[0033] in, The adjustment parameter for the degree of influence of intent complexity on the number of information acquisition rounds. This is a parameter used to adjust the degree of influence of the mean information complexity on the number of information acquisition rounds. It is a constant. For the complexity of intent, This represents the average information complexity.

[0034] Furthermore, all candidate historical question-and-answer information is sorted in reverse chronological order, and a time decay weight is assigned to each sorted candidate historical question-and-answer information i. :

[0035] in, t is the attenuation coefficient, and t is the time interval between historical question and answer information i and its corresponding previous historical question and answer information i-1.

[0036] Furthermore, based on the time decay weights and information retrieval rounds of candidate historical question-and-answer information, target candidate historical question-and-answer information to be stored in the cache pool is selected from all mutually selected historical question-and-answer information. Specifically, high-weight candidate historical question-and-answer information with a time decay weight greater than a preset threshold is identified from all candidate historical question-and-answer information. Then, target candidate historical question-and-answer information for the corresponding information retrieval rounds is determined from the high-weight candidate historical question-and-answer information. For example, if there are 5 information retrieval rounds, then 5 rounds of question-and-answer information are selected from the high-weight candidate historical question-and-answer information as target candidate historical question-and-answer information and stored in the cache pool. Finally, multi-round historical question-and-answer information is directly retrieved from the target candidate historical question-and-answer information stored in the cache pool. For example, if the cache pool stores 5 rounds of target candidate historical question-and-answer information, then 3 rounds of historical question-and-answer information are randomly selected from the cache pool as multi-round historical question-and-answer information.

[0037] The intent complexity in this invention reflects the difficulty and scope of the current user intent, while the average information complexity reflects the uncertainty and amount of information in historical question-and-answer information. By comprehensively considering these two factors to determine the information retrieval rounds, the system's need for historical information can be more accurately matched. For example, for complex intents and high historical information complexity, increasing the retrieval rounds can ensure that enough relevant information is obtained to deeply understand the user intent, thereby generating answers that better meet the user's actual needs. The cache pool has fast read and write characteristics, storing historical question-and-answer information therein, which can be quickly retrieved when historical information is needed, reducing the time spent searching from the original database and improving the system's response speed. The time decay weight considers the time factor of historical question-and-answer information. As time goes by, the relevance and importance of information may gradually decrease. By introducing the time decay weight, the value of historical information can be more reasonably evaluated, thereby reasonably and accurately selecting historical question-and-answer information to participate in the generation process of response content.

[0038] 202. Determine user sentiment based on user question information, and determine the response style for user question information based on user intent and user sentiment.

[0039] User emotions include, but are not limited to, happiness, sadness, frustration, calmness, confusion, and anger. For this embodiment of the invention, to enhance the user's interactive experience, it is also necessary to determine the user's emotions based on the user's question information. Specifically, keywords are extracted from the user's question information, and user emotions are determined based on these keywords. For example, if the user's question information contains keywords such as "happy," "excited," and "satisfied," the user is judged to be in a positive emotional state, i.e., happy or excited. If keywords such as "angry," "disappointed," and "frustrated" appear, the user is judged to be in a negative emotional state, i.e., angry or frustrated. Neutral words, such as "inquire" and "understand," do not inherently carry a clear emotional tendency and require further analysis of the user's emotions in the context. Furthermore, keywords can be combined with punctuation and tone in the user's question information to determine user emotions. For example, the extensive use of exclamation marks usually indicates strong user emotions, while questioning or complaining tones often indicate that the user is in a negative emotional state. The context of multiple rounds of historical question-and-answer information can also be used to further verify and determine user emotions.

[0040] Furthermore, in order to improve the user interaction experience, it is also necessary to determine the response style for the user. Based on this, step 202 specifically includes: determining the intent intensity of the user's intent and the emotional intensity of the user's emotion; and determining the response style for the user's question information based on the user's intent and its corresponding intent intensity, and the user's emotion and its corresponding emotional intensity.

[0041] Among these, response styles include, but are not limited to, a professional and rigorous style, a gentle and patient style, a concise and clear style, a friendly and lively style, an encouraging style, and a humorous and witty style.

[0042] Specifically, intent keywords related to the user's intent are identified from the user's query information. Based on the importance of these keywords to the user's intent, each intent keyword is categorized into three levels: core, important, and general. Core keywords are those that play a decisive role in clarifying the user's intent; important keywords are those closely related to the intent and help clarify it; and general keywords are those that are somewhat related but less important. Further, the position of each intent keyword in the user's query information is determined, such as at the beginning, middle, or end. Then, for each intent keyword, its corresponding intent contribution value is determined based on its level and position. For example, a level score is assigned based on the intent keyword's classification. Based on the intent keyword's position in the user's query information, a corresponding position weight is assigned, such as a higher weight for the beginning and a lower weight for the middle. The product of the position weight and the level score is then used as the intent contribution value for the corresponding intent keyword. The intent contribution value of each intent keyword can be calculated using the above method. By summing the intent contribution values ​​of all intent keywords, the intent intensity value of the user's question can be obtained.

[0043] Furthermore, emotional keywords related to the user's emotions are identified in the user's question information, and corresponding emotional intensity values ​​are assigned to different emotional keywords. For example, the intensity value of "happy" is 0.3, the intensity value of "excited" is 0.6, and the intensity value of "ecstatic" is 0.9. Then, the intensity values ​​of all identified emotional keywords are added together to obtain the total emotional intensity value. For example, if a user says, "I am very angry, this service is terrible," the intensity value of "angry" is 0.7, and "very," as an adverb of degree, can add an additional intensity value of 0.2, for a total emotional intensity of 0.9. In another embodiment of the invention, the punctuation and tone of the user's question information can also be determined. If exclamation marks are used extensively, the emotional intensity value can be appropriately increased; when the tone of questioning or complaining is obvious, the emotional intensity is also increased; if the tone is calm, the emotional intensity may be decreased. For example, if a user says in a strong questioning tone, "How could you do this!", the emotional intensity can be increased by an appropriate amount on top of the original emotional intensity.

[0044] Furthermore, by comprehensively considering the user intent, intent strength, user emotion, and emotion strength obtained from the analysis, different response style strategies should be adopted under different combinations of intent and emotion.

[0045] Based on experience and an understanding of user needs, a mapping table is developed that corresponds response styles to intent intensity and emotional intensity. The calculated user intent intensity and emotional intensity are then matched against this mapping table to determine the final response style. For example, when intent intensity is high and emotional intensity is low, a professional and detailed response style is used to ensure the user fully understands the relevant information. If both intent and emotional intensity are high, the response style should be gentler and more patient, first calming the user before providing a solution. When the user is emotionally positive and has high intent intensity, such as excitedly inquiring about how to purchase a popular product, a friendly, lively, and enthusiastic response style is used. The reply can include cheerful interjections and emoticons, such as, "Wow, you're so interested in this product! The purchase method is very simple...", making the user feel valued and cared for. If the user is emotionally positive but has low intent intensity, such as casually mentioning that a product is good, a gentle, friendly, and relaxed response style is used. The reply could be, "It seems you really like this product; it certainly has many advantages...", communicating with the user in a friendly manner.

[0046] 203. When the user's question information is multimodal, determine the multimodal feature vector corresponding to the multimodal information, wherein the multimodal information includes at least two types of information such as text information, video information, image information, and voice information.

[0047] 204. Determine the vector magnitude of the multimodal feature vector, determine the vector weight of the multimodal feature vector based on the vector magnitude, and perform vector fusion on the multimodal feature vector based on the vector weight to obtain the multimodal fused feature vector.

[0048] Specifically, feature vectors corresponding to text, video, image, and audio information can be determined separately using feature extraction models, such as CNN models. For each feature vector, taking the text feature vector as an example, the text feature vector is calculated using the following formula. vector magnitude :

[0049] in, , ... These are vector elements in the text feature vector. Therefore, the vector magnitude of each feature vector can be determined in the above manner. If the multimodal feature vector includes text feature vectors, video feature vectors, image feature vectors, and speech feature vectors, firstly, the vector magnitudes corresponding to the text feature vectors, video feature vectors, image feature vectors, and speech feature vectors are added together to obtain the total vector magnitude. The ratio of the vector magnitude of the text feature vector to the total vector magnitude is used as the vector weight of the text feature vector; the ratio of the vector magnitude of the video feature vector to the total vector magnitude is used as the vector weight of the video feature vector; the ratio of the vector magnitude of the image feature vector to the total vector magnitude is used as the vector weight of the image feature vector; and the difference between 1 and the sum of the above vector weights is used as the vector weight of the speech feature vector. Then, based on the vector weights of each feature vector, each feature vector is weighted and fused to obtain the multimodal fused feature vector. The multimodal fused feature vector of this embodiment integrates information from different modalities, comprehensively and holistically capturing various information in user questions, avoiding misunderstandings caused by missing information. For example, in a medical consultation scenario, a patient may describe their symptoms in text and upload relevant medical images at the same time. Multimodal fusion can utilize both the symptom description in text and the pathological features in the images to more accurately determine the condition.

[0050] 205. Based on multimodal fusion feature vectors and multi-round historical question-and-answer information, determine the response content for user questions.

[0051] In this embodiment of the invention, after determining the multimodal fusion feature vector, it is necessary to comprehensively analyze the multimodal fusion feature vector and multi-round historical question-and-answer information to determine the content. Based on this, step 205 specifically includes: determining the historical question-and-answer feature vector of the multi-round historical question-and-answer information; determining the time interval between the question time of the user's question and the question-and-answer time of each historical question-and-answer information, and determining the time weight of each historical question-and-answer feature vector based on the time interval; dividing the multimodal fusion feature vector and each historical question-and-answer feature vector into multiple subspaces, and determining the multimodal fusion feature vector and each historical question-and-answer feature vector in each subspace. The similarity of the subspaces is used to determine the spatial weight of each historical question-and-answer feature vector based on the similarity of each subspace. The multiple subspaces include at least two of the following: entity subspace, intent subspace, and sentiment subspace. Based on the temporal weight and the spatial weight, the vector weight of the corresponding historical question-and-answer feature vector is determined. Based on the vector weight, each historical question-and-answer feature vector is weighted and fused to obtain a historical question-and-answer fused feature vector. The multimodal fused feature vector and the historical question-and-answer fused feature vector are interactively processed, and the interaction processing result is input into a preset response model to predict the response content, thereby obtaining the response content for the user's question information.

[0052] Specifically, a feature extraction model, such as a CNN model, is used to extract historical question-and-answer feature vectors. Further, historical question-and-answer information with a shorter time interval from the user's question is assigned a higher time weight, while historical question-and-answer information with a longer time interval is assigned a lower time weight. This allows each historical question-and-answer feature vector to be assigned an appropriate time weight. Simultaneously, the multimodal fusion feature vector and each historical question-and-answer feature vector are divided into multiple subspaces. In this embodiment, the multiple subspaces include at least two of the following: an entity subspace, an intent subspace, and a sentiment subspace. The entity subspace represents the entity objects in the information, such as product names or person names; the intent subspace reflects the intent expressed by the information, such as inquiries, suggestions, or complaints; and the sentiment subspace reflects the emotional tendency of the information, such as positive, negative, or neutral. In the entity subspace, the entity similarity between the multimodal fusion feature vector and each historical question-and-answer feature vector is calculated separately. In the intent subspace, the intent similarity between the multimodal fusion feature vector and each historical question-and-answer feature vector is calculated separately. In the sentiment subspace, the sentiment similarity between the multimodal fusion feature vector and each historical question-and-answer feature vector is calculated separately. Then, for each historical question-and-answer feature vector, the entity similarity, intent similarity, and sentiment similarity of that historical question-and-answer feature vector are weighted and summed to obtain the comprehensive similarity between the corresponding historical question-and-answer feature vectors. The higher the comprehensive similarity, the higher the spatial weight of the corresponding historical question-and-answer feature vector; the lower the comprehensive similarity, the lower the spatial weight. Thus, the spatial weight of each historical question-and-answer feature vector can be determined in the above manner. Further, for each historical question-and-answer feature vector, its corresponding time weight and spatial weight are weighted and summed to obtain the vector weight of the corresponding historical question-and-answer feature vector. Finally, based on the vector weights, each historical question-and-answer feature vector is weighted and fused to obtain the historical question-and-answer fusion feature vector.

[0053] Furthermore, to extract more latent features, the multimodal fusion feature vector and the historical question-answering fusion feature vector need to be interactively processed. This interaction can be achieved through weighted fusion or cross-processing. Next, a pre-defined response model is used to predict the response content. This model consists of an input layer, a backbone network, and a prediction network. The backbone network uses an improved version of the Transformer-XL architecture, which, compared to the Transformer, introduces a "segment-level recurrent mechanism" and "relative position encoding." The interaction processing results are input into the backbone network through the input layer for feature processing, and the prediction network uses this result to predict the response content. Prior to this, in order to improve the prediction accuracy of the model, it is first necessary to train and construct a preset response model. Based on this, the method includes: constructing a preset initial response model; obtaining a sample dataset, wherein the sample dataset includes sample user question information with response content labels and historical question and answer information with the agreement graph of the sample user question information; dividing the sample dataset into a training set and a test set, using the training set to train the preset initial response model, and using the test set to test the trained preset initial response model, and finally using the trained preset initial response model that meets the test conditions as the preset response model.

[0054] Specifically, during model training, a pre-defined initial response model is first constructed, followed by the acquisition of a sample dataset. The dataset is ensured to contain all necessary files. The data is then converted to a format understandable by the pre-defined initial response model. Finally, the model is trained and tested. Specifically, the dataset can be divided first: using randomness or a specific strategy (such as stratified sampling), the sample dataset is divided into training and test sets. The model is then trained using the training set, and tested using the test set to evaluate its performance on unseen data. Precision, recall, and other metrics on the test set are calculated and recorded. If the model performance does not meet requirements, it can return to the training phase for further iterations or adjustments. This process yields a pre-defined response model that meets the requirements. Finally, the interaction results of the multimodal fusion feature vector and the historical question-and-answer fusion feature vector are input into the pre-defined response model for response content prediction. This invention uses a model to predict response content without human intervention, thus improving the accuracy and efficiency of response content prediction. Furthermore, by interactively processing multimodal fusion feature vectors and historical question-and-answer fusion feature vectors, this invention can extract more latent features, making fuller use of information and improving the rationality of subsequent response content determination, thereby meeting the actual needs of users.

[0055] 206. Respond to user questions based on response style and content.

[0056] In this embodiment of the invention, the response content is displayed to the user according to the response style, thereby realizing the response to the user's question.

[0057] In another embodiment of the present invention, in order to improve the response efficiency of subsequent users based on intelligent question-and-answer interaction and thus improve the user experience, the subsequent question information can be directly answered based on the response content and response style of similar historical questions. Based on this, the method includes: determining user feature data of multiple historical users who have performed historical intelligent question-and-answer interactions; classifying each historical user into different user categories based on the user feature data; in response to the input operation of the target user for the target question information, obtaining the target user's target feature data; determining the target user category to which the target user belongs in different user categories based on the target feature data and the user feature data; matching the target historical question information similar to the target question information in the historical question information of historical users under the target user category; and answering the target question information according to the historical response strategy of the target historical question information, wherein the historical response strategy includes historical response content and historical response style.

[0058] The user characteristic data and target characteristic data include, but are not limited to, user-authorized information such as age, gender, and interests. Specifically, data related to multiple historical users who have engaged in past intelligent question-and-answer interactions are collected from past records. These data sources include, but are not limited to, personal information forms filled out by users during question-and-answer interactions, language habits and preferences expressed during the interaction, and the subject areas covered by past questions. The collected data is then analyzed in depth to extract user characteristic data that represents the characteristics of each historical user. User characteristic data covers multiple dimensions, such as age, gender, occupation, interests, question frequency, and question complexity. Similarly, target characteristic data for target users is obtained using the same method. It should be noted that the above characteristic data is authorized by the users and does not involve user privacy information. Based on the extracted user characteristic data, cluster analysis or other suitable classification algorithms are used to classify each historical user into different user categories. For example, based on interests, users can be divided into categories such as technology enthusiasts, sports enthusiasts, and art enthusiasts; based on question frequency and complexity, users can be divided into categories such as high-frequency professional questioners and low-frequency simple questioners. This classification allows for better management and analysis of users with similar characteristics. Furthermore, when a target user inputs information related to the target question, the system acquires the target user's target feature data in real time. The similarity between the target feature data and historical user feature data in each user category is calculated, and the target user is categorized into the user category with the highest similarity. For example, if the target user's interests are primarily focused on the technology field, and their questioning style is similar to historical users in the technology enthusiast category, then the system will categorize the target user into the technology enthusiast category. After determining the target user's category, the system searches and matches historical questions from users within that category to find similar historical questions. The matching process comprehensively considers multiple aspects of the target question, including keywords, semantics, and theme. For example, for the target question "How to use the camera function of a new smartphone," the system searches historical questions in the technology enthusiast category for questions containing keywords or semantically similar phrases such as "smartphone," "camera function," and "usage method," and selects these as matching historical questions. Furthermore, by drawing on valid information from historical responses to the target's questions, and appropriately adjusting and supplementing it according to the specific circumstances of the target's questions, the current response content is used. The historical response style of the target's questions is referenced to ensure consistency between the response style and the historical response style. Finally, a response is sent to the target user based on the response style and content of the target's questions. This embodiment of the invention uses historical response strategies to determine the response strategy for the current user's question, thereby improving the efficiency of response strategy determination.

[0059] In this embodiment of the invention, to further improve the response effect, it is also necessary to adjust the response content and response style based on feedback information from historical users regarding their responses. Based on this, the method includes: determining historical user feedback information that responded to the target historical question; extracting style evaluation information for the historical response style and content evaluation information for the historical response content from the historical user feedback information; adjusting the style of the historical response style based on the style evaluation information; adjusting the content of the historical response content based on the content evaluation information; determining the current response strategy for the target question based on the style adjustment result and the content adjustment result; and responding to the target question based on the current response strategy.

[0060] Specifically, after responding to a target historical question, the intelligent question-answering system initiates a feedback collection mechanism to obtain feedback from historical users regarding that response. Feedback can be collected through various means. For example, a dedicated feedback entry can be set up on the question-answering interface, guiding historical users to click the feedback button after viewing the response and enter a feedback page to fill in specific opinions. Feedback surveys can also be sent to historical users via email or SMS, covering questions about satisfaction with the response and suggestions for improvement. Furthermore, the system's built-in automatic feedback monitoring function can be used to analyze the subsequent behavior of historical users after receiving a response, such as whether they continue to ask follow-up questions or end the interaction within a short period, thereby indirectly obtaining feedback information. Further, the collected historical user feedback information undergoes natural language processing and semantic analysis. Through preset keywords and semantic rules, descriptions and evaluations of the historical response style are identified in the feedback information. For example, if the feedback information contains expressions such as "the tone of the answer is too stiff," "the expression is not easy to understand," or "the style is very friendly and professional," these are extracted as style evaluation information. Simultaneously, sentiment analysis is performed on these evaluations to determine whether they are positive, negative, or neutral, in order to more accurately understand the acceptance level of historical users towards the response style. Natural language processing technology is also used to filter evaluations related to historical responses from historical user feedback. For example, feedback mentioning "incomplete answer, missing key steps," "incorrect information, inconsistent with reality," or "detailed and accurate content" is extracted as content evaluation information. Further analysis of the specific aspects involved in the content evaluation information, such as completeness, accuracy, and practicality, provides a detailed basis for subsequent content adjustments. Furthermore, the historical response style is optimized based on the extracted style evaluation information. If there are many negative evaluations, particularly focusing on harsh tone, the response style parameters are adjusted to make the tone more gentle and friendly. If the style evaluation indicates that the expression is not easily understood, simpler vocabulary and sentence structures are introduced to reduce language complexity. For example, for responses to scientific and technological questions, which originally used many technical terms, adjustments are made based on feedback to appropriately increase the use of colloquial explanations, making it understandable for users with different knowledge levels. Meanwhile, taking into account the ratio of positive to negative feedback, appropriate adjustments are made while maintaining the original style and characteristics to avoid excessive changes that could lead to stylistic distortion. Simultaneously, the responses are modified based on the feedback information. If the feedback indicates an incomplete answer, the response is reviewed, and missing key information, such as operational steps and theoretical basis, is added. If errors are found, they are promptly verified and corrected. For content lacking practicality, optimization and simplification are implemented based on the actual needs and usage scenarios of past users. For example, when answering product usage questions, common troubleshooting methods and practical examples are added based on feedback information to improve the practicality and operability of the content.After adjusting the historical response style and content, the results are integrated to form a current response strategy for the target question. This current strategy clarifies the style characteristics and key content points of the response, ensuring that subsequent responses accurately and effectively meet the needs of the target user. For example, the current strategy specifies a friendly, professional, and easy-to-understand response style, with content covering comprehensive answers to the question, practical case analysis, and expansions on common questions. Finally, based on the determined current response strategy, the target question is answered, and feedback is provided to the target user through appropriate channels. Through these specific implementation methods, this invention can fully utilize historical user feedback, continuously optimize the response strategy, provide users with higher-quality and more tailored response services, and improve the performance and user experience of the intelligent question-and-answer system.

[0061] According to another intelligent question-answering method provided by the present invention, compared with the current method of generating answers solely based on the user's current question, the present invention receives user question information, determines the user's intent based on the user question information, and obtains multi-round historical question-and-answer information with the user's intent agreement graph; then, it determines the user's emotion based on the user question information, determines the response style for the user question information based on the user's intent and the user's emotion, and determines the response content for the user question information based on the user question information and the multi-round historical question-and-answer information; finally, it responds to the user question information based on the response style and the response content. Therefore, by comprehensively analyzing the user question information and the multi-round historical question-and-answer information with the user question information agreement graph, the present invention can comprehensively and accurately understand the user's intent, making the answers more targeted and practical; by analyzing the user's emotion and intent to determine the response style, it can adapt to the needs of different user emotional states, thereby enhancing the user's emotional identification and satisfaction.

[0062] Furthermore, as Figure 1 In specific implementation, embodiments of the present invention provide an intelligent question-answering device, such as... Figure 3 As shown, the device includes: an acquisition unit 31, a determination unit 32, and a response unit 33.

[0063] The acquisition unit 31 can be used to receive user question information, determine user intent based on the user question information, and acquire multi-round historical question and answer information that agrees with the user intent graph.

[0064] The determining unit 32 can be used to determine the user's emotion based on the user's question information, determine the response style for the user's question information based on the user's intention and the user's emotion, and determine the response content for the user's question information based on the user's question information and multi-round historical question and answer information.

[0065] The response unit 33 can be used to respond to the user's question information based on the response style and the response content.

[0066] In specific application scenarios, in order to determine the response content, such as Figure 4 As shown, the determining unit 32 includes a vector determining module 321, a fusion module 322, and a content determining module 323.

[0067] The vector determination module 321 can be used to determine the multimodal feature vector corresponding to the multimodal information when the user's question information is multimodal information, wherein the multimodal information includes at least two types of information such as text information, video information, image information, and voice information.

[0068] The fusion module 322 can be used to determine the vector magnitude of the multimodal feature vector, determine the vector weight of the multimodal feature vector based on the vector magnitude, and perform vector fusion on the multimodal feature vector based on the vector weight to obtain a multimodal fused feature vector.

[0069] The content determination module 323 can be used to determine the response content for the user's question information based on the multimodal fusion feature vector and multi-round historical question and answer information.

[0070] In specific application scenarios, to determine the response content, the content determination module 323 can be used to determine the historical question-and-answer feature vectors of multi-round historical question-and-answer information; determine the time interval between the question time of the user's question and the question-and-answer time of each historical question-and-answer information; determine the time weight of each historical question-and-answer feature vector based on the time interval; divide the multimodal feature vector and each historical question-and-answer feature vector into multiple subspaces; determine the subspace similarity of the multimodal feature vector and each historical question-and-answer feature vector in each subspace; and determine the subspace similarity of each subspace based on the subspace similarity. The spatial weights of each historical question-and-answer feature vector are determined, wherein the multiple subspaces include at least two of the following: entity subspace, intent subspace, and sentiment subspace. Based on the temporal weights and spatial weights, the vector weights of the corresponding historical question-and-answer feature vectors are determined, and each historical question-and-answer feature vector is weighted and fused based on the vector weights to obtain a historical question-and-answer fused feature vector. The multimodal fused feature vector and the historical question-and-answer fused feature vector are interactively processed, and the interaction processing result is input into a preset response model to predict the response content, thereby obtaining the response content for the user's question information.

[0071] In specific application scenarios, in order to obtain multi-round historical question-and-answer information that aligns with the user's intent agreement graph, the acquisition unit 31 includes a complexity determination module 311 and an acquisition module 312.

[0072] The complexity determination module 311 can be used to determine the intent complexity of the user intent and the average information complexity of all candidate historical question and answer information that agrees with the user intent graph.

[0073] The acquisition module 312 can be used to determine the information acquisition round of candidate historical question and answer information based on the intent complexity and the average information complexity, and store the candidate historical question and answer information of the corresponding information acquisition round in a cache pool, and acquire multi-round historical question and answer information that agrees with the user intent graph from the candidate historical question and answer information in the cache pool.

[0074] In specific application scenarios, in order to determine the response style, the content determination module 323 can also be used to determine the intent intensity of the user's intent and the emotional intensity of the user's emotion; based on the user's intent and its corresponding intent intensity, and the user's emotion and its corresponding emotional intensity, the response style for the user's question information is determined.

[0075] In specific application scenarios, in order to respond to subsequent user questions, the response unit 33 can also be used to determine the user feature data of multiple historical users who have engaged in historical intelligent question-and-answer interactions, and classify each historical user into different user categories based on the user feature data; in response to the input operation of the target user for the target question, the target user's target feature data is obtained, and based on the target feature data and the user feature data, the target user category to which the target user belongs is determined in different user categories, and target historical question information similar to the target question information is matched in the historical question information of historical users under the target user category; the target question information is responded to according to the historical response strategy of the target historical question information, wherein the historical response strategy includes historical response content and historical response style.

[0076] In specific application scenarios, in order to respond to subsequent user inquiries, the response unit 33 includes an information extraction module 331 and a response adjustment module 332.

[0077] The information extraction module 331 can be used to determine historical user feedback information that responded to the target historical question information, and extract style evaluation information for the historical response style and content evaluation information for the historical response content from the historical user feedback information.

[0078] The response adjustment module 332 can be used to adjust the style of the historical response based on the style evaluation information, adjust the content of the historical response based on the content evaluation information, determine the current response strategy for the target question based on the style adjustment result and the content adjustment result, and respond to the target question based on the current response strategy.

[0079] It should be noted that other corresponding descriptions of the functional modules involved in the intelligent question-answering device provided in this embodiment of the invention can be found in [reference]. Figure 1 The corresponding description of the method shown will not be repeated here.

[0080] Based on the above, Figure 1 Accordingly, this embodiment of the invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, performs the following steps: receiving user question information; determining user intent based on the user question information and obtaining multi-round historical question-and-answer information with agreement graphs with the user intent; determining user emotion based on the user question information; determining a response style for the user question information based on the user intent and the user emotion; determining response content for the user question information based on the user question information and the multi-round historical question-and-answer information; and responding to the user question information based on the response style and the response content.

[0081] Based on the above, Figure 1 The method shown and as Figure 3 The embodiment of the device shown in the invention also provides a physical structure diagram of a computer device, such as... Figure 5 As shown, the computer device includes: a processor 41, a memory 42, and a computer program stored in the memory 42 and executable on the processor. Both the memory 42 and the processor 41 are mounted on a bus 43. When the processor 41 executes the program, it performs the following steps: receiving user question information; determining user intent based on the user question information; and acquiring multi-round historical question-and-answer information that aligns with the user intent; determining user emotion based on the user question information; determining a response style for the user question information based on the user intent and the user emotion; determining response content for the user question information based on the user question information and the multi-round historical question-and-answer information; and responding to the user question information based on the response style and the response content.

[0082] The present invention, through its technical solution, receives user questions, determines user intent based on the questions, and obtains multi-round historical question-and-answer information consistent with the user intent. Then, it determines user emotion based on the questions, determines a response style based on the user intent and emotion, and determines response content based on the questions and historical question-and-answer information. Finally, it responds to the user questions based on the response style and content. This comprehensive analysis of user questions and historical question-and-answer information consistent with the user's intent allows for a more accurate and targeted understanding of user intent, resulting in more practical and relevant answers. Furthermore, by analyzing user emotions and intent to determine the response style, the invention adapts to the needs of users in different emotional states, thereby enhancing user emotional identification and satisfaction.

[0083] It is obvious to those skilled in the art that the modules or steps of the present invention described above can be implemented using general-purpose computing devices. They can be centralized on a single computing device or distributed across a network of multiple computing devices. Optionally, they can be implemented using computer-executable program code, thereby storing them in a storage device for execution by a computing device. In some cases, the steps shown or described can be performed in a different order than those presented herein, or they can be fabricated as separate integrated circuit modules, or multiple modules or steps can be fabricated as a single integrated circuit module. Thus, the present invention is not limited to any particular combination of hardware and software.

[0084] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. An intelligent question answering method, characterized by, include: Receive user question information, determine user intent based on the user question information, and obtain multi-round historical question and answer information with the agreement graph of the user intent; Based on the user's question information, determine the user's emotion; based on the user's intent and the user's emotion, determine the response style for the user's question information; based on the user's question information and multi-round historical question and answer information, determine the response content for the user's question information. The system responds to the user's question based on the response style and the response content.

2. The method of claim 1, wherein, The step of determining the response content for the user's question based on the user's question information and multi-round historical question-and-answer information includes: When the user's question information is multimodal information, determine the multimodal feature vector corresponding to the multimodal information, wherein the multimodal information includes at least two types of information selected from text information, video information, image information, and voice information; The vector magnitude of the multimodal feature vector is determined, the vector weight of the multimodal feature vector is determined based on the vector magnitude, and the multimodal feature vector is fused based on the vector weight to obtain a multimodal fused feature vector. Based on the multimodal fusion feature vector and multi-round historical question-and-answer information, the response content for the user's question is determined.

3. The method of claim 2, wherein, The step of determining the response content for the user's question based on the multimodal fusion feature vector and multi-round historical question-and-answer information includes: Determine the historical question-and-answer feature vector of multi-round historical question-and-answer information; The time interval between the question time of the user's question and the question and answer time of each historical question and answer is determined, and the time weight of each historical question and answer feature vector is determined based on the time interval. The multimodal feature vector and each historical question-and-answer feature vector are divided into multiple subspaces respectively. The subspace similarity of the multimodal feature vector and each historical question-and-answer feature vector in each subspace is determined respectively. Based on the subspace similarity, the spatial weight of each historical question-and-answer feature vector is determined respectively. The multiple subspaces include at least two of the following: entity subspace, intent subspace, and sentiment subspace. Based on the time weight and the spatial weight, the vector weight of the corresponding historical question and answer feature vector is determined, and based on the vector weight, each historical question and answer feature vector is weighted and fused to obtain a historical question and answer fused feature vector. The multimodal fusion feature vector and the historical question-and-answer fusion feature vector are interactively processed, and the interaction processing result is input into a preset response model to predict the response content, thereby obtaining the response content for the user's question information.

4. The method according to claim 1, characterized in that, The acquisition of multi-turn historical question-and-answer information related to the user intent agreement graph includes: Determine the intent complexity of the user intent and the average information complexity of all candidate historical question-and-answer information that agree with the user intent graph, respectively. Based on the intent complexity and the average information complexity, the information acquisition rounds of candidate historical question and answer information are determined, and the candidate historical question and answer information of the corresponding information acquisition rounds is stored in a cache pool. Multi-round historical question and answer information that agrees with the user intent graph is obtained from the candidate historical question and answer information in the cache pool.

5. The method according to claim 1, characterized in that, The step of determining the response style for the user's question based on the user's intent and the user's emotion includes: Determine the intensity of the user's intent and the intensity of the user's emotion; Based on the user's intent and its corresponding intent intensity, and the user's emotion and its corresponding emotion intensity, determine the response style for the user's question.

6. The method according to claim 1, characterized in that, The method further includes: Determine the user characteristic data of multiple historical users who have engaged in historical intelligent question-and-answer interactions, and classify each historical user into different user categories based on the user characteristic data; In response to the input operation of the target user for the target question information, the target user's target feature data is obtained. Based on the target feature data and the user feature data, the target user category to which the target user belongs is determined in different user categories. And the target historical question information that is similar to the target question information is matched in the historical question information of the historical users under the target user category. The target question information is responded to based on the historical response strategy of the target historical question information, wherein the historical response strategy includes historical response content and historical response style.

7. The method according to claim 6, characterized in that, The step of responding to the target question information based on the historical response strategy of the target historical question information includes: Determine historical user feedback information that responded to the target historical question information, and extract style evaluation information for the historical response style and content evaluation information for the historical response content from the historical user feedback information; Based on the style evaluation information, the historical response style is adjusted; based on the content evaluation information, the historical response content is adjusted; based on the style adjustment results and content adjustment results, the current response strategy for the target question is determined; and based on the current response strategy, the target question is answered.

8. An intelligent question-and-answer device, characterized in that, include: The acquisition unit is used to receive user question information, determine user intent based on the user question information, and acquire multi-round historical question and answer information that agrees with the user intent. The determining unit is configured to determine the user's emotion based on the user's question information, determine the response style for the user's question information based on the user's intent and the user's emotion, and determine the response content for the user's question information based on the user's question information and multi-round historical question and answer information. The response unit is used to respond to the user's question information based on the response style and the response content.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

10. A computer device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.