Intelligent question and answer method and device, equipment and medium

By extracting semantic features from natural language questions and combining them with large-scale model intent recognition, multimodal card information is constructed, solving the problem of single information presentation in intelligent question answering systems. This enables accurate cross-domain answers and visual displays, improving user experience and interaction efficiency.

CN121071103APending Publication Date: 2025-12-05CHINA PING AN PROPERTY INSURANCE CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511472301.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-14
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing intelligent question-answering systems lack multimodal interaction capabilities, making it impossible to achieve accurate answers and visual displays across different fields. This results in a single way of presenting information, requiring users to repeatedly search and manually extract information to obtain directly usable results.

Method used

By acquiring natural language questions input by users, extracting semantic feature information, using large models for intent recognition, determining multi-source data, constructing multimodal card information, and generating question-and-answer results that integrate text, charts, audio, or video, it supports flexible configuration and configurable features.

Benefits of technology

It enables users to quickly obtain intuitive, accurate, and personalized multimodal answers in insurance, healthcare, and financial service scenarios, adapting to different scenario needs and possessing scalability and continuous evolution capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121071103A_ABST
    Figure CN121071103A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, and discloses an intelligent question answering method and device, equipment and a medium, and the method comprises the steps: obtaining a natural language question input by a user, extracting semantic feature information based on the natural language question, and obtaining a semantic feature set; inputting the semantic feature set into a large model for intention recognition to obtain intention category information, and determining corresponding multi-source data in a preset domain knowledge base based on the intention category information; based on the intention category information and the multi-source data, constructing and obtaining multi-modal card information; wherein the multi-modal card information comprises unique identification information, model reasoning data and project side business interface data; and transmitting the multi-modal card information to a front-end page, and calling a corresponding multi-modal card component according to the unique identification information, so that the multi-modal card generates and displays a question and answer result based on the natural language question. The method can be applied to financial science and technology or medical care service program systems, and cross-domain accurate answering and visual display are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence, and in particular to an intelligent question answering method, device, equipment and medium. BACKGROUND

[0002] Artificial intelligence is increasingly widely used in key industries such as finance and medicine, and in particular in insurance sales, medical health consultation and financial service scenarios. A question answering system has become an important tool for improving efficiency and user experience. However, existing intelligent question answering systems are mainly based on single text output, lack multi-modal interaction capabilities, and result in a single information presentation method, requiring repeated searching and manual extraction by users to obtain directly usable results.

[0003] In addition, related technologies also lack a configurable mechanism, and cannot flexibly switch and customize according to different insurance types, regulatory regions, medical standards or financial rules, which seriously limits the practicality and expansibility of intelligent question answering. Therefore, there is an urgent need for an intelligent question answering method that can integrate multi-modal information such as text, charts, audio and video, implement accurate answers and visual display across domains, and improve the interactive experience and decision support value in insurance sales, medical health and financial services. SUMMARY

[0004] The present application provides an intelligent question answering method, device, equipment and medium to solve the technical problem that the question answering system in related technologies cannot integrate multi-modal information such as text, charts, audio and video, and thus cannot implement accurate answers and visual display across domains.

[0005] In a first aspect, an intelligent question answering method is provided, the method comprising: obtaining a natural language question input by a user, extracting semantic feature information based on the natural language question to obtain a semantic feature set; inputting the semantic feature set into a large model for intent recognition to obtain intent category information, and determining corresponding multi-source data in a preset domain knowledge base based on the intent category information; based on the intent category information and the multi-source data, constructing multi-modal card information; wherein the multi-modal card information includes unique identification information, model inference data and engineering side business interface data; transmitting the multi-modal card information to a front-end page, and calling corresponding multi-modal card components according to the unique identification information, so that the multi-modal card generates and displays a question answering result based on the natural language question; wherein the question answering result is an answer integrating at least one modality of text, chart, audio and video.

[0006] In a second aspect, an intelligent question answering device is provided, comprising: An acquisition module is configured to acquire a natural language question input by a user, extract semantic feature information based on the natural language question, and obtain a semantic feature set; An intent recognition module is configured to input the semantic feature set to a large model to perform intent recognition, obtain intent category information, and determine corresponding multi-source data in a preset domain knowledge base based on the intent category information; A construction module is configured to construct multi-modal card information based on the intent category information and the multi-source data, wherein the multi-modal card information includes unique identification information, model inference data, and engineering-side business interface data. A question and answer module is configured to transmit the multi-modal card information to a front-end page, call a corresponding multi-modal card component according to the unique identification information, and enable the multi-modal card to generate and display a question and answer result based on the natural language question, wherein the question and answer result is an answer that fuses at least one modality of text, a chart, audio, and a video.

[0007] In a third aspect, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the intelligent question and answer method when executing the computer program.

[0008] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the intelligent question and answer method when executed by a processor.

[0009] The scheme realized by the intelligent question and answer method, device, computer device and storage medium includes: obtaining a natural language question input by a user, extracting semantic feature information based on the natural language question to obtain a semantic feature set; inputting the semantic feature set to a large model for intent recognition to obtain intent category information, and determining corresponding multi-source data in a preset domain knowledge base based on the intent category information; based on the intent category information and the multi-source data, multi-modal card information is constructed; the multi-modal card information includes unique identification information, model reasoning data and engineering side business interface data; the multi-modal card information is transmitted to a front-end page, and a corresponding multi-modal card component is called according to the unique identification information, so that the multi-modal card generates and displays a question and answer result based on the natural language question; and the question and answer result is an answer that fuses at least one mode of text, chart, audio and video. In the present application, the natural language question is converted into a semantic feature set, and the intent recognition capability of the large model is combined to realize the calling of multi-source data of a preset domain knowledge base such as insurance, medical health and financial services, so that not only text but also multi-modal results such as charts, audios or videos can be dynamically generated in the question and answer process. Compared with the existing question and answer system with single text output, the present application can flexibly configure card components according to different scenes, automatically extract insurance clauses, visualize medical image interpretation or compare financial risk indicators, so that users can quickly obtain intuitive, accurate and personalized answers. At the same time, the multi-modal card component has the characteristics of scalability and configurability, and can be quickly adjusted according to regulatory rules or business needs to ensure the adaptability and continuous evolution capability in the question and answer process. BRIEF DESCRIPTION OF DRAWINGS

[0010] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.

[0011] Figure 1 is an application environment schematic diagram of the intelligent question and answer method in an embodiment of the present application; Figure 2 is a flow schematic diagram of the intelligent question and answer method in an embodiment of the present application; Figure 3 is Figure 2 is a specific implementation flow schematic diagram of step S10 in the embodiment; Figure 4 is a structure schematic diagram of the intelligent question and answer device in an embodiment of the present application; Figure 5 is a structure schematic diagram of the computer device in an embodiment of the present application; Figure 6Fig. 2 is another structural schematic diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0012] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of the present application.

[0013] The intelligent question and answer method provided by the embodiments of the present application can be applied in an application environment as shown in Fig. 1. Figure 1 In the application environment, a client communicates with a server through a network. The server can obtain a natural language question input by a user through the client, extract semantic feature information based on the natural language question to obtain a semantic feature set, input the semantic feature set into a large model to perform intent recognition to obtain intent category information, and determine corresponding multi-source data in a preset domain knowledge base based on the intent category information. The multi-modal card information is constructed based on the intent category information and the multi-source data. The multi-modal card information includes unique identification information, model reasoning data, and engineering side business interface data. The multi-modal card information is transmitted to a front-end page, and a corresponding multi-modal card component is called according to the unique identification information, so that the multi-modal card generates and displays a question and answer result based on the natural language question. The question and answer result is an answer that fuses at least one mode of text, chart, audio, and video. In the present application, the natural language question is converted into a semantic feature set, and the intent recognition capability of the large model is combined to realize the calling of multi-source data in a preset domain knowledge base such as insurance, medical health, and financial services, so that not only text but also multi-modal results such as charts, audios, or videos are dynamically generated in the question and answer process. Compared with the existing question and answer system that outputs single text, the present application can flexibly configure card components according to different scenes, automatically extract insurance clauses, visualize medical image interpretation, or compare financial risk indicators, so that users can quickly obtain intuitive, accurate, and personalized answers. At the same time, the multi-modal card component has the characteristics of scalability and configurability, and can be quickly adjusted according to regulatory rules or business needs to ensure the adaptability and continuous evolution capability in the question and answer process. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers, and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The present application will be described in detail through specific embodiments.

[0014] Please refer to Fig. 1, Figure 2 Figure 2 ​A flowchart of an intelligent question answering method provided by an embodiment of the present application is shown in the figure. The method comprises the following steps: S10: obtaining a natural language question input by a user, extracting semantic feature information based on the natural language question, and obtaining a semantic feature set.

[0015] It should be noted that the natural language question can be in the form of text, voice, or image, and the present application does not limit this.

[0016] For example, after being received through a unified input interface, the natural language question can first be preprocessed, converting voice into text, converting images into recognizable text or structured symbols, and performing word segmentation, error correction, and standardization processing on the text. Subsequently, the semantic parsing capability of a large model is used to extract features from the processed natural language question, identify keywords, key entities, syntax dependency relationships, and context logical structures therein, and perform semantic expansion in combination with a domain dictionary of a medical knowledge base, a financial rule base, or an insurance clause base. Finally, these semantic information is uniformly represented as a semantic feature set, providing basic data support for subsequent intent recognition and multi-modal card construction.

[0017] As shown in the figure, step S10, i.e., the step of extracting semantic feature information based on the natural language question and obtaining a semantic feature set, comprises the following steps: Figure 3 S11: performing standardization processing on the natural language question to obtain an input information set.

[0018] S12: sequentially performing semantic conversion operations on the input information set to obtain semantic representation information.

[0019] The semantic conversion operations comprise voice-to-text conversion, image recognition, and text error correction.

[0020] S13: extracting semantic feature information from the semantic representation information to obtain the semantic feature set.

[0021] The semantic feature information comprises keywords, entities, and context logical relationships.

[0022] For example, the natural language question can first be standardized. Specifically, after receiving the natural language question through a unified interface, format normalization and character encoding conversion can be performed for standardization processing, obtaining an input information set for subsequent recognition. In a medical scenario, the natural language question can include a symptom description, and the image input can be a medical image segment; in a financial scenario, the natural language question can include an investment consultation question, and the image can be a bill or contract segment. Through standardization processing, differences in processing caused by different input formats can be avoided, thereby forming a unified input information set. ​

[0023] Further, after obtaining the input information set, semantic conversion operations can be sequentially performed to obtain semantic representation information. The semantic conversion operations include converting speech into text, recognizing images into structured descriptions, and correcting text to ensure the completeness and accuracy of the input information.

[0024] On this basis, feature extraction can be performed on the semantic representation information to identify and extract keywords, key entities, and context logical relationships therein. For example, in the medical field, disease names, symptoms, and drug entities can be identified, and in the financial field, investment varieties, risk levels, and policy provisions can be identified. Finally, these extracted semantic features are unified into a semantic feature set, laying a foundation for subsequent intent recognition and multi-modal card construction.

[0025] S20: inputting the semantic feature set into a large model for intent recognition to obtain intent category information, and determining corresponding multi-source data in a preset domain knowledge base based on the intent category information.

[0026] For example, after receiving the semantic feature set, the large model first comprehensively analyzes the keywords, entities, and context logic using deep semantic understanding capabilities to identify the core intent category of the user's question, such as insurance consultation, medical diagnosis assistance, or financial investment analysis. Subsequently, based on the identified intent category, the corresponding preset domain knowledge base, including the insurance clause library, the medical health guideline library, and the financial regulation rule library, can be called to filter and determine multi-source data matching the intent category. Thus, the user's input question can be quickly mapped to a precise business scenario, and the knowledge base can provide authoritative, up-to-date, and multi-angle information support, providing a basis for subsequent multi-modal card construction.

[0027] In some embodiments, inputting the semantic feature set into the large model for intent recognition to obtain intent category information includes: inputting the semantic feature set into the large model to obtain a candidate intent set; comparing the candidate intent set with a preset domain knowledge base through the large model to obtain a target intent category; calling an external data interface to determine supplementary data corresponding to the target intent category, and determining the supplementary data as external enhancement data; and fusing the external enhancement data with the target intent category to obtain the intent category information.

[0028] For example, the semantic feature set can be first input to the large model to obtain a candidate intent set. The large model generates multiple possible intent categories through comprehensive analysis of keywords, entities, and context logic. For example, if the user input is "Please help me compare the differences in protection between two critical illness insurance", the candidate intents can include "insurance clause comparison" and "risk assessment"; in the medical scenario, if the input is "cough accompanied by chest pain", the candidate intents can be "symptom diagnosis" and "image examination suggestion". Subsequently, the large model further compares the candidate intent set with the preset domain knowledge base to filter out the target intent category with the highest matching degree, ensuring accurate positioning of the core problem that the user is really concerned about.

[0029] Further, after obtaining the target intent category, an external data interface can be called to determine the supplementary data corresponding to the category and use it as external enhancement data. For example, in the financial field, real-time market quotations or regulatory announcements can be obtained through the interface; in the medical field, the latest diagnosis and treatment specifications can be obtained by calling an external guideline database. Finally, the external enhancement data and the target intent category can be fused to generate complete intent category information. This not only ensures the authority of the result, but also enhances the timeliness of the information, so that the subsequent multi-modal card can simultaneously reflect static knowledge and dynamic data when displaying the question and answer results, thereby improving the practical value of the question and answer process in the medical health and financial business scenarios.

[0030] S30: Based on the intent category information and the multi-source data, a multi-modal card information is constructed.

[0031] The multi-modal card information includes unique identification information, model inference data, and engineering-side business interface data.

[0032] In some embodiments, based on the intent category information and the multi-source data, the multi-modal card information is constructed, including: determining the unique representation information according to the intent category information and generating an initial card framework; injecting the inference data of the large model into the initial card framework to obtain the model inference data; determining the engineering-side business interface data corresponding to the multi-source data, and fusing the engineering-side business interface data and the model inference data to obtain a fusion data structure; introducing a preset configured component template into the fusion data interface to obtain the multi-modal card information.

[0033] Exemplarily, first, unique identification information can be determined according to the intention category information, and an initial card frame is generated. The unique identification information is used to indicate a specific business scenario and display requirements, such as comparison of insurance types in insurance sales, symptom interpretation in medical health, or risk assessment in financial services. Subsequently, the semantic understanding result obtained by the large model reasoning is injected into the initial card frame to obtain model reasoning data, so that the card has a deep semantic response capability to the user's question. In this process, the model reasoning data can reflect the core answer logic, such as the difference in guarantee responsibility, drug use recommendation, or investment risk trend.

[0034] Further, after obtaining the model reasoning data, engineering-side business interface data corresponding to the multi-source data can be determined and fused with the model reasoning data to obtain a fusion data structure. The engineering-side business interface data can come from an insurance clause library interface, a medical image analysis interface, or a financial regulatory information interface, which is introduced to ensure that the card information not only contains reasoning results, but also integrates authoritative and real-time business data. On this basis, the system introduces a preset configured component template in the fusion data structure, so that different types of answers are intuitively displayed in the form of text, charts, or audio and video. The finally generated multi-modal card information can not only adapt to the diversified needs of the medical health and financial insurance fields, but also provide a unified data structure and display logic for subsequent front-end rendering and display of the question and answer results.

[0035] S40: transmit the multi-modal card information to a front-end page, and call a corresponding multi-modal card component according to the unique identification information, so that the multi-modal card generates and displays a question and answer result based on the natural language question; wherein the question and answer result is an answer that fuses at least one modality of text, charts, audio, and video.

[0036] In some embodiments, the transmitting the multi-modal card information to the front-end page and calling the corresponding multi-modal card component according to the unique identification information to make the multi-modal card generate and display a question and answer result based on the natural language question includes: transmitting the multi-modal card information to the front-end page and calling the multi-modal card component according to the unique identification information to obtain a set of to-be-rendered components; binding the set of to-be-rendered components with the multi-modal card information to generate a front-end renderable object; performing rendering processing on the front-end renderable object, and making the multi-modal card generate and display a question and answer result based on the natural language question.

[0037] For example, the backend-generated multi-modal card information can be first completely transmitted to the front-end interface, and a card component corresponding to the unique identification information is called based on the unique identification information to obtain a set of components to be rendered. The unique identification information can ensure that different intent categories call matched display templates, such as calling a medical record interpretation or medication guide component in a medical health scenario, or calling a yield comparison or risk warning component in a financial scenario. Through this matching method, the front end can quickly locate the required functional module to ensure that the display logic is highly consistent with the user's question.

[0038] After obtaining the set of components to be rendered, the components can be further bound with the corresponding multi-modal card information to generate a rendering object that can be directly displayed by the front end. Subsequently, the front-end renderable object is subjected to rendering processing to generate a final question and answer result. The question and answer result can not only be directly answered in the form of text, but also can be multi-modal displayed by fusing charts, audio and video content, such as displaying a lesion position through image annotation in a medical health scenario, or displaying a yield fluctuation trend through a dynamic chart in a financial field. Through this process, the user can obtain intuitive, complete and personalized multi-modal answers based on natural language questions, which significantly improves the interactive experience and information utilization value of the question and answer system.

[0039] In some embodiments, the method further comprises: obtaining a training data set and a pre-trained model; wherein the training data set comprises a plurality of historical semantic feature sets; labeling the training data set to obtain a labeled result, wherein the labeled result comprises historical intent category information corresponding to the historical semantic feature set; training the pre-trained model based on the training data set and the labeled result to obtain the large model.

[0040] Based on the above embodiments, after obtaining the large model, the method comprises: iteratively training the large model based on the training data set and the labeled result to extract data features, and calculating the classification loss function; iteratively training the classification loss function using a preset method to reduce the value of the classification loss function until the value of the classification loss function is less than an expected threshold; and obtaining the large model after iteration based on the classification loss function after iteration.

[0041] Specifically, a training data set containing a plurality of historical semantic feature sets can be collected for training. For example, the training data set can be obtained by manual collection, web crawling or public data set, and the like, which are not limited in the present application.

[0042] Further, each of the plurality of historical semantic feature sets can be labeled to obtain a label corresponding to each of the plurality of historical semantic feature sets. Then, the label is taken as a label of the group of input data, and each group of training data set carrying the label is input into the pre-trained model for supervised learning. When a training end condition is met, such as a number of training times reaching a threshold or an output accuracy of the model reaching an accuracy threshold, the training is ended, and a trained large model is obtained.

[0043] In the embodiments of the present application, the training data set and the label are input into the pre-trained model for supervised learning, and then a large model is trained. Thus, the label can be output based on the large model.

[0044] The above embodiments can enhance data quality and diversity during training of the large model, and improve generalization ability and actual application effect of the model.

[0045] It can be understood that, in order to train a large model with higher accuracy, the large model can be iteratively trained in a manner such that the classification loss function is continuously reduced until the classification loss function meets a predetermined expected threshold, and then a more accurate label can be obtained based on the large model after iteration.

[0046] It should be noted that the present application does not limit the above-mentioned preset method and expected threshold, for example, the preset method can be a gradient descent algorithm, a batch gradient descent algorithm, a stochastic gradient descent algorithm, etc., and the present application takes the gradient descent algorithm as an example for illustration.

[0047] The purpose of the gradient descent algorithm is to find the minimum value of the classification loss function by iteration, or to converge to the minimum value. Geometrically, the gradient descent algorithm is to find the minimum value of the function by moving in the opposite direction of the vector with the fastest increase in the function, so that the gradient decreases the fastest and the function minimum value is easier to find. Based on this, in the embodiments of the present application, the gradient descent algorithm is used to iteratively train the large model to continuously reduce the classification loss function, thereby reducing the error of the calculation result.

[0048] In the embodiments of the present application, the gradient descent algorithm is used to iteratively train the large model to continuously reduce the classification loss function, so as to obtain the large model after iteration, and then a more accurate label can be obtained based on the large model after iteration.

[0049] As can be seen, the above solution transforms natural language questions into a set of semantic features and combines this with the intent recognition capabilities of a large model to access multi-source data from pre-defined knowledge bases in areas such as insurance, healthcare, and financial services. This allows for the output of not only text during the question-and-answer process but also the dynamic generation of multimodal results such as charts, audio, or video. Compared to existing question-and-answer systems that only output text, this invention can flexibly configure card components according to different scenarios, automatically extracting insurance terms, visualizing medical image interpretations, or comparing financial risk indicators, enabling users to quickly obtain intuitive, accurate, and personalized answers. Furthermore, the multimodal card components are scalable and configurable, allowing for rapid adjustments based on regulatory rules or business needs, ensuring adaptability and continuous evolution during the question-and-answer process.

[0050] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0051] In one embodiment, an intelligent question-answering device is provided, which corresponds one-to-one with the intelligent question-answering methods described in the above embodiments. For example... Figure 4 As shown, the intelligent question-answering device includes an acquisition module 101, an intent recognition module 102, a construction module 103, and a question-answering module 104. Detailed descriptions of each functional module are as follows: The acquisition module 101 is used to acquire the natural language question input by the user, extract semantic feature information based on the natural language question, and obtain a semantic feature set; The intent recognition module 102 is used to input the semantic feature set into the large model for intent recognition, obtain intent category information, and determine the corresponding multi-source data in the preset domain knowledge base based on the intent category information; The construction module 103 is used to construct multimodal card information based on the intent category information and the multi-source data; wherein, the multimodal card information includes unique identification information, model inference data, and engineering-side business interface data; The question-and-answer module 104 is used to transmit the multimodal card information to the front-end page and call the corresponding multimodal card component according to the unique identifier information, so that the multimodal card generates and displays the question-and-answer results based on the natural language question; wherein the question-and-answer results are answers that integrate at least one modality of text, charts, audio and video.

[0052] The acquisition module 101 is configured to perform standardization processing on the natural language question to obtain an input information set; sequentially perform semantic conversion operations on the input information set to obtain semantic representation information; the semantic conversion operations include voice-to-text conversion, image recognition, and text correction; extract semantic feature information in the semantic representation information to obtain a semantic feature set; the semantic feature information includes keywords, entities, and context logical relationships.

[0053] The intent recognition module 102 is configured to input the semantic feature set into the large model to obtain a candidate intent set; compare the candidate intent set with a preset domain knowledge base through the large model to obtain a target intent category; call an external data interface to determine supplementary data corresponding to the target intent category, and determine the supplementary data as external enhancement data; fuse the external enhancement data and the target intent category to obtain the intent category information.

[0054] The construction module 103 is configured to determine the unique representation information according to the intent category information, and generate an initial card framework; inject inference data of the large model into the initial card framework to obtain the model inference data; determine engineering-side business interface data corresponding to the multi-source data, and fuse the engineering-side business interface data and the model inference data to obtain a fusion data structure; introduce a preset configured component template into the fusion data interface to obtain the multi-modal card information.

[0055] The question and answer module 104 is configured to transmit the multi-modal card information to the front-end page, and call the multi-modal card component according to the unique representation information to obtain a set of to-be-rendered components; bind the set of to-be-rendered components and the multi-modal card information to generate a front-end renderable object; perform rendering processing on the front-end renderable object, and enable the multi-modal card to generate and display a question and answer result based on the natural language question.

[0056] In an embodiment, the acquisition module 101 is further configured to acquire a training data set and a pre-training model; the training data set includes a plurality of historical semantic feature sets; label the training data set to obtain a labeling result, wherein the labeling result includes historical intent category information corresponding to the historical semantic feature sets; train the pre-training model through the training data set and the labeling result to obtain the large model.

[0057] In an embodiment, the acquisition module 101 is further configured to: based on the training data set and the labeling result, iteratively train the large model to extract data features and calculate the classification loss function; iteratively train the classification loss function by using a preset method to reduce the value of the classification loss function until the value of the classification loss function is less than an expected threshold; and based on the iteratively trained classification loss function, obtain the iteratively trained large model.

[0058] The present application provides an intelligent question answering device, which converts natural language questions into a semantic feature set and combines the intention recognition capability of a large model to realize multi-source data calling of a preset domain knowledge base in insurance, medical health and financial services, etc., so as to not only output text but also dynamically generate multi-modal results such as charts, audio and video in the question answering process. Compared with the existing single text output question answering system, the present application can flexibly configure card components according to different scenarios, automatically extract insurance clauses, visualize medical image interpretation or compare financial risk indicators, so that users can quickly obtain intuitive, accurate and personalized answers. At the same time, the multi-modal card component has the characteristics of extensibility and configurability, which can be quickly adjusted according to regulatory rules or business needs to ensure adaptability and continuous evolution capability in the question answering process.

[0059] The specific limitations of the intelligent question answering device can be referred to the limitations of the intelligent question answering method in the above, which will not be repeated here. Each module in the above intelligent question answering device can be realized by software, hardware and their combination in whole or in part. The above modules can be embedded in or independent of the processor in the computer device in hardware form, or can be stored in the memory in the computer device in software form, so as to be called and executed by the processor to perform the operations corresponding to each module.

[0060] In one embodiment, a computer device is provided, which can be a server, and its internal structure diagram can be as shown in Figure 5 The computer device includes a processor, a memory, a network interface and a database connected by a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with the external client through the network connection. The computer program is executed by the processor to realize the functions or steps of the intelligent question answering method server side.

[0061] In one embodiment, a computer device is provided, which can be a client, and its internal structure diagram can be as shown in Figure 6As shown. The computer device includes a processor, a memory, a network interface, a display screen and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium, an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operating system and the computer program in the non-volatile storage medium to run. The network interface of the computer device is used to communicate with the external server through the network connection. The computer program is executed by the processor to realize the function or step of the client side of the intelligent question and answer method.

[0062] In one embodiment, a computer device is provided, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, the processor executing the computer program to implement the following steps: Obtaining a natural language question input by a user, extracting semantic feature information based on the natural language question to obtain a semantic feature set; Inputting the semantic feature set into a large model for intent recognition to obtain intent category information, and determining corresponding multi-source data in a preset domain knowledge base based on the intent category information; Based on the intent category information and the multi-source data, a multi-modal card information is constructed; wherein the multi-modal card information includes unique identification information, model reasoning data and engineering side business interface data; Transmitting the multi-modal card information to a front-end page, and calling a corresponding multi-modal card component according to the unique identification information, so that the multi-modal card generates and displays a question and answer result based on the natural language question; wherein the question and answer result is an answer that fuses at least one modality of text, chart, audio and video.

[0063] In one embodiment, a computer readable storage medium is provided, having a computer program stored thereon, the computer program being executed by a processor to implement the following steps: Obtaining a natural language question input by a user, extracting semantic feature information based on the natural language question to obtain a semantic feature set; Inputting the semantic feature set into a large model for intent recognition to obtain intent category information, and determining corresponding multi-source data in a preset domain knowledge base based on the intent category information; Based on the intent category information and the multi-source data, a multi-modal card information is constructed; wherein the multi-modal card information includes unique identification information, model reasoning data and engineering side business interface data; The multi-modal card information is transmitted to a front-end page, and a corresponding multi-modal card component is called according to the unique identification information, so that the multi-modal card generates and displays a question and answer result based on the natural language question; wherein the question and answer result is an answer fusing at least one mode of text, chart, audio and video.

[0064] It should be noted that the functions or steps described above with respect to the computer-readable storage medium or the computer device can correspond to the related descriptions of the server side and the client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0065] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiments can be completed by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer readable storage medium. When the computer program is executed, it can include the processes of the above-mentioned embodiments. Any reference to memory, storage, database or other medium used in the embodiments provided by the present application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0066] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above-mentioned functional units and modules is exemplified. In actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above.

[0067] The above-described embodiments are only used to illustrate the technical solutions of the present application, and are not intended to limit the present application; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.

Claims

1. An intelligent question answering method, characterized by, The method comprises: acquiring a natural language question input by a user, extracting semantic feature information based on the natural language question to obtain a semantic feature set; inputting the semantic feature set into a large model for intent recognition to obtain intent category information, and determining corresponding multi-source data in a preset domain knowledge base based on the intent category information; based on the intent category information and the multi-source data, constructing multi-modal card information; wherein the multi-modal card information comprises unique identification information, model inference data and engineering side business interface data; transmitting the multi-modal card information to a front-end page, and calling a corresponding multi-modal card component according to the unique identification information, so that the multi-modal card generates and displays a question and answer result based on the natural language question; wherein the question and answer result is an answer that fuses at least one modality of text, chart, audio and video.

2. The method of claim 1, wherein, The method comprises: standardizing the natural language question to obtain an input information set; performing semantic conversion operations on the input information set in sequence to obtain semantic representation information; wherein the semantic conversion operations comprise voice-to-text conversion, image recognition and text correction operations; extracting semantic feature information from the semantic representation information to obtain the semantic feature set; wherein the semantic feature information comprises keywords, entities and context logical relationships.

3. The method of claim 1, wherein, The method comprises: inputting the semantic feature set into the large model to obtain a candidate intent set; comparing the candidate intent set with a preset domain knowledge base through the large model to obtain a target intent category; calling an external data interface to determine supplementary data corresponding to the target intent category, and determining the supplementary data as external enhancement data; fusing the external enhancement data with the target intent category to obtain the intent category information.

4. The method of claim 1, wherein, The method comprises: determining the unique representation information according to the intent category information and generating an initial card framework; injecting inference data of the large model into the initial card framework to obtain the model inference data; determining engineering side business interface data corresponding to the multi-source data, and fusing the engineering side business interface data with the model inference data to obtain a fusion data structure; introducing a preset configured component template into the fusion data interface to obtain the multi-modal card information.

5. The method of claim 1, wherein, The method comprises: transmitting the multi-modal card information to the front-end page, and calling the multi-modal card component according to the unique representation information to obtain a set of components to be rendered; binding the set of components to be rendered with the multi-modal card information to generate a front-end renderable object; The front-end renderable object is rendered, and the multi-modal card is caused to generate and display a question and answer result based on the natural language question.

6. The method of claim 1, wherein, The method further includes: obtaining a training data set and a pre-trained model; wherein the training data set includes a plurality of historical semantic feature sets; annotating the training data set to obtain an annotation result, wherein the annotation result includes historical intent category information corresponding to the historical semantic feature set; training the pre-trained model based on the training data set and the annotation result to obtain the large model.

7. The method of claim 6, wherein, After obtaining the large model, the method includes: iteratively training the large model based on the training data set and the annotation result to extract data features and calculate the classification loss function; iteratively training the classification loss function using a preset method to reduce the value of the classification loss function until the value of the classification loss function is less than an expected threshold value; obtaining an iteratively trained large model based on the iteratively trained classification loss function.

8. An intelligent question answering apparatus, characterized by comprising: includes: an obtaining module configured to obtain a natural language question input by a user, extract semantic feature information based on the natural language question, and obtain a semantic feature set; an intent recognition module configured to input the semantic feature set to a large model for intent recognition, obtain intent category information, and determine corresponding multi-source data in a preset domain knowledge base based on the intent category information; a constructing module configured to construct multi-modal card information based on the intent category information and the multi-source data; wherein the multi-modal card information includes unique identification information, model inference data, and engineering side business interface data; a question and answer module configured to transmit the multi-modal card information to a front-end page, call a corresponding multi-modal card component according to the unique identification information, and cause the multi-modal card to generate and display a question and answer result based on the natural language question; wherein the question and answer result is an answer that fuses at least one modality of text, a chart, audio, and a video.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to implement the steps of the intelligent question and answer method of any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 9. The computer program is executed by the processor to implement the steps of the intelligent question and answer method of any one of claims 1 to 7.

Citation Information

Cited By

  • Business request processing method, electronic equipment and program product

    CN121350231A