Intelligent question answering method and device based on large model, medium, equipment and product
Through the classification model, user questions are classified at a level, and the Q&A big model with the smallest parameter volume is selected to generate answers, which solves the problem of high demand for large model computing resources and achieves the improvement of user experience of efficient response and accuracy retention.
Patent Information
- Application Number
- CN202510887829.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-07-29
AI Technical Summary
The huge amount of parameters of the large model leads to a high demand for computing resources, affecting the response speed and user experience. Although the existing technology improves efficiency through model quantization or cache systems, it loses accuracy or occupies storage resources.
Through the classification model, classify user questions, select the Q&A big model with the smallest parameter quantity to generate answers, avoid calling large models with large scale parameters and reduce concurrency.
While ensuring the accuracy of answers, it improves the processing efficiency and response speed of the large model, reduces the waiting time for users, and improves the user experience.
Smart Images

Figure CN120386850A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of large models, agents, artificial intelligence, and computer technology. Specifically, it relates to an intelligent question - answering method, device, medium, equipment, and product based on a large model. Background Art
[0002] With the rapid development of generative large - model technology, more and more applications and software are connected to large models. Generally speaking, the larger the number of parameters of a large model, the better its performance, for example, the higher the accuracy of answering questions.
[0003] However, due to the huge number of parameters, the storage and computing resources required for the operation of large models are extremely large. To improve the response speed and answering efficiency of large models, in related technologies, the processing efficiency of large models is usually improved by model quantization or building a caching system. However, model quantization will lose model accuracy, resulting in a decrease in the accuracy of answering questions, and the caching system will occupy storage resources and the computational amount gradually increases with the increase of the cache, thus reducing the speed and efficiency of the output answer. Summary of the Invention
[0004] This content is provided to briefly introduce concepts that will be described in detail in the subsequent detailed implementation section. This content is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, the present disclosure provides an intelligent question - answering method based on a large model. The intelligent question - answering method includes: Display an intelligent interaction page, which is used to answer questions sent by users through a question - answering large model. The question - answering large model includes multiple question - answering large models with different parameter scales, and one question - answering large model corresponds to one level classification; In response to a target question sent by a first user in the intelligent interaction page, classify the target question through a classification model, determine the target level corresponding to the target question, and determine the first question - answering large model with the smallest parameter scale among the multiple question - answering large models that can answer questions of the target level; Generate a target answer corresponding to the target question through the first question - answering large model.
[0006] In a second aspect, the present disclosure provides an intelligent question - answering device based on a large model. The intelligent question - answering device includes: A display module, configured to display an intelligent interaction page, which is used to answer questions sent by users through a question - answering large model. The question - answering large model includes multiple question - answering large models with different parameter scales, and one question - answering large model corresponds to one level classification; A classification module, configured to classify a target question in response to a target question sent by a first user in the intelligent interaction page through a classification model, determine a target level corresponding to the target question, and determine a first question-answering large model that can answer questions of the target level and has the smallest parameter scale among the multiple question-answering large models; A generation module, configured to generate a target answer corresponding to the target question through the first question-answering large model.
[0007] In a third aspect, the present disclosure provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the method described in the first aspect are implemented.
[0008] In a fourth aspect, the present disclosure provides an electronic device, including: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method described in the first aspect.
[0009] In a fifth aspect, the present disclosure provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method described in the first aspect are implemented.
[0010] Through the above technical solution, in response to a target question sent by a first user in the intelligent interaction page, the target question is classified through a classification model, the target level corresponding to the target question is determined, and a first question-answering large model that can answer questions of the target level and has the smallest parameter scale is determined among multiple question-answering large models, and then a target answer corresponding to the target question is generated through the first question-answering large model. By using this method, the level corresponding to the user's question is determined through a classification model, so as to select a question-answering large model that can answer questions of this level and has the smallest parameter scale among multiple question-answering large models to generate an answer, that is, while ensuring the accuracy of the answer, avoiding all questions from calling large models with large parameter scales, thereby reducing the concurrency of large models with large parameter scales, improving the processing efficiency of the large model, further improving the response speed and answering efficiency of the large model, reducing the user's waiting time, and enhancing the user experience.
[0011] Other features and advantages of the present disclosure will be described in detail in the subsequent specific implementation part. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Combined with the drawings and referring to the following specific implementation manners, the above and other features, advantages and aspects of each embodiment of the present disclosure will become more obvious. Throughout the drawings, the same or similar reference numerals represent the same or similar elements. It should be understood that the drawings are schematic, and the original components and elements are not necessarily drawn to scale. In the drawings: Figure 1It is a schematic diagram of the process of large model question answering shown according to an exemplary embodiment of the present disclosure; Figure 2 It is a schematic diagram of model quantization shown according to an exemplary embodiment of the present disclosure; Figure 3 It is a schematic diagram of the question answering process based on a cache system shown according to an exemplary embodiment of the present disclosure; Figure 4 It is a schematic flowchart of an intelligent question answering method based on a large model shown according to an exemplary embodiment of the present disclosure; Figure 5 It is a schematic framework diagram of large model processing based on a classification model shown according to an exemplary embodiment of the present disclosure; Figure 6 It is a schematic diagram of training samples of a classification model shown according to an exemplary embodiment of the present disclosure; Figure 7 It is a schematic diagram of the process of sample annotation shown according to an exemplary embodiment of the present disclosure; Figure 8 It is a schematic diagram of the process of model processing for annotating a large model shown according to an exemplary embodiment of the present disclosure; Figure 9 It is a schematic structural diagram of a classification model shown according to an exemplary embodiment of the present disclosure; Figure 10 It is a structural block diagram of an intelligent question answering device based on a large model shown according to an exemplary embodiment of the present disclosure; Figure 11 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment of the present disclosure. Detailed implementation manners
[0013] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.
[0014] It should be understood that the steps recited in the method embodiments of the present disclosure can be executed in a different order and / or in parallel. In addition, the method embodiments may include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this regard.
[0015] As used herein, the term "including" and its variations are open-ended, i.e., "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". The relevant definitions of other terms will be given in the following description.
[0016] It should be noted that the concepts such as "first", "second", etc. mentioned in this disclosure are only used to distinguish different devices, modules or units, and are not used to limit the order or interdependence of the functions performed by these devices, modules or units.
[0017] It should be noted that the modifications of "one" and "multiple" mentioned in this disclosure are illustrative rather than restrictive. Those skilled in the art should understand that, unless otherwise clearly specified in the context, it should be understood as "one or more".
[0018] The names of the messages or information exchanged between multiple devices in the embodiments of this disclosure are only for illustrative purposes and are not used to limit the scope of these messages or information.
[0019] It can be understood that before using the technical solutions disclosed in the embodiments of this disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0020] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that performs the operations of the technical solutions of this disclosure according to the prompt message.
[0021] As an optional but non-limiting implementation manner, the way of sending a prompt message to the user in response to receiving an active request from the user can be, for example, in the form of a pop-up window, and the prompt message can be presented in text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0022] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not constitute a limitation on the implementation manner of this disclosure. Other ways that meet relevant laws and regulations can also be applied to the implementation manner of this disclosure.
[0023] Meanwhile, it can be understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of data) should comply with the requirements of relevant laws, regulations and related provisions.
[0024] With the rapid development of generative large language model technology, more and more applications and software are connected to large models. Large models follow the Scaling Law, that is, the larger the number of parameters of a large model, the better its performance, such as more accurate answers and better meeting user needs.
[0025] However, large models have extremely high requirements for storage and computing resources. Whether it is a large model service deployed privately or a large model service provided by cloud providers, due to limited computing resources, the inference efficiency of the model is low and cannot meet the usage requirements of the production environment.
[0026] Take the Figure 1 question-and-answer scenario based on large models shown as an example. A large model service refers to deploying a large model inference service on computing resources, generally an HTTP (HyperText Transfer Protocol) service, and providing an API (Application Programming Interface) externally. Users send query requests through the specified service IP (Internet Protocol) and port. The large model service obtains the answer to the question through inference and then returns it to the user. The more complex the question, the longer it takes for the large model to generate an answer.
[0027] This is because the number of parameters of large models is huge. Even when computing resources are sufficient, due to the token-by-token inference process of autoregressive models, the actual inference efficiency is not high. In actual usage scenarios, a large model service supports multiple users, that is, concurrent use of the large model service. Therefore, when there are too many users, the large model service will experience problems such as lag, waiting, and slow response, which will affect the user experience.
[0028] In related technologies, post-quantization is used to reduce the size of large models, and the large models perform inference with lower precision, thereby improving the processing efficiency of the models. Such as Figure 2As shown, quantize the BF16 (Brain Floating Point 16-bit) large model into In4 (Integer 4-bit) and Int8, which can effectively reduce the video memory occupancy during the deployment and inference of the large model, improve the concurrency rate of the large model service, and to a certain extent, enhance the inference efficiency of the model. However, model quantization will largely lose the accuracy of the large model. Compared with the original large model, the question-and-answer effect of the model will deteriorate, and the accuracy of the answers will decrease, thus affecting the user experience. In addition, in usage scenarios with high precision requirements, the priority of precision is generally higher than that of performance. Therefore, this method cannot be applied to all scenarios.
[0029] Alternatively, by constructing a caching system to cache the questions queried by users and the answers output by the large model, when the user inputs the same or similar questions, directly obtain the corresponding answers from the caching system for output, saving the inference calculation of the large model, thereby improving the response speed and answering efficiency. The architecture of the large model cache is as Figure 3 shown. The questions input by the user are successively calculated for semantic similarity with the questions in the cache. Here, a pre-trained representation model can be used to extract feature vectors. After calculating the semantic similarity, filter according to the similarity threshold, and judge whether the questions in the cache are hit based on the filtered results. If the filtered result is empty, it means that no similar or identical questions are matched, and the large model service can be called for inference, and the generated answers and corresponding questions are inserted into the caching system to update the caching system. If similar or identical questions are matched, an answer filtered out is output, such as the answer corresponding to the question with the highest similarity.
[0030] The above caching system needs to occupy a large amount of storage resources to build a cache index. In addition, for the user's input, it is necessary to calculate the similarity with the cache index in turn to determine whether it hits. Therefore, this method has high requirements for storage resources, and it takes a certain amount of time to process the calculation logic of cache hits. And as the caching system gradually expands, the computational complexity of similarity calculation will also become larger and larger. And if the cache is not hit, this part of the calculation time is additional waiting time for the user.
[0031] In view of this, the embodiments of the present disclosure provide an intelligent question-and-answer method, device, medium, device and product based on a large model to solve the above technical problems.
[0032] The following further explains and illustrates the embodiments of the present disclosure in conjunction with the accompanying drawings.
[0033] Figure 4 is a flowchart of an intelligent question-and-answer method based on a large model shown according to an exemplary embodiment of the present disclosure. Referring to Figure 4 , the intelligent question-and-answer method may include the following steps: S401: Display an intelligent interaction page.
[0034] Among them, the intelligent interaction page is used to answer questions sent by users through a question-and-answer large model. The question-and-answer large model includes multiple question-and-answer large models with different parameter scales, and one question-and-answer large model corresponds to one level classification.
[0035] Exemplarily, with the development of large model technology, more and more applications and software are connected to intelligent objects associated with large models. In response to the user triggering the entrance of the intelligent object displayed on the page, the intelligent interaction page is displayed. The user can interact with the intelligent object on the intelligent interaction page. For example, in a question-and-answer scenario, the question-and-answer large model can be called based on the question sent by the user to generate a corresponding answer and display it on the intelligent interaction page.
[0036] Exemplarily, the level of the question-and-answer large model corresponds one-to-one with the level of the question, indicating that the question-and-answer large model of this level is the one with the smallest scale of parameter quantity that can handle questions of this level.
[0037] S402: In response to the target question sent by the first user on the intelligent interaction page, classify the target question through a classification model, determine the target level corresponding to the target question, and determine the first question-and-answer large model with the smallest parameter scale that can answer questions of the target level among multiple question-and-answer large models.
[0038] Exemplarily, the questions can be classified from the perspective of the difficulty of answering by the large model. The questions input by the user are classified through the trained classification model, so as to use large models with corresponding scale of parameter quantity for different levels of questions for model inference. As much as possible, without reducing the answer accuracy, select the large model with the smallest parameter quantity from the large models that can solve the problem for inference, which not only improves the inference efficiency, but also reduces the inference concurrency of large models with large or even extremely large parameter quantities, not only improving the user experience, but also improving the throughput rate of the large model service under the same computing resources.
[0039] S403: Generate a target answer corresponding to the target question through the first question-and-answer large model.
[0040] Such as Figure 5As shown, multiple question-and-answer large models and classification models are pre-determined according to computing resources and inference frameworks respectively. During actual use, the classification model service and multiple large model services corresponding one by one to the multiple question-and-answer large models are started. Then, the classification model determines the corresponding level of the question queried by the user, and then the question-and-answer large model service of the corresponding level can be selected to generate the corresponding answer. Among them, the level of the question-and-answer large model indicates that the question-and-answer large model can answer the questions corresponding to this level. Generally speaking, the higher the level of the question, the greater the difficulty of the model's answer, and the higher the level of the question-and-answer large model, the larger the scale of the parameter quantity, and the more complex the questions it can answer. Therefore, the question-and-answer large model with a high level can be called to answer high-level questions, and the question-and-answer large model with a low level can be called to answer low-level questions, so as to reduce the concurrency of large models with a large scale of parameter quantity.
[0041] Using the above method, the classification model determines the level corresponding to the user's question, so as to select the question-and-answer large model with the smallest parameter quantity scale that can answer the questions of this level among multiple question-and-answer large models to generate an answer. That is, while ensuring the answer accuracy, it is avoided that all questions are called by large models with a large scale of parameter quantity, thereby reducing the concurrency of large models with a large scale of parameter quantity, improving the processing efficiency of the large model, further improving the response speed and answering efficiency of the large model, reducing the user's waiting time, and enhancing the user experience.
[0042] It should be understood that simple questions actually do not require super large models to perform inference responses. Some models with relatively small parameter quantities, such as models with a parameter quantity of 7B or 2B, have very good effects when dealing with user questions in general scenarios, and are not worse than super large models. Users are also very satisfied with the answer effects.
[0043] Moreover, since the inference accuracy of large models with relatively small parameter quantities is equivalent to that of super large models in some scenarios, the consumption of computing resources is extremely low, and the inference efficiency is increased by dozens of times. Therefore, classifying the user's questions and sending questions of different levels to question-and-answer large models with different parameter quantity scales can, while ensuring the inference accuracy of the model, adhere to the principle of the smallest model inference, reduce the concurrency of super large models, and thus improve the inference and answering efficiency.
[0044] In a possible way, classifying the target question through the classification model to determine the target level corresponding to the target question includes: classifying the target question through the first classification model to determine the target level corresponding to the target question.
[0045] Among them, the first classification model is trained in the following way: Obtain the first sample questions, and respectively generate sample answers corresponding to the first sample questions through multiple question-answering large models to obtain multiple sample answers; Determine the first answer whose answer quality meets the preset conditions among the multiple sample answers, and determine the second question-answering large model with the smallest parameter scale among the question-answering large models that generate the first answer; Label the first sample questions based on the level corresponding to the second question-answering large model to obtain the first training samples; Train the initial first classification model based on the training samples to obtain the trained first classification model.
[0046] Exemplarily, as Figure 6 shown, question-answering large models with different parameter scales can be determined, such as multiple question-answering large models like 2B, 7B, 13B, 26B, 72B, etc., and the corresponding large model services are started on the computing resources. And a large number of user questions are collected as the first sample questions. To ensure the training effect, questions in different fields can be covered here, such as data and scenario user data in fields like history, literature, physics, programming, mathematics, etc. to construct a question dataset, and the present disclosure does not limit this. Additionally, the larger the data volume and the higher the data quality of the samples, the more beneficial it is to the training of the classification model. For example, a data volume of 10B can be selected, and the present disclosure does not limit this.
[0047] Then, as Figure 7 shown, use the large model services with different parameter scales to respectively generate corresponding sample answers for multiple sample questions, and screen out the first answers whose answer quality meets the preset conditions from them. For example, sample answers with reasonable answers can be screened out. Then, for each question, generate the label of the question according to the level corresponding to the question-answering large model that can reasonably answer the question and has the smallest parameter scale. For example, the level of the question-answering large model with the smallest parameter scale is 1, increasing sequentially, and the level of the question-answering large model with the largest parameter scale is N. For example, if for a certain question among the answers generated by N question-answering large models, the question-answering large model with reasonable answers and the smallest parameter scale corresponds to level 2, then the question level of this question is also 2.
[0048] After screening, each question corresponds to a level label, such as (question 1, level 2), (question 2, level 3). This label represents the level of the question-answering large model with the smallest parameter scale that can answer the question, and can be specifically set according to requirements. For example, the larger the question level, the larger the parameter scale of the question-answering large model required to answer the question.
[0049] Construct training samples based on the labeled sample questions, and train the classification model based on the training samples. Furthermore, through the classification service, bridge the classification model and the large model services with different parameter scales.
[0050] Therefore, through the second classification model completed by training, the questions in the user query can be classified from the dimension processed by the model. For example, the questions in the user query can be classified according to the difficulty of the model's answer, so as to more accurately call the corresponding large question-answering model to generate answers, avoiding waste of computing resources while improving the response speed and answer efficiency of the model, and further enhancing the user experience.
[0051] In a possible way, determining the first answer among multiple sample answers whose answer quality meets the preset conditions includes: for each sample answer among the multiple sample answers, annotating the sample answer according to the answer quality corresponding to the sample answer, where the sample answer whose answer quality meets the preset conditions is annotated with a first label; determining the sample answer annotated with the first label as the first answer.
[0052] Exemplarily, the sample answers can be annotated according to the answer quality corresponding to the sample answers. For example, a reasonable answer is annotated as 1, and an unreasonable answer is annotated as 0, so as to subsequently screen out the sample answers with reasonable answers, and further screen out the large question-answering model that can handle questions of this level and has the smallest scale parameter quantity.
[0053] The above annotation process can be carried out in the way of manual annotation, or can be automatically annotated based on an annotation large model, or can also combine manual annotation and model automatic annotation. Specifically, it can be selected according to requirements, and the present disclosure does not limit this.
[0054] As Figure 8 shown, for the answers generated by each large question-answering model, use the annotation large model for annotation, construct a prompt word based on the sample question and the sample answer, and this prompt word is used to prompt the annotation large model to annotate a reasonable answer as 1 and an unreasonable answer as 0, or an answer with good quality as 1 and an answer with poor quality as 0, so as to obtain the annotated question-answer sample pair, and further screen out the large question-answering model that can handle questions of this level and has the smallest scale parameter quantity.
[0055] Among them, it is possible to determine whether the sample answer is reasonable or the answer quality is good or bad according to the matching degree between the sample question and the sample answer, the similarity between the sample answer and the standard answer, and by combining the matching degree between the sample question and the sample answer and the similarity between the sample answer and the standard answer, and then perform automatic annotation.
[0056] In a possible way, annotating the sample answer according to the answer quality corresponding to the sample answer includes: constructing a first prompt word based on the first sample question and the sample answer; inputting the first prompt word into the annotation large model to obtain the annotation result of the sample answer output by the annotation large model, and the annotation large model is at least used to annotate the sample answer with the first label when the matching degree between the sample answer and the first sample question is greater than a first threshold.
[0057] Exemplarily, sample annotation can be performed based on the matching degree between the sample answer and the sample question, that is, by annotating the large model to perform semantic understanding on the sample answer and the sample question, and judging whether the sample answer can answer the sample question and the logical rationality of the answer. Generally speaking, the higher the matching degree, the higher the answer quality. Thus, the answer quality can be determined by the annotation large model based on the semantic logic between the sample answer and the sample question, improving the accuracy and efficiency of sample annotation, and further improving the efficiency and effect of the training of the Q&A large model.
[0058] In a possible way, the sample answer is annotated according to the answer quality corresponding to the sample answer, including: constructing a second prompt word based on the sample answer and the preset standard answer corresponding to the first sample question; inputting the second prompt word into the annotation large model to obtain the annotation result of the sample answer output by the annotation large model. The annotation large model is at least used to annotate the first label for the sample answer when the similarity between the sample answer and the preset standard answer is greater than the second threshold.
[0059] Exemplarily, sample annotation can be performed based on the similarity between the sample answer and the standard answer of the sample question, that is, by annotating the large model to perform semantic understanding on the sample answer and the standard answer, and judging whether the sample answer is semantically similar to the standard answer. Generally speaking, the higher the similarity, the higher the answer quality. Thus, the answer quality can be determined by the annotation large model based on the semantic similarity between the sample answer and the standard answer, improving the accuracy and efficiency of sample annotation, and further improving the efficiency and effect of the training of the Q&A large model.
[0060] In a possible way, the sample answer is annotated according to the answer quality corresponding to the sample answer, including: constructing a third prompt word based on the sample answer, the first sample question, and the preset standard answer corresponding to the first sample question; inputting the third prompt word into the annotation large model to obtain the annotation result of the sample answer output by the annotation large model. The annotation large model is at least used to annotate the first label for the sample answer when the matching degree between the sample answer and the first sample question is greater than the third threshold and the similarity between the sample answer and the preset standard answer is greater than the fourth threshold.
[0061] Exemplarily, sample annotation can be performed based on the similarity between the sample answer and the standard answer of the sample question and the matching degree between the sample answer and the sample question, that is, by annotating the large model to perform semantic understanding on the sample answer and the standard answer, the sample question and the sample answer respectively, and judging whether the sample answer is semantically similar to the standard answer and whether the sample answer is logically reasonable with the sample question. Thus, the answer quality can be determined by the annotation large model based on the semantic similarity between the sample answer and the standard answer and the matching degree between the sample answer and the sample question, further improving the accuracy and efficiency of sample annotation, and further improving the efficiency and effect of the training of the Q&A large model.
[0062] Among them, the above-mentioned first threshold, second threshold, third threshold, and fourth threshold can be set according to requirements, and the present disclosure places no restrictions thereon.
[0063] In the embodiments of the present disclosure, the above-mentioned question-and-answer large model based on the minimum scale of parameter quantity generates answers to questions of corresponding levels. Essentially, it selects a question-and-answer large model with corresponding capabilities for model inference based on the difficulty level of model processing of the question, that is, the more complex the question, the larger the scale of the parameter quantity of the selected question-and-answer large model.
[0064] In a possible manner, classifying the target question through a classification model to determine the target level corresponding to the target question includes: classifying the target question through a second classification model to determine the initial level corresponding to the target question, and determining the mapping level corresponding to the initial level in the level mapping relationship corresponding to the first user as the target level corresponding to the target question; wherein, the level mapping relationship corresponding to the first user is determined based on the historical interaction data of the first user, and the level mapping relationship corresponding to the first user includes the mapping relationship between the first level and the second level. The first level is the classification level determined by the first classification model according to the historical questions sent by the first user, and the second level is the classification level corresponding to the question-and-answer large model that generates the second answer and has the smallest scale of parameter quantity. The second answer represents the historical answer in the historical answers corresponding to the historical questions of the first user with a satisfaction degree higher than the preset satisfaction degree. The first classification model is used to output the corresponding classification level according to the input question.
[0065] In the embodiments of the present disclosure, in addition to selecting a question-and-answer large model based on the difficulty level of model processing of the question, a question-and-answer large model can also be selected in combination with the user's needs under the condition of user authorization. For example, the level mapping relationship corresponding to the user can be constructed according to the user's historical interaction data, such as the questions sent by the user history and the satisfaction degree feedback on the historical answers corresponding to the historical questions.
[0066] Exemplarily, assume that the level of the question sent by the user is 2, and the satisfaction degree feedback by the user on the answer generated by the question-and-answer large model of level 2 is lower than the preset threshold. Then, when the user sends a question of level 2 next time, a question-and-answer large model of a higher level can be selected, such as a question-and-answer large model of level 3, to generate an answer until the satisfaction degree feedback by the user on the answer generated by the question-and-answer large model is higher than the preset threshold. Record the level corresponding to the question-and-answer large model that the user is satisfied with, for example, it is level 4. Then, the level mapping relationship between the model prediction level 2 and the user satisfaction level 4 can be constructed.
[0067] Furthermore, when the question sent by the user next time is initially determined by the classification model to be at level 2, the Q&A large model with level 4 can be selected in combination with the user's level mapping relationship to generate an answer. This is to ensure that while meeting the user's needs, the Q&A large model with the smallest parameter scale is selected as much as possible, thereby reducing the concurrency of large models with large parameter scales, improving the processing efficiency of large models, further improving the response speed and answering efficiency of large models, reducing the user's waiting time, and enhancing the user's usage experience.
[0068] It should be noted that for users without historical interaction data, the above first classification model can be directly used for question classification. For users with historical interaction data, with the user's authorization, the user's historical interaction data can be collected to construct the user's level mapping relationship, and then question classification can be performed based on the above second classification model. Specifically, it can be selected according to requirements, and the present disclosure does not limit this.
[0069] In a possible way, the second classification model is trained as follows: Obtain the second sample question sent by the second user, the third level corresponding to the second sample question determined by the first classification model, and the historical satisfaction of the second user with the answers to the second sample question generated by multiple Q&A large models; Determine the fourth level corresponding to the Q&A large model with the historical satisfaction greater than the preset satisfaction and the smallest parameter scale among the multiple Q&A large models, and construct the level mapping relationship corresponding to the second user based on the third level and the fourth level; Based on the second sample question and the level mapping relationship corresponding to the second user, construct a second training sample; Train the first classification model based on the second training sample to obtain the trained second classification model.
[0070] Exemplarily, the historical interaction data can be collected with the user's authorization to construct the user's level mapping relationship, and then the training sample can be constructed based on the user's sample question and the level mapping relationship to obtain the trained second classification model. Here, the second classification model can be trained on the basis of the trained first classification model, so that the second classification model can not only classify questions but also classify questions in combination with the user's needs to select a Q&A large model that satisfies the user.
[0071] During the training process of the above first classification model and second classification model, supervised training can be performed through the grading labels of the questions, the cross-entropy loss can be calculated, and then the parameters of the model can be updated by backpropagation. The present disclosure will not elaborate on this here.
[0072] It is worth noting that as Figure 9As shown, the classification model can be a Transformer structure (a deep learning architecture based on self-attention mechanism). M layers of Transformers can be stacked, for example, 12 layers, with the number of parameters less than 200M and extremely fast inference efficiency. The main function of this classification model is to infer a problem level from the question input by the user, so as to find the Q&A large model with the smallest number of parameters that can solve this problem for inference, avoiding directly using a super-large model for inference, reducing the consumption of computing resources, and thus improving the Q&A response efficiency.
[0073] A comparative experiment is conducted between the multi-model processing architecture provided in the embodiments of the present disclosure and the single large model processing architecture in the related art. The embodiments of the present disclosure can effectively improve the throughput rate and inference efficiency of the Q&A large model without losing the Q&A accuracy of the large model service, thereby being able to solve problems such as tight computing resources, high concurrency of user requests, slow response and lag of the large model service in the related art.
[0074] Based on the same concept, the embodiments of the present disclosure provide an intelligent Q&A device based on a large model, as Figure 10 shown. The intelligent Q&A device 100 includes: A display module 101 for displaying an intelligent interaction page, which is used to answer questions sent by users through a Q&A large model. The Q&A large model includes multiple Q&A large models with different parameter scales, and one Q&A large model corresponds to one level classification; A classification module 102 for classifying a target question through a classification model in response to a target question sent by a first user in the intelligent interaction page, determining the target level corresponding to the target question, and determining the first Q&A large model with the smallest parameter scale that can answer questions of the target level among the multiple Q&A large models; A generation module 103 for generating a target answer corresponding to the target question through the first Q&A large model.
[0075] Optionally, the classification module 102 is used to: Classify the target question through a first classification model to determine the target level corresponding to the target question; The first classification model is obtained by training through a first training module, and the first training module is used to: Obtain a first sample question, and respectively generate sample answers corresponding to the first sample question through the multiple Q&A large models to obtain multiple sample answers; Determine a first answer whose answer quality meets a preset condition among the multiple sample answers, and determine the second Q&A large model with the smallest parameter scale from the Q&A large model that generates the first answer; Annotate the first sample question based on the level corresponding to the second question-and-answer large model to obtain a first training sample; Train the initial first classification model based on the training sample to obtain a trained first classification model.
[0076] Optionally, the first training module is used to: For each sample answer among the multiple sample answers, annotate the sample answer according to the answer quality corresponding to the sample answer, where the sample answer whose answer quality meets the preset condition is annotated with a first label; Determine the sample answer annotated with the first label as the first answer.
[0077] Optionally, the first training module is used to: Construct a first prompt word based on the first sample question and the sample answer; Input the first prompt word into the annotation large model to obtain the annotation result of the sample answer output by the annotation large model. The annotation large model is at least used to annotate the sample answer with the first label when the matching degree between the sample answer and the first sample question is greater than a first threshold.
[0078] Optionally, the first training module is used to: Construct a second prompt word based on the sample answer and the preset standard answer corresponding to the first sample question; Input the second prompt word into the annotation large model to obtain the annotation result of the sample answer output by the annotation large model. The annotation large model is at least used to annotate the sample answer with the first label when the similarity between the sample answer and the preset standard answer is greater than a second threshold.
[0079] Optionally, the first training module is used to: Construct a third prompt word based on the sample answer, the first sample question, and the preset standard answer corresponding to the first sample question; Input the third prompt word into the annotation large model to obtain the annotation result of the sample answer output by the annotation large model. The annotation large model is at least used to annotate the sample answer with the first label when the matching degree between the sample answer and the first sample question is greater than a third threshold and the similarity between the sample answer and the preset standard answer is greater than a fourth threshold.
[0080] Optionally, the classification module 102 is used to: Classify the target question through a second classification model, determine the initial level corresponding to the target question, and determine the mapped level corresponding to the initial level in the level mapping relationship of the first user as the target level corresponding to the target question; Among them, the level mapping relationship of the first user is determined based on the historical interaction data of the first user. The level mapping relationship of the first user includes the mapping relationship between the first level and the second level. The first level is the classification level determined by the first classification model according to the historical questions sent by the first user, and the second level is the classification level corresponding to the question-answering large model that generates the second answer and has the smallest parameter scale. The second answer represents the historical answer among the historical answers corresponding to the historical questions of the first user with a satisfaction greater than the preset satisfaction. The first classification model is used to output the corresponding classification level according to the input question.
[0081] Optionally, the second classification model is trained by a second training module, and the second training module is used for: Obtain the second sample question sent by the second user, the third level corresponding to the second sample question determined by the first classification model, and the historical satisfaction of the second user with respect to the answers to the second sample question generated by the multiple question-answering large models; Determine the fourth level corresponding to the question-answering large model with the historical satisfaction greater than the preset satisfaction and the smallest parameter scale among the multiple question-answering large models, and construct the level mapping relationship of the second user based on the third level and the fourth level; Construct a second training sample based on the second sample question and the level mapping relationship of the second user; Train the first classification model based on the second training sample to obtain the trained second classification model.
[0082] Based on the same concept, an embodiment of the present disclosure also provides a computer-readable medium, on which a computer program is stored, and when the program is executed by a processing device, the steps of the above-mentioned intelligent question-answering method based on a large model are implemented.
[0083] Based on the same concept, an embodiment of the present disclosure also provides an electronic device, which may include: A storage device, on which a computer program is stored; A processing device for executing the computer program in the storage device to implement the steps of the above-mentioned intelligent question-answering method based on a large model.
[0084] Based on the same concept, an embodiment of the present disclosure also provides a computer program product, including a computer program, which when executed by a processor implements the steps of the above-mentioned large model-based intelligent question answering method.
[0085] Reference is made below Figure 11 , which shows a schematic structural diagram of an electronic device 110 suitable for implementing the embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 11 The electronic device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present disclosure.
[0086] As Figure 11 shown, the electronic device 110 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 111, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 112 or the program loaded from the storage device 118 into the random access memory (RAM) 113. In the RAM 113, various programs and data required for the operation of the electronic device 110 are also stored. The processing device 111, the ROM 112, and the RAM 113 are connected to each other through a bus 114. The input / output (I / O) interface 115 is also connected to the bus 114.
[0087] Generally, the following devices may be connected to the I / O interface 115: an input device 116 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 117 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 118 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 119. The communication device 119 can allow the electronic device 110 to communicate with other devices wirelessly or wirelesly to exchange data. Although Figure 11 shows the electronic device 110 having various devices, it should be understood that it is not required to implement or include all the shown devices. Instead, more or fewer devices may be implemented or included.
[0088] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present disclosure include a computer program product that includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 119, or installed from the storage device 118, or installed from the ROM 112. When the computer program is executed by the processing device 111, the above-described functions defined in the methods of the embodiments of the present disclosure are performed.
[0089] It should be noted that the above-mentioned computer-readable medium in the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in the present disclosure, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0090] In some embodiments, any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol) can be used for communication, and it can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0091] The above computer-readable medium may be included in the above electronic device; or it may exist separately and not be assembled into the electronic device.
[0092] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: display an intelligent interaction page for answering questions sent by a user through a question-and-answer large model, the question-and-answer large model including multiple question-and-answer large models with different parameter scales, and one question-and-answer large model corresponding to one level classification; in response to a target question sent by a first user in the intelligent interaction page, classify the target question through a classification model, determine the target level corresponding to the target question, and determine a first question-and-answer large model that can answer questions of the target level and has the smallest parameter scale among the multiple question-and-answer large models; generate a target answer corresponding to the target question through the first question-and-answer large model.
[0093] Computer program code for performing the operations of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the “C” language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).
[0094] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that, in some alternative implementations, the functions noted in the blocks may occur in a different order than noted in the drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.
[0095] The modules described in the embodiments of the present disclosure can be implemented in software or in hardware. In some cases, the name of a module does not constitute a limitation on the module itself.
[0096] The functions described above herein can be performed, at least in part, by one or more hardware logic components. By way of example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip (SOCs), complex programmable logic devices (CPLDs), and the like.
[0097] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0098] The above description is only a preferred embodiment of the present disclosure and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by the mutual replacement of the above features and the technical features (but not limited to) with similar functions disclosed in the present disclosure.
[0099] In addition, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although a number of specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features described in the context of separate embodiments may also be implemented combinatorially in a single embodiment. Conversely, the various features described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0100] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. On the contrary, the specific features and acts described above are merely example forms for implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.
Claims
1. An intelligent question-answering method based on a large model, characterized in that The intelligent question answering method includes: Displaying an intelligent interaction page, which is used to answer questions sent by users through a question answering large model. The question answering large model includes multiple question answering large models with different parameter scales, and one question answering large model corresponds to one level classification; In response to a target question sent by a first user in the intelligent interaction page, classifying the target question through a classification model, determining the target level corresponding to the target question, and determining a first question answering large model with the smallest parameter scale among the multiple question answering large models that can answer questions of the target level; Generating a target answer corresponding to the target question through the first question answering large model.
2. The intelligent question and answer method based on a large model according to claim 1, wherein The classifying the target question through the classification model to determine the target level corresponding to the target question includes: Classifying the target question through a first classification model to determine the target level corresponding to the target question; The first classification model is trained in the following way: Obtaining a first sample question, and respectively generating sample answers corresponding to the first sample question through the multiple question answering large models to obtain multiple sample answers; Determining a first answer whose answer quality meets a preset condition among the multiple sample answers, and determining a second question answering large model with the smallest parameter scale among the question answering large models that generate the first answer; Labeling the first sample question based on the level corresponding to the second question answering large model to obtain a first training sample; Training an initial first classification model based on the training sample to obtain a trained first classification model.
3. The intelligent question-answering method based on a large model according to claim 2, wherein, The determining the first answer whose answer quality meets a preset condition among the multiple sample answers includes: For each sample answer among the multiple sample answers, labeling the sample answer according to the answer quality corresponding to the sample answer, where the sample answer whose answer quality meets the preset condition is labeled with a first label; Determining the sample answer labeled with the first label as the first answer.
4. The intelligent question-answering method based on a large model according to claim 3, wherein The labeling the sample answer according to the answer quality corresponding to the sample answer includes: Constructing a first prompt word based on the first sample question and the sample answer; Inputting the first prompt word into a labeling large model to obtain a labeling result of the sample answer output by the labeling large model. The labeling large model is at least used to label the first label for the sample answer when the matching degree between the sample answer and the first sample question is greater than a first threshold.
5. The intelligent question-answering method based on a large model according to claim 3, wherein, The labeling the sample answer according to the answer quality corresponding to the sample answer includes: Constructing a second prompt word based on the sample answer and a preset standard answer corresponding to the first sample question; Inputting the second prompt word into a labeling large model to obtain a labeling result of the sample answer output by the labeling large model. The labeling large model is at least used to label the first label for the sample answer when the similarity between the sample answer and the preset standard answer is greater than a second threshold.
6. The intelligent question-answering method based on a large model according to claim 3, wherein The labeling the sample answer according to the answer quality corresponding to the sample answer includes: Construct a third prompt based on the sample answer, the first sample question, and the preset standard answer corresponding to the first sample question; Input the third prompt into the annotation large model to obtain the annotation result of the sample answer output by the annotation large model. The annotation large model is at least used to annotate the first label for the sample answer when the matching degree between the sample answer and the first sample question is greater than a third threshold and the similarity between the sample answer and the preset standard answer is greater than a fourth threshold.
7. The intelligent question-answering method based on a large model according to claim 1, wherein, The classifying the target question by the classification model to determine the target level corresponding to the target question includes: Classify the target question through a second classification model to determine the initial level corresponding to the target question, and determine the mapped level corresponding to the initial level in the level mapping relationship corresponding to the first user as the target level corresponding to the target question; Among them, the level mapping relationship corresponding to the first user is determined based on the historical interaction data of the first user. The level mapping relationship corresponding to the first user includes the mapping relationship between the first level and the second level. The first level is the classification level determined by the first classification model according to the historical questions sent by the first user. The second level is the classification level corresponding to the question and answer large model that generates the second answer and has the smallest parameter scale. The second answer represents the historical answer with a satisfaction greater than the preset satisfaction among the historical answers corresponding to the historical questions of the first user. The first classification model is used to output the corresponding classification level according to the input question.
8. The intelligent question-answering method based on a large model according to claim 7, wherein, The second classification model is trained in the following way: Obtain the second sample question sent by the second user, the third level corresponding to the second sample question determined by the first classification model, and the historical satisfaction of the second user with respect to the answers to the second sample question generated by the multiple question and answer large models; Determine the fourth level corresponding to the question and answer large model with the historical satisfaction greater than the preset satisfaction and the smallest parameter scale among the multiple question and answer large models, and construct the level mapping relationship corresponding to the second user based on the third level and the fourth level; Construct a second training sample based on the second sample question and the level mapping relationship corresponding to the second user; Train the first classification model based on the second training sample to obtain the trained second classification model.
9. An intelligent question-answering device based on a large model, characterized in that, The intelligent question answering device includes: A display module for displaying an intelligent interaction page, which is used to answer questions sent by users through a question and answer large model. The question and answer large model includes multiple question and answer large models with different parameter scales, and one question and answer large model corresponds to one level classification; A classification module for responding to a target question sent by a first user on the intelligent interaction page, classifying the target question through a classification model to determine the target level corresponding to the target question, and determining the first question and answer large model with the smallest parameter scale that can answer the questions of the target level among the multiple question and answer large models; A generation module, configured to generate a target answer corresponding to the target question through the first question-and-answer large model.
10. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processing device, it implements the steps of the method according to any one of claims 1-8.
11. An electronic device, characterized in that, Comprising: A storage device, on which a computer program is stored; A processing device, configured to execute the computer program in the storage device to implement the steps of the method according to any one of claims 1-8.
12. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.
Citation Information
Patent Citations
Question and answer method and device, equipment and medium
CN118585619A
Question and answer processing method and device based on intelligent agent, medium, equipment and product
CN119357338A
Classification and question and answer large model joint parameter adjustment method and device
CN119621912A
Task processing method and device, electronic equipment and computer readable storage medium
CN119718573A
Cited By
Hierarchical answering method, electronic equipment and computer readable storage medium
CN121279451A
Large model screening method, device, equipment, medium and product
CN121743048A