User question response method and device based on AI large model and storage medium

By sending user questions to multiple dedicated large models and displaying the aggregated response results, the problems of slow response speed and incomplete information of large AI models are solved, achieving fast and complete response to user questions.

CN120994382APending Publication Date: 2025-11-21SHANGHAI DAJIAYING INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511109125.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-08
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

Existing large AI models suffer from slow response times and incomplete response information when responding to user questions. The model with the least load is not necessarily the fastest and may even result in missing information.

Method used

Send the user's question to at least two dedicated large models. First, display the response of the first dedicated large model and store the responses of the other dedicated large models. Obtain the summary results, perform deduplication, and display supplementary information to ensure the integrity of the response.

Benefits of technology

It improved the response speed to user issues, reduced user waiting time, ensured the integrity of response results, and avoided information loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994382A_ABST
    Figure CN120994382A_ABST
Patent Text Reader

Abstract

The invention discloses a user question response method and device based on an AI large model, electronic equipment and a storage medium, and relates to the field of AI large models.The method comprises the steps that in response to the obtained user question, the user question is sent to at least two special large models, and the first response of the first special large model obtained at first is displayed to a current user; storing a second response of the at least one second special large model to obtain a matched summarized response result, and performing duplicate removal processing on the summarized response result; and if it is determined that the display of the first response is completed, obtaining supplementary information except the first response in the summarized response result, so as to display the supplementary information to the current user. According to the technical scheme provided by the embodiment of the invention, the timely response of the user problem is ensured, the response speed of the user problem is improved, the waiting time of the user is shortened, the integrity of the response result is improved, and the phenomenon of information missing is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence large models, in particular to a user question response method and device based on an AI large model, an electronic device and a storage medium. BACKGROUND

[0002] With the continuous development of large model technology, processing user questions with the help of artificial intelligence (AI) large models has become a common solution for application software.

[0003] In the prior art, application software is usually communicatively connected to multiple AI large models, so that when communication failure occurs in an AI large model, other large models can be switched to continue completing problem solving. Specifically, the user question is sent to the AI large model with the smallest load based on load balancing, so as to improve the response speed of the current user question through the AI large model with the smallest load.

[0004] However, such a user question processing method does not mean that the AI large model with the smallest load necessarily has the fastest response speed, and there is a risk of response delay for the user question. At the same time, the AI large model cannot ensure the completeness of the displayed information, and there is a possibility of information loss in the response result. SUMMARY

[0005] The present application provides a user question response method and device based on an AI large model, an electronic device and a storage medium to solve the problem of slow response speed and incomplete response result information of the current AI large model.

[0006] According to an aspect of the present application, a user question response method based on an AI large model is provided, comprising:

[0007] In response to obtaining a user question, the user question is sent to at least two special-purpose large models, and a first response of a first special-purpose large model obtained first is displayed to the current user;

[0008] The second response of at least one second special-purpose large model is stored to obtain a matching summary response result, and the summary response result is de-duplicated;

[0009] If it is determined that the first response display is complete, supplementary information in the summary response result other than the first response is obtained to display the supplementary information to the current user.

[0010] The sending of the user question to the at least two special large models comprises: acquiring, according to a question category of the user question, average response time and average task achievement degree of each special large model under the current question category; wherein the average task achievement degree is acquired based on content completeness and / or user satisfaction; and the user question is sent to the special large model with the shortest average response time and the special large model with the highest average task achievement degree.

[0011] The acquiring of the average response time and the average task achievement degree of each special large model under the current question category comprises: acquiring, according to information closed-loop, logical endpoint and element coverage of historical responses of each special large model under the current question category, content completeness of each special large model under the current question category.

[0012] After the average response time and the average task achievement degree of each special large model under the current question category are acquired, the method further comprises: judging whether a predicted response time of the special large model with the highest average task achievement degree is greater than or equal to a habitual waiting time of the current user; if it is determined that the predicted response time is less than the habitual waiting time of the current user, the user question is sent to the special large model with the highest average task achievement degree; and the sending of the user question to the special large model with the shortest average response time and the special large model with the highest average task achievement degree comprises: if it is determined that the predicted response time is greater than or equal to the habitual waiting time of the current user, the user question is sent to the special large model with the shortest average response time and the special large model with the highest average task achievement degree.

[0013] The sending of the user question to the at least two special large models comprises: sending the user question to a specified special large model and judging whether a response of the specified special large model is acquired within a predicted response time; wherein the predicted response time is related to a response time threshold of the specified special large model, and the predicted response time is less than the response time threshold; and if it is determined that the response of the specified special large model is not acquired within the predicted response time, the user question is sent to at least one other special large model except the specified special large model.

[0014] The acquiring of the supplementary information in the summary response result except the first response to display the supplementary information to the current user comprises: acquiring the supplementary information in the summary response result except the first response, and sending prompt information to the current user to guide the user to send a supplementary display instruction; and in response to the acquisition of the supplementary display instruction sent by the current user, the supplementary information is displayed to the current user.

[0015] According to another aspect of the present application, a user question response device based on an AI large model is provided, comprising:

[0016] The question sending module is configured to, in response to obtaining a user question, send the user question to at least two special large models, and display a first response of a first special large model obtained first to a current user.

[0017] The summary execution module is configured to store second responses of at least one second special large model, obtain a matched summary response result, and perform deduplication processing on the summary response result.

[0018] The response display module is configured to, if it is determined that the first response display is completed, obtain supplementary information other than the first response in the summary response result, and display the supplementary information to the current user.

[0019] According to another aspect of the present application, an electronic device is provided, which comprises:

[0020] at least one processor; and

[0021] a memory connected to the at least one processor in communication; wherein

[0022] The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the AI large model-based user question response method according to any of the embodiments of the present application.

[0023] According to another aspect of the present application, a computer readable storage medium is provided, which stores computer instructions for enabling a processor to implement the AI large model-based user question response method according to any of the embodiments of the present application when executed by the processor.

[0024] According to another aspect of the present application, a computer program product is provided, which comprises a computer program for implementing the AI large model-based user question response method according to any of the embodiments of the present application when executed by a processor.

[0025] The technical solution of the embodiments of the present application sends a user question to at least two special large models in response to obtaining the user question, and displays a first response of a first special large model obtained first to a current user; stores second responses of at least one second special large model, obtains a matched summary response result, and performs deduplication processing on the summary response result; if it is determined that the first response display is completed, obtains supplementary information other than the first response in the summary response result, and displays the supplementary information to the current user. Thus, not only is the timely response of the user question ensured, the response speed of the user question is improved, and the user waiting time is reduced, but also the completeness of the response result is improved, and the information missing phenomenon is avoided.

[0026] It should be understood that the matters described in this detailed description are intended to be illustrative and are not intended to limit or restrict the scope or applicability of the embodiments of the application in any way. This description set in examples is intended to be illustrative, and is not intended to be limiting on the scope, applicability or configuration of the disclosure. Various modifications can be made to the disclosure without departing from its scope. BRIEF DESCRIPTION OF DRAWINGS

[0027] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort based on these drawings.

[0028] Figure 1 is a flow chart of a user question response method based on an AI large model according to an embodiment of the present application;

[0029] Figure 2 is a flow chart of another user question response method based on an AI large model according to an embodiment of the present application;

[0030] Figure 3 is a flow chart of still another user question response method based on an AI large model according to an embodiment of the present application;

[0031] Figure 4 is a structural schematic diagram of a user question response device based on an AI large model according to an embodiment of the present application;

[0032] Figure 5 is a structural schematic diagram of an electronic device implementing a user question response method based on an AI large model according to an embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to make the technical personnel in the art better understand the present application scheme, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative effort should be within the scope of protection of the present application.

[0034] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and in the above-described drawings are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments of the application described herein can be implemented in an order other than that illustrated or described herein. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion, for example, a process, method, system, product or device that includes a series of steps or units does not have to be limited to only those steps or units clearly listed, but can include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0035] Embodiment one

[0036] Figure 1 A flowchart of a user question response method based on an AI large model is provided for the first embodiment of the present application. The present embodiment can be applied to the case of displaying response text to a user through at least two special large models. The method can be executed by the user question response device based on the AI large model in any embodiment of the present application. The user question response device based on the AI large model can be realized in the form of hardware and / or software. The user question response device based on the AI large model can be configured in an electronic device such as a server. The server is a background server corresponding to a client. For example, the client can be used to solve the user's job-seeking problem, and to recommend suitable employment information to the user, such as Figure 1 As shown in the figure, the method comprises:

[0037] S101, in response to obtaining a user question, sending the user question to at least two special large models, and displaying the first response of the first special large model obtained first to the current user.

[0038] The special large model refers to a large language model (LLM) based on deep learning. It is a large language model trained and optimized based on training samples in the business field of the current client, on the basis of the general large models provided by different suppliers, so as to significantly improve the performance in the business field while retaining the general ability. Different special large models can be accessed through different interfaces provided by different suppliers, or each special large model can be saved locally.

[0039] Taking a job client as an example, the special large model is an AI large model used for accurately identifying user needs and matching them with recruitment information, which matches the demand information (i.e., user questions) issued by the user with the recruitment information still effective at the current moment, to recommend matching employment information to the user; for example, the user question raised by the user can be "I want to find a job in the A area with a monthly salary of X yuan, and my sleep quality is not good, preferably day shift.", and the special large model displays the matching employment information to the user based on the above user question.

[0040] The user can raise questions in the form of text or voice in the electronic customer service of the client, and the server sends the user questions to two or more special large models after obtaining the user questions; since the large language model generates response text by predicting content token by token, that is, the generation of each new token depends on all previous tokens, and the special large model feeds back the first token after the first token, indicating that the response is issued, so the special large model that first feeds back the first token among the above two or more special large models is defined as the first special large model, and the response (i.e., the first response) of the first special large model is successively displayed to the current user.

[0041] S102, store the second response of at least one second special large model to obtain a matching summary response result, and perform deduplication processing on the summary response result.

[0042] Among the above two or more special large models, the remaining special large models except the first special large model that first issues the response are defined as second special large models, and after the first special large model issues the response, the second special large models will successively return their respective responses (i.e., the second response), and since the first response related information is displayed in the user interface at this time, the second responses returned by each second special large model are stored locally.

[0043] Since the first special large model issues the response first, it may return the complete response text at a faster speed first, and after the first response is completely displayed, one or more second special large models may not have fed back the complete response text, but this does not affect the continued feedback of the second special large model, because although the response text of the first special large model is completely displayed, the user's reading speed lags far behind the token feedback speed of the special large model, so even if the first special large model has fed back the complete response text completely, the user has not finished reading the above response text, and the server can still continue to obtain the second responses of each second special large model during this period.

[0044] The aggregated response result refers to the result after the second responses of the respective second special large models are aggregated. The aggregated response result includes all response information fed back by all second special large models. At this time, the second response of any second special large model can be taken as the benchmark text. For example, according to the pre-set use priority of each special large model, the respective second special large models are sorted, the second special large model with the highest priority is determined, and the second response (i.e., the target second response) issued by the second special large model is taken as the initial benchmark text.

[0045] First, one second response other than the target second response is compared with the current benchmark text in terms of semantic similarity to obtain text content (i.e., a text sentence or a text paragraph) other than the benchmark text, and the text content is merged with the current benchmark text to form a new benchmark text. Then, the next second response is compared with the current benchmark text in terms of semantic similarity to continue to obtain text content other than the benchmark text, and the text content is merged with the current benchmark text again to form a new benchmark text. This process is repeated until all second responses are compared with the current benchmark text in terms of semantic similarity and the text content is merged to obtain the aggregated response result after the deduplication process is completed.

[0046] The deduplication process of the aggregated response result can be performed by a pre-training language model based on the Transformer architecture. The current benchmark text and the current second response are converted into feature vectors, the cosine similarity between the feature vectors is calculated, and then the text sentence or the text paragraph with a similarity greater than a first similarity threshold (e.g., 0.85) is regarded as a semantically repeated text. The semantically repeated text is deleted, and the remaining text in the current second response is retained to complete the deletion of the repeated text in the aggregated response result.

[0047] S103, if it is determined that the first response has been displayed, the supplementary information other than the first response in the aggregated response result is obtained to display the supplementary information to the current user.

[0048] After the first response has been displayed, the response speed of the first special large model is faster, but the accuracy and completeness of the response text are not necessarily the best, as described in the above technical solution. At this time, the aggregated response result of at least one second special large model has been obtained. The first response is compared with the aggregated response result to obtain text content contained in the aggregated response result and not contained in the first response as supplementary information of the first response. After the first response is displayed, the supplementary information is continued to be displayed.

[0049] In particular, the supplementary information in the summary response result except the first response can be obtained by using the similarity comparison method in the above technical solution, that is, by comparing the first response with the text sentence or text segment of the summary response result respectively to obtain the text content with the same semantics according to the second similarity threshold (for example, 0.95), and the text content in the summary response result except the above text content with the same semantics is the supplementary information.

[0050] Therefore, not only the timely response of the user's question is ensured, the response speed of the user's question is improved, and the user's waiting time is reduced, but also the completeness of the response result is improved, and the information missing phenomenon is avoided; in particular, as described in the above technical solution, since the reading speed of the user is far behind the token feedback speed of the special large model, when the supplementary information is displayed, the user has not finished reading the first response, and accordingly the user can continue to read the above supplementary information after reading the first response, and in the user's view, the first token returned from the first response starts until the last character of the supplementary information, and the entire reading process does not have any pause or waiting, ensuring the smoothness of the user's reading.

[0051] In particular, the second similarity threshold can be configured to be greater than the first similarity threshold, so that the first similarity threshold with a lower value can ensure that the text content in the summary response result is as complete as possible to avoid missing of the response text content, and the second similarity threshold with a higher value can effectively avoid the occurrence of repeated information to prevent the occurrence of repeated information from affecting the user's reading experience.

[0052] Optionally, in the embodiment of the application, the obtaining of the supplementary information in the summary response result except the first response to display the supplementary information to the current user specifically includes: obtaining the supplementary information in the summary response result except the first response, and sending a prompt information to the current user to guide the user to issue a supplementary display instruction; and in response to obtaining the supplementary display instruction issued by the current user, the supplementary information is displayed to the current user.

[0053] Specifically, after the first response is displayed to the user, a prompt information can be sent to the user to decide whether to continue to display the supplementary information, if the user does not trigger the supplementary display instruction, the supplementary information is not displayed to the current user, and the locally stored summary response result is deleted, and if the user triggers the supplementary display instruction, the supplementary information is continued to be displayed to the current user, so that the supplementary information display strategy is configured according to the actual needs of the user, which not only ensures the completeness of the displayed information under the user's needs, but also avoids redundant display under non-user needs.

[0054] The technical scheme of the embodiment of the application responds to the acquired user question, sends the user question to at least two special large models, and displays the first response of the first special large model acquired first to the current user; the second response of at least one second special large model is stored to acquire a matched summary response result, and the summary response result is de-duplicated; if it is determined that the first response display is completed, the supplementary information in the summary response result except the first response is acquired to display the supplementary information to the current user. Thus, not only the timely response of the user question is ensured, the response speed of the user question is improved, and the user waiting time is reduced, but also the completeness of the response result is improved, and the information missing phenomenon is avoided.

[0055] Embodiment two

[0056] Figure 2 A flowchart of a user question response method based on an AI large model provided by the second embodiment of the application, the relationship between the present embodiment and the above-mentioned embodiments is that the user question is first sent to one special large model, and whether to continue sending the user question to other special large models is determined according to the response time of the special large model, for example, as shown in the following formula: Figure 2 The method comprises the following steps:

[0057] S201, in response to the acquired user question, the user question is sent to a specified special large model, and whether the response of the specified special large model is acquired within a predicted response time is judged; wherein the predicted response time is related to the response time threshold of the specified special large model, and the predicted response time is less than the response time threshold.

[0058] The specified special large model can be the special large model with the highest frequency of use, the special large model with the smallest load or the highest user evaluation, or an arbitrarily specified special large model; after the user raises a question, the special large model calculates the question and feeds back an answer after the calculation is completed; the time period from when the user raises the question to when the special large model feeds back the answer is the actual waiting time of the user and the response time of the special large model; in order to avoid long-time occupation of the interface resources of the special large model, the special large model allocates a response time threshold, i.e. a maximum response time, to each question; once the response time threshold is exceeded and no token is fed back, it is determined that the current question is timed out and will not be responded to.

[0059] The predicted response time can be a time result calculated according to a percentage of the response time threshold, for example, 80% of the response time threshold as the predicted response time, or an average response time obtained based on historical feedback records of the special large model under the current problem category, or a time when a preset number of problems (more than half, for example, more than 80%) have been responded. Obviously, the average response time or the time when the preset number of problems have been responded (i.e., the predicted response time) is less than the maximum response time (i.e., the response time threshold).

[0060] S202, if it is determined that the response of the specified special large model is not obtained within the predicted response time, the user question is sent to at least one other special large model other than the specified special large model.

[0061] The predicted response time represents the time when the user question should have been responded; if the response of the specified special large model to the current user question has been obtained after the predicted response time, the current user question does not need to be sent to other special large models; if the response of the specified special large model to the current user question has not been obtained after the predicted response time, it indicates that there is a risk that the response result cannot be obtained for the current user question, and at this time, the user question needs to be sent to other special large models, so that when the response of the specified special large model is timed out and the user question cannot be answered, the subsequent other special large models can continue to feed back the response text, thereby reducing the possibility of failure of the user question.

[0062] S203, the first response of the first special large model obtained first is displayed to the current user.

[0063] S204, the second response of at least one second special large model is stored to obtain a matched summary response result, and the summary response result is de-duplicated.

[0064] S205, if it is determined that the first response display is completed, the supplementary information in the summary response result other than the first response is obtained to display the supplementary information to the current user.

[0065] The technical scheme of the embodiment of the application sends the user question to the specified special large model and judges whether the response of the specified special large model is obtained within the predicted response time; if it is determined that the response of the specified special large model is not obtained within the predicted response time, the user question is sent to at least one other special large model other than the specified special large model. Therefore, when the response of one special large model is timed out and the user question cannot be answered, the subsequent other special large models can continue to feed back the response text, thereby reducing the possibility of failure of the user question.

[0066] Embodiment three

[0067] Figure 3 A flowchart of a user question response method based on an AI large model is provided for Embodiment Three of the present application. The relationship between this embodiment and the above-mentioned embodiments is that the user question is sent to the special large model with the shortest average response time and the special large model with the highest average task completion rate, as shown in Figure 3 The method comprises the following steps:

[0068] S301, in response to obtaining the user question, obtaining the average response time and the average task completion rate of each special large model under the current question category according to the question category of the user question; wherein the average task completion rate is obtained based on the content completeness and / or user satisfaction.

[0069] Due to the differences in model structure, model size and inference characteristics, different special large models may have different processing capabilities and processing methods when dealing with the same category of questions, and the same special large model may also have different processing capabilities and processing methods when dealing with different categories of questions. The question category can be obtained through a classification model constructed and trained based on neural network technology to determine the specific category of the question currently raised by the user.

[0070] For example, taking the client used to solve the user's job-seeking problem in the above-mentioned technology as an example, which recommends suitable employment information to the user, the question category can include job-seeking questions and non-job-seeking questions. The response time refers to the time period from when the user question is sent to the special large model to when the first token is returned by the special large model; under the current question category, the historical responses of each special large model are obtained to calculate the average response time for answering each question through the response time of the historical responses.

[0071] The average task completion rate reflects the quality of the text response of the special large model; wherein the content completeness can be calculated based on the character length of the response text, and after obtaining all the response texts of the special large model for the current category of questions, the average character length of each response text is calculated, the longer the average character length, the higher the content completeness; the content completeness objectively reflects the reply quality of each special large model based on the response text.

[0072] The user satisfaction degree is given by the user after the user is shown the response text of each question, and after the special large model obtains all response texts of the current category question and the corresponding satisfaction scores, the average satisfaction score of each response text is calculated. The higher the average satisfaction score, the higher the user satisfaction degree. The user satisfaction degree reflects the reply quality of each special large model from the user's subjective perspective based on the user's perception. Accordingly, the average task completion degree can be obtained according to the content completeness degree and / or the user satisfaction degree. For example, the average task completion degree is calculated according to the content completeness degree, the user satisfaction degree, the weight coefficient of the content completeness degree, and the weight coefficient of the user satisfaction degree.

[0073] S302, the user question is sent to the special large model with the shortest average response time and the special large model with the highest average task completion degree.

[0074] The user question is sent to the special large model with the shortest average response time and the special large model with the highest average task completion degree. Not only does it ensure that the special large model with the shortest average response time responds to the user question in a timely manner, but it also ensures that the special large model with the highest content completeness degree improves the reply quality of the response result.

[0075] S303, the first response of the first special large model obtained first is displayed to the current user.

[0076] S304, the second response of at least one second special large model is stored to obtain a matched summary response result, and the summary response result is de-duplicated.

[0077] S305, if it is determined that the first response display is complete, supplementary information in the summary response result except the first response is obtained to display the supplementary information to the current user.

[0078] Optionally, in the embodiments of the present application, the average response time and the average task completion degree of each special large model under the current question category are obtained by obtaining the content completeness degree of each special large model under the current question category according to the information closed-loop, the logical end point, and the element coverage of the historical response of each special large model under the current question category.

[0079] Specifically, the information closure is to check whether the response text of the special large model covers all explicit and implicit requirements of the input question, which can be obtained by element extraction based on information pairs (i.e., information pairs composed of questions and answers) or knowledge graph verification; the logical end point is to evaluate whether the response text ends naturally or is suddenly interrupted, which can be obtained by linguistic feature engineering, pre-trained classification model or streaming generation monitoring; the element coverage is to quantify the coverage ratio of key information points, which can be obtained by constructing a domain checklist, multi-modal verification or dynamic weight adjustment.

[0080] In particular, the above dynamic weight adjustment includes a weighting method based on Term Frequency-Inverse Document Frequency (TF-IDF); compared with the evaluation method based on the length of text characters, the content completeness of each special large model under the current question category is obtained by the information closure, the logical end point and the element coverage of the historical responses of each special large model under the current question category, the text quality evaluation based on the reply content is realized, and the accuracy of the content completeness evaluation result is further improved.

[0081] Optionally, in the embodiments of the present application, after obtaining the average response time and the average task achievement degree of each special large model under the current question category, it further includes: judging whether the predicted response time of the special large model with the highest average task achievement degree is greater than or equal to the habitual waiting time of the current user; if it is determined to be less than the habitual waiting time of the current user, the user question is sent to the special large model with the highest average task achievement degree; the sending of the user question to the special large model with the shortest average response time and the special large model with the highest average task achievement degree includes: if it is determined to be greater than or equal to the habitual waiting time of the current user, the user question is sent to the special large model with the shortest average response time and the special large model with the highest average task achievement degree.

[0082] Specifically, the habitual waiting time is obtained according to the historical question records of the current user. After the user raises a question, if the user voluntarily stops the current request before the special large model responds, the waiting time of the voluntary stop behavior is recorded, and the average waiting time of all voluntary stop behaviors is taken as the habitual waiting time of the current user. If the predicted response time of the special large model with the highest average task achievement degree is less than or equal to the habitual waiting time of the current user, it means that the user is patient enough to wait for the response of the special large model with the highest average task achievement degree, at this time, the user question is only sent to the special large model to reduce the number of called special large models and save model resources.

[0083] If the predicted response time of the special large model with the highest average task completion degree is greater than the habitual waiting time of the current user, it indicates that the user may not have enough patience to wait for the current special large model to respond, at this time, the user question is sent to the special large model with the shortest average response time and the special large model with the highest average task completion degree. Thus, the user question is ensured to be responded in time by the special large model with the shortest average response time, and the completeness of the response result is improved by the special large model with the highest average task completion degree.

[0084] The technical scheme of the embodiment of the application, after obtaining the user question, first acquires the average response time and the average task completion degree of each special large model under the current question category according to the question category of the user question, and then sends the user question to the special large model with the shortest average response time and the special large model with the highest average task completion degree. Accordingly, not only the special large model with the shortest average response time is ensured to respond to the user question in time, the response speed of the user question is improved, and the user waiting time is reduced, but also the special large model with the highest average task completion degree is ensured to improve the reply quality of the response result.

[0085] Embodiment Four

[0086] Figure 4 is a structural block diagram of a user question response device based on an AI large model provided by the embodiment four of the application, and specifically comprises:

[0087] The question sending module 401 is configured to send the user question to at least two special large models in response to obtaining the user question, and display the first response of the first special large model obtained first to the current user.

[0088] The summary executing module 402 is configured to store the second response of at least one second special large model to obtain a matched summary response result, and perform deduplication processing on the summary response result.

[0089] The response displaying module 403 is configured to obtain supplementary information in the summary response result except the first response if it is determined that the first response display is completed, and display the supplementary information to the current user.

[0090] The technical scheme of the embodiment of the application responds to the acquisition of the user question, sends the user question to at least two special large models, and displays the first response of the first special large model acquired first to the current user; the second response of at least one second special large model is stored to acquire a matched summary response result, and the summary response result is de-duplicated; if it is determined that the first response display is completed, the supplementary information in the summary response result except the first response is acquired to display the supplementary information to the current user. Thus, not only the timely response of the user question is ensured, the response speed of the user question is improved, and the user waiting time is reduced, but also the completeness of the response result is improved, and the information missing phenomenon is avoided.

[0091] Optionally, the summary execution module 402 is specifically configured to acquire the average response time and the average task achievement degree of each special large model under the current question category according to the question category of the user question; the average task achievement degree is acquired based on the content completeness and / or the user satisfaction; the user question is sent to the special large model with the shortest average response time and the special large model with the highest average task achievement degree.

[0092] Optionally, the summary execution module 402 is specifically further configured to acquire the content completeness of each special large model under the current question category according to the information closed-loop, the logical end point and the element coverage of the historical response of each special large model under the current question category.

[0093] Optionally, the summary execution module 402 is specifically further configured to judge whether the predicted response time of the special large model with the highest average task achievement degree is greater than or equal to the habitual waiting time of the current user; if it is determined that the predicted response time is less than the habitual waiting time of the current user, the user question is sent to the special large model with the highest average task achievement degree; if it is determined that the predicted response time is greater than or equal to the habitual waiting time of the current user, the user question is sent to the special large model with the shortest average response time and the special large model with the highest average task achievement degree.

[0094] Optionally, the summary execution module 402 is specifically further configured to send the user question to a specified special large model, and judge whether the response of the specified special large model is acquired within the predicted response time; the predicted response time is related to the response time threshold of the specified special large model, and the predicted response time is less than the response time threshold; if it is determined that the response of the specified special large model is not acquired within the predicted response time, the user question is sent to at least one other special large model except the specified special large model.

[0095] Optionally, in response to the display module 403, specifically configured to acquire supplementary information in the summary response result except the first response, and send prompt information to the current user to guide the user to issue a supplementary display instruction; in response to acquiring the supplementary display instruction issued by the current user, the supplementary information is displayed to the current user.

[0096] The device described above can perform the AI large model-based user question response method provided by any embodiment of the application, has the corresponding function modules and beneficial effects of the execution method. Technical details not described in detail in the present embodiment can be referred to the AI large model-based user question response method provided by any embodiment of the application.

[0097] Embodiment five

[0098] Figure 5 A structural schematic diagram of an electronic device 10 that can be used to implement embodiments of the present application is shown. The electronic device is intended to represent various forms of digital computers, such as laptops, desktops, tablets, personal digital assistants, electronic devices, blade servers, mainframes, and other appropriate computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular telephones, smart phones, wearable devices (e.g., headsets, glasses, watches, etc.), and other similar computing devices. The components shown here, their connections and relationships, and their functions, are meant to be examples only, and are not meant to limit implementations of the applications described and / or claimed in this document.

[0099] As shown in Figure 5 The electronic device 10 includes at least one processor 11, and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is in communication with the at least one processor 11, wherein the memory stores a computer program that can be executed by the at least one processor, and the processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or loaded into the random access memory (RAM) 13 from the storage unit 18. In the RAM 13, various programs and data required for the operation of the electronic device 10 can also be stored. The processor 11, the ROM 12, and the RAM 13 are connected to each other through a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0100] A plurality of components in the electronic device 10 are connected to the I / O interface 15, including: an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, a speaker, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices through a computer network, such as the Internet, and / or various telecommunication networks.

[0101] The processor 11 can be various general and / or special-purpose processing components having processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 performs various methods and processes described above, such as the AI large model-based user question response method.

[0102] In some embodiments, the AI large model-based user question response method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit. In some embodiments, part or all of the computer program can be loaded and / or installed onto the heterogeneous hardware accelerator via the ROM and / or the communication unit. When the computer program is loaded into the RAM and executed by the processor, one or more steps of the AI large model-based user question response method described above can be performed. Alternatively, in other embodiments, the processor can be configured to perform the AI large model-based user question response method by any other appropriate means, such as by means of firmware.

[0103] Various implementations of the systems and techniques described above can be realized in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a programmable logic device (PLD), a computer hardware, firmware, software, and / or combinations thereof. These various implementations can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0104] Computer programs for implementing the methods of the present application can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the computer program, when executed, implements the functions / acts specified in the flowcharts and / or block diagrams. The computer program can be executed entirely on a machine, partially on a machine, partially on a machine as part of a standalone software package, or entirely on a remote machine or electronic device.

[0105] In the context of the present application, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. A computer-readable storage medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. Alternatively, a computer-readable storage medium can be a machine-readable signal medium. More specific examples of a machine-readable storage medium will include one or more lines of a program of instructions in a transitory signal form, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0106] To provide for interaction with a user, the systems and techniques described here can be implemented on a heterogeneous hardware accelerator having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the heterogeneous hardware accelerator. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.

[0107] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a client computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), blockchain network, and the Internet.

[0108] The computing system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server arises by virtue of computer programs running on the respective computers and having a client-server relationship to each other. The server is a cloud server, also known as a cloud computing server or cloud host, which is a host product in the cloud computing service system to solve the defects of large management difficulty and weak business scalability in traditional physical hosts and VPS services.

[0109] It should be understood that the steps shown in the above various forms of flowcharts can be reordered, added, or deleted. For example, each step described in the present application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions of the present application can be achieved, and the present application is not limited herein.

[0110] The above detailed description does not constitute a limitation on the protection scope of the present application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A method for responding to a user question based on an AI large model, the method comprising: The method comprises the following steps: in response to obtaining a user question, sending the user question to at least two special large models, and displaying a first response of a first special large model obtained first to the current user; storing second responses of at least one second special large model to obtain a matched summary response result, and performing deduplication processing on the summary response result; if it is determined that the first response display is completed, obtaining supplementary information in the summary response result except the first response to display the supplementary information to the current user.

2. The method of claim 1, wherein, The step of sending the user question to at least two special large models comprises the following steps: According to the problem category of the user question, the average response time and the average task achievement degree of each special large model under the current problem category are obtained; wherein the average task achievement degree is obtained based on the content completeness and / or user satisfaction; send the user question to the special large model with the shortest average response time and the special large model with the highest average task achievement degree.

3. The method of claim 2, wherein, The step of obtaining the average response time and the average task achievement degree of each special large model under the current problem category comprises the following steps: According to the information closure, logical endpoint and element coverage of the historical response of each special large model under the current problem category, the content completeness of each special large model under the current problem category is obtained.

4. The method of claim 2, wherein, After obtaining the average response time and the average task achievement degree of each special large model under the current problem category, the following steps are further included: determine whether the predicted response time of the special large model with the highest average task achievement degree is greater than or equal to the habitual waiting time of the current user; if it is determined to be less than the habitual waiting time of the current user, send the user question to the special large model with the highest average task achievement degree; The step of sending the user question to the special large model with the shortest average response time and the special large model with the highest average task achievement degree comprises the following steps: if it is determined to be greater than or equal to the habitual waiting time of the current user, send the user question to the special large model with the shortest average response time and the special large model with the highest average task achievement degree.

5. The method of claim 1, wherein, The step of sending the user question to at least two special large models comprises the following steps: send the user question to a specified special large model and determine whether the response of the specified special large model is obtained within the predicted response time; wherein the predicted response time is related to the response time threshold of the specified special large model, and the predicted response time is less than the response time threshold; if it is determined that the response of the specified special large model is not obtained within the predicted response time, send the user question to at least one other special large model except the specified special large model.

6. The method of claim 1, wherein, The step of obtaining supplementary information in the summary response result except the first response to display the supplementary information to the current user comprises the following steps: obtain supplementary information in the summary response result except the first response, and send a prompt information to the current user to guide the user to issue a supplementary display instruction; in response to obtaining the supplementary display instruction issued by the current user, display the supplementary information to the current user. 7.A user question response apparatus based on an AI large model, characterized by The method comprises the following steps: The question sending module is configured to, in response to obtaining a user question, send the user question to at least two special large models, and display a first response of a first special large model obtained first to a current user; The summary execution module is configured to store second responses of at least one second special large model, obtain a matched summary response result, and perform deduplication processing on the summary response result; The response display module is configured to, if it is determined that the first response display is completed, obtain supplementary information in the summary response result other than the first response, and display the supplementary information to the current user.

8. An electronic device, comprising: The electronic device comprises: at least one processor; and a memory connected in communication with the at least one processor; wherein The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor to enable the at least one processor to execute the AI large model-based user question response method of any one of claims 1-6.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions for enabling the processor to execute the AI large model-based user question response method of any one of claims 1-6 when executed by the processor.

10. A computer program product comprising a computer program that, when executed by a processor, implements the AI large model-based user question response method of any one of claims 1-6.