Multi-Model Based Service Q&A Response Method, Device, Equipment and Storage Medium
By selecting multiple service scheduling models in the public large model container, a dual-model mechanism is generated, and the target response results are determined based on the response time, and an unanswered questions caused by the differences in processing capabilities of different large models are solved, achieving efficient and accurate question-and-answer responses.
Patent Information
- Application Number
- CN202510294315.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-03-13
AI Technical Summary
The differences in processing capabilities of different large models lead to the situation where answers are not asked in Q&A response. How to improve the accuracy of Q&A response while ensuring the efficiency of Q&A response.
By storing several service scheduling models in the public large model container, the first and second service scheduling models are selected based on the model static metadata, a dual-model mechanism is generated, the response time is compared to determine the target response result, and feedback to the user.
It realizes the accuracy of Q&A response while taking into account the response efficiency, and avoids the generative deviation and answer imbalance of single model prediction.
Smart Images

Figure CN119808965B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of data processing, and in particular, to a service question and answer response method, device, equipment, and storage medium based on multiple models. Background Art
[0002] With the increasing popularity of AIGC (Artificial Intelligence Generated Content) in all walks of life, large model platforms emerge in an endless stream. However, due to the different processing capacity characteristics of different large models, there are differences in the accuracy of knowledge responses to different services, and there may be situations of "answering off-topic". Therefore, how to improve the accuracy of question and answer responses to different business service requirements while ensuring the efficiency of question and answer responses has become an urgent problem to be solved at present. Summary of the Invention
[0003] The present invention provides a service question and answer response method, device, equipment, and storage medium based on multiple models to improve the accuracy of question and answer responses to different business service requirements while ensuring the efficiency of question and answer responses.
[0004] According to one aspect of the present invention, there is provided a service question and answer response method based on multiple models, the method comprising:
[0005] Obtain a service question and answer statement input by a target user. If there is no model operation metadata in the current public large model container, determine a first service scheduling model and a second service scheduling model according to the model static metadata and establish a current service session; a plurality of service scheduling models are stored in the public large model container;
[0006] Input the service question and answer statement into the first service scheduling model and the second service scheduling model respectively to obtain a first response result output by the first service scheduling model and a second response result output by the second service scheduling model;
[0007] Determine a target response result of the current service session according to a first response time of the first response result and a second response time of the second response result, and feedback the target response result to the target user.
[0008] According to another aspect of the present invention, there is provided a service question and answer response device based on multiple models, the device comprising:
[0009] The scheduling model determination module is used to obtain the service question and answer statement input by the target user. If the model running metadata does not exist in the current public large model container, the first service scheduling model and the second service scheduling model are determined according to the model static metadata and the current service session is established; the public large model container stores several service scheduling models;
[0010] a response result generating module, used for inputting the service question and answer statement into the first service scheduling model and the second service scheduling model respectively, and obtaining a first response result output by the first service scheduling model and a second response result output by the second service scheduling model;
[0011] A response result feedback module is used to determine a target response result of the current service session according to a first response time of the first response result and a second response time of the second response result, and feed back the target response result to the target user.
[0012] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising:
[0013] at least one processor; and
[0014] a memory communicatively connected to the at least one processor; wherein,
[0015] The memory stores a computer program that can be executed by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multi-model based service question and answer response method described in any embodiment of the present invention.
[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the multi-model-based service question and answer response method described in any embodiment of the present invention when executed.
[0017] The technical solution of the embodiment of the present invention determines the first service scheduling model and the second service scheduling model and establishes the current service session according to the model static metadata when receiving the service question and answer statement input by the user and determining that there is no model running metadata in the public large model container, and determines the target response result of the current service session according to the first response time when the first service scheduling model generates the first response result and the second response time when the second service scheduling model generates the second response result. This achieves the goal of balancing the response efficiency to the service question and answer statement and improving the response accuracy to the service question and answer statement, thereby achieving better adaptation to user questions and avoiding the generative bias and answer imbalance during single model prediction.
[0018] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0020] Figure 1 is a flowchart of a multi-model-based service question and answer response method according to Embodiment 1 of the present invention;
[0021] Figure 2 is a flowchart of a multi-model-based service question and answer response method according to Embodiment 2 of the present invention;
[0022] Figure 3 is a schematic structural diagram of a multi-model-based service question and answer response device according to Embodiment 3 of the present invention;
[0023] Figure 4 is a schematic structural diagram of an electronic device for implementing the multi-model-based service question and answer response method of the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the following clearly and completely describes the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some, rather than all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0025] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of the present invention are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0026] Embodiment 1
[0027] Figure 1 As shown in the flowchart of a multi-model-based service question and answer response method provided in Embodiment 1 of the present invention, this embodiment is applicable to the situation of accurately responding to user questions and answers with different service requirements. This method can be executed by a multi-model-based service question and answer response device, which can be implemented in the form of hardware and / or software, and the multi-model-based service question and answer response device can be configured in an electronic device. As Figure 1 shown, the method includes:
[0028] S110. Obtain the service question and answer statement input by the target user. If there is no model operation metadata in the current public large model container, determine the first service scheduling model and the second service scheduling model according to the model static metadata and establish the current service session; there are several service scheduling models stored in the public large model container.
[0029] S120. Input the service question and answer statement into the first service scheduling model and the second service scheduling model respectively, and obtain the first response result output by the first service scheduling model and the second response result output by the second service scheduling model.
[0030] S130. Determine the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, and feedback the target response result to the target user.
[0031] Among them, the target user can be a user with question and answer needs in the same or different service fields. The public large model container is pre-constructed by relevant technical personnel and is used to store public large models of several different model development platforms, that is, service scheduling models.
[0032] Exemplarily, a public large model container can be constructed through the Spring AI (Spring Artificial Intelligence) framework, and the model scheduling code of the service scheduling model under different large model development platforms can be sliced and developed, and a runtime container for different service scheduling models can be constructed. Specifically, Docker (containerization technology) can be used to encapsulate different service scheduling models into independent containers, and each container can contain model inference code, dependent environments, call interfaces, etc. To achieve the flexibility of model scheduling, the scheduling logic of the service scheduling model is split into multiple slices or modules, and each slice or module can be responsible for different functions. The core of the runtime container of each service scheduling model is to integrate the above slices and provide a unified scheduling service. When the service scheduling model is scheduled, the corresponding runtime container is started.
[0033] Among them, the model runtime metadata can include the model call time, the model service domain, the model response accuracy, etc.; the model static metadata can include the tocken quantity and the model service price. Different service scheduling models each correspond to their own model static metadata and model runtime metadata, and are uniformly stored in the public large model container. It should be noted that if a certain service scheduling model stored in the public large model container has not been called in the current and historical time periods, there is no model runtime metadata for it; the model runtime metadata is dynamically updated with the response results of the service scheduling model.
[0034] When the target user conducts a large model service question and answer, a scheduling mechanism for the large model is constructed, and it is judged whether there is model runtime data in the public large model container in the current time period; if it exists, the service scheduling model is selected according to the model runtime metadata; if it does not exist, the service scheduling model is selected according to the model static metadata.
[0035] For the scenario where there is no model runtime data in the public large model container in the current time period, according to the model static metadata, the first service scheduling model and the second service scheduling model are determined. Specifically, according to the model service price in the model static metadata, several service scheduling models are sorted from low to high in price, and the two service scheduling models with lower prices are selected as the first service scheduling model and the second service scheduling model, a dual-model mechanism scheduling is constructed, and at the same time, the current service session is established, and the runtime containers corresponding to the first service scheduling model and the second service scheduling model are mounted respectively. At the same time, the first service scheduling model and the second service scheduling model are triggered to run.
[0036] The service question and answer statements are input into the first service scheduling model and the second service scheduling model respectively, and the first response result output by the first service scheduling model and the second response result output by the second service scheduling model are obtained. At the same time, the runtime containers of the two service scheduling models are monitored in real time, and the running time of the runtime containers is recorded as the response time of the models. According to the first response time of the first response result and the second response time of the second response result, the target response result of the current service session is determined.
[0037] Optionally, a target response result of the current service session is determined based on the first response time of the first response result and the second response time of the second response result, including: if the first response time of the first response result is greater than the second response time of the second response result, using the second response result as the target response result of the current service session; or, if the first response time of the first response result is not greater than the second response time of the second response result, using the first response result as the target response result of the current service session.
[0038] Specifically, by comparing the first response time of the first response result and the second response time of the second response result, the response result with the shortest response time in the two service scheduling models is used as the target response result of the current service session, and the target response result is fed back to the target user.
[0039] Furthermore, if the target response result is the first response result, the second response result is stored as a temporary result, and when the target user initiates the same service question and answer statement in the current service session, the second response result is directly fed back to the target user.
[0040] The technical solution of the embodiment of the present invention determines the first service scheduling model and the second service scheduling model and establishes the current service session according to the model static metadata when receiving the service question and answer statement input by the user and determining that there is no model running metadata in the public large model container, and determines the target response result of the current service session according to the first response time when the first service scheduling model generates the first response result and the second response time when the second service scheduling model generates the second response result. This achieves the goal of balancing the response efficiency to the service question and answer statement and improving the response accuracy to the service question and answer statement, thereby achieving better adaptation to user questions and avoiding the generative bias and answer imbalance during single model prediction.
[0041] It should be noted that there can be different Q&A scenarios for the current service session. For example, the same service Q&A statement can be initiated within a certain time period, or other Q&A statements similar to the service Q&A statement can be initiated within a certain time period, or other Q&A statements different from the service Q&A statement can be initiated within a certain time period, etc. For different Q&A scenarios, the present application adopts different Q&A statement processing methods to ensure the response efficiency and response accuracy of the service Q&A statement.
[0042] In an alternative embodiment, after determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, it further includes: determining a cached response result according to the target response result and storing the cached response result; and generating at least one associated Q&A statement corresponding to the service Q&A statement; and generating a Q&A waiting time according to the target response result.
[0043] If the target response result is the first response result, then the second response result is determined as the cached response result; if the target response result is the second response result, then the first response result is determined as the cached response result. The cached response result is temporarily stored in the public large model container in the current service session.
[0044] After obtaining the target response result, a question chain of the service Q&A statement is constructed, where the question chain includes at least one associated Q&A statement corresponding to the service Q&A statement. For example, associated Q&A statements can be generated based on context awareness or knowledge graph association, etc.
[0045] Among them, the Q&A waiting time is the time to wait for the target user to input a new Q&A statement after the target response result is fed back to the target user. For example, the Q&A waiting time can be preset by relevant technical personnel according to actual needs, actual experience, experiments, etc. Or, to further improve the rationality of the Q&A waiting time, it can also be generated according to the target response result. Specifically, the Q&A waiting time is calculated based on the content size of the target response result at a rate of 25 characters per second. For example, if the target response result includes 100 characters, then the Q&A waiting time is 4 seconds.
[0046] Since the model expertise areas and model processing capabilities of different service scheduling models are different, the response results of different service scheduling models for the same service question-and-answer statement also vary. When the user initiates the same service question-and-answer statement, it may indicate a lower satisfaction or acceptance of the previous response result. Therefore, in an alternative embodiment, after determining the target response result of the current service session based on the first response time of the first response result and the second response time of the second response result, it further includes: if an intermediate question-and-answer statement input by the target user is received within the question-and-answer waiting time and the intermediate question-and-answer statement is the same as the service question-and-answer statement, then the cached response result is fed back to the target user in the current service session.
[0047] In the current service session, if an intermediate question-and-answer statement that is the same as the service question-and-answer statement is received within the question-and-answer waiting time, the cached response result is fed back to the target user.
[0048] Through the question-and-answer response method based on the dual-model mechanism, the above technical solution directly feeds back the cached response result to the target user when the same service question-and-answer statement is received, improving the response accuracy of the service question-and-answer service, avoiding falling into the response deviation of a single model, and being able to achieve fast response and improve the response efficiency.
[0049] In an alternative embodiment, after determining the target response result of the current service session based on the first response time of the first response result and the second response time of the second response result, it further includes: if an intermediate question-and-answer statement input by the target user is received within the question-and-answer waiting time and the intermediate question-and-answer statement is the same as the associated question-and-answer statement, then the previous historical question-and-answer statement of the intermediate question-and-answer statement is determined; the service scheduling model that responded to the historical question-and-answer statement is used to respond to the intermediate question-and-answer statement to obtain an intermediate response result, and the intermediate response result is fed back to the target user in the current service session.
[0050] Exemplarily, if the previous historical question-and-answer statement of the intermediate question-and-answer statement is statement a and the service scheduling model that responded to statement a is model A, if an intermediate question-and-answer statement that is the same as any one of several associated question-and-answer statements is received within the question-and-answer waiting time, then model A is used to generate the intermediate response result of the intermediate question-and-answer statement, and the intermediate response result is fed back to the target user in the current service session.
[0051] It can be understood that if, in the current service session, the target user asks an associated question-and-answer statement in the question chain, it may indicate that the hit rate of the service question-and-answer model that responded to the previous historical question-and-answer statement is relatively high, and the target user has a high satisfaction or acceptance of the model response result. Therefore, the previous service question-and-answer model can still be used to continue to respond to the result of the intermediate question-and-answer statement to ensure the response accuracy and response efficiency of the question-and-answer statement.
[0052] In an optional embodiment, after determining the target response result of the current service session based on the first response time of the first response result and the second response time of the second response result, it further includes: if intermediate question-and-answer statements identical to the service question-and-answer statement are obtained within a preset number of loops, then determine a third service scheduling model and a fourth service scheduling model according to the model static metadata; input the intermediate question-and-answer statements into the third service scheduling model and the fourth service scheduling model respectively, and obtain a third response result output by the third service scheduling model and a fourth response result output by the fourth service scheduling model; determine the intermediate response result according to the third response time of the third response result and the fourth response time of the fourth response result, and feedback the intermediate response result to the target user in the current service session.
[0053] If intermediate question-and-answer statements identical to the service question-and-answer statement are obtained within a preset number of loops, it means that the response result hit rates of the service scheduling models used within a certain number of times for the service question-and-answer statement are not high, and the satisfaction and acceptance of the target user for the response result are relatively low. Therefore, a dual service scheduling model is reselected based on the model static metadata. Among them, the number of loops can be preset by relevant technical personnel according to actual needs or experience. For example, the number of loops can be set to 3 times or 4 times, etc.
[0054] Specifically, select the third service scheduling model and the fourth service scheduling model from the remaining service scheduling models according to the model static metadata stored in the public large model container. Among them, the remaining service scheduling models are the service scheduling models in the public large model container except for the first service scheduling model and the second service scheduling model.
[0055] Select the two service scheduling models with lower prices as the third service scheduling model and the fourth service scheduling model according to the model service price in the model static metadata. Record in real time the third response time of the third service scheduling model to respond to the intermediate question-and-answer statement, and the fourth response time of the fourth service scheduling model to respond to the intermediate question-and-answer statement. Feed back the intermediate response result with the shortest time used among the third response time and the fourth response time to the target user.
[0056] It can be understood that when the target user repeatedly asks the same service question-and-answer statement, it means that the satisfaction or acceptance of the target user for the current dual service scheduling model is relatively low, and the hit rate of the dual models is low. Therefore, re-determine the response results of the dual service scheduling model for generating the intermediate question-and-answer statements of the repeated questions to ensure the response accuracy of the question-and-answer statements of the target user.
[0057] In an alternative embodiment, after determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, the method further includes: if no intermediate question-and-answer statement identical to the service question-and-answer statement is received within the question-and-answer waiting time, or if the received intermediate question-and-answer statement is different from the associated question-and-answer statement, end the current service session and update the model operation metadata.
[0058] It should be noted that if no intermediate question-and-answer statement identical to the service question-and-answer statement is received within the question-and-answer waiting time; or if the received intermediate question-and-answer statement is completely different from the associated question-and-answer statement, it indicates that the target user starts a new round of service session, that is, a new question is raised. Then, terminate the current service session, and update the model operation metadata and model static metadata of the involved service scheduling model according to at least one round of question-and-answer statements and the response results of each question-and-answer statement in the previous service session.
[0059] This embodiment further provides a method for updating the model operation metadata. In an alternative embodiment, updating the model operation metadata includes: determining the model input question-and-answer statement, model response result, model response time, associated question hit result, and main question hit result of the same service scheduling model in the current service session; determining the model service domain of the service scheduling model according to the model input question-and-answer statement and model response result of the service scheduling model; determining the accuracy rate of the model service domain of the service scheduling model according to the associated question hit result and main question hit result; and updating the model operation metadata of the service scheduling model according to the model service domain and its accuracy rate of the service scheduling model, and the model response time.
[0060] It should be noted that for any service session, the listener records the relevant information of the question-and-answer response of the service scheduling model involved in the service session, such as the response time of the model to the question-and-answer statement, and the associated question hit rate and main question hit rate can be calculated according to the question-and-answer statement, question-and-answer response, and the question-and-answer statement raised again by the user for the question-and-answer response.
[0061] Specifically, in the current service session, determine the model input statements of the same service scheduling model and the model response results for these model input statements. Among them, the associated question hit result is the hit rate of the associated Q&A statements on the question chain; the main question hit result is the hit rate of the service Q&A statements initiated by the target user. The average of the associated question hit rate and the main question hit rate can be used as the accuracy rate of the model service area of this service scheduling model. The model service area indicates the technical area that this service scheduling model is good at. When the accuracy rate of this service scheduling model for the current model service area is high, it means that this service scheduling model is better at predicting questions in this model service area and has a high accuracy rate for predicting questions in this model service area. When the accuracy rate of this service scheduling model for the current model service area is low, it means that this service scheduling model is not good at predicting questions in this model service area and has a low accuracy rate for predicting questions in this model service area.
[0062] Use the model service area and its accuracy rate of this service scheduling model in the current service session, as well as the model response time, to update the model operation metadata of this service scheduling model. Specifically, if the model operation data of this service scheduling model is not stored in the current public large model container, then store the model service area and its accuracy rate of this service scheduling model in the current service session, as well as the model response time, in the public large model container as the model operation metadata of this service scheduling model. If the model operation data of this service scheduling model is stored in the current public large model container, then use the model service area and its accuracy rate of this service scheduling model in the current service session, as well as the model response time, to update the model operation metadata of this service scheduling model stored in the public large model container.
[0063] By updating the model operation metadata of the model service scheduling model after each service session ends, the dynamic update of the model operation metadata is realized, which facilitates the more accurate selection of the service scheduling model for response when selecting the dual model, and thus improves the response accuracy of the service Q&A statements.
[0064] This embodiment also provides a method for updating the model static metadata. In an alternative embodiment, updating the model static metadata includes: determining the Tocken consumption of this service session according to the model response result of this service scheduling model; obtaining the remaining Tocken amount of this service scheduling model, and determining the current Tocken amount of this service scheduling model according to the Tocken consumption and the remaining Tocken amount; and updating the model static metadata of this service scheduling model according to the current Tocken amount.
[0065] Among them, the remaining amount of Tokens in the service scheduling model is the number of Tokens remaining in the previous service session cycle.
[0066] Specifically, the difference between the remaining amount of Tokens and the consumed amount of Tokens is determined as the current number of Tokens in the service scheduling model. The model static metadata of the service scheduling model is updated using the current number of Tokens.
[0067] By updating the model static metadata of the model service scheduling model after each service session ends, the dynamic update of the model static metadata is realized, which facilitates more accurate selection of the service scheduling model for response when selecting between dual models, thereby improving the response accuracy of service question-and-answer statements.
[0068] Embodiment 2
[0069] Figure 2 The flowchart of a service question-and-answer response method based on multiple models provided by Embodiment 2 of the present invention is optimized and improved on the basis of the above technical solutions.
[0070] Further, after the step of "obtaining the service question-and-answer statement input by the target user", add the step of "if there is model operation metadata in the current public large model container, determine the field to which the service question-and-answer statement belongs; according to the model service fields and their accuracies of each service scheduling model in the model operation metadata, based on the field to which the service question-and-answer statement belongs, determine the first hit scheduling model and the second hit scheduling model; generate the target response result of the service question-and-answer statement using the first hit scheduling model and the second hit scheduling model." to improve the selection method of the service scheduling model.
[0071] It should be noted that for parts not detailed in the embodiments of the present invention, reference can be made to the descriptions of other embodiments. As Figure 2 shown, the method includes the following specific steps:
[0072] S210. Obtain the service question-and-answer statement input by the target user. If there is no model operation metadata in the current public large model container, execute S220 - S240; if there is model operation metadata in the current public large model container, execute S250 - S270.
[0073] S220. Determine the first service scheduling model and the second service scheduling model according to the model static metadata and establish the current service session; several service scheduling models are stored in the public large model container.
[0074] S230. Input the service Q&A statement into the first service scheduling model and the second service scheduling model respectively, to obtain the first response result output by the first service scheduling model and the second response result output by the second service scheduling model.
[0075] S240. Determine the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, and feedback the target response result to the target user.
[0076] S250. Determine the domain to which the service Q&A statement belongs.
[0077] Specifically, a pre-trained domain prediction model can be used to predict the domain of the service Q&A statement, or text analysis can also be performed on the service Q&A statement to determine the belonging domain. This embodiment does not limit this.
[0078] S260. Based on the model service domains and their accuracies of the service scheduling models in the model operation metadata, determine the first hit scheduling model and the second hit scheduling model according to the domain to which the service Q&A statement belongs.
[0079] Specifically, select at least one service scheduling model from the service scheduling models whose model service domains are the same as or similar to the domain to which the service Q&A statement belongs, and determine the two service scheduling models with higher accuracies according to the accuracies of the model service domains of the selected service scheduling models, as the first hit scheduling model and the second hit scheduling model.
[0080] S270. Use the first hit scheduling model and the second hit scheduling model to generate the target response result of the service Q&A statement.
[0081] Specifically, input the service Q&A statement into the first hit scheduling model and the second hit scheduling model respectively, to obtain the first response result output by the first hit scheduling model and the second response result output by the second service scheduling model. Take the one with the shortest response time among the two service scheduling models as the target response result and feedback it to the target user. Exemplarily, if the first response time of the first response result is greater than the second response time of the second response result, then take the second response result as the target response result of the current service session; or, if the first response time of the first response result is not greater than the second response time of the second response result, then take the first response result as the target response result of the current service session.
[0082] It should be noted that among the two service scheduling models selected based on the model operation metadata, there may be models with a relatively small remaining number of Tokens or a relatively high model service price, and such models are no longer applicable to the model service scheduling of this session service. Therefore, in response to such a situation, the first hit scheduling model and the second hit scheduling model obtained through screening can be further screened to ensure that the model used for answering question statements is the optimal model.
[0083] In an alternative embodiment, after determining the first hit scheduling model and the second hit scheduling model based on the model service fields and their accuracies of each service scheduling model in the model operation metadata and the field to which the service question statement belongs, it further includes:
[0084] Step a: According to the model operation metadata and model static metadata respectively corresponding to the first hit scheduling model and the second hit scheduling model, determine whether the first hit scheduling model and / or the second hit scheduling model meet the preset model replacement conditions.
[0085] Step b: If so, determine a candidate scheduling model according to the operation metadata and model static metadata of other service scheduling models, and use the candidate scheduling model to replace the first hit scheduling model and / or the second hit scheduling model that meet the model replacement conditions.
[0086] Among them, other service scheduling models are the service scheduling models in the current public large model container except the first hit scheduling model and the second hit scheduling model.
[0087] Among them, the model replacement conditions can be preset by relevant technical personnel according to actual needs. For example, the model replacement conditions can be that the remaining amount of model Tokens is lower than the preset remaining amount threshold, the model response time is higher than the preset response duration, and / or the model service price is higher than the preset price threshold, etc.
[0088] If the first hit scheduling model and / or the second hit scheduling model meet the preset model replacement conditions, determine a candidate scheduling model according to the operation metadata and model static metadata of other service scheduling models. Specifically, a service scheduling model with sufficient remaining Tokens, a relatively low model service price, and a relatively short model response duration determined according to the model static metadata can be selected from the model operation metadata of other service scheduling models as the candidate scheduling model. Use the candidate scheduling model to replace the first hit scheduling model and / or the second hit scheduling model that meet the model replacement conditions.
[0089] When the technical solution of this embodiment determines that there is model operation metadata in the current public large model container, it determines the field to which the service question-and-answer statement belongs, and based on the model service fields and their accuracies of each service scheduling model in the model operation metadata, determines the first hit scheduling model and the second hit scheduling model according to the field to which the service question-and-answer statement belongs; uses the first hit scheduling model and the second hit scheduling model to generate the target response result of the service question-and-answer statement, realizes the selection of dual models based on the model operation metadata in the scenario where there is model operation metadata, thereby performing question-and-answer statement prediction, taking into account the response efficiency of the service question-and-answer statement while improving the response accuracy of the service question-and-answer statement, so as to better adapt to the user's questions and avoid situations such as generative bias and answer imbalance when performing single-model prediction.
[0090] Embodiment III
[0091] Figure 3 FIG. 7 is a schematic structural diagram of a service question-and-answer response device based on multiple models provided by Embodiment III of the present invention. A service question-and-answer response device based on multiple models provided by an embodiment of the present invention can be applied to the situation of accurately responding to user questions with different service requirements. The service question-and-answer response device based on multiple models can be implemented in the form of hardware and / or software, such as Figure 3 shown, the device specifically includes: a scheduling model determination module 301, a response result generation module 302, and a response result feedback module 303. Among them,
[0092] The scheduling model determination module 301 is configured to obtain a service question-and-answer statement input by a target user. If there is no model operation metadata in the current public large model container, determine a first service scheduling model and a second service scheduling model according to the model static metadata and establish the current service session; several service scheduling models are stored in the public large model container;
[0093] The response result generation module 302 is configured to input the service question-and-answer statement into the first service scheduling model and the second service scheduling model respectively, and obtain a first response result output by the first service scheduling model and a second response result output by the second service scheduling model;
[0094] The response result feedback module 303 is configured to determine the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, and feedback the target response result to the target user.
[0095] The technical solution of the embodiment of the present invention determines the first service scheduling model and the second service scheduling model and establishes the current service session according to the model static metadata when receiving the service question and answer statement input by the user and determining that there is no model running metadata in the public large model container, and determines the target response result of the current service session according to the first response time when the first service scheduling model generates the first response result and the second response time when the second service scheduling model generates the second response result. This achieves the goal of balancing the response efficiency to the service question and answer statement and improving the response accuracy to the service question and answer statement, thereby achieving better adaptation to user questions and avoiding the generative bias and answer imbalance during single model prediction.
[0096] Optionally, the response result feedback module 303 is specifically used to:
[0097] If the first response time of the first response result is greater than the second response time of the second response result, the second response result is used as the target response result of the current service session; or,
[0098] If the first response time of the first response result is not greater than the second response time of the second response result, the first response result is used as the target response result of the current service session.
[0099] Optionally, the device further comprises:
[0100] A result storage module is used to determine a cache response result according to the target response result after determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, and store the cache response result; and
[0101] a related statement generating module, configured to generate at least one related question and answer statement corresponding to the service question and answer statement; and
[0102] The waiting time generation module is used to generate the question-answering waiting time according to the target response result.
[0103] Optionally, the device further comprises:
[0104] A cached result response module is used to feed back the cached response result to the target user in the current service session after determining the target response result of the current service session based on the first response time of the first response result and the second response time of the second response result, if an intermediate question and answer statement input by the target user is received within the question and answer waiting time, and the intermediate question and answer statement is the same as the service question and answer statement.
[0105] Optionally, the device further comprises:
[0106] A historical statement determination module, configured to, after determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, if an intermediate Q&A statement input by the target user is received within the Q&A waiting time and the intermediate Q&A statement is the same as the associated Q&A statement, determine the previous historical Q&A statement of the intermediate Q&A statement;
[0107] A first intermediate result response module, configured to use a service scheduling model for responding to the historical Q&A statement to respond to the intermediate Q&A statement, obtain an intermediate response result, and feedback the intermediate response result to the target user in the current service session.
[0108] Optionally, the apparatus further includes:
[0109] A model screening module, configured to, after determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, if intermediate Q&A statements identical to the service Q&A statement are obtained within a preset number of loops, determine a third service scheduling model and a fourth service scheduling model according to model static metadata;
[0110] An intermediate statement processing module, configured to input the intermediate Q&A statement into the third service scheduling model and the fourth service scheduling model respectively, to obtain a third response result output by the third service scheduling model and a fourth response result output by the fourth service scheduling model;
[0111] A second intermediate result response module, configured to determine an intermediate response result according to the third response time of the third response result and the fourth response time of the fourth response result, and feedback the intermediate response result to the target user in the current service session.
[0112] Optionally, the apparatus further includes:
[0113] An operation metadata update module, configured to, after determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, if no intermediate Q&A statement identical to the service Q&A statement is received within the Q&A waiting time, or an intermediate Q&A statement received is different from the associated Q&A statement, end the current service session and update model operation metadata.
[0114] Optionally, the operation metadata update module is specifically configured to:
[0115] Determine the model input Q&A statement, model response result, model response time, associated question hit result, and main question hit result of the same service scheduling model in the current service session;
[0116] Determine the model service domain of the service scheduling model according to the model input Q&A statement and model response result of the service scheduling model;
[0117] Determine the accuracy rate of the model service domain of the service scheduling model according to the associated question hit result and main question hit result;
[0118] Update the model operation metadata of the service scheduling model according to the model service domain of the service scheduling model, its accuracy rate, and the model response time.
[0119] Optionally, the device further includes:
[0120] Tocken consumption rate determination module, configured to determine the Tocken consumption amount of the current service session according to the model response result of the service scheduling model;
[0121] Current Tocken amount determination module, configured to obtain the remaining Tocken amount of the service scheduling model, and determine the current Tocken quantity of the service scheduling model according to the Tocken consumption amount and the remaining Tocken amount;
[0122] Static metadata update module, configured to update the model static metadata of the service scheduling model according to the current Tocken quantity.
[0123] Optionally, the device further includes:
[0124] Model domain determination module, configured to, after obtaining the service Q&A statement input by the target user, if there is model operation metadata in the current public large model container, determine the domain to which the service Q&A statement belongs;
[0125] Hit scheduling model determination module, configured to determine the first hit scheduling model and the second hit scheduling model based on the domain to which the service Q&A statement belongs according to the model service domain of each service scheduling model in the model operation metadata and its accuracy rate;
[0126] Response module, configured to generate the target response result of the service Q&A statement by using the first hit scheduling model and the second hit scheduling model.
[0127] Optionally, the device further includes:
[0128] A replacement condition judgment module, configured to, after determining a first hit scheduling model and a second hit scheduling model based on the model service fields and their accuracies of the service scheduling models in the model operation metadata and the field to which the service question-and-answer statement belongs, judge whether the first hit scheduling model and / or the second hit scheduling model meet a preset model replacement condition according to the model operation metadata and model static metadata respectively corresponding to the first hit scheduling model and the second hit scheduling model;
[0129] A candidate model determination module, configured to, if the first hit scheduling model and / or the second hit scheduling model meet the preset model replacement condition, determine a candidate scheduling model according to the operation metadata and model static metadata of other service scheduling models, and use the candidate scheduling model to replace the first hit scheduling model and / or the second hit scheduling model that meet the model replacement condition; the other service scheduling models are the service scheduling models in a current public large model container except the first hit scheduling model and the second hit scheduling model.
[0130] The service question-and-answer response device based on multiple models provided by the embodiments of the present invention can execute the service question-and-answer response method based on multiple models provided by any embodiment of the present invention, and has functional modules and beneficial effects corresponding to the execution of the method.
[0131] Embodiment 4
[0132] Figure 4 FIG. shows a schematic structural diagram of an electronic device 40 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processors, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described herein and / or claimed.
[0133] As Figure 4As shown, the electronic device 40 includes at least one processor 41 and a memory communicatively connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. The memory stores computer programs executable by the at least one processor. The processor 41 can execute various appropriate actions and processes according to the computer programs stored in the read-only memory (ROM) 42 or the computer programs loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0134] Multiple components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a disk, an optical disc, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0135] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include but are not limited to a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the multi-model-based service question-and-answer response method.
[0136] In some embodiments, the multi-model-based service question-and-answer response method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the multi-model-based service question-and-answer response method described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the multi-model-based service question-and-answer response method by any other appropriate means (e.g., by means of firmware).
[0137] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.
[0138] The computer program for implementing the methods of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the computer programs, when executed by the processor, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The computer program can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.
[0139] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0141] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected to each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0142] The computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0143] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is imposed herein.
[0144] The above specific embodiments do not constitute a limitation on the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A service question-answering response method based on multiple models, characterized in that, Including: Obtain the service Q&A statement input by the target user. If there is no model operation metadata in the current public large model container, determine the first service scheduling model and the second service scheduling model according to the model static metadata and establish the current service session; There are several service scheduling models stored in the public large model container; Input the service Q&A statement into the first service scheduling model and the second service scheduling model respectively, and obtain the first response result output by the first service scheduling model and the second response result output by the second service scheduling model; Determine the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, and feedback the target response result to the target user; Generate at least one associated Q&A statement corresponding to the service Q&A statement; And, Generate the Q&A waiting time according to the target response result; If an intermediate Q&A statement input by the target user is received within the Q&A waiting time and the intermediate Q&A statement is the same as the associated Q&A statement, determine the previous historical Q&A statement of the intermediate Q&A statement; use the service scheduling model that responded to the historical Q&A statement to respond to the intermediate Q&A statement to obtain an intermediate response result, and feedback the intermediate response result to the target user in the current service session; Or, If an intermediate Q&A statement input by the target user is received within the Q&A waiting time and intermediate Q&A statements identical to the service Q&A statement are obtained within the preset number of loops, determine the third service scheduling model and the fourth service scheduling model according to the model static metadata; input the intermediate Q&A statement into the third service scheduling model and the fourth service scheduling model respectively, and obtain the third response result output by the third service scheduling model and the fourth response result output by the fourth service scheduling model; Determine the intermediate response result according to the third response time of the third response result and the fourth response time of the fourth response result, and feedback the intermediate response result to the target user in the current service session.
2. The method according to claim 1, characterized in that, The determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result includes: If the first response time of the first response result is greater than the second response time of the second response result, use the second response result as the target response result of the current service session; or, If the first response time of the first response result is not greater than the second response time of the second response result, use the first response result as the target response result of the current service session.
3. The method according to claim 1, wherein The method further includes: after ending the current service session, updating the model operation metadata includes: Determine the model input Q&A statement, model response result, model response time, associated question hit result, and main question hit result of the same service scheduling model in the current service session; Determine the model service area of the service scheduling model according to the model input Q&A statement and the model response result of the service scheduling model; Determine the accuracy rate of the model service area of the service scheduling model according to the associated question hit result and the main question hit result; Update the model operation metadata of the service scheduling model according to the model service area of the service scheduling model, its accuracy rate, and the model response time.
4. The method according to claim 1, wherein The method further includes: Determine the Token consumption of the current service session according to the model response result of the corresponding service scheduling model; Obtain the remaining Token amount of the corresponding service scheduling model, and determine the current Token amount of the service scheduling model according to the Token consumption and the remaining Token amount; Update the model static metadata of the service scheduling model according to the current Token amount.
5. The method according to claim 1, wherein After obtaining the service Q&A statement input by the target user, it further includes: If there is model operation metadata in the current public large model container, determine the domain to which the service Q&A statement belongs; Based on the domain to which the service Q&A statement belongs, determine the first hit scheduling model and the second hit scheduling model according to the model service areas of the service scheduling models in the model operation metadata and their accuracy rates; Generate the target response result of the service Q&A statement by using the first hit scheduling model and the second hit scheduling model.
6. The method according to claim 5, wherein After determining the first hit scheduling model and the second hit scheduling model based on the domain to which the service Q&A statement belongs according to the model service areas of the service scheduling models in the model operation metadata and their accuracy rates, it further includes: Judge whether the first hit scheduling model and / or the second hit scheduling model meet the preset model replacement conditions according to the model operation metadata and the model static metadata corresponding to the first hit scheduling model and the second hit scheduling model; If so, determine the candidate scheduling model according to the operation metadata and the model static metadata of other service scheduling models, and use the candidate scheduling model to replace the first hit scheduling model and / or the second hit scheduling model that meet the model replacement conditions; the other service scheduling models are the service scheduling models in the current public large model container except the first hit scheduling model and the second hit scheduling model.
7. A service question-answering response device based on multiple models, characterized in that, It includes: A scheduling model determination module, configured to obtain the service Q&A statement input by the target user. If there is no model operation metadata in the current public large model container, determine the first service scheduling model and the second service scheduling model according to the model static metadata and establish the current service session; There are several service scheduling models stored in the public large model container; A response result generation module, configured to input the service Q&A statement into the first service scheduling model and the second service scheduling model respectively to obtain the first response result output by the first service scheduling model and the second response result output by the second service scheduling model; A response result feedback module, configured to determine a target response result of the current service session according to a first response time of the first response result and a second response time of the second response result, and feedback the target response result to the target user; An associated statement generation module, configured to generate at least one associated Q&A statement corresponding to the service Q&A statement; And, A waiting time generation module, configured to generate a Q&A waiting time according to the target response result; A historical statement determination module, configured to, after determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, if an intermediate Q&A statement input by the target user is received within the Q&A waiting time and the intermediate Q&A statement is the same as the associated Q&A statement, determine the previous historical Q&A statement of the intermediate Q&A statement; use a service scheduling model that responds to the historical Q&A statement to respond to the intermediate Q&A statement to obtain an intermediate response result, and feedback the intermediate response result to the target user in the current service session; Or, A model screening module, configured to, after determining the target response result of the current service session according to the first response time of the first response result and the second response time of the second response result, if intermediate Q&A statements that are the same as the service Q&A statement are obtained within a preset number of loop times, determine a third service scheduling model and a fourth service scheduling model according to model static metadata; input the intermediate Q&A statement into the third service scheduling model and the fourth service scheduling model respectively to obtain a third response result output by the third service scheduling model and a fourth response result output by the fourth service scheduling model; Determine an intermediate response result according to a third response time of the third response result and a fourth response time of the fourth response result, and feedback the intermediate response result to the target user in the current service session.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the multi-model-based service Q&A response method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a processor to implement the multi-model-based service Q&A response method according to any one of claims 1-6 when executed.
Citation Information
Patent Citations
Information processing method and device, electronic equipment and storage medium
CN119271781A
Question answer obtaining method and device based on distillation model, equipment and medium
CN119357358A