Multi-model cooperative service method and device, equipment and storage medium
Through the multi-model collaborative service method, dynamically match the task processing model, the problem of imbalance in resource allocation of large language models is solved, and the optimal allocation of resources and the maximum utilization of computing resources is realized.
Patent Information
- Application Number
- CN202510797403.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2045-06-16
AI Technical Summary
In the prior art, there is an imbalance in resource allocation of large language models, high-performance computing resources are over-occupied, while small and medium-sized computing resources are idle, making it difficult to achieve optimal resource allocation.
Through the multi-model collaborative service method, the task difficulty is used to judge the difficulty of the large model to analyze the current question data, dynamically match the optimal task processing model for content generation, and adjust the output results based on the feedback results to avoid resource mismatch.
The maximum utilization of computing resources is achieved, the average computing cost is reduced, the resource allocation optimization rate is improved, and the quality of answers is ensured.
Smart Images

Figure CN120297428A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of model services, and particularly to a multi-model collaborative service method, apparatus, device, and storage medium. Background Art
[0002] With the rapid development of large language model technology, the inference costs of models of different scales show significant differences. From small models deployed on a single card to ultra-large models that require multi-machine cluster support, the resource requirements span a wide range. Taking the series of models released by DeepSeek as an example, its distilled model Qwen1.5B only requires 3GB of video memory to achieve stable deployment, while the R1 model requires up to 1.4TB of video memory resources. In actual application scenarios, users generally tend to choose the model with the strongest computing power to handle all types of tasks. This usage pattern has led to a serious problem of unbalanced resource allocation: high-performance computing resources are over-occupied, while medium and small computing resources are idle.
[0003] To address this problem, in related technologies, intention recognition is mainly achieved by analyzing a large amount of user behavior data, and accordingly, a large model of the corresponding scale is called to execute the corresponding task. However, under this static allocation mechanism, once the model is selected, it cannot be dynamically adjusted, making it difficult to achieve the optimal allocation of resources, and still difficult to achieve the optimized allocation of computing resources. Summary of the Invention
[0004] The main purpose of the embodiments of this application is to propose a multi-model collaborative service method, apparatus, device, and storage medium to improve the resource allocation optimization rate of large language models.
[0005] To achieve the above object, the first aspect of the embodiments of this application proposes a multi-model collaborative service method, including: Input the current question data in the target dialogue into a task difficulty judgment large model, and perform difficulty analysis on the current question data based on the historical question data of the target dialogue to obtain the task difficulty; Perform performance matching among at least one task processing model according to the task difficulty to obtain a target task processing model. At least two of the task processing models have different computing power requirements, and different computing power requirements indicate matching different task difficulties; Send the current question data to the target task processing model for content generation to obtain answer data, and obtain the feedback result corresponding to the answer data. Based on the feedback result and the answer data, obtain the output result for the current question data.
[0006] In some embodiments, the performing difficulty analysis on the current question data based on the historical question data of the target dialogue to obtain the task difficulty includes: Parse the current question data to obtain a parsing result; If the parsing result indicates that the current question data includes at least one continuous action identifier, obtain the previous sentence of the current question data in the historical question data as the current question data until the current question data does not include the continuous action identifier, and use the task difficulty judgment large model to perform difficulty parsing on the last current question data to obtain the task difficulty.
[0007] In some embodiments, the at least one task processing model is divided into model groups corresponding to different preset difficulties respectively, and each model group includes at least one of the task processing models. The performance matching in at least one task processing model according to the task difficulty to obtain a target task processing model includes: Select the preset difficulty that is consistent with the task difficulty, and use the corresponding model group as the target model group; Select one of the task processing models from the target model group as the target task processing model according to a preset selection strategy.
[0008] In some embodiments, the obtaining the output result for the current question data based on the feedback result and the answer data includes: If the feedback result is a positive feedback, use the answer data as the output result; If the feedback result is a negative feedback, update the task difficulty and the corresponding answer data until the output result is obtained.
[0009] In some embodiments, the updating the task difficulty and the corresponding answer data until the output result is obtained includes: Update the task difficulty to the next level of difficulty; Based on the next level of difficulty, re-determine the target task processing model, update the corresponding answer data and the feedback result until the next level of difficulty is the highest difficulty or the feedback result is the positive feedback, and use the last answer data as the output result.
[0010] In some embodiments, after obtaining the feedback result corresponding to the answer data, the method further includes: Obtain the task difficulty, the feedback result, and the last current question data as fine-tuning sample data; Use the fine-tuning sample data to perform parameter fine-tuning on the task difficulty judgment large model and update the task difficulty judgment large model.
[0011] In some embodiments, selecting one of the task processing models from the target model group according to a preset selection strategy to be the target task processing model includes: Obtaining the model interface corresponding to each task processing model in the target model group; Selecting the model interface according to a random selection strategy or a polling selection strategy, and invoking the model interface to obtain the corresponding target task processing model.
[0012] To achieve the above object, a second aspect of the embodiments of the present application proposes a multi-model collaborative service device, including: A task difficulty acquisition module: configured to input the current question data in the target dialogue into a task difficulty judgment large model, perform difficulty analysis on the current question data based on the historical question data of the target dialogue, and obtain the task difficulty; A task model selection module: configured to perform performance matching among at least one task processing model according to the task difficulty to obtain a target task processing model, where the computing power requirements of at least two task processing models are different, and different computing power requirements indicate matching different task difficulties; A feedback update module: configured to send the current question data to the target task processing model for content generation to obtain answer data, and obtain a feedback result corresponding to the answer data, and obtain an output result for the current question data based on the feedback result and the answer data.
[0013] To achieve the above object, a third aspect of the embodiments of the present application proposes an electronic device, where the electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.
[0014] To achieve the above object, a fourth aspect of the embodiments of the present application proposes a storage medium, where the storage medium is a storage medium, the storage medium stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.
[0015] The multi-model collaborative service method, device, equipment, and storage medium proposed in the embodiments of this application input the current question data in the target conversation into the task difficulty judgment large model, and analyze the difficulty of the current question data based on the historical question data of the target conversation to obtain the task difficulty. Perform performance matching among at least one task processing model according to the task difficulty to obtain the target task processing model, where at least two task processing models have different computing power requirements, and different computing power requirements indicate matching different task difficulties. Then send the current question data to the target task processing model for content generation to obtain answer data, and obtain the feedback result corresponding to the answer data, and obtain the output result for the current question data based on the feedback result and the answer data. The embodiments of this application perform difficulty judgment on the current question data, analyze the current task difficulty in combination with the context relationship of the historical question data in the conversation, so as to identify the real computing requirements of the problem and avoid resource misallocation. In addition, task processing models with different computing power requirements are pre-deployed, and the optimal model is dynamically matched according to the task difficulty evaluated in real time to achieve accurate on-demand invocation and maximize the utilization rate of computing power resources. At the same time, the output result is dynamically adjusted based on the feedback result of the answer data to avoid the problem of low result accuracy caused by misjudgment, significantly reduce the average computing cost on the premise of ensuring the answer quality, and achieve the purpose of improving the resource allocation optimization rate. Description of the Drawings
[0016] Figure 1 is a flowchart of the multi-model collaborative service method provided by the embodiments of this application.
[0017] Figure 2 is a flowchart of analyzing the difficulty of the current question data based on the historical question data of the target conversation to obtain the task difficulty provided by the embodiments of this application.
[0018] Figure 3 is a schematic diagram of the process of forward tracing of the current question data provided by the embodiments of this application.
[0019] Figure 4 is a flowchart of performing performance matching among at least one task processing model according to the task difficulty to obtain the target task processing model provided by the embodiments of this application.
[0020] Figure 5 is a flowchart of selecting a task processing model as the target task processing model from the target model group according to the preset selection strategy provided by the embodiments of this application.
[0021] Figure 6 is a schematic diagram of selecting the target task processing model provided by the embodiments of this application.
[0022] Figure 7It is a flowchart provided by an embodiment of the present application for obtaining an output result for the current question data based on feedback results and answer data.
[0023] Figure 8 It is a flowchart provided by an embodiment of the present application for updating the task difficulty and corresponding answer data until an output result is obtained.
[0024] Figure 9 It is a flowchart provided by an embodiment of the present application for fine-tuning the task difficulty judgment large model.
[0025] Figure 10 It is an example of fine-tuning sample data provided by an embodiment of the present application.
[0026] Figure 11 It is an overall flowchart of the multi-model collaborative service method provided by an embodiment of the present application.
[0027] Figure 12 It is a structural block diagram of a multi-model collaborative service device provided by another embodiment of the present application.
[0028] Figure 13 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners
[0029] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0030] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the flowchart.
[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0032] First, several nouns involved in the present application are analyzed: Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; Artificial intelligence is a branch of computer science. It attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results in terms of theories, methods, technologies, and application systems.
[0033] With the rapid development of large language model technology, there are significant differences in the inference costs of models of different scales. From small models deployed on a single card to extremely large models that require multi-machine cluster support, the resource requirements span a wide range. Taking the series of models released by DeepSeek as an example, its distilled model Qwen1.5B only requires 3GB of video memory to achieve stable deployment, while the R1 model requires up to 1.4TB of video memory resources. In actual application scenarios, users generally tend to choose the model with the strongest computing power to handle all types of tasks. This usage pattern has led to a serious problem of unbalanced resource allocation: high-performance computing resources are over-occupied, while medium and small computing resources are idle.
[0034] To address this problem, in related technologies, intention recognition is mainly achieved by analyzing a large amount of user behavior data, and accordingly, a large model of the corresponding scale is called to execute the corresponding task. However, under this static allocation mechanism, once the model is selected, it cannot be dynamically adjusted, making it difficult to achieve the optimal allocation of resources. After calling the large model, it cannot be dynamically adjusted, resulting in it still being difficult to achieve the optimal allocation of computing resources.
[0035] Based on this, the embodiments of this application provide a multi-model collaborative service method, device, equipment, and storage medium, which determine the difficulty of the current question data, analyze the context relationship of the historical question data in the conversation to parse the current task difficulty, thereby identifying the true computing requirements of the question and avoiding resource misallocation. Additionally, task processing models with different computing power requirements are pre-deployed, and the optimal model is dynamically matched according to the real-time evaluated task difficulty to achieve accurate on-demand invocation, maximizing the utilization rate of computing power resources. At the same time, the output result is dynamically adjusted based on the feedback result of the answer data to avoid the problem of low result accuracy caused by misjudgment. On the premise of ensuring the quality of the answer, the average computing cost is significantly reduced, achieving the goal of improving the resource allocation optimization rate.
[0036] The embodiments of the present application provide a multi-model collaborative service method, apparatus, device, and storage medium, which will be specifically described through the following embodiments. First, the multi-model collaborative service method in the embodiments of the present application will be described.
[0037] The embodiments of the present application can acquire and process relevant data based on artificial intelligence technology. Among them, Artificial Intelligence (AI) is a theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0038] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level technologies and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0039] The multi-model collaborative service method provided by the embodiments of the present application relates to the field of model services. The multi-model collaborative service method provided by the embodiments of the present application can be applied to terminals, can also be applied to the server side, or can be a computer program running on the terminal or the server side. For example, the computer program can be a native program or software module in the operating system; it can be a local (Native) application (Application, APP), that is, a program that needs to be installed in the operating system to run, such as a client that supports multi-model collaborative services, that is, a program that only needs to be downloaded to the browser environment to run; it can also be a small program that can be embedded into any APP. In short, the above computer program can be any form of application program, module, or plug-in. Among them, the terminal communicates with the server through the network. The multi-model collaborative service method can be executed by the terminal or the server, or jointly executed by the terminal and the server.
[0040] In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart watch, etc. In addition, the terminal can also be an intelligent vehicle-mounted device. The intelligent vehicle-mounted device applies the multi-model collaborative service method of this embodiment to provide relevant services and enhance the driving experience. The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms; it can also be a service node in a blockchain system, and the service nodes in this blockchain system form a Peer To Peer (P2P) network. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). The terminal and the server can be connected through communication connection methods such as Bluetooth, Universal Serial Bus (USB), or network, and this embodiment does not limit this here.
[0041] This application can be used in many general or specific computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet-type devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.
[0042] It should be noted that in each specific embodiment of the present application, when it comes to relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of such data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of the present application will be obtained.
[0043] The multi-model collaborative service method in the embodiments of the present application will be described below.
[0044] Figure 1 It is an optional flowchart of the multi-model collaborative service method provided by the embodiments of the present application. Figure 1 The method in [the figure] may include but is not limited to steps 110 to 130. At the same time, it can be understood that the present embodiment does not specifically limit the order of steps 110 to 130 in [the figure], and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added. Figure 1 Steps 110 to 130 in [the figure] are not specifically limited in order, and the order of steps can be adjusted according to actual needs, or some steps can be reduced or added.
[0045] Step 110: Input the current question data in the target dialogue into the task difficulty judgment large model, and perform difficulty analysis on the current question data based on the historical question data of the target dialogue to obtain the task difficulty.
[0046] In one embodiment, when the user uses the large language model service, the specific interaction form is a dialogue-based interaction. The user asks questions in a dialogue manner, and the large language model will generate corresponding answers according to the user's questions. During the entire interaction process, there will be at least one round of question-answer process. In the embodiments of the present application, a user's one dialogue process is called a target dialogue, each round of question is called question data, and all the question data before the current round constitutes the historical question data. At this time, the question data of the current round is called the current question data. It can be understood that the current question data can be the question data of the first round or the intermediate round question data with context. In addition, if the current round is the first round of dialogue, the historical question data is the current question data itself.
[0047] After obtaining the current question data in the target conversation, input the current question data into the task difficulty judgment large model. Analyze the historical question data generated by the target conversation in the task difficulty judgment large model, and perform difficulty analysis on the current question data based on the historical question data to obtain the task difficulty. The historical question data here refers to all the question data before the current round. It can be seen that in this process, the difficulty analysis of the current question data is not limited to itself, but is comprehensively analyzed in the entire target conversation, and the obtained task difficulty is more accurate.
[0048] In one embodiment, refer to Figure 2 , Figure 2 FIG. is a flowchart for performing difficulty analysis on the current question data based on the historical question data of the target conversation in the embodiment of the present application to obtain the task difficulty, which specifically includes the following steps: Step 210: Parse the current question data to obtain a parsing result.
[0049] In one embodiment, the purpose of parsing the current question data is to analyze whether the question data is a continuation of the previous question. Therefore, the parsing here is to parse the continuous action identification words in the current question data. Among them, the continuous action identification words can be words with continuous action meanings such as "continue", "continue from the previous one", "continue to answer", etc. The parsing result of the current question data is obtained through the process of semantic analysis. In addition, all the continuous action identification words can be formed into a special word list for storage.
[0050] For example, the first question of the user in the target conversation is "Write a 5000-word science and technology paper". After getting the answer, the user is not satisfied and the number of words needs to be increased. So, the second question in the second round of questioning is "continue", and its purpose is to let the large language model continue to write. Therefore, when "continue" is used as the current question data, the parsing result contains continuous action identification words.
[0051] Step 220: If the parsing result indicates that the current question data includes at least one continuous action identification word, obtain the previous sentence of the current question data in the historical question data as the current question data until the current question data does not include continuous action identification words, and use the task difficulty judgment large model to perform difficulty analysis on the last current question data to obtain the task difficulty.
[0052] In one embodiment, if the parsing result indicates that the current question data includes at least one consecutive action identifier, it means that the current question data continues the question process of the previous round. At this time, trace back, select the previous sentence of the current question data in the historical question data as the current question data, and continue the parsing process. Judge whether the parsing result still contains consecutive action identifiers. If it contains, continue to trace back until the selected current question data does not include consecutive action identifiers, and perform difficulty parsing on the final current question data to obtain the task difficulty.
[0053] In one embodiment, a multi-round dialogue parser is set to perform the update process of the current question data. Refer to Figure 3 , Figure 3 FIG. is a schematic diagram of the forward traceback process for the current question data provided by the embodiments of the present application. First, for the target dialogue, disassemble it, and obtain a list of question data for each round according to the historical question data. For the current question data, after obtaining the parsing result, perform one-by-one matching in the special word list. After the matching is completed, judge whether it contains at least one consecutive action identifier. If it does not contain consecutive action identifiers, perform difficulty parsing on the current question data to obtain the task difficulty. If it contains consecutive action identifiers, move one forward along the historical question data, obtain the previous question data as the current question data for judgment, until a question data that does not contain consecutive action identifiers is selected as the latest current question data. In this process, judge whether there is a list overflow. If there is no overflow, it means that it is still possible to continue tracing back, otherwise select the question data corresponding to the overflow state.
[0054] After selecting the current question data that does not contain consecutive action identifiers, input it into the task difficulty judgment large model for difficulty parsing. At this time, multiple difficulty levels can be set according to actual needs. According to the input current question data, the task difficulty judgment large model can directly give the task difficulty. For example, the difficulty is divided into 7 levels in total, and the task difficulty for "writing a 5000-word scientific and technological paper" can be set as level 6. If the recognition of consecutive action identifiers is not performed, the task difficulty of "continue" may only be level 1. At this time, directly generating a further answer based on "continue" cannot meet the actual needs of users. At this time, after using the multi-round dialogue parser to perform a rollback, the question data of "continue" is rolled back to "help me write a scientific and technological paper", so as to more accurately judge the question difficulty.
[0055] Step 120: Perform performance matching in at least one task processing model according to the task difficulty to obtain a target task processing model.
[0056] In one embodiment, a computing power network usually includes different computing power chips, such as Ascend 910A, Ascend 910B, A100, etc. These computing power chips have different computing capabilities and can support different large language model services. Embodiments of the present application can deploy large language models with different computing power requirements on different computing power chips as task processing models according to the computing capabilities of the computing power chips. The computing power requirements of at least two task processing models are different, and different computing power requirements indicate matching different task difficulties.
[0057] For example, a task processing model with high performance is deployed on a computing power chip with high computing power to process tasks with high task difficulty, and a task processing model with low performance is deployed on a computing power chip with low computing power to process tasks with low task difficulty. Different performance task processing models can also be deployed on the same computing power chip. Among them, at least one task processing model corresponding to the same task difficulty is regarded as a model group, and there are as many corresponding model groups as there are task difficulties. At this time, the same model group can be located on the same computing power chip or on different computing power chips. The specific deployment of the model group is not limited here.
[0058] For example, suppose there are 20 task processing models for matching different task difficulties and there are 7 kinds of task difficulties in total. At this time, these 20 task processing models are divided into 7 different model groups, and each model group corresponds to one kind of task difficulty.
[0059] In one embodiment, referring to Figure 4 , Figure 4 is a flowchart of performing performance matching in at least one task processing model according to the task difficulty to obtain a target task processing model provided by an embodiment of the present application, which specifically includes the following steps: Step 410: Select a preset difficulty consistent with the task difficulty, and regard the corresponding model group as the target model group.
[0060] In one embodiment, the preset difficulty is all levels of task difficulty, for example, 7 kinds of preset difficulties. The difficulty obtained according to the current question data at this time is called the task difficulty. For example, the task difficulty = level 6. Select the preset difficulty representing level 6, and regard the model group corresponding to this preset difficulty as the target model group.
[0061] Step 420: Select a task processing model from the target model group as the target task processing model according to a preset selection strategy.
[0062] In one embodiment, a model group may include more than one task processing model. Therefore, it is necessary to select from the target model group according to a preset selection strategy. Referring to Figure 5 , Figure 5It is a flowchart for selecting a task processing model as the target task processing model from a target model group according to a preset selection strategy provided by an embodiment of the present application, specifically including the following steps: Step 510: Obtain the model interfaces corresponding to each task processing model in the target model group.
[0063] In one embodiment, each task processing model in the target model group has a corresponding model interface for invocation. Encapsulating each task processing model in the form of a model interface can achieve standardized management and improve the flexibility of invocation and the switching efficiency.
[0064] Step 520: Select a model interface with a random selection strategy or a polling selection strategy, and call the model interface to obtain the corresponding target task processing model.
[0065] In one embodiment, it can be randomly selected or polled. Select one interface from at least one model interface, and call the model interface to access the corresponding task processing model, and use it as the target task processing model.
[0066] In one embodiment, refer to Figure 6 , Figure 6 It is a schematic diagram for selecting a target task processing model provided by an embodiment of the present application. First, each preset difficulty indicates a different difficulty level, each difficulty level corresponds to a model group, and each model group includes at least one task processing model. A task processing model can be selected from the model group as the target task processing model by means of random scheduling or polling scheduling.
[0067] Step 130: Send the current question data to the target task processing model for content generation to obtain answer data, and obtain the feedback result corresponding to the answer data, and obtain the output result for the current question data based on the feedback result and the answer data.
[0068] In one embodiment, after having the target task processing model for processing the current question data, send the current question data to the target task processing model for content generation to obtain answer data. After having the answer data, the generation quality of the answer data can be evaluated by means of user feedback or other artificial intelligence analysis methods, and the evaluation result is used as the feedback result corresponding to the answer data. After having the feedback result, the subsequent fine-tuning process can be carried out.
[0069] In one embodiment, refer to Figure 7 , Figure 7 It is a flowchart for obtaining the output result for the current question data based on the feedback result and the answer data provided by an embodiment of the present application, specifically including the following steps: Step 710: If the feedback result is a positive feedback, use the answer data as the output result.
[0070] In one embodiment, if the feedback result is a positive feedback, that is to say, according to the evaluation result, the matching degree between the answer data and the current question data is high and the generated result is accurate. Therefore, the answer data can be directly used as the final output result.
[0071] Step 720: If the feedback result is a negative feedback, update the task difficulty and the corresponding answer data until an output result is obtained.
[0072] In one embodiment, if the feedback result is a negative feedback, that is to say, the answer data is not the content that the user wants for the current question data. Therefore, the user can give a negative feedback through "regenerate" or other means. If the feedback result is a negative feedback, it indicates that the task difficulty needs to be updated at this time, and new answer data is obtained according to the updated task difficulty until the answer data corresponds to a positive feedback to obtain the final output result.
[0073] In one embodiment, referring to Figure 8 , Figure 8 is a flowchart for updating the task difficulty and the corresponding answer data until an output result is obtained provided by an embodiment of the present application, which specifically includes the following steps: Step 810: Update the task difficulty to the next level of difficulty.
[0074] Step 820: Based on the next level of difficulty, re-determine the target task processing model, update the corresponding answer data and feedback result until the next level of difficulty is the highest difficulty or the feedback result is a positive feedback, and use the final answer data as the output result.
[0075] In one embodiment, if the user is not satisfied with the generated answer data, it can be considered that there is a problem with the task difficulty judgment of the large model for the current question data. Therefore, increase the task difficulty and update the task difficulty to the next level of difficulty. For example, if the original task difficulty is level 6, increase it to level 7. It can be understood that the maximum value of the next level of difficulty is the maximum value of the set preset difficulty.
[0076] Next, based on the new task difficulty, select a new target processing task model to generate content, update the corresponding answer data, and obtain the feedback result for the new answer data. Analyze the feedback result to determine whether to stop increasing the task difficulty. Until the next level of difficulty is the highest difficulty, that is, reaching the difficulty limit, or the feedback result of the obtained answer data is a positive feedback, that is, the answer meets the requirements. After either of these two situations occurs, the final answer data is used as the output result.
[0077] In the embodiments of the present application, the difficulty of the current question data is judged through the above steps, and the context relationship of the historical question data in the conversation is combined to analyze the current task difficulty, so as to identify the true computing requirements of the question and avoid resource misallocation. In addition, task processing models with different computing power requirements are pre-deployed, and the optimal model is dynamically matched according to the task difficulty evaluated in real time, realizing accurate call on demand and maximizing the utilization rate of computing power resources. At the same time, the output result is dynamically adjusted based on the feedback result of the answer data to avoid the problem of low result accuracy caused by misjudgment. On the premise of ensuring the quality of the answer, the average computing cost is significantly reduced, achieving the purpose of improving the resource allocation optimization rate.
[0078] In one embodiment, after obtaining the feedback result corresponding to the answer data, the task difficulty judgment large model can also be fine-tuned accordingly. Refer to Figure 9 , Figure 9 which is a flowchart for fine-tuning the task difficulty judgment large model provided by the embodiments of the present application, specifically including the following steps: Step 910: Obtain the task difficulty, feedback result, and the last current question data as fine-tuning sample data.
[0079] In one embodiment, the task difficulty judgment large model is an instruction fine-tuning model, and its system instruction is expressed as: system_prompts = "You are an expert in judging the difficulty of questions. The difficulty level of the questions is from 1 to 7, and the difficulty levels from low to high are 1, 2, 3, 4, 5, 6, 7 respectively. You only need to directly give the difficulty level of the question.", the input current question data is: prompts = prompts + "What is the difficulty level of this question?", and the output result is the task difficulty: the task difficulty of this question is level 7.
[0080] According to the above format, the fine-tuning sample data is jointly composed of the last current question data, task difficulty, and feedback result, and the specific format is as follows: "Write a 5000-word scientific and technological paper", task difficulty = 6, feedback result = False (negative feedback); "Write a 5000-word scientific and technological paper", task difficulty = 7, feedback result = True (positive feedback).
[0081] Step 920: Use the fine-tuning sample data to perform parameter fine-tuning on the task difficulty judgment large model and update the task difficulty judgment large model.
[0082] In one embodiment, after having the fine-tuning sample data, the parameter fine-tuning of the task difficulty judgment large model can be performed using the fine-tuning sample data, and the task difficulty judgment large model is updated. In one embodiment, refer to Figure 10 , Figure 10This is an example of the fine-tuning sample data provided by the embodiments of the present application. Suppose there are task processing models corresponding to 7 task difficulties respectively. At this time, for the current question data of each task processing model, input it into the task processing model corresponding to the task difficulty for processing to obtain the corresponding answer data. If the feedback result is positive feedback, use the task difficulty as the correct output of the task difficulty judgment large model. If the feedback result is negative feedback, increase the task difficulty one by one, send the current question data into the task processing model corresponding to a higher difficulty level for processing, and re-obtain the feedback result until the task difficulty corresponding to the positive feedback is obtained. Match this task difficulty with the current question data, and accordingly fine-tune the task difficulty judgment large model so that it can correctly output the task difficulty corresponding to the question data and improve the accuracy of the judgment.
[0083] In one embodiment, refer to Figure 11 , Figure 11 This is the overall flowchart of the multi-model collaborative service method provided by the embodiments of the present application.
[0084] First, obtain the target dialogue, and disassemble each question data from the target dialogue as a task. Use the multi-round dialogue parser to determine the actual corresponding question data for the current question data, which can be the current question data itself or a certain historical question data in the target dialogue for more accurate single-task difficulty judgment. Then input the updated current question data into the task difficulty judgment large model for difficulty analysis to obtain the task difficulty. Next, input the task difficulty into the task dispatcher to select the target task processing model, select one of the task processing models corresponding to this task difficulty from multiple task processing models, and use this task processing model to generate the corresponding answer data.
[0085] Then send the answer data to the user feedback module to obtain the feedback result generated by the user for this answer data. If the feedback result is positive feedback, use the answer data as the output result. If the feedback result is negative feedback, update the task difficulty to the next level of difficulty, re-determine the target task processing model based on the next level of difficulty, update the corresponding answer data and feedback result until the next level of difficulty is the highest difficulty or the feedback result is positive feedback, and use the final answer data as the output result. Obtain the output result for the current question data based on the feedback result and the answer data.
[0086] Finally, obtain the task difficulty, the feedback result, and the final current question data as the fine-tuning sample data, perform data processing on it, and accordingly perform instruction fine-tuning on the task difficulty judgment large model so that it can correctly output the task difficulty corresponding to the question data.
[0087] The above process realizes multi-model collaborative services through multi-round dialogue parsers, task difficulty judgment expert big models, task distributors, etc. Users can make corresponding feedback based on the answer data, and then based on the feedback results, the task difficulty judgment expert big model is adaptively updated to make the judgment results more accurate, ensuring the stable and accurate operation of the entire multi-model collaborative system, solving the model service concurrency problem, greatly reducing the concurrency pressure, and making the use of computing power resources more balanced.
[0088] According to Moore's Law of computing chip development: the processing power of a chip can double every 18 months. Users have more powerful high-end computing chips every 18 months, which makes low-end chips idle due to obsolescence. The multi-model collaborative service method provided in the embodiment of the present application can use both low-end and high-end computing chips, maximizing resource utilization while ensuring service quality.
[0089] The technical solution provided by the embodiment of the present application is to input the current question data in the target dialogue into the task difficulty judgment model, and analyze the difficulty of the current question data based on the historical question data of the target dialogue to obtain the task difficulty. According to the task difficulty, performance matching is performed in at least one task processing model to obtain the target task processing model, wherein the computing power requirements of at least two task processing models are different, and different computing power requirements indicate matching different task difficulties. Then the current question data is sent to the target task processing model for content generation to obtain answer data, and the feedback result corresponding to the answer data is obtained, and the output result for the current question data is obtained based on the feedback result and the answer data. The embodiment of the present application judges the difficulty of the current question data, and analyzes the current task difficulty in combination with the contextual relationship of the historical question data in the dialogue, so as to identify the real computing requirements of the problem and avoid resource mismatch. In addition, task processing models with different computing power requirements are pre-deployed, and the optimal model is dynamically matched according to the task difficulty of real-time evaluation, so as to realize on-demand precise calling and maximize the utilization of computing power resources. At the same time, the output result is dynamically adjusted based on the feedback result of the answer data to avoid the problem of low result accuracy caused by misjudgment, and the average computing cost is significantly reduced under the premise of ensuring the quality of the answer, so as to achieve the purpose of improving the resource allocation optimization rate.
[0090] The present application also provides a multi-model collaborative service device, which can implement the multi-model collaborative service method. Figure 12 , the device comprises: Task difficulty acquisition module 1210: used to input the current question data in the target dialogue into the task difficulty judgment model, perform difficulty analysis on the current question data based on the historical question data of the target dialogue, and obtain the task difficulty.
[0091] Task model selection module 1220: It is used to perform performance matching among at least one task processing model according to the task difficulty to obtain a target task processing model. The computing power requirements of at least two task processing models are different, and different computing power requirements indicate matching different task difficulties.
[0092] Feedback update module 1230: It is used to send the current question data to the target task processing model for content generation to obtain answer data, and obtain the feedback result corresponding to the answer data, and obtain the output result for the current question data based on the feedback result and the answer data.
[0093] The specific implementation manner of the multi-model collaborative service device in this embodiment is basically the same as that of the above multi-model collaborative service method, and will not be elaborated here.
[0094] This application embodiment also provides an electronic device, including: At least one memory; At least one processor; At least one program; The program is stored in the memory, and the processor executes the at least one program to implement the multi-model collaborative service method described above in this application. This electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), an in-vehicle computer, etc.
[0095] Please refer to Figure 13 , Figure 13 which illustrates the hardware structure of the electronic device in another embodiment. The electronic device includes: Processor 1301 can be implemented in the form of a general-purpose central processing unit (CPU), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in this application embodiment; Memory 1302 can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. Memory 1302 can store an operating system and other application programs. When implementing the technical solutions provided in this specification embodiment through software or firmware, the relevant program codes are stored in memory 1302, and are called by processor 1301 to execute the multi-model collaborative service method of this application embodiment; An input / output interface 1303 for implementing information input and output; A communication interface 1304 for implementing communication and interaction between this device and other devices, which can achieve communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.); A bus 1305 for transmitting information between various components of the device (such as the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304); Among them, the processor 1301, the memory 1302, the input / output interface 1303, and the communication interface 1304 achieve communication connections with each other inside the device through the bus 1305.
[0096] The embodiment of this application also provides a storage medium. The storage medium is a storage medium that stores a computer program, and when the computer program is executed by a processor, it implements the above multi-model collaborative service method.
[0097] As a non-transitory storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory can include high-speed random access memory, and can also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory optionally includes a memory remotely set relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0098] The multi-model collaborative service method, device, equipment, and storage medium proposed in the embodiments of the present application input the current question data in the target conversation into the task difficulty judgment large model, and analyze the difficulty of the current question data based on the historical question data of the target conversation to obtain the task difficulty. Perform performance matching among at least one task processing model according to the task difficulty to obtain the target task processing model, where the computing power requirements of at least two task processing models are different, and different computing power requirements indicate matching different task difficulties. Then send the current question data to the target task processing model for content generation to obtain answer data, and obtain the feedback result corresponding to the answer data, and obtain the output result for the current question data based on the feedback result and the answer data. The embodiments of the present application perform difficulty judgment on the current question data, analyze the current task difficulty in combination with the context relationship of the historical question data in the conversation, so as to identify the true computing requirements of the problem and avoid resource misallocation. In addition, task processing models with different computing power requirements are pre-deployed, and the optimal model is dynamically matched according to the task difficulty evaluated in real time to achieve accurate on-demand invocation and maximize the utilization rate of computing power resources. At the same time, the output result is dynamically adjusted based on the feedback result of the answer data to avoid the problem of low result accuracy caused by misjudgment. On the premise of ensuring the quality of the answer, the average computing cost is significantly reduced, and the purpose of improving the resource allocation optimization rate is achieved.
[0099] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.
[0100] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine some steps, or different steps.
[0101] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0102] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0103] In the description of this application and the above-mentioned drawings, the terms "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the application described here can be implemented in an order different from those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that comprises a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0104] It should be understood that in this application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the relationship between associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Here, A and B can be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression means any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0105] In several embodiments provided by this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0106] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0107] In addition, in each embodiment of the present application, each functional unit can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.
[0108] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. The foregoing storage medium includes: various media that can store programs, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0109] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of the rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the rights of the embodiments of the present application.
Claims
1. A multi-model collaborative service method, characterized in that, Including: Input the current question data in the target conversation into the task difficulty judgment large model, and perform difficulty analysis on the current question data based on the historical question data of the target conversation to obtain the task difficulty; Perform performance matching among at least one task processing model according to the task difficulty to obtain a target task processing model. The computing power requirements of at least two of the task processing models are different, and different computing power requirements indicate matching different task difficulties; Send the current question data to the target task processing model for content generation to obtain answer data, and obtain the feedback result corresponding to the answer data. Based on the feedback result and the answer data, obtain the output result for the current question data.
2. The multi-model collaborative service method according to claim 1, wherein, The performing difficulty analysis on the current question data based on the historical question data of the target conversation to obtain the task difficulty includes: Parse the current question data to obtain a parsing result; If the parsing result indicates that the current question data includes at least one continuous action identifier, obtain the previous sentence of the current question data in the historical question data as the current question data until the current question data does not include the continuous action identifier, and use the task difficulty judgment large model to perform difficulty analysis on the last current question data to obtain the task difficulty.
3. The multi-model collaborative service method according to claim 1, wherein The at least one task processing model is divided into model groups corresponding to different preset difficulties respectively. Each model group includes at least one of the task processing models. The performing performance matching among at least one task processing model according to the task difficulty to obtain a target task processing model includes: Select the preset difficulty that is consistent with the task difficulty, and use the corresponding model group as the target model group; Select one of the task processing models from the target model group as the target task processing model according to a preset selection strategy.
4. The multi-model collaborative service method according to claim 1, wherein The obtaining the output result for the current question data based on the feedback result and the answer data includes: If the feedback result is a positive feedback, use the answer data as the output result; If the feedback result is a negative feedback, update the task difficulty and the corresponding answer data until the output result is obtained.
5. The multi-model collaborative service method according to claim 4, wherein The updating the task difficulty and the corresponding answer data until the output result is obtained includes: Update the task difficulty to the next level of difficulty; Based on the next level of difficulty, re-determine the target task processing model, update the corresponding answer data and the feedback result until the next level of difficulty is the highest difficulty or the feedback result is the positive feedback, and use the last answer data as the output result.
6. The multi-model collaborative service method according to claim 2, wherein After obtaining the feedback result corresponding to the answer data, the method further includes: Obtain the task difficulty, the feedback result, and the last current question data as fine-tuning sample data; Use the fine-tuning sample data to perform parameter fine-tuning on the task difficulty judgment large model, and update the task difficulty judgment large model.
7. The multi-model collaborative service method according to claim 3, wherein Selecting one of the task processing models from the target model group according to a preset selection strategy as the target task processing model includes: Obtaining the model interface corresponding to each task processing model in the target model group; Selecting the model interface according to a random selection strategy or a polling selection strategy, and calling the model interface to obtain the corresponding target task processing model.
8. A multi-model collaborative service device, characterized in that, Including: Task difficulty acquisition module: configured to input the current question data in the target conversation into a task difficulty judgment large model, perform difficulty analysis on the current question data based on the historical question data of the target conversation, and obtain the task difficulty; Task model selection module: configured to perform performance matching among at least one task processing model according to the task difficulty to obtain a target task processing model, where the computing power requirements of at least two of the task processing models are different, and different computing power requirements indicate matching different task difficulties; Feedback update module: configured to send the current question data to the target task processing model for content generation to obtain answer data, obtain a feedback result corresponding to the answer data, and obtain an output result for the current question data based on the feedback result and the answer data.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the multi-model collaborative service method according to any one of claims 1 to 7 is implemented.
10. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the multi-model collaborative service method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Clinical test questionnaire generation method and device, equipment and storage medium
CN119007900A
Answer matching method and system of artificial intelligence large model
CN119311843A
Digital human question and answer method and device, electronic equipment and storage medium
CN119474320A
Task processing method and task processing system
WO2025007892A1
Cited By
Multi-architecture feedback and collaborative big language model prompt data synthesis method, device and medium
CN121436189A