Model calling method and device, storage medium and program product
Through the single-model and multi-model link design of the model service framework, combined with access conditions, multiple AI models are used to process target tasks in parallel, which solves the hallucination problem of generative models and achieves a balance between accuracy and cost in different scenarios.
Patent Information
- Application Number
- CN202510963756.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-11
- Publication Date
- 2025-10-14
AI Technical Summary
Existing generative models are prone to producing inaccurate, incomplete, or misleading outputs when faced with certain inputs, leading to hallucination problems, affecting user experience and the application of the model in important scenarios.
A model service framework is provided, which provides single-model links and multi-model links, sets the access conditions for multi-model links, uses multiple AI models to process target tasks in parallel, improves the accuracy of the reasoning process, and selects single-model links in scenarios with lower accuracy requirements to balance model call costs and service quality.
It effectively alleviates the model hallucination problem and improves the inference accuracy in application scenarios with high accuracy requirements, while controlling the model call cost to meet the needs of different application scenarios.
Smart Images

Figure CN120780504A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of Internet technology, and in particular to a model calling method, device, storage medium and program product. Background Art
[0002] With the continuous development of artificial intelligence (AI) technology, AI models are not only able to process massive amounts of data, but also have the ability to learn, reason, and make decisions independently. Through continuous learning and practice, they optimize algorithm models and can adapt to complex and changing environments and tasks. Therefore, they are widely used in smart homes, smart transportation, smart customer service and other fields.
[0003] In practical applications, generative models that primarily focus on thinking, reasoning, and decision-making, such as large language models, are quite popular. These models are typically pre-trained using large amounts of unlabeled samples. Pre-training focuses on text coherence, reasoning and decision-making capabilities, and knowledge generalization, but lacks coverage of specific domain knowledge. Due to the model's training mechanism and data characteristics, generative models may experience hallucinations in applications. This means that when presented with certain inputs, the reasoning process produces inaccurate, incomplete, or misleading outputs.
[0004] Model hallucinations can have serious consequences. For example, in customer service scenarios, if a model frequently outputs false information or makes incorrect judgments, it can degrade the user experience, leading to user churn and financial losses. Furthermore, model hallucinations can limit their application in critical scenarios, such as those requiring high accuracy and robustness. Therefore, addressing model hallucinations is urgent. Summary of the Invention
[0005] The embodiments of the present application provide a model calling method, device, storage medium and program product to more effectively alleviate the hallucination problem.
[0006] An embodiment of the present application provides a model calling method, which is applied to a model service framework, wherein the model service framework is located between a target application and an AI model, and can provide a single-model link and a multi-model link for the target application. The method includes: receiving task description information of a target task submitted by a target application, where the target task is generated by the target application based on the user's service request information; if the target task meets the access conditions of the multi-model link, calling multiple AI models on the multi-model link based on the task description information, and processing the target task in parallel to obtain multiple initial task results; performing decision processing on the multiple initial task results according to the task type of the target task to obtain a first target task result; and returning the first target task result to the target application so that the target application can provide the user with a service adapted to the service request information based on the first target task result.
[0007] An embodiment of the present application also provides an electronic device, comprising: a memory and a processor; the memory is used to store computer programs / instructions; the processor is coupled to the memory and is used to execute the computer programs / instructions to implement the steps in the model calling method.
[0008] An embodiment of the present application also provides a computer-readable storage medium storing a computer program / instruction, wherein the computer program / instruction is executed by a processor to implement the steps in the model calling method.
[0009] An embodiment of the present application also provides a computer program product, including: a computer program / instructions, which are executed by a processor to implement the steps in the model calling method.
[0010] An embodiment of the present application also provides a computer program product, including: a computer program / instructions, which are executed by a processor to implement the steps in the model calling method.
[0011] In this embodiment, on the one hand, the model service framework located between the target application and the AI model can provide single-model links and multi-model links for the target application, breaking through the traditional model application mode in which the target application directly calls the model; on the other hand, based on the model service framework, the two model links can coexist and the access conditions of the multi-model link are set. It can not only process the target task in parallel based on multiple AI models on the multi-model link, improve the accuracy of the reasoning process, meet the model requirements in application scenarios with higher accuracy requirements, and more effectively alleviate the hallucination problem, but also meet the model requirements of application scenarios with lower accuracy requirements based on a single model link, thereby balancing the model call cost and model service quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:
[0013] Figure 1 A flowchart of a model calling method provided by an exemplary embodiment of the present application;
[0014] Figure 2 A schematic diagram of the architecture of the model service framework provided by an exemplary embodiment of the present application in an actual application scenario;
[0015] Figure 3 A flowchart of a model calling method in a refund scenario provided by an exemplary embodiment of the present application;
[0016] Figure 4 A schematic diagram of the structure of a model calling device provided by an exemplary embodiment of the present application;
[0017] Figure 5 A schematic structural diagram of an electronic device provided as an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0018] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0019] It should be noted that, in the case of user information involved in the embodiments of the present application, the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in the embodiments of the present application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data need to comply with the relevant laws, regulations and standards of the relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse. In addition, the various models involved in this application (including but not limited to language models or large models) are in compliance with relevant laws and standards.
[0020] In addition, it should be noted that when the embodiments of the present application involve user interaction operations or triggering operations, the user interaction operations or triggering operations involved in the embodiments of the present application include but are not limited to: touch operations, gesture operations, voice operations, head movement operations, eye movement operations and other interactive operations in various ways; among which, touch operations include but are not limited to: click operations, double-click operations, long press operations, sliding operations, pinch operations or mouse hover operations, etc. Sliding operations include but are not limited to: straight sliding, curved sliding, etc.
[0021] Hallucinations are a common problem in current practical applications of models, and these problems often stem from the training mechanisms and data characteristics of these models. First, these models typically learn statistical laws through large-scale, unlabeled samples, lacking mechanisms for verifying authenticity, resulting in the generation of content that is inconsistent with the facts. Second, pre-training objectives focus on text coherence rather than accuracy, and models may arbitrarily fabricate details in pursuit of fluency. Furthermore, models are prone to generating fabricated information when errors accumulate during multi-step reasoning or when they lack comprehensive coverage of specific domain knowledge. In short, when faced with certain inputs, the model's reasoning process often produces inaccurate, incomplete, or misleading outputs.
[0022] In order to solve the problem of model hallucination, in an embodiment of the present application, a new model service framework is provided. The model service framework provides model services for the target application as a service tool for the target application, breaking through the model application mode in which the traditional target application directly calls the model. The target application refers to an application that relies on an AI model, which can also be referred to as an AI application. In the embodiment of the present application, the focus is on applications that rely on generative models. Generative models refer to a type of model in the field of machine learning and artificial intelligence that can learn the intrinsic distribution of data and generate new samples, including but not limited to: large language models, image generation models, text-based graph models, etc. Among them, the model service framework can provide single-model links and multi-model links for the target application. A single AI model can be deployed on the single-model link, and multiple AI models can be deployed on the multi-model link.
[0023] Furthermore, the AI model in the embodiments of the present application refers to a model that provides AI capabilities for the target application. It is formed based on data and algorithm training, and can simulate human intelligent behaviors (such as perception, reasoning, decision-making, etc.). It can serve as the core technical component of AI applications, and outputs specific results by processing input data, providing intelligent capability support for various AI applications.
[0024] Of course, the embodiments of the present application do not limit the specific implementation methods of the target application and the AI model. The target application may also refer to an application on any model such as a discriminant model, a causal model or a reinforcement learning model, and may be any type of application such as an intelligent recommendation application, an autonomous driving application, a medical diagnosis application, an online shopping application or an instant messaging application; the AI model may be any model such as a generative model, a discriminant model, a causal model or a reinforcement learning model. Among them, the discriminant model is a model that directly learns the mapping relationship between sample features and target labels, the causal model is a mathematical framework for describing the causal relationship between variables, and the reinforcement learning model is a model that learns the optimal behavior strategy through the interaction between the agent and the environment.
[0025] The embodiment of the present application further provides a model calling method based on the model service framework. Based on the model service framework, the two model links coexist and the access conditions of the multi-model link are set. It can not only process the target task in parallel based on multiple AI models on the multi-model link, improve the accuracy of the reasoning process, meet the model requirements in application scenarios with higher accuracy requirements, and more effectively alleviate the hallucination problem, but also meet the model requirements of application scenarios with lower accuracy requirements based on a single model link, thereby balancing the model calling cost and model service quality.
[0026] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0027] Figure 1 The flow chart of the model calling method provided by the exemplary embodiment of the present application is as follows. Figure 1 As shown, the method may include the following steps:
[0028] Step 11: Receive task description information of the target task submitted by the target application. The target task is generated by the target application according to the user's service request information.
[0029] Step 12: If the target task meets the access conditions of the multi-model link, call multiple AI models on the multi-model link based on the task description information, and process the target task in parallel to obtain multiple initial task results.
[0030] Step 13: Perform decision processing on the multiple initial task results according to the task type of the target task to obtain a first target task result.
[0031] Step 14: Return the first target task result to the target application, so that the target application can provide the user with a service adapted to the service request information according to the first target task result.
[0032] In this embodiment, users can upload service request information on the target application. This service request information reflects the user's service needs and can be information uploaded by the user on a platform such as an AI application to describe the service they require, such as service content and service requirements. Based on the user's service request information, the target application can generate a target task and task description information for the target task. The target task can be any type of task, depending on the service request information. The task description information can be used to describe the details of the task. For example, if the user's service request information indicates that the user needs a refund for a product, the target task generated based on the service request information can be a refund task to refund the user.
[0033] In this embodiment, considering that some application scenarios have higher requirements for the accuracy of inference results, while some application scenarios have relatively low requirements for the accuracy of inference results, if only multi-model links are used to process the target tasks, although the model service quality / task processing quality can be improved, the model call cost will be significantly increased; based on the above considerations, in this embodiment, the access conditions of multi-model links are set to reasonably choose from single-model links and multi-model links to balance the model call cost and model service quality. This will be explained in detail below.
[0034] After receiving the task description information of the target task, it can be determined whether the target task meets the access conditions of the multi-model link. The embodiment of the present application does not limit the specific implementation method of the judgment process. In some exemplary embodiments, it can be judged from multiple access dimensions whether the target task meets the access conditions of the multi-model link, such as the scenario dimension, the user dimension or the task dimension. The access conditions can be configured using the default configuration rules. The default configuration rules have preset access conditions corresponding to target tasks of different task types. The developer can also customize the configuration according to the target task. For example, if the target task is a refund task, the access condition may be that the refund amount is greater than 1,000. For example, if the target task is a product recommendation task, the access condition may be that the number of recommended products is greater than 10 or the product type of the recommended product is a specific type, and so on.
[0035] If the target task meets the access conditions of the multi-model link, multiple AI models on the multi-model link can be called based on the task description information to process the target task in parallel to obtain multiple initial task results. Among them, one AI model can usually output one initial task result. In this case, the number of AI models on the multi-model link is the same as the number of initial task results. In the event that some AI models fail to call, respond or time out during task processing, the number of AI models on the multi-model link can be greater than the number of initial task results.
[0036] Afterwards, a decision process can be performed on the multiple initial task results based on the task type of the target task to obtain a first target task result. Different target task types require different decision-making methods for the multiple initial task results. For example, when the target task is a classification task, a result voting method can be used. For example, if some AI models determine that a refund of 1,500 yuan is allowed to be made to the user, while other AI models determine that a refund of 1,500 yuan is not allowed to be made to the user, a result voting method can be used to determine whether to refund the user 1,500 yuan.
[0037] When the target task is a regression task, a result integration approach can be used. This approach allows multiple AI models on a multi-model chain to process the target task in parallel, improving the accuracy of the inference process, meeting the model requirements of application scenarios with high accuracy requirements, and effectively alleviating the problem of hallucinations.
[0038] After obtaining the first target task result, it can be returned to the target application. The target application can then provide the user with a service adapted to the service request information based on the first target task result. For example, if the target task is a refund task to refund the user, the first target task result obtained through parallel processing and decision-making based on the target task may include result information related to the refund task, such as the refund time, amount, and policy. The target application can then provide the user with a refund service adapted to the service request information based on the first target task result.
[0039] In this embodiment, on the one hand, the model service framework located between the target application and the AI model can provide single-model links and multi-model links for the target application, breaking through the traditional model application mode in which the target application directly calls the model; on the other hand, the two model links coexist, and the access conditions of the multi-model link are set, which can not only process the target task in parallel based on multiple AI models on the multi-model link, improve the accuracy of the reasoning process, meet the model requirements in application scenarios with higher accuracy requirements, and more effectively alleviate the hallucination problem, but also meet the model requirements of application scenarios with lower accuracy requirements based on a single model link, thereby balancing the model call cost and model service quality.
[0040] In some optional embodiments, the access conditions of the multi-model link in the aforementioned embodiments may include at least one of the following: access conditions of the scenario dimension, access conditions of the user dimension, and access conditions of the task dimension. When judging whether the target task meets the access conditions of the multi-model link, at least one of the application scenario, user attributes, and task attributes corresponding to the target task may be obtained as access judgment information. User attributes may include at least one of user identification, user level, and user category; task attributes may include at least one of task type, task parameters, and task cost. If at least one piece of information meets the access conditions of the corresponding dimension, it is determined that the target task meets the access conditions of the multi-model link.
[0041] For example, assuming that the admission judgment information is: the user level targeted by the target task is level 4, and the admission condition of the user dimension is: the user level targeted by the target task is greater than level 3, then the target task meets the admission conditions of the multi-model link.
[0042] Assume that the admission judgment information is: the application scenario of the target task is the navigation scenario, and the admission condition of the scenario dimension is: the target task is applicable to the online shopping scenario, then the target task does not meet the admission conditions of the multi-model link.
[0043] Assume that the admission judgment information is: the task cost of the target task is to consume two computing power units, and the admission condition of the scenario dimension is: the computing power consumed by the target task does not exceed three computing power units, then the target task meets the admission conditions of the multi-model link.
[0044] In this way, it is possible to more accurately judge whether the target task meets the access conditions of the multi-model link from multiple access dimensions; no matter from which access dimension the target task is judged whether it meets the access conditions of the multi-model link, as long as the target task has high requirements for the accuracy of the model, it can meet the access conditions of the multi-model link and solve the problem of model hallucination.
[0045] Following the above embodiment, the embodiment of the present application does not limit the specific method for obtaining the access judgment information. For example, the access judgment information can be obtained in the following ways:
[0046] Method 1: Based on the access conditions of the multi-model link, determine to request at least one of the application scenario, user attributes, and task attributes corresponding to the target task from the target application as access judgment information. Send an information acquisition request to the target application and receive at least one piece of information returned by the target application based on the information acquisition request.
[0047] Method 2: Receive at least one of the following information: application scenario, user attributes, and task attributes corresponding to the target task, which is delivered by the target application along with the target task.
[0048] In this way, based on method 1, the model service framework can determine what information needs to be obtained as admission judgment information according to the admission conditions; based on method 2, the target application can determine what information needs to be obtained as admission judgment information.
[0049] After judging the access conditions of the target task, if the target task does not meet the access conditions of the multi-model link, a single AI model on the single-model link can be called to process the target task based on the task description information to obtain a second target task result. Afterwards, the second target task result can be returned to the AI model so that the target application can provide the user with a service adapted to the service request information based on the second target task result. For example, when the target task is a refund task and the access condition is that the refund amount is greater than 1,000, the target task is a refund task to refund 950 yuan to the user. In this case, the target task does not meet the access conditions, and a single AI model on the single-model link can be called to process the target task based on the task description information to obtain the second target task result.
[0050] In this way, based on the model service framework, the two model links coexist and the access conditions of the multi-model link are set. It can not only process the target tasks in parallel based on multiple AI models on the multi-model link, improve the accuracy of the reasoning process, meet the model requirements in application scenarios with higher accuracy requirements, and more effectively alleviate the hallucination problem, but also meet the model requirements of application scenarios with lower accuracy requirements based on a single model link, thereby balancing the model call cost and model service quality.
[0051] In some optional embodiments, in order to improve the scalability and flexibility of the aforementioned embodiments, the embodiments of the present application provide an optional architecture of a model service framework, which may include an access layer and a model capability layer. The access layer may provide a configuration interface for developers to configure information, and the model capability layer may be used to process target tasks in parallel. The following is an exemplary description of an optional architecture of the model service framework. Of course, this does not limit the specific architecture of the model service framework. The model service framework may also include any layers such as a data management layer and a task result evaluation layer according to actual design requirements.
[0052] In this embodiment, in response to configuration operations on the configuration interface, the service identifier corresponding to the target application, the access conditions for the multi-model link, and the mapping relationship between the task type and the decision method can be configured. The service identifier is bound to the access conditions and the mapping relationship. The configuration interface can be implemented as a configuration interface or in any other form, such as a configuration button.
[0053] When the configuration interface is implemented as a configuration interface, the configuration interface can be displayed to developers. The configuration interface may include service identification configuration items, access condition configuration items and task type configuration items. Developers can configure the above configuration items. Correspondingly, in response to the configuration operation of the service identification configuration item, a service identification list can be displayed. In response to the selection operation of the service identification list, the selected service identification can be used as the service identification corresponding to the target application.
[0054] Based on the target application's corresponding service identifier, the access condition list corresponding to the access condition configuration item and the task type list corresponding to the task type configuration item can be determined. Different service identifiers correspond to different access condition lists and task type lists. In other words, the service identifier can further determine the range of available access conditions and task types, allowing developers to configure access conditions and task types within a reasonable range.
[0055] Developers can configure the access condition configuration items. Correspondingly, they can respond to the configuration operation of the access condition configuration items and display the access condition list for developers to select access conditions from the access condition list; in response to the selection operation of the access condition list, the selected access condition will be used as the access condition of the multi-model link.
[0056] Developers can configure task type configuration items. In response to these configuration operations, a list of task types is displayed. In response to a developer's selection of a task type from the list, a mapping between the task type and the decision-making method is generated based on the selected task type. Different task types correspond to different decision-making methods.
[0057] In this way, the access layer based on the model service framework allows developers to customize the service identification corresponding to the target application, the access conditions of the multi-model link, and the mapping relationship between the task type and the decision-making method, thereby improving the scalability and flexibility of the aforementioned implementation schemes.
[0058] Optionally, after the developer configures the service identifier of the target application through the configuration interface, he or she can respond to the configuration operation on the configuration interface and configure multiple AI models on the multi-model link based on the service identifier of the target application.
[0059] Among them, multiple AI models are at least part of the AI model set accessed by the model service framework, and the multiple AI models on the multi-model link corresponding to different service identifiers are different. For example, the AI model set accessed by the model service framework may include 5 AI models, namely AI model M1, AI model M2, AI model M3, AI model M4 and AI model M5, and the number of multiple AI models may be 1, 2, 3 or 4, or 5; the multiple AI models corresponding to the service identifier H1 are: AI model M2 and AI model M3, the multiple AI models corresponding to the service identifier H2 are: AI model M2, AI model M3 and AI model M4, the multiple AI models corresponding to the service identifier H3 are: AI model M1 and AI model M4, and the multiple AI models corresponding to the service identifier H4 are: AI model M1, AI model M2, AI model M3, AI model M4 and AI model M5.
[0060] In this way, the access layer based on the model service framework allows developers to flexibly configure multiple AI models on the multi-model link, further improving the scalability and flexibility of the aforementioned implementation schemes.
[0061] In some optional embodiments, the model capability layer in the model service framework can provide single-model links, multi-model links, and thread pools. Based on this, step 12 of the aforementioned embodiment, "based on the task description information, calling multiple AI models on the multi-model link to process the target task in parallel to obtain multiple initial task results," can be implemented based on the following steps 121-122:
[0062] Step 121: Select multiple threads from the thread pool and establish a correspondence between the multiple threads and the multiple AI models on the multi-model link. This application does not limit the thread selection method. Optionally, the mapping relationship between the threads to be allocated and the models can be fixed according to the static allocation rules (i.e., pre-binding rules), and each thread can specifically handle the corresponding specific AI model; in addition, according to the load balancing scheduling rules, the threads can be dynamically allocated according to the current load of the AI model to avoid overloading of certain AI models. In this way, thread selection can be performed more flexibly and accurately.
[0063] Step 122: Send the task description information to multiple threads and start multiple threads. Specifically, multiple threads can be started in multiple independent operating environments, and the multiple independent operating environments are isolated from each other to achieve service isolation for multiple AI models. The embodiment of the present application does not limit the specific implementation method of the independent operating environment. It can be a server, that is, separated by physical hardware, each server has an independent processor, memory, storage and network interface, and runs an independent operating system; it can also be a virtual machine (VM), that is, multiple virtual operating system instances are created on a physical server through a virtual machine monitor, and each virtual machine has independent virtual hardware; it can also be a container, that is, process-level isolation is achieved based on the namespace and resource restrictions of the operating system kernel, and multiple containers share the host kernel, but have independent file systems, process spaces and resource quotas. In this way, on the one hand, based on independent operating environments such as servers / virtual machines / containers, physical / logical isolation of different AI model services is achieved, avoiding the overall system crash due to single model anomalies (such as memory leaks, high loads). On the one hand, each independent operating environment can independently configure resource quotas such as processors or memory to avoid performance jitter caused by multiple AI models competing for resources.
[0064] After the thread is started, the task description information can be converted into model prompt words according to the prompt word structure of the corresponding AI model. The task description information can be structurally converted into model prompt words with a prompt word structure, or the semantic features of the task description information can be extracted and the model prompt words with a prompt word structure can be generated based on the semantic features. This is not limited in this embodiment. Afterwards, the corresponding AI model can be called to perform task processing based on the model prompt words to obtain the initial task result. The model prompt words can be provided with identification information of the corresponding AI model, so that the corresponding AI model can be called to perform task processing based on the identification information.
[0065] In this way, multiple AI models on the multi-model link can be called based on the task description information, and the target tasks can be processed in parallel more efficiently, and the accuracy of the multiple initial task results obtained is higher.
[0066] In some optional embodiments, after the target task is processed in parallel to obtain multiple initial task results, the multiple initial task results can be decision-processed based on the task type of the target task based on the following implementation method to obtain a first target task result.
[0067] Implementation method 1: If the target task is a classification task, multiple initial task results are voted on, and the initial task result pointed to by the voting result is used as the first target task result. This embodiment does not limit the number of classification results for the classification task, and can be 2, 3, or 4, or any integer greater than 2, depending on the application scenario.
[0068] The following details the voting logic for classification tasks with two classification results (i.e., binary classification tasks). The voting logic for classification tasks with three or four classification results (i.e., three- and four-class classification tasks) is similar to that for binary classification tasks and will not be further detailed.
[0069] For the "two-classification" classification task, multiple initial task results may include two classification results. In this case, the multiple initial task results can be voted to obtain the voting results; if the voting results indicate that the number of votes for the two classification results is different, the classification result with more votes in the voting results will be used as the first target task result. For example, assuming that the target task is to classify pictures, 5 AI models can obtain two results after processing the target task respectively: cats and dogs, where cats are voted 3 times and dogs are voted 2 times. Since cats have more votes, the result of the first target task is: the picture is a picture of a cat. For another example, 5 AI models are used to process the target task (emotional classification of a text) in parallel; wherein, calling 5 AI models to process the target task can obtain two classification results, the first classification result is that the text emotion is "positive", and the second classification result is that the text emotion is "negative". Among them, the number of votes for the first classification result is 2, and the number of votes for the second classification result is 1. At this time, the first classification result has more votes, and the first classification result with more votes in the voting results can be directly used as the first target task result.
[0070] If the voting result indicates that the two classification results have the same number of votes, and the number of the multiple initial task results is less than the number of the multiple AI models, then one of the AI models that did not output the initial task result is selected as the retry model. The fact that the voting result indicates that the two classification results have the same number of votes indicates that the number of initial task results is an even number. The embodiments of the present application do not limit the specific method for selecting the retry model. In one exemplary embodiment, the AI model with the highest performance or largest scale can be selected as the retry model based on the performance or scale information of the AI model. A retry model can also be randomly selected from the AI models, or the AI model with the highest priority can be selected as the retry model. Based on the task description information, the retry model is invoked to process the target task and obtain a retry task result, which is any classification result. The retry task result and the multiple initial task results are re-voted to obtain a first re-voting result. The classification result with the largest number of votes in the first re-voting result is used as the first target task result. In this way, based on the retry model selection and the subsequent re-voting process, the tie situation described above, where "the number of initial task results is an even number," can be more effectively resolved.
[0071] For example, using 8 AI models to process the target task (emotion classification of a text) in parallel, two classification results can be obtained. The first classification result is that the text sentiment is "positive" with 3 votes, and the second classification result is that the text sentiment is "negative" with 3 votes. The total number of votes for the two classification results, 6 (i.e., the number of multiple initial task results), is less than the number of multiple AI models, 8, indicating that there are two AI models that have never output an initial task result. Therefore, one of the two AI models that have never output an initial task result can be selected as a retry model. Based on the task description information of the target task, the retry model is called for processing, and the retry task result is "positive". The retry result of this task is re-voted with the 6 initial task results, and the first re-voting result is: the first classification result "positive" has 3 votes, and the second classification result "negative" has 3 votes. Therefore, the first classification result "positive" can be finally used as the first target task result.
[0072] The above introduces the case where the retry model is called to successfully obtain the retry task result. Considering the actual application scenario, the retry model may not be able to obtain the retry task result even if it is retried due to more serious failure problems; therefore, the embodiment of the present application can also discard the initial task results output by some AI models if the retry task result is not obtained.
[0073] Optionally, the initial task results output by some AI models may be randomly selected from the initial task results output by multiple AI models and discarded. This embodiment does not impose any restrictions on this.
[0074] Optionally, based on the attribute information of multiple AI models that output initial task results, the initial task results output by AI models whose attribute information meets set conditions may be discarded; the attribute information of the AI model includes at least one of parameter scale, service performance, model type, and priority. For example, the AI model with the lowest priority, the smallest parameter scale, or the worst service performance may be selected from multiple AI models and its output initial task results discarded.
[0075] The above describes in detail the implementation method 1 in which the target task is a classification task. In this way, when the target task is a classification task, multiple initial task results can be voted on more accurately to obtain more accurate task results.
[0076] Implementation method 2: If the target task is a regression task, multiple initial task results are integrated to obtain a second target task result.
[0077] Among them, if multiple initial task results are numerical results, numerical calculation is performed on the numerical values of the multiple initial task results, and the result of the numerical calculation is used as the second target task result. Among them, the numerical calculation can be a direct addition method, a weighted summation method, a division or multiplication method, etc. This embodiment does not limit the specific numerical calculation method, and can be flexibly set according to the application scenario. For example, 3 AI models predict the size of an object in a picture and obtain 3 initial task results. The 3 initial task results are numerical results and are specific centimeters, namely 12.5cm, 13.1cm, and 12.8cm. These values are calculated, such as taking the average value: (12.5+13.1+12.8)÷3=12.8cm, and finally 12.8cm is used as the second target task result.
[0078] If multiple initial task results are text-based results, semantic understanding is performed on the texts of the multiple initial task results to obtain multiple semantic information; the logical relationship between the multiple semantic information is analyzed, and the texts of the multiple initial task results are integrated and optimized based on the logical relationship to obtain the second target task result. The integration and optimization process may include operations such as content merging, redundant information elimination, and structural optimization.
[0079] Among them, content merging refers to: reorganizing the texts of multiple initial task results according to the logical relationship between multiple semantic information, and adding conjunctions between the texts of multiple initial task results to make the expression effect more coherent. Redundant information elimination refers to unifying specific words (such as terms or other preset words) that appear in the texts of multiple initial task results. For example, if the text of the first initial task result contains the term "unit test" and the text of the second initial task result contains the term "automated test", after eliminating redundant information of the two initial task results, "unit test" and "automated test" can be replaced by "test strategy". Structural optimization refers to optimizing the structure between the texts of multiple initial task results according to the logical relationship between multiple semantic information. In the above manner, the texts of multiple initial task results can be integrated and optimized more accurately to obtain a more accurate second target task result.
[0080] Based on the above implementations 1 and 2, when the target task is a regression task, multiple initial task results can be integrated and processed more accurately to obtain more accurate task results.
[0081] The following will further explain the model calling method provided by the embodiment of this application in combination with actual scenarios. In actual scenarios, the model service framework can be Figure 3The architecture shown includes: an application layer, an access layer, and a model capability layer. Of course, this is only an optional example and does not limit the specific framework of the model service framework of this solution.
[0082] The application layer is used to access the target application or its functions. The target application can receive service request information uploaded by the user, generate a target task based on the user's service request information, and submit the task description information of the target task to the application layer. The target application may have a judgment function, a quality inspection function, or a complaint function, etc., which is not limited in this embodiment. The number of target applications can be one or more, which is not limited.
[0083] Among them, the access layer can be oriented towards business personnel or program developers, allowing them to configure information such as service identification, access conditions, and decision conditions. The access layer may include a service configuration module, a link access module, and a service monitoring module. The service configuration module is used to configure the service identification corresponding to the target application and the mapping relationship between the task type and the decision method; the link access module is used to configure the access conditions of the multi-model link; the service monitoring module is used to monitor the execution of the target task, the user-oriented service status of the target application, or the operation of the AI model.
[0084] The model capability layer provides two model links: a single-model link and a multi-model link, as well as the scheduling, execution, and decision-making capabilities of the models on the links. The model capability layer includes a model access layer, an intelligent execution layer, and a decision-making processing layer.
[0085] The model access layer can be used to call multiple AI models. The model configuration module within the model access layer configures the key, version, and model type of the connected AI model, such as whether it is an inference model or a multimodal model. The model call module within the model access layer can call the corresponding AI model based on the above information configured by the model configuration module. The priority definition module in the model access layer can eliminate or select AI models in the scenario of equal votes.
[0086] The intelligent execution layer may include: a thread pool task creation module, used to allocate a corresponding thread to each AI model for the AI model to execute the target task; an exception handling module, used to retry one or more times when the AI model call fails; a timeout control module, used to selectively discard the task execution results output by the AI model whose response time or processing time exceeds the preset time threshold, so as to avoid waiting for a long time for the task execution result of a certain AI model.
[0087] The decision processing layer may include: a rule definition configuration module, a rule execution module and a result summary module. Among them, the rule definition configuration module can be used for developers to configure the decision processing method for target tasks of different task types. For example, the target task is a classification task or a regression task; the classification task can be processed by voting; the regression task can be processed by integration. The rule execution module can be used to call multiple AI models on the multi-model link to process the target task in parallel to obtain multiple initial task results. The result summary module can be used to perform decision processing on multiple initial task results to obtain the first target task result.
[0088] It should be noted that the AI models mentioned in the above embodiments of this application can be generative AI models. Further, optionally, the generative AI model can be a large AI model, such as a large language model. A large AI model refers to a model whose model parameters meet the set parameter quantity requirements. The parameter quantity requirements are not limited here and may vary in different scenarios or fields.
[0089] The detailed implementation and beneficial effects of each step in the method of this embodiment have been described in detail in the aforementioned embodiments and will not be elaborated here.
[0090] In this way, based on the model service framework and the access conditions of the set multi-model link, it is possible to process the target task in parallel based on multiple AI models on the multi-model link, improve the accuracy of the reasoning process, meet the model requirements in application scenarios with higher accuracy requirements, and more effectively alleviate the hallucination problem. At the same time, it is possible to meet the model requirements of application scenarios with lower accuracy requirements based on a single model link, thereby balancing the model call cost and model service quality.
[0091] The following will take the refund application scenario as an example to illustrate the model calling method. Figure 3As shown, the task description information of the refund task submitted by the target application can be received, and it can be determined whether the refund task meets the access conditions (for example, the refund amount is greater than 1000). If the refund task does not meet the access conditions, a single AI model on the single model link can be called to process the refund task. If the refund task meets the access conditions, corresponding threads can be allocated to multiple AI models, and multiple AI models on the multi-model link (AI model M1, AI model M2 and AI model M3) can be called to process the refund task in parallel to obtain multiple initial task results. Among them, the initial task result output by AI model M1 is {Whether to refund: yes; reason: XXX}, the initial task result output by AI model M2 is {Whether to refund: yes; reason: XXX}, and the initial task result output by AI model M3 is {Whether to refund: no; reason: XXX}. Afterwards, the multiple initial task results can be decision-processed to obtain the first target task result and returned to the target application.
[0092] The detailed implementation and beneficial effects of each step in the method of this embodiment have been described in detail in the aforementioned embodiments and will not be elaborated here.
[0093] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 11 to 14 can be device A; for another example, the execution entity of steps 11 and 12 can be device A, and the execution entity of steps 13 and 14 can be device B; and so on.
[0094] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The sequence numbers of the operations, such as 11, 12, etc., are only used to distinguish between different operations, and the sequence numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.
[0095] Figure 4 This is a structural diagram of a model calling device provided by another exemplary embodiment of the present application. Figure 4As shown, the device includes: a receiving module 401, used to receive task description information of a target task submitted by a target application, where the target task is generated by the target application according to the service request information of the user; a parallel processing module 402, used to: if the target task meets the access conditions of the multi-model link, call multiple AI models on the multi-model link based on the task description information, and perform parallel processing on the target task to obtain multiple initial task results; a decision processing module 403, used to: perform decision processing on the multiple initial task results according to the task type of the target task to obtain a first target task result; a result return module 404, used to return the first target task result to the target application, so that the target application can provide the user with a service adapted to the service request information according to the first target task result.
[0096] Optionally, the model calling device further includes: a single model link calling module 405, the single model link calling module 405 being configured to: if the target task does not meet the access conditions of the multi-model link, call a single AI model on the single model link based on the task description information to process the target task to obtain a second target task result;
[0097] The second target task result is returned to the AI model so that the target application can provide the user with a service adapted to the service request information according to the second target task result.
[0098] Optionally, the access conditions of the multi-model link include at least one of the following: access conditions of the scene dimension, access conditions of the user dimension and access conditions of the task dimension; the model calling device also includes: a condition judgment module 405, the condition judgment module 405 is used to obtain at least one information of the application scenario, user attributes and task attributes corresponding to the target task as access judgment information; if the at least one information meets the access conditions on the corresponding dimension, it is determined that the target task meets the access conditions of the multi-model link.
[0099] Optionally, when obtaining at least one of the application scenarios, user attributes and task attributes corresponding to the target task, the conditional judgment module 405 is specifically used to: determine, based on the access conditions of the multi-model link, to request at least one of the application scenarios, user attributes and task attributes corresponding to the target task from the target application as access judgment information; send an information acquisition request to the target application, and receive the at least one of the information returned by the target application according to the information acquisition request; or, receive at least one of the application scenarios, user attributes and task attributes corresponding to the target task issued by the target application together with the target task.
[0100] Optionally, the model calling device also includes: a configuration module 406, the model service framework includes an access layer, the access layer provides a configuration interface to the outside, and the configuration module 406 is used to: respond to the configuration operation on the configuration interface, configure the service identifier corresponding to the target application, the access conditions of the multi-model link, and the mapping relationship between the task type and the decision-making method; wherein, the service identifier is bound to the access conditions and the mapping relationship.
[0101] Optionally, the configuration interface is implemented as a configuration interface. Then, when the configuration module 406 configures the service identifier corresponding to the target application, the access conditions of the multi-model link, and the mapping relationship between the task type and the decision-making method in response to the configuration operation on the configuration interface, it is specifically used to: display the configuration interface, which includes a service identifier configuration item, an access condition configuration item, and a task type configuration item; in response to the configuration operation of the service identifier configuration item, display a service identifier list, and in response to the selection operation of the service identifier list, use the selected service identifier as the service identifier corresponding to the target application; based on the service identifier corresponding to the target application, determine the The access condition list corresponding to the access condition configuration item and the task type list corresponding to the task type configuration item; the access condition list and task type list corresponding to different service identifiers are different; in response to the configuration operation of the access condition configuration item, the access condition list is displayed, and in response to the selection operation of the access condition list, the selected access condition is used as the access condition of the multi-model link; in response to the configuration operation of the task type configuration item, the task type list is displayed, and in response to the selection operation of the task type list, a mapping relationship between the task type and the decision-making method is generated according to the selected task type; wherein different task types correspond to different decision-making methods.
[0102] Optionally, the configuration module 406 is also used to: respond to the configuration operation on the configuration interface, configure multiple AI models on the multi-model link based on the service identifier of the target application, and the multiple AI models are at least part of the AI model set accessed by the model service framework; wherein, the multiple AI models on the multi-model link corresponding to different service identifiers are different.
[0103] Optionally, the model service framework includes a model capability layer, which provides a single-model link, a multi-model link and a thread pool; the parallel processing module 402 calls multiple AI models on the multi-model link based on the task description information, and performs parallel processing on the target task to obtain multiple initial task results. It is specifically used to: select multiple threads from the thread pool, establish a correspondence between the multiple threads and the multiple AI models on the multi-model link; send the task description information to the multiple threads, and start the multiple threads; wherein, after the thread is started, the task description information is generated into a model prompt word according to the prompt word structure of the corresponding AI model, and the corresponding AI model is called according to the model prompt word to perform task processing to obtain the initial task result.
[0104] Optionally, when the parallel processing module 402 starts the multiple threads, it is specifically used to: start the multiple threads in multiple independent operating environments, and the multiple independent operating environments are isolated from each other to achieve service isolation of the multiple AI models, and the independent operating environments are servers, virtual machines or containers.
[0105] Optionally, the decision processing module 403 performs decision processing on the multiple initial task results according to the task type of the target task to obtain a first target task result, specifically for: if the target task is a classification task, then voting on the multiple initial task results, and taking the initial task result pointed to by the voting result as the first target task result; if the target task is a regression task, then integrating the multiple initial task results to obtain the second target task result.
[0106] Optionally, the multiple initial task results include two classification results. When the decision processing module 403 votes on the multiple initial task results and takes the initial task result pointed to by the voting result as the first target task result, it is specifically used to: vote on the multiple initial task results to obtain a voting result; if the voting result indicates that the votes for the two classification results are different, take the classification result with more votes in the voting result as the first target task result; if the voting result indicates that the votes for the two classification results are the same, and the number of the multiple initial task results is less than the number of the multiple AI models, then select one from the AI models that have not output the initial task result as a retry model; call the retry model based on the task description information to process the target task, and obtain a retry task result, which is any classification result; re-vote the retry task result and the multiple initial task results to obtain a first re-voting result; and take the classification result with more votes in the first re-voting result as the first target task result.
[0107] Optionally, the decision processing module 403 is also used to: if no retry task result is obtained, then based on the attribute information of multiple AI models that output the initial task result, discard the initial task result output by the AI model whose attribute information meets the set conditions, and re-vote the remaining initial task results to obtain a second re-voting result; use the classification result with more votes in the second re-voting result as the first target task result; the attribute information of the AI model includes at least one of parameter scale, service performance, model type and priority.
[0108] Optionally, when the decision processing module 403 integrates the multiple initial task results to obtain the second target task result, it is specifically used to: if the multiple initial task results are numerical results, perform numerical calculation on the numerical values of the multiple initial task results, and use the results of the numerical calculation as the second target task result; if the multiple initial task results are text results, perform semantic understanding on the texts of the multiple initial task results to obtain multiple semantic information; analyze the logical relationship between the multiple semantic information, and integrate and optimize the texts of the multiple initial task results according to the logical relationship to obtain the second target task result.
[0109] The detailed implementation methods and beneficial effects of each step in the method of this embodiment have been described in detail in the aforementioned embodiments and will not be elaborated on here. In this embodiment, on the one hand, the model service framework located between the target application and the AI model can provide single-model links and multi-model links for the target application, breaking through the traditional model application mode in which the target application directly calls the model; on the other hand, based on the model service framework, the two model links coexist and the access conditions of the multi-model link are set, which can not only process the target task in parallel based on multiple AI models on the multi-model link, improve the accuracy of the reasoning process, meet the model requirements in application scenarios with higher accuracy requirements, and more effectively alleviate the hallucination problem, but also meet the model requirements of application scenarios with lower accuracy requirements based on a single model link, thereby balancing the model call cost and model service quality.
[0110] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. Figure 5 As shown, in practice, the electronic device includes: a memory 501 and a processor 502 .
[0111] Memory 501 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, data structures, contact data, phone book data, messages, images, videos, etc.
[0112] The processor 502 is coupled to the memory 501 and is used to execute the computer program in the memory 501, so as to: receive task description information of a target task submitted by a target application, where the target task is generated by the target application according to the service request information of the user; if the target task meets the access conditions of the multi-model link, call multiple AI models on the multi-model link based on the task description information, and perform parallel processing on the target task to obtain multiple initial task results; perform decision processing on the multiple initial task results according to the task type of the target task to obtain a first target task result; and return the first target task result to the target application so that the target application can provide the user with a service adapted to the service request information according to the first target task result.
[0113] Optionally, the processor 502 is also used to: if the target task does not meet the access conditions of the multi-model link, call a single AI model on the single model link based on the task description information to process the target task to obtain a second target task result; return the second target task result to the AI model so that the target application can provide the user with a service adapted to the service request information based on the second target task result.
[0114] Optionally, the access conditions of the multi-model link include at least one of the following: access conditions of the scene dimension, access conditions of the user dimension, and access conditions of the task dimension; the processor 502 is also used to: obtain at least one piece of information among the application scenario, user attributes, and task attributes corresponding to the target task as access judgment information; if the at least one piece of information meets the access conditions on the corresponding dimension, it is determined that the target task meets the access conditions of the multi-model link.
[0115] Optionally, when the processor 502 obtains at least one of the application scenarios, user attributes, and task attributes corresponding to the target task, it is specifically used to: determine, based on the access conditions of the multi-model link, to request at least one of the application scenarios, user attributes, and task attributes corresponding to the target task from the target application as access judgment information; send an information acquisition request to the target application, and receive the at least one of the information returned by the target application based on the information acquisition request; or, receive at least one of the application scenarios, user attributes, and task attributes corresponding to the target task issued by the target application together with the target task.
[0116] Optionally, the model service framework includes an access layer, which provides a configuration interface to the outside world, and the processor 502 is also used to: in response to the configuration operation on the configuration interface, configure the service identifier corresponding to the target application, the access conditions of the multi-model link, and the mapping relationship between the task type and the decision-making method; wherein the service identifier is bound to the access conditions and the mapping relationship.
[0117] Optionally, the configuration interface is implemented as a configuration interface. When the processor 502 configures the service identifier corresponding to the target application, the access conditions of the multi-model link, and the mapping relationship between the task type and the decision-making method in response to the configuration operation on the configuration interface, it is specifically used to: display the configuration interface, which includes a service identifier configuration item, an access condition configuration item, and a task type configuration item; in response to the configuration operation of the service identifier configuration item, display a service identifier list, and in response to the selection operation of the service identifier list, use the selected service identifier as the service identifier corresponding to the target application; based on the service identifier corresponding to the target application, determine the An access condition list corresponding to the access condition configuration item and a task type list corresponding to the task type configuration item; different service identifiers correspond to different access condition lists and task type lists; in response to the configuration operation of the access condition configuration item, the access condition list is displayed, and in response to the selection operation of the access condition list, the selected access condition is used as the access condition of the multi-model link; in response to the configuration operation of the task type configuration item, the task type list is displayed, and in response to the selection operation of the task type list, a mapping relationship between the task type and the decision-making method is generated according to the selected task type; wherein different task types correspond to different decision-making methods.
[0118] Optionally, the processor 502 is also used to: respond to the configuration operation on the configuration interface, configure multiple AI models on the multi-model link based on the service identifier of the target application, and the multiple AI models are at least part of the AI model set accessed by the model service framework; wherein, the multiple AI models on the multi-model link corresponding to different service identifiers are different.
[0119] Optionally, the model service framework includes a model capability layer, which provides a single-model link, a multi-model link and a thread pool; the processor 502 calls multiple AI models on the multi-model link based on the task description information, and performs parallel processing on the target task to obtain multiple initial task results, specifically for: selecting multiple threads from the thread pool, establishing a correspondence between the multiple threads and the multiple AI models on the multi-model link; sending the task description information to the multiple threads, and starting the multiple threads; wherein, after the thread is started, the task description information is generated into a model prompt word according to the prompt word structure of the corresponding AI model, and the corresponding AI model is called according to the model prompt word to perform task processing to obtain the initial task result.
[0120] Optionally, when the processor 502 starts the multiple threads, it is specifically used to: start the multiple threads in multiple independent operating environments, and the multiple independent operating environments are isolated from each other to achieve service isolation of the multiple AI models, and the independent operating environments are servers, virtual machines or containers.
[0121] Optionally, when the processor 502 performs decision processing on the multiple initial task results according to the task type of the target task to obtain a first target task result, it is specifically used to: if the target task is a classification task, vote on the multiple initial task results, and use the initial task result pointed to by the voting result as the first target task result; if the target task is a regression task, integrate the multiple initial task results to obtain the second target task result.
[0122] Optionally, the multiple initial task results include two classification results. When the processor 502 votes on the multiple initial task results and takes the initial task result pointed to by the voting result as the first target task result, it is specifically used to: vote on the multiple initial task results to obtain a voting result; if the voting result indicates that the votes for the two classification results are different, take the classification result with more votes in the voting result as the first target task result; if the voting result indicates that the votes for the two classification results are the same, and the number of the multiple initial task results is less than the number of the multiple AI models, then select one from the AI models that have not output the initial task result as a retry model; call the retry model based on the task description information to process the target task, and obtain a retry task result, which is any classification result; re-vote the retry task result and the multiple initial task results to obtain a first re-voting result; and take the classification result with more votes in the first re-voting result as the first target task result.
[0123] Optionally, the processor 502 is also used to: if no retry task result is obtained, then based on the attribute information of multiple AI models that output the initial task result, discard the initial task result output by the AI model whose attribute information meets the set conditions, and re-vote the remaining initial task results to obtain a second re-voting result; use the classification result with more votes in the second re-voting result as the first target task result; the attribute information of the AI model includes at least one of parameter scale, service performance, model type and priority.
[0124] Optionally, when the processor 502 integrates the multiple initial task results to obtain the second target task result, it is specifically used to: if the multiple initial task results are numerical results, perform numerical calculation on the numerical values of the multiple initial task results, and use the results of the numerical calculation as the second target task result; if the multiple initial task results are text results, perform semantic understanding on the texts of the multiple initial task results to obtain multiple semantic information; analyze the logical relationship between the multiple semantic information, and integrate and optimize the texts of the multiple initial task results according to the logical relationship to obtain the second target task result.
[0125] Further, if Figure 5 As shown, the electronic device also includes: a communication component 503, a display 504, a power component 505, an audio component 506 and other components. Figure 5 Only some components are shown schematically, which does not mean that the electronic device only includes Figure 5 In addition, Figure 5 The components in the dotted box are optional components, not mandatory components, and the specific components depend on the product form of the working node. The working node of this embodiment can be implemented as a terminal device such as a desktop computer, laptop computer, smart phone or IOT device, or a server device such as a conventional server, cloud server or server array. If the working node of this embodiment is implemented as a terminal device such as a desktop computer, laptop computer, smart phone, etc., it can include Figure 5 If the working node of this embodiment is implemented as a server device such as a conventional server, a cloud server or a server array, it may not include Figure 5 Components within the dotted box.
[0126] The above-mentioned memory can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.
[0127] The communication component is configured to facilitate wired or wireless communication between the device in which the communication component resides and other devices. The device in which the communication component resides can access a wireless network based on a communication standard, such as a 2G, 3G, 4G / LTE, 5G, or other mobile communication network, or a combination thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel.
[0128] The above-mentioned display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundary of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.
[0129] The power supply assembly provides power to various components of the device in which the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.
[0130] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as call mode, recording mode, and voice recognition mode, the microphone is configured to receive external audio signals. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.
[0131] Accordingly, an embodiment of the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the steps in the above method embodiment. The computer-readable storage medium includes volatile or non-volatile or a combination thereof, and may be removable or non-removable. Examples of computer-readable storage media include, but are not limited to, phase-change random access memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), flash memory or other memory technology, CD-ROM, digital versatile disc (DVD) or other optical storage, magnetic cassette, tape disk storage or other magnetic storage device or any other non-transmission medium.
[0132] Accordingly, an embodiment of the present application further provides a computer program product, which includes a computer program or instructions, and when the computer program or instructions are executed by a processor, the processor is enabled to implement the steps in the above-mentioned method embodiment. It should be understood that each process or a combination of multiple processes in the above-mentioned method flow can be implemented by a computer program or instruction. In addition, these computer programs or instructions can be applied to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device, so that the processor of the general-purpose computer, the special-purpose computer, the embedded processor or other programmable data processing device can be implemented as a device for implementing the corresponding functions in the above-mentioned method embodiment.
[0133] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0134] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.
Claims
1. A model calling method, characterized in that: Applied to a model service framework, the model service framework is located between the target application and the AI model, and can provide single-model links and multi-model links for the target application. The method includes: receiving task description information of a target task submitted by a target application, wherein the target task is generated by the target application according to service request information of a user; If the target task meets the admission conditions of the multi-model link, multiple AI models on the multi-model link are called based on the task description information to process the target task in parallel to obtain multiple initial task results; Performing decision processing on the multiple initial task results according to the task type of the target task to obtain a first target task result; The first target task result is returned to the target application, so that the target application can provide the user with a service adapted to the service request information according to the first target task result.
2. The method according to claim 1, characterized in that Also includes: If the target task does not meet the admission conditions of the multi-model link, calling a single AI model on the single-model link based on the task description information to process the target task to obtain a second target task result; The second target task result is returned to the AI model so that the target application can provide the user with a service adapted to the service request information based on the second target task result.
3. The method according to claim 1, characterized in that The access condition of the multi-model link includes at least one of the following: an access condition of the scene dimension, an access condition of the user dimension, and an access condition of the task dimension; the method further includes: Obtaining at least one of an application scenario, user attributes, and task attributes corresponding to the target task as admission judgment information; If the at least one information meets the access conditions on the corresponding dimensions respectively, it is determined that the target task meets the access conditions of the multi-model link.
4. The method according to claim 3, characterized in that Obtaining at least one of the application scenario, user attributes, and task attributes corresponding to the target task, including: According to the access conditions of the multi-model link, determine to request at least one of the application scenario, user attributes and task attributes corresponding to the target task from the target application as access judgment information; send an information acquisition request to the target application, and receive the at least one information returned by the target application according to the information acquisition request; or Receive at least one of an application scenario, a user attribute, and a task attribute corresponding to the target task, which is sent by the target application along with the target task.
5. The method according to any one of claims 1 to 4, characterized in that The model service framework includes an access layer, and the access layer provides a configuration interface to the outside. The method further includes: In response to the configuration operation on the configuration interface, the service identifier corresponding to the target application, the access conditions of the multi-model link, and the mapping relationship between the task type and the decision-making method are configured; wherein the service identifier is bound to the access conditions and the mapping relationship.
6. The method according to claim 5, characterized in that The configuration interface is implemented as a configuration interface. Then, in response to the configuration operation on the configuration interface, the service identifier corresponding to the target application, the access conditions of the multi-model link, and the mapping relationship between the task type and the decision method are configured, including: Displaying the configuration interface, which includes a service identification configuration item, an access condition configuration item, and a task type configuration item; In response to a configuration operation of the service identification configuration item, a service identification list is displayed, and in response to a selection operation on the service identification list, the selected service identification is used as the service identification corresponding to the target application; Based on the service identifier corresponding to the target application, determining the access condition list corresponding to the access condition configuration item and the task type list corresponding to the task type configuration item; different service identifiers correspond to different access condition lists and task type lists; In response to a configuration operation of the access condition configuration item, displaying the access condition list, and in response to a selection operation on the access condition list, using the selected access condition as the access condition of the multi-model link; In response to the configuration operation of the task type configuration item, the task type list is displayed. In response to the selection operation of the task type list, a mapping relationship between the task type and the decision method is generated according to the selected task type; wherein different task types correspond to different decision methods.
7. The method according to claim 5, characterized in that Also includes: In response to the configuration operation on the configuration interface, multiple AI models on the multi-model link are configured based on the service identifier of the target application, and the multiple AI models are at least part of the AI model set accessed by the model service framework; wherein, the multiple AI models on the multi-model link corresponding to different service identifiers are different.
8. The method according to any one of claims 1 to 4, 6 and 7, characterized in that The model service framework includes a model capability layer, which provides single-model links, multi-model links and a thread pool; Based on the task description information, multiple AI models on the multi-model link are called to process the target task in parallel to obtain multiple initial task results, including: Selecting multiple threads from the thread pool and establishing a correspondence between the multiple threads and the multiple AI models on the multi-model link; The task description information is sent to the multiple threads, and the multiple threads are started; wherein, after the threads are started, the task description information is generated into model prompt words according to the prompt word structure of the corresponding AI model, and the corresponding AI model is called according to the model prompt words to perform task processing to obtain an initial task result.
9. The method according to claim 8, characterized in that Starting the multiple threads includes: starting the multiple threads in multiple independent operating environments, the multiple independent operating environments are isolated from each other to achieve service isolation of the multiple AI models, and the independent operating environments are servers, virtual machines or containers.
10. The method according to any one of claims 1 to 4, 6 and 7, characterized in that According to the task type of the target task, decision processing is performed on the multiple initial task results to obtain a first target task result, including: If the target task is a classification task, voting is performed on the multiple initial task results, and the initial task result pointed to by the voting result is used as the first target task result; If the target task is a regression task, the multiple initial task results are integrated to obtain the second target task result.
11. The method according to claim 10, characterized in that The multiple initial task results include two classification results, and voting on the multiple initial task results is performed, and the initial task result pointed to by the voting result is used as the first target task result, including: Voting the multiple initial task results to obtain voting results; If the voting result indicates that the votes for the two classification results are different, the classification result with more votes in the voting result is used as the first target task result; If the voting result indicates that the two classification results have the same number of votes, and the number of the multiple initial task results is less than the number of the multiple AI models, then one is selected from the AI models that have not output the initial task result as a retry model; based on the task description information, the retry model is called to process the target task, and a retry task result is obtained, and the retry task result is any classification result; the retry task result and the multiple initial task results are re-voted to obtain a first re-voting result; and the classification result with more votes in the first re-voting result is used as the first target task result.
12. The method according to claim 11, characterized in that Also includes: If no retry task result is obtained, the initial task result output by the AI model whose attribute information meets the set conditions is discarded based on the attribute information of multiple AI models that output the initial task result, and the remaining initial task results are re-voted to obtain a second re-voting result; the classification result with more votes in the second re-voting result is used as the first target task result; the attribute information of the AI model includes at least one of parameter scale, service performance, model type and priority.
13. The method according to claim 10, characterized in that Integrating the multiple initial task results to obtain the second target task result includes: If the multiple initial task results are numerical results, numerical calculation is performed on the numerical values of the multiple initial task results, and the results of the numerical calculation are used as the second target task result; If the multiple initial task results are text-type results, semantic understanding is performed on the texts of the multiple initial task results to obtain multiple semantic information; the logical relationship between the multiple semantic information is analyzed, and the texts of the multiple initial task results are integrated and optimized according to the logical relationship to obtain the second target task result.
14. An electronic device, characterized in that: include: memory and processor; The memory is used to store computer programs / instructions; The processor is coupled to the memory and configured to execute the computer program / instructions to implement the steps of the method according to any one of claims 1 to 13.
15. A computer-readable storage medium, characterized in that A computer program / instruction is stored, and the computer program / instruction is executed by a processor to implement the steps in the method according to any one of claims 1 to 13.
16. A computer program product, characterized in that include: A computer program / instruction, wherein the computer program / instruction is executed by a processor to implement the steps of the method according to any one of claims 1 to 13.