Model scheduling method and device and electronic equipment

By identifying the type of computing task and selecting the most suitable model for processing, the model scheduling problem in multi-model mixed deployment scenarios is solved, the efficiency and accuracy of computing tasks are improved, and resource allocation and interface adaptation are optimized.

CN120688357APending Publication Date: 2025-09-23BEITAI ZHENHUAN (CHONGQING) TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510794387.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-13
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

In the mixed deployment scenario of multiple models, existing technologies cannot achieve accurate model scheduling, resulting in inefficient computing task processing and unreasonable resource allocation.

Method used

By obtaining the domain information of the target request, using the feature vocabulary and machine learning algorithm to identify the computing task type, combining historical performance data and response time, the most suitable model is selected to process the computing task, and the request format is converted to adapt to the interface requirements of different models. The dynamic model selection algorithm and load balancing strategy are used to optimize model scheduling.

Benefits of technology

It achieves precise model scheduling in multi-model hybrid deployment scenarios, improves computing task processing efficiency and accuracy, optimizes resource allocation, and reduces expansion and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120688357A_ABST
    Figure CN120688357A_ABST
Patent Text Reader

Abstract

The invention discloses a model scheduling method and device and electronic equipment. The method comprises the steps that a target request sent by a target object is acquired, and the target request comprises a target calculation task; field information to which the target calculation task belongs is determined, and the field information is used for indicating a calculation field type of the target calculation task; determining a target model corresponding to the domain information from a plurality of models; and converting a request message of the target request into a request format corresponding to the target model, and executing the target request by adopting the target model. According to the method and the device, the technical problem that accurate model scheduling cannot be realized in a multi-model mixed deployment scene when a calculation task is carried out by utilizing a model in the related technology is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and more specifically, to a method, device, and electronic device for model scheduling. Background Art

[0002] In the current field of scientific computing software, with the development of artificial intelligence technology, the use of models to process computing tasks has become an important method. However, when using models for computing tasks, related technologies cannot achieve accurate model scheduling in scenarios where multiple models are mixed and deployed.

[0003] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0004] The embodiments of the present application provide a method, device and electronic device for model scheduling to at least solve the technical problem that when the related technology uses models to perform computing tasks, accurate model scheduling cannot be achieved in a mixed deployment scenario of multiple models.

[0005] According to one aspect of an embodiment of the present application, a model scheduling method is provided, comprising: obtaining a target request sent by a target object, wherein the target request includes a target computing task; determining domain information to which the target computing task belongs, wherein the domain information is used to indicate the computing domain type of the target computing task; determining a target model corresponding to the domain information from multiple models; converting a request message of the target request into a request format corresponding to the target model, and executing the target request using the target model.

[0006] In some embodiments of the present application, determining the domain information to which the target computing task belongs includes: obtaining description information of the target computing task from the target request, and extracting keywords corresponding to the description information; matching the keywords with a feature vocabulary to obtain a matching result, wherein the feature vocabulary is used to store feature words corresponding to multiple computing fields; and determining the domain information to which the target computing task belongs based on the matching result.

[0007] In some embodiments of the present application, the feature words of each sub-field of multiple computing fields in the feature vocabulary correspond to a feature word subset; matching the keywords with the feature vocabulary to obtain matching results includes: determining the word frequency of each keyword in the feature word subset of the corresponding sub-field, wherein the word frequency is used to indicate the number of times the keyword appears in the sub-field; determining the inverse document frequency of each keyword in the feature vocabulary, wherein the inverse document frequency is used to indicate the uniqueness of the keyword to the feature vocabulary; and determining the matching result based on the word frequency and inverse document frequency corresponding to each keyword.

[0008] In some embodiments of the present application, a matching result is determined based on the word frequency and inverse document frequency corresponding to each keyword, including: comparing a target keyword with a feature vocabulary to obtain a comparison result, wherein the target keyword is any keyword in the keywords; when the comparison result indicates that the target keyword belongs to a target sub-field of the feature vocabulary, a first weight is used to determine a first score for the target keyword and the target sub-field, wherein the first weight includes a weight determined based on the word frequency and inverse document frequency of the target keyword, and the first score is used to represent the correlation between the target keyword and the target sub-field; determining a score corresponding to each sub-field in the feature vocabulary based on the first score, and determining the score as the matching result.

[0009] In some embodiments of the present application, it also includes: when the comparison result indicates that the target keyword does not belong to any sub-field of the feature vocabulary, using a second weight to determine a second score of the target keyword, wherein the second weight is less than the first weight, and the second score is used to represent the correlation between the target keyword and the sub-field obtained by fuzzy matching between the target keyword and the feature vocabulary.

[0010] In some embodiments of the present application, determining a target model corresponding to domain information from multiple models includes: obtaining historical performance data of the multiple models under the computing domain type; determining matching scores corresponding to the multiple models under the computing domain type based on the historical performance data, wherein the matching scores are used to quantitatively represent the degree of matching between the target computing task and the multiple models; and determining the target model based on the matching scores.

[0011] In some embodiments of the present application, before determining the target model based on the matching score, it also includes: obtaining the response time corresponding to multiple models within a preset time period, wherein the response time is used to quantify the processing speed of multiple models processing computing tasks; determining the computing costs corresponding to multiple models, wherein the computing cost is used to quantify the resource consumption cost of multiple models processing computing tasks; determining the target model based on the matching score, response time and computing cost.

[0012] In some embodiments of the present application, a request message of a target request is converted into a request format corresponding to a target model, including: determining interface characteristics corresponding to the target model, wherein the interface characteristics include request header characteristics and request parameter characteristics; converting the request header in the request message according to the request header characteristics to obtain a target request header; converting the request parameters in the request message according to the request parameter characteristics to obtain target request parameters; and determining a unified interface request corresponding to the target request header and target request parameters.

[0013] In some embodiments of the present application, it also includes: obtaining historical conversation data of the target object; constructing a customer profile of the target object based on the historical conversation data, wherein the customer profile is used to describe the model and computing tasks preferred by the target object; predicting the prediction computing task of the target object based on the current conversation data; preloading the model corresponding to the prediction computing task to the target location, wherein the target location includes a local cache or a server node whose distance from the target object meets a preset threshold.

[0014] According to another aspect of an embodiment of the present application, a model scheduling device is also provided, including: an acquisition module for acquiring a target request sent by a target object, wherein the target request includes a target computing task; a determination module for determining domain information to which the target computing task belongs, wherein the domain information is used to indicate the computing domain type of the target computing task; a matching module for determining a target model corresponding to the domain information from multiple models; and an execution module for converting a request message of the target request into a request format corresponding to the target model, and executing the target request using the target model.

[0015] According to another aspect of the embodiments of the present application, an electronic device is provided, including: a memory and a processor, the memory being used to store program instructions; the processor being connected to the memory and being used to execute the method for implementing the above-mentioned model scheduling.

[0016] According to another aspect of the embodiments of the present application, a non-volatile storage medium is provided, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the above-mentioned model scheduling method by running the computer program.

[0017] According to another aspect of the embodiments of the present application, a computer program product is provided, including computer instructions, which implement the above-mentioned model scheduling method when executed by a processor.

[0018] In an embodiment of the present application, a model scheduling method is adopted to obtain the target request and determine its domain information, and then select the target model corresponding to the domain information, and convert the request message into the request format of the target model for execution, thereby achieving the purpose of accurately scheduling the model, thereby achieving the technical effect of improving the efficiency and accuracy of computing task processing, and thus solving the technical problem that the relevant technology cannot achieve accurate model scheduling in a mixed deployment scenario of multiple models when using models to perform computing tasks. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0020] Figure 1 This is a hardware structure block diagram of a computer terminal according to a model scheduling method according to an embodiment of the present application;

[0021] Figure 2 is a system architecture diagram of a model scheduling method according to an embodiment of the present application;

[0022] Figure 3 is a flow chart of a dynamic selection algorithm of a model scheduling method according to an embodiment of the present application;

[0023] Figure 4 is a flow chart of a model scheduling method according to an embodiment of the present application;

[0024] Figure 5 It is a structural diagram of a model scheduling device according to an embodiment of the present application. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the present invention, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequential order. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in a sequence other than those illustrated or described herein. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device comprising a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0027] In order to better understand the embodiments of the present application, the technical terms involved in the embodiments of the present application are explained as follows:

[0028] Dynamic Model Selection Algorithm (DMSA): An algorithm that dynamically selects the optimal model based on real-time data and preset rules. It is typically used in a multi-model environment to adapt to different task requirements. In the embodiment of the present application, DMSA dynamically selects the model that best suits the target computing task by constructing a multidimensional evaluation matrix (including response delay, computational complexity, and domain matching) and a preloading mechanism based on user behavior prediction, thereby optimizing model scheduling efficiency and response speed.

[0029] In the current field of scientific computing software, with the development of artificial intelligence technology, using models to process computing tasks has become an important method. However, related technologies have many shortcomings:

[0030] (1) Contradictions of single model services: Cloud-based single model services face the problem of high response latency, as data transmission between the cloud and local devices takes a lot of time; while local single model services are limited by local computing power and are inefficient in processing complex computing tasks, making it difficult to meet users' requirements for computing speed and accuracy.

[0031] (2) Cost issues caused by differences in API interfaces: There are significant differences in the API interfaces provided by different manufacturers, including differences in authentication methods and request parameters. This makes the expansion and maintenance costs extremely high when integrating models from multiple manufacturers, and requires separate adaptation and development for each manufacturer's interface.

[0032] (3) Deficiencies of traditional load balancing strategies: Traditional load balancing strategies do not fully consider the knowledge characteristics of the scientific computing field. For example, in scientific computing, matrix operation requests have different model requirements from other types of requests. Specific models are needed to ensure the accuracy and efficiency of calculations. However, traditional strategies do not distinguish between request types when allocating tasks, resulting in unreasonable allocation of computing resources and affecting computing results.

[0033] (4) After analyzing mainstream scientific computing software such as Matlab, its AI module only supports single cloud model calling and does not solve the model scheduling problem in hybrid deployment scenarios.

[0034] In order to solve the above technical problems, the embodiments of the present application provide corresponding solutions, which are described in detail below.

[0035] The model scheduling method embodiment provided in the embodiment of the present application can be executed in a mobile terminal, a computer terminal or a similar computing device. Figure 1 FIG1 shows a hardware structure block diagram of a computer terminal for implementing a method for model scheduling. Figure 1As shown, the computer terminal 10 may include one or more (illustrated by 102a, 102b, ..., 102n in the figure) processors (the processor may include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission module 106 for communication functions connected via a wired and / or wireless network. In addition, it may also include: a display, a keyboard, a cursor control device, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the I / O interface), a network interface, and a BUS bus. It will be understood by those skilled in the art that Figure 1 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 1 More or fewer components than shown, or with Figure 1 Different configurations shown.

[0036] It should be noted that the one or more processors and / or other data processing circuits described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, independent processing module, or may be incorporated in whole or in part into any of the other components of the computer terminal 10. As described in the embodiments of the present application, the data processing circuitry serves as a processor control (e.g., selection of a variable resistor terminal path connected to an interface).

[0037] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the model scheduling method in the embodiment of the present application. The processor executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implementing the above-mentioned model scheduling method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 104 may further include a memory remotely located relative to the processor, and these remote memories may be connected to the computer terminal 10 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0038] The transmission module 106 is configured to receive or transmit data via a network. A specific example of the aforementioned network may include a wireless network provided by the communications provider of the computer terminal 10. In one embodiment, the transmission module 106 includes a network interface controller (NIC), which can be connected to other network devices via a base station to enable communication with the Internet. In another embodiment, the transmission module 106 may be a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0039] The display may be, for example, a touch screen liquid crystal display (LCD) that enables a user to interact with a user interface of the computer terminal 10 .

[0040] It should be noted that, in some optional embodiments, the above Figure 1 The computer terminal shown may include hardware elements (including circuits), software elements (including computer code stored on a computer-readable medium), or a combination of hardware elements and software elements. Figure 1 This is merely one example of a particular embodiment and is intended to illustrate the types of components that may be present in the computer terminal described above.

[0041] In the above-mentioned operating environment, an embodiment of the present application provides a method embodiment of model scheduling. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0042] Figure 2 is a flow chart of a model scheduling method according to an embodiment of the present application. Figure 2 As shown, the method includes the following steps:

[0043] Step S202: Obtain a target request sent by a target object, wherein the target request includes a target computing task.

[0044] In step S202, the target object refers to a user or application using scientific computing software. The target object sends a request to perform a specific scientific computing task. A target request is a request from a user or application to initiate or execute a specific computing task, such as matrix inversion or differential equation solving. The target computing task, the core content of the target request, determines the specific direction of subsequent model selection and task processing.

[0045] In some embodiments of the present application, the system first obtains a target request sent by a target object, for example, receives user input, such as through a graphical user interface (GUI), a command line interface (CLI), or an API call; parses and understands the intent of the request, that is, extracts the target computing task that the user wants to perform. For example, if the user input is "solve the inverse matrix of matrix A", the system needs to recognize that this is a computing task request for matrix inversion.

[0046] After receiving the target request, it can also be pre-processed, including standardizing the request format and parameters, to ensure that the request can be correctly parsed and processed by the unified interface adaptation layer within the system. For example, keywords in the user request can be converted into a standard format recognized by the system, such as converting "matrix inversion" into a matrix operation request type within the system.

[0047] Step S204: Determine the domain information to which the target computing task belongs, wherein the domain information is used to indicate the computing domain type of the target computing task.

[0048] In step S204, the domain information refers to the scientific computing domain or type to which the target computing task belongs, such as descriptive information about matrix operations, differential equation solving, and fast Fourier transforms (FFTs). Computing domain types are common classifications in scientific computing, used to describe the nature and requirements of computing tasks. These include specific domains such as mathematics, physics, chemistry, and bioinformatics, as well as more fine-grained domains such as linear algebra and dynamical systems.

[0049] In some embodiments of this application, a feature word library for the scientific computing field can be established, containing key words that represent the characteristics of each field. For example, "matrix" and "inversion" are associated with the field of matrix operations; "differential equation" and "initial value problem" are associated with the field of dynamics. It should be noted that the weight of the feature word in the feature word library can also be dynamically adjusted based on the frequency and importance of the feature word in the user request, ensuring that the system can more accurately and quickly identify the field of the computing task.

[0050] Based on the feature word library, a real-time classifier can be developed to analyze the keywords in the target request and determine the domain type of the computing task. That is, the user's computing request is analyzed in real time, and the type of scientific computing field to which it belongs is determined based on the feature words contained in the request. Then, the request is assigned to the most appropriate model for processing based on the domain characteristics, thereby improving the processing efficiency and accuracy of the computing task. For example, the TF-IDF (Term Frequency-Inverse Document Frequency) algorithm can be used to weight the feature words in the request to calculate the similarity between the request and the features of each domain, and ultimately determine the type of computing field to which the request belongs. Alternatively, a machine learning classification model, such as a support vector machine (SVM) or a deep learning model, can be trained using historical request data as a training set. The model can automatically learn the association between feature words and fields, thereby improving the accuracy of request classification.

[0051] In some embodiments of the present application, the domain information to which the target computing task belongs can be determined by the following steps: obtaining description information of the target computing task from the target request, and extracting keywords corresponding to the description information; matching the keywords with a feature vocabulary to obtain a matching result, wherein the feature vocabulary is used to store feature words corresponding to multiple computing fields; and determining the domain information to which the target computing task belongs based on the matching result.

[0052] The description information is a detailed description of the computing task provided by the user in the target request. It can be a mathematical formula, a specific numerical value or a text description, such as "solve the inverse matrix of matrix A". Keywords are words extracted from the description information that can characterize the characteristics of the computing task, such as "matrix", "inverse matrix", etc., which are the key basis for the system to determine the field to which the computing task belongs.

[0053] In order to accurately identify the specific requirements of computing tasks from the unstructured or semi-structured descriptions provided by users, natural language processing (NLP) technology can be used to analyze the text of user requests, use part-of-speech tagging, entity recognition and other technologies to extract key information related to the computing task, or use predefined patterns or regular expressions to match specific keywords or phrases from user requests.

[0054] Methods for matching keywords with feature lexicons may include, but are not limited to: directly matching keywords with words in the feature lexicon to find the most relevant computing domain, or using text similarity algorithms in NLP (such as cosine similarity and Jaccard similarity) to measure the similarity between keywords and entries in the feature lexicon, and selecting the domain with the highest similarity as the domain information for the target computing task.

[0055] Based on the keyword matching results, combined with domain priority, the complexity of the computing task, and historical data, a decision engine is used to determine the domain information for the computing task. It should be noted that machine learning can also be used to learn preference patterns in historical user requests. Even if the keywords don't match exactly, the most likely domain information can be inferred based on the user's behavior patterns.

[0056] In some embodiments of the present application, the feature word of each sub-field of multiple computing fields in the feature word library corresponds to a feature word subset. Based on this, the matching result can be determined by: determining the word frequency of each keyword in the feature word subset of the corresponding sub-field, where the word frequency is used to represent the number of times the keyword appears in the sub-field; determining the inverse document frequency of each keyword in the feature word library, where the inverse document frequency is used to represent the uniqueness of the keyword with respect to the feature word library; and determining the matching result based on the word frequency and inverse document frequency corresponding to each keyword.

[0057] A feature word subset is a set of words in the feature word library that corresponds to a specific computing subfield. These words can significantly describe and distinguish the computing tasks in that subfield. For example, the matrix operation subfield may contain feature words such as "matrix", "determinant", and "eigenvalue".

[0058] Term Frequency (TF) indicates how often a keyword appears in a subset of characteristic words within a particular computing subfield. It measures the relevance of the keyword to that subfield. A higher TF indicates a stronger relevance. Inverse Document Frequency (IDF) measures the rarity of a keyword within the entire characteristic vocabulary, specifically its contribution to distinguishing different computing subfields. A higher IDF value indicates a greater ability of the keyword to distinguish different computing subfields.

[0059] Specifically, for each keyword extracted from the target request, the system needs to determine its frequency within the subset of characteristic words in each computing subfield. For example, this can be accomplished by analyzing a large number of historical computing requests, counting the frequency of each keyword in different subfields, and constructing a frequency table.

[0060] For each keyword, the system also needs to calculate its inverse document frequency in the entire feature vocabulary to evaluate the uniqueness and distinguishing ability of the keyword. For example, the number of sub-domain feature word subsets in the feature vocabulary that contain a certain keyword, that is, the document frequency (DF), is counted. The inverse document frequency is calculated using the IDF formula: IDF = log (total number of sub-domain feature word subsets in the feature vocabulary / number of sub-domain feature word subsets that contain the keyword).

[0061] Some keywords may have different meanings in different sub-fields, resulting in inaccurate calculations of word frequency and IDF. For example, "linear" has different meanings in the sub-fields of "linear algebra" and "linear programming." To address this issue, when calculating word frequency, we can not only consider the keyword itself, but also comprehensively analyze the context in which the keyword appears to ensure that the word frequency can reflect the actual meaning of the keyword in a specific context. In addition, for polysemous words, we can also introduce an adjustment mechanism to calculate the IDF separately according to the specific sub-field represented by the keyword in different contexts, rather than calculating it uniformly. For example, the IDF calculation of "linear" in "linear algebra" will be different from that in "linear programming."

[0062] In some embodiments of the present application, a matching result is determined based on the word frequency and inverse document frequency corresponding to each keyword, specifically: a target keyword is compared with a feature vocabulary to obtain a comparison result, wherein the target keyword is any keyword in the keywords; when the comparison result indicates that the target keyword belongs to a target sub-field of the feature vocabulary, a first weight is used to determine a first score for the target keyword and the target sub-field, wherein the first weight includes a weight determined based on the word frequency and inverse document frequency of the target keyword, and the first score is used to represent the correlation between the target keyword and the target sub-field; a score corresponding to each sub-field in the feature vocabulary is determined based on the first score, and the score is determined as the matching result.

[0063] The target keyword is extracted from the computing task description submitted by the user and is used to identify the keywords of the specific nature or type of the computing task. The target sub-domain is a sub-domain in the feature vocabulary. When the target keyword matches the feature word under the sub-domain, it is considered that the target keyword belongs to the target sub-domain.

[0064] The first weight is calculated by combining the target keyword's term frequency (TF) and inverse document frequency (IDF) to measure the target keyword's importance and uniqueness within the target sub-field. The first score, calculated using the first weight, represents the degree of relevance between the target keyword and the target sub-field. A higher first score indicates a stronger relevance between the keyword and the sub-field, and a higher likelihood of belonging to that sub-field.

[0065] After determining that the target keyword belongs to a certain sub-field, the system will calculate the first score of the target keyword in the sub-field based on the TF-IDF algorithm to measure the correlation between the keyword and the sub-field. Based on the first score of each keyword, the score of each sub-field in the feature vocabulary is comprehensively determined to finally form a matching result. In some embodiments of the present application, the first scores of all target keywords in each sub-field can be weighted averaged, and the weights can be adjusted based on the word frequency and inverse document frequency of the keywords in the field, and the scores of all sub-fields are sorted, and the field with the highest score is selected as the field information of the calculation task.

[0066] For example, when a user submits a computing task request involving "matrix inversion", the system first identifies and extracts "matrix" and "inversion" as target keywords. Subsequently, the system compares these keywords with the subset of feature words in the feature vocabulary to determine that they belong to the "matrix operation" sub-field. Then, based on the word frequency and inverse document frequency of the keywords in the "matrix operation" sub-field, the system calculates the first score of "matrix" and "inversion" in the "matrix operation" sub-field. Finally, the system comprehensively considers the first scores of all keywords through a weighted average method, and determines that the "matrix operation" sub-field has the highest score, thereby marking the computing task as the "matrix operation" field.

[0067] When the comparison result indicates that the target keyword does not belong to any sub-field of the feature vocabulary, a second weight can be used to determine a second score of the target keyword, wherein the second weight is less than the first weight, and the second score is used to represent the correlation between the target keyword and the sub-field obtained by fuzzy matching between the target keyword and the feature vocabulary.

[0068] The second weighting is a lighter scoring criterion used when the target keyword does not directly match any sub-fields. It is used to score fuzzy matches between the keyword and potentially related sub-fields to compensate for the limitations of exact matches. The second score is a correlation score between the target keyword and the sub-fields obtained by fuzzy matching, calculated based on the second weighting, reflecting any indirect connection between the keyword and the sub-field.

[0069] Specifically, if the target keyword fails to be accurately matched with any sub-domain in the feature vocabulary, that is, no matching feature word subset is found in direct comparison, the system will enter the fuzzy matching stage. For example, using a dictionary or a pre-established synonym library, look for words with similar meanings to the target keyword to see if they belong to a sub-domain in the feature vocabulary. Alternatively, abstract the task description and try to match a broader computing field. For example, if there is no direct match for "matrix inversion", the system may generalize it to the field of "matrix operations".

[0070] It should be noted that the second weight is typically set to a certain percentage of the first weight, such as 70%-80%, to reduce the impact of indirect connections between keywords and sub-fields. In addition to the keywords themselves, the system can also consider the context surrounding the keywords, indirectly inferring the sub-field to which the target keyword belongs by analyzing other keywords that appear in the context.

[0071] For example, suppose a user submits a request containing the keyword "feature vector decomposition," but this specific expression is not yet included in the feature vocabulary. In this case, our system will not give up on task identification immediately, but will take the following steps:

[0072] (1) Through synonym search, the system may match "eigenvector decomposition" with broader domain keywords such as "matrix operations" or "linear algebra".

[0073] (2) The system will use the second weight to calculate the second score between the “eigenvector decomposition” and these generalized sub-fields. Although this score will be lower than the first score, it reflects the indirect correlation between the keyword and the field.

[0074] (3) The system comprehensively considers the scoring results of all keywords, including the first score obtained by exact matching and the second score obtained by fuzzy matching, to determine the domain information of the computing task.

[0075] Step S206: determining a target model corresponding to the domain information from the multiple models.

[0076] In step S206 above, the target model is the optimal model selected by the system after a comprehensive evaluation of all candidate models. It is used to process the current target computing task and ensure that the computing task is executed efficiently and accurately. For example, the target model can be determined by constructing a multidimensional evaluation matrix: an evaluation matrix containing three dimensions: response delay, computational complexity, and domain matching. By monitoring the model's response delay in real time, the speed at which it processes the task is evaluated; the complexity of the computing task is analyzed to select a model suitable for processing the task; and based on the scientific computing field to which the task belongs, the degree of match between the model and the field is determined, thereby comprehensively evaluating and selecting the optimal model.

[0077] In some embodiments of the present application, a target model corresponding to domain information can be determined from multiple models in the following manner: historical performance data of multiple models under the computing domain type is obtained; matching scores corresponding to the multiple models under the computing domain type are determined based on the historical performance data, wherein the matching scores are used to quantitatively represent the degree of matching between the target computing task and the multiple models; and the target model is determined based on the matching scores.

[0078] Historical performance data refers to the past records of models in the system executing tasks in different computing domains. It primarily includes key metrics such as the model's accuracy, response time, and resource consumption when handling specific types of computing tasks. The matching score quantifies the degree of fit between the model and the computing task based on the model's historical performance in handling tasks in a specific computing domain type. A higher score indicates a more suitable model for handling that type of computing task.

[0079] Specifically, the historical processing records of each model under different computing domain types are obtained from the database, and historical performance data is extracted from the historical processing records, including but not limited to indicators such as accuracy, response time, and computing resource consumption. Based on the historical performance data, the system scores the matching degree of each model in a specific computing domain, and quantifies the degree of adaptation of the model to the computing task. For example, weights are set for the performance of the model on different indicators. For example, when processing the "matrix inversion" task, the weights of accuracy and response time may be higher, and then the total score is calculated based on these weighted indicators. Based on the matching score of each model under the computing domain type, the system selects the model with the highest score as the target model for processing the current computing task.

[0080] It should be noted that if the scores of multiple models are similar, the system will consider distributing tasks to the clusters of these high-scoring models and optimize the overall computing performance through load balancing strategies.

[0081] For example, when receiving a "matrix inversion" task, the system first determines that the task falls into the "matrix operations" domain based on keyword analysis. The system then retrieves historical performance data for all models in the model library when handling "matrix operations" tasks, including accuracy and response time. Based on this data, the system uses a weighted scoring method and pre-trained machine learning models to calculate a match score for each model. Suppose model A has an accuracy of 95% and a response time of 1 second in the "matrix operations" domain; while model B has a slightly lower accuracy of 90%, its response time is only 0.5 seconds. Based on the urgency of the task (assuming response time is given a higher weight), the system may assign a higher match score to model B. Ultimately, based on the match scores of all models, the system automatically selects the model with the highest score to handle the "matrix inversion" task, ensuring efficient and accurate completion of the computational task.

[0082] Before determining the target model based on the matching score, the following steps can also be performed: obtaining the response time corresponding to multiple models within a preset time period, wherein the response time is used to quantify the processing speed of the multiple models in processing computing tasks; determining the computing costs corresponding to the multiple models, wherein the computing cost is used to quantify the resource consumption cost of the multiple models in processing computing tasks; determining the target model based on the matching score, response time and computing cost.

[0083] Response time refers to the time between a model receiving a computational task request and returning the results. It is a key metric for evaluating model processing speed and directly impacts user experience and computational efficiency. Computational costs include hardware resources required to run the model (such as CPU and GPU time), network bandwidth consumption, and API call fees.

[0084] To accurately assess a model's processing speed, the system needs to continuously monitor its response time over a period of time. For example, built-in monitoring tools can record the model's processing time and response status periodically or in real time to ensure the timeliness and reliability of the data. Furthermore, by analyzing the model's average response time over a period of time, performance fluctuations and trends can be identified, providing dynamic data support for model selection.

[0085] To evaluate the economic feasibility of a model, it is necessary to calculate the overall resource cost (i.e., computing cost) consumed by the model in processing computing tasks. In some embodiments of the present application, the hardware resources such as CPU time, GPU time, memory usage, and network data transmission volume used when the model processes tasks can be recorded to form a resource consumption list. For cloud-based model services, the system integrates and parses the service provider's billing rules (such as charging by the number of times, charging by resource usage), and calculates the actual API call fee.

[0086] Based on the domain information of the computational task, the model's processing speed, and its economic efficiency, the system ultimately determines a target model that both efficiently completes the task and balances cost-effectiveness. For example, different weights are assigned to the matching score, response time, and computational cost. Each model is comprehensively evaluated using weighted and / or more complex decision tree algorithms, and the model with the highest weighted score is selected. During the weighted decision-making process, the weight of computational cost is dynamically adjusted based on the user's or system's cost sensitivity. For example, in a resource-constrained environment, the system may prioritize computational cost in order to conserve resources.

[0087] Step S208: convert the request message of the target request into a request format corresponding to the target model, and execute the target request using the target model.

[0088] In the above step S208 , the request message is the original data representation of the target request, and may include, for example, a data format and parameters specific to a certain API or model.

[0089] Request format conversion converts the original request message of the target request into a request format that matches the target model, ensuring that the target model can correctly understand and process the computation task. Because different models or API services may have different requirements for the format and parameters of request messages, the system must perform format conversion before model execution to ensure that the target model can correctly process it.

[0090] In some embodiments of the present application, a standardized API adaptation layer can be constructed, which can automatically convert the request message into the corresponding format and parameters according to the specific requirements of the target model. The system maintains a parameter mapping table that records the parameter differences between different models and the corresponding mapping rules. When the target request is assigned to a certain model, the request parameters are converted according to the mapping table.

[0091] After converting the request message format, the system uses the target model to perform the actual computations on the request message. To improve model response speed and execution efficiency, the system uses model container technology, deploying the model close to the user or on a node with abundant computing resources. Request messages are then sent directly to the target model within the container for processing.

[0092] After the model is executed, the system can also verify the calculation results to ensure the accuracy and validity of the results. If problems are found in the results, such as calculation errors or format inconsistencies, the system can automatically provide feedback and reschedule until the correct calculation results are obtained.

[0093] In some embodiments of the present application, the request message of the target request is converted into a request format corresponding to the target model, specifically: the interface characteristics corresponding to the target model are determined, wherein the interface characteristics include request header characteristics and request parameter characteristics; the request header in the request message is converted according to the request header characteristics to obtain the target request header; the request parameters in the request message are converted according to the request parameter characteristics to obtain the target request parameters; and a unified interface request corresponding to the target request header and target request parameters is determined.

[0094] Request header features refer to the header information characteristics of the target model's communication interface, such as authentication method, version number, content type, etc., which are used to identify the request source, permissions, and data format. For example, based on the target model's requirements for data content type, adjust the "Content-Type" field in the request header, such as switching from JSON to XML (data encoding adjustment).

[0095] Request parameter characteristics refer to the specific requirements of the target model interface for input parameters, including parameter name, type, format, and default value, to ensure that the model can correctly receive and parse computation requests. For example, the "temperature" parameter of model A is mapped to the "top_p" parameter of model B to ensure semantic consistency of parameter transfer (parameter mapping). If the target model requires an integer and the request message provides a floating-point number, type conversion is performed. Furthermore, format adjustments are made based on the target model's required input format, such as string or list (data type and format conversion).

[0096] A unified interface request refers to a request message that has undergone format conversion and can be accepted and forwarded to the target model in a consistent manner by the intelligent scheduling system, regardless of the differences in the target model's native interface characteristics. For example, the converted request header and request parameters are encapsulated into a standardized request object, such as a Python dictionary or HTTP request, according to the target model's interface specifications.

[0097] In some embodiments of the present application, a preloading mechanism based on user behavior prediction can be developed, that is, by analyzing the user's historical session data, a user profile can be established. For example, the types of computing tasks frequently performed by the user, the model preferences used, etc. can be analyzed to predict the user's next possible computing needs, and the relevant models can be loaded in advance to a suitable location (such as a local cache or a server node close to the user), thereby reducing user waiting time and improving system response speed.

[0098] The preloading mechanism based on user behavior prediction includes: obtaining historical conversation data of the target object; building a customer profile of the target object based on the historical conversation data, wherein the customer profile is used to describe the model and computing tasks preferred by the target object; predicting the prediction computing task of the target object based on the current conversation data; preloading the model corresponding to the prediction computing task to the target location, wherein the target location includes a local cache or a server node whose distance from the target object meets a preset threshold.

[0099] Historical conversation data records all past requests and responses between the target user (user) and the intelligent scheduling system. It serves as the foundation for building user profiles and predicting their computing needs. A customer profile is a collection of personalized information, including the target user's preference model, common computing task types, and operating habits, derived from analyzing historical conversation data.

[0100] Predictive computing tasks use customer profile analysis to predict potential future computing tasks for the target object, helping the system prepare computing resources in advance. Model preloading involves preloading the AI ​​models that may be needed for predictive computing tasks into a local cache or a server node close to the target object. This reduces response delays during model loading and improves system responsiveness.

[0101] Specifically, each time a user interacts with the system, key information such as request content, model selection, task results, and response time is recorded; machine learning algorithms are used to analyze the user's historical request patterns, identify user preference models, common computing task types, and operating habits, and collect statistics on the user's request frequency and task type distribution; based on the user's current conversation content and historical behavior, the system predicts the computing tasks that the user is about to propose. For example, the system monitors the user's current conversation content, identifies keywords and computing demand expressions, and makes real-time predictions based on customer profiles. The current conversation is matched with historical conversation data for similarity to find the most similar conversation pattern and predict similar tasks that the user may propose; based on the predicted computing tasks and customer profile information, the system loads the model into the local cache or a server node close to the user in advance. For example, if the predicted computing task is of a higher frequency type, the system preloads the model into the user's local cache to ensure that the model can respond to the user's request immediately, or, based on the user's location information, the system dynamically selects a server node within a preset threshold from the user for model deployment to reduce network latency.

[0102] In a specific embodiment, assume the target user (user) is a researcher who frequently performs matrix operations. By chronically recording the user's computational requests, the system discovers that the user prefers to use the GLM4 model for operations such as matrix inversion, and that these requests occur very frequently. When constructing a customer profile, the system not only records the user's preferred model but also analyzes the user's request timing patterns, task type distribution, and operational habits, forming a detailed user profile. When the user begins a new computational session, the system analyzes the conversation in real time, identifies tasks likely related to matrix operations, and, based on the user's profile information, predicts that the user will request a matrix inversion computational task. The system then selects the server node closest to the user based on the user's location information, preloads the GLM4 model onto that node, and stores some of the model's initialization information in the user's local cache. This allows the system to quickly retrieve the model from the preloaded location when the user submits a formal matrix computation request, significantly reducing the latency associated with model loading and response times. This enables rapid computational task processing and improves user computing efficiency and satisfaction.

[0103] It's important to note that a server node is a physical or virtual server responsible for executing computing tasks. It can be a local server, an edge computing node, or a cloud server. Distance refers to the physical or network distance between the user and the server node. When selecting a server node, the intelligent scheduling system not only considers the distance between the user and the node, but also key factors such as computing resources and load.

[0104] For example, real-time monitoring of the resource status of all server nodes, such as CPU and GPU utilization and memory usage, ensures that models are preloaded onto nodes with sufficient resources. Advanced load balancing strategies are employed, taking into account the current load of each node and the resource requirements of the predicted computing task, selecting nodes with lower loads for model preloading to avoid resource contention. This implementation overcomes the limitations of selecting server nodes based solely on distance, ensuring that model execution is not only responsive but also runs in an environment with ample resources and moderate load, thereby improving computing efficiency and user satisfaction.

[0105] In order to select server nodes more flexibly, the preset thresholds can also be dynamically adjusted based on the network status, the urgency of the computing task, and resource requirements. Specifically, network bandwidth and latency are monitored in real time. For areas with ample bandwidth, the preset thresholds (including distance thresholds, computing resource thresholds, load thresholds, etc.) can be appropriately relaxed to select nodes that are slightly farther away but have better resources and load. For urgent or resource-intensive tasks, the system increases the computing resources and low load weights in the preset thresholds, giving priority to nodes that can respond quickly and have sufficient resources, even if they are slightly farther away.

[0106] In a specific embodiment, suppose the target object (user) is located in an area with medium network bandwidth and is performing a complex scientific computing task, such as large-scale matrix operations, which requires high computing resources and fast response. When selecting server nodes, the intelligent scheduling system first filters out all nodes that meet the conditions based on a preset distance threshold. The system discovers that although there are multiple nodes close to the user, these nodes are currently highly loaded and have limited computing resources. However, slightly further away from the threshold edge, there is an edge computing node with abundant computing resources and low load. At this point, the system dynamically adjusts the preset threshold based on the current network bandwidth and the resource requirements of the task, increasing the weight of computing resources and low load, and selecting this edge computing node for model preloading. Although this node is slightly farther away from the user, due to its abundant computing resources and low load, model preloading and computing task execution can be more efficient, and the user experience is actually improved.

[0107] Through the above steps S202 to S208, a model scheduling method is adopted to obtain the target request and determine its domain information, and then select the target model corresponding to the domain information, and convert the request message into the request format of the target model for execution, thereby achieving the purpose of accurately scheduling the model, thereby achieving the technical effect of improving the efficiency and accuracy of computing task processing, and thus solving the technical problem that related technologies cannot achieve accurate model scheduling in a mixed deployment scenario of multiple models when using models to perform computing tasks.

[0108] To facilitate understanding of the above process, the following is an explanation in conjunction with a specific embodiment. Taking a matrix operation request as an example, the implementation process of the system and method of the present application is as follows:

[0109] (1) Request analysis: The request analysis module detects the keyword “matrix inversion” and determines that the request belongs to the field of matrix operations by comparing it with the characteristic vocabulary of the scientific computing field.

[0110] (2) Model selection: Based on historical records, the scheduler finds that the GLM4 model has an accuracy rate of up to 92% in matrix operations, so the GLM4 model is selected first to process the request.

[0111] (3) Load monitoring and switching: If the local GLM4 model load is too high, the system will automatically switch to the Qwen-Max service of other manufacturers to ensure that the request can be processed in time and avoid response delays and calculation errors caused by excessive model load.

[0112] (4) Result Verification: After the model returns the calculation results, code verification is performed to check the dimension matching. For example, the result dimension of the matrix inversion is checked to see if it meets the expectations to ensure the accuracy of the calculation results. If the result does not meet the requirements, the system will take corresponding measures, such as reselecting the model for calculation or prompting the user to check the input parameters.

[0113] Figure 3 This is a system architecture diagram of a model scheduling method according to an embodiment of the present application. Figure 3 As shown in the figure, the local model container is used to store and manage locally available models. These models can run directly locally, reducing dependence on cloud resources and improving response speed. In addition, the system also includes models of cloud service providers (i.e., cloud first service provider, cloud second service provider, etc.). The API gateway serves as the unified interface layer of the system, responsible for communicating with models from different vendors, and converting requests into the format required by each vendor's model through the unified interface adaptation layer. The intelligent scheduler is responsible for intelligently scheduling computing requests based on the dynamic model selection algorithm, the unified interface adaptation layer, and the domain-aware load balancing strategy. The intelligent scheduler is the brain of the system. It dynamically selects the most suitable model to process computing requests based on the domain information of the task, the performance of the model, and the load situation. The system also includes a request analysis module, which is responsible for analyzing the user's computing requests, extracting key information (such as domain feature words), and determining the domain type to which the task belongs.

[0114] Figure 4 This is a flow chart of a dynamic selection algorithm for a model scheduling method according to an embodiment of the present application. Figure 4As shown in the figure, the decision tree includes three parts: QoS (Quality of Service) evaluation, cost calculation, and domain matching. QoS evaluation is used to determine the model's response delay and other quality of service indicators; cost calculation analyzes the computational cost of the model's processing tasks; and domain matching determines the degree of match between the model and the domain to which the computational task belongs. Ultimately, the optimal model is selected by combining information from these three aspects.

[0115] The model scheduling method used in this application can significantly reduce response time. Compared with a single model solution, the dynamic model selection algorithm and preloading mechanism of this application can quickly select the optimal model and prepare computing resources in advance, greatly reducing the waiting time for user requests. In addition, it can also significantly improve the accuracy of complex computing tasks. Through domain matching optimization, using domain-aware load balancing and the domain matching dimension in the multi-dimensional evaluation matrix, computing tasks can be assigned to the most suitable model for processing, significantly improving the accuracy of complex computing tasks.

[0116] Figure 5 is a structural diagram of a model scheduling device according to an embodiment of the present application, such as Figure 5 As shown, the device includes:

[0117] An acquisition module 502 is configured to acquire a target request sent by a target object, wherein the target request includes a target computing task;

[0118] Determining module 504, configured to determine domain information to which the target computing task belongs, wherein the domain information is used to indicate a computing domain type of the target computing task;

[0119] A matching module 506 is used to determine a target model corresponding to the domain information from multiple models;

[0120] The execution module 508 is configured to convert the request message of the target request into a request format corresponding to the target model, and execute the target request using the target model.

[0121] It should be noted that Figure 5 The model shown schedules the device for executing Figure 2 The model scheduling method shown is therefore Figure 2 The explanations in the model scheduling method also apply to Figure 5 The device for model scheduling shown will not be described in detail here.

[0122] An embodiment of the present application also provides an electronic device, which includes a memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the steps of the model scheduling method in each embodiment of the present application.

[0123] An embodiment of the present application also provides a non-volatile storage medium, which includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the steps of the model scheduling method in each embodiment of the present application by running the computer program.

[0124] An embodiment of the present application also provides a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the model scheduling method in each embodiment of the present application.

[0125] An embodiment of the present application also provides a computer program, which, when executed by a processor, implements the steps of the model scheduling method in each embodiment of the present application.

[0126] The serial numbers of the above-mentioned embodiments of the present application are for description only and do not represent the advantages or disadvantages of the embodiments.

[0127] In the above embodiments of the present application, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, please refer to the relevant description of other embodiments.

[0128] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only exemplary. For example, the division of the units can be a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.

[0129] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0130] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0131] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.

[0132] The above is only a preferred embodiment of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.

Claims

1. A model scheduling method, characterized in that: include: Obtaining a target request sent by a target object, wherein the target request includes a target computing task; Determining domain information to which the target computing task belongs, wherein the domain information is used to indicate a computing domain type of the target computing task; determining a target model corresponding to the domain information from a plurality of models; The request message of the target request is converted into a request format corresponding to the target model, and the target request is executed using the target model.

2. The method according to claim 1, characterized in that Determine the domain information to which the target computing task belongs, including: Obtaining description information of the target computing task from the target request, and extracting keywords corresponding to the description information; Matching the keyword with a feature word library to obtain a matching result, wherein the feature word library is used to store feature words corresponding to multiple computing fields; The domain information to which the target computing task belongs is determined according to the matching result.

3. The method according to claim 2, characterized in that The characteristic words of each sub-field of the plurality of computing fields in the characteristic word library correspond to a characteristic word subset; the keyword is matched with the characteristic word library to obtain a matching result, including: Determine the word frequency of each keyword in the characteristic word subset of the corresponding sub-field, wherein the word frequency is used to represent the number of times the keyword appears in the sub-field; Determining an inverse document frequency of each keyword in the feature lexicon, wherein the inverse document frequency is used to represent the uniqueness of the keyword with respect to the feature lexicon; The matching result is determined based on the word frequency and the inverse document frequency corresponding to each keyword.

4. The method according to claim 3, characterized in that Determining the matching result based on the word frequency and the inverse document frequency corresponding to each keyword includes: Comparing a target keyword with the feature word library to obtain a comparison result, wherein the target keyword is any one of the keywords; When the comparison result indicates that the target keyword belongs to the target sub-domain of the feature lexicon, determining a first score between the target keyword and the target sub-domain using a first weight, wherein the first weight includes a weight determined based on a word frequency and an inverse document frequency of the target keyword, and the first score is used to represent a correlation between the target keyword and the target sub-domain; A score corresponding to each sub-field in the feature vocabulary is determined according to the first score, and the score is determined as the matching result.

5. The method according to claim 4, characterized in that The method further includes: when the comparison result indicates that the target keyword does not belong to any sub-field of the feature vocabulary, using a second weight to determine a second score of the target keyword, wherein the second weight is less than the first weight, and the second score is used to represent the correlation between the target keyword and the sub-field obtained by fuzzy matching between the target keyword and the feature vocabulary.

6. The method according to claim 1, characterized in that Determining a target model corresponding to the domain information from a plurality of models includes: Obtaining historical performance data of the multiple models under the computing domain type; Determining, based on the historical performance data, matching scores corresponding to the multiple models under the computing domain type, wherein the matching scores are used to quantitatively represent the degree of matching between the target computing task and the multiple models; The target model is determined according to the matching score.

7. The method according to claim 6, characterized in that Before determining the target model according to the matching score, the method further includes: Obtaining response times corresponding to the multiple models within a preset time period, wherein the response times are used to quantitatively represent the processing speed of the multiple models in processing computing tasks; Determining computing costs corresponding to the multiple models respectively, wherein the computing costs are used to quantify resource consumption costs of processing computing tasks by the multiple models; The target model is determined according to the matching score, the response time, and the computational cost.

8. The method according to claim 1, characterized in that Converting the target request message into a request format corresponding to the target model includes: Determining interface features corresponding to the target model, wherein the interface features include request header features and request parameter features; Converting the request header in the request message according to the request header characteristics to obtain a target request header; Converting the request parameters in the request message according to the request parameter characteristics to obtain target request parameters; Determine a unified interface request corresponding to the target request header and the target request parameters.

9. The method according to claim 1, characterized in that The method further comprises: Obtaining historical conversation data of the target object; Constructing a customer profile of the target object based on the historical conversation data, wherein the customer profile is used to describe the model and computing tasks preferred by the target object; A prediction computing task for predicting the target object based on the current conversation data; The model corresponding to the prediction calculation task is preloaded to a target location, wherein the target location includes a local cache or a server node whose distance from the target object meets a preset threshold.

10. A model scheduling device, characterized in that: include: An acquisition module, configured to acquire a target request sent by a target object, wherein the target request includes a target computing task; a determination module, configured to determine domain information to which the target computing task belongs, wherein the domain information is used to indicate a computing domain type of the target computing task; a matching module, configured to determine a target model corresponding to the domain information from a plurality of models; An execution module is used to convert the request message of the target request into a request format corresponding to the target model, and execute the target request using the target model.

11. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store program instructions; the processor is connected to the memory and is used to execute the method for implementing model scheduling as described in any one of claims 1 to 9.

12. A non-volatile storage medium, characterized in that: The non-volatile storage medium includes a stored computer program, wherein the device where the non-volatile storage medium is located executes the model scheduling method according to any one of claims 1 to 9 by running the computer program.

13. A computer program product comprising computer instructions, characterized in that When the computer instructions are executed by a processor, the model scheduling method according to any one of claims 1 to 9 is implemented.

Citation Information

Cited By

  • Routing method and device for model call request, equipment, storage medium and product

    CN121478455A

  • A model calling request routing method, device, equipment, storage medium and product

    CN121478455B