Model reasoning request scheduling method and device, equipment and medium
By building a unified data index and cache hit rate management, the problem of excessive hardware resource usage during large model inference is solved, more efficient model inference request scheduling is achieved, and model inference efficiency is improved.
Patent Information
- Application Number
- CN202510838880.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-20
- Publication Date
- 2025-10-24
AI Technical Summary
In the field of large models, as the scale of model parameters expands and the context length increases, the computing power consumption of the model inference process increases significantly. The existing technology takes a long time to calculate when scheduling model inference requests between multiple model instances, occupies a lot of hardware resources, and affects efficiency.
Build a unified data index, deduplicate historical data items cached by multiple model instances, use data identifiers to manage cache hit rates, determine target model instances for request scheduling, reduce hardware resource usage, and improve efficiency.
Through unified data indexing and cache hit rate management, the hit rate of model inference requests in multiple model instances can be efficiently calculated, saving hardware resources and achieving more efficient model inference request scheduling.
Smart Images

Figure CN120832232A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to the field of large models and distributed model service technology, and specifically to a scheduling method, device, electronic device, computer-readable storage medium and computer program product for model inference requests. Background Art
[0002] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.
[0003] In the large-scale model domain, computing power consumption increases significantly with the expansion of model parameter size and context length. By having large model instances cache the intermediate encoding results of previous requests when processing inference requests, the cached intermediate encoding results of the large model instance can be reused as much as possible when receiving new requests, thus reducing the hardware resource usage of the encoding process and improving model inference efficiency.
[0004] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art. Summary of the Invention
[0005] The present disclosure provides a method, apparatus, electronic device, computer-readable storage medium, and computer program product for scheduling model inference requests.
[0006] According to an aspect of the present disclosure, a model inference request scheduling method is provided, comprising: determining at least one first data item to be encoded based on a model inference request to be scheduled; determining, for each model instance in a plurality of model instances, a cache hit rate corresponding to the model instance based on a first data index and the at least one first data item, wherein each model instance caches a plurality of historical data items and encoding results of the plurality of historical data items, and the cache hit rate corresponding to each model instance indicates a proportion of an intersection of the plurality of historical data items cached by the model instance and the at least one first data item in the at least one first data item, and wherein the first data index comprises: a plurality of second data items, wherein the plurality of second data items are determined based on a deduplication processing performed on a plurality of historical data items respectively cached by the plurality of model instances; and a data identifier corresponding to each second data item in the plurality of second data items, wherein the data identifier corresponding to each second data item comprises a plurality of sub-identifiers respectively corresponding to the plurality of model instances, and each sub-identifier indicates whether the second data item is cached by the model instance corresponding to the sub-identifier; determining a target model instance from the plurality of model instances based on the cache hit rate corresponding to each model instance; and scheduling the model inference request to the target model instance to perform inference.
[0007] According to another aspect of the present disclosure, a model inference request scheduling apparatus is provided, comprising: a first determining unit configured to determine at least one first data item to be encoded based on a model inference request to be scheduled;
[0008] a second determining unit configured to determine, for each model instance in a plurality of model instances, a cache hit rate corresponding to the model instance based on a first data index and the at least one first data item, wherein each model instance caches a plurality of historical data items and encoding results of the plurality of historical data items, and the cache hit rate corresponding to each model instance indicates a proportion of an intersection of the plurality of historical data items cached by the model instance and the at least one first data item in the at least one first data item, and wherein the first data index comprises: a plurality of second data items, wherein the plurality of second data items are determined based on a deduplication processing performed on a plurality of historical data items respectively cached by the plurality of model instances; and a data identifier corresponding to each second data item in the plurality of second data items, wherein the data identifier corresponding to each second data item comprises a plurality of sub-identifiers respectively corresponding to the plurality of model instances, and each sub-identifier indicates whether the second data item is cached by the model instance corresponding to the sub-identifier; a third determining unit configured to determine a target model instance from the plurality of model instances based on the cache hit rate corresponding to each model instance; and an inference unit configured to schedule the model inference request to the target model instance to perform inference.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above-mentioned model inference request scheduling method.
[0010] According to another aspect of the present disclosure, a non-transitory computer readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to perform the above-mentioned model inference request scheduling method.
[0011] According to another aspect of the present disclosure, a computer program product is provided, comprising a computer program, wherein the computer program, when executed by a processor, can implement the above-mentioned model inference request scheduling method.
[0012] According to one or more embodiments of the present disclosure, the cache hit rate of the request to be processed in the plurality of model instances can be calculated more efficiently, hardware resources are saved, and more efficient scheduling of the model inference request is achieved.
[0013] It should be understood that the content described in this section is not intended to identify key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become apparent through the following description. BRIEF DESCRIPTION OF DRAWINGS
[0014] The accompanying drawings, which are included to provide a further understanding of the application and are incorporated in and constitute a part of this specification, illustrate embodiments and together with the description serve to explain exemplary implementations of the application. The illustrated embodiments are merely examples and do not limit the scope of the claims. In all the drawings, like reference numerals refer to like elements throughout the accompanying drawings.
[0015] Figure 1 A schematic diagram of an exemplary system in which the various methods described herein can be implemented according to an exemplary embodiment of the present disclosure is shown;
[0016] Figure 2 A flowchart of a model inference request scheduling method according to an exemplary embodiment of the present disclosure is shown;
[0017] Figure 3 A schematic diagram of a model inference process according to an exemplary embodiment of the present disclosure is shown;
[0018] Figure 4 A block diagram of a model inference request scheduling apparatus according to an exemplary embodiment of the present disclosure is shown;
[0019] Figure 5A structural block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0020] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, in which various details of embodiments of the present disclosure are set forth in order to provide an understanding of the present disclosure. It will be apparent to those of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure. Also, the description below is merely illustrative of the principles of the present disclosure, and does not limit the scope of the present disclosure. Thus, it will be apparent to those of ordinary skill in the art that various changes and modifications can be made to the embodiments described herein without departing from the scope of the present disclosure.
[0021] In the present disclosure, the terms "first", "second", and the like are used to describe various elements only for the purpose of distinguishing one element from another, and do not otherwise limit the position, sequence, or importance of the elements. In some examples, a first element and a second element can refer to the same instance of the element, and in some cases, they can refer to different instances of the element based on the context of the description.
[0022] The terms used in the description of various described examples in the present disclosure are only for the purpose of describing particular examples and are not intended to be limiting. Unless specifically defined otherwise, an element that is a singular can be plural and vice versa. Also, the term "and / or" used in the present disclosure encompasses any and all possible combinations of one or more of the associated listed items.
[0023] Embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.
[0024] Figure 1 A schematic diagram of an example system 100 in which various methods and apparatuses described herein can be implemented according to embodiments of the present disclosure is shown. Referring to Figure 1 , the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 that couple the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more application programs.
[0025] In embodiments of the present disclosure, the server 120 can run one or more services or software applications that enable the execution of a method of scheduling model inference requests.
[0026] In certain embodiments, the server 120 can also provide other services or software applications, which can include non-virtual and virtual environments. In certain embodiments, these services can be provided as web-based services or cloud services, for example, to users of the client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.
[0027] In Figure 1 In the illustrated configuration, the server 120 can include one or more components that implement the functionality performed by the server 120. These components can include software components that are executable by one or more processors, hardware components, or combinations thereof. Users operating the client devices 101, 102, 103, 104, 105, and / or 106 can in turn utilize one or more client applications to interact with the server 120 to utilize the services provided by the components. It should be understood that a wide variety of system configurations are possible, which can differ from system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.
[0028] A user can use a client device 101, 102, 103, 104, 105, and / or 106 to send a model inference request. The client device can provide an interface that enables a user of the client device to interact with the client device. The client device can also output information to the user via the interface. Although Figure 1 Only six client devices are depicted, but one of skill in the art will understand that the present disclosure can support any number of client devices.
[0029] Client devices 101, 102, 103, 104, 105, and / or 106 can include various categories of computer devices, such as portable handheld devices, general purpose computers (such as personal computers and laptop computers), workstation computers, wearable devices, smart screen devices, self-service kiosk devices, service robots, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, and the like. These computer devices can run various categories and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-like operating systems (such as GOOGLE Chrome OS); or include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, Android. Portable handheld devices can include cellular phones, smartphones, tablet computers, personal digital assistants (PDAs), and the like. Wearable devices can include head-mounted displays (such as smart glasses) and other devices. Gaming systems can include various handheld gaming devices, Internet-enabled gaming devices, and the like. Client devices are capable of executing a variety of different applications, such as various Internet-related applications, communication applications (such as email applications), short message service (SMS) applications, and can use various communication protocols.
[0030] Network 110 can be any category of network in the art, including, but not limited to, a LAN, an Ethernet-based network, Token Ring, a WAN, the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, WIFI), and / or any combination of these and / or other networks.
[0031] Server 120 can include one or more general purpose computers, special purpose server computers (e.g., PC (personal computer) servers, UNIX servers, midrange servers), blade servers, mainframe computers, server clusters, or any other appropriate arrangement and / or combination. Server 120 can include one or more virtual machines running a virtual operating system, or other computing architectures involving virtualization (such as one or more flexible pools of logical storage devices that can be virtualized to maintain virtual storage devices for servers). In various embodiments, server 120 can run one or more services or software applications that provide the functionality described below.
[0032] The computing units in the server 120 can run one or more operating systems including any of the operating systems described above, as well as any commercially available server operating systems. Server 120 can also run any of a variety of additional server applications and / or mid-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, etc.
[0033] In some embodiments, the server 120 can include one or more applications to analyze and consolidate data feeds and / or event updates from users of the client devices 101, 102, 103, 104, 105, and 106. The server 120 can also include one or more applications to display the data feeds and / or real-time events via one or more display devices of the client devices 101, 102, 103, 104, 105, and 106.
[0034] In some embodiments, the server 120 can be a server of a distributed system, or a server in combination with a blockchain. The server 120 can also be a cloud server, or an intelligent cloud computing server or intelligent cloud host with artificial intelligence technology. The cloud server is a host product in the cloud computing service system, to solve the defects of large management difficulty and weak business scalability in traditional physical host and virtual private server (VPS, Virtual Private Server) services.
[0035] The system 100 can also include one or more databases 130. In certain embodiments, these databases can be used to store data and other information. For example, one or more of the databases 130 can be used to store information such as audio files and video files. The databases 130 can reside in a variety of locations. For example, databases used by the server 120 can reside locally to the server 120, or can be remote from the server 120 and can communicate with the server 120 via a network- or application-specific connection. The databases 130 can be of different categories. In certain embodiments, databases used by the server 120 can be, for example, relational databases. One or more of these databases can store, update, and retrieve data to and from the databases in response to commands.
[0036] In certain embodiments, one or more of the databases 130 can also be used by applications to store application data. Databases used by applications can be databases of different categories, such as key-value stores, object stores, or regular stores backed by file systems.
[0037] Figure 1The system 100 can be configured and operated in various ways to enable the application of various methods and apparatuses described in accordance with the present disclosure.
[0038] In the field of large models, as the scale of model parameters and the length of context increase, the consumption of computing power in the model inference process increases significantly. In the large model inference process, it is usually necessary to first encode the input information and then obtain the final inference result based on the encoding result. For example, in the autoregressive inference of the Transformer model, it is necessary to convert the original input information into an initial encoding result to obtain a plurality of minimum data units, i.e., a plurality of input tokens, which can be processed by the model, and then perform encoding calculation based on the attention mechanism based on the plurality of input tokens to obtain the key vector (Key) and the value vector (Value) corresponding to the input token. This process corresponds to the prefill process of large model inference. Since Key and Value are necessary intermediate results in attention calculation, by caching the Key and Value corresponding to the input token, the cached Key / Value cache result can be reused in the process of processing a large amount of input information by the model, avoiding repeated calculation, saving hardware resources and improving model inference efficiency. Taking the scene of realizing multi-round human-computer dialogue based on a language large model as an example, as the number of dialogue rounds increases, the proportion of repeated calculation of historical tokens continues to grow, directly leading to redundant consumption of computing power.
[0039] In actual application scenarios, in order to improve the model inference efficiency and support large-scale concurrent requests, multiple model instances are usually deployed to process the model inference requests of users. In this case, each model instance caches the intermediate encoding result of the inference request processed by the model instance itself in the inference process. In related technologies, the cache information index of each model instance is usually maintained separately for different model instances. This way occupies a lot of hardware resources and takes a long time to calculate when scheduling the model inference requests to be processed based on the cache information of the model instances, thereby affecting the model inference efficiency.
[0040] Based on this, the present disclosure provides a model inference request scheduling method, which constructs a unified data index for the encoding results of historical data items cached by multiple model instances, removes duplicate data items and manages them using data identifiers, which is conducive to more efficiently calculating the cache hit rate of each model instance, determining the target instance for request scheduling based on this, and saving hardware resources to more efficiently and accurately schedule model inference requests among multiple model instances.
[0041] Figure 2 A flowchart of a model inference request scheduling method 200 according to an example embodiment of the present disclosure is shown. As shown in Figure 2 The method 200 includes:
[0042] Step S201, determining at least one first data item to be encoded based on a model inference request to be scheduled;
[0043] Step S202, for each model instance in the plurality of model instances, determining a cache hit rate corresponding to the model instance based on the first data index and the at least one first data item, wherein each model instance caches a plurality of historical data items and respective encoding results of the plurality of historical data items, the cache hit rate corresponding to each model instance indicates a proportion of an intersection of the plurality of historical data items cached by the model instance and the at least one first data item in the at least one first data item, and wherein the first data index comprises: a plurality of second data items, wherein the plurality of second data items are determined based on performing deduplication processing on the plurality of historical data items respectively cached by the plurality of model instances; and a data identifier corresponding to each second data item in the plurality of second data items, wherein the data identifier corresponding to each second data item comprises a plurality of sub-identifiers respectively corresponding to the plurality of model instances, each sub-identifier indicating whether the second data item is cached by the model instance corresponding to the sub-identifier;
[0044] Step S203, determining a target model instance from the plurality of model instances based on the cache hit rate corresponding to each model instance; and
[0045] Step S204, scheduling the model inference request to the target model instance to perform inference.
[0046] By applying the first data index described in step S202 in the management process of the plurality of model instances, the cache index information of each model instance can be maintained, a unified data index can be constructed for the historical data items and their encoding results cached by the plurality of model instances, repeated data items can be deduplicated and managed by using data identifiers, and storage resources can be saved. Further, by applying the above method 200 to schedule the model inference request, the cache hit rate of the model inference request to be processed in each model instance can be more efficiently calculated, and the target model instance corresponding to the request to be processed can be determined based thereon, thereby improving the model inference efficiency.
[0047] It is understandable that since the intermediate encoding results of historical requests cached by each model instance are different, when a new request to be processed is received, the cache hit rate corresponding to each model instance can be determined by sensing the historical data items of each model instance maintained in the first data index. Specifically, the cache hit rate corresponding to each model instance can be determined based on the repetition rate calculation of multiple second data items and at least one first data item to be encoded, and then based on the data identifier corresponding to the hit result of at least one first data item in multiple second data items. There is no need to calculate the repetition rate of data hits for each model instance cache separately, thereby reducing the occupation of hardware resources and improving request scheduling efficiency.
[0048] In some examples, the model inference request may be in various forms, such as a text processing request including natural language text, or a model inference request including data in various modalities such as pictures and voices. As described above, when using a model to process based on input data, it is necessary to first convert the input data into an initial encoding result that the model can understand and process. In this case, at least one first data item may be an initial encoding result of different forms corresponding to different types of input data, such as a text semantic embedding representation corresponding to a text segmentation result (token). As long as the intermediate results of the model inference process can be cached by maintaining the data items and the corresponding encoding results of the data items, the present disclosure does not limit the specific form of the first data items and the historical data items.
[0049] According to some embodiments, the multiple historical data items cached by each model instance are determined based on text characters, the first data index is an index tree composed of the multiple second data items, and the text character sequence corresponding to the path from the root node to the leaf node in the index tree is determined based on the model inference request historically processed by the multiple model instances. When method 200 is applied to the inference service of a large language model, it is possible to perform word segmentation on the model inference request in the form of a natural language and construct a cache index based on the word segmentation result and the encoding result corresponding to the word segmentation result. By utilizing the nodes of the tree structure to store text characters (word segmentation results), that is, the path nodes in the tree can be used to indicate the text character sequence contained in the model inference request, and the cache hit rate can be calculated more efficiently by querying the path nodes in the tree structure starting from the root node.
[0050] In one example, when the historical requests processed by the model instance include a text sequence A-B-C-D, A can be taken as a root node of the cache index tree, and then B, C, and D are sequentially stored in respective nodes based on the parent-child connection relationship. In this case, when a new request to be processed includes a text sequence A-B-C-E, the branch path in the index tree can be queried based on A as the root node, and then A, B, and C nodes that can be hit in the index tree for the new request to be processed can be efficiently and conveniently determined. In one example, the information stored in each node can be a word or a text in the form of text, or a text semantic embedding representation. When the historical request including the text sequence A-B-C-D is processed by the model instance 01, the data identifiers of the four nodes corresponding to A, B, C, and D include a sub-identifier indicating the model instance 01, and the cache hit rate of different model instances can be determined based on the data identifiers of the nodes in the scheduling process of the model inference request, so as to achieve efficient and accurate request scheduling.
[0051] In some examples, the model instance with the highest cache hit rate in step S203 can be taken as a target model instance, so as to fully utilize the intermediate encoding results cached by the model instance to accelerate the execution of the model inference request and improve the utilization of hardware resources.
[0052] According to some embodiments, the method 200 further includes: dividing the plurality of model instances into a plurality of model instance groups; for each model instance group in the plurality of model instance groups, determining at least one third data item cached by at least one model instance in the model instance group from the plurality of second data items based on the data identifiers corresponding to each second data item in the first data index; and constructing a second data index corresponding to the model instance group based on the at least one third data item and the data identifiers corresponding to each third data item. In step S202, for each model instance in the plurality of model instances, determining the cache hit rate corresponding to the model instance based on the first data index and the at least one first data item includes: for each model instance group in the plurality of model instance groups, using a cache awareness unit corresponding to the model instance group, determining the cache hit rate corresponding to each model instance in the model instance group based on the second data index corresponding to the model instance group and the at least one first data item.
[0053] Thus, the cache-aware routing service (i.e., a service for querying cache information of model instances and calculating the hit rate of the current request in the cache of each model instance) can be horizontally sharded, and the multiple model instances can be divided into different shards (i.e., divided into multiple model instance groups) for separate management, thereby improving the cache-aware efficiency and stability. By querying the identifier of the data item (i.e., the node of the cache information index of the model instance) to extract the instance index information during the construction of the shard of the cache-aware routing service (configuring multiple cache-aware units), the cache index can be easily and efficiently sharded, and the distributed deployment of the cache-aware routing service can be supported.
[0054] In some examples, the number of model instance groups and the specific rules for dividing the multiple model instances into multiple groups can be determined according to actual needs. For example, the multiple model instances can be evenly grouped, or the multiple model instances can be grouped based on the request throughput or request processing capacity of each model instance, so as to achieve load balancing among the multiple model instance groups (i.e., multiple cache-aware units).
[0055] According to some embodiments, the cache-aware unit corresponding to each model instance group includes a primary service module and a backup service module, and the second data index corresponding to each model instance group is stored in the primary service module of the cache-aware unit corresponding to the model instance group. The method 200 further includes: backing up the second data index corresponding to each model instance group to the backup service module of the cache-aware unit corresponding to the model instance group; and in response to determining that there is a faulty primary service module in the cache-aware unit, determining the cache hit rate of each model instance in the model instance group corresponding to the faulty primary service module by using the backup service module corresponding to the faulty primary service module. Thus, a master-slave backup mechanism can be provided for each cache-aware unit to ensure service stability.
[0056] In some examples, the second data index of the model instance group maintained by the primary service module can be asynchronously updated to the backup service module, so as to ensure that the backup service module can accurately determine the cache hit rate of each model instance corresponding to the request to be processed, and accurate request scheduling can be implemented based on this. By first dividing the multiple model instances into multiple groups, and then maintaining the cache index and performing backup of the cache-aware routing service in units of model instance groups, the data traffic pressure of the master-slave asynchronous backup can be reduced, the backup response speed in the event of a fault can be improved, and thus the robustness of the cache-aware routing service can be improved.
[0057] According to some embodiments, the method 200 further includes: in response to receiving an instruction to add a cache-aware unit, wherein the instruction indicates at least one first model instance to be divided into a new model instance group of the new cache-aware unit, determining at least one fourth data item cached by the at least one first model instance based on a data identifier corresponding to each second data item in a second data index corresponding to each model instance group to which the at least one first model instance belongs; and constructing the second data index corresponding to the new model instance group based on the at least one fourth data item and the data identifier corresponding to each fourth data item. In this way, horizontal expansion of the cache-aware unit can be implemented during operation of the model inference service to adapt to changes in the load demand of the model inference service and ensure service stability.
[0058] In some examples, when the cache-aware routing service is divided into multiple cache-aware units and distributedly deployed, load monitoring and dynamic expansion and contraction management of the multiple cache-aware units are implemented by an upper host. When a new cache-aware unit is configured in the cache-aware routing service and it is determined that at least one first model instance needs to be divided into the new cache-aware unit for management, the data items corresponding to each first model instance can be extracted based on the data identifiers corresponding to each first model instance in the second data index maintained by the original cache-aware unit to initialize the cache index information that needs to be managed and maintained by the new cache-aware unit, thereby achieving convenient expansion operation.
[0059] According to some embodiments, the method 200 further includes: determining current occupation information of each model instance in the plurality of model instances, wherein determining, in step S203, the target model instance from the plurality of model instances based on the cache hit rate corresponding to each model instance includes: determining the target model instance from the plurality of model instances based on the cache hit rate corresponding to each model instance and the current occupation information of each model instance. By further combining the current occupation information of the model instance to determine the target model instance for processing the model inference request, the accuracy of request scheduling can be further improved, so that the request to be processed can be processed more efficiently, load balancing of the plurality of model instances is achieved, and model inference efficiency is improved.
[0060] According to some embodiments, the current occupancy information comprises a load rate, and determining the target model instance from the plurality of model instances based on the cache hit rate corresponding to each model instance and the current occupancy information of each model instance comprises: for each model instance in the plurality of model instances, determining a score of the model instance based on a weighted sum of the opposite number of the cache hit rate corresponding to the model instance and the load rate of the model instance; and determining the target model instance based on the score of each model instance. In this way, the cache hit rate and the load rate of the model instance can be linearly weighted to obtain the score of each model instance, and the request scheduling can be implemented more simply and accurately based on this, and the model inference efficiency can be improved.
[0061] In some examples, the calculation weights of the cache hit rate corresponding to each model instance and the load rate of each model instance can be flexibly adjusted according to requirements, so that the scheduling result of the request to be processed is more in line with the requirements of the actual application scenario.
[0062] According to some embodiments, the method 200 further comprises: determining the description information of each model instance in the plurality of model instances, the description information of each model instance indicating the request processing capability of the model instance, wherein the target model instance is determined from the plurality of model instances based on the cache hit rate corresponding to each model instance and the current occupancy information of each model instance comprises: the target model instance is determined from the plurality of model instances based on the cache hit rate corresponding to each model instance, the current occupancy information of each model instance and the description information of each model instance. By further combining the request processing capability of the model instance to determine the target model instance of the model inference request to be processed, the accuracy of the request scheduling can be further improved, so that the request to be processed can be processed more efficiently, the load balancing of the plurality of model instances can be achieved, and the model inference efficiency can be improved.
[0063] Figure 3 A schematic diagram of a model inference process according to an example embodiment of the present disclosure is shown. Referring to FIG. 1, the model inference process comprises the following steps. Figure 3In this example, the gateway 310 is used to connect the receiving port of the model inference request and the plurality of model instances capable of actually processing the model inference request, so as to schedule the model inference request sent by the user to the model instance for processing. After the gateway 310 receives the model inference request, a request for querying the target model instance can be sent to the routing query agent service 320, and the routing query agent service 320 can send a query request to the plurality of shards of the cache-aware routing service (for example, the cache-aware unit 331 and the cache-aware unit 332). On this basis, the cache-aware unit 331 can calculate the cache hit rate corresponding to the model instance 351 and the model instance 352 based on the cache index information of the model instance 351 and the model instance 352 managed by the unit, and at the same time, the cache-aware unit 332 can calculate the cache hit rate corresponding to the model instance 353 and the model instance 354 based on the cache index information of the model instance 353 and the model instance 354 managed by the unit. Each cache-aware unit returns an instance score of each model instance based on the cache hit rate corresponding to each model instance, so that the routing query agent service 320 can determine the target model instance based on the instance score of each model instance and return a query result to the gateway 310 of the upper layer, and the gateway 310 can schedule the model inference request to the target model instance for inference based on the query result. In this example, each cache-aware unit does not directly perceive the information of the plurality of model instances at the bottom layer. In order to manage the mapping relationship between the cache-aware unit and the model instance (that is, to allocate the plurality of model instances to the plurality of cache-aware units for management), the instance registration agent service 340 is used to register the information of each model instance and maintain the mapping relationship between the cache-aware unit and the model instance, so that each cache-aware unit can obtain the information of the model instance at the bottom layer managed by the cache-aware unit and maintain the cache index of the model instance. By applying the above method to schedule the model inference request, the cache hit rate of the model inference request to be processed in each model instance can be calculated more efficiently, and the target model instance corresponding to the request to be processed is determined based on this, thereby improving the model inference efficiency.
[0064] According to an aspect of the present disclosure, a model inference request scheduling apparatus is also provided. Figure 4 A structural block diagram of a model inference request scheduling apparatus 400 according to an exemplary embodiment of the present disclosure is shown. As shown, the apparatus 400 includes: Figure 4
[0065] A first determining unit 401 configured to determine at least one first data item to be encoded based on a model inference request to be scheduled;
[0066] The second determining unit 402 is configured to determine, for each model instance in the plurality of model instances, a cache hit rate corresponding to the model instance based on the first data index and the at least one first data item, wherein each model instance caches a plurality of historical data items and respective encoding results of the plurality of historical data items, and the cache hit rate corresponding to each model instance indicates a proportion of an intersection of the plurality of historical data items cached by the model instance and the at least one first data item in the at least one first data item, and wherein the first data index includes: a plurality of second data items, wherein the plurality of second data items are determined based on a plurality of historical data items respectively cached by the plurality of model instances by performing deduplication processing; and a data identifier corresponding to each second data item in the plurality of second data items, wherein the data identifier corresponding to each second data item includes a plurality of sub-identifiers respectively corresponding to the plurality of model instances, and each sub-identifier indicates whether the second data item is cached by the model instance corresponding to the sub-identifier.
[0067] The third determining unit 403 is configured to determine, based on the cache hit rate corresponding to each model instance, a target model instance from the plurality of model instances.
[0068] The reasoning unit 404 is configured to schedule the model reasoning request to the target model instance to perform reasoning.
[0069] According to some embodiments, the apparatus 400 further includes: a dividing unit configured to divide the plurality of model instances into a plurality of model instance groups; a fourth determining unit configured to determine, for each model instance group in the plurality of model instance groups, at least one third data item cached by at least one model instance in the model instance group from the plurality of second data items based on the data identifier corresponding to each second data item in the first data index; and a first constructing unit configured to construct a second data index corresponding to the model instance group based on the at least one third data item and the data identifier corresponding to each third data item, wherein the second determining unit 402 is configured to: for each model instance group in the plurality of model instance groups, utilize a cache awareness unit corresponding to the model instance group to determine, based on the second data index corresponding to the model instance group and the at least one first data item, a cache hit rate corresponding to each model instance in the model instance group.
[0070] According to some embodiments, the cache-aware unit corresponding to each model instance group comprises a primary service module and a backup service module, the second data index corresponding to each model instance group is stored in the primary service module of the cache-aware unit corresponding to the model instance group, and the apparatus 400 further comprises a backup unit configured to backup the second data index corresponding to each model instance group to the backup service module of the cache-aware unit corresponding to the model instance group, and the second determining unit 402 is further configured to, in response to determining that there is a faulty primary service module in the cache-aware unit, determine the cache hit rate corresponding to each model instance in the model instance group corresponding to the faulty primary service module by using the backup service module corresponding to the faulty primary service module.
[0071] According to some embodiments, the apparatus 400 further comprises a fifth determining unit configured to, in response to receiving an instruction to increase the cache-aware unit, determine at least one fourth data item cached by at least one first model instance to be included in a new model instance group corresponding to a new cache-aware unit based on the data identifier corresponding to each second data item in the second data index corresponding to at least one model instance group to which the at least one first model instance belongs, and a second constructing unit configured to construct the second data index corresponding to the new model instance group based on the at least one fourth data item and the data identifier corresponding to each fourth data item.
[0072] According to some embodiments, the apparatus 400 further comprises a sixth determining unit configured to determine current occupancy information of each model instance in the plurality of model instances, and the third determining unit 403 is configured to determine the target model instance from the plurality of model instances based on the cache hit rate corresponding to each model instance and the current occupancy information of each model instance.
[0073] According to some embodiments, the current occupancy information comprises a load rate, and the third determining unit 403 comprises a first determining sub-unit configured to, for each model instance in the plurality of model instances, determine a score of the model instance based on a weighted sum of the cache hit rate corresponding to the model instance and the inverse of the load rate of the model instance, and a second determining sub-unit configured to determine the target model instance based on the score of each model instance.
[0074] According to some embodiments, the apparatus 400 further comprises a seventh determining unit configured to determine description information of each model instance in the plurality of model instances, the description information of each model instance indicating a request processing capability of the model instance, and the third determining unit 403 is configured to determine the target model instance from the plurality of model instances based on the cache hit rate corresponding to each model instance, the current occupancy information of each model instance, and the description information of each model instance.
[0075] According to some embodiments, the multiple historical data items cached by each model instance are determined based on text characters, the first data index is an index tree composed of the multiple second data items, and the text character sequence corresponding to the path from the root node to the leaf node in the index tree is determined based on the model inference requests historically processed by the multiple model instances.
[0076] According to another aspect of the present disclosure, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above-mentioned scheduling method for model inference requests.
[0077] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is further provided, wherein the computer instructions are used to enable the computer to execute the above-mentioned scheduling method for model inference requests.
[0078] According to another aspect of the present disclosure, a computer program product is further provided, comprising a computer program, wherein the computer program implements the above-mentioned method for scheduling model inference requests when executed by a processor.
[0079] refer to Figure 5 , a block diagram of an electronic device 500 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0080] like Figure 5 As shown, the device 500 includes a computing unit 501, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 502 or a computer program loaded from a storage unit 508 into a random access memory (RAM) 503. Various programs and data required for the operation of the device 500 can also be stored in the RAM 503. The computing unit 501, the ROM 502, and the RAM 503 are connected to each other via a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0081] A plurality of components in the device 500 are connected to the I / O interface 505, including: an input unit 506, an output unit 507, a storage unit 508, and a communication unit 509. The input unit 506 can be any kind of device capable of inputting information to the device 500, which can receive inputted digital or character information, and generate key signal inputs related to user settings and / or function controls of the electronic device, and can include, but is not limited to, a mouse, a keyboard, a touch screen, a track pad, a track ball, a joystick, a microphone, and / or a remote controller. The output unit 507 can be any kind of device capable of presenting information, and can include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. The storage unit 508 can include, but is not limited to, a magnetic disk, an optical disk. The communication unit 509 allows the device 500 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks, and can include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver, and / or a chipset, such as a Bluetooth™ device, an 802.11 device, a WiFi device, a WiMax device, a cellular communication device, and / or the like.
[0082] The computing unit 501 can be various general and / or special purpose processing components with processing and computing capabilities. Some examples of the computing unit 501 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 501 performs various methods and processes described above, such as the model inference request scheduling method. For example, in some embodiments, the model inference request scheduling method can be implemented as a computer software program, which is tangibly embodied in a machine-readable medium, such as the storage unit 508. In some embodiments, part or all of the computer program can be loaded and / or installed onto the device 500 via the ROM 502 and / or the communication unit 509. When the computer program is loaded onto the RAM 503 and executed by the computing unit 501, one or more steps of the model inference request scheduling method described above can be performed. Alternatively, in other embodiments, the computing unit 501 can be configured to perform the model inference request scheduling method by any other appropriate means, such as by means of firmware.
[0083] The various embodiments of the systems and techniques described above can be implemented in digital electronic circuitry, integrated circuitry, a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a system on a chip (SOC), a complex programmable logic device (CPLD), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include implementation in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be special or general purpose, coupled to receive data and instructions from, and to transmit data and instructions to, a storage system, at least one input device, and at least one output device.
[0084] Program code for carrying out methods of the present disclosure can be written in any combination of one or more programming languages. The program code can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the program code, when executed by the processor or controller, produces the functions / operations specified in the flowcharts and / or the block diagrams. The program code can be executed entirely on a machine, partially on a machine, partially on a machine as a stand-alone software package, or entirely on a remote machine or server.
[0085] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical conductors, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0086] To provide for interaction with a user, the systems and techniques described here can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can be used to provide for interaction with a user as well; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form, including acoustic, speech, or tactile input.
[0087] The systems and techniques described here can be implemented in a computing system that includes a back end component (e.g., as a data server), or that includes a middleware component (e.g., an application server), or that includes a front end component (e.g., a user computer having a graphical user interface or a Web browser through which a user can interact with an implementation of the systems and techniques described here), or any combination of such back end, middleware, or front end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0088] The computer system can include clients and servers. A client and server are generally remote from each other and typically interact through a communication network. The relationship of client and server is generally established by computer programs running on the respective computers and having a client-server relationship to each other. The servers can be cloud servers, servers of a distributed system, or servers combined with a blockchain.
[0089] It should be understood that the various forms of flow shown above can be re-ordered, added to, or deleted from without departing from the scope of the present disclosure. For example, the steps recited in the present disclosure can be performed in parallel, in series, or in a different order, as long as the desired results of the technology disclosed in the present disclosure are achieved, which is not limited herein.
[0090] While embodiments or examples of this disclosure have been described with reference to the figures, it will be understood that the methods, systems, and devices described above are merely exemplary embodiments or examples, and the scope of the application is not limited to these embodiments or examples. Various elements of the embodiments or examples can be omitted or substituted by equivalents thereof. Furthermore, the steps can be performed in a different order than described in the disclosure. Further, various elements of the embodiments or examples can be combined in various ways. It is important that as technology evolves, many of the elements described herein can be substituted by equivalents which serve the same function.
Claims
1. A method for scheduling a model inference request, comprising: determining at least one first data item to be encoded based on the model inference request to be scheduled; determining, for each model instance in a plurality of model instances, a cache hit rate corresponding to the model instance based on a first data index and the at least one first data item, wherein each model instance caches a plurality of historical data items and encoding results of the plurality of historical data items respectively, and the cache hit rate corresponding to each model instance indicates a proportion of an intersection of the plurality of historical data items cached by the model instance and the at least one first data item in the at least one first data item, and wherein the first data index comprises: a plurality of second data items, wherein the plurality of second data items are determined based on a deduplication processing performed on the plurality of historical data items cached by the plurality of model instances respectively; and a data identifier corresponding to each second data item in the plurality of second data items, wherein the data identifier corresponding to each second data item comprises a plurality of sub-identifiers corresponding to the plurality of model instances respectively, and each sub-identifier indicates whether the second data item is cached by the model instance corresponding to the sub-identifier; determining, from the plurality of model instances, a target model instance based on the cache hit rate corresponding to each model instance; and scheduling the model inference request to the target model instance to perform inference.
2. The method of claim 1, further comprising: dividing the plurality of model instances into a plurality of model instance groups; for each model instance group in the plurality of model instance groups, determining, from the plurality of second data items, at least one third data item cached by at least one model instance in the model instance group based on the data identifier corresponding to each second data item in the first data index; and constructing a second data index corresponding to the model instance group based on the at least one third data item and the data identifier corresponding to each third data item, wherein the determining, for each model instance in the plurality of model instances, the cache hit rate corresponding to the model instance based on the first data index and the at least one first data item comprises: for each model instance group in the plurality of model instance groups, determining, by a cache awareness unit corresponding to the model instance group, the cache hit rate corresponding to each model instance in the model instance group based on the second data index corresponding to the model instance group and the at least one first data item. The cache awareness unit corresponding to each model instance group comprises a primary service module and a backup service module, and the second data index corresponding to each model instance group is stored in the primary service module of the cache awareness unit corresponding to the model instance group, and the method further comprises: backing up the second data index corresponding to each model instance group to the backup service module of the cache awareness unit corresponding to the model instance group; and in response to determining that there is a faulty primary service module in the cache awareness unit, determining, by the backup service module corresponding to the faulty primary service module, the cache hit rate corresponding to each model instance in the model instance group corresponding to the faulty primary service module.
4. The method of claim 2 or 3, further comprising: 3. The method of claim 2, wherein, In response to receiving an instruction to add a cache-aware unit, wherein the instruction indicates at least one first model instance to be assigned to a new model instance group corresponding to the new cache-aware unit, determining at least one fourth data item cached by the at least one first model instance based on a data identifier corresponding to each second data item in a second data index corresponding to each of the at least one model instance groups to which the at least one first model instance belongs; and Based on the at least one fourth data item and the data identifier corresponding to each fourth data item, a second data index corresponding to the new model instance group is constructed.
5. The method according to any one of claims 1 to 4, further comprising: determining current occupancy information for each of the plurality of model instances, The determining of the target model instance from the multiple model instances based on the cache hit rate corresponding to each model instance includes: A target model instance is determined from the multiple model instances based on a cache hit rate corresponding to each model instance and current occupancy information of each model instance.
6. The method of claim 5, wherein, The current occupancy information includes a load rate, and determining a target model instance from the multiple model instances based on a cache hit rate corresponding to each model instance and the current occupancy information of each model instance includes: For each model instance in the plurality of model instances, determining a score for the model instance based on a weighted sum of a cache hit rate corresponding to the model instance and an inverse of a load rate of the model instance; and The target model instance is determined based on the score of each model instance.
7. The method according to claim 5 or 6, further comprising: determining descriptive information for each model instance in the plurality of model instances, the descriptive information for each model instance indicating a request processing capability of the model instance; The determining of the target model instance from the multiple model instances based on the cache hit rate corresponding to each model instance and the current occupancy information of each model instance includes: A target model instance is determined from the multiple model instances based on a cache hit rate corresponding to each model instance, current occupancy information of each model instance, and description information of each model instance.
8. The method of any one of claims 1-7, wherein, The multiple historical data items cached by each model instance are determined based on text characters, the first data index is an index tree composed of the multiple second data items, and the text character sequence corresponding to the path from the root node to the leaf node in the index tree is determined based on the model inference requests of the historical processing of the multiple model instances.
9. A scheduling device for model inference requests, comprising: A first determining unit is configured to determine at least one first data item to be encoded based on the model inference request to be scheduled; a second determining unit, configured to determine, for each model instance in the plurality of model instances, a cache hit rate corresponding to the model instance based on the first data index and the at least one first data item, wherein each model instance caches a plurality of historical data items and respective encoding results of the plurality of historical data items, and the cache hit rate corresponding to each model instance indicates a proportion of an intersection of the plurality of historical data items cached by the model instance and the at least one first data item in the at least one first data item, and wherein the first data index comprises: a plurality of second data items, wherein the plurality of second data items are determined based on a deduplication processing performed on the plurality of historical data items respectively cached by the plurality of model instances; and a data identifier corresponding to each second data item in the plurality of second data items, wherein the data identifier corresponding to each second data item comprises a plurality of sub-identifiers respectively corresponding to the plurality of model instances, and each sub-identifier indicates whether the model instance corresponding to the sub-identifier caches the second data item; a third determining unit, configured to determine, based on the cache hit rate corresponding to each model instance, a target model instance from the plurality of model instances; and an inference unit, configured to schedule the model inference request to the target model instance to perform inference.
10. An electronic device, comprising: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-8.
11. A non-transitory computer readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to make the computer perform the method according to any one of claims 1-8.
12. A computer program product comprising a computer program, wherein, The computer program, when executed by the processor, implements the method according to any one of claims 1-8. The computer program, when executed by the processor, implements the method according to any one of claims 1-8.