Calling method and device of pre-training model service, equipment, medium and product
By deploying model edge gateways in the edge area, near-access and near-access cache are achieved, solving the problem of high delay when devices on the edge area call model services, and improving response speed and resource usage efficiency.
Patent Information
- Application Number
- CN202510812225.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-17
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2045-06-17
AI Technical Summary
The end-side device in the edge area has a high delay when calling model services, resulting in a low application response speed.
Deploy the model edge gateway in the edge area, query the pre-trained model cache information in the cache server through the model edge gateway, and directly return it to the end-side device when there is matching information on the cache server. Otherwise, the target model service will be obtained through the model server in the intranet environment.
It reduces the call time of model services, improves the application response speed of end-side devices, saves model inference calculations, and optimizes resource utilization.
Smart Images

Figure CN120378476A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer technologies, and particularly to a method, apparatus, device, medium, and product for invoking a pre-trained model service. Background Art
[0002] Currently, model gateway services are mainly deployed in central regions. When end-side devices in different edge regions call model services, they all need to go through the central network. Even if the cache is hit, the latency from the end-side device to the central network is relatively high, resulting in a relatively low response speed of the application programs on the end-side devices in the edge regions. Summary of the Invention
[0003] In view of this, the present disclosure provides a method, apparatus, device, medium, and product for invoking a pre-trained model service to solve the problem that the latency of the end-side device in the edge region calling the model service is relatively high, resulting in a relatively low application response speed.
[0004] In a first aspect, the present disclosure provides a method for invoking a pre-trained model service, including: obtaining a model edge gateway corresponding to an end-side device, where the model edge gateway and the end-side device are deployed in the same edge region; when receiving a model service invocation request initiated by the end-side device, querying, through the model edge gateway, whether there is pre-trained model cache information matching the model service invocation request in a cache server, where the cache server and the model edge gateway are deployed in the same edge region; if there is pre-trained model cache information matching the model service invocation request in the cache server, returning the pre-trained model cache information to the end-side device; if there is no pre-trained model cache information matching the model service invocation request in the cache server, obtaining the internal network environment where the model edge gateway is located; sending the model service invocation request to a target model server in the internal network environment, and returning, through the target model server, a target model service matching the model service invocation request.
[0005] Second aspect, the present disclosure provides an apparatus for invoking a pre-trained model service, including: an acquisition module, configured to acquire a model edge gateway corresponding to an edge device, where the model edge gateway and the edge device are deployed in the same edge area; a cache query module, configured to, when receiving a model service invocation request initiated by the edge device, query whether there is pre-trained model cache information matching the model service invocation request in a cache server through the model edge gateway, where the cache server and the model edge gateway are deployed in the same edge area; a model feedback module, configured to, if the cache server has pre-trained model cache information matching the model service invocation request, return the pre-trained model cache information to the edge device; an intranet information acquisition module, configured to, if the cache server does not have pre-trained model cache information matching the model service invocation request, acquire the intranet environment where the model edge gateway is located; a model matching module, configured to send the model service invocation request to a target model server in the intranet environment, and return a target model service matching the model service invocation request through the target model server.
[0006] Third aspect, the present disclosure provides a computer device, including: a memory and a processor, which are communicatively connected to each other, where computer instructions are stored in the memory, and the processor executes the computer instructions to execute the method for invoking a pre-trained model service according to the first aspect or any corresponding implementation manner thereof.
[0007] Fourth aspect, the present disclosure provides a computer-readable storage medium, on which computer instructions are stored, and the computer instructions are used to cause a computer to execute the method for invoking a pre-trained model service according to the first aspect or any corresponding implementation manner thereof.
[0008] Fifth aspect, the present disclosure provides a computer program product, including computer instructions, and the computer instructions are used to cause a computer to execute the method for invoking a pre-trained model service according to the first aspect or any corresponding implementation manner thereof.
[0009] The method, apparatus, device, medium and product for invoking a pre-trained model service provided by the present disclosure deploy a model edge gateway in the edge area where the end-side device is located, so as to query the pre-trained model cache information matching the model service invocation request from the cache server in the edge area where it is located. When it is determined that the cache server has the pre-trained model cache information matching the model service invocation request, the pre-trained model cache information can be directly returned to the end-side device. Thus, the model gateway and the related cache server are sunk from the center to the edge area closer to the end-side device, enabling the end-side device to access nearby, reason nearby, and provide nearby caching. Compared with retrieving the model from the central cache, the invocation time is greatly reduced, and the application response speed of the end-side device is improved. When it is determined that the cache server does not have the pre-trained model cache information matching the model service invocation request, inference is performed on the model server in the same intranet environment as the model edge gateway to obtain the corresponding target model service. Compared with performing inference after returning to the central model server, the latency is greatly reduced. At the same time, the model inference calculation is saved, the resource utilization rate is optimized, the throughput is increased, and in scenarios with many similar queries, accessing nearby can greatly reduce the invocation time and greatly improve the response speed. BRIEF DESCRIPTION OF THE DRAWINGS
[0010] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the prior art, the following will briefly introduce the drawings required for use in the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0011] Figure 1 is a schematic diagram of a model invocation method in the related art; Figure 2 is a flowchart of a method for invoking a pre-trained model service according to an embodiment of the present disclosure; Figure 3 is a specific schematic diagram of model invocation according to an embodiment of the present disclosure; Figure 4 is a flowchart of another method for invoking a pre-trained model service according to an embodiment of the present disclosure; Figure 5 is a flowchart of yet another method for invoking a pre-trained model service according to an embodiment of the present disclosure; Figure 6 is a schematic diagram of model invocation under the shortest delay first strategy according to an embodiment of the present disclosure; Figure 7 is a schematic diagram of model invocation under the maximum throughput first strategy according to an embodiment of the present disclosure; Figure 8 It is a schematic diagram of model invocation under the resource consumption optimization strategy according to an embodiment of the present disclosure; Figure 9 It is a schematic diagram of the distribution of model metrics according to an embodiment of the present disclosure; Figure 10 It is a structural block diagram of a calling device for a pre-trained model service according to an embodiment of the present disclosure; Figure 11 It is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present disclosure. Detailed implementation manners
[0012] To make the objectives, technical solutions and advantages of the embodiments of the present disclosure clearer, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present disclosure. Apparently, the described embodiments are some but not all of the embodiments of the present disclosure. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present disclosure.
[0013] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure shall be informed to the user and the user's authorization shall be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0014] For example, when responding to receiving an active request from a user, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server, or a storage medium that executes the operation of the technical solution of the present disclosure according to the prompt message.
[0015] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving an active request from the user may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0016] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manner of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0017] It can be understood that the data involved in the technical solution of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of corresponding laws, regulations and related provisions.
[0018] With the increasing use of large language models (LLMs) in various industries and applications on various edge devices, it has become increasingly important to improve the response speed and throughput of LLMs so that they can operate on a larger scale and with higher cost-effectiveness. Therefore, the large model gateway has emerged, playing a crucial role in connecting various applications. It provides important functions such as a wider range of LLM support, query acceleration, cache mechanism construction, monitoring, and security protection.
[0019] The large model gateway services in the related technologies are mainly deployed in the center, such as Figure 1 shown, and the capabilities it provides mainly include: 1) proxy large models from different public cloud providers; 2) monitoring and log tracking; 3) API authentication and authorization; 4) traffic limiting policies; 5) security prevention, etc. Based on this, the overall link of its implementation method is roughly: 1) the edge device initiates a large model service call request; 2) the request is forwarded to the central gateway server; 3) after the gateway server passes the authentication and authorization, cache recall is performed based on vector similarity. If a hit occurs, the cached result is directly returned. If not, it goes back to the large models of each different public cloud manufacturer; 4) after each manufacturer's public cloud large model responds to the corresponding request, it is returned to the edge device through the gateway. However, in this process, when edge devices in different edge regions call large models, they all need to pass through the central network. Even if the cache is hit, the latency from the edge device to the central network will be relatively high. As a result, for scenarios with high latency requirements, inferring in the center will, to a certain extent, bring latency impacts and affect the application response speed of edge devices. Moreover, the application programs on edge devices usually use multiple large models from different providers, which will consume a large amount of resources during the call process, and errors of temporarily unable to process requests often occur when frequently calling a specific model of a certain provider.
[0020] Based on this, the technical solution of the present disclosure sinks the model gateway service, related cache infrastructure, and model server from the center to the edge region where the edge device is located to achieve nearby access, nearby inference, and nearby caching, greatly reducing the latency of the edge device calling the model service. In addition, by tracking the latency and resource consumption changes of each model service, automatic optimization of the model service is performed according to actual needs, reducing the call latency and resource consumption of the model service while ensuring the availability of edge device applications.
[0021] According to an embodiment of the present disclosure, an embodiment of a method for calling a pre-trained model service is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0022] In this embodiment, a method for calling a pre-trained model service is provided, which can be used in computer devices such as gateways, servers, etc. Figure 2 is a flowchart of a method for calling a pre-trained model service according to an embodiment of the present disclosure. Figure 2 As shown, the process includes the following steps: Step S101, obtain the model edge gateway corresponding to the end-side device, and the model edge gateway and the end-side device are deployed in the same edge area.
[0023] The model edge gateway is deployed in the edge area where the end-side device is located, that is, the model edge gateway and the end-side device are deployed in the same edge area. The same edge area is an area with a close geographical location, such as the same city A; the end-side device refers to the device running at the edge of data processing, that is, the device close to the data source or the end user. The end-side device can be a personal device such as a tablet computer and a laptop computer, or an IoT device such as a smart home device and a wearable device. A smart application is deployed on the end-side device, and the smart application will use multiple network models (such as a large language model, etc.) from the same manufacturer or different manufacturers.
[0024] For terminal devices in different edge areas, the model edge gateway deployed in the edge area is determined by detecting the edge area where the terminal device is located, so that when the terminal device initiates a model service call request, the model service call request is routed to the model edge gateway through the gateway domain name.
[0025] Step S102, when a model service call request initiated by an end-side device is received, the model edge gateway is used to query whether there is pre-trained model cache information matching the model service call request in the cache server, and the cache server and the model edge gateway are deployed in the same edge area.
[0026] The cache server is a server that is connected to the model edge gateway and is located in the same edge area and the same intranet environment. The server stores the model information that has been called by the end-side device in the past, that is, the pre-trained model cache information. The model service call request is a request initiated when the application in the end-side device needs to call the model.
[0027] Specifically, Figure 3 As shown, when the end-side device initiates a model service call request to the model edge gateway through the gateway domain name, the model edge gateway can receive the model service call request initiated by the end-side device, parse the information carried by the model service call request, and determine the model information requested by the model service call request. Then, according to the information carried by the model service call request, query the existing cache information in the cache server, and determine whether there is pre-trained model cache information that matches it through similarity calculation.
[0028] Step S103, if the pre-trained model cache information matching the model service call request exists in the cache server, return the pre-trained model cache information to the edge device.
[0029] If the pre-trained model cache information matching the model service call request exists in the cache server, it means that the edge device has requested the relevant model service before. At this time, the pre-trained model cache information can be directly retrieved from the cache server and returned to the edge device, as Figure 3 shown. If the pre-trained model cache information matching the model service call request does not exist in the cache server, other operations are performed, such as retrieving from the model server in the edge area and performing matching again, which is not specifically limited here.
[0030] The method for calling the pre-trained model service provided in this embodiment deploys a model edge gateway in the edge area where the edge device is located, so as to query the pre-trained model cache information matching the model service call request from the cache server in the edge area where it is located through the model edge gateway. When it is determined that the pre-trained model cache information matching the model service call request exists in the cache server, the pre-trained model cache information can be directly returned to the edge device. Thus, the model gateway and the relevant cache server are sunk from the center to the edge area closer to the edge device, enabling the edge device to access nearby, perform inference nearby, and provide nearby caching. Compared with retrieving the model from the central cache, the call time is greatly reduced, and the application response speed of the edge device is improved. At the same time, the model inference calculation is saved, the resource utilization rate is optimized, and the throughput is increased. In scenarios with many similar queries, accessing nearby can greatly reduce the call time and significantly improve the response speed.
[0031] In this embodiment, a method for calling the pre-trained model service is provided, which can be used in computer devices such as gateways and servers. Figure 4 It is a flowchart of the method for calling the pre-trained model service according to the embodiments of the present disclosure, as Figure 4 shown, and the process includes the following steps: Step S201, obtain the model edge gateway corresponding to the edge device, and the model edge gateway and the edge device are deployed in the same edge area. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be repeated here.
[0032] Step S202, when receiving a model service call request initiated by the edge device, query through the model edge gateway whether there is pre-trained model cache information in the cache server that matches the model service call request. The cache server and the model edge gateway are deployed in the same edge area. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be repeated here.
[0033] Step S203, if the cache server has pre-trained model cache information that matches the model service call request, return the pre-trained model cache information to the edge device. For details, please refer to the relevant descriptions of the corresponding steps in the above embodiments, which will not be elaborated here.
[0034] Step S204, if the cache server does not have pre-trained model cache information that matches the model service call request, obtain the internal network environment where the model edge gateway is located.
[0035] The internal network environment where the model edge gateway is located refers to the network infrastructure that supports edge computing and data processing. Specifically, network devices such as switches, routers, and firewalls are deployed in the internal network environment to connect the model edge gateway, model server, cache server, edge device, etc. that access the internal network environment through each network device. At the same time, corresponding security measures such as access control, encryption, and defense systems are also deployed in the internal network environment to ensure the data security of the internal network environment.
[0036] If the cache server does not have pre-trained model cache information that matches the model service call request, it is necessary to request the corresponding model information from the nearest model server. At this time, it is necessary to parse the network topology where the model edge gateway is located to obtain the internal network environment where the model edge gateway is located.
[0037] Step S205, send the model service call request to the target model server in the internal network environment, and return the target model service that matches the model service call request through the target model server.
[0038] The target model server refers to the model server that points to the edge device to feedback model information; the target model service is the model information required by the model service call request inferred by the target model server. After determining that there is no corresponding pre-trained model cache information in the cache server, the model edge gateway will route the model service call request to the target model server in the internal network environment, so that the target model server performs inference calculations according to the model service call request and returns the corresponding target model service.
[0039] For the method of calling the pre-trained model service provided in this embodiment, when it is determined that the cache server does not have pre-trained model cache information that matches the model service call request, inference is performed on the model server in the same internal network environment as the model edge gateway to obtain the corresponding target model service. Compared with performing inference after returning to the central model server, the latency is greatly reduced.
[0040] In some alternative embodiments, as Figure 5 shown, the above step S205 includes: Step S2051, obtaining multiple model servers deployed in the intranet environment and the load request amount of each model server.
[0041] The load request volume is the number of pending model service call requests currently loaded by the model server. Multiple model servers are deployed in the intranet environment to provide model inference services. Each model server can receive model service call requests from different end-side devices, and the number of model service call requests received and the processing status of each model server are updated in real time to determine the model service call requests that have been processed and the pending model service call requests, thereby obtaining the load request volume of each model server.
[0042] Step S2052: Based on the load request amount, a target model server is determined from each model server according to load balancing.
[0043] When a new model service call request enters the intranet environment, the load request volume corresponding to each model server is analyzed, and one or more model servers that can currently allocate the model service call request are determined in a load balancing manner. If there is only one model server that can allocate the model service call request, then this model server is determined as the target model server; if there are multiple model servers that can allocate the model service call request, then the optimal model server is determined based on the model server's data processing speed, memory usage and other parameter information, and the optimal model server is determined as the target model server.
[0044] Step S2053, obtaining the model service calling strategy corresponding to the terminal device and multiple candidate model services matching the model service calling request.
[0045] The model service call policy is the policy for the end-side device to call the model service, such as shortest delay priority, maximum throughput priority, etc. The model service call policy corresponding to the end-side device is pre-configured, such as the shortest delay priority policy corresponding to end-side device 1, and the maximum throughput priority policy corresponding to end-side device 2. When the end-side device initiates a model service call request, the model service call policy is encapsulated in the model service call request. Therefore, when the model edge gateway receives the model service call request, it can obtain the model service call policy corresponding to the end-side device by parsing the content carried in the model service call request.
[0046] A candidate model service is a model information that can provide the functions required by a model service call request. For the same model service call request, there may be multiple models that can meet its requirements. Specifically, when the model server performs inference calculations according to the model service call request, it will filter out all model information that meets the model service call request, that is, multiple candidate model services.
[0047] Step S2054, determine the target model service from multiple candidate model services according to the model service invocation policy.
[0048] According to the model service invocation policy, determine the model service parameters of each candidate model service under the model service invocation policy, such as model service latency, model service throughput, etc. Sort the model service parameters corresponding to each candidate model service to generate a corresponding parameter sorting result, and characterize the model service invocation priority through this parameter sorting result. Then, screen the required target model service from multiple candidate model services according to the model service invocation priority.
[0049] The method for invoking the pre-trained model service provided in this embodiment determines the target model server from each model server according to load balancing, so as to process the model service invocation request through the target model server, avoid request overload of the model servers in the intranet environment, optimize the resource utilization of each model server in the intranet environment, and achieve minimum throughput and minimum response time, thereby greatly shortening the latency of the end-side device invoking the model service and improving the application response speed. By configuring multiple candidate model services that match the model service invocation request, it is possible to avoid too many retrieval failures and affect the user experience. When there are multiple candidate model services, by automatically optimizing the target model service according to the model service invocation policy, the application response speed and application availability of the end-side device in the edge area are greatly improved.
[0050] In some alternative embodiments, when the model service invocation policy is the shortest latency first policy, the above step S2054 may include: Step a1, obtain the model service latency of each candidate model service under the shortest latency first policy.
[0051] Step a2, sort the model service latencies of each model service, and determine the first invocation priority of each candidate model service based on the latency sorting result.
[0052] Step a3, determine the target model service according to the first invocation priority.
[0053] The model service latency is the time consumed for the model to perform task inference and output the result. The model service latency of each candidate model service is determined based on the processing speed of the candidate model service. Specifically, the central gateway can monitor the task processing status of each candidate model service, determine the average latency required for each candidate model service to process the task in real time, and synchronize the model service latency corresponding to each candidate model service to the model servers in each edge area in real time.
[0054] Sort the model service latencies of each candidate model service under the shortest latency first policy to obtain a latency sorting result from smallest to largest model service latency. Through this latency sorting result, the call priorities of each candidate model service during invocation can be determined.
[0055] Taking the large model as an example, the model service latency can be characterized by the first token latency (TTFT). As Figure 6 shown, the edge device 1 enables the TTFT shortest first policy, and the candidate model services are LLM1, LLM2, and LLM3 respectively. When the edge device 1 accesses the model edge gateway to call the large model, it sorts based on the shortest average TTFT of the three candidate model services and invokes them in sequence. For example, LLM3 is 100 milliseconds, LLM1 is 150 milliseconds, and LLM2 is 200 milliseconds. Since the TTFT of LLM3 is the shortest, LLM3 is called first. If the call to LLM3 fails, LLM1 with the second shortest TTFT is automatically called, and so on.
[0056] The method for invoking the pre-trained model service provided in this embodiment supports automatically optimizing the target model service according to the shortest latency first policy, thereby ensuring the application availability of the edge device under the shortest latency first policy and improving the application response speed of the edge device.
[0057] In some alternative embodiments, when the model service invocation policy is the maximum throughput first policy, the above step S2054 may include: Step b1, obtain the model service throughput of each candidate model service under the maximum throughput first policy.
[0058] Step b2, sort the throughputs of each model service, and determine the second call priority of each candidate model service based on the throughput sorting result.
[0059] Step b3, determine the target model service according to the second call priority.
[0060] The model service throughput is the amount of data that the model can process per unit time. The model service throughput of each candidate model service is determined based on the processing speed of the candidate model service. Specifically, the central gateway can monitor the task processing status of each candidate model service, determine the model service throughput generated by each candidate model service for task processing in real time, and synchronize the model service throughput corresponding to each candidate model service to the model servers in each edge area in real time.
[0061] Sort the model service throughputs of each candidate model service obtained under the maximum throughput first strategy to obtain a throughput sorting result from largest to smallest model service throughput. Through this throughput sorting result, the call priorities of each candidate model service during invocation can be determined.
[0062] Taking a large model as an example, the model service throughput can be characterized by the number of tokens per second (TPS). As Figure 7 shown, the edge device 2 enables the TPS maximum first strategy, and the candidate model services are LLM1, LLM2, and LLM3 respectively. When the edge device 2 accesses the model edge gateway to invoke the large model, it sorts based on the maximum average TPS of the three candidate model services and invokes them in sequence. For example, the TPS of LLM3 is 20, the TPS of LLM2 is 15, and the TPS of LLM1 is 10. Since the TPS of LLM3 is the highest, LLM3 is invoked first. If the invocation of LLM3 fails, LLM2 with the second highest TPS is automatically invoked, and so on.
[0063] The method for invoking the pre-trained model service provided in this embodiment supports automatically selecting the optimal target model service according to the maximum throughput first strategy, thereby ensuring the application availability of the edge device under the maximum throughput first strategy and ensuring the application response speed of the edge device.
[0064] In some alternative embodiments, when the model service invocation policy is a resource consumption optimization policy, the above step S2054 may include: Step c1, obtain the model service resource consumption of each candidate model service under the resource consumption optimization policy.
[0065] Step c2, sort the model service resource consumptions of each model service, and determine the third call priority of each candidate model service based on the resource consumption sorting result.
[0066] Step c3, determine the target model service according to the third call priority.
[0067] The model service resource consumption is the resources consumed during the use of the model. The model service resource consumption of each candidate model service is determined based on the inherent attributes of the candidate model service. Specifically, the central gateway can parse the attribute information of each candidate model service, determine the model service resource consumption of each candidate model service, and synchronize the model service resource consumption corresponding to each candidate model service to the model servers in each edge area in real time.
[0068] Sort the model service resource consumption of each candidate model service obtained under the resource consumption optimization strategy to obtain the resource consumption sorting result with the model service resource consumption from low to high. Through this resource consumption sorting result, the call priority of each candidate model service during invocation can be determined.
[0069] Taking the large model as an example, the model service resource consumption can be characterized by the model usage cost. As Figure 8 shown, the edge device 3 enables the resource consumption optimization strategy, and the candidate model services are LLM1, LLM2, and LLM3 respectively. When the edge device 3 accesses the model edge gateway to call the large model, sort based on the model service resource consumption of the three candidate model services and call them in sequence. For example, the model service usage cost of LLM3 is 0.0003 yuan per thousand tokens, the model service usage cost of LLM2 is 0.005 yuan per thousand tokens, and the model service usage cost of LLM1 is 0.001 yuan per thousand tokens. Since the model service usage cost of LLM3 is the least, LLM3 is called first. If the call to LLM3 fails, then the model service with the second lowest usage cost, LLM1, is automatically called, and so on.
[0070] The method for invoking the pre-trained model service provided in this embodiment supports automatically optimizing the target model service according to the resource consumption optimization strategy, thereby reducing the resource consumption on the basis of ensuring the application availability of the edge device under the resource consumption optimization strategy.
[0071] In some alternative embodiments, as Figure 9 shown, the model metrics corresponding to the candidate model services are sent in real time through the central gateway corresponding to the model edge gateway, including model service latency, model service throughput, resource consumption, etc.
[0072] When deploying the central gateway service, set corresponding tracking services for the model metrics of the model (such as model service latency TTFT, model service throughput TPS, etc.) and the model service resource consumption of each manufacturer (such as model cost), so as to synchronize and send the model latency, model service throughput, resource consumption and other model metrics of each model to the model servers in the edge areas where each model edge gateway is located in real time through the central gateway configuration service for making decisions on the target model service.
[0073] The method for invoking the pre-trained model service provided in this embodiment sends the model metrics corresponding to the candidate model services in real time through the central gateway corresponding to the model edge gateway, ensuring the accuracy of each model metric in the edge area and improving the decision-making accuracy of the target model service.
[0074] In this embodiment, a device for invoking a pre-trained model service is further provided. This device is used to implement the above-mentioned embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0075] This embodiment provides a device for invoking a pre-trained model service, as Figure 10 shown, including: An acquisition module 301, configured to acquire a model edge gateway corresponding to an end-side device, where the model edge gateway and the end-side device are deployed in the same edge area.
[0076] A cache query module 302, configured to, when receiving a model service invocation request initiated by an end-side device, query whether there is pre-trained model cache information matching the model service invocation request in a cache server through the model edge gateway, where the cache server and the model edge gateway are deployed in the same edge area.
[0077] A model feedback module 303, configured to, if the cache server has pre-trained model cache information matching the model service invocation request, return the pre-trained model cache information to the end-side device.
[0078] An intranet information acquisition module 304, configured to, if the cache server does not have pre-trained model cache information matching the model service invocation request, acquire the intranet environment where the model edge gateway is located.
[0079] A model matching module 305, configured to send the model service invocation request to a target model server in the intranet environment, and return a target model service matching the model service invocation request through the target model server.
[0080] In some optional implementation manners, the model matching module includes: A load acquisition unit, configured to acquire multiple model servers deployed in the intranet environment and the load request volume of each model server.
[0081] A load balancing unit, configured to determine a target model server from each model server based on the load request volume according to load balancing.
[0082] In some optional implementation manners, the model matching module further includes: A policy acquisition unit, configured to acquire a model service invocation policy corresponding to the end-side device, and multiple candidate model services matching the model service invocation request.
[0083] A model determination unit, configured to determine a target model service from multiple candidate model services according to a model service invocation policy.
[0084] In some alternative embodiments, when the model service invocation policy is the shortest delay first policy, the model determination unit includes: A model service delay acquisition subunit, configured to acquire the model service delays of each candidate model service under the shortest delay first policy.
[0085] A first priority determination subunit, configured to sort the model service delays of each candidate model service and determine the first invocation priority of each candidate model service based on the delay sorting result.
[0086] A first determination subunit, configured to determine the target model service according to the first invocation priority.
[0087] In some alternative embodiments, when the model service invocation policy is the maximum throughput first policy, the model determination unit includes: A throughput acquisition subunit, configured to acquire the model service throughputs of each candidate model service under the maximum throughput first policy.
[0088] A second priority determination subunit, configured to sort the model service throughputs of each candidate model service and determine the second invocation priority of each candidate model service based on the throughput sorting result.
[0089] A second determination subunit, configured to determine the target model service according to the second invocation priority.
[0090] In some alternative embodiments, when the model service invocation policy is the resource consumption optimization policy, the model determination unit includes: A resource consumption acquisition subunit, configured to acquire the model service resource consumptions of each candidate model service under the resource consumption optimization policy.
[0091] A third priority determination subunit, configured to sort the model service resource consumptions of each candidate model service and determine the third invocation priority of each candidate model service based on the resource consumption sorting result.
[0092] A third determination subunit, configured to determine the target model service according to the third invocation priority.
[0093] In some alternative embodiments, the above device further includes: An indicator distribution module, configured to distribute the model indicators corresponding to the candidate model services in real time through the central gateway corresponding to the model edge gateway.
[0094] The further function descriptions of the above modules and units are the same as those in the corresponding above embodiments, and will not be elaborated here.
[0095] The calling device for the pre-trained model service in this embodiment is presented in the form of functional units. Here, the unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and a memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0096] The calling device for the pre-trained model service provided in this embodiment deploys a model edge gateway in the edge area where the edge device is located, so as to query the pre-trained model cache information matching the model service call request from the cache server in the edge area where it is located. When it is determined that the cache server has the pre-trained model cache information matching the model service call request, the pre-trained model cache information can be directly returned to the edge device. Thus, the model gateway and the related cache server are sunk from the center to the edge area closer to the edge device, enabling the edge device to access nearby, reason nearby, and provide nearby caching. Compared with retrieving the model from the central cache, the calling time is greatly reduced, and the application response speed of the edge device is improved. At the same time, the model inference calculation is saved, the resource utilization rate is optimized, and the throughput is increased. In scenarios with many similar queries, accessing nearby can greatly reduce the calling time and significantly improve the response speed.
[0097] This disclosure embodiment also provides a computer device having Figure 10 the calling device for the pre-trained model service as shown.
[0098] Please refer to Figure 11 , Figure 11 which is a schematic structural diagram of a computer device provided by an optional embodiment of this disclosure. As shown in Figure 11 , the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including a high-speed interface and a low-speed interface. Each component communicates with each other using different buses and can be installed on a common motherboard or in other ways as needed. The processor can process instructions executed within the computer device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some optional implementation manners, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (such as an array of servers, a set of blade servers, or a multi-processor system). Figure 11 In
[0099] The processor 10 may be a central processing unit, a network processor, or a combination thereof. Among them, the processor 10 may further include a hardware chip. The above-mentioned hardware chip may be an application-specific integrated circuit, a programmable logic device, or a combination thereof. The above-mentioned programmable logic device may be a complex programmable logic device, a field programmable gate array, a generic array logic, or any combination thereof.
[0100] Among them, the memory 20 stores instructions executable by at least one processor 10, so that the at least one processor 10 executes the method shown in the above embodiments.
[0101] The memory 20 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some alternative embodiments, the memory 20 may optionally include a memory remotely provided with respect to the processor 10, and these remote memories may be connected to the computer device through a network. Examples of the above-mentioned network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0102] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk, or a solid-state drive; the memory 20 may also include a combination of the above types of memories.
[0103] The computer device further includes a communication interface 30 for the computer device to communicate with other devices or communication networks.
[0104] The embodiments of the present disclosure also provide a computer-readable storage medium. The method according to the embodiments of the present disclosure may be implemented in hardware, firmware, or may be implemented as computer code that can be recorded on a storage medium, or may be implemented as computer code originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and stored in a local storage medium, so that the method described herein can be stored in such software processed on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium may be a magnetic disk, an optical disk, a read-only memory, a random access memory, a flash memory, a hard disk, or a solid-state drive, etc.; further, the storage medium may also include a combination of the above types of memories. It can be understood that a computer, a processor, a microprocessor controller, or programmable hardware includes a storage component that can store or receive software or computer code, and when the software or computer code is accessed and executed by the computer, the processor, or the hardware, the method shown in the above embodiments is implemented.
[0105] Part of the present disclosure can be applied as a computer program product, for example, computer program instructions. When executed by a computer, through the operations of the computer, it can call or provide the methods and / or technical solutions according to the present disclosure. Those skilled in the art should understand that the forms in which computer program instructions exist in a computer-readable medium include, but are not limited to, source files, executable files, installation package files, etc. Correspondingly, the ways in which computer program instructions are executed by a computer include, but are not limited to: the computer directly executes the instructions, or the computer compiles the instructions and then executes the corresponding compiled program, or the computer reads and executes the instructions, or the computer reads and installs the instructions and then executes the corresponding installed program. Herein, the computer-readable medium can be any available computer-readable storage medium or communication medium accessible to the computer.
[0106] Although the embodiments of the present disclosure have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations all fall within the scope defined by the appended claims.
Claims
1. A method for invoking a pre-trained model service, characterized in that The method includes: Obtaining a model edge gateway corresponding to an edge device, where the model edge gateway and the edge device are deployed in the same edge area; When receiving a model service call request initiated by the edge device, querying whether there is pre-trained model cache information matching the model service call request in a cache server through the model edge gateway, where the cache server and the model edge gateway are deployed in the same edge area; If the cache server has pre-trained model cache information matching the model service call request, returning the pre-trained model cache information to the edge device; If the cache server does not have pre-trained model cache information matching the model service call request, obtaining the internal network environment where the model edge gateway is located; Sending the model service call request to a target model server in the internal network environment, and returning a target model service matching the model service call request through the target model server.
2. The method according to claim 1, characterized in that Sending the model service call request to the target model server in the internal network environment includes: Obtaining multiple model servers deployed in the internal network environment and the load request volume of each model server; Based on the load request volume, determining the target model server from each model server according to load balancing.
3. The method according to claim 1, wherein The returning a target model service matching the model service call request through the target model server includes: Obtaining a model service call policy corresponding to the edge device and multiple candidate model services matching the model service call request; Determining the target model service from the multiple candidate model services according to the model service call policy.
4. The method according to claim 3, wherein When the model service call policy is the shortest delay first policy, the determining the target model service from the multiple candidate model services according to the model service call policy includes: Obtaining the model service delay of each candidate model service under the shortest delay first policy; Sorting the model service delays of each candidate model service, and determining the first call priority of each candidate model service based on the delay sorting result; Determining the target model service according to the first call priority.
5. The method according to claim 3, characterized in that When the model service call policy is the maximum throughput first policy, the determining the target model service from the multiple candidate model services according to the model service call policy includes: Obtaining the model service throughput of each candidate model service under the maximum throughput first policy; Sorting the model service throughputs of each candidate model service, and determining the second call priority of each candidate model service based on the throughput sorting result; Determining the target model service according to the second call priority.
6. The method according to claim 3, wherein When the model service call policy is the resource consumption optimization policy, the determining the target model service from the multiple candidate model services according to the model service call policy includes: Obtaining the model service resource consumption of each candidate model service under the resource consumption optimization policy; Sort the resource consumption of each of the model service resources, and determine the third call priority of each of the candidate model services based on the resource consumption sorting result; Determine the target model service according to the third call priority.
7. The method according to any one of claims 3 to 6, characterized in that, It further includes: Real-time send the model metrics corresponding to the candidate model services through the central gateway corresponding to the model edge gateway.
8. An apparatus for invoking a pre-trained model service, characterized in that, The device includes: An acquisition module, configured to acquire a model edge gateway corresponding to the end-side device, where the model edge gateway and the end-side device are deployed in the same edge area; A cache query module, configured to, when receiving a model service call request initiated by the end-side device, query whether there is pre-trained model cache information matching the model service call request in a cache server through the model edge gateway, where the cache server and the model edge gateway are deployed in the same edge area; A model feedback module, configured to, if the cache server has pre-trained model cache information matching the model service call request, return the pre-trained model cache information to the end-side device; An intranet information acquisition module, configured to, if the cache server does not have pre-trained model cache information matching the model service call request, acquire the intranet environment where the model edge gateway is located; A model matching module, configured to send the model service call request to a target model server in the intranet environment, and return a target model service matching the model service call request through the target model server.
9. A computer device, characterized in that, It includes: A memory and a processor, where the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to execute the method for calling the pre-trained model service according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, Computer instructions are stored on the computer-readable storage medium, and the computer instructions are used to cause a computer to execute the method for calling the pre-trained model service according to any one of claims 1 to 7.
11. A computer program product, characterized in that, It includes computer instructions, and the computer instructions are used to cause a computer to execute the method for calling the pre-trained model service according to any one of claims 1 to 7.
Citation Information
Patent Citations
Model calling method, device and equipment and storage medium
CN112596919A
Edge data caching method
CN115934308A
Model deployment method and system and storage medium
CN116700985A
Service request processing method and device, server and storage medium
CN117278641A
Model reasoning service deployment method and device, equipment and storage medium
CN120104328A