A method, device and system for inference service

By introducing a message bus and subscription publishing mechanism in the online model inference service, the service instances allow them to select and process requests based on their own conditions, solving the problems of uneven allocation of inference requests and insufficient fault tolerance in the prior art, and achieving more efficient and reliable inference service processing.

CN114021052BActive Publication Date: 2025-07-01DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111130073.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-26
Publication Date
2025-07-01
Estimated Expiration
2041-09-26

AI Technical Summary

Technical Problem

Existing online model inference services cannot achieve uniform rationality when allocating inference requests, resulting in overloading or idling of service instances, increasing the average time-consuming of requests, and failing to obtain available inference services in real time and accurately, resulting in request failure and poor fault tolerance.

Method used

The subscription and publishing mechanism of the message bus is adopted to place inference requests into the message queue, and send new request notifications to the service instances subscribed to the queue, allowing the service instance to decide whether to undertake the request based on its own load and availability.

Benefits of technology

By actively selecting and processing requests through service instances, load balancing is achieved, overload and no-load problems are avoided, the success rate and efficiency of request processing are improved, and the fault tolerance of the system is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114021052B_ABST
    Figure CN114021052B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for inference service. In this method, after receiving an inference request sent by a client, the message bus puts it into a message queue corresponding to its service type, and sends a new request notification to a service instance that subscribes to this message queue. After receiving the new request notification, the service instance can determine whether to accept the request according to its actual performance, including the load situation and availability. If it can accept the request, it fetches the inference request from the message bus and processes it. During the processing of this request, the service instance accepts the request according to its actual performance, ensuring the balanced processing of requests; moreover, when the inference request is sent to the message bus, the request can continue to be processed after the network is restored, with high fault tolerance; at the same time, each service instance can accept and process requests simultaneously, and the processing efficiency of requests is high. The present invention also discloses an inference service device and system, which have corresponding technical effects.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of information processing technology, and in particular to an inference service method, device and system. Background Art

[0002] The model is mainly used to calculate the request data (such as text, pictures, videos, etc.) provided by the client and obtain a result (such as classification, numerical value, etc.), including machine learning models, deep neural network models and other types of models. The common model development process requires problem definition, data preparation, feature extraction, modeling, training and deployment. Among them, data preparation, feature extraction, modeling, training and deployment require strong data collection capabilities, data processing capabilities and analysis capabilities, model structure and parameter knowledge, and require strong professionalism. In addition, the performance requirements for deployed equipment are also high, and the development cost is high. Some enterprises or units find it difficult to meet the conditions for model development, but they still need the model's own strong reasoning ability to meet the high-precision requirements of their own data processing. Therefore, model reasoning services came into being.

[0003] Model inference service refers to a service that provides model capabilities to the outside world through a certain network protocol (such as http, grpc, etc.). After the client initiates an inference request, the corresponding service instance (instance, i.e., model) in the model inference service responds to the inference request to perform inference service. In order to provide multiple model services at the same time and meet high concurrency requirements, existing online model inference services usually adopt a proxy structure, where the proxy server is responsible for managing multiple model service instances and using a routing algorithm to send model inference requests to idle service instances. However, in this mode, the proxy server cannot accurately measure the actual capacity and pressure of each inference service, and cannot match the number of requests with the processing capacity of the service instance itself in the distribution of requests, which can easily cause the service instance to be overloaded or idle, increasing the average time consumption of inference requests; at the same time, it is also impossible to obtain available inference services in real time and accurately, and it is very easy for requests to fail due to the use of incorrect addresses; and the proxy service will issue the next inference request after one request is processed, which has low overall request processing efficiency. If the network fails or the inference service fails, the inference request will fail and the next inference request cannot be processed, resulting in poor fault tolerance.

[0004] To sum up, how to ensure the uniformity and rationality of the distribution of inference requests and improve the response efficiency and success rate of inference requests are technical problems that technical personnel in this field urgently need to solve. Summary of the invention

[0005] The object of the present invention is to provide a method, device and system for inference service, so as to ensure the uniform rationality of the distribution of inference requests, and improve the response efficiency and success rate of inference requests.

[0006] To solve the above technical problems, the present invention provides the following technical solutions:

[0007] An inference service method includes:

[0008] After a message bus receives an inference request sent by a client, it determines the service type of the inference request;

[0009] Add the inference request to a message queue whose topic corresponds to the service type;

[0010] Send a new request notification to a service instance subscribing to the message queue, so that the service instance receives or rejects the processing of the inference request according to its own load and service availability.

[0011] Optionally, after sending the new request notification to the service instance subscribing to the message queue, it further includes:

[0012] After receiving a request processing notification sent by a service instance, determine the processed request as a target request;

[0013] Add a file lock to the target request;

[0014] After receiving a request processing completion notification, delete the target request.

[0015] Optionally, after adding the file lock to the target request, it further includes:

[0016] If the processing of the target request is abnormal, unlock the file lock of the target request.

[0017] A message bus includes: a plurality of message queues provided with topics for indicating service types;

[0018] The message bus is configured to: after receiving an inference request sent by a client, determine the service type of the inference request; add the inference request to a message queue whose topic corresponds to the service type; send a new request notification to a service instance subscribing to the message queue, so that the service instance receives or rejects the processing of the inference request according to its own load and service availability.

[0019] An inference service method includes:

[0020] The service instance receives a new request notification sent by the message queue in the subscribed message bus; wherein, the new request notification is triggered after the message bus adds the inference request sent by the client to the message queue; the topic of the message queue corresponds to the inference service type of the service instance;

[0021] Judge whether it can undertake the inference request according to its own load and service availability;

[0022] If it can undertake, read the inference request from the message queue and perform inference processing.

[0023] Optionally, the obtaining the inference request from the message queue and performing inference processing includes:

[0024] Read several inference requests from the message queue according to its own load, and batch process each inference request simultaneously.

[0025] A computer device includes:

[0026] A memory for storing a computer program;

[0027] A processor for implementing the steps of the inference service method based on the service instance when executing the computer program.

[0028] An inference service system includes: a client, a message bus, and several service instances with different service types;

[0029] Wherein, the client is used to receive the inference request initiated by the user and send the inference request to the message bus;

[0030] The message bus contains several message queues. The message bus is used to determine the service type of the inference request after receiving the inference request; add the inference request to the message queue whose topic corresponds to the service type; send a new request notification to the service instance subscribing to the message queue;

[0031] The service instance is used to receive the new request notification; judge whether it can undertake the inference request according to its own load and service availability; if it can undertake, obtain the inference request from the message queue and perform inference processing.

[0032] Optionally, the inference service system further includes: a service manager connected to the message queue;

[0033] The service manager is used to monitor the request processing speed of each message queue in the message bus, generate a request processing monitoring record, and perform scaling processing on the service instance according to the request processing monitoring record.

[0034] Optionally, the service manager is specifically configured to: determine the request processing speed of the target message queue according to the request processing monitoring record; if the request processing speed is lower than the first threshold, add a service instance of the service type corresponding to the topic of the target message queue; if the request processing speed is higher than the second threshold, reduce the service instance of the service type corresponding to the topic of the target message queue; where the first threshold is lower than the second threshold.

[0035] The method provided by the embodiments of the present invention combines the characteristics of stateless model inference services with large fluctuations in the request volume and the advantages of easy scalability and high fault tolerance of the message bus. The allocation of inference requests is changed from the active selection by the proxy service to the active selection by the inference service. After receiving the inference request sent by the client, the message bus puts it into the message queue and sends a new request notification to the service instance subscribing to the message queue, indicating that a new request has arrived at the message queue. After receiving the new request notification, the service instance can determine whether to undertake the request according to its actual performance, including the load situation and availability. If it can undertake the request, it obtains the inference request from the message bus and processes it. During the processing of this request, the service instance undertakes the request according to its actual performance, avoiding the problems of overload and idleness caused by not understanding the actual performance, ensuring load balancing, and also avoiding the problem of undertaking requests when unavailable. The success rate of request processing is high; moreover, when the inference request is sent to the message bus, even if the inference service cannot receive the request due to a network failure, the request can continue to be processed after the network is restored, with high fault tolerance; at the same time, each service instance can undertake and process requests simultaneously, significantly improving the request processing efficiency.

[0036] Correspondingly, the embodiments of the present invention also provide an inference service device and system corresponding to the above inference service method, which have the above technical effects and will not be elaborated here. Description of the Drawings

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0038] Figure 1 It is a schematic diagram of a traditional proxy structure;

[0039] Figure 2 It is a signaling diagram of an inference service method in the embodiments of the present invention;

[0040] Figure 3Schematic structural diagram of a message bus in an embodiment of the present invention;

[0041] Figure 4 Schematic structural diagram of a computer device in an embodiment of the present invention;

[0042] Figure 5 Schematic structural diagram of an inference service system in an embodiment of the present invention;

[0043] Figure 6 Schematic structural diagram of another inference service system in an embodiment of the present invention. Detailed implementation manners

[0044] The core of the present invention is to provide an inference service method, which can ensure the uniform rationality of the distribution of inference requests, and improve the response efficiency and success rate of inference requests.

[0045] In order to enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0046] In order to be able to provide multiple model services simultaneously and meet the requirements of high concurrency for existing online model inference services, a proxy structure method is usually adopted, such as Figure 1 As shown, the system mainly includes: a client, a proxy server, service instances (i.e., service instances, hereinafter simply referred to as service instances), and a service manager.

[0047] The processing process of an inference request based on the above proxy structure is as follows:

[0048] 1. The client sends an inference request to the proxy server;

[0049] 2. The proxy server queries the service manager according to the inference service type of the inference request to obtain all service instances of the inference service (there may be multiple service instances for each inference service);

[0050] 3. The proxy server uses a routing algorithm to select an idle service instance from all instances for the model inference request according to a load balancing strategy (such as random selection, least number of connections, round robin RoundRobin, etc.);

[0051] 4. The proxy server sends the inference request to the selected service instance;

[0052] 5. After receiving the request, the service instance responds to the request to perform corresponding data inference calculations and generate an inference result;

[0053] 6. The service instance returns the inference result to the proxy server;

[0054] 7. The proxy server returns the inference result to the client, and the processing of an inference request is completed.

[0055] When the proxy server selects an idle service instance, it can usually only roughly judge the load of the service instance based on the pre-set performance specifications (such as the number of requests that can be processed per second) and the number of requests currently being processed. However, during the actual operation of the service instance, its processing capacity will change, which may be lower or higher than the pre-configured performance specifications. Therefore, the proxy server cannot accurately know the actual capacity and pressure of each service instance, and therefore cannot achieve accurate load balancing, resulting in uneven distribution of requests. If the request is forwarded to a fully loaded service instance, it will cause the service instance to be overloaded, while other instances will be unloaded, resulting in reduced processing efficiency.

[0056] Moreover, the proxy server needs to obtain the address of the service instance from the service registry generated by the service manager to monitor the service instance and create and delete the service instance. However, due to factors such as network synchronization delay or failure, or the inference service failing to update its own information in time, the information in the service registry generated by the service manager may be inaccurate or non-real-time. In this case, the service instance may fail to request due to the use of the wrong address returned by the proxy server.

[0057] In addition, the proxy server processes inference requests in a synchronous manner, that is, after sending an inference request to the inference service, it needs to wait for the service instance to complete the processing and return the inference result, and then start the response to the second inference request after returning the inference result to the client. This request processing mechanism results in a long processing time when processing multiple inference requests; and if the network fails or the inference service fails, the inference request will fail and the next inference request cannot be processed, which has poor fault tolerance.

[0058] In order to solve the problems of uneven distribution of inference requests, inaccurate service instance information, long processing time, and poor fault tolerance in traditional methods, this paper proposes an inference service method, which adopts a subscription and publishing method based on a message bus. Please refer to Figure 1 , Figure 1 The signaling diagram of a reasoning service method in an embodiment of the present invention includes the following steps:

[0059] S110, the client receives the inference request initiated by the user and sends the inference request to the message bus;

[0060] The inference request is initiated by the user on the client side. The inference request contains the data objects (such as text, images, data) that the user needs to call the model for inference calculation and the type of model to be called. Of course, other types of information can also be included. In this embodiment, the information included in the inference request is not limited, as long as it is sufficient to instruct the service instance to complete the inference service.

[0061] After the client receives the inference request initiated by the user, it sends it to the message bus. Combining the characteristics of the model inference service (stateless, large fluctuations in the request volume) and the advantages of the message bus (easy scalability, high fault tolerance) can effectively meet the dynamic changes of the model inference service, improve the resource utilization rate of the model service, and at the same time enhance the robustness of the model inference service. Among them, the process of sending the request can refer to the implementation of related technologies and will not be elaborated here.

[0062] S120. After the message bus receives the inference request sent by the client, it determines the service type of the inference request;

[0063] The message bus can receive inference requests from any client. The message bus contains services of multiple message queues, and each message queue corresponds to a topic. The producer (i.e., the client) publishes the inference request to the message queue of a certain topic. The consumer (i.e., the inference service) subscribes to the message queue of a certain topic. When there is a request in the topic queue, it obtains the request and processes it.

[0064] Specifically, after receiving the inference request sent by a certain client, determine the type of service that the current inference request is going to request according to the information in the inference request, that is, the type of the model or the type of the service instance. Determining the service type of the inference request can be further analyzed and determined by the type of model to be called, or a corresponding service type field can be set in the inference request, and then the content of this field can be directly read. The determination method of the service type is not limited in this embodiment and can be set according to actual usage needs. Among them, the configuration of the service type needs to be set corresponding to the topic of the message queue in the message bus, so that the topic corresponding to the service type can be matched according to the service type, and then a certain message queue can be located. It should be noted that generally, the service type and the topic of the message queue can be set one-to-one. Of course, there can also be more than one topic (or message queue) matching the service type, and there can also be more than one service type matching the topic (or message queue). Specifically, the matching relationship between the service type and the topic can be set according to the actual model service call requirements and will not be elaborated here.

[0065] S121. The message bus adds the inference request to the message queue corresponding to the topic and the service type;

[0066] After matching the service type that can satisfy the current inference request, add the inference request to the message queue corresponding to the matched service type. For example, if the matched topic is image feature extraction, and the message queue with image feature extraction as the topic is Message Queue 1, then add the inference request to Message Queue 1. Generally, adding information to the message queue follows the first-in, first-out rule. After adding the inference request to the message queue, if there are other unprocessed inference requests stored in the message queue, the current inference request is placed at the end of the queue and will be processed after other requests are completed. This mechanism can ensure the average processing time of each inference request and avoid the problem of excessive waiting time.

[0067] S122. The message bus sends a new request notification to the service instance subscribing to the message queue.

[0068] If the service instance establishes a subscription relationship with the message queue corresponding to the topic, after a new request is deposited in the message queue subscribed by the service instance, a new request notification will be sent to the subscribed service instance. The new request notification indicates that there is a new inference request in the message queue, but there may not be only new inference requests, and historical unprocessed inference requests may also be arranged among them. Then, each idle service instance takes the inference request in the order of the requests in the queue for processing in turn.

[0069] S130. The service instance receives the new request notification sent by the message queue in the subscribed message bus, and determines whether it can undertake the inference request based on its own load and service availability.

[0070] After receiving the new request notification, the service instance subscribing to the topic (topic) will determine whether to undertake the inference request according to its actual operating state. The specific operating state is mainly judged from two aspects: its own load and service availability.

[0071] S131. If the service instance can undertake it, read the inference request from the message queue and perform inference processing.

[0072] If the inference service of the service instance for this inference request is available, it indicates that the service instance has the ability to undertake this service request. Further, if the service instance is in an idle state (including no pending tasks and processing capacity exceeding the current pending task volume), that is, its own load is low, then the service instance can determine that it can undertake this inference request.

[0073] Conversely, if the inference service of the service instance subscribing to the topic is unavailable for the current inference request, or if the instance itself has a high load (its processing capacity has reached beyond the current volume of tasks to be processed), then it can choose not to undertake the current inference request. And if all inference services are busy and unable to undertake the inference request, the number of requests for this topic will pile up, and the processing speed of this queue will slow down.

[0074] In this method, the allocation of inference requests changes the traditional proxy service allocation method to an active selection by the inference service. Under this method, the inference service can obtain and process inference requests from the message bus according to its actual performance, with uniform traffic allocation and no problems of overload and idleness.

[0075] It should be noted that after the service instance completes the inference process and obtains the inference result, the inference result will also be fed back to the message bus and then returned to the client by the message bus. The process of result return can refer to the process of request sending and will not be elaborated here.

[0076] Based on the above introduction, the technical solution provided by the embodiments of the present invention combines the characteristics of stateless model inference services with large fluctuations in the number of requests and the advantages of easy scalability and high fault tolerance of the message bus. It changes the allocation of inference requests from an active selection by the proxy service to an active selection by the inference service. After receiving the inference request sent by the client, the message bus puts it into the message queue and sends a new request notification to the service instances subscribing to this message queue, indicating that a new request has arrived at the message queue. After receiving the new request notification, the service instance can determine whether to undertake the request according to its actual performance, including the load situation and availability. If it can undertake the request, it will obtain the inference request from the message bus and process it. During the processing of this request, the service instance undertakes the request according to its actual performance, without problems of overload and idleness caused by not understanding the actual performance, nor problems of undertaking requests when unavailable. The success rate of request processing is high; moreover, when the inference request is sent to the message bus, even if the inference service cannot receive the request due to a network failure, the request can continue to be processed after the network is restored, with high fault tolerance; at the same time, each service instance can undertake and process requests simultaneously, significantly improving the request processing efficiency.

[0077] It should be noted that based on the above embodiments, the embodiments of the present invention also provide corresponding improvement solutions. In the preferred / improved embodiments, the same steps or corresponding steps involved in the above embodiments can be referred to each other, and the corresponding beneficial effects can also be referred to each other, and will not be elaborated one by one in the preferred / improved embodiments of this article.

[0078] On the basis of the above embodiment, in order to further improve the standardization of each service instance in taking over tasks from the message queue and avoid the waste of computing resources caused by multiple service instances processing one inference request at the same time, after sending a new request notification to the service instance subscribed to the message queue, the following steps can be further performed:

[0079] (1) After receiving the request processing notification sent by the service instance, determine the request to be processed as the target request;

[0080] If the service instance determines that it can accept the inference request, it sends a request processing notification to the message bus, instructing service instance A to process inference request B. After receiving the request processing notification, the message bus determines the request (inference request B) to be processed by the sender of the request processing notification as the target request.

[0081] (2) Add a file lock to the target request;

[0082] To avoid problems such as resource waste caused by multiple idle service instances processing an inference request at the same time, a file lock can be added to the target request immediately after receiving the request processing notification, to ensure that the first service instance that initiates the request processing notification can process the inference request alone, while other idle service instances cannot process the inference request, thus ensuring the uniqueness of the inference request processing.

[0083] (3) After receiving the request processing completion notification, delete the target request.

[0084] After the service instance completes the processing of a certain reasoning request, it sends a request processing completion notification to the message bus. The request processing completion notification indicates that service instance A has completed the processing of the target request (reasoning request B). At this time, the target request can be deleted from the message queue to avoid the accumulation of service requests in the message queue.

[0085] Furthermore, in order to prevent the processing of the target request from falling into an infinite loop, and to improve the processing efficiency of the request and avoid requests piling up in the queue, after adding a file lock to the target request, the following steps can be further performed: if the processing of the target request is abnormal, unlock the file lock of the target request.

[0086] If an exception occurs in the processing of the target request, the file lock of the target request can be released. After the file lock of the target request is released, the target request can be re-accepted for processing by other service instances (other service instances also need to add a file lock to it after taking over), so as to accelerate the processing flow of the target request. Among them, the determination method of the target request processing exception is not limited in this embodiment. It can be determined by the service instance sending a processing exception notification to the message bus, or the message bus or other devices can monitor the processing process of the service request. If certain exceptions occur or the processing time exceeds the maximum threshold, it is determined as an exception. In this embodiment, only the above two exception determination methods are used as examples for introduction, and other exception determination methods can refer to the introduction of this embodiment and will not be elaborated here. Of course, the above steps may not be executed, which is not limited here.

[0087] In the above embodiments, each service instance can simultaneously obtain inference requests from the message queue and perform inference processing simultaneously. The asynchronous processing of inference requests among service instances can significantly improve the overall processing speed of inference requests. To further improve the processing speed of inference requests, the process of obtaining inference requests from the message queue and performing inference processing can specifically be: reading several inference requests from the message queue according to its own load and batch-processing each inference request simultaneously.

[0088] In the traditional method, the proxy service (according to the type) sends the inference request to the corresponding inference service, and (using the method of load balancing) sends the request to the inference service in a uniform manner (there are multiple instances of each type of inference service). The result of this method is that each inference service processes only one request at a time. However, the actual processing capacity of the inference service itself can process a batch (multiple) of requests simultaneously, and the time-consuming is almost the same as processing one request at a time. Therefore, in this embodiment, it is proposed that when inference requests accumulate in the queue, the service instance simultaneously obtains multiple inference requests in batches (batch) and processes them simultaneously, which can make full use of the model service resources, well alleviate the problem of occasional sudden increase in the request volume, and significantly improve the processing efficiency of the inference service.

[0089] Corresponding to the above method embodiment, the embodiment of the present invention also provides a message bus, and the message bus described below can be mutually corresponding and referred to with the inference service method described above.

[0090] As Figure 3 shown is a schematic diagram of a message bus provided in this embodiment. The message bus mainly includes several message queues for storing service requests, and each message bus has a unique topic, and the topic is used to indicate the service type of the stored inference requests.

[0091] Specifically, the message bus in this setting is specifically used for: after receiving an inference request sent by a client, determining the service type of the inference request; adding the inference request to a message queue corresponding to the service type; and sending a new request notification to a service instance subscribing to the message queue, so that the service instance can receive or reject the processing of the inference request according to its own load and service availability. The introduction of this part can refer to the introduction of the above method embodiments and will not be elaborated here.

[0092] Corresponding to the above method embodiments, an embodiment of the present invention further provides a computer device, which is mainly used to carry service instances. A computer device described below can be correspondingly referred to with an inference service method described above.

[0093] The computer device can specifically be a server, a computer, etc. The computer device includes:

[0094] A memory for storing computer programs;

[0095] A processor for implementing the steps of the inference service method in the above method embodiments when executing the computer programs.

[0096] Specifically, please refer to Figure 4 , which is a schematic structural diagram of a computer device provided in this embodiment. The computer device may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 322 (for example, one or more processors) and a memory 332. The memory 332 stores one or more computer application programs 342 or data 344. Among them, the memory 332 can be short-term storage or persistent storage. The programs stored in the memory 332 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the data processing device. Further, the central processor 322 can be set to communicate with the memory 332 and execute a series of instruction operations in the memory 332 on the computer device 301.

[0097] The computer device 301 may further include one or more power supplies 326, one or more wired or wireless network interfaces 350, one or more input / output interfaces 358, and / or one or more operating systems 341.

[0098] The steps in the above-described inference service method with a service instance as the execution entity can be implemented by the structure of the computer device provided in this embodiment.

[0099] Corresponding to the above device embodiment, an embodiment of the present invention further provides an inference service system. An inference service system described below can be correspondingly referred to the message bus and the computer device described above.

[0100] An inference service system specifically includes: a client, a message bus, and several service instances with different service types, such as Figure 5 is a schematic structural diagram of an inference service system.

[0101] Among them, the client is mainly used to interact with users, receive inference requests initiated by users, and send the inference requests to the message bus;

[0102] The message bus contains several message queues. After receiving an inference request, the message bus is used to determine the service type of the inference request; add the inference request to the message queue corresponding to the theme and service type; send a new request notification to the service instance subscribing to the message queue. For the specific structure and working process of the message bus, reference can be made to the introduction of the above message bus embodiment, which will not be elaborated here.

[0103] The service instance is used to receive the new request notification; judge whether it can undertake the inference request according to its own load and service availability; if it can undertake, obtain the inference request from the message queue and perform inference processing. The service instance is carried on a computer device. Specifically, the working process of the service instance and the structure of the carried device can also be referred to the introduction of the above method embodiment and computer device embodiment, which will not be elaborated here.

[0104] In the inference service system provided in this embodiment, the client obtains the user's inference request and sends it to the message bus. After receiving the inference request sent by the client, the message bus puts it into the message queue corresponding to the service type of the inference request and sends a new request notification to the service instance subscribing to the message queue, indicating that a new request has arrived at the message queue. After receiving the new request notification, the service instance can determine whether to undertake the request according to its actual performance, including the load situation and availability. If it can undertake, it obtains the inference request from the message bus and processes it. This process can be compared with the above method embodiment. By setting up an inference service system composed of a client, a message bus, and several service instances with different service types, the message bus can implement inference service processing dominated by service instances, ensuring the load balance and efficiency of requests.

[0105] In one embodiment, the inference service system may further include: a service manager connected to the message queue. The service manager is a device responsible for managing the creation, destruction, and information query of inference services according to certain external conditions (such as request volume monitoring data, queue queuing depth, etc.), such as Figure 6The figure shows a schematic structural diagram of an inference service system provided in this embodiment.

[0106] The service manager is connected to the message bus and each service instance, can monitor the request processing speed of each message queue in the message bus, generate a request processing monitoring record, and perform scaling processing on the service instance according to the request processing monitoring record.

[0107] Currently, under the traditional method, the service manager increases or decreases the number of inference services according to the request volume and request processing time of each inference service in the proxy service. However, due to the uneven distribution of inference requests in the proxy service, the average time-consuming of inference requests is relatively high, resulting in the service manager being unable to accurately judge the size of the request volume and unable to accurately and real-time scale (i.e., increase or decrease) the number of inference services. Moreover, the proxy service needs to obtain the address of the inference service through the service registry. However, due to factors such as network synchronization delay or failure, or the inference service failing to update its own information in a timely manner, the information in the service registry may be inaccurate or non-real-time. In such a case, the scaling cycle of the inference service instance is relatively long, unable to respond to the change of the inference request volume in real time, and may complete the expansion after the request volume has decreased, which has lost its meaning at this time.

[0108] However, in the service manager provided in this embodiment, according to the change trend of the number of pending requests on the message bus, the model service instances can be accurately and real-time elastically increased or decreased. When there is a trend of request accumulation, increase the number of instances of the model service to improve the processing ability. When the request processing speed speeds up, reduce the number of instances of the model service to reduce the processing ability. Since it is based on the change trend of the request volume, this can fully and truly reflect the real processing ability of each model service instance (including request volume and request processing time-consuming, etc.), so that the service manager can make accurate scaling. Moreover, when the service manager adds an inference instance, as long as the inference service starts up and completes, it can participate in the processing of inference requests, which reduces the time-consuming of registering information to the service manager compared with the proxy mode, reduces the time-consuming of synchronizing data table updates between the service manager and the proxy service, shortens the time-consuming cycle of elastic scaling, and the real-time performance of scaling is well improved.

[0109] For the implementation of the service instance number adjustment function of the service manager, this embodiment does not make a limitation. Optionally, the service manager can specifically be used to: determine the request processing speed of the target message queue according to the request processing monitoring record; if the request processing speed is lower than the first threshold; add service instances of the service type corresponding to the topic of the target message queue; if the request processing speed is higher than the second threshold; reduce service instances of the service type corresponding to the topic of the target message queue; where the first threshold is lower than the second threshold.

[0110] If all inference services are busy, the processing speed of the topic request slows down. In this case, the following scaling process will be carried out:

[0111] When the service manager notices that the processing speed of a certain topic queue slows down (or) requests pile up, the manager will dynamically increase the number of instances corresponding to the relevant inference service;

[0112] After the service instances increase, there will be more service instances to handle the requests of this topic.

[0113] If the inference services of a certain topic are mostly idle, or the request processing speed of the topic queue becomes faster, the following scaling-down process will be carried out:

[0114] When the service manager notices that the request processing of a certain topic queue becomes faster, the manager will dynamically reduce the number of instances corresponding to the relevant inference service;

[0115] After the service instances decrease, the idle time of the inference service can be reduced.

[0116] In this scaling process, the processing speed of the message queue can truly reflect the request processing situation and is easy to monitor, thus ensuring the real-time nature of scaling. Of course, service instance scaling can also be carried out through other monitoring means, which are not limited here.

[0117] Those skilled in the art can further realize that, combining the units and algorithm steps of each example described in the embodiments disclosed herein, they can be implemented by electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

Claims

1. A method for inference service, characterized in that, It includes: After the message bus receives an inference request sent by a client, it determines the service type of the inference request; Add the inference request to a message queue whose topic corresponds to the service type; Send a new request notification to the service instance subscribing to the message queue, so that the service instance can receive or reject the processing of the inference request according to its own load and service availability; Wherein, after sending the new request notification to the service instance subscribing to the message queue, the method further includes: After receiving a request processing notification sent by a service instance, determine the processed request as the target request; Add a file lock to the target request; After receiving a request processing completion notification, delete the target request; After adding the file lock to the target request, it further includes: If the processing of the target request is abnormal, unlock the file lock of the target request; Wherein, each service instance reads several inference requests from the message queue according to its own load and processes each inference request in batch simultaneously.

2. A message bus, characterized in that, It includes: Several message queues with topics for indicating service types; The message bus is used to: after receiving an inference request sent by a client, determine the service type of the inference request; add the inference request to a message queue whose topic corresponds to the service type; Send a new request notification to the service instance subscribing to the message queue, so that the service instance can receive or reject the processing of the inference request according to its own load and service availability; Wherein, after sending the new request notification to the service instance subscribing to the message queue, the method further includes: After receiving a request processing notification sent by a service instance, determine the processed request as the target request; Add a file lock to the target request; After receiving a request processing completion notification, delete the target request; After adding the file lock to the target request, it further includes: If the processing of the target request is abnormal, unlock the file lock of the target request; Wherein, each service instance reads several inference requests from the message queue according to its own load and processes each inference request in batch simultaneously.

3. A method for inference service, characterized in that, It includes: The service instance receives a new request notification sent by the message queue in the subscribed message bus; wherein, the new request notification is triggered after the message bus adds an inference request sent by a client to the message queue; the topic of the message queue corresponds to the inference service type of the service instance; Judge whether it can undertake the inference request according to its own load and service availability; If it can undertake, read the inference request from the message queue and perform inference processing; Wherein, after the service instance receives a new request notification sent by the message queue in the subscribed message bus, the method further includes: After receiving a request processing notification sent by a service instance, determine the processed request as the target request; Add a file lock to the target request; After receiving a request processing completion notification, delete the target request; After adding the file lock to the target request, it further includes: If the processing of the target request is abnormal, unlock the file lock of the target request; Obtaining the inference request from the message queue and performing inference processing includes: Reading a number of the inference requests from the message queue according to its own load, and batch-processing each of the inference requests simultaneously.

4. A computer device, characterized in that, Including: A memory for storing a computer program; A processor for implementing the steps of the inference service method as described in claim 1 or 3 when executing the computer program.

5. An inference service system, characterized in that, Including: A client, a message bus, and a number of service instances with different service types; Wherein, the client is used for receiving an inference request initiated by a user and sending the inference request to the message bus; The message bus contains a number of message queues. The message bus is used for, after receiving the inference request, determining the service type of the inference request; adding the inference request to the message queue corresponding to the service type in terms of topic; sending a new request notification to the service instances subscribing to the message queue; The service instance is used for receiving the new request notification; judging whether it can undertake the inference request according to its own load and service availability; if it can undertake, obtaining the inference request from the message queue and performing inference processing; Wherein, after sending the new request notification to the service instances subscribing to the message queue, the method further includes: After receiving a request processing notification sent by a service instance, determining the processed request as a target request; Adding a file lock to the target request; After receiving a request processing completion notification, deleting the target request. After adding the file lock to the target request, it further includes: If the processing of the target request is abnormal, unlocking the file lock of the target request; Wherein, each of the service instances reads a number of inference requests from the message queue according to its own load, and batch-processes each of the inference requests simultaneously.

6. The inference service system according to claim 5, wherein Further including: A service manager connected to the message queue; The service manager is used for the service manager to monitor the request processing speed of each message queue in the message bus, generate a request processing monitoring record, and perform scaling processing on the service instances according to the request processing monitoring record.

7. The inference service system according to claim 6, wherein The service manager is specifically used for: determining the request processing speed of a target message queue according to the request processing monitoring record; if the request processing speed is lower than a first threshold; adding a service instance of the service type corresponding to the topic of the target message queue; if the request processing speed is higher than a second threshold; reducing the service instance of the service type corresponding to the topic of the target message queue; wherein, the first threshold is lower than the second threshold.

Citation Information

Patent Citations

  • Transaction information push method and system

    CN107105064A

  • Service request distribution method and device, receiving method and device, and server cluster

    CN109660607A