A model switching method and apparatus, electronic device, storage medium, and product
Patent Information
- Application Number
- CN202610880904.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-17
- Publication Date
- 2026-09-15
Smart Images

Figure CN122756884A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a model switching method, apparatus, electronic device, storage medium, and product. Background Technology
[0002] Large-scale model inference services have become a core foundational service in the field of artificial intelligence. The high-efficiency virtual large language model (vLLM) inference engine, with its efficient key-value (KV) caching and scheduling capabilities, is widely used for online deployment of large models. vLLM is a high-performance inference service engine for large language models. It achieves efficient memory management through a pagination attention mechanism, offering advantages such as high throughput, low latency, and high concurrency inference, and has been widely applied in online deployment scenarios for large models. However, vLLM's compatibility with Ascend processors is generally limited, and some functions lack robustness and stability, failing to meet relevant business requirements. To meet the concurrent demands of multiple services, servers often need to deploy multiple large models of different types. However, the memory and computing power resources of hardware such as graphics processing units (GPUs) and Ascend AI chips are limited. Long-term online operation of a large number of models can lead to resource redundancy and high system power consumption.
[0003] Currently, mainstream model scheduling methods mostly manage model status through manual start / stop, overall restart, or direct forced offline. When a model is temporarily not accessed by any business, the model process is generally terminated to release resources; the model is then reloaded and the service is restarted when needed later. This approach has significant shortcomings. On the one hand, forcibly stopping a model directly interrupts the ongoing inference task, causing request loss, abnormal inference results, poor business continuity, and a high probability of model server crashes due to instruction conflicts or abnormal status. On the other hand, frequent full loading and unloading of models generates significant time consumption, high service response latency, low scheduling efficiency, and may also lead to communication domain conflicts, preventing the realization of multi-card task reuse of computing power cards. Furthermore, existing solutions lack a unified identification and branching mechanism for model running status, cannot distinguish between different operating conditions such as online and dormant models, and do not set up dedicated initialization procedures for newly launched models, resulting in insufficient system fault tolerance and operational stability.
[0004] Therefore, there is an urgent need for a model switching method that can achieve smooth model state switching, ensure normal execution of inference tasks, and improve hardware resource utilization and service stability, in order to solve the above-mentioned problems existing in the current technology. Summary of the Invention
[0005] This invention provides a model switching method, apparatus, electronic device, storage medium, and product to solve the problem that inference tasks are easily interrupted and business anomalies are caused during model switching in the prior art.
[0006] According to one aspect of the present invention, a model switching method is provided, wherein the method includes: In response to the target model initiation command sent by the user, determine the current running status of the current model; When the current running state is the first running state, a business flow reset operation is performed on the current model, the current model is switched to a sleep state, and a sleep success response is sent to the user terminal. Obtain the wake-up command generated by the user based on the successful sleep response, determine the target running state of the target model, perform a wake-up operation on the target model according to the target running state and the wake-up command, and perform inference operation through the target model.
[0007] According to another aspect of the present invention, a model switching device is provided, wherein the device comprises: The status determination module is used to determine the current running status of the current model in response to the target model initiation command sent by the user terminal. The sleep response module is used to perform a business flow reset operation on the current model when the current running state is the first running state, switch the current model to the sleep state, and send a sleep success response to the user terminal. The model switching module is used to obtain the wake-up command generated by the user terminal based on the successful sleep response, determine the target running state of the target model, perform a wake-up operation on the target model according to the target running state and the wake-up command, and perform inference operation through the target model.
[0008] According to another aspect of the present invention, an electronic device is provided, the electronic device comprising: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the model switching method according to any embodiment of the present invention.
[0009] According to another aspect of the present invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to execute and implement the model switching method described in any embodiment of the present invention.
[0010] According to another aspect of the present invention, embodiments of the present invention also provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the model switching method of any embodiment of the present invention.
[0011] The technical solution of this invention, in response to the target model initiation command sent by the user terminal, determines the current running state of the current model. When the current running state is the first running state, a business flow reset operation is performed on the current model, switching the current model to a sleep state, and a sleep success response is sent back to the user terminal. This ensures that the existing inference tasks are executed in an orderly manner, avoiding the problems of task interruption and request anomalies caused by forced service shutdown, and effectively ensuring the continuity of business operation. By obtaining the wake-up command generated by the user terminal based on the sleep success response, the target running state of the target model is determined. A wake-up operation is performed on the target model according to the target running state and the wake-up command, and inference operations are performed through the target model. This achieves seamless switching between multiple models, ensuring the continuity and response efficiency of the inference service.
[0012] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of the present invention, nor is it intended to limit the scope of the invention. Other features of the invention will become readily apparent from the following description. Attached Figure Description
[0013] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0014] Figure 1 This is a flowchart of a model switching method provided in Embodiment 1 of the present invention; Figure 2 This is a flowchart of a model switching method provided in Embodiment 2 of the present invention; Figure 3 This is a flowchart of a model initiation method provided in Embodiment 3 of the present invention; Figure 4 This is a flowchart of a model switching method provided in Embodiment 3 of the present invention; Figure 5 This is a schematic diagram of the structure of a model switching device according to Embodiment 4 of the present invention; Figure 6 This is a schematic diagram of the structure of an electronic device that implements the model switching method of this invention. Detailed Implementation
[0015] To enable those skilled in the art to better understand the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.
[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0017] Example 1 Figure 1 This is a flowchart of a model switching method according to Embodiment 1 of the present invention. This embodiment is applicable to model sleep and wake-up control based on the vLLM framework. The method can be executed by a model switching device, which can be implemented in hardware and / or software and can be configured in an electronic device, such as a server. Figure 1 As shown, the method includes: S110. In response to the target model initiation command sent by the user terminal, determine the current running status of the current model.
[0018] In this context, "user client" refers to the client that calls the model inference service. For example, the user client can be a WorldWide Web (Web) application, a client device, or a business front-end system. Generally, the user client can send model switching requests and inference requests to the server. The target model initiation command can be understood as a scheduling command issued by the user client, used to request state switching and business scheduling for a specified model. In actual operation, the target model initiation command can be sent via Hypertext Transfer Protocol (HTTP). Generally, the target model initiation command can include the name or identifier (ID) of the target model for easy identification. "Current model" refers to the model that the server is currently processing for inference requests. "Current running state" refers to the current running state of the model; in actual operation, the current running state can include online and sleep states.
[0019] In this embodiment, the client can send a target model initiation command to the server. The server can receive the target model initiation command and query the current running status of the current model. In actual operation, the server can query the current model's status record to determine the current running status as the current running status. In one embodiment, the current running status can be read from the status dictionary maintained by the server; alternatively, the current running status can be queried by calling the status query interface provided by the vLLM engine. In one embodiment, before determining the current running status of the current model in response to the target model initiation command sent by the client, the server can add an additional startup parameter: enable-sleep-mode, to the regular vLLM start command. Additionally, the following environment variables need to be added: VLLM_SERVER_DEV_MODE=1, to enable the development endpoint; VLLM_WORKER_MULTIPROC_METHOD="spawn", to improve system stability and avoid deadlocks; and VLLM_ASCEND_ENABLE_NZ="0", to avoid incompatibility with some operators.
[0020] S120. When the current running state is the first running state, perform a business flow reset operation on the current model, switch the current model to sleep state, and send a sleep success response to the user.
[0021] The first running state refers to the online state where the current model is in normal working mode and can continuously receive and process external inference requests. The business flow reset operation refers to the operation of orderly terminating existing inference tasks and clearing the interaction context before the current model enters the sleep state. Generally, the business flow reset operation can include pausing and resuming. The sleep state can be understood as a low-power state where the model has released memory resources, stopped receiving and processing inference requests, suspended providing inference services, and entered a low-resource-occupancy state. The sleep success response can be understood as the confirmation information returned by the server to the user after the current model successfully enters the sleep state.
[0022] In one embodiment, when the current running state is the first running state, i.e., the online state, the server can send a pause request to the current model, controlling the current model to stop receiving new inference requests and only continue processing the already received pending inference tasks. During this period, the server will not assign any new inference tasks to the current model. The server continuously monitors the current model's task queue and waits until all pending inference tasks for the current model have been completed before sending a resume request to the current model to complete the business flow reset operation. The server sends a sleep request to the current model through a preset service link, calling the vLLM sleep mode interface to switch the current model from the first running state to the sleep state. After the model successfully enters the sleep state, the server generates a sleep success response and pushes this response message to the user terminal, informing the user terminal that the model has completed the hibernation switch and can initiate subsequent wake-up operations. In one embodiment, the server resets the current model's business flow by calling the target interface of the vLLM engine, clearing the internal request queue and cache state, restoring the model to its initial state, and completing the current model's business flow reset operation.
[0023] S130. Obtain the wake-up command generated by the user terminal based on the successful sleep response, determine the target running state of the target model, perform a wake-up operation on the target model according to the target running state and the wake-up command, and perform inference operation through the target model.
[0024] The wake-up command can be understood as a request sent by the user to the server to wake up the target model after receiving a successful sleep response. In practice, the wake-up command can be sent via HTTP, and it may include the identifier of the target model to be woken up. The target model refers to the model that the user wants to switch to, i.e., the model that needs to be woken up. The target running status refers to the running status of the target model. Generally, the target model may currently be in a sleep state or may already be online.
[0025] In this embodiment, after receiving a wake-up command, the server can parse the command, determine the target model corresponding to the wake-up command, and query the running status of the target model. If the target model is online, inference operations can be performed directly through the target model; if the target model is in a sleep state, it can be woken up by the wake-up command, switching its running status to online to perform inference operations. Once the target model is online, the client can send an inference request to the target model's service link. The target model receives the request, performs inference calculations, and returns the inference results to the client.
[0026] In this embodiment of the invention, by responding to the target model initiation command sent by the user terminal, the current running state of the current model is determined. When the current running state is the first running state, a business flow reset operation is performed on the current model, switching the current model to a sleep state, and a sleep success response is sent back to the user terminal. This ensures that the existing inference tasks are executed in an orderly manner, avoiding the problems of task interruption and request anomalies caused by forced service shutdown, and effectively ensuring the continuity of business operation. By obtaining the wake-up command generated by the user terminal based on the sleep success response, the target running state of the target model is determined. A wake-up operation is performed on the target model according to the target running state and the wake-up command, and inference operations are performed through the target model. This achieves seamless switching between multiple models, ensuring the continuity and response efficiency of the inference service.
[0027] In one embodiment, before determining the current running state of the current model, the method further includes: Check if the target model corresponding to the command initiated by the target model exists in the preset service list; If the model does not exist, the query results will be sent back to the user so that the user can select a new target model.
[0028] The default service list refers to a list of deployed models maintained by the server, recording the names of all available models for switching and invocation, along with their corresponding service links. Generally, the default service list can be initialized when the service starts and dynamically updated when a new model is deployed.
[0029] In this embodiment, after receiving the instruction from the target model, the server can parse the target model name and query whether the target model exists in the preset service list. If it exists, the server determines that the target model can be switched. If it does not exist, the server generates a query result indicating that the target model does not exist in the preset service list and sends the query result back to the user to prompt the user to reselect the target model. This proactively intercepts invalid requests, avoids invalid operations and resource waste, and improves system efficiency and user experience.
[0030] In one embodiment, after obtaining the wake-up command generated by the user terminal based on the successful sleep response, determining the target running state of the target model, and performing a wake-up operation on the target model according to the target running state and the wake-up command, the method further includes: Determine if the target model is being used for the first time. If so, perform sleep and wake-up operations on the target model in sequence to complete the initialization of the target model's state.
[0031] In this embodiment, after performing a wake-up operation on the target model, the activation record of the target model can be queried. If the target model is being woken up and used for the first time after deployment on the server, the server can send a sleep request to the target model and call the vLLM engine's interface to perform a sleep operation on the target model. Generally, the sleep operation can release GPU memory resources, release communication domain resources, and update the state to sleep. After the sleep operation is completed, the server immediately sends a wake-up request to the target model, reloads the model weights to GPU memory, resumes the inference process, and re-requests communication domain resources. Generally, the server can update the target model's activation record to "initialized" or increment the activation count by 1 to avoid repeating the initialization operation next time.
[0032] In one embodiment, after determining the current running state of the current model in response to the target model initiation command sent by the user terminal, the method further includes: When the current running state is the second running state, the system receives the wake-up command from the user terminal, determines the target running state of the target model, performs a wake-up operation on the target model based on the target running state and the wake-up command, and performs inference operations through the target model.
[0033] The second running state refers to the running state where the current model is in a sleep state. In this state, the model has released GPU memory resources, does not process any inference requests, and is in a state of waiting to be woken up.
[0034] In this embodiment, when the current running state is the second running state, that is, there is no model currently in an online state, the wake-up command from the user terminal can be received directly to determine the target running state of the target model. When the target running state is a sleep state, the wake-up command is sent to the target model to switch the target model to an online state. When the target running state is an online state, inference operations are performed through the target model, avoiding repetitive operations and improving model switching efficiency.
[0035] Example 2 Figure 2 This is a flowchart of a model switching method according to Embodiment 2 of the present invention. This embodiment is a further optimization and extension based on the above embodiments, and can be combined with various optional technical solutions in the above embodiments. Figure 2 As shown, the method includes: S210. In response to the target model initiation command sent by the user terminal, determine the current running status of the current model.
[0036] S220. When the current running state is the first running state, send a pause request to the current model through a preset service link so that the current model stops receiving inference requests.
[0037] Here, the preset service link refers to the fixed service address pre-assigned by the server to each deployed model. The pause request refers to the instruction sent by the server to the current model to stop receiving new inference requests; generally, this can be achieved by calling the pause interface.
[0038] In one embodiment, when the server confirms that the current model is in the first running state, it can obtain the preset service link of the current model, construct a pause request through the preset service link, and send it to the current model. Upon receiving the pause request, the current model can stop receiving new inference requests, continue processing requests already in progress, and wait for all ongoing requests to complete before entering a pause state. In one embodiment, the pause request can carry the parameter `wait_for_inflight_requests=True`, which requires the current model to wait for all currently processed requests to complete before confirming the pause.
[0039] S230. After the pending inference task belonging to the current model is completed, a recovery request is sent to the current model through a preset service link to complete the business flow reset operation of the current model.
[0040] In this context, "pending inference tasks" refers to inference requests that have been received by the current model but have not yet been completed, or are currently being executed. A "resume request" is a command sent by the server to the model to resume receiving inference requests. Generally, this can be achieved by calling the `resume` interface.
[0041] In this embodiment, once it is confirmed that all pending inference tasks belonging to the current model have been completed, a recovery request can be constructed through a preset service link and sent to the current model. Upon receiving the recovery request, the current model can clear its internal request queue and restore its internal state to the initial baseline state. In one embodiment, the recovery request can be sent via HTTP.
[0042] S240. Send a sleep request to the current model through a preset service link, switch the current running state of the current model to sleep state, generate a sleep success response, and send the sleep success response to the user terminal.
[0043] In this context, a sleep request refers to a command sent by the server to the model to switch the model from a running state to a sleep state. In the vLLM engine, a sleep request is implemented by calling the sleep interface.
[0044] In this embodiment, a preset service link for the current model can be obtained, a sleep request can be constructed, and the sleep request can be sent to the current model. After receiving the sleep request, the current model can perform operations such as releasing video memory resources, suspending the inference process, and releasing communication domain resources, switching its current running state to a sleep state. When the current model successfully enters sleep mode, a sleep success response can be generated and sent to the user terminal.
[0045] S250. Send a sleep request to the current model through a preset service link, switch the current running state of the current model to sleep state, generate a sleep success response, and send the sleep success response to the user terminal.
[0046] In this embodiment, a wake-up command generated by the user terminal based on a successful sleep response can be received, and the target running status of the target model can be read from the state dictionary maintained by the server; alternatively, the target running status of the target model can be queried by calling the state query interface provided by the vLLM engine.
[0047] S260. When the target is in a sleep state, a wake-up command is sent to the target model to switch the target model to an online state.
[0048] The online state refers to the state in which the model is running and can normally receive and process inference requests.
[0049] In this embodiment, when the target is in a sleep state, the target model can be woken up by a wake-up command, and the target model's target running state can be adjusted to an online state.
[0050] S270. When the target is in an online state, perform inference operations through the target model.
[0051] In this embodiment, if the target is in an online state, inference operations can be performed directly through the target model.
[0052] In this embodiment of the invention, by responding to a target model initiation command sent by the user terminal, the current running state of the current model is determined. When the current running state is the first running state, a pause request is sent to the current model through a preset service link to cause the current model to pause receiving inference requests. After the pending inference tasks belonging to the current model are completed, a recovery request is sent to the current model through the preset service link to complete the business flow reset operation of the current model, clearing the residual state of the model, preventing state conflicts, and improving operational reliability. A sleep request is sent to the current model through the preset service link to switch the current running state of the current model to a sleep state and generate a sleep success response, which is sent to the user terminal. The user terminal sends a wake-up command, queries the running state of the target model as the target running state, and when the target running state is a sleep state, a wake-up command is sent to the target model to switch the target model to an online state. When the target running state is online, inference operations are performed through the target model, realizing rapid model recovery and seamless switching, and ensuring the continuity of inference services.
[0053] Example 3 Figure 3 This is a flowchart of a model startup method according to Embodiment 3 of the present invention. This embodiment takes the Ascend processor as an example, based on the vLLM native application programming interface (API) and the Huawei Collective Communication Library (HCCL) management interface, to further illustrate model switching. Figure 3 As shown, the model initiation phase includes: Step 1: Start the model service: On the server side, in addition to the standard vLLM start command, add the startup parameter: `--enable-sleep-mode`. Also, add the following environment variables: `VLLM_SERVER_DEV_MODE=1` to enable the development endpoint; `VLLM_WORKER_MULTIPROC_METHOD="spawn"` to improve system stability and avoid deadlocks; and `VLLM_ASCEND_ENABLE_NZ="0"` to avoid incompatibility issues with some operators.
[0054] Step 2: The client sends a sleep request to the model: After the server successfully restarts the model, the client executes `curl -XPOST`. <url>` / sleep?level=1`, where `url` is the service link set when the model starts.
[0055] Step 3: Determine if the model is in its first sleep: Perform a sleep-wake process on all deployed models before service-oriented architecture to ensure that the execution time of related operations is shortened after formal service-oriented architecture is implemented.
[0056] Step 4: Send a model start-up request from the client: If the model is sleeping for the first time, execute curl -XPOST on the client. <url>' / wake_up', the URL is the service link set when the model starts.
[0057] Step 5: Determine whether to continue starting other models: The current model has completed two sleep cycles and one wake-up (currently in a sleep state). If you need to continue starting other models, return to step 1; otherwise, exit directly to complete the entire server-side process.
[0058] In one embodiment, Figure 4 This is a flowchart of a model switching method provided in Embodiment 3 of the present invention, as shown below. Figure 4 As shown, taking the model selection process as the user's choice of the target model as an example, the model inference service switching phase includes: S1. User Model Selection: The user selects the model to be used.
[0059] S2. Determine if the target model is in the service list (sleep or start): Traverse the model service list and check if the model is in the list. If not, please select the model again.
[0060] S3. Determine if the model is sleeping: If the model is not sleeping, you can skip directly to S7 and use the model to process requests normally.
[0061] S4. Pause Current Model: Pause the currently active model. The user executes `curl -X POST '` on the client side. <url>` / pause?wait_for_inflight_requests=True`, where `url` is the service connection set when the model starts. After the command succeeds, execute `curl -X POST '`. <url>' / resume', the URL is the service link set when the model starts.
[0062] S5. The current model is sleeping: The user executes curl -X POST ' <url>` / sleep?level=1`, where `url` is the service link set when the current model starts.
[0063] S6. Launch the demand model: Execute `curl -X POST` on the user's end. <url>' / wake_up', where url is the service link set when the model starts.
[0064] S7. Model accepts and processes requests: The requirement model is in normal status, and users can send requests from the user end to the server normally.
[0065] Before the client sends the model sleep or wake-up command, this invention provides a way to query the model status (is_sleeping) and re-encode the model business flow (pause) through the native vLLM interface. This ensures that there are no abnormalities on the model server when using vLLM to sleep or wake up, thus guaranteeing the stability of the entire system.
[0066] Example 4 Figure 5 This is a schematic diagram of a model switching device according to Embodiment 4 of the present invention. Figure 5 As shown, the device includes: a state determination module 51, a sleep response module 52, and a model switching module 53.
[0067] The status determination module 51 is used to determine the current running status of the current model in response to the target model initiation command sent by the user terminal.
[0068] The sleep response module 52 is used to perform a business flow reset operation on the current model when the current running state is the first running state, switch the current model to the sleep state, and send a sleep success response to the user terminal.
[0069] The model switching module 53 is used to obtain the wake-up command generated by the user terminal based on the successful sleep response, determine the target running state of the target model, perform wake-up operation on the target model according to the target running state and the wake-up command, and perform inference operation through the target model.
[0070] The technical solution of this invention involves a state determination module responding to a target model initiation command sent by the user terminal to determine the current running state of the current model. When the current running state is the first running state, the sleep response module performs a business flow reset operation on the current model, switches the current model to a sleep state, and sends a sleep success response back to the user terminal. This ensures that existing inference tasks are executed in an orderly manner, avoids task interruption and request anomalies caused by forced service shutdown, and effectively guarantees the continuity of business operations. The model switching module obtains the wake-up command generated by the user terminal based on the sleep success response, determines the target running state of the target model, performs a wake-up operation on the target model according to the target running state and the wake-up command, and performs inference operations through the target model. This achieves seamless switching between multiple models, ensuring the continuity and response efficiency of the inference service.
[0071] In one embodiment, the sleep response module 52 includes: The model pause unit is used to send a pause request to the current model via a preset service link, so that the current model stops receiving inference requests. The model recovery unit is used to send a recovery request to the current model through a preset service link after the pending inference task belonging to the current model has been completed, so as to complete the business flow reset operation of the current model. The sleep response unit is used to send a sleep request to the current model through a preset service link, switch the current running state of the current model to a sleep state, generate a sleep success response, and send the sleep success response to the user terminal.
[0072] In one embodiment, the model switching module 53 includes: The status query unit is used to receive the wake-up command sent by the user terminal and query the running status of the target model as the target running status. The model wake-up unit is used to send a wake-up command to the target model when the target is in a sleep state, so that the target model switches to an online state. The model application unit is used to perform inference operations through the target model when the target is in an online state.
[0073] In one embodiment, before determining the current running state of the current model, the method further includes: The model query module is used to query whether the target model corresponding to the command initiated by the target model exists in the preset service list; The result feedback module is used to return the query result to the user if the model does not exist, so that the user can select a different target model.
[0074] In one embodiment, after obtaining the wake-up command generated by the user terminal based on the successful sleep response, determining the target running state of the target model, and performing a wake-up operation on the target model according to the target running state and the wake-up command, the method further includes: The model initialization module is used to determine whether the target model is being used for the first time. If so, it performs sleep and wake-up operations on the target model in sequence to complete the state initialization of the target model.
[0075] In one embodiment, the model switching device further includes: The model inference module is used to receive the wake-up command from the user terminal when the current running state is the second running state, determine the target running state of the target model, perform a wake-up operation on the target model according to the target running state and the wake-up command, and perform inference operation through the target model.
[0076] The model switching device provided in this embodiment of the invention can execute the model switching method provided in any embodiment of the invention, and has the corresponding functional modules and beneficial effects of the method execution.
[0077] Example 5 Figure 6 This is a schematic diagram of the structure of an electronic device implementing the model switching method of an embodiment of the present invention. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices (such as helmets, glasses, watches, etc.), and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the invention described and / or claimed herein.
[0078] like Figure 6 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12 or a random access memory (RAM) 13, communicatively connected to the at least one processor 11. The memory stores computer programs executable by the at least one processor. The processor 11 can perform various appropriate actions and processes based on the computer program stored in the ROM 12 or loaded from storage unit 18 into the RAM 13. The RAM 13 can also store various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.
[0079] Multiple components in electronic device 10 are connected to I / O interface 15, including: input unit 16, such as keyboard, mouse, etc.; output unit 17, such as various types of displays, speakers, etc.; storage unit 18, such as disk, optical disk, etc.; and communication unit 19, such as network card, modem, wireless transceiver, etc. Communication unit 19 allows electronic device 10 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.
[0080] Processor 11 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, digital signal processors (DSPs), and any suitable processor, controller, microcontroller, etc. Processor 11 performs the various methods and processes described above, such as model switching methods.
[0081] In some embodiments, the model switching method may be implemented as a computer program tangibly contained in a computer-readable storage medium, such as storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on electronic device 10 via ROM 12 and / or communication unit 19. When the computer program is loaded into RAM 13 and executed by processor 11, one or more steps of the model switching method described above may be performed. Alternatively, in other embodiments, processor 11 may be configured to perform the model switching method by any other suitable means (e.g., by means of firmware).
[0082] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.
[0083] Computer programs used to implement the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the processor, the computer programs cause the functions / operations specified in the flowcharts and / or block diagrams to be performed. The computer programs may be executed entirely on a machine, partially on a machine, or as a standalone software package, partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0084] In the context of this invention, a computer-readable storage medium can be a tangible medium that may contain or store a computer program for use by or in conjunction with an instruction execution system, apparatus, or device. A computer-readable storage medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination thereof. Alternatively, a computer-readable storage medium may be a machine-readable signal medium. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0085] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0086] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as data servers), or middleware components (e.g., application servers), or frontend components (e.g., user computers with graphical user interfaces or web browsers through which users can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (LANs), wide area networks (WANs), blockchain networks, and the Internet.
[0087] A computing system can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or cloud host, which is a hosting product within the cloud computing service system to address the shortcomings of traditional physical hosts and VPS services, such as high management difficulty and weak business scalability.
[0088] In one embodiment, the present invention further includes a computer program product, which includes a computer program that, when executed by a processor, implements the model switching method of any embodiment of the present invention.
[0089] In implementing the computer program product, computer program code for performing the operations of this invention can be written in one or more programming languages or a combination thereof. Programming languages include object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0090] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.
[0091] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.< / url> < / url> < / url> < / url> < / url> < / url>
Claims
1. A model switching method characterized by comprising: include: In response to the target model initiation command sent by the user, determine the current running status of the current model; When the current running state is the first running state, a business flow reset operation is performed on the current model, the current model is switched to a sleep state, and a sleep success response is sent to the user terminal. Obtain the wake-up command generated by the user based on the successful sleep response, determine the target running state of the target model, perform a wake-up operation on the target model according to the target running state and the wake-up command, and perform inference operation through the target model.
2. The method according to claim 1, characterized in that, The step of performing a business flow reset operation on the current model, switching the current model to a sleep state, and sending a successful sleep response to the user terminal includes: A pause request is sent to the current model via a preset service link to cause the current model to pause receiving inference requests. Once the pending inference task belonging to the current model is completed, a recovery request is sent to the current model through the preset service link to complete the business flow reset operation of the current model; A sleep request is sent to the current model via a preset service link to switch the current model's current running state to sleep state, and a sleep success response is generated and sent to the user terminal.
3. The method according to claim 1, characterized in that, The process of obtaining a wake-up command generated by the user terminal based on the successful sleep response, determining the target running state of the target model, performing a wake-up operation on the target model according to the target running state and the wake-up command, and performing inference operations through the target model includes: Receive the wake-up command sent by the user terminal and query the running status of the target model as the target running status; When the target is in a sleep state, the wake-up command is sent to the target model to switch the target model to an online state. When the target is in an online state, inference operations are performed through the target model.
4. The method according to claim 1, characterized in that, Before determining the current running state of the current model, the following steps are also included: Check if the target model corresponding to the command initiated by the target model exists in the preset service list; If the model does not exist, the query result will be sent back to the user so that the user can select a new target model.
5. The method according to claim 1, characterized in that, After obtaining the wake-up command generated by the user terminal based on the successful sleep response, determining the target operating state of the target model, and performing a wake-up operation on the target model according to the target operating state and the wake-up command, the method further includes: Determine whether the target model is being used for the first time. If so, perform sleep and wake-up operations on the target model in sequence to complete the state initialization of the target model.
6. The method according to claim 1, characterized in that, After determining the current running state of the current model in response to the target model initiation command sent by the user terminal, the method further includes: When the current running state is the second running state, a wake-up command from the user terminal is received, the target running state of the target model is determined, a wake-up operation is performed on the target model according to the target running state and the wake-up command, and an inference operation is performed through the target model.
7. A model switching device, characterized in that, include: The status determination module is used to determine the current running status of the current model in response to the target model initiation command sent by the user terminal. The sleep response module is used to perform a business flow reset operation on the current model when the current running state is the first running state, switch the current model to the sleep state, and send a sleep success response to the user terminal. The model switching module is used to obtain the wake-up command generated by the user terminal based on the successful sleep response, determine the target running state of the target model, perform a wake-up operation on the target model according to the target running state and the wake-up command, and perform inference operation through the target model.
8. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor, the computer program being executed by the at least one processor to enable the at least one processor to perform the model switching method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions that cause a processor to execute the model switching method of any one of claims 1-6.
10. A computer program product, characterized in that, The computer program product includes a computer program that, when executed by a processor, implements the model switching method according to any one of claims 1-6.