Task processing method and electronic equipment
By loading and running multiple processing models in large language models and isolating their use in different running spaces, the problem of question-and-answer or response chaos during multitasking concurrency is solved, and efficient task isolation and stability is achieved.
Patent Information
- Application Number
- CN202411999757.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
When multitasking concurrent access to the same large language model, the memory mechanism of the model may lead to questions that are confused by Q&A or response.
By responding to the task request, the remaining time for the first processing model to enter the idle state is obtained, and when the target condition is satisfied, at least one second processing model is loaded and run based on the first processing model, and the task request is processed using the second processing model. The second processing model and the first processing model are located in different operating spaces, realizing operation isolation and task isolation.
It effectively solves the problem of question-and-answer or response chaos caused by memory mechanism during multi-task concurrency, realizes efficient isolation of task processing, and improves the stability and response speed of the system.
Smart Images

Figure CN119938269A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to task processing technology, and more particularly to a task processing method and electronic equipment. Background Art
[0002] As large language models (LLM) are deployed on personal computers (PCs), the number of applications using LLM is increasing, resulting in multiple tasks accessing the same model concurrently. However, when multiple tasks access the same model concurrently, the model's memory mechanism may cause confusion in questions and answers. Summary of the invention
[0003] The present application provides a task processing method and an electronic device.
[0004] The technical solution of this application is implemented as follows:
[0005] In a first aspect, a task processing method is provided, the method comprising:
[0006] In response to obtaining at least one task request, obtaining a remaining time for the first processing model to enter an idle state;
[0007] When the remaining time meets the target condition, loading and running at least one second processing model based on the first processing model;
[0008] Processing the task request using the at least one second processing model;
[0009] The operating space where the second processing model is located is different from the operating space where the first processing model is located.
[0010] In a second aspect, a task processing device is provided, the device comprising:
[0011] an obtaining unit, configured to obtain, in response to obtaining at least one task request, a remaining time for the first processing model to enter an idle state;
[0012] A processing unit, configured to load and run at least one second processing model based on the first processing model when the remaining time meets a target condition;
[0013] a processing unit, configured to process the task request using the at least one second processing model;
[0014] The operating space where the second processing model is located is different from the operating space where the first processing model is located.
[0015] In a third aspect, an electronic device is provided, including a processor and at least one first processing model capable of running on the processor, wherein the first processing model can be called by a target application to perform at least one of the following:
[0016] In response to obtaining at least one task request, obtaining a remaining time for the first processing model to enter an idle state;
[0017] When the remaining time meets the target condition, loading and running at least one second processing model based on the first processing model;
[0018] Processing the task request using the at least one second processing model;
[0019] The operating space where the second processing model is located is different from the operating space where the first processing model is located.
[0020] According to a fourth aspect, a computer-readable storage medium is provided, on which a computer program is stored, wherein the computer program implements the steps of the method according to the first aspect when executed by a processor.
[0021] According to a fifth aspect, a computer program product is provided, comprising a computer program, wherein the computer program implements the steps of the method according to the first aspect when executed by a processor.
[0022] The present application provides a task processing method and electronic device, including: in response to obtaining at least one task request, obtaining the remaining time for a first processing model to enter an idle state; when the remaining time meets the target condition, loading and running at least one second processing model based on the first processing model; processing the task request using at least one second processing model; wherein the operating space where the second processing model is located is different from the operating space where the first processing model is located. In this way, the second processing model and the first processing model are respectively located in different operating spaces, so that the operation isolation is achieved between the first processing model and the second processing model, and further the task isolation is achieved between the task processed by the first processing model and the task processed by the second processing model, thereby solving the question and answer or response confusion problem caused by the memory mechanism when multiple tasks are concurrent. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 The process diagram of the task processing method in the embodiment of the present application is as follows: Figure 1 ;
[0024] Figure 2 The process diagram of the task processing method in the embodiment of the present application is as follows: Figure 2 ;
[0025] Figure 3 The process diagram of the task processing method in the embodiment of the present application is as follows: Figure 3 ;
[0026] Figure 4 The process diagram of the task processing method in the embodiment of the present application is as follows: Figure 4 ;
[0027] Figure 5 The process diagram of the task processing method in the embodiment of the present application is as follows: Figure 5 ;
[0028] Figure 6 A task processing block diagram of an example of an embodiment of the present application;
[0029] Figure 7 This is a schematic diagram of the structure of the task processing device in the embodiment of the present application;
[0030] Figure 8 It is a schematic diagram of the structure of the electronic device in the embodiment of the present application. DETAILED DESCRIPTION
[0031] In order to enable a more detailed understanding of the features and technical contents of the embodiments of the present application, the implementation of the embodiments of the present application is described in detail below in conjunction with the accompanying drawings. The attached drawings are for reference only and are not used to limit the embodiments of the present application.
[0032] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application belongs. The terms used herein are only for the purpose of describing this embodiment and are not intended to limit this application.
[0033] In the following description, references to “some embodiments,” “this embodiment,” “this embodiment,” and examples, etc., describe a subset of all possible embodiments, but it can be understood that “some embodiments” may be the same subset or different subsets of all possible embodiments, and may be combined with each other without conflict.
[0034] If similar descriptions of "first / second" appear in the application documents, the following instructions are added. In the following description, the terms "first\second\third" involved are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It can be understood that "first\second\third" can be interchanged in a specific order or sequence where permitted, so that the present embodiment described here can be implemented in an order other than that illustrated or described here.
[0035] In this embodiment, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, object A and / or object B may represent three situations: object A exists alone, object A and object B exist at the same time, and object B exists alone.
[0036] The present application embodiment provides a task processing method, Figure 1 The process diagram of the task processing method in the embodiment of the present application is as follows: Figure 1 , applied to electronic devices, including but not limited to smartphones, tablets and desktop computers. Figure 1 As shown, the task processing method includes the following steps S101 to S103:
[0037] Step S101: in response to obtaining at least one task request, obtaining a remaining time for a first processing model to enter an idle state.
[0038] Exemplarily, when at least one task request is obtained, the inference time of the first processing model is obtained, and the average processing time of the first processing model for processing the current number of historical task requests is obtained, and then based on the difference between the average processing time and the inference time, the remaining time required for the first processing model to complete the current task (i.e., enter the idle state) is obtained. The first processing model can be a large language model (LLM) for processing a variety of text processing tasks such as text generation, translation, summarization, question and answer, or image processing tasks or other processing tasks.
[0039] If there is no task being executed currently, the remaining time for the first processing model to enter the idle state is zero.
[0040] The task request may be a call request for the same model from at least one different application, or may be multiple call model requests from the same user or different users to the same application. Exemplarily, the task request includes but is not limited to: a request for text generation or processing, a request for image generation or processing, a request for video generation or processing, and a request for audio processing or processing.
[0041] Step S102: when the remaining time meets the target condition, at least one second processing model is loaded and run based on the first processing model.
[0042] In an embodiment of the present application, it is determined whether the remaining time for the first processing model to enter the idle state meets the target condition; if the remaining time meets the target condition, at least one virtual model or backup model, i.e., the second processing model, is initialized, started, and run based on the first processing model; if the remaining time does not meet the target condition, at least one task request is added to the pending task queue of the first processing model, and the first processing model is used to perform serial or parallel processing on at least one task request in the pending task queue. Among them, whether it is serial or parallel processing depends on the number of current first processing models. If the number of current first processing models is one, the first processing model is used to perform serial processing on at least one task request in the pending task queue. If the number of current first processing models is multiple, multiple first processing models are used to perform parallel processing on at least one task request in the pending task queue.
[0043] The target condition is used to determine whether to load and run at least one second processing model based on the first processing model.
[0044] The second processing model may be completely identical to the first processing model, or may be only a partial model of the first processing model, that is, a processing model created based on some operators and model weight configuration files of the first processing model.
[0045] Exemplarily, the number of the second processing models depends on the number of tasks in the task request, that is, the number of the second processing models is consistent with the number of tasks in the task request. Alternatively, the number of the second processing models is less than the number of tasks in the task request, which is because there is a top-down relationship between some task requests, that is, the second processing model serially processes at least one task request with a top-down relationship.
[0046] Step S103: Process the task request using at least one second processing model; wherein the operating space of the second processing model is different from the operating space of the first processing model.
[0047] In an embodiment of the present application, at least one second processing model is used to process at least one task request of the batch. At the same time, the first processing model is used to process the task being executed, that is, the first processing model and the second processing model process different task requests in parallel.
[0048] In an embodiment of the present application, the second processing model and the first processing model are located in different operating spaces, respectively, so that operation isolation is achieved between the first processing model and the second processing model, and further task isolation is achieved between the tasks processed by the first processing model and the tasks processed by the second processing model, thereby solving the question and answer confusion problem or response confusion problem caused by the memory mechanism during multi-tasking concurrency.
[0049] In some embodiments of the present application, when the remaining time meets the target condition, loading and running at least one second processing model based on the first processing model includes the following steps S201 to S202:
[0050] Step S201: Obtain a target reference duration, where the target reference duration is the duration required to load and run the second processing model.
[0051] In the embodiment of the present application, the target reference duration can be understood as the total duration required to independently start at least one second processing model.
[0052] Exemplarily, there is a corresponding relationship between the number of second processing models started and the reference duration, so the target reference duration can be determined based on the number of second processing models to be loaded and run here. That is, the target reference duration can be a fixed duration. Alternatively, the target reference duration is dynamically determined based on the task to be executed. Alternatively, the target reference duration is dynamically determined based on the computing resource configuration of the electronic device.
[0053] Step S202: when the remaining duration is greater than the target reference duration, at least one second processing model is loaded and run based on the model file of the first processing model, and the first processing model is the same as or different from the second processing model.
[0054] It should be noted that when the remaining time for the first processing model to enter the idle state is greater than the target reference time required to load and run at least one second processing model, it means that there is no need to wait for the first processing model to complete the task. At this time, it is necessary to load and run at least one second processing model based on the model file of the first processing model, and then use at least one second processing model to process at least one task request, thereby realizing rapid processing of at least one task request.
[0055] The second processing model is the same as the first processing model, which means that the model weights and configuration files of the second processing model are exactly the same as the model weights and configuration files of the first processing model. Alternatively, the second processing model is different from the first processing model, which means that the model weights and configuration files of the second processing model are only the same as part of the model weights and configuration files of the first processing model, and the rest are different. For different parts, the corresponding calculation or operation capabilities are different.
[0056] In some embodiments of the present application, obtaining the target reference duration includes at least one of the following:
[0057] Obtaining a processing time required by the first processing model to process the at least one task request, and processing the processing time based on a target weight coefficient to obtain the target reference time;
[0058] In the embodiment of the present application, it is assumed that if at least one task request is processed by the first processing model, the total processing time required for the first processing model to execute at least one task request is calculated. Further, the total processing time required to execute at least one task request is multiplied by the target weight coefficient to obtain the total time required to load and run at least one second processing model, that is, the target reference time. Among them, the target weight coefficient is a pre-set value, which can be set by the developer based on experience or experiment.
[0059] Alternatively, assuming that at least one task request is processed by the first processing model, the processing time required for the first processing model to execute each task request is calculated. Based on the processing time required for each task request and the corresponding target weight coefficient, a weighted sum is performed to obtain the total time required to load and run at least one second processing model, that is, the target reference time. Among them, the target weight coefficients corresponding to different task requests or second processing models may be the same or different.
[0060] Obtaining task information of the at least one task request, determining configuration information of the second processing model based on the task information, and determining the target reference duration based on the configuration information;
[0061] In the embodiment of the present application, the task information includes but is not limited to: the task content, the task complexity and the length of the instruction or information including the task request input to the second processing model. The configuration information includes but is not limited to: the configuration parameters and the complexity of the configuration parameters. Among them, the configuration parameters may include configuration files and model weights.
[0062] Exemplarily, the configuration parameters of at least one second processing model are determined based on the task content of at least one task request, and the target reference duration required to load and run the at least one second processing model is determined based on the configuration parameters of the at least one second processing model. Alternatively, the complexity of the configuration parameters of at least one second processing model is determined based on the task complexity of at least one task request, and the target reference duration required to load and run the at least one second processing model is determined based on the complexity of the configuration parameters of the at least one second processing model. The higher the complexity of the configuration parameters, the longer the corresponding target reference duration.
[0063] Reading the target reference duration from a preset configuration file;
[0064] Exemplarily, the preset configuration file may include a configuration table of models and reference durations for processing different tasks, then the reference duration corresponding to at least one second processing model is found from the configuration table, and the at least one reference duration is added to obtain the target reference duration. Alternatively, the preset configuration file may include a configuration table of processing model size and reference duration, and the processing model size is determined by the model weight and the configuration file, then the reference duration corresponding to the size of at least one second processing model is found from the configuration table, and the at least one reference duration is added to obtain the target reference duration.
[0065] The target reference duration is determined based on configuration data and usage data of a processor of the electronic device.
[0066] Exemplarily, there may be a relationship equation with the configuration data and usage data of the electronic device's processor as independent variables and the target reference duration as the dependent variable. By substituting the configuration data and usage data of the electronic device's processor into the relationship equation, the corresponding target reference duration can be obtained.
[0067] In some embodiments of the present application, obtaining the processing time required for the first processing model to process the at least one task request includes the following steps S301 to S303:
[0068] Step S301: obtaining the number of units to be processed obtained by identifying at least one task request.
[0069] In an embodiment of the present application, at least one task request is identified and processed (eg, word segmentation processing) to obtain a plurality of units to be processed, that is, a plurality of tokens, so that the number of the plurality of units to be processed, that is, the number of tokens, can be known.
[0070] Step S302: Obtain the processing capability of the first processing model.
[0071] In the embodiment of the present application, the processing capability of the first processing model includes the reasoning speed of the model, wherein the reasoning speed of the model includes but is not limited to: the reasoning speed when executing the last task request, and the average reasoning speed of executing multiple task requests within a period of time.
[0072] Step S303: Calculate the processing time based on the quantity and the processing capacity of the first processing model.
[0073] In an embodiment of the present application, the ratio of the number of to-be-processed units corresponding to at least one task request to the processing capacity of the first processing model is calculated to obtain the processing time required for the first processing model to process at least one task request.
[0074] Exemplarily, the ratio of the number of to-be-processed units corresponding to at least one task request to the inference speed of the first processing model is calculated to obtain the processing time required for the first processing model to process at least one task request.
[0075] In some embodiments of the present application, obtaining the remaining time for the first processing model to enter the idle state includes the following steps S401 to S403:
[0076] Step S401: In response to obtaining at least one task request, obtaining the number of tasks in the task request and the processing capacity of the first processing model.
[0077] In the embodiment of the present application, the number of tasks of at least one task request within T1 time is counted. The processing capacity of the first processing model includes but is not limited to: the number of tasks within the previous T1 time, and the average number of tasks within multiple consecutive T1 times.
[0078] After executing step S401, it is determined whether the number of tasks matches the processing capability of the first processing model; if so, step S402 is executed; if not, step S403 is executed.
[0079] Step S402: when the number of tasks does not match the processing capacity of the first processing model, determine to execute a step of obtaining a remaining time for the first processing model to enter an idle state.
[0080] Exemplarily, when the number of tasks exceeds the average number of tasks in a plurality of consecutive T times corresponding to the first processing model, it is determined that the processing capacity of the first processing model does not match, that is, the current state is a high concurrent task state. Further, a step of obtaining the remaining time for the first processing model to enter an idle state is performed.
[0081] Step S403: When the number of tasks matches the processing capability of the first processing model, at least one task request is added to the to-be-processed task queue of the first processing model.
[0082] Exemplarily, when the number of tasks does not exceed the average number of tasks within a plurality of consecutive T times corresponding to the first processing model, it is determined that the processing capacity matches the first processing model, that is, the current state is not a high concurrent task state. Further, at least one task request is added to the pending task queue of the first processing model, and the pending tasks in the pending task queue are subsequently processed serially or in parallel using the first processing model.
[0083] In some embodiments of the present application, loading and running at least one second processing model based on the first processing model includes at least one of the following:
[0084] Obtaining task information of the task request, determining a model weight and a configuration file matching the task information from a model file of the first processing model, and initializing a second processing model in a second operating space different from a first operating space where the first processing model is located by using the model weight and the configuration file;
[0085] In an embodiment of the present application, the task information includes but is not limited to the task content, the configuration requirements of the task processing model, and the requirements of the task processing specifications. Based on this, the model weights and configuration files of the task content are matched from the model file of the first processing model, so as to use the model weights and configuration files to initialize the second processing model in a second operating space different from the first operating space where the first processing model is located. Alternatively, the model weights and configuration files of the configuration requirements of the task processing model are matched from the model file of the first processing model, so as to use the model weights and configuration files to initialize the second processing model in a second operating space different from the first operating space where the first processing model is located. Alternatively, the model weights and configuration files of the configuration requirements of the task processing model are matched from the model file of the first processing model, so as to use the model weights and configuration files to initialize the second processing model in a second operating space different from the first operating space where the first processing model is located.
[0086] The first operation space can be understood as the working storage space of the first processing model, and the second operation space can be understood as the working storage space of the second processing model.
[0087] In the case where there are multiple task requests, obtaining the number of tasks of the task request, loading the model file of the first processing model into multiple running spaces corresponding to the number of tasks, so as to initialize multiple second processing models in parallel in the multiple running spaces;
[0088] In the embodiment of the present application, the number of running spaces of the second processing model depends on the number of tasks requested by the task. For example, if the number of tasks requested by the task is 6, the model file of the first processing model is loaded into the corresponding 6 running spaces to initialize 6 second processing models in parallel in the 6 running spaces.
[0089] The model file of the first processing model may be a mirror model file of the first processing model, or may be a partial model file that matches the respective task requests and is determined based on the respective task information.
[0090] When there are multiple task requests, the association relationship between the multiple task requests is obtained, and based on the association relationship, the model file of the first processing model is loaded into at least one operating space different from the first processing model, and at least one second processing model is initialized.
[0091] In an embodiment of the present application, the association relationship between multiple task requests is used to determine whether there is a necessary upper and lower relationship between tasks, or whether serial processing is required. If serial processing is required between certain tasks, there is no need to create a second processing model for the corresponding number of tasks, that is, the number of second processing models is less than the number of tasks.
[0092] For example, if task request 1 and task request 5 are connected in a vertical relationship, and task request 2 and task request 3 are connected in a vertical relationship, based on this, the model file of the first processing model is loaded into four running spaces different from the first processing model, and four second processing models are initialized, and then the second processing model 1 is used to serially process task request 1 and task request 5, the second processing model 2 is used to serially process task request 2 and task request 3, and the remaining two second processing models 3 and 4 are used to process task request 4 and task request 6 respectively. Among them, each second processing model performs parallel processing of tasks.
[0093] In some embodiments of the present application, at least one of the following is also included:
[0094] When the remaining duration is not greater than the target reference duration, adding the at least one task request to a queue of pending tasks of the first processing model;
[0095] It should be noted that, when the remaining time for the first processing model to enter the idle state is not greater than the target reference time required to load and run the second processing model, there is no need to create a second processing model based on the first processing model, add at least one task request to the pending task queue of the first processing model, wait for the first processing model to complete the task execution, and then, the first processing model processes the pending tasks serially or in parallel based on the identification information of the pending tasks in the pending task queue.
[0096] When there are multiple task requests, use multiple second processing models to process the multiple task requests in parallel;
[0097] Exemplarily, there are 6 task requests, including task request 1, task request 2, task request 3, task request 4, task request 5 and task request 6. Task request 1 is processed using the second processing model 1. At the same time, task request 2 is processed in parallel using the second processing model 2, task request 3 is processed in parallel using the second processing model 3, task request 4 is processed in parallel using the second processing model 4, task request 5 is processed in parallel using the second processing model 5, and task request 6 is processed in parallel using the second processing model 6, that is, the 6 task requests are processed in parallel using 6 second processing models.
[0098] When the task request is unique, the second processing model is used in parallel to process the unique task request while the first processing model is used to process the current task.
[0099] Exemplarily, there is a unique task request 1. While the current task is processed using the first processing model, the task request 1 is processed in parallel using the second processing model, that is, the first processing model and the second processing model process different task requests in parallel.
[0100] In some embodiments of the present application, obtaining at least one task request includes at least one of the following:
[0101] Obtaining a plurality of task configuration operations acting on a plurality of applications of the electronic device, and obtaining a plurality of task requests, wherein the plurality of applications can all call the first processing model to execute the task requests;
[0102] In an embodiment of the present application, multiple task configuration operations are performed on multiple applications of an electronic device, that is, a user publishes the same or similar task requests on different applications, so that the electronic device obtains the published task requests. Exemplarily, a user publishes a task for generating images on Lenovo's creator zone or Xiaotian, so that the electronic device obtains a task request for generating images from Lenovo's creator zone or Xiaotian. Alternatively, a user publishes a task for generating text or document processing on Lenovo's learning zone and Xiaotian at the same time, so that the electronic device obtains a task request for generating text or document processing from Lenovo's learning zone and Xiaotian.
[0103] Obtaining a plurality of task requests input to a first application of an electronic device, wherein the plurality of task requests all need to call the first processing model to execute corresponding processing actions;
[0104] In the embodiment of the present application, the user inputs multiple task requests in the first application, so that the electronic device obtains the multiple task requests from the first application.
[0105] A plurality of task requests are obtained, which are sent from a plurality of terminals through a communication connection with an electronic device, wherein the electronic device is a device configured with the first processing model and capable of executing the task requests.
[0106] In an embodiment of the present application, multiple terminals are respectively connected to electronic devices for communication, and the same user or different users input task requests in applications of multiple terminals. Multiple terminals send their respective task requests to the electronic device, so that the electronic device obtains multiple task requests from multiple terminals.
[0107] For example, in a home scenario, terminal devices of different users send task requests to a home center device (such as AI Center) to enable non-AI devices to use the functional services of AI devices.
[0108] In some embodiments of the present application, at least one of the following is also included:
[0109] In response to the electronic device configured with the first processing model establishing a communication connection with the first processing device, providing at least one of the at least one task request to the first processing device for processing;
[0110] In an embodiment of the present application, the first processing device may be an AI device in an edge network (such as a notebook, a chassis, an all-in-one (All In One, AIO) or an AI computing card, etc.), or a server providing AI services in the cloud.
[0111] Exemplarily, an electronic device configured with a first processing model establishes a communication connection with the AIO. When the remaining time of the first processing model entering the idle state meets the target condition, at least one second processing model is loaded and run based on the first processing model, and the number of the second processing models is less than the number of task requests, at least one second processing model is used to process part of the task requests in parallel, and the electronic device sends the remaining task requests to the AIO, and the AIO processes the remaining task requests.
[0112] In case the electronic device is configured with a third processing model, at least one of the at least one task request is given to the third processing model for processing.
[0113] In the embodiment of the present application, the electronic device is configured with not only the first processing model but also a third processing model, and the function of the third processing model is the same as that of the first processing model. The number of the third processing model is at least one.
[0114] Based on this, when the remaining time of the first processing model entering the idle state meets the target conditions, at least one second processing model is loaded and run based on the first processing model, and the number of second processing models is less than the number of task requests, at least one second processing model is used to process some task requests in parallel, and the third processing model is used to process the remaining task requests.
[0115] Based on the above embodiments, this application specifically illustrates a task processing method. Figure 5 The process diagram of the task processing method in the embodiment of the present application is as follows: Figure 5 ,like Figure 5 As shown, the following steps are included:
[0116] Step S501: upon receiving at least one task request, obtaining the current number of tasks.
[0117] Step S502: Obtain the average number of tasks in a plurality of consecutive T1 times corresponding to the first processing model.
[0118] Step S503: Determine whether the number of tasks exceeds the average number of tasks.
[0119] If yes, go to step S504; if no, go to step S508.
[0120] If the number of tasks exceeds the average number of tasks, it indicates that the current task is in a high-concurrency task state. If the number of tasks does not exceed the average number of tasks, it indicates that the current task is not in a high-concurrency task state.
[0121] Step S504: Obtain the number of units to be processed obtained by identifying at least one task request.
[0122] Step S505: Whether (the ratio of quantity / the inference speed of the first processing model*the target weight coefficient) is less than (T2-the inference time of the first processing model).
[0123] If yes, go to step S506; if no, go to step S508.
[0124] Among them, the ratio of quantity / inference speed of the first processing model*target weight coefficient is the target reference time, and T2-the inference time of the first processing model is the remaining time for the first processing model to enter the idle state. T2 can be the average processing time of the first processing model to process the current number of historical task requests.
[0125] Step S506: Load and run at least one second processing model based on the model file of the first processing model.
[0126] The model files include model weights and configuration files.
[0127] Step S507: using at least one second processing model to perform parallel processing on at least one task request.
[0128] Step S508: adding at least one task request to the to-be-processed task queue of the first processing model.
[0129] Step S509: using the first processing model to process at least one task request in the pending task queue.
[0130] Specifically, at least one task request in the pending task queue is processed serially or in parallel using the first processing model. Whether it is processed serially or in parallel depends on the number of current first processing models. If the number of current first processing models is one, at least one task request in the pending task queue is processed serially using the first processing model. If the number of current first processing models is multiple, at least one task request in the pending task queue is processed in parallel using multiple first processing models.
[0131] Based on this, the second processing model and the first processing model are located in different running spaces, respectively, so that the first processing model and the second processing model are isolated from each other, and further the tasks processed by the first processing model and the tasks processed by the second processing model are isolated from each other, thereby solving the question-answer or response confusion problem caused by the memory mechanism during multi-task concurrency. In addition, for high-concurrency tasks, task isolation is performed, and the application side switches to multi-model reasoning without perception. There is no need to perform special internal reasoning optimization for different models, such as adjusting kvcache, layering or operators. For hierarchical scheduling, for scenarios where concurrency demands are not very large, task queues are used to cache tasks and support concurrency without reducing the reasoning speed.
[0132] Based on the above embodiment, the present application illustrates a task processing block diagram. Figure 6 A task processing block diagram of an example of an embodiment of the present application is shown in FIG. Figure 6 As shown, there are multiple applications, including but not limited to AI agent application 1, AI browser application 2, and AI presentation application 3;
[0133] The load balancing module is used to execute the task processing method of the present application, that is, receiving a request from each application to call a first processing model, the first processing model can be an LLM model 4, or an Automatic Speech Recognition (ASR) model 5, or a Text-To-Speech (TTS) model 6;
[0134] The load balancing module 7 responds to at least one task request, and based on the number of tasks in the task request and the average number of tasks in the continuous T1 time corresponding to the first processing model, determines that the current state is a high-concurrency task, and obtains the remaining time for the first processing model to enter the idle state. When the remaining time is greater than the time required to load and run the second processing model, at least one working storage space, i.e., a running space, is opened up, and based on the first processing model through the vllm63 architecture, partial or all model weights and configuration files, i.e., model files 64 (stored in the working folder 65) corresponding to each task request are obtained, thereby loading and running at least one second processing model in at least one working storage space, and using at least one second processing model to perform parallel processing on at least one task request. On the contrary, if it is determined that the current state is not a high-concurrency task state based on the number of tasks requested and the average number of tasks within the continuous T1 time corresponding to the first processing model, or if the remaining time is not greater than the time required to load and run the second processing model, a container orchestration tool is used to open up a long-term storage space 62 for storing the pending task queue 61 of the first processing model, and at least one task request is added to the pending task queue of the first processing model, waiting for the first processing model to perform serial or parallel processing on the at least one task request.
[0135] Among them, the trained model or pruned model needs to be stored in the model repository for subsequent deployment, testing or sharing. The model repository provides unified management of models, allowing users to easily find and use trained or pruned models. The model repository needs to rely on persistent volumes to ensure the long-term preservation and reliability of model data. Persistent volumes provide a high-performance, high-availability, and scalable storage solution for the model repository, which helps improve the overall performance and user experience of the model repository.
[0136] In order to implement the method of the embodiment of the present application, based on the same inventive concept, the embodiment of the present application also provides a task processing device, Figure 7 Schematic diagram of the structure of the task processing device in the embodiment of the present application. Figure 7 As shown, the task processing device 70 includes:
[0137] The obtaining unit 701 is used for obtaining a remaining time length of the first processing model entering an idle state in response to obtaining at least one task request;
[0138] A processing unit 702 is configured to load and run at least one second processing model based on the first processing model when the remaining time meets a target condition;
[0139] The processing unit 702 is used to process the task request using the at least one second processing model;
[0140] The operating space where the second processing model is located is different from the operating space where the first processing model is located.
[0141] In an embodiment of the present application, the second processing model and the first processing model are located in different operating spaces, respectively, so that the operation isolation is achieved between the first processing model and the second processing model, and further the task isolation is achieved between the tasks processed by the first processing model and the tasks processed by the second processing model, thereby solving the question and answer or response confusion problem caused by the memory mechanism when multiple tasks are concurrent.
[0142] In some embodiments of the present application, the processing unit 702 is specifically used to obtain a target reference duration, which is the duration required to load and run the second processing model; when the remaining duration is greater than the target reference duration, at least one second processing model is loaded and run based on the model file of the first processing model, and the first processing model is the same as or different from the second processing model.
[0143] In some embodiments of the present application, the processing unit 702 is configured to include at least one of the following:
[0144] Obtain the processing time required for the first processing model to process the at least one task request, process the processing time based on the target weight coefficient to obtain the target reference time; obtain task information of the at least one task request, determine the configuration information of the second processing model based on the task information, and determine the target reference time based on the configuration information; read the target reference time from a preset configuration file; determine the target reference time based on the configuration data and usage data of the processor of the electronic device.
[0145] In some embodiments of the present application, the processing unit 702 is used to obtain the number of units to be processed obtained by identifying the at least one task request; obtain the processing capacity of the first processing model; and calculate the processing time based on the number and the processing capacity of the first processing model.
[0146] In some embodiments of the present application, the obtaining unit 701 is used to obtain the number of tasks of the task request and the processing capacity of the first processing model in response to obtaining at least one task request; when the first task state is determined based on the number of tasks and the processing capacity of the first processing model, determine to execute the step of obtaining the remaining time for the first processing model to enter the idle state; when the second task state is determined based on the number of tasks and the processing capacity of the first processing model, add the at least one task request to the queue of tasks to be processed of the first processing model.
[0147] In some embodiments of the present application, the processing unit 702 is configured to include at least one of the following:
[0148] Obtaining task information of the task request, determining a model weight and a configuration file matching the task information from a model file of the first processing model, and initializing a second processing model in a second operating space different from a first operating space where the first processing model is located by using the model weight and the configuration file;
[0149] In the case where there are multiple task requests, obtaining the request quantity of the task requests, loading the model file of the first processing model into multiple running spaces corresponding to the request quantity, so as to initialize multiple second processing models in parallel in the multiple running spaces;
[0150] When there are multiple task requests, the association relationship between the multiple task requests is obtained, and based on the association relationship, the model file of the first processing model is loaded into at least one operating space different from the first processing model, and at least one second processing model is initialized.
[0151] In some embodiments of the present application, the processing unit 702 is further configured to add the at least one task request to a to-be-processed task queue of the first processing model when the remaining duration is not greater than the target reference duration;
[0152] When there are multiple task requests, use multiple second processing models to process the multiple task requests in parallel;
[0153] When the task request is unique, the second processing model is used in parallel to process the unique task request while the first processing model is used to process the current task.
[0154] In some embodiments of the present application, the obtaining unit 701 is configured to include at least one of the following:
[0155] Obtaining a plurality of task configuration operations acting on a plurality of applications of the electronic device, and obtaining a plurality of task requests, wherein the plurality of applications can all call the first processing model to execute the task requests;
[0156] Obtaining a plurality of task requests input to a first application of an electronic device, wherein the plurality of task requests all need to call the first processing model to execute corresponding processing actions;
[0157] A plurality of task requests are obtained, which are sent from a plurality of terminals through a communication connection with an electronic device, wherein the electronic device is a device configured with the first processing model and capable of executing the task requests.
[0158] In some embodiments of the present application, the processing unit 702 is further configured to include at least one of the following:
[0159] In response to the electronic device configured with the first processing model establishing a communication connection with the first processing device, providing at least one of the at least one task request to the first processing device for processing;
[0160] In case the electronic device is configured with a third processing model, at least one of the at least one task request is given to the third processing model for processing.
[0161] The present application also provides another electronic device, Figure 8 Schematic diagram of the structure of the electronic device in the embodiment of the present application. Figure 8 As shown, the electronic device 80 includes: a processor 801 and at least one first processing model 802 capable of running on the processor, wherein the first processing model 802 can be called by a target application to perform at least one of the following: in response to obtaining at least one task request, obtaining a remaining time for the first processing model to enter an idle state;
[0162] When the remaining time meets the target condition, loading and running at least one second processing model based on the first processing model;
[0163] Processing the task request using the at least one second processing model;
[0164] The operating space where the second processing model is located is different from the operating space where the first processing model is located.
[0165] In practical applications, the processor may be at least one of an application-specific integrated circuit (ASIC), a digital signal processing device (DSPD), a programmable logic device (PLD), a field-programmable gate array (FPGA), a controller, a microcontroller, and a microprocessor. It is understandable that for different devices, the electronic device used to implement the functions of the processor may also be other, and the embodiments of the present application do not specifically limit this.
[0166] The above-mentioned memory can be a volatile memory (volatile memory), such as a random access memory (RAM); or a non-volatile memory (non-volatile memory), such as a read-only memory (ROM), a flash memory, a hard disk (HDD) or a solid-state drive (SSD); or a combination of the above-mentioned types of memory, and provide instructions and data to the processor.
[0167] In an exemplary embodiment, the present application also provides a computer-readable storage medium for storing a computer program.
[0168] Optionally, the computer-readable storage medium can be applied to any one of the methods in the embodiments of the present application, and the computer program enables the computer to execute the corresponding processes implemented by the processor in each method in the embodiments of the present application. For the sake of brevity, they are not repeated here.
[0169] Illustratively, an embodiment of the present application further provides a computer program product, including a computer program, which can be executed by a processor of an electronic device to complete the steps of any of the aforementioned methods.
[0170] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as: multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be electrical, mechanical or other forms.
[0171] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units; some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0172] In addition, all functional units in the embodiments of the present invention may be integrated into one processing module, or each unit may be separately used as a unit, or two or more units may be integrated into one unit; the above-mentioned integrated unit may be implemented in the form of hardware or in the form of hardware plus software functional units. A person of ordinary skill in the art may understand that all or part of the steps of implementing the above-mentioned method embodiments may be completed by hardware related to program instructions, and the aforementioned program may be stored in a computer-readable storage medium, which, when executed, executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium includes various media that can store program codes, such as mobile storage devices, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks or optical disks.
[0173] The methods disclosed in several method embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments.
[0174] The features disclosed in several product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0175] The features disclosed in several method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method embodiments or device embodiments.
[0176] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.
Claims
1. A task processing method, comprising: In response to obtaining at least one task request, obtaining a remaining time for the first processing model to enter an idle state; When the remaining time meets the target condition, loading and running at least one second processing model based on the first processing model; Processing the task request using the at least one second processing model; The operating space where the second processing model is located is different from the operating space where the first processing model is located.
2. The method according to claim 1, wherein: When the remaining time meets the target condition, loading and running at least one second processing model based on the first processing model includes: Obtaining a target reference duration, where the target reference duration is the duration required to load and run the second processing model; When the remaining duration is greater than the target reference duration, at least one second processing model is loaded and run based on the model file of the first processing model, and the first processing model is the same as or different from the second processing model.
3. The method according to claim 2, wherein: Obtaining the target reference duration includes at least one of the following: Obtaining a processing time required by the first processing model to process the at least one task request, and processing the processing time based on a target weight coefficient to obtain the target reference time; Obtaining task information of the at least one task request, determining configuration information of the second processing model based on the task information, and determining the target reference duration based on the configuration information; Reading the target reference duration from a preset configuration file; The target reference duration is determined based on configuration data and usage data of a processor of the electronic device.
4. The method according to claim 3, wherein: Obtaining the processing time required for the first processing model to process the at least one task request includes: Obtaining the number of units to be processed obtained by identifying the at least one task request; Obtaining the processing capability of the first processing model; The processing duration is calculated based on the number and a processing capability of the first processing model.
5. The method according to claim 1, wherein: Obtaining the remaining time for the first processing model to enter the idle state includes: In response to obtaining at least one task request, obtaining the number of tasks in the task request and the processing capacity of the first processing model; In the case where the number of tasks does not match the processing capacity of the first processing model, determining to execute a step of obtaining a remaining time for the first processing model to enter an idle state; When the number of tasks matches the processing capability of the first processing model, the at least one task request is added to a queue of pending tasks of the first processing model.
6. The method according to claim 1 or 2, wherein: Loading and running at least one second processing model based on the first processing model includes at least one of the following: Obtaining task information of the task request, determining a model weight and a configuration file matching the task information from a model file of the first processing model, and initializing a second processing model in a second operating space different from a first operating space where the first processing model is located by using the model weight and the configuration file; In the case where there are multiple task requests, obtaining the number of tasks of the task request, loading the model file of the first processing model into multiple running spaces corresponding to the number of tasks, so as to initialize multiple second processing models in parallel in the multiple running spaces; When there are multiple task requests, the association relationship between the multiple task requests is obtained, and based on the association relationship, the model file of the first processing model is loaded into at least one operating space different from the first processing model, and at least one second processing model is initialized.
7. The method according to claim 2, further comprising at least one of the following: When the remaining duration is not greater than the target reference duration, adding the at least one task request to a queue of pending tasks of the first processing model; When there are multiple task requests, use multiple second processing models to process the multiple task requests in parallel; When the task request is unique, the second processing model is used in parallel to process the unique task request while the first processing model is used to process the current task.
8. The method according to claim 1, wherein: Obtain at least one task request, including at least one of the following: Obtaining a plurality of task configuration operations acting on a plurality of applications of the electronic device, and obtaining a plurality of task requests, wherein the plurality of applications can all call the first processing model to execute the task requests; Obtaining a plurality of task requests input to a first application of an electronic device, wherein the plurality of task requests all need to call the first processing model to execute corresponding processing actions; A plurality of task requests are obtained, which are sent from a plurality of terminals through a communication connection with an electronic device, wherein the electronic device is a device configured with the first processing model and capable of executing the task requests.
9. The method according to claim 1, further comprising at least one of the following: In response to the electronic device configured with the first processing model establishing a communication connection with the first processing device, providing at least one of the at least one task request to the first processing device for processing; In case the electronic device is configured with a third processing model, at least one of the at least one task request is given to the third processing model for processing.
10. An electronic device, comprising a processor and at least one first processing model capable of running on the processor, wherein the first processing model can be called by a target application to perform at least one of the following: In response to obtaining at least one task request, obtaining a remaining time for the first processing model to enter an idle state; When the remaining time meets the target condition, loading and running at least one second processing model based on the first processing model; Processing the task request using the at least one second processing model; in, The operating space where the second processing model is located is different from the operating space where the first processing model is located.