Model reasoning method and device for processor cluster, electronic equipment and storage medium
By obtaining relevant information of each target processor in the processor cluster, determining the processing order of inference tasks, and loading preset modules of the inference model during idle or efficient resource utilization periods, the problem of the processor cluster affecting other processing tasks when executing inference tasks is solved, and user experience and resource utilization efficiency are improved.
Patent Information
- Application Number
- CN202411953296.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-27
- Publication Date
- 2025-06-03
AI Technical Summary
When performing inference tasks, processor clusters will affect other processing tasks, resulting in poor user experience.
By obtaining relevant information of each target processor, including available resource information and idle periods, the processing order of each target processor performs inference tasks is determined, and the target processor is instructed to load preset modules of the inference model separately to perform inference task processing during idle or efficient resource utilization periods.
Reduces the processor's execution impact on other processing tasks, improves the user experience, and effectively utilizes the resources of the processor cluster.
Smart Images

Figure CN120085975A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer processing, and in particular, to a method, device, electronic device and storage medium for model inference of a processor cluster. Background Art
[0002] With the rapid development of artificial intelligence, more and more scenarios utilize artificial intelligence to process inference tasks, such as processing image tasks, text tasks, or voice tasks, etc.
[0003] In the related art, multiple preset modules of an inference model can be respectively loaded by multiple processors in a processor cluster, and then each processor can process an inference task based on the preset module it loads.
[0004] However, the processors in the processor cluster also need to execute other processing tasks besides the inference tasks. In this way, when using the processors in the processor cluster to execute inference tasks, it will affect other processing tasks of the processors, resulting in a poor user experience. Summary of the Invention
[0005] In view of this, an object of the present invention is to provide a method, device, electronic device and storage medium for model inference of a processor cluster, which can reduce the impact of processors on the execution of other processing tasks when using the processors in the processor cluster to execute inference tasks, thereby improving the user experience.
[0006] In a first aspect, an embodiment of the present invention provides a method for model inference of a processor cluster. The processor cluster includes multiple processors. The method includes: obtaining first data related to an inference task; obtaining relevant information of each target processor among multiple target processors, where the relevant information includes available resource information and idle periods, or the relevant information includes target periods that meet data processing capacity requirements, and the available resource information indicates the data processing capacity of the target processor. The multiple target processors are at least part of the multiple processors; determining the order of relevant processing for each target processor to execute the inference task based on the idle periods or target periods of each target processor; and sequentially instructing the multiple target processors to respectively load multiple preset modules of a target inference model based on the order of relevant processing for each target processor to execute the inference task, so as to perform data processing by each target processor based on the preset module it loads during its corresponding idle period or target period, where the first data is used as the input of the target processor indicated for the first time.
[0007] In a possible implementation, based on the order of relevant processing for each target processor to execute the inference task, multiple target processors are successively instructed to load multiple preset modules of the target inference model, including: determining a first target processor and a second target processor among the multiple target processors based on the order of relevant processing for each target processor to execute the inference task, where the order of the first target processor is prior to that of the second target processor; when reaching the idle period or the target period corresponding to the first target processor, instructing the first target processor to load the first part of the modules of the target inference model, so that the first target processor processes the first data based on the first part of the modules it loads to obtain second data, and the first part of the modules is a part of the multiple preset modules; when reaching the idle period or the target period corresponding to the second target processor, instructing the second target processor to load the second part of the modules of the target inference model, so that the second target processor processes the second data based on the second part of the modules it loads to obtain third data, and the second part of the modules is the modules other than the first part of the modules among the multiple preset modules. In the target inference model, the processing data order of the second part of the modules is after that of the first part of the modules.
[0008] In a possible implementation, the method further includes: initiating a consensus vote to the processor cluster to obtain the voting results from each processor in the processor cluster, where the voting results include the participating processors voted by each processor for participating in the inference task, and the participating processors are at least part of the multiple processors; and selecting multiple target processors from the multiple processors based on the voting results of each processor.
[0009] In a possible implementation, selecting multiple target processors from the multiple processors based on the voting results of each processor includes: determining the participating processors in the voting results from each processor as the target processors; or, the voting results further include the target times corresponding to each participating processor, and the target times are used to indicate the number of times the voting processor collaborates with the participating processor to process the same inference task. Determining the participating processors with target times greater than the times threshold in the voting results from each processor as the target processors.
[0010] In a possible implementation, obtaining the relevant information of each target processor among the multiple target processors includes: obtaining the historical operation information of each target processor among the multiple target processors, and determining the relevant information of each target processor based on the historical operation information of each target processor; or, sending a relevant information acquisition request to each target processor among the multiple target processors, and receiving the relevant information sent by each target processor, where the relevant information is determined by each target processor based on its historical operation information in response to the relevant information acquisition request.
[0011] In a possible implementation, the historical operation information includes: the resource utilization rate of the target processor corresponding to each preset time period. The step of determining the relevant information of the target processor based on the historical operation information of the target processor includes: determining the idle period of the target processor based on the resource utilization rate of the target processor corresponding to each preset time period, where the idle period is a preset time period in which the corresponding resource utilization rate is less than the resource utilization rate threshold; obtaining the maximum resource information of the target processor, where the maximum resource information is used to indicate the maximum data processing capacity of the target processor; and determining the available resource information of the target processor corresponding to the idle period based on the maximum resource information and the resource utilization rate corresponding to the idle period.
[0012] In a possible implementation, the method further includes: determining the total resource information of the processor cluster, where the total resource information is used to indicate the total data processing capacity of the processor cluster; and determining a target inference model from multiple preset inference models based on the total resource information, where the target inference model is the preset inference model with the largest corresponding number of parameters among the preset inference models that can be run by the total data processing capacity.
[0013] In a possible implementation, the method further includes: in the case where the total data processing capacity cannot run the preset inference model corresponding to the smallest number of parameters among the multiple preset inference models, broadcasting a target processor addition request, where the target processor addition request is used to request to add a new processor to the processor cluster; receiving an agreement message from the new processor, where the agreement message is used to indicate that the new processor agrees to join the processor cluster; and in response to the agreement message, adding the new processor to the processor cluster to use the new processor to perform part of the processing of the inference task.
[0014] In a second aspect, an embodiment of the present invention provides a model inference device for a processor cluster. The processor cluster includes multiple processors. The device includes: an acquisition module, configured to acquire first data related to an inference task; acquire the relevant information of each target processor among multiple target processors, where the relevant information includes available resource information and an idle period, or the relevant information includes a target period that meets the data processing capacity requirement, and the available resource information indicates the data processing capacity of the target processor; a determination module, configured to determine the order of relevant processing for each target processor to execute the inference task based on the idle period or the target period of each target processor; and a processing module, configured to sequentially instruct the multiple target processors to load multiple preset modules of the target inference model based on the order of relevant processing for each target processor to execute the inference task, so as to perform data processing by each target processor based on the preset module loaded by it during its corresponding idle period or target period, where the first data is used as the input of the target processor indicated for the first time.
[0015] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method of the first aspect.
[0016] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing computer-executable instructions. When the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method of the first aspect.
[0017] The embodiments of the present invention bring the following beneficial effects: obtaining first data related to an inference task; obtaining relevant information of each target processor among multiple target processors, where the relevant information includes available resource information and idle periods, or the relevant information includes target periods that meet data processing capacity requirements, and the available resource information indicates the data processing capacity of the target processor, and the multiple target processors are at least part of multiple processors; determining the order of relevant processing for each target processor to execute the inference task based on the idle periods or target periods of each target processor; and sequentially instructing the multiple target processors to respectively load multiple preset modules of a target inference model based on the order of relevant processing for each target processor to perform data processing on the basis of the preset modules loaded by each target processor during its corresponding idle period or target period, where the first data is used as the input of the target processor indicated for the first time. That is to say, part of the processing of the inference task can be performed by the target processor only when the target processor has redundant data processing capacity for the inference task. Thus, when using the target processors in a processor cluster to execute the inference task, the influence of the target processors on the execution of other processing tasks can be reduced, thereby improving the user experience.
[0018] Other features and advantages of the present invention will be described in the following specification, and part of them will become obvious from the specification, or be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.
[0019] To make the above objectives, features, and advantages of the present invention more obvious and understandable, the following specific preferred embodiments are given in conjunction with the accompanying drawings and described in detail as follows. Description of the Drawings
[0020] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the following drawings are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.
[0021] Figure 1 Schematic diagram of the architecture of a multi-processor inference task processing system provided by an embodiment of the present invention;
[0022] Figure 2 Schematic diagram of the process of a model inference method for a processor cluster provided by an embodiment of the present invention;
[0023] Figure 3 Schematic diagram of the architecture of an inference model provided by an embodiment of the present application;
[0024] Figure 4 Schematic diagram of the structure of a model inference device for a processor cluster provided by an embodiment of the present invention;
[0025] Figure 5 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed implementation manners
[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Apparently, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0027] In the related art, multiple preset modules of an inference model can be loaded by multiple processors in a processor cluster, and each processor can process inference tasks based on the preset module it loads.
[0028] However, the processors in the processor cluster also need to execute other processing tasks besides inference tasks. For example, the processors in the processor cluster also need to perform processing such as video encoding and decoding on the recorded video. At this time, if the processor calls the processor to jointly execute the inference task when performing other processing tasks, it will cause the computing power of the processor to be insufficient, thereby affecting other processing tasks executed by the processor. For example, the encoding and decoding speed becomes slower, resulting in a lag in the displayed picture.
[0029] It should be understood that other processing tasks are not limited to the above examples. Any task different from the inference task that needs to be processed can be considered as other tasks, and no limitation is made here.
[0030] In view of this, the embodiments of the present invention provide a model inference method, device, electronic device, and storage medium for a processor cluster, which can reduce the impact of the processor on the execution of other processing tasks when using the processors in the processor cluster to execute inference tasks, thereby improving the user experience.
[0031] Next, an exemplary description of the application scenarios of the technical solutions of the embodiments of the present application will be given:
[0032] Application Scenario 1: Invoke the remaining computing power of the processors in each vehicle. When the processor in the vehicle is idle or can meet the computing power requirements, the remaining computing power provided by the processors in multiple vehicles is used to jointly process the inference task.
[0033] Application Scenario 2: Invoke the remaining computing power of the processors of each training model. When the processor of the training model is idle (such as after training ends or waiting for training) or can meet the computing power requirements (such as when training a model with a small number of training parameters or performing data processing with a small amount of computing), the remaining computing power of the processors of multiple training models is used to jointly process the inference task.
[0034] To facilitate the understanding of this embodiment, first, the system framework disclosed in the embodiments of the present invention will be described. As Figure 1 shown, the inference task processing system with multiple processors may include a management node (master), multiple processors, and a database. Among them, the management node can be used for the management of multiple processors, such as instructing the model modules loaded by multiple processors, etc. The multiple processors are used to jointly execute the inference task. The database is used to store data related to the inference task, such as storing the parameters of the model, storing the first data for inference, and storing the intermediate data generated during the inference process, etc. Among them, the database can include a local database or a database stored in the cloud. The multiple processors can form a processor cluster. It should be noted that the multiple processors can be set in one terminal device. In addition, the multiple processors can be set in multiple different terminal devices, and no limitation is made here. The terminal device can also be called an electronic device. The processor can include a graphics processor, a central processor, or a neural network processor, etc.
[0035] Optionally, the system may further include a scheduling node (scheduler). The scheduling node can run a scheduling program to schedule the master, such as periodically polling the Master to check whether there are idle nodes, so as to submit the inference task.
[0036] It should be noted that the management node and the multiple processors are independent of each other, and one of the processors can also be used as the management node, and no limitation is made here. In addition, the scheduling program and the management node can be independent or the same node, and no limitation is made here.
[0037] Next, a detailed introduction to a model inference method for a processor cluster will be given. Please refer to Figure 2 , Figure 2The flowchart of a model inference method for a processor cluster provided by an embodiment of the present invention. The model inference method for a processor cluster in one embodiment of the present disclosure can run on a management node. The method includes the following steps:
[0038] S210. Obtain first data related to an inference task.
[0039] Among them, an inference task refers to a process in which a model derives a conclusion or data through internal logic and algorithms based on input information or data. In this embodiment, the inference task may include, but is not limited to, text processing tasks, image processing tasks, or speech processing tasks, etc. Optionally, the text processing task may be, for example, the matching between texts. For example, the model finds matching texts in a database based on the text input by the user. The image processing task may be, for example, after the model processes the input image, outputs a decision for autonomous driving or a denoised image, etc. The speech processing task may be, for example, the model outputs a speech text from the input speech. Optionally, one application scenario may be, for example: generating a description text of a picture, matching the text input by the user with the description text, so as to find the picture matching the text input by the user. Then the inference task may be, for example, generating description texts of multiple pictures, or matching the text input by the user with the description text, so as to find the picture matching the text input by the user, etc., which is not limited herein.
[0040] In this embodiment, the first data may be the data input to the model, or the intermediate data obtained after the model processes the data input to the model. Taking the text processing task as an example, the first data may be the text input to the model, or the text vector obtained by processing the text. Taking image processing as an example, the first data may be the input image data, or the image features obtained by feature extraction of the image, which is not limited herein.
[0041] S220. Obtain the relevant information of each target processor among multiple target processors. The relevant information includes available resource information and idle periods, or the relevant information includes target periods that meet the data processing capacity requirements. The available resource information indicates the data processing capacity of the target processor.
[0042] Among them, the multiple target processors are at least part of the multiple processors. The available resource information may include, but is not limited to, the number of cores, operating frequency, etc. If the target processor includes a graphics processor, the available resource information may also include the video memory size. It should be noted that the available resource information may be the remaining resource information, and the available resources are less than or equal to the maximum resources, and the maximum resources indicate the maximum data processing capacity of the target processor. In this embodiment, the relevant information may include the associated available resource information and the idle period. That is to say, the relevant information includes the idle period and the available resource information corresponding to the idle period. In this way, it can be known how much data processing capacity can be used during the idle period. Alternatively, the relevant information includes the target period that meets the data processing capacity requirement. That is to say, through the relevant information, it can be known when the target processor can meet the data processing capacity. In this way, it can be known when the target processor can meet the data processing capacity requirement, so as to perform part of the inference task processing through the target processor during the period when the target processor can meet the data processing requirement.
[0043] S230. Determine the order of the relevant processing for each target processor to execute the inference task based on the idle period or target period of each target processor.
[0044] In this embodiment, the order of the relevant processing for each target processor to execute the inference task is related to the idle period or target period of each target processor. Specifically, the earlier the idle period is, the earlier the order of the relevant processing for executing the inference task is. Exemplarily, the idle period or target period corresponding to target processor A is from 9:00 to 10:00, and the idle period or target period corresponding to target processor B is from 11:00 to 12:00. Then the order of the relevant processing for target processor A to execute the inference task is earlier than the order of the relevant processing for target processor B to execute the inference task. That is to say, first perform the relevant processing for target processor A to execute the inference task, and then perform the relevant processing for target processor B to execute the inference task.
[0045] S240. Based on the order of the relevant processing for each target processor to execute the inference task, sequentially instruct the multiple target processors to respectively load multiple preset modules of the target inference model, so as to perform data processing based on the loaded preset modules by each target processor during its corresponding idle period or target period, where the first data is used as the input of the target processor indicated for the first time.
[0046] In this embodiment, multiple target processors respectively load multiple preset modules of a target inference model, and at least some of the preset modules loaded by the target processors are different. Specifically, since the multiple target processors are successively instructed to load the preset modules of the target inference model according to the order of relevant processing for executing the inference task by each target processor, the target processor instructed for the first time may be the target processor that first executes the inference task. Therefore, the first data can be used as the input of the target processor instructed for the first time, so that the target processor instructed for the first time processes based on the first data.
[0047] In this embodiment, by obtaining first data related to an inference task; obtaining relevant information of each target processor among multiple target processors, where the relevant information includes available resource information and idle time periods, or the relevant information includes target time periods that meet the data processing capacity requirements, and the available resource information indicates the data processing capacity of the target processor; determining the order of relevant processing for each target processor to execute the inference task based on the idle time periods or target time periods of each target processor; and successively instructing multiple target processors to respectively load multiple preset modules of the target inference model based on the order of relevant processing for each target processor to execute the inference task, so as to perform data processing by each target processor based on the preset module loaded by it during its corresponding idle time period or target time period, where the first data is used as the input of the target processor instructed for the first time. That is to say, when the target processor has redundant data processing capacity for the inference task, a part of the processing for the inference task can be performed by using the target processor. Thus, when using the target processors in the processor cluster to execute the inference task, the influence of the target processors on the execution of other processing tasks can be reduced, thereby improving the user experience.
[0048] In a possible implementation manner, successively instructing multiple target processors to respectively load multiple preset modules of the target inference model based on the order of relevant processing for each target processor to execute the inference task includes:
[0049] Determine a first target processor and a second target processor among multiple target processors based on the order of relevant processing for the inference task executed by each target processor, where the order of the first target processor is prior to that of the second target processor; when reaching the idle period or target period corresponding to the first target processor, instruct the first target processor to load the first part of modules of the target inference model, so as to process the first data based on the first part of modules loaded by the first target processor to obtain second data, and the first part of modules is a part of multiple preset modules; when reaching the idle period or target period corresponding to the second target processor, instruct the second target processor to load the second part of modules of the target inference model, so as to process the second data based on the second part of modules loaded by the second target processor to obtain third data, and the second part of modules is the modules other than the first part of modules among the multiple preset modules. In the target inference model, the processing data order of the second part of modules is after that of the first part of modules.
[0050] In this embodiment, the first part of modules may include one or more modules. The second part of modules may include one or more modules. In the inference model, the processing data order of the second part of modules is after that of the first part of modules. That is to say, in the process of running the inference model to process the inference task, the data is first processed by the first part of modules and then by the second part of modules. Therefore, in the inference model, the first part of modules and the second part of modules may be adjacent modules, that is, the output of the first part of modules can be used as the input of the second part of modules. In this embodiment, the inference model may include the first part of modules and the second part of modules, and may also include other modules. That is to say, the first part of modules plus the second part of modules may be part or all of the modules of the inference model, which is not limited here. Among them, the second data may be obtained by processing the first data, so the second data can be understood as intermediate data. The third data may be the data finally output by the inference model, such as the decision result of autonomous driving or the denoised image, etc., or may be the intermediate data result of the inference model, that is, the inference model needs to continue to process the third data to output the final data.
[0051] Exemplarily, assume that the idle period or target period corresponding to the target processor A is from 9:00 to 10:00, and the idle period or target period corresponding to the target processor B is from 11:00 to 12:00. Then, the target processor A is used as the first target processor, and the target processor B is used as the second target processor. The first part of the module is loaded by the target processor A, so that the target processor A processes the first data based on the first part of the module it loads from 9:00 to 10:00 to obtain the second data. Moreover, the second part of the module is loaded by the target processor B, so that the target processor B processes the second data based on the second part of the module it loads from 11:00 to 12:00 to obtain the third data.
[0052] It should be noted that in this embodiment, the target processor can obtain and load the corresponding preset module before the minimum time of its idle period or target period. Exemplarily, the time difference between the minimum time and the time when the target processor obtains and loads the preset module can be used as the time when the target processor starts to obtain the preset model. In this way, the preset module can be loaded for relevant processing when the corresponding idle period or target period is reached.
[0053] It should be understood that multiple preset modules of the inference model can be stored locally in the terminal device or on the server, and the target processor can request the preset module from the server.
[0054] It should be noted that in this embodiment, if the order of the target processors to execute the inference task is determined based on the idle period, the modules loaded by each processor can be indicated based on the available resource information of each processor. Specifically, the larger the available resource information, the larger the number of parameters of the loaded module. If the order of the target processors is determined based on the target period, the module that can be loaded and run according to the data processing ability requirement can be used as the module loaded by the target processor. Among them, the data processing ability requirement can be represented by a resource threshold.
[0055] For ease of understanding, the following uses a specific model to exemplarily illustrate this solution. Please refer to Figure 3 , Figure 3 which is a schematic diagram of the architecture of an inference model provided by an embodiment of the present application. As Figure 3 shown, the inference model can be a model with a transformer architecture.
[0056] Among them, models with a transformer architecture all include a language model, and different models have different numbers of layers. The models for image processing include the following modules:
[0057] Vision Tower: Responsible for extracting image features. The number of layers in the vision block varies in different models. Vision Tower generally refers to a neural network module dedicated to processing visual data, which can extract and encode features of images. In a multi-modal model, Vision Tower is combined with modules responsible for processing text data (such as language models) to jointly achieve cross-modal information processing and generation.
[0058] Multi-modal Projector: Responsible for projecting the encoded image modal features into the text feature space to obtain aligned features, usually consisting of 2 layers. The main function of the multi_modal_projector block is to project data of different modalities (such as images, text, etc.) into the same feature space for cross-modal information processing and fusion. In a multi-modal model, the multi_modal_projector block is usually responsible for converting visual features into a format compatible with text features, thus enabling cross-modal interaction and generation.
[0059] Next, an Figure 3 exemplary description of the data processing process will be given.
[0060] 1. Input Embedding:
[0061] First, the input data (which may be text, image, or other sequential data) is converted into embedding vectors. These embedding vectors are usually learned and can capture the features of the input data.
[0062] 2. Positional Encoding:
[0063] Since the transformer model itself does not have the ability to process the sequence order, positional encoding needs to be added to identify the position of each element in the input sequence. The positional encoding can be added to the input embedding or directly included in the input embedding.
[0064] 3. Multi-Head Attention:
[0065] Next, the data passes through the multi-head attention layer. This layer divides the input vector into multiple heads, and each head independently calculates the attention weights and performs weighted summation on the input vector based on these weights. The multi-head attention mechanism allows the model to simultaneously focus on different positions in the input sequence, thereby capturing more complex dependencies.
[0066] 4. Add & Normalization (add&nNorm):
[0067] After the multi-head attention layer, residual connections (i.e., adding the input and the output) and layer normalization are usually performed to improve the stability and performance of the model.
[0068] 5. Feed Forward Layer:
[0069] The data then passes through the feed forward layer, which is a simple fully-connected neural network that performs a non-linear transformation on the vectors at each position.
[0070] 6. Add & Normalize Again:
[0071] Similar to after the multi-head attention layer, residual connections and layer normalization are also performed after the feed forward layer.
[0072] 7. Linear Layer:
[0073] After being processed by the feed forward layer, the data is transformed by the linear layer in preparation for output.
[0074] 8. Softmax Layer:
[0075] Finally, the data is normalized by the softmax layer to obtain the output probability distribution. The softmax layer converts the output of the linear layer into probability values such that the sum of all output values is 1.
[0076] 9. Output Layer (output probabilities):
[0077] The output layer provides the final prediction results, i.e., the probability distribution for each class.
[0078] Referring to Figure 3 In the model architecture shown, by way of example, the first data can be, for example, the input data such as text, image or other sequence data, the second data can be, for example, the data output by the multi-head attention, and the third data can be, for example, the data output by the output layer. Again by way of example, the first data can be, for example, the input data, the second data can be, for example, the data output by the positional encoding, and the third data can be, for example, the data output by the multi-head attention. Again by way of example, the first data can be, for example, the data output by the positional encoding, the second data can be, for example, the data output by the multi-head attention, and the third data can be, for example, the data output by the output layer.
[0079] It should be understood that the way of dividing the actual first - part module and second - part module is not limited to the examples shown above. The first - part module and second - part module can be divided as needed. For example, they can be divided according to functions, and no restrictions are imposed here. In a possible implementation, the model can also be divided according to the performance of the target processor. Taking the target processor as a GPU, when facing a GPU with limited video memory, it is crucial to reasonably divide the model. The following is a detailed strategy for dividing the model based on the size of the GPU video memory, especially for the two key modules of vision_tower and language_model. It should be noted that the occupancy of GPU video memory mainly comes from two aspects:
[0080] 1. Loading model parameters: This includes static parameters such as network weights and biases, which are loaded into the video memory during model initialization.
[0081] 2. Runtime calculations: This includes dynamic processes such as forward propagation, backward propagation, and gradient updates, which occupy video memory during model training or inference.
[0082] To effectively utilize the GPU video memory, a control parameter can be introduced to define the upper limit of the proportion of video - memory occupancy during model runtime. This control parameter can be, for example, gpu_memory_utilization, which defines the upper limit of the proportion of video - memory occupancy during model runtime. By default, this proportion is set to 0.5, meaning that the video - memory occupancy during model runtime should not exceed half of the total GPU video memory.
[0083] For a GPU with 24G of video memory, if the model is loaded using the float32 data type, the number of parameters per model block should be less than 6 billion (6 * 10^9). This is because each parameter of the float32 type occupies 4 bytes, and other factors (such as gradients and activation values) also need to be considered for video - memory occupancy.
[0084] When splitting the vision_tower block, the vision_tower block is split into multiple small blocks, and each small block contains k layers. The specific value of k depends on the number of parameters and computational complexity of each layer. To ensure that each small block can run within the video - memory limit, the value of k can be gradually adjusted through experiments. When splitting, layers with larger numbers of parameters (such as convolutional layers and fully - connected layers) can be considered to be dispersed into different small blocks to balance video - memory occupancy.
[0085] When splitting the language_model block, similarly, the language_model block is also split into multiple small blocks, and each small block contains m layers. The value of m also needs to be adjusted according to the number of parameters and computational complexity of each layer. For language models, especially Transformer-based models, key components such as the attention mechanism and feed-forward network can be considered to be dispersed into different small blocks to optimize the use of video memory.
[0086] Next, an exemplary description of how to determine the target processor among multiple processors will be given.
[0087] In one possible implementation, the method further includes:
[0088] Initiate a consensus vote to the processor cluster to obtain the voting results from each processor in the processor cluster. The voting results include the participating processors voted by each processor for participating in the inference task. The participating processors are at least part of the multiple processors. Based on the voting results of each processor, select multiple target processors from the multiple processors.
[0089] In this embodiment, the purpose of the consensus vote is to vote for the target processors for collaborative inference tasks. In this embodiment, the participating processors in the voting results sent by the voting processors do not include the voting processors themselves. Optionally, the way to initiate the consensus vote can be, for example, to broadcast a consensus vote request to each processor in the processor cluster, and then the processors that receive the consensus vote request can send the voting results to the management node.
[0090] In one possible implementation, based on the voting results of each processor, selecting multiple target processors from the multiple processors includes:
[0091] Determine the participating processors in the voting results from each processor as the target processors.
[0092] Exemplarily, assume that the multiple processors include processor A, processor B, and processor C. Then assume that the voting result of processor A includes processor B, the voting result of processor B includes processor C, and the voting result of processor C includes processor B. Then the multiple target processors include processor B and processor C.
[0093] In this embodiment, the participating processors in the voting results of each processor can be determined as the target processors. In this way, the number of target processors is larger, so a larger inference model with more parameters can be loaded.
[0094] In another possible implementation, based on the voting results of each processor, selecting multiple target processors from the multiple processors includes:
[0095] The voting result also includes the target number corresponding to each participating processor. The target number is used to indicate the number of times the processor that votes collaborates with the participating processors to process the same inference task. Among the participating processors in the voting results from each processor, the participating processors with a target number greater than the number threshold are determined as target processors.
[0096] Exemplarily, assume that multiple processors include processor A, processor B, and processor C. Then assume that the voting result of processor A includes processor B, and the target number corresponding to processor B is 2. The voting result of processor B includes processor C and processor A, and the target number corresponding to processor A is 2. The target number corresponding to processor C is 4. The voting result of processor C includes processor B, and the target number corresponding to processor B is 4. Assume that the number threshold is 3. Then the multiple target processors include processor B and processor C.
[0097] In this embodiment, the more times two processors collaborate on an inference task together, the greater the likelihood that both processors can meet the data processing capacity requirements. Therefore, first, multiple target processors are selected through consensus voting, and then, based on the relevant information corresponding to each of the multiple target processors, each part module of the model is respectively indicated to be loaded by the multiple target processors, which can reduce the computing power resources required to execute the inference task.
[0098] In a possible implementation manner, before obtaining the relevant information of each target processor among the multiple target processors, it includes:
[0099] Determine the target processors among the multiple processors that meet the data processing capacity requirements.
[0100] In this embodiment, not all processors can meet the data processing capacity requirements to load the preset module for relevant processing. Therefore, based on the available resource information corresponding to each processor, the target processors among the multiple processors can be determined, and the target processors are the processors whose data processing capacity meets the data processing capacity requirements. In this way, the relevant processing of the inference task can be performed by the target processors that can meet the data processing capacity requirements, thereby improving the success rate of executing the inference task and reducing the situation where the relevant processing of the inference task cannot be normally performed due to insufficient processor performance.
[0101] It should be noted that the available resource information can be the available resource information corresponding to the idle period. Taking the processor as the GPU and the available resource information including the video memory as an example, the video memory threshold is determined according to the data processing capacity requirement. If the video memory is less than the video memory threshold, the data processing capacity of the GPU cannot meet the data processing capacity requirement, and the processor does not participate in the relevant processing of the inference task. If the video memory is greater than or equal to the video memory threshold, the data processing capacity of the GPU can meet the data processing capacity requirement, and the processor participates in the relevant processing of the inference task as the target processor.
[0102] In another possible implementation, all processors can also be determined as target processors.
[0103] In a possible implementation, the relevant information of each target processor among the multiple target processors is obtained, including:
[0104] Obtain the historical operation information of each target processor among the multiple target processors, and determine the relevant information of each target processor based on the historical operation information of each target processor; or, send a relevant information acquisition request to each target processor among the multiple target processors, and receive the relevant information sent by each target processor, where the relevant information is determined by each target processor based on its historical operation information in response to the relevant information acquisition request.
[0105] In this embodiment, the relevant information of the target processor can be determined by the management node according to the historical operation information of each target processor, or the management node can request the target processor to request the target processor to inform its relevant information, that is, the relevant information of the target processor is determined by the target processor itself.
[0106] Optionally, the historical operation information includes the resource utilization rate corresponding to each preset time period, and the preset time period corresponding to the resource utilization rate lower than the resource utilization rate threshold is determined as the idle time period of the target processor. It should be understood that the resource utilization rate threshold can be set as needed, for example, set to 10% etc., and is not limited here. Exemplarily, the preset time period can be set as needed. For example, a day is divided into 24 time periods, the minimum time and the maximum time in each time period are separated by one hour, and the time between any two time periods does not overlap.
[0107] In a possible implementation, the historical operation information includes: the resource utilization rate of the target processor corresponding to each preset time period. The step of determining the relevant information of the target processor based on the historical operation information of the target processor includes:
[0108] Determine the idle period of the target processor based on the resource utilization rate of the target processor corresponding to each preset time period, where the idle period is a preset time period in which the corresponding resource utilization rate is less than the resource utilization rate threshold; obtain the maximum resource information of the target processor, and the maximum resource information is used to indicate the maximum data processing capacity of the target processor; determine the available resource information of the target processor corresponding to the idle period based on the maximum resource information and the resource utilization rate corresponding to the idle period.
[0109] In this embodiment, the maximum resource information may be obtained by querying based on the identification information of the target processor, and the identification information may be, for example, the model of the target processor.
[0110] In a possible implementation manner, the method further includes:
[0111] Determine the total resource information of the processor cluster, and the total resource information is used to indicate the total data processing capacity of the processor cluster; determine the target inference model from multiple preset inference models based on the total resource information, where the target inference model is the preset inference model with the largest corresponding number of parameters among the preset inference models that can be run by the total data processing capacity.
[0112] Exemplarily, taking the target processor as a GPU and the resource information as the video memory, the total resource information may be the sum of the video memories of multiple GPUs. If the sum of the video memories of multiple GPUs is greater than the number of parameters of the preset inference model globally, then multiple GPUs can run the preset inference model. Then, select the preset inference model with the largest number of parameters as the target inference model from the preset inference models that can be run.
[0113] It should be understood that generally, the larger the number of parameters, the better the processing effect of the inference task. For example, the more accurate the text matching, or for another example, the better the effect of image denoising.
[0114] In this embodiment, by determining the target inference model from multiple preset inference models based on the total resource information, where the target inference model is the preset inference model with the largest corresponding number of parameters among the preset inference models that can be run by the total data processing capacity, that is, based on the total resource information of multiple target processors, select the preset inference model with the best processing effect from the preset inference models that can be run. In this way, the processing effect of the inference task can be improved.
[0115] In a possible implementation manner, the method further includes:
[0116] In the case where the total data processing capacity cannot run the preset inference model corresponding to the minimum number of parameters among multiple preset inference models, the broadcast processor adds a request, where the processor addition request is used to request adding a new processor to the processor cluster; receives an approval message from the new processor, where the approval message is used to indicate that the new processor agrees to join the processor cluster; in response to the approval message, adds the new processor to the processor cluster to utilize the new processor to perform part of the processing of the inference task.
[0117] Taking the target processor as a GPU and the resource information as the video memory as an example, the total resource information can be the sum of the video memories of multiple GPUs. If the sum of the video memories of multiple GPUs is less than the number of parameters of the preset inference model globally, then multiple GPUs cannot run this preset inference model.
[0118] In this embodiment, in the case where the total data processing capacity cannot run the preset inference model corresponding to the minimum number of parameters among multiple preset inference models, the broadcast processor adds a request, so that a new processor is added to the processor cluster. Then, the processor cluster with the new processor added may be able to meet at least one preset inference model, which can improve the data processing capacity of the processor cluster.
[0119] In this embodiment, by broadcasting a processor addition request in the case where the total data processing capacity cannot run the preset inference model corresponding to the minimum number of parameters among multiple preset inference models, receiving an approval message from the new processor, and in response to the approval message, adding the new processor to the processor cluster to utilize the new processor to perform part of the processing of the inference task, in this way, the data processing capacity of the processor cluster can be improved.
[0120] In a possible implementation manner, for the new processor added to the processor cluster, virtual resources can be sent to this new processor to encourage other processors to also be able to join the processor cluster to improve the global data processing capacity of the processor cluster.
[0121] It should be noted that in this embodiment, the processing method of the target processor's inference task can include an online mode and an offline mode.
[0122] In the online mode, for the same batch of data, it is hoped to execute the complete inference process and generate inference results in the shortest time. In this mode, the task submission strategy is as follows: The target processor remains in the online mode for a long time to avoid performance loss caused by frequently starting to process the inference task.
[0123] In the offline mode, after the target processor starts processing the inference task, and when there is sufficient storage space, all the data to be processed in the current layer is fully processed, avoiding frequent model loading by the target processor, thereby improving the inference performance.
[0124] It should be noted that in this embodiment, different inference tasks can be distinguished by a unique task identifier, and subsequent intermediate data files generated by the inference tasks are named with the same task identifier. Specifically, in the way of "model name" + "stage Index", it can be determined which model to load when scheduling tasks, and the loaded module is determined according to the stage index. Each data processing performs inference calculations based on the video memory of the GPU and the number of parameters of the module. After each data processing is completed, the generated intermediate data is placed in the directory of the last layer index + 1 of the module loaded this time, so as to know which module the intermediate data is output from, and then when loading the next module, the intermediate data is obtained from this directory for further processing.
[0125] In this embodiment, each inference task is distinguished by a unique task identifier. This identifier can be a task ID, a task name, or other unique identifier. For example, the task identifier can be "task_001". Each task may involve multiple models or different stages (modules) of a model. The way of "model name" + "stage Index" is used to determine which model or module to load specifically. For example, if there is a model called "resnet" and the task is its first stage, it can be expressed as "resnet_0". Then, before performing the inference calculation, it is necessary to evaluate the current video memory situation of the GPU and the number of parameters of the model. Based on the video memory and the number of model parameters, the number of models that can be processed in parallel or the amount of data to be processed in batches is determined. Then, the specified model or module is loaded, and the input data is processed. After the inference calculation is completed, an intermediate data file is generated. The name or path of the intermediate data file should be able to reflect its source (which module) and order. The way of "module name_stage Index_output Index" can be used to name it, where "output Index" represents the number of times the module outputs. The intermediate data is stored in a general directory, and this directory is named according to the task identifier. The output data of each module is stored in a subdirectory named with its "last layer index + 1". For example, if the last layer index of the module "resnet_0" is 3, then its output data is stored in the directory "task_001 / 4 / " (because the index starts from 0, so +1). Then a processing process can be as follows:
[0126] 1. Load the model and process the data:
[0127] Load the "resnet_0" model. Then, process the input data and store the intermediate data in the directory "task_001 / 4 / " (assuming the index of the last layer of "resnet_0" is 3).
[0128] 2. Load the next module:
[0129] Before loading the "resnet_1" model, read the intermediate data from the directory "task_001 / 4 / ". Then, process the intermediate data and store the newly generated intermediate data in the directory "task_001 / 5 / " (assuming the index of the last layer of "resnet_1" is 4).
[0130] 3. Continue with subsequent processing:
[0131] If there are more modules to be processed, repeat the above steps.
[0132] Finally, all intermediate data and the final result are stored in the total directory named after the task identifier, and the output data of each module is stored in its corresponding sub-directory.
[0133] Please refer to Figure 4 , Figure 4 which is a schematic structural diagram of a model inference device for a processor cluster provided in an embodiment of the present invention. As Figure 4 shown, the model inference device of the processor cluster is applied to a management node. As Figure 4 shown, the device may include:
[0134] An acquisition module 410, configured to acquire first data related to an inference task; acquire relevant information of each target processor among a plurality of target processors, where the relevant information includes available resource information and idle periods, or the relevant information includes target periods that meet data processing capacity requirements, the available resource information indicates the data processing capacity of the target processor, and the plurality of target processors are at least part of the plurality of processors; a determination module 420, configured to determine the order of relevant processing for each target processor to execute the inference task based on the idle periods or target periods of each target processor; a processing module 430, configured to sequentially instruct the plurality of target processors to respectively load a plurality of preset modules of a target inference model based on the order of relevant processing for each target processor to execute the inference task, so as to perform data processing by each target processor based on the preset module loaded by it during its corresponding idle period or target period, where the first data is used as the input of the target processor indicated for the first time.
[0135] The model inference device for a processor cluster provided in an embodiment of the present invention has the same technical features as the model inference method for a processor cluster provided in the above embodiment, so it can also solve the same technical problems and achieve the same technical effects. The model inference device for a processor cluster in this embodiment may refer to the description of the model inference method for a processor cluster and will not be elaborated herein.
[0136] This embodiment also provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned model inference method for a processor cluster. This electronic device can be a server or a terminal device.
[0137] See Figure 5 As shown, this electronic device includes a processor 100 and a memory 101. The memory 101 stores computer-executable instructions that can be executed by the processor 100, and the processor 100 executes the computer-executable instructions to implement the above-mentioned model inference method for a processor cluster.
[0138] Furthermore, Figure 5 the electronic device shown also includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.
[0139] Among them, the memory 101 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 5 only a bidirectional arrow is used in
[0140] The processor 100 may be an integrated circuit chip with the ability to process signals. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 100 or instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or executed and completed by a combination of hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.
[0141] The processor in the above electronic device can implement the steps in the model inference method of the above processor cluster by executing computer-executable instructions.
[0142] This embodiment also provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the model inference method of the above processor cluster.
[0143] The computer-executable instructions stored in the above computer-readable storage medium can implement the steps in the model inference method of the above processor cluster by executing the computer-executable instructions.
[0144] This embodiment also provides a computer program product, including program code, and the instructions included in the program code can be used to execute the methods in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated here.
[0145] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0146] In addition, in the description of the embodiments of the present invention, unless otherwise clearly specified and limited, the terms "install", "connect", and "couple" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two components. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.
[0147] If the functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or a part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0148] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore cannot be understood as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0149] Finally, it should be noted that the above embodiments are only specific implementation manners of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the technical field can still modify the technical solutions described in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.
Claims
1. A model reasoning method for a processor cluster, characterized in that: The processor cluster includes a plurality of processors, and the method includes: Acquiring first data related to the reasoning task; Acquire relevant information of each target processor among a plurality of target processors, the relevant information including available resource information and an idle period, or the relevant information including a target period that meets a data processing capability requirement, the available resource information indicating the data processing capability of the target processor, the plurality of target processors being at least part of the plurality of processors; Determining the order in which each target processor performs related processing of the inference task based on an idle period or a target period of each target processor; Based on the order in which each target processor performs related processing of the reasoning task, the multiple target processors are instructed in turn to load multiple preset modules of the target reasoning model respectively, so that each target processor performs data processing based on the preset modules loaded during its corresponding idle period or target period, wherein the first data is used as the input of the target processor instructed for the first time.
2. The method according to claim 1, characterized in that The step of sequentially instructing the plurality of target processors to load a plurality of preset modules of the target reasoning model respectively based on the order in which the target processors perform the relevant processing of the reasoning task comprises: Determine a first target processor and a second target processor among the plurality of target processors based on an order in which each target processor performs related processing of the inference task, wherein the order of the first target processor is before the order of the second target processor; When the idle time period or the target time period corresponding to the first target processor is reached, instructing the first target processor to load a first partial module of the target reasoning model, so as to process the first data based on the first partial module loaded by the first target processor to obtain second data, where the first partial module is a partial module of the multiple preset modules; When the idle period or target period corresponding to the second target processor is reached, the second target processor is instructed to load the second part module of the target reasoning model, so that the second target processor processes the second data based on the second part module loaded by it to obtain third data, wherein the second part module is a module other than the first part module among the multiple preset modules, and in the target reasoning model, the processing data order of the second part module is located after the processing data order of the first part module.
3. The method according to claim 1, characterized in that The method further comprises: Initiate a consensus vote to the processor cluster to obtain voting results from each processor in the processor cluster, wherein the voting results include participating processors voted by each processor for participating in the reasoning task, and the participating processors are at least part of the multiple processors; The plurality of target processors are selected from the plurality of processors based on the voting results of the respective processors.
4. The method according to claim 3, characterized in that The selecting the plurality of target processors from the plurality of processors based on the voting results of each processor comprises: Determine the participating processor from the voting results of each processor as the target processor; or, The voting result also includes a target number corresponding to each participating processor, and the target number is used to indicate the number of times the voting processor and the participating processor jointly process the same reasoning task. Among the participating processors in the voting results from each processor, the participating processor whose corresponding target number is greater than the number threshold is determined as the target processor.
5. The method according to any one of claims 1 to 4, characterized in that The method further comprises: Determining total resource information of the processor cluster, where the total resource information is used to indicate a total data processing capability of the processor cluster; Based on the total resource information, the target reasoning model is determined from a plurality of preset reasoning models, wherein the target reasoning model is a preset reasoning model having the largest corresponding parameter amount among the preset reasoning models that can be run by the total data processing capability.
6. The method according to claim 5, characterized in that The method further comprises: In a case where the total data processing capacity cannot run a preset inference model corresponding to a minimum parameter amount among multiple preset inference models, broadcasting a target processor addition request, wherein the target processor addition request is used to request to add a new processor to the processor cluster; receiving a consent message from a new processor, wherein the consent message is used to indicate that the new processor agrees to join the processor cluster; In response to the consent message, the new processor is added to the processor cluster to utilize the new processor to perform a portion of processing of the inference task.
7. The method according to any one of claims 1 to 4, characterized in that The obtaining of relevant information of each target processor among the multiple target processors includes: Acquire historical operation information of each target processor among the multiple target processors, and determine relevant information of each target processor based on the historical operation information of each target processor; or, A related information acquisition request is sent to each of the multiple target processors, and related information sent by each target processor is received, wherein the related information is determined by each target processor in response to the related information acquisition request based on its historical operation information.
8. A model reasoning device for a processor cluster, characterized in that: The processor cluster includes a plurality of processors, and the device includes: an acquisition module, configured to acquire first data related to the reasoning task; acquire relevant information of each target processor among a plurality of target processors, wherein the relevant information includes available resource information and an idle period, or the relevant information includes a target period that meets the data processing capability requirement, wherein the available resource information indicates the data processing capability of the target processor, and the plurality of target processors are at least part of the plurality of processors; A determination module, used for determining the order in which each target processor performs the related processing of the reasoning task based on the idle period or target period of each target processor; A processing module is used to instruct the multiple target processors to load multiple preset modules of the target reasoning model respectively based on the order in which each target processor performs related processing of the reasoning task, so that each target processor performs data processing based on the preset modules loaded during its corresponding idle period or target period, wherein the first data is used as input of the target processor instructed for the first time.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 7.