Multiprocessor reasoning task processing method and device and electronic equipment

The modules of artificial intelligence models are loaded and processed by multiple processors, and the success rate of inference tasks is solved.

CN120066697APending Publication Date: 2025-05-30SHANGHAI WERIDE AUTOMOTIVE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411953299.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-27
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

As the amount of parameters of the artificial intelligence model increases, the processing failure rate of a single processor when performing inference tasks increases, resulting in a decrease in the processing success rate of inference tasks.

Method used

By obtaining the first data related to the inference task and instructing the multiple processors to load multiple modules of the inference model, the third data is obtained by processing the first data based on the modules loaded by each processor in the multiple processors. The method can load modules sequentially or separately through multiple processors, and assign module loading tasks according to the available resource information of the processor.

Benefits of technology

Using multiple processors to jointly execute inference tasks reduces processing failures caused by a single processor to perform inference tasks, thereby improving the success rate of inference task processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066697A_ABST
    Figure CN120066697A_ABST
Patent Text Reader

Abstract

The invention provides a multiprocessor reasoning task processing method and device and electronic equipment, and relates to the technical field of artificial intelligence. First data related to a reasoning task is obtained; and indicating the plurality of processors to load a plurality of preset modules of a reasoning model, so that each processor in the plurality of processors processes the first data based on the loaded preset module to obtain third data, and the reasoning model is used for executing the reasoning task, so that the plurality of processors can be utilized to jointly execute the reasoning task, and the reasoning efficiency is improved. Therefore, the situation that reasoning task processing fails due to the fact that the reasoning task is executed through a single processor can be reduced, and the reasoning task processing success rate can be increased.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method, device, electronic device, and storage medium for processing inference tasks of multiple processors. Background Art

[0002] With the rapid development of artificial intelligence, more and more scenarios utilize artificial intelligence to process inference tasks, such as processing image tasks, text tasks, or voice tasks, etc.

[0003] Currently, when using artificial intelligence to process inference tasks, it is necessary to load a trained model, and then use the trained model to process the inference task.

[0004] However, as the number of parameters of the model increases, there are situations where the inference task processing fails. Summary of the Invention

[0005] In view of this, an object of the present invention is to provide a method, device, electronic device, and storage medium for processing inference tasks of multiple processors, so as to improve the success rate of inference task processing.

[0006] In a first aspect, an embodiment of the present invention provides a method for processing inference tasks of multiple processors, including: obtaining first data related to an inference task; instructing multiple processors to load multiple modules of an inference model, so as to process the first data by each processor among the multiple processors based on the module it loads, and obtain third data, where the inference model is used to execute the inference task.

[0007] In a possible implementation manner, instructing multiple processors to load multiple modules of an inference model, so as to process the first data by each processor among the multiple processors based on the module it loads, and obtain third data, includes: instructing each processor among the multiple processors to sequentially load multiple modules of the inference model, so as to process the first data by each processor among the multiple processors based on the modules sequentially loaded by it, and obtain third data; or, instructing multiple processors to respectively load multiple modules of the inference model, so as to process the first data by the modules respectively loaded by the multiple processors, and obtain third data.

[0008] In a possible implementation, it is indicated that each of multiple processors sequentially loads multiple modules of an inference model to process first data through the modules sequentially loaded by each of the multiple processors, so as to obtain third data, including: indicating that each of the multiple processors loads a first partial module of the inference model to process the first data through the first partial module loaded by each of the multiple processors, so as to obtain second data, where the first partial module is part of the multiple modules; indicating that each of the multiple processors loads a second partial module of the inference model to process the second data through the second partial module loaded by each of the multiple processors, so as to obtain third data, where the second partial module is the module other than the first partial module among the multiple modules, and in the inference model, the data processing order of the second partial module is after the data processing order of the first partial module.

[0009] In a possible implementation, indicating that each of the multiple processors loads a first partial module of the inference model includes: obtaining the available resource information of each of the multiple processors, where the available resource information is used to indicate the data processing capability of the processor; based on the available resource information of each processor, indicating that each of the multiple processors loads a first partial module of the inference model, where the first partial module is the module that the processor with the worst data processing capability among the multiple processors can load and run.

[0010] In a possible implementation, indicating that multiple processors respectively load multiple modules of an inference model to process first data through the modules respectively loaded by the multiple processors, so as to obtain third data, includes: indicating that a first processor among the multiple processors loads a third partial module of the inference model to process the first data through the third partial module loaded by the first processor, so as to obtain second data, where the third partial module is part of the multiple modules; indicating that a second processor among the multiple processors loads a fourth partial module of the inference model to process the second data through the fourth partial module loaded by the second processor, so as to obtain third data, where the fourth partial module is the module other than the third partial module among the multiple modules, and in the inference model, the data processing order of the fourth partial module is after the data processing order of the third partial module, and the first processor and the second processor are different processors.

[0011] In a possible implementation, the method further includes: determining the idle periods of each of the multiple processors, where the idle periods include a first idle period and a second idle period, and the second idle period is the idle period after the first idle period; using the processor corresponding to the first idle period as the first processor, and using the processor corresponding to the second idle period as the second processor.

[0012] In a possible implementation, the method of instructing multiple processors to load multiple modules of an inference model includes: querying multiple idle processors from the registered processors, and instructing the multiple idle processors to load the multiple modules of the inference model; the method further includes: in response to a first instruction to register a new processor, adding the new processor indicated by the first instruction as a registered processor to execute an inference task through the new processor; or, in response to a second instruction to remove a registered processor, removing the registered processor indicated by the second instruction from the list of registered processors, where the list of registered processors includes the registered processors, and the processor removed from the list of registered processors does not execute the inference task.

[0013] In a second aspect, an embodiment of the present invention provides an inference task processing apparatus for multiple processors, including: an acquisition module, configured to acquire first data related to an inference task; a data processing module, configured to instruct multiple processors to load multiple modules of an inference model, so that each processor among the multiple processors processes the first data based on the module it loads to obtain third data, and the inference model is used to execute the inference task.

[0014] In a third aspect, an embodiment of the present invention provides an electronic device, including a processor and a memory, where the memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the method of the first aspect.

[0015] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, where the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions cause the processor to implement the method of the first aspect.

[0016] The embodiments of the present invention bring the following beneficial effects: by acquiring first data related to an inference task; instructing multiple processors to load multiple preset modules of an inference model, so that each processor among the multiple processors processes the first data based on the preset module it loads to obtain third data, and the inference model is used to execute the inference task, in this way, it is possible to use multiple processors to jointly execute the inference task, thereby reducing the situation of inference task processing failure caused by executing the inference task through a single processor, and thus improving the success rate of inference task processing.

[0017] Other features and advantages of the present invention will be described in the following specification, and, in part, will be obvious from the specification, or will be understood by implementing the present invention. The objectives and other advantages of the present invention are achieved and obtained by the structures specifically pointed out in the specification, claims, and drawings.

[0018] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following specifically describes preferred embodiments in conjunction with the accompanying drawings as follows. Description of the Drawings

[0019] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those skilled in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 Schematic diagram of the architecture of a multi-processor inference task processing system provided by an embodiment of the present invention;

[0021] Figure 2 Schematic diagram of the flow of a multi-processor inference task processing method provided by an embodiment of the present invention;

[0022] Figure 3 Schematic diagram of the architecture of an inference model provided by an embodiment of the present application;

[0023] Figure 4 Schematic diagram of the architecture of another multi-processor inference task processing system provided by an embodiment of the present invention;

[0024] Figure 5 Schematic diagram of the architecture of another multi-processor inference task processing system provided by an embodiment of the present invention;

[0025] Figure 6 Schematic diagram of the structure of a multi-processor inference task processing device provided by an embodiment of the present invention;

[0026] Figure 7 Schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. Detailed Embodiments

[0027] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0028] Currently, more and more scenarios utilize artificial intelligence to process inference tasks, such as processing image tasks, text tasks, or speech tasks, etc. When using artificial intelligence to process inference tasks, it is necessary to load a trained model and then use the trained model to process the inference task.

[0029] However, as the number of parameters of the model increases, the performance of the processor may not be able to keep up with the requirements of the number of parameters, which may lead to the failure of inference task processing. In order to be able to normally execute the processing of inference tasks, it is necessary to replace the processor with more powerful performance. However, the cost of replacing the processor with more powerful performance is also higher.

[0030] In view of this, the embodiments of the present application provide a method, device, electronic device, and storage medium for processing inference tasks with multiple processors to improve the success rate of inference task processing. And even with a processor with relatively poor performance, it can also run a model with a large number of parameters, thereby realizing the normal processing of inference tasks.

[0031] For the convenience of understanding this embodiment, first, the system framework disclosed in the embodiments of the present invention will be described. As Figure 1 shown, the inference task processing system with multiple processors may include a management node (master), multiple processors, and a database. Among them, the management node can be used for the management of multiple processors, such as instructing the model modules loaded by multiple processors, etc. The multiple processors are used to jointly execute inference tasks. The database is used to store data related to inference tasks, such as storing the parameters of the model, storing the first data for inference, and storing the intermediate data generated during the inference process, etc. Among them, the database can include a local database or a database stored in the cloud.

[0032] Optionally, the system may further include a scheduling node (scheduler). The scheduling node may run a scheduling program to schedule the master, such as periodically polling the Master to check if there are idle nodes, so as to submit inference tasks.

[0033] It should be noted that the management node and the multiple processors are independent of each other, and one of the processors can also be used as the management node, which is not limited here. In addition, the scheduling program and the management node can be independent or on the same node, which is not limited here.

[0034] Next, a method for processing inference tasks with multiple processors will be introduced in detail. Please refer to Figure 2 , Figure 2Schematic flowchart of a method for processing inference tasks of a multi-processor provided by an embodiment of the present invention. The method for processing inference tasks of a multi-processor in one embodiment of the present disclosure can run on a management node. The method includes the following steps:

[0035] S210. Obtain first data related to the inference task.

[0036] Among them, the inference task refers to the process in which the model derives a conclusion or data based on the input information or data through internal logic and algorithms. In this embodiment, the inference task may include, but is not limited to, text processing tasks, image processing tasks, or speech processing tasks, etc. Optionally, the text processing task may be, for example, the matching between texts. For example, the model finds the matching text in the database based on the text input by the user. The image processing task may be, for example, after the model processes the input image, outputs the decision for autonomous driving or the denoised image, etc. The speech processing task may be, for example, the model outputs the speech text from the input speech. Optionally, one application scenario may be, for example: generating a description text of a picture, matching the text input by the user with the description text, so as to find the picture matching the text input by the user. Then the inference task may be, for example, generating description texts of multiple pictures, or matching the text input by the user with the description text, so as to find the picture matching the text input by the user, etc., which is not limited herein.

[0037] In this embodiment, the first data may be the data input to the model, or the intermediate data obtained after the model processes the data input to the model. Taking the text processing task as an example, the first data may be the text input to the model, or the text vector obtained by processing the text. Taking image processing as an example, the first data may be the input image data, or the image features obtained by feature extraction of the image, which is not limited herein.

[0038] S220. Instruct multiple processors to load multiple modules of the inference model, so that each processor among the multiple processors processes the first data based on the module it loads, and obtain third data. The inference model is used to execute the inference task.

[0039] The third data may be the data finally output by the inference model, such as the decision result of autonomous driving or the denoised image, etc., or it may be the intermediate data result of the inference model, that is, the inference model needs to continue to process the third data to output the final data. In this embodiment, the processor may be at least one of a graphics processing unit (GPU), a central processing unit (CPU), or a neural processing unit (NPU) with data processing capabilities. Taking the inference task as a text processing task as an example, the inference model may include a text processing model; taking the inference task as an image processing task as an example, the inference model may include an image processing model; taking the inference task as a voice processing task as an example, the inference model may include a voice processing model.

[0040] In this embodiment, by obtaining the first data related to the inference task; instructing multiple processors to load multiple preset modules of the inference model, so that each processor among the multiple processors processes the first data based on the preset module it loads to obtain the third data, and the inference model is used to execute the inference task. In this way, multiple processors can be used to jointly execute the inference task, thereby reducing the situation where the inference task processing fails due to executing the inference task by a single processor, and thus improving the success rate of inference task processing.

[0041] It should be noted that multiple processors may load the same or different modules when executing the inference task. Hereinafter, the cases where multiple processors may load the same or different modules when executing the inference task will be described separately.

[0042] First, the case where multiple processors may load the same module when executing the inference task will be described.

[0043] In a possible implementation manner, instructing multiple processors to load multiple modules of the inference model, so that each processor among the multiple processors processes the first data based on the module it loads to obtain the third data, includes:

[0044] Instructing each processor among the multiple processors to sequentially load multiple modules of the inference model, so that each processor among the multiple processors processes the first data based on the modules sequentially loaded by it to obtain the third data.

[0045] In this embodiment, if there are multiple inference tasks, the first data also includes the first data of each of the multiple inference tasks. Optionally, each processor among the multiple processors may sequentially load the same module, and different processors may process the first data of different inference tasks, and then obtain the third data corresponding to different inference tasks.

[0046] In this embodiment, by instructing each of the multiple processors to sequentially load multiple modules of the inference model, so as to process the first data through the modules sequentially loaded by each of the multiple processors to obtain the third data. In this way, it is possible to enable each of the multiple processors to obtain the corresponding module for processing through a single instruction, which can improve the convenience of processing the inference task.

[0047] In a possible implementation, instructing each of the multiple processors to sequentially load multiple modules of the inference model, so as to process the first data through the modules sequentially loaded by each of the multiple processors to obtain the third data, includes:

[0048] Instructing each of the multiple processors to load the first part of the modules of the inference model, so as to process the first data through the first part of the modules loaded by each of the multiple processors to obtain the second data, where the first part of the modules is part of the multiple modules; instructing each of the multiple processors to load the second part of the modules of the inference model, so as to process the second data through the second part of the modules loaded by each of the multiple processors to obtain the third data, where the second part of the modules is the modules other than the first part of the modules among the multiple modules, and in the inference model, the processing data order of the second part of the modules is after the processing data order of the first part of the modules.

[0049] In this embodiment, the first part of the modules may include one or more modules. The second part of the modules may include one or more modules. In the inference model, the processing data order of the second part of the modules is after the processing data order of the first part of the modules. That is to say, in the process of running the inference model to process the inference task, the data is first processed by the first part of the modules, and then processed by the second part of the modules. Therefore, in the inference model, the first part of the modules and the second part of the modules may be adjacent modules in the inference model, that is to say, the output of the first part of the modules can be used as the input of the second part of the modules. In this embodiment, the inference model may include the first part of the modules and the second part of the modules, and may also include other modules. That is to say, the first part of the modules plus the second part of the modules may be part or all of the modules of the inference model, which is not limited here. Among them, the second data may be obtained by processing the first data, so the second data can be understood as intermediate data. The third data may be the data finally output by the inference model, such as the decision result of autonomous driving, or the denoised image, etc., or may be the intermediate data result of the inference model, that is, the inference model needs to continue to use this third data for processing to output the final data.

[0050] Exemplarily, assume that multiple processors include processor A and processor B, and the inference tasks include inference task A and inference task B. Then, it is indicated that processor A and processor B respectively load the first part of the module. Then, processor A processes the first data corresponding to inference task A to obtain the second data corresponding to inference task A, and processor B processes the first data corresponding to inference task B to obtain the second data corresponding to inference task B. Then, it is indicated that processor A and processor B replace the module they loaded from the first part of the module with the second part of the module. Then, processor A uses the second part of the module it loaded to process the second data corresponding to inference task A to obtain the third data corresponding to inference task A; and processor B uses the second part of the module it loaded to process the second data corresponding to inference task B to obtain the third data corresponding to inference task B.

[0051] In a possible implementation manner, indicating that each processor among multiple processors loads the first part of the inference model includes:

[0052] Obtain the available resource information of each processor among the multiple processors. The available resource information is used to indicate the data processing capability of the processor; based on the available resource information of each processor, indicate that each processor among the multiple processors loads the first part of the inference model, where the first part of the module is a module that the processor with the worst data processing capability among the multiple processors can load and run.

[0053] Among them, the available resource information may include, but is not limited to, the number of cores, operating frequency, etc. If the processor includes a graphics processor, the available resource information may also include the video memory size. It should be noted that the available resource information may be the remaining resource information, and the available resource is less than or equal to the maximum resource, and the maximum resource indicates the maximum data processing capability of the processor.

[0054] In this embodiment, by obtaining the available resource information based on each processor, indicating that each processor among the multiple processors loads the first part of the inference model, where the first part of the module is a module that the processor with the worst data processing capability among the multiple processors can load and run. In this way, the indicated first part of the module can be loaded and run by each processor, which can reduce the situation that some processors cannot load and run due to the selected part of the module being too large, thereby improving the processing efficiency of the multi-processor joint inference task.

[0055] It should be understood that the description of indicating to load the second part of the inference model can refer to the description of indicating to load the first part of the inference model, and no limitation is made here.

[0056] Next, an explanation will be given on how multiple processors can load different modules when performing inference tasks.

[0057] In a possible implementation, it is to instruct multiple processors to load multiple modules of an inference model, so that each of the multiple processors processes first data based on the modules it loads to obtain third data, including:

[0058] Instruct multiple processors to separately load multiple modules of an inference model, so that the multiple processors process first data through the modules they separately load to obtain third data.

[0059] In this embodiment, the inference task can be one or more.

[0060] In this embodiment, by instructing multiple processors to separately load multiple modules of an inference model, so that the multiple processors process first data through the modules they separately load to obtain third data, in this way, the different performances of different processors can be utilized, and thus the performance utilization rate of inference task processing can be improved.

[0061] In a possible implementation, it is to instruct multiple processors to separately load multiple modules of an inference model, so that the multiple processors process first data through the modules they separately load to obtain third data, including:

[0062] Instruct the first processor among the multiple processors to load the third part of the modules of the inference model, so that the first processor processes first data through the third part of the modules it loads to obtain second data, and the third part of the modules is part of the multiple modules;

[0063] Instruct the second processor among the multiple processors to load the fourth part of the modules of the inference model, so that the second processor processes the second data through the fourth part of the modules it loads to obtain third data, and the fourth part of the modules is the modules other than the third part of the modules among the multiple modules. In the inference model, the data processing order of the fourth part of the modules is after the data processing order of the third part of the modules, and the first processor and the second processor are different processors.

[0064] In this embodiment, the third part of the module may include one or more modules. The fourth part of the module may include one or more modules. In the inference model, the processing data order of the fourth part of the module is after that of the third part of the module. That is to say, in the process of running the inference model to process the inference task, the data is first processed by the third part of the module, and then processed by the fourth part of the module. Therefore, in the inference model, the third part of the module and the fourth part of the module may be adjacent modules in the inference model. That is to say, the output of the third part of the module can be used as the input of the fourth part of the module. In this embodiment, the inference model may include the third part of the module and the fourth part of the module, and may also include other modules. That is to say, the third part of the module plus the fourth part of the module may be part or all of the modules of the inference model, which is not limited here.

[0065] Exemplarily, assuming that the multiple processors include processor A and processor B, processor A can be instructed to load the third part of the module, and processor B can be instructed to load the fourth part of the module. Then, processor A can use the third part of the module it loads to process the first data to obtain the second data, and then processor B can use the fourth part of the module it loads to process the second data to obtain the third data.

[0066] It should be noted that which is the first processor and which is the second processor can be determined according to the idle periods of the processors or the available resource information of the processors.

[0067] In a possible implementation manner, the method further includes:

[0068] Determine the idle periods of the processors in the multiple processors. The idle periods include the first idle period and the second idle period, and the second idle period is the idle period after the first idle period; use the processor corresponding to the first idle period as the first processor, and use the processor corresponding to the second idle period as the second processor.

[0069] In this embodiment, use the processor with an earlier idle period as the first processor to load the first part of the module that is more forward, and use the processor with a later idle period as the second processor to load the second part of the module that is more backward. In this way, the modules can be loaded separately in combination with the idle periods of the processors, which is beneficial to improving the resource utilization rate of the processors.

[0070] In a possible implementation, the idle period of the processor can be determined based on the historical running information of the processor. Optionally, the historical running information includes the resource utilization rate corresponding to each preset time period, and the preset time period corresponding to the resource utilization rate lower than the resource utilization rate threshold is determined as the idle time period of the processor. It should be understood that the resource utilization rate threshold can be set as needed, for example, set to 10% etc., and there is no limitation here. Exemplarily, the preset time period can be set as needed. For example, a day is divided into 24 time periods, the minimum time and the maximum time in each time period are separated by one hour, and there is no overlap in time between any two time periods.

[0071] In another possible implementation, the first processor and the second processor can also be determined based on the available resource information of each processor. Exemplarily, if the number of parameters of the first part of the module is greater than the number of parameters of the second part of the module, the processor with greater available resource information among the multiple processors is used as the first processor, and the processor with smaller available resource information is used as the second processor. Another example is that if the number of parameters of the first part of the module is less than the number of parameters of the second part of the module, the processor with smaller available resource information among the multiple processors is used as the first processor, and the processor with greater available resource information is used as the second processor.

[0072] For ease of understanding, the following uses a specific model to exemplarily illustrate this solution. Please refer to Figure 3 , Figure 3 which is a schematic diagram of the architecture of an inference model provided by an embodiment of this application. As Figure 3 shown, the inference model can be a model with a transformer architecture.

[0073] Among them, models with a transformer architecture all include a language model, and different models have different numbers of layers. The model for image processing includes the following modules:

[0074] Vision Tower: Responsible for extracting image features (features). The number of layers included in the vision block is different in different models. Vision Tower usually refers to a neural network module dedicated to processing visual data, which can extract and encode image features. In a multimodal model, Vision Tower is combined with a module responsible for processing text data (such as a language model) to jointly achieve cross-modal information processing and generation.

[0075] Multi-modal Projector: Responsible for projecting the encoded image modal features into the text feature space to obtain aligned features, usually consisting of two layers. The main function of the multi-modal projector block is to project data of different modalities (such as images, text, etc.) into the same feature space for cross-modal information processing and fusion. In a multi-modal model, the multi-modal projector block is usually responsible for converting visual features into a format compatible with text features, thus enabling cross-modal interaction and generation.

[0076] Next, an Figure 3 exemplary description of the data processing process will be given.

[0077] 1. Input Embedding:

[0078] First, the input data (which may be text, images, or other sequential data) is converted into embedding vectors. These embedding vectors are usually learned and can capture the features of the input data.

[0079] 2. Positional Encoding:

[0080] Since the transformer model itself does not have the ability to handle sequence order, positional encoding needs to be added to identify the position of each element in the input sequence. The positional encoding can be added to the input embedding or directly included in the input embedding.

[0081] 3. Multi-Head Attention:

[0082] Next, the data passes through the multi-head attention layer. This layer splits the input vector into multiple heads, and each head independently calculates the attention weights and performs a weighted sum of the input vector based on these weights. The multi-head attention mechanism allows the model to simultaneously focus on different positions in the input sequence, thereby capturing more complex dependencies.

[0083] 4. Add & Normalization (add&nNorm):

[0084] After the multi-head attention layer, residual connections (i.e., adding the input to the output) and layer normalization are usually performed to improve the stability and performance of the model.

[0085] 5. Feed Forward Layer:

[0086] The data then passes through the feed-forward layer, which is a simple fully-connected neural network that performs a non-linear transformation on the vectors at each position.

[0087] 6. Add & normalize again:

[0088] Similar to after the multi-head attention layer, residual connections and layer normalization are also performed after the feed-forward layer.

[0089] 7. Linear layer:

[0090] After being processed by the feed-forward layer, the data is transformed through the linear layer in preparation for output.

[0091] 8. Softmax layer:

[0092] Finally, the data is normalized through the softmax layer to obtain the output probability distribution. The softmax layer converts the output of the linear layer into probability values such that the sum of all output values is 1.

[0093] 9. Output layer (output probabilities):

[0094] The output layer provides the final prediction result, that is, the probability distribution for each class.

[0095] Refer to Figure 3 In the model architecture shown, by way of example, the first data can be, for example, the input data, such as text, image, or other sequence data, the second data can be, for example, the data output by the multi-head attention, and the third data can be, for example, the data output by the output layer. Also by way of example, the first data can be, for example, the input data, the second data can be, for example, the data output by the positional encoding, and the third data can be, for example, the data output by the multi-head attention. Also by way of example, the first data can be, for example, the data output by the positional encoding, the second data can be, for example, the data output by the multi-head attention, and the third data can be, for example, the data output by the output layer.

[0096] It should be understood that the way of dividing the actual first - part module and second - part module is not limited to the examples shown above. The first - part module and second - part module can be divided as needed, for example, divided by function, and there is no limitation here. In a possible implementation, the model can also be divided according to the performance of the processor. Taking the processor as the GPU, when facing a GPU with limited video memory, it is crucial to reasonably divide the model. The following is a detailed strategy for dividing the model based on the size of the GPU video memory, especially for the two key modules of vision_tower and language_model. It should be noted that the occupancy of the GPU video memory mainly comes from two aspects:

[0097] 1. Loading model parameters: This includes static parameters such as network weights and biases, which are loaded into the video memory during model initialization.

[0098] 2. Runtime calculations: This includes dynamic processes such as forward propagation, backward propagation, and gradient updates, which occupy the video memory during model training or inference.

[0099] To effectively utilize the GPU video memory, a control parameter can be introduced to define the upper limit of the proportion of video memory occupancy during model runtime. This control parameter can be, for example, gpu_memory_utilization, which defines the upper limit of the proportion of video memory occupancy during model runtime. By default, this proportion is set to 0.5, meaning that the video memory occupancy during model runtime should not exceed half of the total GPU video memory.

[0100] For a GPU with 24G of video memory, if the model is loaded using the float32 data type, the number of parameters per model block should be less than 6 billion (6 * 10^9). This is because each parameter of the float32 type occupies 4 bytes, and other factors (such as gradients and activation values) also need to be considered for video memory occupancy.

[0101] When splitting the vision_tower block, the vision_tower block is split into multiple small blocks, and each small block contains k layers. The specific value of k depends on the number of parameters and computational complexity of each layer. To ensure that each small block can run within the video memory limit, the value of k can be gradually adjusted through experiments. When splitting, layers with larger numbers of parameters (such as convolutional layers and fully - connected layers) can be considered to be dispersed into different small blocks to balance the video memory occupancy.

[0102] When splitting the language_model block, similarly, the language_model block is also split into multiple small blocks, and each small block contains m layers. The value of m also needs to be adjusted according to the number of parameters and computational complexity of each layer. For language models, especially Transformer-based models, key components such as the attention mechanism and feed-forward network can be considered to be dispersed into different small blocks to optimize the use of video memory.

[0103] In a possible implementation, it is indicated that multiple processors load multiple modules of the inference model, including:

[0104] Query multiple idle processors from the registered processors, and indicate the multiple idle processors to load multiple modules of the inference model; the method further includes: in response to a first instruction to register a new processor, adding the new processor indicated by the first instruction as a registered processor to perform the inference task through the new processor; or, in response to a second instruction to remove a registered processor, removing the registered processor indicated by the second instruction from the list of registered processors, where the list of registered processors includes the registered processors, and the processors removed from the list of registered processors do not perform the inference task.

[0105] In this embodiment, each node joins the scheduling node by registering with the master node and waits for the inference task to be executed.

[0106] Next, an example will be given in combination with the registration of the inference task system of multiple processors. Please refer to Figure 4 , Figure 4 FIG. is a schematic diagram of the system architecture of another inference task processing system of multiple processors provided by the embodiment of the present invention. As Figure 4 shown in the inference task processing system of multiple processors, the multiple processors include processor 1 (Node-1), processor 2 (Node-2), and processor 3 (Node-3). Then, the multiple processors can request registration, and then the master adds the newly registered processors to the list. The processors in the list are the nodes that can be scheduled. Then, if a certain node is not required to process the inference task, the processor is removed from the list.

[0107] In this embodiment, by registering new processors or removing registered processors, the processors for executing the inference task can be dynamically determined, thereby improving the flexibility of the execution of the inference task.

[0108] It should be noted that in this embodiment, the processing mode of the processor inference task can include an online mode and an offline mode.

[0109] In the online mode, for the data of the same batch, it is hoped that the complete inference process can be executed in the shortest time to generate inference results. In this mode, the task submission strategy is as follows: The processor remains in the online mode for a long time in multiple cases to avoid the performance loss caused by frequently starting the processing of inference tasks.

[0110] In the offline mode, after the processor starts the processing of the inference task, and when the storage space is sufficient, it completely processes all the data that needs to be processed in the current layer, avoiding the processor from frequently loading the model, thereby improving the inference performance.

[0111] It should be noted that in this embodiment, different inference tasks can be distinguished by a unique task identifier, and the subsequent modified inference tasks are named after the same task identifier for the generated intermediate data files. Specifically, in the way of "model name" + "stage Index", it can be determined which model is loaded when scheduling tasks, and the loaded module is determined according to the stage index. Each data processing performs inference calculations based on the GPU's video memory and the number of parameters of the module. And after each data processing is completed, the generated intermediate data is placed in the directory of the last layer index + 1 of the module loaded this time, so as to know which module outputs the intermediate data, and then when loading the next module, the intermediate data is obtained from this directory for further processing.

[0112] In this embodiment, each inference task is distinguished by a unique task identifier. This identifier can be a task ID, a task name, or other unique identifier. For example, the task identifier can be "task_001". Each task may involve multiple models or different stages (modules) of a model. The method of using "model name" + "stage Index" is used to determine which specific model or module to load. For example, if there is a model named "resnet" and the task is its first stage, it can be represented as "resnet_0". Then, before performing the inference calculation, it is necessary to evaluate the video memory situation of the current GPU and the number of parameters of the model. Based on the video memory and the number of model parameters, the number of models that can be processed in parallel or the amount of data to be processed in batches is determined. Then, the specified model or module is loaded, and the input data is processed. After the inference calculation is completed, an intermediate data file is generated. The name or path of the intermediate data file should be able to reflect its source (which module) and order. The method of using "module name_stage Index_output Index" can be used for naming, where "output Index" represents the number of times this module outputs. The intermediate data is stored in a general directory, and this directory is named according to the task identifier. The output data of each module is stored in a subdirectory named after its "last layer index + 1". For example, if the last layer index of the module "resnet_0" is 3, then its output data is stored in the "task_001 / 4 / " directory (because the index starts from 0, so +1). Then, a single processing process can be as follows:

[0113] 1. Load the model and process the data:

[0114] Load the "resnet_0" model. Then, process the input data and store the intermediate data in the "task_001 / 4 / " directory (assuming the last layer index of "resnet_0" is 3).

[0115] 2. Load the next module:

[0116] Before loading the "resnet_1" model, read the intermediate data from the "task_001 / 4 / " directory. Then, process the intermediate data and store the newly generated intermediate data in the "task_001 / 5 / " directory (assuming the last layer index of "resnet_1" is 4).

[0117] 3. Continue subsequent processing:

[0118] If there are more modules to be processed, repeat the above steps.

[0119] Finally, all intermediate data and final results are stored in the general directory named after the task identifier, and the output data of each module is stored in its corresponding sub-directory.

[0120] Next, an example will be given in combination with the model of the inference task system of multiple processors, as well as the extraction and storage of data. Please refer to Figure 5 , Figure 5 FIG. [FIGURE NUMBER] is a schematic diagram of the system architecture of another inference task processing system of multiple processors provided by an embodiment of the present invention. As Figure 5 shown, the Scheduler will periodically poll the Master to check if there are idle nodes and submit tasks. Then, the master obtains the number of GPU cards and the video memory of the GPU on the node by querying the node information, so as to determine the processor for executing the inference task. How to determine the processor for executing the inference task can refer to the description of the above embodiment and will not be elaborated here. Then, the Scheduler will scan the public data directory and perform task scheduling according to the data of the corresponding layer.

[0121] Optionally, Ray can be selected as the distributed scheduling framework, including a Ray head node and multiple Ray worker nodes. The Ray head node can be a management node, and the Ray worker node can be a processor node. The Ray worker node connects to the Ray head node through a client. The Scheduler applies for the number of workers required to execute the inference task. One worker node runs part of the model stage task, and the worker node loads the model parameters from the shared storage.

[0122] Please refer to Figure 6 , Figure 6 FIG. [FIGURE NUMBER] is a schematic diagram of the structure of an inference task processing device of multiple processors provided by an embodiment of the present invention. As Figure 6 shown, this data processing device is applied to the management node. As Figure 6 shown, the device may include:

[0123] An acquisition module 610 for acquiring first data related to the inference task; a data processing module 620 for instructing multiple processors to load multiple modules of the inference model, so as to process the first data by each processor in the multiple processors based on the modules it loads to obtain third data, and the inference model is used to execute the inference task.

[0124] The inference task processing device for multiple processors provided by the embodiments of the present invention has the same technical features as the inference task processing method for multiple processors provided by the above embodiments, so it can also solve the same technical problems and achieve the same technical effects. The inference task processing device for multiple processors can refer to the description of the inference task processing method for multiple processors and will not be elaborated here.

[0125] This embodiment also provides an electronic device, including a processor and a memory. The memory stores computer-executable instructions that can be executed by the processor, and the processor executes the computer-executable instructions to implement the above-mentioned inference task processing method for multiple processors. This electronic device can be a server or a terminal device.

[0126] See Figure 7 As shown, this electronic device includes a processor 100 and a memory 101. The memory 101 stores computer-executable instructions that can be executed by the processor 100, and the processor 100 executes the computer-executable instructions to implement the above-mentioned inference task processing method for multiple processors.

[0127] Furthermore, Figure 7 the electronic device shown also includes a bus 102 and a communication interface 103. The processor 100, the communication interface 103, and the memory 101 are connected through the bus 102.

[0128] Among them, the memory 101 may include a high-speed random access memory (RAM, Random Access Memory), and may also include a non-volatile memory, such as at least one disk memory. Through at least one communication interface 103 (which can be wired or wireless), a communication connection is established between this system network element and at least one other network element, and the Internet, wide area network, local area network, metropolitan area network, etc. can be used. The bus 102 can be an ISA bus, a PCI bus, an EISA bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of representation, Figure 7 only a bidirectional arrow is used in [description] but it does not mean that there is only one bus or one type of bus.

[0129] The processor 100 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit of the hardware in the processor 100 or the instructions in the form of software. The above-mentioned processor 100 may be a general-purpose processor, including a central processing unit (CPU for short), a network processor (NP for short), etc.; it may also be a digital signal processor (DSP for short), an application specific integrated circuit (ASIC for short), a field-programmable gate array (FPGA for short), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute each method, step and logic block diagram disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present invention can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 101, and the processor 100 reads the information in the memory 101 and combines its hardware to complete the steps of the method in the foregoing embodiments.

[0130] The processor in the above electronic device can implement the steps in the above multi-processor inference task processing method by executing computer-executable instructions.

[0131] This embodiment also provides a computer-readable storage medium storing computer-executable instructions, which, when called and executed by a processor, cause the processor to implement the above multi-processor inference task processing method.

[0132] The computer-executable instructions stored in the above computer-readable storage medium can implement the steps in the above multi-processor inference task processing method by executing the computer-executable instructions.

[0133] This embodiment also provides a computer program product, including program code, and the instructions included in the program code can be used to execute the methods in the foregoing method embodiments. For specific implementation, reference can be made to the method embodiments, which will not be elaborated here.

[0134] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems and devices described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.

[0135] In addition, in the description of the embodiments of the present invention, unless otherwise clearly defined and limited, the terms "installed", "connected", and "coupled" should be understood in a broad sense. For example, it can be a fixed connection, a detachable connection, or an integral connection; it can be a mechanical connection or an electrical connection; it can be directly connected or indirectly connected through an intermediate medium, and it can be the communication inside two elements. For those skilled in the art, the specific meanings of the above terms in the present invention can be understood according to specific situations.

[0136] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or part of this technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods of the various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0137] In the description of the present invention, it should be noted that the orientation or positional relationship indicated by the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation to the present invention. In addition, the terms "first", "second", and "third" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance.

[0138] Finally, it should be noted that the above embodiments are only specific embodiments of the present invention, used to illustrate the technical solutions of the present invention, rather than limiting it. The protection scope of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the technical field can still modify the technical solutions recorded in the foregoing embodiments, or can easily think of changes, or perform equivalent replacements on some of the technical features; and these modifications, changes or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A method for processing inference tasks of a multi-processor, characterized in that: include: Acquiring first data related to the reasoning task; Instruct multiple processors to load multiple modules of the reasoning model so that each processor in the multiple processors processes the first data based on the modules loaded by the processor to obtain third data, and the reasoning model is used to execute the reasoning task.

2. The method according to claim 1, characterized in that The instructing multiple processors to load multiple modules of the inference model so that each processor in the multiple processors processes the first data based on the modules loaded thereby to obtain third data includes: Instructing each of the multiple processors to sequentially load multiple modules of the inference model, so as to process the first data through the modules sequentially loaded by each of the multiple processors to obtain third data; or, Instruct multiple processors to load multiple modules of the reasoning model respectively, so as to process the first data through the modules loaded by the multiple processors respectively, so as to obtain third data.

3. The method according to claim 2, characterized in that The instructing each of the plurality of processors to sequentially load the plurality of modules of the inference model, so as to process the first data through the modules sequentially loaded by each of the plurality of processors to obtain third data, comprises: Instructing each of the multiple processors to load a first part module of the inference model, so as to process the first data through the first part module loaded by each of the multiple processors to obtain second data, wherein the first part module is a part of the multiple modules; Instruct each of the multiple processors to load the second part module of the reasoning model, so as to process the second data through the second part module loaded by each of the multiple processors to obtain third data, wherein the second part module is a module among the multiple modules except the first part module, and in the reasoning model, the processing data sequence of the second part module is located after the processing data sequence of the first part module.

4. The method according to claim 3, characterized in that The instructing each of the plurality of processors to load a first part of modules of the inference model comprises: Acquire available resource information of each processor among the multiple processors, where the available resource information is used to indicate the data processing capability of the processor; Based on the available resource information of each processor, each of the multiple processors is instructed to load a first part of modules of the inference model, wherein the first part of modules is a module that can be loaded and run by a processor with the worst data processing capability among the multiple processors.

5. The method according to claim 2, characterized in that: The instructing a plurality of processors to respectively load a plurality of modules of the reasoning model, so as to process the first data through the modules respectively loaded by the plurality of processors to obtain third data, comprises: Instructing a first processor among the multiple processors to load a third partial module of the inference model, so as to process the first data through the third partial module loaded by the first processor to obtain second data, wherein the third partial module is a part of the multiple modules; Instruct a second processor among the multiple processors to load the fourth part module of the reasoning model, so as to process the second data through the fourth part module loaded by the second processor to obtain third data, the fourth part module is a module among the multiple modules except the third part module, in the reasoning model, the processing data sequence of the fourth part module is located after the processing data sequence of the third part module, and the first processor and the second processor are different processors.

6. The method according to claim 5, characterized in that The method further comprises: Determine an idle period of each processor among the plurality of processors, the idle period comprising a first idle period and a second idle period, the second idle period being an idle period after the first idle period; The processor corresponding to the first idle period is used as the first processor, and the processor corresponding to the second idle period is used as the second processor.

7. The method according to any one of claims 1 to 6, characterized in that The instructing multiple processors to load multiple modules of the inference model includes: Querying multiple idle processors from registered processors, and instructing the multiple idle processors to load multiple modules of the inference model; The method further comprises: In response to a first instruction for registering a new processor, adding the new processor indicated by the first instruction as a registered processor, so as to execute the inference task through the new processor; or, In response to a second instruction for removing a registered processor, the registered processor indicated by the second instruction is removed from a registered processor list, the registered processor list including the registered processors, and the processor removed from the registered processor list does not execute the inference task.

8. A multi-processor inference task processing device, characterized in that: include: An acquisition module, used to acquire first data related to the reasoning task; A data processing module is used to instruct multiple processors to load multiple modules of the reasoning model, so that each processor in the multiple processors processes the first data based on the modules loaded by it to obtain third data, and the reasoning model is used to execute the reasoning task.

9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer executable instructions that can be executed by the processor, and the processor executes the computer executable instructions to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are called and executed by a processor, the computer-executable instructions prompt the processor to implement the method according to any one of claims 1 to 7.