Model Inference Method, Computing Device, Medium, and Program Product Applied to a Computing Device

By configuring multiple data subsets in multiple computing units and implementing parallel processing, the pressure on the hyperscale deep learning model in terms of memory and inference speed is solved, and task processing efficiency is improved.

CN119902901BActive Publication Date: 2025-06-10INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510390615.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-10
Estimated Expiration
2045-03-31

AI Technical Summary

Technical Problem

Hyperscale deep learning models are under pressure in terms of memory and inference speed, and the memory of a single computing unit is limited, making it difficult to handle tasks efficiently.

Method used

By configuring at least two subsets of data in multiple computing units and processing the pending tasks sequentially on each computing unit, the subprocessing results are obtained and sent to other computing units simultaneously to realize parallel processing of computing and communication.

Benefits of technology

The communication time and data amount after the calculation unit is calculated is reduced, the idle time of computing resources is reduced, the computing efficiency and resource utilization are improved, and task processing efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119902901B_ABST
    Figure CN119902901B_ABST
Patent Text Reader

Abstract

The present invention provides a model inference method, a computing device, a medium, and a program product applied to a computing device, which can be applied to the field of artificial intelligence technology. The method includes: a control unit sends a task to be processed to a plurality of computing units respectively, and each computing unit is configured with at least two data subsets, where the data subsets are obtained by splitting a weight set of a target inference model required for executing the task to be processed; the plurality of computing units respectively execute the task to be processed to obtain a plurality of sub-processing results; the control unit obtains a target execution result according to the plurality of sub-processing results; any one of the plurality of computing units executes the task to be processed in the following manner: using the configured at least two data subsets, sequentially processes the task to be processed to obtain at least two sub-processing results, and when any one of the sub-processing results is obtained, synchronously sends the any one of the sub-processing results to other computing units among the plurality of computing units that are processing the task to be processed in parallel.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and particularly relates to a model inference method, a computing device, a medium, and a program product applied to a computing device. Background Art

[0002] With the successful application of deep learning models in various fields, people have begun to focus on how to scale deep learning models to a larger scale to improve the data processing ability, accuracy, and performance of the models. Based on this, ultra-large-scale deep learning models have emerged. Ultra-large-scale deep learning models face the pressure of memory and inference speed. However, the memory of a single computing unit is very limited. Therefore, a computing strategy in which multiple computing units jointly execute inference has been proposed. For example, in the computing strategy where multiple computing units jointly execute inference, it is crucial to reasonably utilize the computing resources of the computing units to achieve a quick response to users. Summary of the Invention

[0003] In view of the above problems, the present invention provides a model inference method, a computing device, a medium, and a program product applied to a computing device.

[0004] According to a first aspect of the present invention, there is provided a model inference method applied to a computing device. The computing device includes: a control unit and a plurality of computing units for parallelly executing tasks to be processed. The model inference method includes: the control unit separately sends the tasks to be processed to the plurality of computing units. Any one of the plurality of computing units is configured with at least two data subsets, and the data subsets are obtained by splitting the weight set of the target inference model required for executing the tasks to be processed; the plurality of computing units respectively execute the tasks to be processed to obtain a plurality of sub-processing results; the control unit obtains the target execution result of the tasks to be processed according to the plurality of sub-processing results; wherein, any one of the plurality of computing units executes the tasks to be processed in the following manner: using the configured at least two data subsets, sequentially processes the tasks to be processed to obtain at least two sub-processing results, and wherein, when obtaining any one of the sub-processing results, synchronously sends the any one of the sub-processing results to other computing units among the plurality of computing units that are parallelly processing the tasks to be processed.

[0005] A second aspect of the present invention provides a computing device, including: a control unit configured to send a task to be processed to a plurality of computing units respectively; a plurality of computing units, any one of the plurality of computing units is configured with at least two data subsets, and the data subsets are obtained by splitting a weight set of a target inference model required for executing the task to be processed; the plurality of computing units are configured to execute the task to be processed respectively to obtain a plurality of sub-processing results; the control unit is further configured to obtain a target execution result of the task to be processed according to the plurality of sub-processing results; wherein, any one of the plurality of computing units is configured to execute the task to be processed in the following manner: using the configured at least two data subsets, processing the task to be processed in sequence to obtain at least two sub-processing results, wherein, when any one of the sub-processing results is obtained, synchronously sending the any one of the sub-processing results to other computing units among the plurality of computing units that are processing the task to be processed in parallel.

[0006] A third aspect of the present invention further provides a computer-readable storage medium, on which a computer program or instruction is stored, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0007] A fourth aspect of the present invention further provides a computer program product, including a computer program or instruction, and when the computer program or instruction is executed by a processor, the steps of the above method are implemented.

[0008] According to an embodiment of the present invention, since at least two data subsets are configured in the computing unit, and the configured at least two data subsets are used to process the task to be processed in sequence, and when any one of the sub-processing results is obtained, the any one of the sub-processing results is synchronously sent to other computing units among the plurality of computing units that are processing the task to be processed in parallel, therefore, the calculation and communication of the computing unit are parallelized. When traversing the configured at least two data subsets, only the sub-processing result obtained by processing the task to be processed for the last time needs to be sent. Compared with the serial calculation and communication of the computing unit, since the sub-processing result obtained by processing the task to be processed for the last time is only a small part of the sub-processing results of the computing unit executing the task to be processed, therefore, when the task to be processed is completed, the communication time and communication data volume after the computing unit finishes the calculation can be reduced, thereby reducing the idle time of the computing resources of the computing unit, improving the calculation efficiency and resource utilization rate of the computing unit, and further improving the task processing efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Through the following description of the embodiments of the present invention with reference to the accompanying drawings, the above content and other objects, features and advantages of the present invention will become clearer.

[0010] Figure 1 FIG. shows an application scenario diagram of a model inference method, a computing device, a medium and a program product applied to a computing device according to an embodiment of the present invention.

[0011] Figure 2 The flowchart of the model inference method applied to a computing device according to an embodiment of the present invention is shown.

[0012] Figure 3A The schematic diagram of parallel execution of tasks to be processed according to an embodiment of the present invention is shown.

[0013] Figure 3B The schematic diagram of any computing unit executing a task to be processed according to an embodiment of the present invention is shown.

[0014] Figure 4 The schematic diagram of determining the target execution result of a task to be processed according to an embodiment of the present invention is shown.

[0015] Figure 5 The structural block diagram of a computing device according to an embodiment of the present invention is shown. Detailed implementation manners

[0016] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. However, it should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. In the following detailed description, for the sake of explanation, many specific details are set forth to provide a comprehensive understanding of the embodiments of the present invention. However, it is obvious that one or more embodiments can also be implemented without these specific details. In addition, in the following description, descriptions of well-known structures and technologies are omitted to avoid unnecessarily confusing the concepts of the present invention.

[0017] The terms used herein are merely for describing specific embodiments and are not intended to limit the present invention. The terms "including", "comprising", etc. used herein indicate the presence of features, steps, operations, and / or components, but do not exclude the presence or addition of one or more other features, steps, operations, or components.

[0018] All terms used herein (including technical and scientific terms) have the meanings commonly understood by those skilled in the art, unless otherwise defined. It should be noted that the terms used herein should be interpreted as having a meaning consistent with the context of this specification and should not be interpreted in an idealized or overly rigid manner.

[0019] In the case of using expressions such as "at least one of A, B, and C", generally, it should be interpreted according to the meaning commonly understood by those skilled in the art (for example, "a system having at least one of A, B, and C" should include, but is not limited to, a system having only A, only B, only C, having A and B, having A and C, having B and C, and / or having A, B, and C, etc.).

[0020] In the technical solution of the present invention, the user information involved (including but not limited to user personal information, user image information, user device information, such as location information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) are all information and data authorized by the user or fully authorized by all parties. Moreover, the processing of relevant data, such as collection, storage, use, processing, transmission, provision, invention, and application, all comply with relevant laws, regulations, and standards, adopt necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0021] In the scenario of making automated decisions using personal information, the methods, devices, and systems provided by the embodiments of the present invention all provide corresponding operation entrances for users to choose to agree or refuse the results of automated decisions; if the user chooses to refuse, the expert decision-making process will be entered. The expression "automated decision" here refers to the activity of automatically analyzing and evaluating an individual's behavior habits, hobbies, or economic, health, credit status, etc. through a computer program and making a decision. The expression "expert decision" here refers to the activity of making a decision by a person who specializes in a certain field, has specialized experience, knowledge, and skills, and has reached a certain professional level.

[0022] In the process of implementing the embodiments of the present invention, it is found that in the computing strategy where each computing unit jointly performs inference, how to reasonably utilize the computing power of the computing unit to achieve a quick response to the user is crucial. Although tensor parallelism realizes simultaneous inference on multiple cards through multi-card or multi-machine parallel computing to meet the deployment and inference requirements of ultra-large-scale models, the calculation and data communication of a single card in tensor parallelism are executed serially. After the calculation is completed, the calculation results are transmitted through inter-card communication to complete the reduction operation. Since the computing resources of the card are idle during the inter-card communication transmission, it will cause waste of the computing resources of the card. In addition, after the calculation is completed, if the calculation results need to occupy a large amount of transmission resources when transmitting the calculation results through inter-card communication, the user cannot be responded to in time, which will affect the user experience.

[0023] Based on this, an embodiment of the present invention provides a model inference method applied to a computing device. The computing device includes: a control unit and a computing unit for parallelly executing tasks to be processed. The model inference method includes: the control unit sends the tasks to be processed to a plurality of computing units respectively. Any one of the plurality of computing units is configured with at least two data subsets, and the data subsets are obtained by splitting the weight set of the target inference model required for executing the tasks to be processed; the plurality of computing units execute the tasks to be processed respectively to obtain a plurality of sub-processing results; the control unit obtains the target execution result of the tasks to be processed according to the plurality of sub-processing results; wherein, any one of the plurality of computing units executes the tasks to be processed in the following manner: using the configured at least two data subsets, processing the tasks to be processed in sequence to obtain at least two sub-processing results, and wherein, when obtaining any one of the sub-processing results, synchronously sending the any one of the sub-processing results to other computing units among the plurality of computing units that are parallelly processing the tasks to be processed.

[0024] Figure 1 Fig. shows an application scenario diagram of a model inference method, a computing device, a medium and a program product applied to a computing device according to an embodiment of the present invention.

[0025] As Figure 1 shown, the application scenario 100 according to this embodiment may include a control unit 101, a first computing unit 102_1, a second computing unit 102_2, …, an Nth computing unit 102_N. N is an integer greater than or equal to 2. There is network interaction between any one of the control unit 101 and the first computing unit 102_1, the second computing unit 102_2, …, the Nth computing unit 102_N, and between any two of the first computing unit 102_1, the second computing unit 102_2, …, the Nth computing unit 102_N. The network may be a medium providing a communication link. The network may include various connection types, such as wired, wireless communication links or fiber optic cables, etc.

[0026] The user may use a terminal device to interact with the control unit 101 to receive the target execution result or send tasks to be processed, etc. The first computing unit 102_1, the second computing unit 102_2, …, the Nth computing unit 102_N may be installed on the terminal device. The terminal device may include various communication client applications, such as a web browser application, a search application, an instant messaging tool, an email client, a social platform software, etc. (only as an example). The terminal device may be various electronic devices having a display screen and supporting web browsing, including but not limited to smart phones, tablet computers, laptop portable computers and desktop computers, etc.

[0027] The control unit 101 can be a central processing unit or other devices in a server. The control unit 101 may also include at least one of the following: a Baseboard Management Controller (BMC), a Complex Programmable Logic Device (CPLD), a single-chip microcomputer, a Field-Programmable Gate Array (FPGA), etc.

[0028] For example, the control unit 101 can separately send the to-be-processed tasks input by the user to the first computing unit 102_1, the second computing unit 102_2, …, the Nth computing unit 102_N. After the first computing unit 102_1, the second computing unit 102_2, …, the Nth computing unit 102_N respectively execute the to-be-processed tasks to obtain multiple sub-processing results, the multiple sub-processing results are processed to obtain the target execution result of the to-be-processed task. The target execution result is fed back to the terminal device.

[0029] The first computing unit 102_1, the second computing unit 102_2, …, the Nth computing unit 102_N can each separately include a graphics processing unit or a tensor processing unit, etc.

[0030] It should be understood that Figure 1 the numbers of the control unit and the computing units in

[0031] are merely illustrative. According to the implementation requirements, there can be any number of control units and computing units. Figure 1 Based on the scenario described below Figures 2 to 4 a model inference method applied to a computing device according to an embodiment of the invention will be described in detail.

[0032] Figure 2 shows a flowchart of a model inference method applied to a computing device according to an embodiment of the present invention.

[0033] As Figure 2 shown, the model inference method applied to a computing device in this embodiment includes operation S210 to operation S230, and this model inference method can be executed by a computing device. The computing device can include: a control unit and computing units for parallelly executing to-be-processed tasks.

[0034] In operation S210, the control unit separately sends the to-be-processed tasks to multiple computing units.

[0035] In operation S220, the multiple computing units respectively execute the to-be-processed tasks to obtain multiple sub-processing results.

[0036] In operation S230, the control unit obtains the target execution result of the task to be processed based on multiple sub - processing results.

[0037] In an embodiment of the present invention, the task to be processed can be determined according to user requirements. For example, the task to be processed may include, but is not limited to, image classification tasks, text generation tasks, question - answering tasks, query tasks, object detection, semantic segmentation, sentiment analysis, etc.

[0038] In an embodiment of the present invention, any one of the multiple computing units can be configured with at least two data subsets, and the data subsets are obtained by splitting the weight set of the target inference model required to execute the task to be processed.

[0039] For example, for a text generation task, the target inference model can be a large - language model. The weight set can be determined according to the weight file of the large - language model, and then the weight set is split to obtain multiple data subsets. The weight set of the large - language model can refer to the set of all parameters in the model, and these parameters can include weight matrices, bias terms, normalization parameters, etc. These parameters can be split to obtain multiple non - related data subsets, and the multiple data subsets are configured to the computing units for parallel execution of the task to be processed, and at least two data subsets are configured on each computing unit.

[0040] In an embodiment of the present invention, any one of the multiple computing units executes the task to be processed in the following manner: using the at least two configured data subsets, sequentially processes the task to be processed to obtain at least two sub - processing results. Among them, when obtaining any sub - processing result, synchronously sends the any sub - processing result to other computing units among the multiple computing units that are parallel - processing the task to be processed.

[0041] For example, sequentially processing the task to be processed can be to perform calculations on any one of the at least two data subsets and the task to be processed, such as matrix multiplication, activation functions, convolution operations, pooling operations, normalization operations, and other operations.

[0042] For example, the control unit can splice the multiple sub - processing results to obtain the target execution result of the task to be processed. In special cases, the control unit can also determine the main computing unit among the multiple computing units, and then the main computing unit returns the target execution result of the task to be processed obtained based on the multiple sub - processing results to the control unit.

[0043] According to an embodiment of the present invention, since at least two data subsets are configured in a computing unit, and the at least two configured data subsets are used to sequentially process a task to be processed, when any sub-processing result is obtained, the any sub-processing result is synchronously sent to other computing units among the multiple computing units that are processing the task to be processed in parallel. Therefore, the computing and communication of the computing unit reach parallelism. When traversing the at least two configured data subsets, only the sub-processing result obtained by processing the task to be processed for the last time needs to be sent. Compared with the serial computing and communication of the computing unit, since the sub-processing result obtained by processing the task to be processed for the last time is only a small part of the sub-processing results of the computing unit executing the task to be processed, when the task to be processed is completed, the communication time and communication data volume after the computing unit finishes computing can be reduced, thereby reducing the idle time of the computing resources of the computing unit, improving the computing efficiency and resource utilization rate of the computing unit, and further improving the task processing efficiency.

[0044] According to an embodiment of the present invention, the at least two data subsets may include N data subsets, where N is an integer greater than or equal to 2. Using the at least two configured data subsets to sequentially process the task to be processed, the operations for obtaining at least two sub-processing results may include: using the first data subset among the N data subsets to process the task to be processed to obtain the first sub-processing result. When n is greater than or equal to 2 and less than or equal to N, using the nth data subset among the N data subsets to process the task to be processed to obtain the nth sub-processing result, and at the same time sending the (n - 1)th sub-processing result to other computing units.

[0045] Figure 3A FIG. shows a schematic diagram of parallel execution of a task to be processed according to an embodiment of the present invention; Figure 3B FIG. shows a schematic diagram of any computing unit executing a task to be processed according to an embodiment of the present invention.

[0046] In an embodiment of the present invention, taking 3 computing units, each of which is configured with 3 data subsets as an example, it should be noted that the number of computing units and data subsets in this example is only exemplary, and according to the implementation requirements, any number of computing units and data subsets can be provided.

[0047] Such as Figure 3AAs shown in the figure, the control unit can send the task to be processed 301 to the first computing unit 102_1, the second computing unit 102_2, and the third computing unit 102_3 respectively. The first computing unit 102_1, the second computing unit 102_2, and the third computing unit 102_3 execute the task to be processed 301 in parallel to obtain multiple sub-processing results 303. For example, the first computing unit 102_1 uses the three configured data subsets to process the task to be processed 301 in sequence to obtain the first sub-processing result, the second sub-processing result, and the third sub-processing result. The second computing unit 102_2 uses the three configured data subsets to process the task to be processed 301 in sequence to obtain the fourth sub-processing result, the fifth sub-processing result, and the sixth sub-processing result. The third computing unit 102_3 uses the three configured data subsets to process the task to be processed 301 in sequence to obtain the seventh sub-processing result, the eighth sub-processing result, and the ninth sub-processing result. It should be noted that the three data subsets configured for the first computing unit 102_1, the second computing unit 102_2, and the third computing unit 102_3 are all different.

[0048] For any one of the first computing unit 102_1, the second computing unit 102_2, and the third computing unit 102_3, when obtaining any sub-processing result, synchronously send the any sub-processing result to other computing units that are processing the task to be processed in parallel.

[0049] For example, as Figure 3B shown, taking the first computing unit 102_1 as an example, when receiving the task to be processed 301, use the first data subset 302_1 configured to process the task to be processed 301 to obtain the first sub-processing result 303_1. Send the first sub-processing result 303_1 to the second computing unit 102_2 and the third computing unit 102_3 respectively. At the same time, use the second data subset 302_2 configured to process the task to be processed 301 to obtain the second sub-processing result 303_2. Send the second sub-processing result 303_2 to the second computing unit 102_2 and the third computing unit 102_3 respectively. At the same time, use the third data subset 302_3 configured to process the task to be processed 301 to obtain the third sub-processing result 303_3, and send the third sub-processing result 303_3 to the second computing unit 102_2 and the third computing unit 102_3 respectively.

[0050] According to an embodiment of the present invention, since the nth data subset is used to process the task to be processed to obtain the nth sub-processing result, and at the same time the (n - 1)th sub-processing result is sent to other computing units, when the Nth data subset is used to process the task to be processed to obtain the Nth sub-processing result, only the Nth sub-processing result needs to be sent to other computing units. The Nth sub-processing result is only 1 / N part of the target execution result. Therefore, after the computing unit finishes processing the task to be processed using the Nth data subset, the time for transmitting the processing result is shortened, thereby shortening the time for executing the task to be processed and improving the inference efficiency.

[0051] According to an embodiment of the present invention, any one of the multiple computing units is configured with the number of execution rounds of the task to be processed.

[0052] Figure 4 A schematic diagram showing the determination of the target execution result of the task to be processed according to an embodiment of the present invention is shown.

[0053] As Figure 4 shown, determining the target execution result of the task to be processed may include operations S401 to S406.

[0054] In operation S401, any one of the multiple computing units processes the task to be processed.

[0055] Exemplarily, when any one of the multiple computing units processes the task to be processed, the nth sub-processing result is obtained and sent to other computing units.

[0056] In operation S402, it is determined whether the Nth sub-processing result sent by other computing units has been received.

[0057] Exemplarily, in the case of determining that the Nth sub-processing result sent by other computing units has been received, multiple sub-processing results can be spliced to obtain the initial execution result, and operation S403 is executed. In the case of determining that the Nth sub-processing result sent by other computing units has not been received, S401 is continued to be executed.

[0058] In operation S403, it is determined whether the number of execution rounds M of the task to be processed is single.

[0059] Exemplarily, in the case of determining that the number of execution rounds of the task to be processed is single, operation S406 is executed, and the initial execution result can be used as the target execution result. In the case of determining that the number of execution rounds of the task to be processed is not single, operation S404 is executed.

[0060] In operation S404, based on the multiple sub-processing results from other computing units in the (m - 1)th execution of the task to be processed, at least two configured data subsets are used to process the mth task to be processed in sequence to obtain the mth execution result.

[0061] Exemplarily, m is an integer greater than 1, and M is an integer greater than 1.

[0062] Exemplarily, any computing unit can splice at least two sub - processing results obtained in the (m - 1) - th execution of the task to be processed and multiple sub - processing results from other computing units in the (m - 1) - th execution of the task to obtain the (m - 1) - th execution result obtained in the (m - 1) - th execution of the task. Take the (m - 1) - th execution result as the m - th task to be processed, and use at least two configured data subsets to process the m - th task to be processed in sequence. Since when executing the task to be processed for the m - th time, any computing unit can use the sub - processing results sent by other computing units to determine the (m - 1) - th execution result, when repeatedly executing the task to be processed, there is no need to obtain the task to be processed from the control unit again, saving the transmission resources and transmission time for the control unit to send data to the computing unit.

[0063] In operation S405, determine whether m = M.

[0064] Exemplarily, in the case of m = M, perform operation S406, and the m - th execution result can be used as the target execution result. In the case of m ≠ M, let m' = m + 1, and then repeat operation S404.

[0065] In operation S406, determine the target execution result.

[0066] According to the embodiments of the present invention, by determining the number of execution rounds of the task to be processed, different inference processes are corresponding to different processing tasks, which has universality. In addition, for the task to be processed with a large number of execution rounds, since the computing idle time of each round of computing units is reduced, it can quickly and accurately respond to user needs and improve the user experience.

[0067] According to another embodiment of the present invention, the model inference method applied to a computing device may further include operations in addition to the operations S210 - S230 as shown above Figure 2 The control unit determines the model structure of the target inference model according to the configuration file of the target inference model. Based on the model structure and the task to be processed, determine the number of rounds for parallel execution of the task to be processed.

[0068] Exemplarily, for different target inference models, the model structure and the weight set are different. When the target inference model is trained, the model structure can be stored in the configuration file, and the weight set can be stored in the weight file.

[0069] Exemplarily, due to different model structures and different tasks to be processed, the number of rounds for executing the task to be processed is different. For example, for tasks to be processed involving complex logical reasoning, multi-step decision-making, or generating long texts, and models with complex model structures, the task to be processed can be executed in multiple rounds. For simple logical tasks or models with simple structures, the task to be processed can be executed once. Models with complex model structures can include, for example, but are not limited to, recurrent neural networks, Transformer models, autoregressive models, etc. Models with simple model structures can include, for example, but are not limited to, convolutional neural networks, fully connected networks, etc.

[0070] For example, semantic parsing can be performed on the task to be processed to determine the type of the task to be processed, and based on the type of the task to be processed and the model structure, the number of rounds for parallelly executing the task to be processed can be determined.

[0071] For example, for a task to be processed that is a text processing task, based on the text sequence of the text processing task, it can be obtained that the type of the text sequence is a long text generation type. The model structure of the target inference model is a Transformer model. Then, the number of rounds for parallelly executing the task to be processed can be determined according to the length of the text sequence and the model padding. For example, when the model structure of the target inference model is a convolutional neural network, a fully connected network, etc., and the task to be processed is a simple logical task, such as a simple question and answer, the number of rounds for parallelly executing the task to be processed is determined to be single time.

[0072] According to an embodiment of the present invention, by the model structure and the task to be processed, reasonably configuring the number of rounds for parallelly executing the task to be processed can reduce the idle time of the computing unit, maximize the resource utilization efficiency, and improve the throughput and response speed of the processing task.

[0073] According to another embodiment of the present invention, the model inference method applied to a computing device may further include: for a high-concurrency scenario, if at the same moment, when multiple users send multiple tasks to be processed, by performing semantic analysis on the multiple tasks to be processed, determining the types and target inference models of the multiple tasks to be processed respectively, the multiple tasks to be processed can be sorted according to the types of the multiple tasks to be processed respectively and the model structure of the target inference model, and the tasks to be processed are parallelly executed in sequence according to the sorting.

[0074] For example, first, based on the types of the multiple tasks to be processed respectively, they can be sorted in descending order according to the complexity of logical reasoning, and then, based on the model structure of the target inference model, they can be sorted again in descending order according to the complexity of the model structure to obtain the finally sorted multiple tasks to be processed.

[0075] According to an embodiment of the present invention, since multiple tasks to be processed are sorted based on the model structure of the target inference model and the types of the multiple tasks to be processed respectively, simple tasks to be processed can be preferentially processed, reducing the response waiting time of simple tasks. In addition, combined with the parallel strategy of computing and communication when executing tasks to be processed in parallel, the response waiting time of tasks can be further reduced. The combination of the two can quickly respond to user requirements and improve the user experience.

[0076] According to an embodiment of the present invention, the time required for multiple computing units to process tasks to be processed in the same round is the same.

[0077] In an embodiment of the present invention, the time required to process a task to be processed may be the total time for any computing unit to sequentially process the task to be processed by using at least two configured data subsets, obtain at least two sub-processing results, and send both sub-processing results to other computing units.

[0078] According to an embodiment of the present invention, the time required for multiple computing units to process tasks to be processed in the same round is the same, which can balance the loads of multiple computing units, significantly reduce the system latency, and thus improve the user experience.

[0079] According to another embodiment of the present invention, in addition to the operations S210~S230 shown above, the model inference method applied to a computing device may further include an operation: the control unit splits the weight set of the target inference model based on the number of units of multiple computing units to execute tasks to be processed, the respective performance parameters of the multiple computing units, and the attributes of the weight set, to obtain multiple data subsets. Figure 2 Specifically, the performance parameters may include computing power, memory capacity, etc. The attributes of the weight set may include the dimension of the weight matrix, the dimension of the bias term, the dimension of the normalization parameter, etc.

[0080] Exemplarily, the performance parameters may include computing power, memory capacity, etc. The attributes of the weight set may include the dimension of the weight matrix, the dimension of the bias term, the dimension of the normalization parameter, etc.

[0081] Exemplarily, the weight set of the target inference model may be split according to a splitting strategy for load balancing among multiple computing units when distributing multiple data subsets to multiple computing units.

[0082] According to an embodiment of the present invention, splitting the weight set of the target inference model based on the number of units of multiple computing units to execute tasks to be processed, the respective performance parameters of the multiple computing units, and the attributes of the weight set is beneficial to load balancing among multiple computing units when the splitting granularity of the weight set is small.

[0083] According to an embodiment of the present invention, the weight set may include a weight matrix, and the attributes of the weight set may include the dimensions of the weight matrix. The control unit splits the weight set based on the number of units of multiple computing units, the respective performance parameters of the multiple computing units, and the attributes of the weight set, and obtains a plurality of data subsets, which may include operations: initially splitting the weight matrix based on the number of units of multiple computing units, the respective performance parameters of the multiple computing units, and the dimensions of the weight matrix to obtain the number of units of initial sub-matrices; and then splitting the number of units of initial sub-matrices based on the task types executed by the computing units to obtain a plurality of data subsets.

[0084] In an embodiment of the present invention, when splitting the initial sub-matrix again, those with the same task types executed by the computing units may be split into a group.

[0085] For example, if the task type executed by the computing unit is a long text generation task, the matrix rows in the initial sub-matrix regarding the same attention head may be in a group during splitting. In this way, each attention head will focus on an independent subspace of the task to be processed, without the need to increase communication.

[0086] According to an embodiment of the present invention, by initially splitting the weight matrix into initial sub-matrices and then splitting again based on the task types executed by the computing units, it can not only ensure that each data subset is used to independently process the task to be processed, but also achieve a finer-grained division of the weight matrix, which is beneficial to implementing the strategy of parallel computing and communication.

[0087] According to an embodiment of the present invention, the performance parameter may include computing power. Based on the number of units of multiple computing units, their respective performance parameters, and the dimensions of the weight matrix, initially splitting the weight matrix may include operations: when the dimensions of the weight matrix are divisible by the number of units and the computing powers of the multiple computing units are the same, splitting the weight matrix proportionally into the number of units of initial sub-matrices. When the dimensions of the weight matrix are not divisible by the number of units, splitting the weight matrix based on the computational complexity of the tasks executed by the computing units so that the computational complexities of the tasks executed by the multiple computing units are consistent.

[0088] In an embodiment of the present invention, the computational complexities of the tasks executed by the multiple computing units being consistent may mean that the computational complexities of the tasks executed by the multiple computing units are exactly the same or the difference in the computational complexities of the tasks executed by any two of the multiple computing units satisfies a predetermined error.

[0089] Exemplarily, when the dimensions of the weight matrix are divisible by the number of units and the computing powers of the multiple computing units are the same, the weight matrix may be proportionally split by rows into the number of units of initial sub-matrices.

[0090] For example, there is a weight matrix A with dimensions (p, k). The number of cells in the computing unit is 2. The weight matrix A can be initially split into 2 initial sub-matrices, such as A1 and A2. The dimensions of each initial sub-matrix are (p / 2, k). The initial sub-matrices, such as A1 and A2, are further split respectively. A1 and A2 are each split into 4 data subsets, such as A11, A12, A13, A14, B11, B12, B13, and B14. The dimensions of A11, A12, A13, A14, B11, B12, B13, and B14 are all (p / 8, k). A11, A12, A13, and A14 can be configured on computing unit 1. B11, B12, B13, and B14 can be configured on computing unit 2.

[0091] According to an embodiment of the present invention, since the weight matrix is split based on the computing power, the number of units, the dimensions of the weight matrix, and the computational complexity of the tasks executed by the computing unit, the waiting time between devices can be reduced.

[0092] According to another embodiment of the present invention, each of the multiple computing units can be configured with a buffer pool. The buffer pool is configured by the control unit based on the task to be processed, the target inference model, and the memory resources of the computing unit. The buffer pool is used to cache the sub-processing results from other computing units.

[0093] Exemplarily, the estimated memory occupancy of the target execution result generated after executing the task to be processed can be predicted according to the task to be processed and the target inference model. Then, the size of the buffer pool is configured according to the estimated memory occupancy and the memory resources of the computing unit. For example, historical processing tasks matching the task type of the task to be processed and the model structure of the target inference model can be determined, and then the estimated memory occupancy of the target execution result generated after executing the task to be processed can be determined according to the memory occupancy of the historical target execution results obtained from the historical processing tasks.

[0094] Exemplarily, when caching the sub-processing results from other computing units in the buffer pool, for the case where the number of execution rounds of the task to be processed is M times, when the control unit determines that multiple computing units have sent the Nth sub-processing result, the control unit can control the computing device to perform a reduction operation, such as matrix addition or matrix multiplication, etc.

[0095] According to an embodiment of the present invention, since a buffer pool is built in the computing unit, when executing the next task to be processed, there is no need to obtain the task to be executed from the control unit again, which can reduce the communication time and save communication resources.

[0096] According to another embodiment of the present invention, the model inference method applied to the computing device further includes the above-mentioned such as Figure 2In addition to the operations S210 to S230 shown, the operations may further include: updating the execution result cached in the buffer pool using the m-th execution result.

[0097] Exemplarily, the m-th execution result may be used to overwrite the (m - 1)-th execution result cached in the buffer pool.

[0098] According to an embodiment of the present invention, by updating the execution result in the cache, the situation of buffer pool memory overflow caused by the large memory occupation of the execution result is avoided. In addition, the outdated execution result may lead to invalid calculations. Updating the execution result in the cache in real time can save unnecessary computing resources, improve computing efficiency, and save storage resources.

[0099] According to another embodiment of the present invention, the model inference method applied to a computing device, in addition to including the operations S210 to S230 shown above Figure 2 may further include the operation of: deleting the target execution result cached in the buffer pool when it is determined that the target execution result is obtained and the next task to be processed is to be executed.

[0100] According to an embodiment of the present invention, since the memory in the buffer pool is limited, when the next task to be processed is to be executed, timely deleting the target execution result cached in the buffer pool can prevent buffer pool memory overflow.

[0101] According to another embodiment of the present invention, the model inference method applied to a computing device, in addition to including the operations S210 to S230 shown above Figure 2 may further include the operation of: when the next task to be processed is to be executed, it may be determined whether the next task to be processed and the current task to be processed are the same task. When the next task to be processed and the current task to be processed are the same task, the target execution result stored in the buffer pool is used as the target execution result of the next task to be processed.

[0102] Exemplarily, the results may be stored in the buffer pool in the form of key-value pairs, where the key is the task to be processed and the value is the target execution result. When the next task to be processed is received, the target execution result of the task to be processed that matches the next task to be processed can be determined from the buffer pool and used as the target execution result for executing the next task to be processed.

[0103] According to an embodiment of the present invention, since the execution result of the task to be processed is stored in the buffer pool, the reuse rate of the execution result can be improved.

[0104] Based on the above model inference method applied to a computing device, the present invention further provides a computing device. The following will be combined with Figure 5 to describe this computing device in detail.

[0105] Figure 5 The structural block diagram of a computing device according to an embodiment of the present invention is shown.

[0106] As Figure 5 shown, the computing device 500 of this embodiment includes a control unit 101, a first computing unit 102_1, a second computing unit 102_2, ……, and an Nth computing unit 102_N.

[0107] The control unit 101 is used to send the tasks to be processed to multiple computing units respectively.

[0108] Each of the first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N is configured with at least two data subsets, and the data subsets are obtained by splitting the weight set of the target inference model required for executing the task to be processed.

[0109] The first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N are used to execute the tasks to be processed respectively to obtain multiple sub-processing results.

[0110] The control unit 101 is further used to obtain the target execution result of the task to be processed according to the multiple sub-processing results.

[0111] Any one of the first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N is used to execute the task to be processed in the following manner: using the at least two configured data subsets to process the task to be processed in sequence to obtain at least two sub-processing results, wherein, when any sub-processing result is obtained, the any sub-processing result is synchronously sent to other computing units among the multiple computing units that are processing the task to be processed in parallel.

[0112] According to an embodiment of the present invention, the at least two data subsets may include N data subsets, where N is an integer greater than or equal to 2; any one of the first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N using the at least two configured data subsets to process the task to be processed in sequence to obtain at least two sub-processing results may include: using the first data subset among the N data subsets to process the task to be processed to obtain the first sub-processing result; when n is greater than or equal to 2 and less than or equal to N, using the nth data subset among the N data subsets to process the task to be processed to obtain the nth sub-processing result, and at the same time sending the (n - 1)th sub-processing result to other computing units.

[0113] According to an embodiment of the present invention, any one of the first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N is configured with the number of execution rounds of the task to be processed.

[0114] According to an embodiment of the present invention, the control unit 101 is further configured to obtain an initial execution result when it is determined that multiple computing units have sent the Nth sub-processing result; when the number of execution rounds of the task to be processed is single, use the initial execution result as the target execution result; when the number of execution rounds of the task to be processed is M times, and when m is greater than 1 and less than M, based on multiple sub-processing results from other computing units in the (m - 1)th execution of the task, use at least two configured data subsets to process the mth task to be processed in sequence to obtain the mth execution result, where M is an integer greater than 1; when m = M, use the mth execution result as the target execution result.

[0115] According to an embodiment of the present invention, the control unit 101 is further configured to determine the model structure of the target inference model according to the configuration file of the target inference model; based on the model structure and the task to be processed, determine the number of rounds for parallel execution of the task to be processed.

[0116] According to an embodiment of the present invention, the time required for multiple computing units to process the task to be processed in the same round is the same.

[0117] According to an embodiment of the present invention, the control unit 101 is further configured to split the weight set of the target inference model based on the number of units of multiple computing units to execute the task to be processed, the respective performance parameters of the multiple computing units, and the attributes of the weight set, to obtain multiple data subsets.

[0118] According to an embodiment of the present invention, the weight set includes a weight matrix, and the attributes of the weight set include the dimension of the weight matrix; the control unit 101 splits the weight set based on the number of units of multiple computing units, the respective performance parameters of the multiple computing units, and the attributes of the weight set, to obtain multiple data subsets, including: initially splitting the weight matrix based on the number of units of multiple computing units, the respective performance parameters of the multiple computing units, and the dimension of the weight matrix, to obtain the number of initial sub-matrices equal to the number of units; and then splitting the number of initial sub-matrices based on the task types executed by the computing units, to obtain multiple data subsets.

[0119] According to an embodiment of the present invention, the performance parameter includes computing power. The control unit 101 initially splits the weight matrix based on the number of units of multiple computing units, the respective performance parameters of the multiple computing units, and the dimension of the weight matrix, including: when the dimension of the weight matrix is divisible by the number of units and the computing power of each of the multiple computing units is the same, equally splitting the weight matrix into the number of initial sub-matrices equal to the number of units; when the dimension of the weight matrix is not divisible by the number of units, splitting the weight matrix based on the computational complexity of the computing unit for executing the task, so that the computational complexity of each of the multiple computing units for executing the task is consistent.

[0120] According to an embodiment of the present invention, any one of the first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N is configured with a buffer pool, and the buffer pool is configured by the control unit based on the task to be processed, the target inference model, and the memory resources of the computing unit, and the buffer pool is used to cache the sub-processing results from other computing units.

[0121] According to an embodiment of the present invention, the control unit 101 is further configured to update the cached (m - 1)th execution result in the buffer pool by using the mth execution result.

[0122] According to an embodiment of the present invention, the control unit 101 is further configured to delete the cached target execution result in the buffer pool when it is determined to obtain the target execution result and the next task to be processed is to be executed.

[0123] According to an embodiment of the present invention, any plurality of units among the control unit 101, the first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N may be combined and implemented in one unit, or any one of them may be split into multiple units. Alternatively, at least part of the functions of one or more of these units may be combined with at least part of the functions of other units and implemented in one unit. According to an embodiment of the present invention, at least one of the control unit 101, the first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N may be at least partially implemented as a hardware circuit, such as a field programmable gate array (FPGA), a programmable logic array (PLA), a system on chip, a system on substrate, a system on package, an application specific integrated circuit (ASIC), or may be implemented by any other reasonable means such as hardware or firmware through integration or packaging of circuits, or may be implemented in any one of the three implementation manners of software, hardware, and firmware or in any suitable combination of several of them. Alternatively, at least one of the control unit 101, the first computing unit 102_1, the second computing unit 102_2, ……, and the Nth computing unit 102_N may be at least partially implemented as a computer program module, and when the computer program module is run, the corresponding functions may be executed.

[0124] The present invention also provides a computer-readable storage medium, which may be included in the device / apparatus / system described in the above embodiments; or may exist alone without being assembled into the device / apparatus / system. The above computer-readable storage medium carries one or more programs, and when the one or more programs are executed, the method according to the embodiments of the present invention is implemented.

[0125] According to an embodiment of the present invention, the computer-readable storage medium may be a non-volatile computer-readable storage medium, for example, may include but is not limited to: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program may be used by or combined with an instruction execution system, apparatus, or device. For example, according to an embodiment of the present invention, the computer-readable storage medium may include one or more memories.

[0126] Embodiments of the present invention also include a computer program product, which includes a computer program that contains program code for executing the method shown in the flowchart. When the computer program product runs in a computer system, the program code is used to cause the computer system to implement the method provided by the embodiments of the present invention.

[0127] When the computer program is executed by a processor, it executes the above functions defined in the system / apparatus of the embodiments of the present invention. According to the embodiments of the present invention, the above-described systems, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0128] In one embodiment, the computer program can rely on tangible storage media such as optical storage devices and magnetic storage devices. In another embodiment, the computer program can also be transmitted and distributed in the form of a signal on a network medium, and be downloaded and installed through the communication part, and / or be installed from a removable medium. The program code included in the computer program can be transmitted using any suitable network medium, including but not limited to: wireless, wired, etc., or any suitable combination of the above.

[0129] In such an embodiment, the computer program can be downloaded and installed from the network through the communication part, and / or be installed from a removable medium. When the computer program is executed by a processor, it executes the above functions defined in the system of the embodiments of the present invention. According to the embodiments of the present invention, the above-described systems, devices, apparatuses, modules, units, etc. can be implemented by computer program modules.

[0130] According to the embodiments of the present invention, the program code for executing the computer program provided by the embodiments of the present invention can be written in any combination of one or more programming languages. Specifically, these computing programs can be implemented using high-level procedures and / or object-oriented programming languages, and / or assembly / machine languages. Programming languages include but are not limited to, such as Java, C++, python, the "C" language, or similar programming languages. The program code can be executed entirely on the user's computing device, partially on the user's device, partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device can be connected to the user's computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computing device (for example, by using an Internet service provider to connect through the Internet).

[0131] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, or they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram or flowchart, as well as combinations of blocks in the block diagram or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or by a combination of dedicated hardware and computer instructions.

[0132] Those skilled in the art will appreciate that the features described in the various embodiments of the present invention can be combined and / or combined in various ways, even if such combinations or combinations are not explicitly described in the present invention. In particular, without departing from the spirit and teachings of the present invention, the features described in the various embodiments of the present invention can be combined and / or combined in various ways. All such combinations and / or combinations fall within the scope of the present invention.

[0133] The embodiments of the present invention have been described above. However, these embodiments are for illustrative purposes only and are not intended to limit the scope of the present invention. Although the embodiments have been described separately above, this does not mean that the measures in the various embodiments cannot be used advantageously in combination. Without departing from the scope of the present invention, those skilled in the art can make various substitutions and modifications, and all such substitutions and modifications should fall within the scope of the present invention.

Claims

1. A model reasoning method applied to a computing device, characterized in that: The computing device comprises: a control unit and a plurality of computing units for executing tasks to be processed in parallel, and the method comprises: The control unit sends the tasks to be processed to the multiple computing units respectively, and any computing unit among the multiple computing units is configured with at least two data subsets, the at least two data subsets are not related to each other, and the data subsets are obtained by splitting the weight set of the target inference model required to execute the tasks to be processed; The multiple computing units respectively execute the tasks to be processed to obtain multiple sub-processing results; The control unit obtains the target execution result of the task to be processed according to the multiple sub-processing results; wherein any one of the multiple computing units executes the task to be processed in the following manner: Using the at least two configured data subsets, the tasks to be processed are processed in sequence to obtain at least two sub-processing results, wherein when any sub-processing result is obtained, the any sub-processing result is synchronously sent to other computing units in the multiple computing units that process the tasks to be processed in parallel.

2. The method according to claim 1, characterized in that: The at least two data subsets include N data subsets, where N is an integer greater than or equal to 2; the at least two data subsets configured are used to sequentially process the tasks to be processed, and at least two sub-processing results are obtained, including: Using a first data subset of the N data subsets to process the task to be processed, to obtain a first sub-processing result; When n is greater than or equal to 2 and less than or equal to N, the task to be processed is processed using the nth data subset of the N data subsets to obtain the nth sub-processing result, and the n-1th sub-processing result is sent to the other computing units.

3. The method according to claim 2, characterized in that Obtaining the number of execution rounds of the task to be processed configured by any computing unit among the multiple computing units; The method further comprises: In the case of determining that the Nth sub-processing result sent by the other computing unit has been received, obtaining an initial execution result; In the case where the number of execution rounds of the task to be processed is single, taking the initial execution result as the target execution result; When the number of execution rounds of the task to be processed is M times, In the case where m is greater than 1 and less than M, based on the multiple sub-processing results from the other computing units in the m-1th execution of the task to be processed, the at least two configured data subsets are used to sequentially process the mth task to be processed to obtain the mth execution result, the mth task to be processed is the m-1th execution result, and M is an integer greater than 1; When m=M, the mth execution result is used as the target execution result.

4. The method according to claim 3, characterized in that The method further comprises: The control unit determines the model structure of the target reasoning model according to the configuration file of the target reasoning model; Based on the model structure and the tasks to be processed, the number of rounds for executing the tasks to be processed in parallel is determined.

5. The method according to claim 3, characterized in that: The time required for the multiple computing units to process the tasks to be processed in the same round is the same.

6. The method according to claim 1, characterized in that The method further comprises: The control unit splits the weight set of the target inference model based on the number of the multiple computing units to execute the tasks to be processed, the performance parameters of the multiple computing units respectively, and the attributes of the weight set to obtain the multiple data subsets.

7. The method according to claim 6, characterized in that The weight set includes a weight matrix, and the attributes of the weight set include the dimension of the weight matrix; The control unit splits the weight set of the target reasoning model based on the number of the plurality of computing units to execute the task to be processed, the performance parameters of the plurality of computing units, and the attributes of the weight set to obtain the plurality of data subsets, including: Based on the number of the plurality of computing units, the performance parameters of the plurality of computing units, and the dimension of the weight matrix, initially splitting the weight matrix to obtain initial sub-matrices of the number of units; Based on the task type executed by the computing unit, the initial sub-matrices of the number of units are split again to obtain the multiple data subsets.

8. The method according to claim 7, characterized in that The performance parameters include computing power; The initial splitting of the weight matrix based on the number of the plurality of computing units, the performance parameters of the plurality of computing units, and the dimension of the weight matrix includes: When the dimension of the weight matrix is ​​divisible by the number of units and the computing capabilities of the plurality of computing units are the same, the weight matrix is ​​divided into initial sub-matrices of the number of units in equal proportion; When the dimension of the weight matrix cannot be divided by the number of units, the weight matrix is ​​split based on the computational complexity of the computing unit executing the task to be processed, so that the computational complexity of each of the multiple computing units executing the task to be processed is consistent.

9. The method according to claim 3, characterized in that: Each of the multiple computing units is configured with a buffer pool, which is configured by the control unit based on the tasks to be processed, the target reasoning model and the memory resources of the computing unit, and is used to cache the sub-processing results from the other computing units.

10. The method according to claim 9, characterized in that The method further includes: using the mth execution result to update the (m-1)th execution result cached in the buffer pool.

11. The method according to claim 9, characterized in that The method further comprises: When it is determined that the target execution result is obtained and the next task to be processed is to be executed, the target execution result cached in the buffer pool is deleted.

12. A computing device comprising: A control unit, used for sending the tasks to be processed to the multiple computing units respectively; The multiple computing units, any computing unit among the multiple computing units is configured with at least two data subsets, the at least two data subsets are not related to each other, and the data subsets are obtained by splitting the weight set of the target reasoning model required to perform the task to be processed; The multiple computing units are used to respectively execute the tasks to be processed to obtain multiple sub-processing results; The control unit is further used to obtain the target execution result of the task to be processed according to the multiple sub-processing results; Wherein, any computing unit among the multiple computing units is used to execute the task to be processed in the following manner: Using the at least two configured data subsets, the tasks to be processed are processed in sequence to obtain at least two sub-processing results, wherein when any sub-processing result is obtained, the any sub-processing result is synchronously sent to other computing units in the multiple computing units that process the tasks to be processed in parallel.

13. The computing device according to claim 12, characterized in that The at least two data subsets include N data subsets, where N is an integer greater than or equal to 2; the computing unit uses the configured at least two data subsets to sequentially process the tasks to be processed, and obtains at least two sub-processing results including: Using a first data subset of the N data subsets to process the task to be processed, to obtain a first sub-processing result; When n is greater than or equal to 2 and less than or equal to N, the task to be processed is processed using the nth data subset of the N data subsets to obtain the nth sub-processing result, and the n-1th sub-processing result is sent to the other computing units.

14. A computer-readable storage medium having a computer program or instruction stored thereon, characterized in that: When the computer program or instruction is executed by a processor, the steps of the method according to any one of claims 1 to 11 are implemented.

15. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 11.

Citation Information

Patent Citations

  • Data processing method, data processor, electronic equipment and storage medium

    CN118313458A