Task execution method and apparatus used for large scale model, electronic device, storage medium, and program

By dividing large-scale model weights into basic and collaborative components and applying a matrix multiplication mechanism, the method addresses high computational costs and resource demands, enhancing efficiency and adaptability for large-scale model deployment.

JP2025098208AActive Publication Date: 2025-07-01BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2025054153
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-19
Filing Date
2025-03-27
Publication Date
2025-07-01
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The high training and inference costs, along with the difficulty in deploying large-scale models due to their large model parameters and resource consumption, pose challenges in efficiently executing computational tasks.

Method used

The method divides the weight parameters of a large-scale model into basic and collaborative weights, using a matrix multiplication mechanism to determine sub-weights for collaborative computing tasks, reducing computational overhead and improving flexibility and adaptability by integrating these weights with a general basic model.

Benefits of technology

This approach reduces computational energy consumption and enhances computing efficiency, enabling deployment of large-scale models on devices with lower computing performance while maintaining inference capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098208000001_ABST
    Figure 2025098208000001_ABST
Patent Text Reader

Abstract

To provide a task execution method and apparatus for a large scale model, an electronic device, a storage medium, and a program.SOLUTION: A method executes, according to a target feature to be processed, a collaborative computing task using a target computing unit, to obtain a target collaborative feature. The collaborative computing task includes a first collaborative task for processing the target feature to be processed and a first collaborative sub-weight to obtain an intermediate collaborative feature, and a second collaborative task for processing the intermediate collaborative feature and a second collaborative sub-weight to obtain a target collaborative feature. The method further processes collaborative weights according to a common matrix multiplication mechanism to determine the first collaborative sub-weight and the second collaborative-sub weight, and fuses a target basic feature obtained by executing a basic computing task using the target computing unit and the target collaborative feature to obtain a next target feature to be processed.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of deep learning technology and large model technology.

Background Art

[0002] Due to the rapid development of artificial intelligence technology, in scenarios such as intelligent customer service and knowledge Q&A, large models can be used to process data in various scenarios.

Summary of the Invention

[0003] The present disclosure provides a task execution method, apparatus, task execution device, electronic device, storage medium and program for use in a large model.

[0004] According to one aspect of the present disclosure, according to the features to be processed, a target computing unit is used to execute a cooperative computing task to obtain a target cooperative feature, where the cooperative computing task includes a first cooperative task for processing the features to be processed and a first cooperative sub-weight to obtain an intermediate cooperative feature, and a second cooperative task for processing the intermediate cooperative feature and a second cooperative sub-weight to obtain the target cooperative feature, and the first cooperative sub-weight and the second cooperative sub-weight are determined by processing the cooperative weights according to the matrix multiplication mechanism of a general matrix, and the target basic feature and the target cooperative feature are fused to obtain the next feature to be processed, where the target basic feature is obtained by using the target computing unit to execute a basic computing task, and the basic computing task is used to process the basic weight and the features to be processed. A task execution method for use in a large model is provided.

[0005] According to another aspect of the present disclosure, a target storage unit for storing collaborative computing tasks, and according to the features to be targeted for processing, executes the collaborative computing tasks, obtains target collaborative features, where the collaborative computing tasks include a first collaborative task for processing the features to be targeted for processing and a first collaborative sub-weight to obtain intermediate collaborative features, and a second collaborative task for processing the intermediate collaborative features and a second collaborative sub-weight to obtain the target collaborative features, the first collaborative sub-weight and the second collaborative sub-weight are determined by processing collaborative weights according to the matrix multiplication mechanism of a general matrix, fuses the target basic features and the target collaborative features to obtain the next features to be targeted for processing, where the target basic features are obtained by executing basic computing tasks in the target computing unit, and the basic computing tasks are used to process the basic weights and the features to be targeted for processing, a target computing unit arranged as such, a task execution device used in a large-scale model including the above is provided.

[0006] According to another aspect of the present disclosure, a task execution device used in a large-scale model including the task execution device used in the large-scale model provided by the embodiments of the present disclosure is provided.

[0007] According to another aspect of the present disclosure, an electronic device including at least one processor and a memory communicably connected to the at least one processor, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor such that the at least one processor executes the method provided by the embodiments of the present disclosure.

[0008] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method provided by the embodiments of the present disclosure is provided.

[0009] According to another aspect of the present disclosure, a computer program for realizing the method provided by the embodiments of the present disclosure when executed by a processor is provided.

[0010] It should be understood that the content described in this part is not intended to indicate the key points or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be easily understood from the following description.

Brief Description of the Drawings

[0011] The drawings are for a better understanding of the present technical solution and do not limit the present application.

[0012]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

DETAILED DESCRIPTION OF THE INVENTION

[0013] Exemplary embodiments of the present disclosure will be described below with reference to the drawings. Here, various details of the embodiments of the present disclosure are included for easier understanding, and they should be considered exemplary. Therefore, those skilled in the art should understand that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and brevity, descriptions of well-known functions and configurations are omitted in the following description.

[0014] In the technical solution of the present disclosure, the acquisition, storage, and application of such personal information of the user all comply with the provisions of relevant laws and regulations, adopt necessary confidentiality measures, and do not violate public order and good customs.

[0015] In the field of deep learning, the application of large-scale models has been continuously expanding. However, the training and inference costs of large-scale models are high, and it is difficult to deploy. For example, large-scale models may include large language models (LLMs), large-scale image models, large-scale audio models, etc. Large language models exhibit powerful text processing capabilities. However, since the scale of the model parameters of large-scale models is large and can reach hundreds of millions or billions, usually, large-scale models need to consume a large amount of computing resources in the process of executing computational tasks.

[0016] Embodiments of the present disclosure provide a task execution method, apparatus, task execution device, electronic device, storage medium, and program for use in large-scale models. The task execution method for use in large-scale models executes a collaborative computing task using a target computing unit according to characteristics to be processed, to obtain target collaborative characteristics, where the collaborative computing task includes a first collaborative task for processing the characteristics to be processed and a first collaborative sub-weight to obtain intermediate collaborative characteristics, and a second collaborative task for processing the intermediate collaborative characteristics and a second collaborative sub-weight to obtain target collaborative characteristics, the first collaborative sub-weight and the second collaborative sub-weight being determined by processing collaborative weights according to a matrix multiplication mechanism of a general matrix, and fusing the target basic characteristics and the target collaborative characteristics to obtain the next characteristics to be processed, where the target basic characteristics are obtained by executing a basic computing task using the target computing unit, and the basic computing task is used to process a basic weight and the characteristics to be processed.

[0017] According to embodiments of the present disclosure, by dividing the weight parameters of a large-scale model into a basic weight and a collaborative weight, the large-scale model uses the basic weight of a general basic model and the collaborative weight related to a specified personalization task to execute the computing tasks that need to be executed, and improves the flexibility and adaptability of the large-scale model in the inference process for executing personalization needs during the training or application process using the general basic model. At the same time, according to the matrix multiplication mechanism based on a general matrix, the collaborative weight is determined as a first collaborative sub-weight and a second collaborative sub-weight, whereby both the matrix dimension of the first collaborative sub-weight and the matrix dimension of the second collaborative sub-weight are smaller than the matrix dimension of the collaborative weight, and the total amount of computation of the first collaborative task and the second collaborative task is made smaller than the amount of computation when directly using the collaborative weight to execute the collaborative computing task, reducing the computational overhead of the target computing unit during the process of the large-scale model executing the computing task, reducing the computational energy consumption, and improving the computing efficiency of the large-scale model.

[0018] FIG. 1 schematically shows an example of a system architecture to which the task execution method and apparatus according to embodiments of the present disclosure can be applied.

[0019] It should be noted that FIG. 1 is only an example of a system architecture to which embodiments of the present disclosure can be applied to help those skilled in the art understand the technical content of the present disclosure, and it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments, or scenarios.

[0020] As shown in FIG. 1, the system architecture according to this embodiment can include a terminal device 101, a network 102, and a server cluster 103. The network 102 is used as a medium to provide a communication link between the terminal device 101 and the server cluster 103. The network 102 can also be used as a medium to provide a communication link within the server cluster 103. The network 102 can include various connection types such as wired and / or wireless communication links.

[0021] The user can interact with the server cluster 103 via the network 102 using the terminal device 101 to send and receive messages, etc. For example, the terminal device 101 may send a request to train a deep learning model to the server cluster 103 via the network 102.

[0022] Various communication client applications such as (mere examples) a knowledge reading application, a web browser application, a search application, an instant messaging tool, an email client, and / or social platform software can be installed on the terminal device 101.

[0023] The terminal device 101 may be various electronic devices having a display screen and supporting web browsing, including but not limited to smartphones, tablet computers, laptop computers, desktop computers, etc.

[0024] The server cluster 103 may be a server that provides various services, such as a background management server that supports requests sent by users using the terminal device 101 (only one example).

[0025] The server cluster 103 is a cloud server also known as a cloud computing server or cloud host, which is the host product of a cloud computing service system and solves the deficiencies of traditional physical hosts and VPS services ("Virtual Private Server" or abbreviated as "VPS") that are difficult to manage and have weak business scalability. The server can also be a distributed system server or a server combined with a blockchain.

[0026] The server cluster 103 includes a plurality of server nodes 1031, 1032, 1033, 1034, and each server node includes one or more hardware devices. Use the server cluster 103 or server nodes to execute the task execution method used in the large-scale model provided by this disclosure, thereby enabling the deployment, inference, or training of the large-scale model with fewer computing resources and storage resources.

[0027] It should be understood that the system architecture of this disclosure has been described above, and the method of this disclosure will be described below.

[0028] It should be understood that the numbers of the terminal devices, networks, and servers in FIG. 1 are merely illustrative. Depending on the needs of implementation, any number of terminal devices, networks, and servers can be provided.

[0029] FIG. 2 schematically shows a flowchart of a task execution method used in a large model according to an embodiment of the present disclosure.

[0030] As shown in FIG. 2, the task execution method used in the large model includes operations S210 to S220.

[0031] In operation S210, according to the features to be processed, a cooperative computing task is executed using a target computing unit to obtain target cooperative features.

[0032] In operation S220, the target basic features and the target cooperative features are fused to obtain the next features to be processed.

[0033] According to an embodiment of the present disclosure, the target computing unit may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and an artificial intelligence computing unit. The artificial intelligence computing unit may include at least one of a neural network processing unit (NPU), a tensor processing unit (TPU), and a Kunlun core.

[0034] According to an embodiment of the present disclosure, the computing task may include a neuron computing task executed by a large model. Alternatively, for example, a computing task executed by a processing layer of a large model, such as an attention task that the large model needs to execute, may also be included.

[0035] According to an embodiment of the present disclosure, the computing tasks of the large model may include cooperative computing tasks and basic computing tasks. The basic computing task can be understood as a computing task executed based on the basic weights of a basic model (also called a base model) within the large model. The cooperative computing task can be understood as a computing task executed based on the cooperative weights of a cooperative model in the large model.

[0036] According to an embodiment of the present disclosure, the collaborative weights can be model parameters obtained by fine-tuning a large-scale model including a basic model, and a computing task can be executed based on the basic weights and the fine-tuned collaborative weights, thereby realizing inference tasks such as prediction of the large-scale model and image generation.

[0037] According to an embodiment of the present disclosure, the collaborative weights may be model parameters to be fine-tuned in a large-scale model constructed based on a basic model. In the process of fine-tuning the large-scale model, fine-tuning the large-scale model is realized by adjusting only the collaborative weights, and inference tasks such as text prediction and image generation of the large-scale model are realized based on the fine-tuned collaborative weights and the basic weights.

[0038] According to an embodiment of the present disclosure, the collaborative computing task includes a first collaborative task for processing a feature to be processed and a first collaborative sub-weight to obtain an intermediate collaborative feature, and a second collaborative task for processing the intermediate collaborative feature and a second collaborative sub-weight to obtain a target collaborative feature.

[0039] According to an embodiment of the present disclosure, the first collaborative sub-weight and the second collaborative sub-weight are determined by processing the collaborative weights according to the matrix multiplication mechanism of a general matrix.

[0040] According to an embodiment of the present disclosure, the first collaborative sub - weight and the second collaborative sub - weight are determined according to the matrix multiplication mechanism of a general matrix. Based on the GEMM (GEneral Matrix to Matrix Multiplication) mechanism, the collaborative weight W_LoRA is divided into a first collaborative sub - weight W_LoRA_A and a second collaborative sub - weight W_LoRA_B. The matrix dimensions of the first collaborative sub - weight W_LoRA_A and the matrix dimensions of the second collaborative sub - weight W_LoRA_B are both smaller than the matrix dimensions of the collaborative weight W_LoRA. Therefore, the calculation task of multiplying the collaborative weight W_LoRA by the features to be processed is determined as a first collaborative task of multiplying the first collaborative sub - weight W_LoRA_B by the features to be processed, and a second collaborative task of multiplying the intermediate collaborative features by the second collaborative sub - weight W_LoRA_B.

[0041] Since the collaborative calculation task can be understood as a LoRA (LowRank Adaptation) calculation task, the collaborative calculation task can be used to assist the basic calculation task to speed up the calculation efficiency of the large - scale model in executing the calculation task, and save the calculation overhead of the target calculation unit.

[0042] According to an embodiment of the present disclosure, the target basic features can be obtained by executing a basic calculation task using a target calculation unit, and the basic calculation task is used to process the basic weight and the features to be processed.

[0043] According to an embodiment of the present disclosure, the basic calculation task can include performing a matrix multiplication operation on the basic weight and the features to be processed to obtain the target basic features. The fusion result between the target basic features and the target collaborative features can represent the combined calculation result of the relevant basic weight and collaborative weight in the current large - scale model.

[0044] According to the embodiments of the present disclosure, the target basic feature and the target coordination feature are fused to obtain the following features to be processed by the target. It may include adding the target basic feature and the target coordination feature to obtain the following features to be processed by the target. It should be understood that the following features to be processed by the target can also be used to execute the following coordination calculation task and basic calculation task of the large-scale model.

[0045] Note that in the embodiments of the present disclosure, the coordination task can be understood as a coordination calculation task, and the basic task can be understood as a basic calculation task.

[0046] According to the embodiments of the present disclosure, the features to be processed by the target can be determined based on the initial features. For example, when the coordination calculation task and the basic calculation task are the first calculation tasks of the large-scale model, the features to be processed by the target can be obtained based on the initial features. The initial features are obtained based on the input data. The input data may be text data. The input text data can be processed such as tokenization and embedding to obtain the initial features.

[0047] According to the embodiments of the present disclosure, alternatively, the features to be processed by the target may be obtained by the target calculation unit executing the previous coordination calculation task and the previous basic calculation task. For example, when the coordination calculation task and the basic calculation task are the nth calculation tasks among the N calculation tasks of the large-scale model, the features to be processed by the target may be obtained by the target calculation unit executing the (n - 1)th calculation task (the (n - 1)th coordination calculation task and the (n - 1)th basic calculation task). n may be an integer greater than 1 and less than N.

[0048] To facilitate the description of the task execution method provided by the embodiments of the present disclosure, in this embodiment, the first collaborative sub-weight W_LoRA_A can be represented as matrix A, the dimension of matrix A is m*LoRA_rank, the second collaborative sub-weight W_LoRA_B can be represented as matrix B, the dimension of matrix B is n*LoRA_rank, and LoRA_rank can be represented as a rank matrix related to the collaborative weight. It should be understood that the collaborative weight W_LoRA = A * B. m*n can represent the collaborative weight matrix dimension. The feature to be target-processed can be represented as matrix MB1, the dimension of matrix MB1 is m*n, and the basic weight can be represented as W_Base.

[0049] According to the embodiments of the present disclosure, the target calculation unit can perform a matrix multiplication operation on the first collaborative sub-weight and the feature to be target-processed to obtain an intermediate collaborative feature. For example, the first collaborative sub-weight (with dimension m*LoRA_rank) is matrix-multiplied by the feature to be target-processed (matrix dimension is m*k), and the dimension of the resulting intermediate collaborative feature can be m*lora_rank.

[0050] According to the embodiments of the present disclosure, the target calculation unit can perform a matrix multiplication operation on the second collaborative sub-weight and the intermediate collaborative feature to obtain the next feature to be target-processed. For example, the second collaborative sub-weight B (matrix dimension is n*lora_rank) can be matrix-multiplied by the intermediate collaborative feature (matrix dimension is m*lora_rank), and the dimension of the obtained next feature to be target-processed may be m*n.

[0051] According to an embodiment of the present disclosure, since the dimension lora_rank of the ranks of the first collaborative sub-weight and the second collaborative sub-weight may be smaller than m or n, the first collaborative task and the second collaborative task that sequentially execute the collaborative task are determined. Since the dimensions of the matrix multiplication of the first collaborative task and the second collaborative task are reduced, the computational overhead when the target computing unit executes the collaborative task is reduced, the computing efficiency of the target computing unit is improved, and the energy consumption level generated in the process of the target computing unit executing the computing task of the large-scale model is reduced. As a result, a large-scale model can be loaded into an electronic device with low computing performance.

[0052] FIG. 3 schematically shows a schematic diagram of the principle of a task execution method used for a large-scale model according to an embodiment of the present disclosure.

[0053] As shown in FIG. 3, the i-th target feature to be processed may be the target feature to be processed by the large-scale model. The target computing unit may multiply the i-th target feature to be processed by the first collaborative calibration parameter to obtain the calibrated i-th target feature to be processed. The calibrated i-th target feature to be processed and the calibrated first collaborative sub-weight W_Lora_A can be input to the first collaborative task module 311. The first collaborative task module 311 can execute the first collaborative task using the target computing unit and obtain intermediate collaborative features. The calibrated first collaborative sub-weight W_Lora_A can be obtained offline based on the division result between the initial first collaborative sub-weight and the first collaborative calibration parameter after the initial first collaborative sub-weight W_Lora_A is calculated.

[0054] As shown in FIG. 3, the intermediate cooperation feature can be multiplied by the second cooperation calibration parameter to obtain a calibrated intermediate cooperation feature. The calibrated intermediate cooperation feature and the calibrated second cooperation sub-weight W_LoRA_B can be input to the second cooperation task module 312. The second cooperation task module 312 can execute the second cooperation task using a target calculation unit and obtain a calibrated target cooperation feature. The calibrated second cooperation sub-weight W_LoRA_B can be obtained offline based on the division result between the initial second cooperation sub-weight and the second cooperation calibration parameter after the initial second cooperation sub-weight is calculated.

[0055] As shown in FIG. 3, the target calculation unit can also multiply the basic calibration parameter by the i-th feature to be processed to obtain a calibrated i-th basic calibration feature. The i-th basic calibration feature and the calibrated basic weight W_Base are input to the basic calculation task module 320, and the basic calculation task module 320 can execute the basic calculation task using the target calculation unit and obtain a calibrated target basic feature. By adding the calibrated target cooperation feature and the calibrated target basic feature, the i-th feature to be processed can be obtained.

[0056] According to an embodiment of the present disclosure, the first cooperation task includes a plurality of first cooperation subtasks.

[0057] According to an embodiment of the present disclosure, executing a cooperation calculation task using a target calculation unit according to the feature to be processed can include reading, from a target storage unit using the target calculation unit, a first cooperation sub-subtask, a corresponding first cooperation sub-weight, and a sub-feature to be processed.

[0058] According to an embodiment of the present disclosure, the sub - features to be target - processed can be obtained by splitting the features to be target - processed. For example, the features to be target - processed can be split by row and / or column according to a preset splitting strategy to obtain the sub - features to be target - processed. For example, the features to be target - processed may be a feature matrix of 1000 * 1000 dimensions, and the sub - matrix to be target - processed may be a matrix of 1000 * 100 dimensions.

[0059] According to an embodiment of the present disclosure, based on the first collaborative sub - weight and the sub - features to be target - processed, the target calculation unit is used to execute the first collaborative subtask to obtain the first intermediate collaborative sub - features.

[0060] According to an embodiment of the present disclosure, a plurality of first collaborative subtasks are associated with the same first collaborative sub - weight. For example, the i - th first collaborative subtask uses the target calculation unit to perform a matrix multiplication operation on the same first collaborative sub - weight and the i - th sub - features to be target - processed to obtain the i - th first intermediate collaborative sub - features. i may be a positive integer greater than 0.

[0061] According to an embodiment of the present disclosure, when N sub - features to be target - processed are split from the features to be target - processed, the target calculation unit is used to perform matrix multiplication of the first collaborative sub - weight with each of the N sub - features to be target - processed to obtain N first intermediate collaborative sub - features. The N first intermediate collaborative sub - weights may respectively correspond to the N first collaborative subtasks.

[0062] In one example, the target calculation unit includes N, and the i - th target calculation unit can execute the i - th first collaborative subtask. For example, the i - th target calculation unit can perform a matrix multiplication operation on the first collaborative sub - weight and the i - th sub - features to be target - processed to obtain the i - th first intermediate collaborative sub - weight.

[0063] According to an embodiment of the present disclosure, the first intermediate cooperation sub-feature obtained by each of a plurality of target computing units is written into a target memory unit, and the intermediate cooperation feature can be obtained without communication interaction among the plurality of target computing units. Thereby, it is avoided that the target computing unit generates communication overhead by transferring the intermediate value of the intermediate cooperation feature obtained by calculation among the plurality of target computing units. At the same time, the feature to be processed with a large tensor dimension is divided into sub-features to be processed with a small tensor dimension, and a part of the first cooperation task is respectively executed using a plurality of target computing units, so that a plurality of target computing units are used to execute a plurality of first cooperation subtasks in parallel and obtain intermediate cooperation features. Thereby, during the task execution process of the large-scale model, the influence of the computing performance such as the number of threads and the memory space of the target computing unit on the computing efficiency of executing the first cooperation task is reduced, and the overall task execution efficiency of the large-scale model is improved.

[0064] According to an embodiment of the present disclosure, the intermediate cooperation feature can be determined based on the first intermediate cooperation sub-feature corresponding to each of the plurality of first cooperation subtasks. The plurality of first intermediate cooperation sub-features can be fused. For example, the plurality of first intermediate cooperation sub-features can be fused (allreduce) according to the split dimension to obtain the intermediate cooperation feature. The intermediate cooperation feature can be stored in the target memory unit.

[0065] According to an embodiment of the present disclosure, the second cooperation task can include a plurality of second cooperation subtasks.

[0066] According to an embodiment of the present disclosure, executing a cooperative computing task using a target computing unit according to the feature to be processed includes reading, from the target memory unit, a second intermediate cooperation sub-feature corresponding to the second cooperation subtask using the target computing unit, and executing the second cooperation subtask using the target computing unit based on the second intermediate cooperation sub-feature and the second cooperation sub-weight to obtain a target cooperation sub-feature.

[0067] According to an embodiment of the present disclosure, the second intermediate cooperation sub-feature is determined based on the intermediate cooperation feature. For example, the intermediate cooperation feature can be divided into rows and / or columns based on a preset division strategy to obtain the second intermediate cooperation sub-feature. For example, the intermediate cooperation feature is a feature matrix of 1000*1000 dimensions, and the second intermediate cooperation sub-feature can be a matrix of 1000*100 dimensions.

[0068] According to an embodiment of the present disclosure, a plurality of second cooperation subtasks are associated with the same second cooperation sub-weight. For example, for the j-th second cooperation subtask, a matrix multiplication operation can be performed on the same second cooperation sub-weight and the j-th second intermediate cooperation sub-weight using a target calculation unit to obtain the j-th target cooperation sub-feature.

[0069] In one example, the target calculation unit includes N units, and the j-th target calculation unit can execute the j-th second cooperation subtask. For example, the j-th target calculation unit can perform a matrix multiplication operation on the second cooperation sub-weight and the j-th second intermediate cooperation sub-feature to obtain the j-th target cooperation sub-feature.

[0070] According to an embodiment of the present disclosure, the target cooperative sub-features obtained by each of a plurality of target computing units are written into a target memory unit, and the target cooperative feature can be obtained without communication interaction among the plurality of target computing units. Thereby, it is avoided that the target computing unit generates communication overhead by transferring the intermediate value of the intermediate cooperative feature obtained by calculation among the plurality of target computing units. At the same time, the intermediate cooperative feature with a large tensor dimension is divided into a second cooperative sub-feature with a small tensor dimension, and a part of the second cooperative task is respectively executed using a plurality of target computing units, so that a plurality of target computing units are used to execute a plurality of second cooperative subtasks in parallel to obtain the target cooperative feature. Thereby, during the task execution process of the large-scale model, the influence of the computing performance such as the number of threads and the memory space of the target computing unit on the computing efficiency of executing the second cooperative task is reduced, and the overall task execution efficiency of the large-scale model is improved.

[0071] According to an embodiment of the present disclosure, the target cooperative feature can be determined based on the second intermediate cooperative sub-features corresponding to each of the plurality of second cooperative subtasks. For example, the target cooperative feature is obtained by calling a fusion function to perform matrix addition on a plurality of second intermediate cooperative sub-features.

[0072] FIG. 4 schematically shows a schematic diagram of the principle of a task execution method used for a large-scale model according to another embodiment of the present disclosure.

[0073] As shown in FIG. 4, the electronic device 400 may include a plurality of target computing units and a target memory unit 440. The plurality of target computing units 411, 412, 421, 422, 431, and 432. The target computing units 411 and 412 may be used to execute the first cooperative subtask, the target computing units 421 and 422 may be used to execute the second cooperative subtask, and the target computing units 431 and 432 may be used to execute the basic computing task. The target memory unit 440 may include a first target memory area 441 and a second target memory area 442.

[0074] As shown in FIG. 4, the target calculation unit 411 can read the first cooperative sub-weight W_Lora_A and the first sub-feature X11 to be target-processed from the target memory unit 440, and the target calculation unit 411 performs a matrix multiplication operation on the first cooperative sub-weight W_Lora_A and the first sub-feature X11 to be target-processed to obtain the first intermediate cooperative sub-feature. The target calculation unit 412 reads the first cooperative sub-weight W_Lora_A and the second sub-feature X12 to be target-processed from the target memory unit 440, and the target calculation unit 412 performs a matrix multiplication operation on the first cooperative sub-weight W_Lora_A and the second sub-feature X12 to be target-processed to obtain the second intermediate cooperative sub-feature. The first intermediate cooperative sub-feature and the second intermediate cooperative sub-feature can be written into the first target memory area 441. By calling the fusion function, the first intermediate cooperative sub-feature and the second intermediate cooperative sub-feature can be processed to obtain the intermediate cooperative feature, and the intermediate cooperative feature is divided based on a pre-set division strategy to obtain the first second intermediate cooperative sub-feature X21 and the second second intermediate cooperative sub-feature X22.

[0075] As shown in FIG. 4, the target calculation unit 421 can read the second cooperative sub-weight W_Lora_B and the first second intermediate cooperative sub-feature amount X21 from the target memory unit 440, and the target calculation unit 421 performs a matrix multiplication operation on the second cooperative sub-weight W_Lora_B and the first second intermediate cooperative sub-feature amount X21 to obtain the first target cooperative sub-feature. The target calculation unit 422 can read the second cooperative sub-weight W_Lora_B and the second second intermediate cooperative sub-feature X22 from the target memory unit 440, and the target calculation unit 422 performs a matrix multiplication operation on the second cooperative sub-weight W_Lora_B and the second second intermediate cooperative sub-feature X22 to obtain the second target cooperative sub-feature. The first target cooperative sub-feature and the second target cooperative sub-feature can be written into the second target memory area 442. By calling the fusion function, the first target cooperative sub-feature and the second target cooperative sub-feature can be processed to obtain the target cooperative feature.

[0076] As shown in FIG. 4, the target calculation unit 431 may read the first basic sub-weight W_Base1 and the feature MB1 to be target-processed from the target memory unit 440. The target calculation unit 431 can execute a matrix multiplication operation on the first basic sub-weight W_Base1 and the feature MB1 to be target-processed to obtain an intermediate basic feature MB2. The intermediate basic feature can be written into the first target memory area 441. The target calculation unit 432 can read the second basic sub-weight W_Base2 and the intermediate basic feature MB2 from the first target memory area 441. The target calculation unit 432 can execute a matrix multiplication operation on the second basic sub-weight W_Base2 and the intermediate basic feature MB2 to obtain a target basic feature. The target basic feature can be written into the second target memory area 442. By calling a fusion function to process the target basic feature and the target cooperation feature, the next feature to be target-processed can be obtained.

[0077] According to an embodiment of the present disclosure, the basic weight W_Base can be processed based on a matrix multiplication mechanism of a general matrix to obtain a first basic sub-weight W_Base1 and a second basic sub-weight W_Base2.

[0078] FIG. 5 schematically shows a flowchart of a task execution method used for a large-scale model according to another embodiment of the present disclosure.

[0079] As shown in FIG. 5, the task execution method may include operations S501 to S506.

[0080] In operation S501, a feature to be target-processed is obtained. The tensor dimension of the feature to be target-processed can be set to m*k.

[0081] In operation S502, target basic feature calculation is executed. For example, a target calculation unit is used to process a basic weight and a feature to be target-processed to obtain a target basic feature.

[0082] In operation S503, it is determined whether the dimension of the feature to be processed is greater than a predetermined dimension threshold. For example, it can be determined whether the dimension m of the feature to be processed is greater than the predetermined dimension threshold.

[0083] If the determination result of operation S503 is "yes", operation S504 can be executed, and the Tensor Core performs the calculation. For example, the Tensor Core (also known as Tensor Core or tensor computing core) in the target computing unit executes the first cooperative task and the second cooperative task to obtain the target cooperative feature.

[0084] In operation S506, a feature fusion operation is performed. For example, the target cooperative feature and the target basic feature are added to obtain the next feature to be processed.

[0085] If the determination result of operation S503 is "no", operation S505 is executed, and the calculation can be performed using the CUDA (Compute Unified Device Architecture) computing core. For example, the CUDA core in the target computing unit can be used to execute the first cooperative task and the second cooperative task to obtain the target cooperative feature. Next, operation S506 is executed to obtain the next feature to be processed.

[0086] According to the embodiments of the present disclosure, by determining the dimension of the feature to be processed and calling the corresponding type of computing core in the target computing unit according to the determination result to execute the cooperative computing task, the architecture of the computing core of the target computing unit (such as GPU) can be utilized to the maximum extent, and the execution efficiency of the large-scale model task can be improved.

[0087] FIG. 6 schematically shows an application scenario diagram of a task execution method used in a large-scale model according to an embodiment of the present disclosure.

[0088] As shown in FIG. 6, the task execution method used for the large-scale model can be realized by installing a task management process 610 and a task execution process 620. The task execution process 620 can execute an operation S601 to apply for video memory space to the task management process 610.

[0089] The task management process 620 can execute an operation S602 to share the weight memory address of the cooperative weight with the task execution process 620. The task execution process 620 can execute an operation S603 to execute a cooperative computing task. For example, the first cooperative sub-weight or the second cooperative sub-weight can be called from the video memory of the target computing unit via the shared weight memory address, and the first cooperative sub-task can be executed, or the second cooperative sub-task can be executed.

[0090] The task management process 610 can also execute an operation S604 to update the cooperative weight. For example, at least one weight matrix of the first cooperative sub-weight or the second cooperative sub-weight may be updated. When the task execution process 620 executes an operation S605 and executes a subsequent cooperative computing task, the updated first cooperative sub-weight or the updated second cooperative sub-weight updated from the video memory of the target computing unit through the weight memory address shared based on the previous task management process can be called, thereby executing the first cooperative sub-task or the second cooperative sub-task.

[0091] According to the embodiments of the present disclosure, by sharing the weight memory address in the task management process and updating the first cooperative sub-weight and the second cooperative sub-weight, the task execution process can asynchronously load the updated first cooperative sub-weight and execute a cooperative computing task under unrecognized conditions, or load the updated second cooperative sub-weight and execute a cooperative computing task, thereby realizing hot updates of the first cooperative sub-weight and the second cooperative sub-weight during the process of the large-scale model executing a task, and shortening the computing time generated for updating the weight memory.

[0092] FIG. 7 schematically shows a schematic diagram of the principle of a task execution method used for a large-scale model according to another embodiment of the present disclosure.

[0093] As shown in FIG. 7, the features to be processed by the large-scale model can be input to a basic task execution module 711 and a first cooperative task execution module 721 respectively. The basic task execution module 711 processes the features to be processed by executing basic calculation tasks, and the obtained calculation results can be input to a basic integration module 712 for fusion to obtain target basic features. The first cooperative task execution module 721 can execute a plurality of first cooperative subtasks to process the features to be processed and obtain a plurality of first intermediate cooperative features. The plurality of first intermediate cooperative sub-features can be input to a cooperative integration module 722 and fused to obtain a plurality of second intermediate cooperative sub-features. The plurality of second intermediate cooperative sub-features are input to a second cooperative task execution module 723. The second cooperative task execution module 723 can execute a plurality of second cooperative subtasks to obtain a plurality of second intermediate cooperative sub-features. The target basic features and the plurality of second intermediate sub-features can be sent to a target integration module 730 to obtain the next features to be processed.

[0094] As shown in FIG. 7, based on the basic task execution module 711 and the basic integration module 712, a basic task flow for the basic model to execute basic calculation tasks can be realized. Also, based on the first cooperative task execution module 721, the cooperative integration module 722, and the second cooperative task execution module 723, a cooperative task flow for the cooperative calculation task to execute can also be executed.

[0095] According to an embodiment of the present disclosure, the features to be processed include text features to be processed, the initial features are determined based on the initial text, and the execution results of the target calculation unit executing the basic calculation task and the cooperative settlement task are the output text corresponding to the initial text.

[0096] For example, the initial text may be the question text input by the user, and the output text may be the answer text corresponding to the question text.

[0097] It will be understood that the present disclosure has been described above by taking, as an example, the input data of the large model as text. However, the present disclosure is not limited thereto, and the input data of the large model may be an image or audio.

[0098] In some embodiments, the feature to be target - processed is the image feature to be target - processed, the initial feature is obtained based on the initial image, and the execution result of the target calculation unit executing a plurality of basic calculation tasks and a plurality of cooperative settlement tasks is the output result corresponding to the initial image. As a result, it may be an adjusted image or text. When the input data is an image, the input image can be processed based on the patch embedding operation to obtain the image feature to be target - processed. The above - mentioned feature to be target - processed may be, for example, edges or colors in the image.

[0099] FIG. 8 schematically shows a block diagram of a task execution device used for a large model according to an embodiment of the present disclosure.

[0100] As shown in FIG. 8, the task execution device 80 for the large model may include a target storage unit 810 and a target calculation unit 820.

[0101] The target storage unit 810 stores the cooperative calculation tasks.

[0102] The target calculation unit 820 executes a collaborative calculation task according to the features to be target - processed, obtains target collaborative features. Here, the collaborative calculation task includes a first collaborative task for processing the features to be target - processed and a first collaborative sub - weight to obtain intermediate collaborative features, and a second collaborative task for processing the intermediate collaborative features and a second collaborative sub - weight to obtain target collaborative features. The first collaborative sub - weight and the second collaborative sub - weight are determined by processing collaborative weights according to the matrix multiplication mechanism of a general matrix. The target basic features and the target collaborative features are fused to obtain the next features to be target - processed. Here, the target basic features are obtained by the target calculation unit executing a basic calculation task, and the basic calculation task is used to process the basic weight and the features to be target - processed, and is arranged as such.

[0103] According to an embodiment of the present disclosure, the first collaborative task includes a plurality of first collaborative subtasks.

[0104] According to an embodiment of the present disclosure, the target calculation unit reads, from the target storage unit, a first collaborative sub - weight corresponding to the first collaborative subtask and sub - features to be target - processed according to the features to be target - processed so as to execute the collaborative calculation task. The sub - features to be target - processed are obtained by dividing the features to be target - processed. Based on the first collaborative sub - weight and the sub - features to be target - processed, the target calculation unit is used to execute the first collaborative subtask to obtain a first intermediate collaborative sub - feature. Here, the intermediate collaborative features are determined based on the first intermediate collaborative sub - features corresponding to each of the plurality of first collaborative subtasks, and is arranged to execute as such.

[0105] According to an embodiment of the present disclosure, the plurality of first collaborative subtasks are associated with the same first collaborative sub - weight.

[0106] According to an embodiment of the present disclosure, the second collaborative task includes a plurality of second collaborative subtasks.

[0107] According to an embodiment of the present disclosure, the target computing unit reads, from the target memory unit, a second intermediate cooperation sub-feature corresponding to a second cooperation subtask so as to execute a cooperation computing task according to the feature to be processed by the target, where the second intermediate cooperation sub-feature is determined based on an intermediate cooperation feature, and based on the second intermediate cooperation sub-feature and a second cooperation sub-weight, executes the second cooperation subtask to obtain a target cooperation sub-feature, where the target cooperation feature is determined according to the target cooperation sub-features corresponding to each of a plurality of second cooperation subtasks, and is arranged to execute the above.

[0108] According to an embodiment of the present disclosure, a plurality of second cooperation subtasks are associated with the same second cooperation sub-weight.

[0109] According to an embodiment of the present disclosure, the feature to be processed by the target is determined based on an initial feature.

[0110] According to an embodiment of the present disclosure, the feature to be processed by the target may also be obtained by the target computing unit executing a previous cooperation computing task and a previous basic computing task.

[0111] According to an embodiment of the present disclosure, the feature to be processed by the target includes a text feature to be processed by the target, is determined based on an initial feature, and the execution result of the target computing unit executing a basic computing task and a cooperation settlement task is an output text corresponding to the initial text.

[0112] FIG. 9 schematically shows a block diagram of a task execution device used for a large-scale model according to an embodiment of the present disclosure.

[0113] As shown in FIG. 9, a task execution device 9000 used for a large-scale model may include a task execution device 80 used for the large-scale model.

[0114] In the technical solution of the present disclosure, any processing such as collection, storage, use, processing, transfer, provision, disclosure, and application of the personal information of such users shall comply with the provisions of relevant laws and regulations, adopt necessary confidentiality measures, and not violate public order and good customs.

[0115] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program.

[0116] According to an embodiment of the present disclosure, the electronic device includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor executes the method.

[0117] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium stores computer instructions, and the computer instructions cause a computer to execute the method.

[0118] According to an embodiment of the present disclosure, the computer program realizes the above method when executed by a processor.

[0119] FIG. 10 schematically shows a block diagram of an electronic device suitable for implementing a code generation method based on a large-scale model according to an embodiment of the present disclosure. The electronic device represents various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can further represent various forms of mobile devices, such as personal digital processes, mobile phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and relationships, and their functions shown in this specification are merely illustrative and do not limit the implementation of the present disclosure described and / or claimed in this specification.

[0120] As shown in FIG. 10, the device 1000 includes a computing unit 1001 and can execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 1002 or a computer program loaded from a storage unit 1008 into a random access memory (RAM) 1003. The RAM 1003 can further store various programs and data necessary for the operation of the device 1000. The computing unit 1001, the ROM 1002, and the RAM 1003 are interconnected by a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.

[0121] A plurality of components in the device 1000 are connected to the I / O interface 1005, including an input unit 1006 such as a keyboard and a mouse, an output unit 1007 such as various types of displays and speakers, a storage unit 1008 such as a magnetic disk and an optical disk, and a communication unit 1009 such as a network card, a modem, and a wireless communication transceiver. The communication unit 1009 allows the device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0122] The computing unit 1001 can be various general-purpose and / or dedicated processing modules with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, computing units for various running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes each of the methods and processes described above, such as the code generation method based on a large-scale model. For example, in some embodiments, the code generation method based on a large-scale model is realized as a computer software program and is tangibly included in a machine-readable medium such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed in the device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the code generation method based on the large-scale model described above can be executed. Alternatively, in other embodiments, the computing unit 1001 may be arranged to execute the code generation method based on a large-scale model in any other suitable manner (e.g., firmware).

[0123] The various embodiments of the systems and techniques described in this specification can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on a chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can be implemented in one or more computer programs that are executed and / or interpreted in a programmable system including at least one programmable processor, where the programmable processor can be a special-purpose or general-purpose programmable processor, and can receive data and instructions from a memory system, at least one input device, and at least one output device, and can transmit the data and instructions to the memory system, the at least one input device, and the at least one output device.

[0124] The program code for implementing the methods of the present disclosure can be created in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations defined in the flowchart and / or block diagram are implemented. The program code may be executed entirely by a machine, partially by a machine, partially by a device as an independent software package and partially by a remote machine, or entirely by a remote machine or server.

[0125] In the context of the present disclosure, a machine-readable medium may be a tangible medium that can contain or store a program for use by or in combination with an instruction execution system, apparatus, or device. The machine-readable medium may be either a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium includes, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium include, but are not limited to, electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0126] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer, which includes a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball), by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with a user. For example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user can be received in any form (including voice input, speech input, or tactile input).

[0127] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser, where the user can interact with embodiments of the systems and techniques described herein via the graphical user interface or the network browser), or a computing system that includes any combination of such background components, middleware components, or front-end components. The components of the system can be connected to each other by digital data communication in any form or medium (e.g., a communication network). Exemplifications of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0128] A computer system can include clients and servers. Clients and servers are generally separated from each other and usually interact via a communication network. A client-server relationship is created by computer programs that run on corresponding computers and have a client-server relationship with each other. The server can be a cloud server, a distributed system server, or a server combined with a blockchain.

[0129] It should be understood that various forms of the flows shown above may be used, and each operation may be sorted again, added, or deleted. For example, each step described in the present disclosure may be executed in parallel, sequentially, or in a different order, and the present specification is not limited here as long as the desired results of the technical solutions disclosed in the present disclosure can be achieved.

[0130] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.

Claims

1. Executing a cooperative computing task using a target computing unit according to a target feature to be processed to obtain a target cooperative feature, where the cooperative computing task includes a first cooperative task for processing the target feature to be processed and a first cooperative sub-weight to obtain an intermediate cooperative feature, and a second cooperative task for processing the intermediate cooperative feature and a second cooperative sub-weight to obtain the target cooperative feature, where the first cooperative sub-weight and the second cooperative sub-weight are determined by processing the cooperative weights according to a general matrix multiplication mechanism; Fusing the target basic feature and the target collaborative feature to obtain a next target feature to be processed, where the target basic feature is obtained by performing a basic calculation task using the target computing unit, and the basic calculation task is used to process a basic weight and the target feature to be processed. Task execution methods used for large scale models.

2. the first cooperative task includes a plurality of first cooperative subtasks; wherein executing a collaborative computing task using a target computing unit according to the target feature to be processed includes: using the target computation unit to read out a first cooperative sub-weight and a target sub-feature corresponding to the first cooperative sub-task from a target storage unit, the target sub-feature being obtained by dividing the target feature; and executing the first cooperative sub-task using the target computing unit based on the first cooperative sub-weights and the target sub-feature to be processed to obtain a first intermediate cooperative sub-feature, where the intermediate cooperative feature is determined based on first intermediate cooperative sub-features corresponding to each of a plurality of the first cooperative sub-tasks. The method of claim 1.

3. A plurality of the first cooperative subtasks are associated with the same first cooperative subweight. The method of claim 2.

4. the second cooperative task includes a plurality of second cooperative subtasks; wherein executing a collaborative computing task using a target computing unit according to the target feature to be processed includes: reading, using the target computing unit, a second intermediate coordination sub-characteristic corresponding to the second coordination sub-task from a target storage unit, the second intermediate coordination sub-characteristic being determined based on the intermediate coordination characteristic; and executing the second cooperative sub-task using the target computation unit based on the second intermediate cooperative sub-feature and the second cooperative sub-weight to obtain a target cooperative sub-feature, where the target cooperative feature is determined according to a target cooperative sub-feature corresponding to each of a plurality of the second cooperative sub-tasks. The method according to any one of claims 1 to 3.

5. A plurality of the second cooperative sub-tasks are associated with the same second cooperative sub-weight. The method according to claim 4.

6. the target feature is determined based on the initial feature; or The target processing feature is obtained by the target computing unit performing the previous collaborative computing task and the previous basic computing task. The method of claim 1.

7. The target features to be processed include target text features to be processed, and the initial features are determined according to an initial text, and an execution result of the target computing unit executing the basic computing task and the collaborative computing task is an output text corresponding to the initial text. The method according to claim 6.

8. a target storage unit for storing a collaborative computing task; execute a cooperative computation task according to a target feature to be processed to obtain a target cooperative feature, where the cooperative computation task includes a first cooperative task for processing the target feature to be processed and a first cooperative sub-weight to obtain an intermediate cooperative feature, and a second cooperative task for processing the intermediate cooperative feature and a second cooperative sub-weight to obtain the target cooperative feature, where the first cooperative sub-weight and the second cooperative sub-weight are determined by processing the cooperative weights according to a general matrix matrix multiplication mechanism; The target basic feature and the target cooperative feature are fused to obtain a next target feature to be processed, where the target basic feature is obtained by executing a basic calculation task in the target calculation unit, and the basic calculation task is used to process a basic weight and the target feature to be processed. A target computing unit; A task execution device used for large-scale models.

9. the first cooperative task includes a plurality of first cooperative subtasks; wherein the target computing unit performs a collaborative computing task according to the target feature to be processed; Reading a first cooperative sub-weight and a target sub-feature corresponding to the first cooperative sub-task from a target storage unit, the target sub-feature being obtained by dividing the target feature; and executing the first cooperative sub-task using the target computation unit based on the first cooperative sub-weights and the target sub-feature to be processed to obtain a first intermediate cooperative sub-feature, where the intermediate cooperative feature is determined based on first intermediate cooperative sub-features corresponding to each of a plurality of the first cooperative sub-tasks.

9. The apparatus of claim 8.

10. A plurality of the first cooperative subtasks are associated with the same first cooperative subweight.

10. The apparatus of claim 9.

11. the second cooperative task includes a plurality of second cooperative subtasks; wherein the target computing unit performs a collaborative computing task according to the target feature to be processed; reading a second intermediate collaboration sub-characteristic corresponding to the second collaboration sub-task from a target storage unit, the second intermediate collaboration sub-characteristic being determined based on the intermediate collaboration characteristic; and performing the second cooperative sub-task based on the second intermediate cooperative sub-feature and the second cooperative sub-weight to obtain a target cooperative sub-feature, where the target cooperative sub-feature is determined according to a target cooperative sub-feature corresponding to each of a plurality of the second cooperative sub-tasks. An apparatus according to any one of claims 8 to 10.

12. A plurality of the second cooperative sub-tasks are associated with the same second cooperative sub-weight.

12. The apparatus of claim 11.

13. the target feature is determined based on the initial feature; or The target processing feature is obtained by the target computing unit performing the previous collaborative computing task and the previous basic computing task.

9. The apparatus of claim 8.

14. The target features to be processed include target text features to be processed, and the initial features are determined according to an initial text, and an execution result of the target computing unit executing the basic computing task and the collaborative computing task is an output text corresponding to the initial text.

14. The apparatus of claim 13.

15. Including a device according to any one of claims 8 to 10 Task execution equipment used for large scale models.

16. At least one processor; a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor performs the method of any one of claims 1 to 3. electronic equipment.

17. A non-transitory computer readable storage medium having stored thereon computer instructions for causing a computer to carry out the method according to any one of claims 1 to 3.

18. A computer program which, when executed by a processor, implements the method according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Neural network adjustment method and device, electronic equipment and readable storage medium

    CN117829228A

  • Dialogue model training method, dialogue processing method and device

    CN117951264A

  • Multi-task large model training method and device

    CN118052999A

  • Figure graph and model training method and device, electronic equipment and storage medium

    CN118155023A

  • Low-Rank Adaptation of Neural Network Models

    US20220383126A1