Task execution method used for large-scale model, device, apparatus, medium, and program
The task execution method for large-scale models addresses the challenge of balancing performance and efficiency by executing tasks based on specific model parameters, resulting in improved inference efficiency and reduced resource requirements.
Patent Information
- Application Number
- JP2025052917
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-06-17
- Filing Date
- 2025-03-27
- Publication Date
- 2025-06-19
- Estimated Expiration
- 2045-03-27
AI Technical Summary
Large-scale models face challenges in balancing high performance and efficient processing, as they often require executing all parameters for each task, leading to redundant computations and low inference efficiency.
The proposed task execution method for large-scale models involves executing modal root tasks, domain root tasks, and feed-forward tasks based on specific model parameters stored in a target storage unit, allowing for targeted processing and reducing the need for redundant computations.
This approach ensures high model performance while improving inference efficiency by only utilizing the necessary parameters for each task, thereby reducing hardware resource requirements and energy consumption.
Smart Images

Figure 2025092576000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular, to the fields of deep learning technology and large-scale model technology. Specifically, it relates to a task execution method, apparatus, electronic device, storage medium, and program used in a large-scale model.
Background Art
[0002] Artificial intelligence technology is an interdisciplinary subject with a wide range of related fields, including both hardware-level technology and software-level technology. Generally, artificial intelligence technology includes large-scale model technology. Large-scale model technology can be widely applied to various fields of artificial intelligence. For example, text processing, semantic understanding, machine translation, man-machine interaction, etc. To execute tasks in different fields by a large-scale model, it is necessary to comprehensively consider factors such as timeliness and cost.
Summary of the Invention
Problems to be Solved by the Invention
[0003] The present disclosure provides a task execution method, apparatus, electronic device, storage medium, and program used in a large-scale model.
Means for Solving the Problems
[0004] According to one aspect of the present disclosure, there is provided a task execution method for a large-scale model. The task execution method includes: executing a modal root task by a target computing unit based on features to be processed, to obtain a modal recognition result; executing a domain root task by the target computing unit based on the features to be processed and target domain gating model parameters, to obtain a domain recognition result, wherein the target domain gating model parameters are domain gating model parameters corresponding to the modal recognition result read from a target storage unit; and executing a feed-forward task by the target computing unit based on the features to be processed and target feed-forward task model parameters, to obtain a task execution result, wherein the target feed-forward task model parameters are feed-forward task model parameters corresponding to the domain recognition result read from the target storage unit.
[0005] According to another aspect of the present disclosure, there is provided a task execution device for a large-scale model, including a target storage unit that stores a plurality of domain gating model parameters and a plurality of feed-forward task model parameters, and a target calculation unit, wherein the target calculation unit is configured as follows: Based on the features to be processed, the target calculation unit executes a modal root task to obtain a modal recognition result; based on the features to be processed and the target domain gating model parameters, the target calculation unit executes a domain root task to obtain a domain recognition result, wherein the target domain gating model parameters are the domain gating model parameters corresponding to the modal recognition result read from the target storage unit; based on the features to be processed and the target feed-forward task model parameters, the target calculation unit executes a feed-forward task to obtain a task execution result, wherein the target feed-forward task model parameters are the feed-forward task model parameters corresponding to the domain recognition result read from the target storage unit.
[0006] According to another aspect of the present disclosure, there is provided an electronic device, including at least one processor and a memory communicatively connected to the at least one processor, wherein instructions executable by the at least one processor are stored in the memory, and the instructions are executed by the at least one processor so that the at least one processor can execute the above method.
[0007] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the above method.
[0008] According to another aspect of the present disclosure, there is provided a computer program that, when executed by a processor, implements the above method.
[0009] It should be understood that the content described in this section is not for identifying the keys or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will be readily understood from the following description.
[0010] The drawings are provided to better understand the present invention and do not limit the present disclosure.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4A
Figure 4B
Figure 5
Figure 6
Modes for Carrying Out the Invention
[0012] Exemplary embodiments of the present disclosure will be described below with reference to the drawings. For ease of understanding, various details of the embodiments of the present disclosure are included, which are merely exemplary. Therefore, those skilled in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0013] In the field of deep learning, the application of large models is constantly expanding. Large models may include, for example, large language models (LLMs), image large models, and audio large models, etc. Large models exhibit super-strong task processing capabilities. However, usually, it is difficult to guarantee both the high model performance of large models and excellent processing efficiency.
[0014] The large model may be a pre-trained large model. The architecture of the large model may be a Transformer (encoding and decoding) architecture. The Transformer architecture has strong data processing capabilities and flexibility, and its application in the field of natural language processing has been extended. The large model of the Transformer architecture may include multiple processing layers. The structure of each processing layer may be the same and may include a module based on the multi-head self-attention mechanism (MHA) and a module based on the feed-forward neural network (FFN).
[0015] All parameters in each processing layer process the input data. Each parameter in the processing layer plays a different role for input data of different field types and different modal types. This means that every time a task is executed, the entire large model is executed, but in fact, only a few parameters can function. If the parameters are redundant, the model inference efficiency may be low.
[0016] Thus, in order to ensure high model performance of the large model while having excellent inference efficiency, the present disclosure provides a task execution method used for the large model. This will be described below.
[0017] FIG. 1 schematically shows an exemplary system architecture to which a task execution method and apparatus used for a large model according to an embodiment of the present disclosure can be applied.
[0018] Note that, for those skilled in the art to understand the technical content of the present disclosure, FIG. 1 is an illustration of a system architecture to which the embodiments of the present disclosure can be applied, and it does not mean that the embodiments of the present disclosure cannot be used in other devices, systems, environments or scenarios.
[0019] As shown in FIG. 1, the system architecture according to this embodiment may include a terminal device 101, a network 102, and a server cluster 103. The network 102 is used to provide a medium for communication links between the terminal device 101 and the server cluster 103. The network 102 may also be used to provide a medium for communication links inside the server cluster 103. The network 102 may include various connection types such as, for example, wired and / or wireless communication links.
[0020] A user can use the terminal device 101 to interact with the server cluster 103 via the network 102 and send and receive messages and the like. For example, the terminal device 101 may send a request to the server cluster 103 via the network 102 to train a deep learning model.
[0021] Various communication client applications such as, for example, a knowledge browsing application, a web page browser application, a search application, an instant messaging tool, an email box client, and / or social platform software may be installed on the terminal device 101 (this is just an example).
[0022] The terminal device 101 may be various electronic devices having a display and supporting web page browsing, including but not limited to smartphones, tablet computers, laptop portable computers, desktop computers, etc.
[0023] The server cluster 103 may be a server that provides various services, for example, a background management server (just an example) that provides support for requests sent by the user using the terminal device 101.
[0024] The server cluster 103 may be a cloud server, also called a cloud computing server or a cloud host, which is one of the host products in the cloud computing service system, and solves the drawbacks of the large management difficulty and weak service scalability existing in the conventional physical host and VPS service (abbreviated as "Virtual Private Server" or "VPS"). The server may be a server of a distributed system or a server combined with a blockchain.
[0025] The server cluster 103 includes a plurality of server nodes 1031, 1032, 1033, 1034, and each server node includes one or more hardware devices. The server cluster 103 or the server node can execute the task execution method used in the large-scale model according to the present disclosure, and realize the deployment, inference or training of the large-scale model with low computing resources.
[0026] The system architecture of the present disclosure has been described above. Next, the method of the present disclosure will be described.
[0027] In the technical solution of the present disclosure, any processing such as the collection, storage, use, processing, transmission, provision, disclosure, and application of such user personal information complies with the provisions of relevant regulations, takes necessary security measures, and does not violate public order and good customs.
[0028] In the technical solution of the present disclosure, before obtaining or collecting the personal information of a user, obtain the approval or consent of the user.
[0029] FIG. 2 schematically shows a flowchart of a task execution method used in a large-scale model according to an embodiment of the present disclosure.
[0030] As shown in FIG. 2, the method includes operations S210-S230.
[0031] In operation S210, based on the features to be processed, a modal root task is executed by a target computing unit to obtain a modal recognition result.
[0032] In an embodiment of the present disclosure, the target computing unit may include at least one of a central processing unit (CPU), a graphics processing unit (GPU), and an artificial intelligence computing unit. The artificial intelligence computing unit may include at least one of a neural network processing unit (NPU), a tensor processing unit (TPU), and a Kunlun chip.
[0033] In an embodiment of the present disclosure, the processing layer of the large-scale model may include a modal router. The modal router performs modal recognition on the features to be processed and is used to obtain a modal recognition result. The modal router may be used to perform modal recognition on the features to be processed. The modal router may be implemented using a linear layer, such as Softmax. Based on the modal root model parameters of the modal router by the target computing unit, the features to be processed can be processed to obtain a modal recognition result and complete the modal root task.
[0034] In an embodiment of the present disclosure, the modal recognition result can represent the modal type of the features to be processed. For example, the modal type of the features to be processed may include at least one of text, image, or audio.
[0035] In operation S220, based on the features to be processed and the target domain gating model parameters, the target calculation unit executes a domain routing task to obtain a domain recognition result.
[0036] In an embodiment of the present disclosure, the domain router included in the processing layer of the large-scale model performs domain recognition on the features to be processed and is used to obtain a domain recognition result. The domain router may be implemented using a linear layer, for example, Softmax. The target calculation unit can execute a domain routing task on the features to be processed based on the target domain gating model parameters to obtain a domain recognition result.
[0037] In an embodiment of the present disclosure, the domain recognition result can represent the type of the domain of the features to be processed. For example, the type of the domain of the features to be processed may include at least one of translation, question answering, search, text generation, and intent recognition.
[0038] In an embodiment of the present disclosure, each processing layer of the large-scale model can arrange a plurality of field gates, such as Router1, Router2, and Router3. Each field gate corresponds to one type of modality. For example, Router1 corresponds to the text modality type, Router2 corresponds to the image modality type, and Router3 corresponds to the audio modality type.
[0039] In an embodiment of the present disclosure, the target domain gating model parameters may be the domain gating model parameters corresponding to the modality recognition result read from the target storage unit. For example, the respective domain gating model parameters of Router1, Router2, and Router3 may be stored in the target storage unit. When the modality recognition result indicates that the modality type of the data to be processed is the text modality type, the domain gating model parameters of Router1 can be read from the target storage unit as the target domain gating model parameters.
[0040] In operation S230, based on the features to be processed and the target feed-forward task model parameters, use the target calculation unit to execute the feed-forward task and obtain the task execution result.
[0041] In an embodiment of the present disclosure, the processing layer of the large-scale model may include a plurality of modal expert modules, and each modal expert module matches one type of modality. Each modal expert module includes a plurality of field expert sub-modules. Each field expert sub-module matches one type of field.
[0042] In an embodiment of the present disclosure, each field expert sub-module may include at least one feed-forward neural network. Each field expert sub-module processes the features to be processed, for example, performs vector mapping, and obtains the task execution result. The target calculation unit can execute the feed-forward task on the features to be processed based on the target feed-forward task model parameters of the target field expert sub-module and obtain the task execution result.
[0043] In an embodiment of the present disclosure, a plurality of modal expert modules corresponding to each type of modality, for example, modal expert modules Expert1, Expert1, and Expert1, may be arranged in each processing layer of the large-scale model. Each modal expert module may include a plurality of field expert sub-modules. For example, modal expert module Expert1 includes field expert sub-modules Expert1-1, Expert1-2, and Expert1-3. Each field expert sub-module corresponds to one type of field. For example, field expert sub-module Expert1-1 corresponds to the type of the question-and-answer field, field expert sub-module Expert1-2 corresponds to the type of the search field, and field expert sub-module Expert1-3 corresponds to the type of the translation field.
[0044] In an embodiment of the present disclosure, the target feedforward task model parameter is a target feedforward task model parameter corresponding to a domain recognition result read from a target memory unit. For example, the respective feedforward task model parameters of the domain expert sub-modules Expert1-1, Expert1-2, and Expert1-3 may be stored in the target memory unit. When the modal type of the data to be processed by the modal recognition result represents the type of the question-and-answer area, the target feedforward task model parameter may be the feedforward task model parameter of the domain expert sub-module Expert1-1.
[0045] According to an embodiment of the present disclosure, in the processing layer of a large-scale model, a plurality of domain expert sub-modules are combined by an integrated learning method to form a plurality of modal expert modules, which can process data of different modal types, thereby improving the application ability and application range of the large-scale model and improving the universality of the large-scale model. In addition, each domain expert sub-module focuses on solving the feed-forward tasks of specific domain types in a specific modal type, thereby improving the application granularity of the large-scale model and improving the responsiveness of the large-scale model. In addition, when executing the feed-forward tasks of specific domain types in a modal type, the target computing unit only needs to execute the feed-forward task based on the target feed-forward task model parameters, without the need to utilize all the feed-forward task model parameters in all domain expert sub-modules, improving the processing efficiency of the target computing unit and reducing the requirements for hardware resources and energy consumption of the target computing unit. In addition, a modal gate, a plurality of domain gates, and a plurality of modal expert modules are combined to form a complete processing layer. The modal gate dynamically adapts the target domain gate based on the data characteristics of the features to be processed, and the target domain gate dynamically adapts the target domain expert sub-module in the target modal expert module based on the data characteristics of the features to be processed, which can improve the inference performance of the activated target feed-forward task model parameters. Thereby, the inference efficiency of the large-scale model is guaranteed and the inference performance is improved.
[0046] The above outlines the method of the present disclosure. Hereinafter, the task execution method used in the large-scale model of the present disclosure will be further described.
[0047] FIG. 3 is a structural schematic diagram of a processing layer of a large-scale model according to an embodiment of the present disclosure.
[0048] As shown in FIG. 3, a single processing layer of the large-scale model includes a modal gate Router0, a first field gate Router1, a second field gate Router2, ···, an nth field gate Routern, a first modal expert module Expert1, a second modal expert module Expert2, ···, an nth modal expert module Expertn. The first modal expert module Expert1 corresponding to the first field gate Router1 includes a first field expert sub-module Expert1-1, a second field expert sub-module Expert1-2, ···, an ith field expert sub-module Expert1-i. The second modal expert module Expert2 corresponding to the second field gate Router2 includes a first field expert sub-module Expert2-1, a second field expert sub-module Expert2-2, ···, an ith field expert sub-module Expert2-i. The nth modal expert module Expertn corresponding to the nth field gate Routern includes a first field expert sub-module Experttn-1, a second field expert sub-module Experttn-2, ···, an ith field expert sub-module Expertn-i.
[0049] The corresponding modal types of the respective field gating model parameters of the first field gate Router1, the second field gate Router2, ···, the nth field gate Routern are different.
[0050] The corresponding field types of the respective feed-forward task model parameters of the multiple field expert sub-modules corresponding to the same field gate are different.
[0051] The modal gating model parameters, the multiple field gating model parameters, and the multiple feed-forward task model parameters may be pre-stored in the target storage unit.
[0052] As shown in FIG. 3, based on the feature 310 to be target - processed and the modal gating model parameters of the modal gates stored from the target memory unit, the target computing unit can execute a modal route task to obtain a modal recognition result. For example, the modal recognition result indicates that the modal type of the feature to be target - processed corresponds to the first - field gate Router1. Based on the feature to be target - processed and the target - field gating model parameters of the first - field gate Router1, the target computing unit can execute a field route task to obtain a field recognition result. For example, the field recognition result represents that the field type of the feature to be target - processed corresponds to the first - field expert sub - module Expert1 - 1. Based on the feature to be target - processed and the target feed - forward task model parameters of the first - field expert sub - module Expert1 - 1, the target computing unit can be used to execute a feed - forward task to obtain a task execution result 320.
[0053] According to an embodiment of the present disclosure, the task execution method used for a large - scale model may further include obtaining target - field gating model parameters from a plurality of field gating model parameters stored in the target memory unit based on the modal recognition result. Based on the feature to be target - processed and the target - field gating model parameters, the target computing unit executes a field route task to obtain a field recognition result.
[0054] According to an embodiment of the present disclosure, the task execution method used for a large - scale model may further include obtaining target feed - forward task model parameters from a plurality of feed - forward task model parameters stored in the target memory unit based on the field recognition result. Based on the feature to be target - processed and the target feed - forward task model parameters, the target computing unit is used to execute a feed - forward task to obtain a task execution result.
[0055] In an embodiment of the present disclosure, the input data of the large-scale model may be data to be processed whose modal type is text and whose field type is question-and-answer. After being processed by the preprocessing layer of the processing layer, the features to be processed are obtained. After passing through the task execution method according to the embodiment of the present disclosure, the obtained task execution result can be used to obtain the output result of the large-scale model. The output result may be an answer corresponding to the data to be processed.
[0056] According to an embodiment of the present disclosure, the modal type may include at least one of image, text, and audio. The field type may include at least one of translation data, question-and-answer data, search data, text generation data, and intention recognition data.
[0057] In some other embodiments, the processing layer of the large-scale model may be composed of two parts: a gating network and an expert module. The expert module includes the first expert module Expert1-1, Expert2-1, ···, Expertn-1, and the second expert module Expert1-2, Expert2-2, ···, Expertn-2. The i-th expert module may include Experti-1, Experti-2, ···, Experti-n. The gating network dynamically determines which expert module should be activated based on the characteristics of the features to be processed, and generates an optimal prediction. Based on the result output by the gating network, a target expert module is determined from a plurality of expert modules. Thereby, the target calculation unit processes the features to be processed based on the model parameters of the target expert module to obtain a task execution result.
[0058] Compared with the task execution method using a single gating network, according to the task execution method according to the embodiments of the present disclosure, by executing the modal route task, it is possible to complete the modal classification for the target processing features corresponding to the target field gating model parameters. By executing the field route task, the field classification of the target processing features corresponding to the target feed-forward task model parameters is completed. Thereby, through two-stage type recognition, by making the target feed-forward task model parameters obtained from a plurality of feed-forward task model parameters accurate and effective, the inference accuracy and inference efficiency of the task execution method used in the large-scale model are improved.
[0059] As described above, the network structure of the processing layer of the large-scale model has been described. Hereinafter, the method for obtaining the model parameters of the processing layer will be further described.
[0060] In the embodiments of the present disclosure, the task execution method used in the large-scale model may be applied to any service node of the server, but is not limited thereto, and may also be applied to the terminal device.
[0061] In the embodiments of the present disclosure, the task execution method used in the large-scale model may further include an operation of obtaining model parameters.
[0062] For example, a model acquisition request is sent to the server through an interface. In response to receiving a plurality of feed-forward task model parameters and a plurality of field gating model parameters corresponding to the model acquisition request, the plurality of feed-forward task model parameters and the plurality of field gating model parameters are stored in a target storage unit.
[0063] In an embodiment of the present disclosure, the terminal device can send a model acquisition request to the server via an interface. After receiving the model acquisition request sent from the terminal device, the server can send, via the interface, model parameters of a large-scale model, such as modal gating model parameters, a plurality of feed-forward task model parameters, and a plurality of field gating model parameters, to the terminal device. In response to receiving the model parameters, the terminal device stores the modal gating model parameters, the plurality of feed-forward task model parameters, and the plurality of field gating model parameters in a target storage unit. During the process of executing a task by the target calculation unit, corresponding model parameters can be acquired from the target storage unit.
[0064] In an embodiment of the present disclosure, the model acquisition request may include a model performance requirement. The model performance requirement may include at least one of model accuracy, model latency, energy consumption of the target calculation unit, and the number of model parameters, but is not limited thereto. It may further include the type of the field of the field expert sub-module and / or the type of the modality of the field gate.
[0065] In an embodiment of the present disclosure, the plurality of feed-forward task model parameters corresponding to the model acquisition request may include feed-forward task model parameters that satisfy the model performance requirement and a plurality of field gating model parameters.
[0066] According to the task execution method for a large-scale model according to an embodiment of the present disclosure, a model acquisition request can be generated for the hardware of the terminal device, such as the resource allocation of the target calculation unit, user needs, etc. Based on the model acquisition request, model parameters that match the hardware resources or user needs of the terminal device are acquired and successfully applied to the terminal device, thereby improving the flexibility and adaptability of the task execution method for the large-scale model and improving the satisfaction of personalized needs.
[0067] The method for obtaining model parameters has been described above. Next, the acquisition and optimization of feed-forward task model parameters will be further described.
[0068] According to an embodiment of the present disclosure, a plurality of feed-forward task model parameters are obtained as follows.
[0069] For example, based on a sample set and a pre-trained large-scale model, a plurality of target compressed large-scale models are obtained. Based on the plurality of target compressed large-scale models, a plurality of feed-forward task model parameters are obtained.
[0070] In an embodiment of the present disclosure, the pre-trained large-scale model may include a large-scale model including a pre-trained Transformer structure. The pre-trained large-scale model may include model parameters to be compressed that have the same function as the feed-forward task model parameters. The model parameters to be compressed may include, for example, model parameters for executing a feed-forward task, such as model parameters of an FFN. The model parameters to be compressed have a larger number of parameters compared to the feed-forward task model parameters.
[0071] In an embodiment of the present disclosure, the sample set may include a plurality of sample subsets with different field types. For example, the sample set may include a text-type question-and-answer sample subset, a text-type search sample subset, a text-type translation sample subset, an audio-type question-and-answer sample subset, an audio-type search sample subset, an audio-type translation sample subset, an image-type question-and-answer sample subset, an image-type search sample subset, and an image-type translation sample subset.
[0072] In an embodiment of the present disclosure, based on a plurality of sample subsets and the same pre-trained large-scale model, a computing unit executes a compression task to obtain a plurality of target compressed large-scale models that correspond one-to-one to the plurality of sample subsets.
[0073] In an embodiment of the present disclosure, the compression task may refer to performing compression, such as truncation or lightweight processing, on predetermined parameters in the pre-trained large-scale model, such as model parameters to be compressed, by a sample set, inheriting the knowledge and capabilities of the pre-trained large-scale model, and obtaining a target compressed large-scale model whose model parameters are smaller than those of the pre-trained large-scale model.
[0074] In an embodiment of the present disclosure, the plurality of target compressed large-scale models each maintain the same backbone network as the pre-trained large-scale model, and only the model parameters of the FFN for performing the feed-forward task are different. For example, the model parameters of the FFN for performing the feed-forward task in the target compressed large-scale model correspond to the corresponding sample subset.
[0075] In an embodiment of the present disclosure, the respective FFN model parameters of the plurality of target compressed large-scale models may be used as a plurality of feed-forward task model parameters.
[0076] According to the compression method of the target compressed large-scale model according to the embodiment of the present disclosure, while the number of parameters of the model parameters of the compressed target compressed large-scale model is reduced, the model parameters to be compressed are compressed so as to inherit the knowledge and inference ability of the model parameters to be compressed well. In addition, due to a plurality of sample subsets with different field types, the target compressed large-scale model inherits the inference ability and knowledge of different field types in different modal types of the pre-trained large-scale model.
[0077] In an embodiment of the present disclosure, obtaining a plurality of target compressed large-scale models based on a sample set and a pre-trained large-scale model can include the following operations.
[0078] For example, for each sample subset, a compressed large-scale model is obtained based on a pruning matrix and a pre-trained large-scale model. A target compressed large-scale model is obtained based on the compressed sample subset, the pre-trained large-scale model, and the compressed large-scale model.
[0079] In an embodiment of the present disclosure, the pruning matrix may be used to indicate a pruning method for model parameters to be compressed. The pruning matrix Z may include a plurality of pruning elements represented by 0 or 1. Each pruning element corresponds to one model parameter of the model parameters to be compressed. 0 represents that the model parameter is pruned, and 1 represents that the model parameter is retained.
[0080] In an embodiment of the present disclosure, a vector multiplication can be performed using the pruning matrix and the model parameters of the pre-trained large-scale model to obtain a compressed large-scale model. A target compressed large-scale model is obtained based on the sample subset, the pre-trained large-scale model, and the compressed large-scale model.
[0081] Obtaining a target compressed large-scale model based on the sample subset, the pre-trained large-scale model, and the compressed large-scale model can include executing a cyclic compression task by a computing unit based on the sample subset, the pre-trained large-scale model, and the compressed large-scale model to obtain the target compressed large-scale model.
[0082] In an embodiment of the present disclosure, the cyclic compression task may include setting the compressed large-scale model as the target compressed large-scale model when determining that the model performance of the pre-trained large-scale model matches the model performance of the compressed large-scale model based on a sample subset. If it is determined that the model performance of the pre-trained large-scale model does not match the model performance of the compressed large-scale model based on the sample subset, update the pruning matrix and obtain a further updated compressed large-scale model. Using the updated compressed large-scale model, the sample subset, and the pre-trained large-scale model, repeatedly execute the cyclic compression task until the model performance of the pre-trained large-scale model matches the model performance of the compressed large-scale model.
[0083] By compressing the pre-trained large-scale model with the pruning matrix according to the embodiment of the present disclosure, the entire compression process can be analyzed, and the compression efficiency can be improved. Further, by respectively compressing the pre-trained large-scale model with multiple sample subsets of different field types, the type of the obtained target compressed large-scale model can be improved, thereby further improving the compression efficiency.
[0084] In an embodiment of the present disclosure, obtaining the target compressed large-scale model based on the sample subset, the pre-trained large-scale model, and the compressed large-scale model may include an operation of determining the reference inference ability of the pre-trained large-scale model and the verification inference ability of the compressed large-scale model based on the sample subset. When the verification inference ability matches the reference inference ability, determine the target compressed large-scale model based on the compressed large-scale model.
[0085] For example, by using a pre-training large-scale model and a compressed large-scale model, sample data subsets in a sample subset can be processed respectively to obtain a first prediction result corresponding to the pre-training large-scale model and a second prediction result corresponding to the compressed large-scale model. Based on the first prediction result and a tag subset corresponding to the sample data subset, a reference inference ability can be obtained. Based on the second prediction result and a tag subset corresponding to the sample data subset, a verification inference ability can be obtained. When the similarity between the verification inference ability and the reference inference ability is greater than a threshold, it is determined that the verification inference ability matches the reference inference ability.
[0086] When the similarity between the verification inference ability and the reference inference ability is less than or equal to the threshold, it is determined that the verification inference ability does not match the reference inference ability. Based on the reference inference ability and the verification inference ability, the model parameters of the pre-training large-scale model and the threshold elements of the threshold matrix can be adjusted to obtain an updated pre-training large-scale model and an updated threshold matrix.
[0087] Until the reference inference ability and the verification inference ability match, based on the sample subset, the updated pre-training large-scale model, and the threshold elements of the updated threshold matrix, the compression task is re-executed. The compressed large-scale model in which the reference inference ability and the verification inference ability match is used as the target compressed large-scale model.
[0088] In other embodiments of the present disclosure, a predetermined sparsity may be further set. The predetermined sparsity is used to indicate a predetermined threshold granularity. The predetermined sparsity may be determined based on the information carried in the model acquisition request.
[0089] Based on the reference inference ability and the verification inference ability, adjusting the model parameters of the pre-trained large-scale model and the decision-making elements of the decision matrix, and obtaining the updated model parameters of the pre-trained large-scale model and the decision-making elements of the updated decision matrix may include determining the sparsity degree based on the decision matrix. Based on the sparsity degree and a predetermined sparsity degree, a sparsity loss value is determined. Based on the reference inference ability, the verification inference ability, and the sparsity loss value, the model parameters of the pre-trained large-scale model and the decision-making elements of the decision matrix are adjusted, and the updated model parameters of the pre-trained large-scale model and the decision-making elements of the updated decision matrix are obtained.
[0090] In an embodiment of the present disclosure, the sparsity loss value can be determined by a loss function based on the sparsity degree and a predetermined sparsity degree. By the loss function, the reference inference ability is obtained based on the first prediction result and the tag subset corresponding to the sample data subset. By the loss function, the verification inference ability is obtained based on the second prediction result and the tag subset corresponding to the sample data subset. The loss function can utilize a cross-entropy loss function, but is not limited thereto, and a function that can represent the matching degree between the other two data to be measured may also be used.
[0091] Based on the sample subset, the updated pre-trained large-scale model, and the decision-making elements of the updated decision matrix, the compression task is re-executed until the reference inference ability and the verification inference ability match and the sparsity degree of the compressed large-scale model meets the predetermined sparsity degree. The compressed large-scale model in which the reference inference ability and the verification inference ability match and the sparsity degree of the compressed large-scale model meets the predetermined sparsity degree is used as the target compressed large-scale model.
[0092] In an embodiment of the present disclosure, the compressed model parameters corresponding to the model parameters to be compressed in the target compressed large-scale model can be used as the feed-forward task model parameters.
[0093] By obtaining a target compressed large-scale model using the compression method according to the embodiments of the present disclosure, the model parameters of the target compressed large-scale model can be the model parameters after trimming and weight reduction processing with respect to the model parameters of the pre-trained large-scale model. Thereby, the number of parameters of the model parameters of the target compressed large-scale model can be reduced. Further, a sample subset is used as a compressed sample, the pre-trained large-scale model is compressed, and the pre-trained large-scale model is trained. The model inference ability of the target compressed large-scale model with respect to the data of the type of the field in the type of the modality of the pre-trained large-scale model updated is inherited, and further the model inference ability of the target compressed large-scale model is improved.
[0094] In an embodiment of the present disclosure, the plurality of feed-forward model parameters obtained by the above method can be directly used as the feed-forward model parameters of the large-scale model.
[0095] In another embodiment of the present disclosure, the plurality of feed-forward model parameters obtained by the above method may be used as a plurality of initial feed-forward model parameters. Based on the pre-trained large-scale model, a plurality of initial domain gating model parameters, a plurality of initial feed-forward task model parameters, and initial modal gating model parameters, an initial large-scale model is obtained. The initial modal gating model parameters are the same as the model parameter function for executing the modal root task, each initial domain gating model parameter is the same as the function of the domain gating model parameter, and each initial feed-forward task model parameter is the same as the function of the feed-forward task model parameter.
[0096] For example, obtaining an initial large-scale model based on a pre-trained large-scale model, a plurality of initial domain gating model parameters, a plurality of initial feed-forward task model parameters, and an initial modal gating model parameter may include replacing the model parameters to be compressed in the pre-trained large-scale model with the plurality of initial domain gating model parameters, the plurality of initial feed-forward task model parameters, and the initial modal gating model parameter, and then obtaining the initial large-scale model. For example, the plurality of initial domain gating model parameters, the plurality of initial feed-forward task model parameters, and the initial modal gating model parameter are used as a mixture-of-experts architecture, and the FFN layer corresponding to the model parameters to be compressed in the pre-trained large-scale model is replaced to obtain the initial large-scale model. The mixture-of-experts architecture includes a plurality of hierarchical networks, and the initial feed-forward task model parameters can be accurately positioned by the multi-stage router in the plurality of hierarchical networks. The initial feed-forward task model parameters can guarantee the inference performance and improve the inference efficiency by the number of small-scale model parameters.
[0097] The initial large-scale model can be optimized and trained to obtain a large-scale model. Thereby, the optimization training can further improve the model combination adaptability of the mixture-of-experts architecture and further improve the model inference performance of the large-scale model.
[0098] The following further describes optimizing and training the initial large-scale model to obtain a large-scale model.
[0099] According to an embodiment of the present disclosure, optimizing and training an initial large-scale model to obtain a large-scale model may include obtaining a model output result set based on a sample data set of a sample set and the initial large-scale model. Based on the model output result set and a tag set, a plurality of loss values are obtained. Based on the plurality of loss values and the initial large-scale model, a large-scale model is obtained. A model parameter having the same function as the feed-forward task model parameter in the large-scale model is used as the feed-forward task model parameter.
[0100] The sample set may further include a tag set matching the sample data set. The sample data set includes a plurality of sample data subsets with different field types. The initial large-scale model includes a model parameter having the same function as the feed-forward task model parameter.
[0101] In an embodiment of the present disclosure, the sample data set of the sample set can be input into the initial large-scale model to obtain a model output result set. The model output result set and the tag set are input into a loss function to obtain a plurality of loss values. Based on the plurality of loss values and the initial large-scale model, a large-scale model is obtained. The plurality of loss values correspond one-to-one with the plurality of sample data subsets.
[0102] The type of the loss function is not limited. It may include a cross-entropy loss function. Any loss function that can identify the loss value is acceptable.
[0103] According to an embodiment of the present disclosure, obtaining a feed-forward task model parameter based on a plurality of loss values and an initial large-scale model may include obtaining a target loss value based on the plurality of loss values. Based on the target loss value and the initial large-scale model, a feed-forward task model parameter is obtained.
[0104] Obtaining a target loss value based on a plurality of loss values may include obtaining the target loss value by weighted addition of the plurality of loss values. When it is determined that the target loss value does not converge, the initial large-scale model is adjusted. When it is determined that the target loss value converges, the initial large-scale model in which the target loss value converges is used as the large-scale model. The model parameters corresponding to the initial feed-forward task model parameters in the large-scale model are used as the feed-forward task model parameters.
[0105] By the above-described hybrid training method, hybrid fine-tuning can be performed on the entire hybrid expert architecture using a sample set, improving the overall adaptability. Furthermore, by improving the inference efficiency and inference performance of the large-scale model feed-forward task model parameters and the classification performance of the modal route task parameters and the domain route task parameters, the overall performance of the large-scale model is improved.
[0106] FIG. 4A schematically shows a schematic diagram for determining a large-scale model according to an embodiment of the present disclosure.
[0107] As shown in FIG. 4A, the pre-trained large-scale model includes a plurality of processing layers, and each processing layer includes a multi-head attention module (MHA), an addition normalization module (Add&Norm), and a feed-forward module (FFN). By compressing the model parameters of the FFN of the pre-trained large-scale model using a sample set, a plurality of target compressed large-scale models can be obtained. Based on the plurality of target compressed large-scale models, a plurality of initial feed-forward task model parameters FFN1, …, FFNX are obtained.
[0108] As shown in FIG. 4A, by a plurality of initial domain gating model parameters Router1, …, RouterX, a plurality of initial feed-forward task model parameters FFN1, …, FFNX, and an initial modal gating model parameter Router0, the model parameters of the FFN in the pre-trained large-scale model are replaced to obtain an initial large-scale model M410.
[0109] As shown in FIG. 4A, an initial large-scale model 410 is optimized and trained using a sample set to obtain a large-scale model M420.
[0110] FIG. 4B schematically shows a schematic diagram for determining a large-scale model according to another embodiment of the present disclosure.
[0111] As shown in FIG. 4B, the task execution method is similar to the task execution method shown in FIG. 4A, and the only difference is as follows. One sample subset in the sample set compresses the model parameters of the FFN of the pre-trained large-scale model to obtain a lightweight large-scale model that matches the sample subset. The FFN parameters in the lightweight large-scale model are replicated to obtain a plurality of initial feed-forward task parameters FFN'.
[0112] A plurality of initial field gating model parameters, a plurality of initial feed-forward task model parameters, and initial modal gating model parameters are replaced with the model parameters of the FFN in the pre-trained large-scale model to obtain an initial large-scale model M410'. The initial large-scale model M410' is trained using a sample set to obtain a large-scale model M420'.
[0113] Compared with the method for determining a large-scale model as shown in FIG. 4B, the method for determining a large-scale model using FIG. 4A enables different types of sample subsets in the sample set to better inherit the knowledge of the pre-trained large-scale model before truncation, and have natural differentiation among the plurality of initial feed-forward task model parameters, which is advantageous for improving the training efficiency and training accuracy of mixed training.
[0114] FIG. 5 schematically shows a block diagram of a task execution apparatus used for a large-scale model according to an embodiment of the present disclosure.
[0115] As shown in FIG. 5, the task execution device 500 used for the large-scale model includes a target memory unit 510 and a target calculation unit 520.
[0116] The target memory unit 510 stores a plurality of domain gating model parameters and a plurality of feed-forward task model parameters.
[0117] The target calculation unit 520 is configured as follows.
[0118] Based on the features to be targeted, the target calculation unit executes the modal route task to obtain a modal recognition result.
[0119] Based on the features to be targeted and the target domain gating model parameters, the target calculation unit executes the domain route task to obtain a domain recognition result. Here, the target domain gating model parameters are the domain gating model parameters corresponding to the modal recognition result read from the target memory unit.
[0120] Based on the features to be targeted and the target feed-forward task model parameters, the target calculation unit executes the feed-forward task to obtain a task execution result. Here, the target feed-forward task model parameters are the feed-forward task model parameters corresponding to the domain recognition result read from the target memory unit.
[0121] According to an embodiment of the present disclosure, the task execution device 500 used for the large-scale model further includes an actuator.
[0122] The actuator is configured as follows.
[0123] Based on the modal recognition result, obtain the target field gating model parameter from a plurality of field gating model parameters stored in the target memory unit, where the types of the corresponding modalities of the plurality of field gating model parameters are different.
[0124] According to an embodiment of the present disclosure, the actuator is further configured as follows.
[0125] Based on the field recognition result, obtain the target feedforward task model parameter from a plurality of feedforward task model parameters stored in the target memory unit, where the types of the corresponding fields of the plurality of feedforward task model parameters are different.
[0126] According to an embodiment of the present disclosure, the plurality of feedforward task model parameters are obtained as follows.
[0127] According to an embodiment of the present disclosure, the target calculation unit is further configured as follows.
[0128] Based on the sample set and the pre-trained large-scale model, obtain a plurality of target compressed large-scale models, where the sample set includes a plurality of sample subsets with different region types, and the pre-trained large-scale model includes model parameters to be compressed with the same function as the feedforward task model parameters.
[0129] Obtain a plurality of feedforward task model parameters based on the plurality of target compressed large-scale models.
[0130] According to an embodiment of the present disclosure, obtaining a plurality of target compressed large-scale models based on the sample set and the pre-trained large-scale model includes the following.
[0131] For each sample subset, obtain a compressed large-scale model based on a pruning matrix and a pre-trained large-scale model, where the pruning matrix is used to indicate a pruning method for model parameters to be compressed.
[0132] Obtain a target compressed large-scale model based on the sample subset, the pre-trained large-scale model, and the compressed large-scale model.
[0133] According to an embodiment of the present disclosure, obtaining a target compressed large-scale model based on a sample subset, a pre-trained large-scale model, and a compressed large-scale model includes the following.
[0134] Based on the sample subset, determine the reference inference ability of the pre-trained large-scale model and the verification inference ability of the compressed large-scale model.
[0135] When the verification inference ability matches the reference inference ability and the sparsity of the compressed large-scale model meets a predetermined sparsity, determine a target compressed large-scale model based on the compressed large-scale model.
[0136] According to an embodiment of the present disclosure, the feed-forward task model parameters are obtained as follows.
[0137] The target computing unit is further configured as follows.
[0138] Based on the sample dataset of the sample set and the initial large-scale model, obtain a model output result set, where the sample set further includes a tag set matching the sample dataset, the sample dataset includes a plurality of sample data subsets with different field types, and the initial large-scale model includes model parameters having the same function as the feed-forward task model parameters.
[0139] Based on the model output result set and the tag set, obtain a plurality of loss values, where the plurality of loss values correspond one-to-one to a plurality of sample data subsets.
[0140] Based on the plurality of loss values and the initial large-scale model, obtain the feed-forward task model parameters.
[0141] According to an embodiment of the present disclosure, obtaining the feed-forward task model parameters based on the plurality of loss values and the initial large-scale model includes the following.
[0142] Based on the plurality of loss values, obtain a target loss value.
[0143] Based on the target loss value and the initial large-scale model, obtain the feed-forward task model parameters.
[0144] According to an embodiment of the present disclosure, the initial large-scale model is obtained as follows.
[0145] The target calculation unit is further configured as follows.
[0146] Based on the pre-trained large-scale model, a plurality of initial domain gating model parameters, a plurality of initial feed-forward task model parameters, and the initial modal gating model parameters, obtain the initial large-scale model, where the initial modal gating model parameters are the same as the model parameter function for executing the modal root task, each initial domain gating model parameter is the same as the function of the domain gating model parameter, and each initial feed-forward task model parameter is the same as the function of the feed-forward task model parameter.
[0147] According to an embodiment of the present disclosure, the type of the modality includes at least one of image, text, and audio.
[0148] The type of the field includes at least one of translation, question answering, search, text generation, and intent recognition.
[0149] According to an embodiment of the present disclosure, the target computing unit is further configured as follows.
[0150] Using an interface, send a model acquisition request to the server.
[0151] In response to receiving a plurality of feed-forward task model parameters and a plurality of field gating model parameters corresponding to the model acquisition request, store the plurality of feed-forward task model parameters and the plurality of field gating model parameters in the target storage unit.
[0152] According to an embodiment of the present disclosure, the model acquisition request includes a model performance requirement.
[0153] The plurality of feed-forward task model parameters corresponding to the model acquisition request include feed-forward task model parameters that satisfy the model performance requirement.
[0154] According to an embodiment of the present disclosure, the model performance requirement includes at least one of the following.
[0155] Model accuracy, model latency, energy consumption of the target computing unit, number of model parameters.
[0156] According to an embodiment of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.
[0157] According to an embodiment of the present disclosure, the electronic device includes at least one processor and a memory communicatively connected to the at least one processor. The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method as described above.
[0158] According to an embodiment of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, which are used to cause a computer to execute the above method.
[0159] According to an embodiment of the present disclosure, when a computer program is executed by a processor, the above method is implemented.
[0160] FIG. 6 shows an exemplary block diagram for implementing an exemplary electronic device 600 according to an embodiment of the present disclosure. The electronic device is intended to represent various forms of digital computers, such as, for example, a laptop computer, a desktop computer, a workbench, a personal digital assistant, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device may further represent various forms of mobile devices, such as, for example, a personal digital assistant, a cellular phone, a smartphone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are exemplary only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0161] As shown in FIG. 6, the device 600 includes a computing unit 601, which can execute various appropriate operations and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 can further store various programs and data necessary for the operation of the device 600. The computing unit 601, the ROM 602, and the RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0162] The multiple components in the device 600 are connected to the input / output (I / O) interface 605 and include an input unit 606 such as a keyboard, a mouse, etc., an output unit 607 such as various types of displays, speakers, etc., a storage unit 608 such as a magnetic disk, an optical disk, etc., and a communication unit 609 such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 609 enables the device 600 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0163] The computing unit 601 may be various general-purpose and / or dedicated processing modules having processing and computing capabilities. Some examples of the computing unit 601 include a central processing unit (CPU), a GPU (Graphics Processing Unit), various dedicated artificial intelligence (AI) computing chips, computing units of various machine learning model algorithms, a DSP (Digital Signal Processor), and any suitable processor, controller, microcontroller, etc., but are not limited thereto. The computing unit 601 executes the processing with each of the methods described above, such as a task execution method used for large-scale models. For example, in some embodiments, the task execution method used for large-scale models may be implemented as a computer software program tangibly included in a machine-readable medium such as the storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed into the device 600 via the ROM 602 and / or the communication unit 609. When the computer program is loaded into the RAM 603 and executed by the computing unit 601, one or more steps of the task execution method used for the large-scale models described above may be executed. Alternatively, in another embodiment, the computing unit 601 may be configured to execute the task execution method used for large-scale models in any other suitable form (e.g., via firmware).
[0164] The various embodiments of the systems and techniques described in this specification may be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may be implemented in one or more computer programs, which may be executed and / or interpreted in a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, and which receives data and instructions from, and transmits data and instructions to, a memory system, at least one input device, and at least one output device.
[0165] The program code for implementing the methods of the present disclosure may be created in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general purpose computer, a dedicated computer, or other programmable data processing apparatus, such that, when the program code is executed by the processor or controller, the functions / operations specified in the flowchart and / or block diagram are implemented. The program code may be executed entirely on the device, may be executed partly on the device, may be executed partly on the device as an independent software package, and may be executed partly on a remote device or entirely on a remote device or server.
[0166] In the context of the present disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use in or in combination with an instruction execution system, apparatus, or device. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium include electrical connections made with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0167] To provide for interaction with a user, a computer may implement the systems and techniques described herein, the computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or trackball), whereby the user can provide input to the computer via the keyboard and the pointing device. Other kinds of devices may further provide for interaction with the user; for example, feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input received from the user may be in any form (including voice input, speech input, or tactile input).
[0168] The systems and techniques described herein can be implemented in a computing system that includes background components (e.g., a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer having a graphical user interface or a web browser, where the user can interact with embodiments of the systems and techniques described herein via the graphical user interface or the network browser), or a computing system that includes any combination of such background components, middleware components, or front-end components. The components of the system can be connected to each other by digital data communication in any form or medium (e.g., a communication network). Exemplary communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0169] The computer system may include clients and servers. The clients and servers are generally remote from each other and typically interact via a communication network. The relationship between the client and the server is generated by a computer program running on the corresponding computer and having a client-server relationship. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0170] It should be understood that various forms of the flows shown above may be used, and the steps may be sorted, added, or deleted again. For example, each step described in the present invention may be executed in parallel, sequentially, or in a different order, and the present specification is not limited herein as long as the desired results of the technical solutions of the present disclosure can be achieved.
[0171] The foregoing specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations and alternatives can be made according to design requirements and other factors. Any modifications, equivalent substitutions and improvements made within the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.
Claims
1. 1. A task execution method for use with large scale models, comprising: Executing a modal root task by a target computing unit according to the target features to be processed to obtain a modal recognition result; According to the target features to be processed and the target field gating model parameters, a field route task is executed by the target calculation unit to obtain a field recognition result, and the target field gating model parameters are field gating model parameters corresponding to the modal recognition result read from the target storage unit; According to the target features to be processed and target feedforward task model parameters, a feedforward task is executed by the target calculation unit to obtain a task execution result, where the target feedforward task model parameters are feedforward task model parameters corresponding to the domain recognition result read from the target storage unit. How to perform the task.
2. The method further includes obtaining the target field gating model parameters from a plurality of field gating model parameters stored in the target storage unit according to the modal recognition result, where the corresponding modal types of the plurality of field gating model parameters are different. The method of claim 1.
3. The method further includes obtaining the target feedforward task model parameter from a plurality of feedforward task model parameters stored in the target storage unit according to the field recognition result, where the corresponding field types of the plurality of feedforward task model parameters are different. The method according to claim 1 or 2.
4. The plurality of feedforward task model parameters are obtained as follows: Obtain a plurality of target compressed large-scale models based on a sample set and a pre-trained large-scale model, the sample set including a plurality of sample subsets with different domain types, and the pre-trained large-scale model including model parameters to be compressed that are the same as a function of the feedforward task model parameters; obtaining the plurality of feedforward task model parameters based on the plurality of target compressed large-scale models; The method according to claim 1 or 2.
5. Based on the sample set and the pre-trained large-scale model, obtaining a multi-objective compressed large-scale model is for each of the sample subsets, a truncation matrix for indicating a truncation manner of the model parameters to be compressed, and obtaining a compressed large model based on the pre-training large model; obtaining the target compressed large scale model based on the sample subset, the pre-training large scale model and the compressed large scale model. The method according to claim 4.
6. Obtaining the target compressed large scale model based on the sample subset, the pre-training large scale model and the compressed large scale model includes: determining a reference inference capability of the pre-trained large-scale model and a validation inference capability of the condensed large-scale model based on the sample subset; If the verification inference ability matches the reference inference ability and the sparsity of the compressed large-scale model satisfies a predetermined sparsity, determining the target compressed large-scale model based on the compressed large-scale model. The method according to claim 5.
7. The feedforward task model parameters are obtained as follows: Obtain a model output result set based on a sample dataset of a sample set and an initial large-scale model, where the sample set further includes a tag set matching the sample dataset, the sample dataset includes a plurality of sample data subsets with different domain types, and the initial large-scale model includes model parameters that are a function of the feedforward task model parameters; obtaining a plurality of loss values based on the model output result set and the tag set, the plurality of loss values corresponding one-to-one to the plurality of sample data subsets; Obtaining the feedforward task model parameters based on the plurality of loss values and the initial large-scale model. The method according to claim 1 or 2.
8. Obtaining the feedforward task model parameters based on the plurality of loss values and the initial large-scale model includes: obtaining a target loss value based on the plurality of loss values; and obtaining the feedforward task model parameters based on the target loss value and the initial large-scale model. The method of claim 7.
9. The initial large-scale model is obtained in the following manner: Obtain the initial large-scale model based on a pre-training large-scale model, a plurality of initial field gating model parameters, a plurality of initial feedforward task model parameters, and an initial modal gating model parameter, where the initial modal gating model parameter is equal to a model parameter function for performing the modal root task, each of the initial field gating model parameters is equal to a function of the field gating model parameter, and each of the initial feedforward task model parameters is equal to a function of the feedforward task model parameter. The method of claim 7.
10. the type of modal includes at least one of image, text, and audio; The domain type includes at least one of translation, question answering, search, text generation, and intent recognition. The method according to claim 1 or 2.
11. sending a model retrieval request to a server via the interface; in response to receiving a plurality of feedforward task model parameters and a plurality of field gating model parameters corresponding to the model acquisition request, storing the plurality of feedforward task model parameters and the plurality of field gating model parameters in the target storage unit. The method according to claim 1 or 2.
12. the model acquisition request includes a model performance request; the plurality of feedforward task model parameters corresponding to the model acquisition request include feedforward task model parameters that satisfy the model performance requirement; Wherein, the model performance requirements include at least one of model accuracy, model delay, energy consumption of the target computing unit, and number of model parameters. The method of claim 11.
13. A task execution device for use in a large scale model, comprising: a target storage unit for storing a plurality of field gating model parameters and a plurality of feedforward task model parameters; a target computing unit; The target computing unit is configured as follows: According to the features to be processed, a modal root task is executed by a target calculation unit to obtain a modal recognition result; According to the target features to be processed and the target field gating model parameters, execute a field route task by the target calculation unit to obtain a field recognition result, and the target field gating model parameters are field gating model parameters corresponding to the modal recognition result read from the target storage unit; According to the target features to be processed and the target feedforward task model parameters, a feedforward task is executed by the target calculation unit to obtain a task execution result, and the target feedforward task model parameters are feedforward task model parameters corresponding to the domain recognition result read from the target storage unit. Task executor.
14. Further comprising an actuator configured to obtain the target field gating model parameters from a plurality of field gating model parameters stored in the target storage unit according to the modal recognition result; The types of modals corresponding to the plurality of field gating model parameters are different.
14. The apparatus of claim 13.
15. The actuator further comprises: The target feedforward task model parameter is obtained from a plurality of feedforward task model parameters stored in the target storage unit according to the field recognition result, where the corresponding field types of the plurality of feedforward task model parameters are different.
15. Apparatus according to claim 13 or 14.
16. The plurality of feedforward task model parameters are obtained as follows: The target computing unit further comprises: Obtain a plurality of target compressed large-scale models based on a sample set and a pre-trained large-scale model, the sample set including a plurality of sample subsets with different domain types, and the pre-trained large-scale model including model parameters to be compressed that are the same as a function of the feedforward task model parameters; configured to derive the plurality of feedforward task model parameters based on the plurality of target compressed large-scale models.
15. Apparatus according to claim 13 or 14.
17. Based on the sample set and the pre-trained large-scale model, obtaining a multi-objective compressed large-scale model is for each of the sample subsets, a truncation matrix for indicating a truncation manner of the model parameters to be compressed, and obtaining a compressed large model based on the pre-training large model; obtaining the target compressed large scale model based on the sample subset, the pre-training large scale model and the compressed large scale model.
17. The apparatus of claim 16.
18. Obtaining the target compressed large scale model based on the sample subset, the pre-training large scale model and the compressed large scale model includes: determining a reference inference capability of the pre-trained large-scale model and a validation inference capability of the condensed large-scale model based on the sample subset; If the verification inference ability matches the reference inference ability and the sparsity of the compressed large-scale model satisfies a predetermined sparsity, determining the target compressed large-scale model based on the compressed large-scale model.
20. The apparatus of claim 17.
19. The feedforward task model parameters are obtained as follows: The target computing unit further comprises: Obtain a model output result set based on a sample dataset of a sample set and an initial large-scale model, where the sample set further includes a tag set matching the sample dataset, the sample dataset includes a plurality of sample data subsets with different domain types, and the initial large-scale model includes model parameters that are a function of the feedforward task model parameters; obtaining a plurality of loss values based on the model output result set and the tag set, the plurality of loss values corresponding one-to-one to the plurality of sample data subsets; configured to derive the feedforward task model parameters based on the plurality of loss values and the initial large-scale model.
15. Apparatus according to claim 13 or 14.
20. Obtaining the feedforward task model parameters based on the plurality of loss values and the initial large-scale model includes: obtaining a target loss value based on the plurality of loss values; and obtaining the feedforward task model parameters based on the target loss value and the initial large-scale model.
20. The apparatus of claim 19.
21. The initial large-scale model is obtained in the following manner: The target computing unit further comprises: Obtain the initial large-scale model based on a pre-training large-scale model, a plurality of initial discipline gating model parameters, a plurality of initial feedforward task model parameters, and an initial modal gating model parameter, where the initial modal gating model parameters are the same as a function of model parameters for performing the modal root task, each of the initial discipline gating model parameters is the same as a function of the discipline gating model parameters, and each of the initial feedforward task model parameters is the same as a function of the feedforward task model parameters.
20. The apparatus of claim 19.
22. the type of modal includes at least one of image, text, and audio; The domain type includes at least one of translation, question answering, search, text generation, and intent recognition.
15. Apparatus according to claim 13 or 14.
23. The target computing unit further comprises: Sending a model retrieval request to the server via the interface; configured to store the plurality of feedforward task model parameters and the plurality of discipline gating model parameters in the target storage unit in response to receiving a plurality of feedforward task model parameters and a plurality of discipline gating model parameters corresponding to the model acquisition request.
15. Apparatus according to claim 13 or 14.
24. the model acquisition request includes a model performance request; the plurality of feedforward task model parameters corresponding to the model acquisition request include feedforward task model parameters that satisfy the model performance requirement; Wherein, the model performance requirements include at least one of model accuracy, model delay, energy consumption of the target computing unit, and number of model parameters.
24. The apparatus of claim 23.
25. At least one processor; a memory in communication with the at least one processor; The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor such that the at least one processor can perform the method of claim 1 or 2. electronic equipment.
26. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: The computer instructions cause the computer to carry out the method according to claim 1 or 2. A non-transitory computer-readable storage medium.
27. A computer program which, when executed by a processor, implements the method according to claim 1 or 2.
Citation Information
Patent Citations
Content recommendation method, device and equipment and readable storage medium
CN113569130A
Data processing method, device and equipment
CN116010545A
Task scheduling method and device, computer equipment, storage medium and program product
CN117742970A
Mixture of Expert Neural Networks
JP2019537133A
Interactive system
JP2023146677A