A task processing method, system, terminal and medium based on multiple expert layers
By building a multi-layer expert network and selecting the appropriate expert layer combination according to task characteristics for processing, the existing models lack flexibility and targeting in different fields and task types is solved, and more efficient and accurate task processing is achieved.
Patent Information
- Application Number
- CN202510013690.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The existing pre-trained language model based on Transformer architecture lacks flexibility and targeting when facing different fields and task types, and cannot effectively adjust the model structure dynamically, resulting in insufficient processing efficiency and accuracy.
Build a multi-layer expert network, including the basic expert layer, the domain expert layer and the task expert layer, and select the corresponding expert layer combination based on the task characteristics (domain, type and complexity) of the input task to process to improve the adaptability and pertinence of the model.
By flexibly adjusting the expert layer combination, we can improve the model's adaptability and accuracy and reliability of output results when handling diverse tasks, and improve task processing efficiency and performance.
Smart Images

Figure CN119416822B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of natural language processing, and particularly relates to a task processing method, system, terminal and medium based on multiple expert layers. Background Art
[0002] Currently, natural language processing (NLP) technology has been widely applied in multiple fields. Among them, pre-trained language models based on the Transformer architecture have become a research hotspot due to their excellent performance. These models can learn rich language representations by pre-training on large-scale text data. However, when facing specific domains or tasks, these models require a large amount of fine-tuning data and computing resources to achieve ideal results.
[0003] Current large model systems can often only process single-domain and single-task types, and cannot effectively adjust the model structure according to task characteristics or domain requirements, resulting in a lack of flexibility and pertinence of the model when dealing with different tasks. Summary of the Invention
[0004] To solve the above problems, the present invention provides a task processing method, system, terminal and medium based on multiple expert layers. By constructing a multi-layer expert network and selecting the corresponding expert layer according to the task characteristics, the adaptation to multiple domains and multiple tasks is realized, the flexibility and pertinence of the model are improved, and the task processing efficiency is increased.
[0005] In a first aspect, the technical solution of the present invention provides a task processing method based on multiple expert layers, including the following steps:
[0006] Construct a multi-layer expert network, including a basic expert layer for processing general domain tasks, a domain expert layer for processing several specific domain tasks, and a task expert layer for processing several task types;
[0007] Receive an input task, analyze the input task to obtain the task characteristics of the input task, including task domain, task type and task complexity;
[0008] Formulate an expert selection strategy according to the task characteristics;
[0009] Input the input task into the target expert layer for processing according to the expert selection strategy to obtain a processing result.
[0010] In an optional embodiment, analyzing the input task to obtain the task characteristics of the input task specifically includes:
[0011] Extract keywords from the input task, and obtain the task domain through a keyword matching algorithm;
[0012] Process the input task based on a pre-trained classification model to obtain the task type;
[0013] Process the input task based on a pre-trained language model to obtain the task complexity.
[0014] In an optional implementation, formulate an expert selection strategy according to the task characteristics, specifically including:
[0015] Pre-configure the expert selection strategy configuration rules;
[0016] Based on the expert selection strategy configuration rules, formulate an expert selection strategy according to the task characteristics.
[0017] In an optional implementation, based on the expert selection strategy configuration rules, formulate an expert selection strategy according to the task characteristics, specifically including:
[0018] Detect the task domain;
[0019] If the task domain is a general domain, detect the task complexity, and select the basic expert layer as the target expert layer according to the task complexity, or select the basic expert layer and the task expert layer as the target expert layer, where the task expert layer is a task expert layer adapted to the task type; including if the task complexity is low complexity, only select the basic expert layer as the target expert layer to process the input task, and if the task complexity is medium or high complexity, select a combination of the basic expert layer and the task expert layer to process the input task;
[0020] If the task domain is a specific domain, detect the task complexity, and select the domain expert layer as the target expert layer according to the task complexity, or select at least one of the basic expert layer and the task expert layer and the domain expert layer as the target expert layer, where the domain expert layer is a domain expert layer adapted to the task domain, and the task expert layer is a task expert layer adapted to the task type; including if the task complexity is low complexity, only select the domain expert layer, if the task complexity is medium complexity, select the basic expert layer and the domain expert layer, and if the task complexity is high complexity, select the basic expert layer, the domain expert layer and the task expert layer.
[0021] In an optional implementation, before inputting the input task into the target expert layer for processing to obtain a processing result according to the expert selection strategy, the following steps are further included:
[0022] Detect the number of target expert layers;
[0023] If the number of target expert layers is greater than 1, configure weights for each target expert layer through a pre-trained gating network based on the task complexity and the task domain;
[0024] Among them, for the i-th expert layer, the weight output by the gating network is calculated by the following formula,
[0025]
[0026] Wherein, are learned parameters, representing the influence degrees of complexity, domain features and constants, is an activation function, is the task complexity, is the domain feature.
[0027] In an optional implementation manner, the input task is input into the target expert layer according to the expert selection strategy for processing to obtain a processing result, specifically including:
[0028] If there is 1 target expert layer, the input task is input into this 1 target expert layer for processing, and the output of this 1 target expert layer is the processing result;
[0029] If the number of target expert layers is greater than 1, according to the levels and weights of the target expert layers, the input task is input into each target expert layer to obtain a processing result.
[0030] In an optional implementation manner, according to the levels and weights of the target expert layers, the input task is input into each target expert layer to obtain a processing result, specifically including:
[0031] The levels of the basic expert layer, domain expert layer, and task expert layer increase in sequence;
[0032] The input task is input into the target expert layer with the lowest level to obtain the current output, and the current output is multiplied by the weight of the target expert layer with the lowest level and used as the input data to be input into the target expert layer of the next level, and so on, until the output is obtained after the processing of the target expert layer with the highest level, and the output of the target expert layer with the highest level is multiplied by the weight of the target expert layer with the highest level as the processing result.
[0033] In a second aspect, the technical solution of the present invention provides a task processing system based on multiple expert layers, including:
[0034] A multi-layer expert network construction module, configured to construct a multi-layer expert network, including a basic expert layer for processing general domain tasks, a domain expert layer for processing several specific domain tasks, and a task expert layer for processing several task types;
[0035] An input task characteristic analysis module, configured to receive an input task and analyze the input task to obtain the task characteristics of the input task, including the task domain, task type, and task complexity;
[0036] An expert selection strategy formulation module, configured to formulate an expert selection strategy according to the task characteristics;
[0037] An input task processing module, configured to input an input task into a target expert layer according to an expert selection strategy for processing to obtain a processing result.
[0038] In a third aspect, a technical solution of the present invention provides a terminal, including:
[0039] A memory, configured to store a task processing program based on multiple expert layers;
[0040] A processor, configured to implement the steps of the task processing method based on multiple expert layers as described in any one of the above when executing the task processing program based on multiple expert layers.
[0041] In a fourth aspect, a technical solution of the present invention provides a computer-readable storage medium, on which a task processing program based on multiple expert layers is stored, and when the task processing program based on multiple expert layers is executed by a processor, the steps of the task processing method based on multiple expert layers as described in any one of the above are implemented.
[0042] A task processing method, system, terminal, and medium based on multiple expert layers provided by the present invention have the following beneficial effects compared with the prior art: First, a multi-layer expert network including a basic expert layer, a domain expert layer, and a task expert layer is constructed, and any task characteristics of the input task are analyzed, and a combination of different expert layers from each expert layer is selected as an expert selection strategy according to the task characteristics. Finally, the target expert layer is used to process the input task based on the expert selection strategy. The present invention realizes flexible adjustment of the expert layer combination according to the task characteristics (task domain, task type, and task complexity) of different input tasks, thereby realizing task processing for different tasks in different fields, improving the adaptability of the model when processing diverse tasks, and improving the accuracy and reliability of the output result by selecting the most appropriate expert layer for processing, and enhancing the task processing efficiency and performance. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] In order to more clearly illustrate the technical solutions of the present invention, the drawings required to be used in the description will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can also obtain other drawings based on these drawings without creative efforts.
[0044] Figure 1 It is a schematic flowchart of a task processing method based on multiple expert layers provided by an embodiment of the present invention.
[0045] Figure 2 It is a schematic block diagram of a task processing system structure based on multiple expert layers provided by an embodiment of the present invention.
[0046] Figure 3A schematic structural diagram of a terminal provided by an embodiment of the present invention. Detailed implementation manners
[0047] To make the objectives, features, and advantages of the present invention more obvious and understandable, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the specific embodiments of the present invention. Obviously, the embodiments described below are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments in this patent, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this patent.
[0048] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs. The terms used in the description of the present invention in this specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0049] Figure 1 A schematic flowchart of a task processing method based on multiple expert layers provided by an embodiment of the present invention. Among them, Figure 1 The execution subject can be a task processing system based on multiple expert layers. According to different requirements, the order of the steps in this flowchart can be changed, and some can be omitted. The task processing method based on multiple expert layers provided by the embodiments of the present invention is executed by a computer device. Correspondingly, the task processing system based on multiple expert layers runs in the computer device.
[0050] As Figure 1 shown, the method includes the following steps.
[0051] S1. Construct a multi-layer expert network.
[0052] The expert network in this embodiment includes three layers of experts, namely a basic expert layer, a domain expert layer, and a task expert layer.
[0053] The basic expert layer is used to process general domain tasks. The BERT architecture can be adopted, and multiple BERT architectures can be included, such as BERT1, BERT2, and BERT3. When the basic expert layer is selected, any one of BERT1, BERT2, and BERT3 can be selected according to needs, for example, according to the processing effect.
[0054] The domain expert layer is used to handle several specific domain tasks, such as in the medical, legal, and financial fields, etc. It can be fine-tuned according to domain knowledge, and models with a BERT-like structure can be adopted, such as BERT, Roberta, etc. Exemplarily, the domain expert layer includes BERT4 in the medical field, BERT4 in the medical field, BERT5 in the legal field, and BERT6 in the legal field. When selecting the domain expert layer, the corresponding BERT in the domain can be selected according to the task domain and complexity, etc.
[0055] The task expert layer is used to handle several task types, such as question answering, text classification, machine translation, etc. Models with structures such as RNN and LSTM can be adopted. Exemplarily, the task expert layer includes lstm+crf for sequence labeling tasks and softmax for classification tasks. When selecting the task expert layer, the corresponding model can be selected according to the task type.
[0056] S2. Receive the input task and analyze the input task to obtain the task characteristics of the input task.
[0057] The task characteristics of this embodiment include task domain, task type, and task complexity, and the task characteristic analysis is achieved through the following steps.
[0058] S2.1. Extract keywords from the input task and obtain the task domain through a keyword matching algorithm.
[0059] Exemplarily, the domain features of medical tasks can be medical terms, case data, etc.; the domain features of legal tasks can be legal provisions, judgments, etc. The domain features can determine which domain the task belongs to through keyword matching.
[0060] S2.2. Process the input task based on a pre-trained classification model to obtain the task type.
[0061] BERT is a pre-trained language representation model based on the Transformer architecture. In an alternative implementation, a BERT fine-tuned text classification model is used.
[0062] First, collect and label the dataset. Each sample contains the input text and the corresponding task type label, and then divide the training set, validation set, and test set. Use the pre-trained BERT model as the basis. Load a library compatible with BERT, such as the Transformers library of Hugging Face. Add a classification head (which can be a fully connected layer + Softmax) to the output layer of BERT to match the number of task types. Use the training data to fine-tune the model.
[0063] S2.3. Process the input task based on a pre-trained language model to obtain the task complexity.
[0064] The task complexity is determined by factors such as the semantic complexity of the input data and the text length. For example, long texts or texts containing a large number of proper nouns may have a higher complexity, while short texts or ordinary conversations may have a lower complexity. The complexity analysis analyzes the features of the input data through a pre-trained language model (such as BERT) to generate a task complexity score.
[0065] S3. Develop an expert selection strategy according to the task characteristics.
[0066] An optional implementation method is to pre-configure the expert selection strategy configuration rules, and based on the expert selection strategy configuration rules, develop an expert selection strategy according to the task characteristics. It should be noted that the expert selection strategy is a combination of each expert layer.
[0067] An optional implementation method is to develop an expert selection strategy according to the task characteristics based on the expert selection strategy configuration rules, which specifically includes the following steps.
[0068] S3.1. Detect the task domain.
[0069] S3.2. If the task domain is a general domain, detect the task complexity, and select the basic expert layer as the target expert layer according to the task complexity, or select the basic expert layer and the task expert layer as the target expert layer, where the task expert layer is the task expert layer adapted to the task type.
[0070] S3.3. If the task domain is a specific domain, detect the task complexity, and select the domain expert layer as the target expert layer according to the task complexity, or select at least one of the basic expert layer and the task expert layer and the domain expert layer as the target expert layer, where the domain expert layer is the domain expert layer adapted to the task domain, and the task expert layer is the task expert layer adapted to the task type.
[0071] In this optional embodiment, it is first determined whether it is necessary to select the domain expert layer through the task domain, and then the combination with other expert layers is selected according to the complexity. Exemplarily, if the task domain is the general domain and the task complexity is low complexity, only the basic expert layer is selected as the target expert layer to process the input task. If the task complexity is medium or high complexity, the combination of the basic expert layer and the task expert layer is selected to process the input task. When selecting the task expert layer, a specific model is selected according to the task type of the input task. Exemplarily, if the task domain is a specific domain, the domain expert layer is first selected as the target expert layer, and then the selection of other expert layers is performed according to the task complexity. For example, if the task complexity is low complexity, only the domain expert layer is selected. If the task complexity is medium complexity, the basic expert layer and the domain expert layer are selected. If the task complexity is high complexity, the basic expert layer, the domain expert layer, and the task expert layer are selected.
[0072] For example, an input task is a classification task in the legal field with high complexity. The formulated expert selection strategy is the basic layer (BERT1) + the domain expert layer (BERT4 in the legal field) + the task expert layer (classification task Softmax). Among them, the selection of each BERT in the basic layer can pre-configure the selection priorities of different domains, and each BERT in the domain expert layer can be selected according to the task complexity.
[0073] S4. Input the input task into the target expert layer for processing according to the expert selection strategy to obtain a processing result.
[0074] In this embodiment, the finally selected expert layer may include multiple. To further improve the accuracy of the processing result, after formulating the expert selection strategy in this embodiment, if there are multiple target expert layers, weights are configured for each target expert layer, and the result is output according to the weights, which specifically includes the following steps.
[0075] S4.1. Detect the number of target expert layers.
[0076] S4.2. If the number of target expert layers is greater than 1, configure weights for each target expert layer through a pre-trained gating network based on the task complexity and the task domain.
[0077] For the i-th expert layer, the weight output by the gating network is calculated by the following formula,
[0078]
[0079] In the formula, are the learned parameters, indicating the influence degrees of complexity, domain features, and constants, is an activation function, is the task complexity. The more complex the task, the greater the weight value may be. is the domain feature, indicating the requirements for specialized knowledge in the domain to which the task belongs. If the task involves a very specific domain (such as law, medicine, etc.), the contribution of domain experts may need to be greater. is the constant term, used to adjust the fixed contribution of certain basic experts (such as the basic layer).
[0080] S4.3. If the target expert layer is 1, input the input task into this 1 target expert layer for processing, and the output of this 1 target expert layer is the processing result.
[0081] S4.4. If the number of target expert layers is greater than 1, according to the levels and weights of the target expert layers, input the input task into each target expert layer to obtain the processing results.
[0082] In this embodiment, the levels of the basic expert layer, the domain expert layer, and the task expert layer increase in sequence, that is, the basic expert layer has the lowest level and the task expert layer has the highest level. Input the input task into the target expert layer with the lowest level to obtain the current output, multiply the current output by the weight of the target expert layer with the lowest level and use it as the input data to input into the target expert layer of the next level, and so on, until the output is obtained after the processing of the target expert layer with the highest level, and multiply the output of the target expert layer with the highest level by the weight of the target expert layer with the highest level as the processing result.
[0083] Exemplarily, the target expert layer includes a basic expert layer, a domain expert layer, and a task expert layer, and the weights of each layer are , , , and the outputs of each layer are as follows.
[0084] Output of the basic layer: .
[0085] Output of the domain expert layer: .
[0086] Output of the task expert layer: .
[0087] The final processing result is the output of the task expert layer.
[0088] In an alternative embodiment, when training the multi-layer expert network, the loss function is designed as a weighted sum of multiple task losses:
[0089]
[0090] where is the task loss, usually the loss of a classification or generation task. is the domain loss, used to evaluate the capture of domain-specific information by the expert layer. is the expert layer loss, which measures the performance of different expert layers and their contribution to the model.
[0091] An embodiment of a multi-expert-layer task processing method is described in detail above. Based on the multi-expert-layer task processing method described in the above embodiment, an embodiment of the present invention further provides a multi-expert-layer task processing system corresponding to the method.
[0092] Figure 2 A schematic block diagram of the structure of a multi-expert layer task processing system provided in an embodiment of the present invention. In this embodiment, the multi-expert layer task processing system 200 can be divided into multiple functional modules according to the functions it performs. The functional modules may include: a multi-layer expert network construction module 210, an input task characteristic analysis module 220, an expert selection strategy formulation module 230, and an input task processing module 240. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory.
[0093] The multi-layer expert network construction module 210 is used to construct a multi-layer expert network, including a basic expert layer for processing general domain tasks, a domain expert layer for processing several specific domain tasks, and a task expert layer for processing several task types.
[0094] The input task characteristic analysis module 220 is used to receive an input task and analyze the input task to obtain the task characteristics of the input task, including task domain, task type and task complexity.
[0095] The expert selection strategy formulation module 230 is used to formulate an expert selection strategy according to task characteristics.
[0096] The input task processing module 240 is used to input the input task into the target expert layer for processing to obtain a processing result according to the expert selection strategy.
[0097] The task processing system based on multiple expert layers of this embodiment is used to implement the aforementioned task processing method based on multiple expert layers. Therefore, the specific implementation method of this system can be seen in the implementation example part of the task processing method based on multiple expert layers in the previous text. Therefore, its specific implementation method can refer to the description of the corresponding embodiments of each part, which will not be elaborated here.
[0098] In addition, since the task processing system based on multiple expert layers of this embodiment is used to implement the aforementioned task processing method based on multiple expert layers, its function corresponds to that of the aforementioned method and will not be described in detail here.
[0099] Figure 3A schematic structural diagram of a terminal 300 provided by an embodiment of the present invention, including: a processor 310, a memory 320, and a communication unit 330. When the processor 310 is used to implement the task processing program based on multiple expert layers stored in the memory 320, the following steps are implemented:
[0100] Construct a multi-layer expert network, including a basic expert layer for processing general domain tasks, a domain expert layer for processing several specific domain tasks, and a task expert layer for processing several task types;
[0101] Receive an input task, analyze the input task to obtain the task characteristics of the input task, including task domain, task type, and task complexity;
[0102] Formulate an expert selection strategy according to the task characteristics;
[0103] Input the input task into the target expert layer for processing according to the expert selection strategy to obtain a processing result.
[0104] The present invention also provides a computer storage medium, and the storage medium here may be a magnetic disk, an optical disk, a read-only memory (abbreviation: ROM), a random access memory (abbreviation: RAM), etc.
[0105] The computer storage medium stores a task processing program based on multiple expert layers. When the task processing program based on multiple expert layers is executed by a processor, the following steps are implemented:
[0106] Construct a multi-layer expert network, including a basic expert layer for processing general domain tasks, a domain expert layer for processing several specific domain tasks, and a task expert layer for processing several task types;
[0107] Receive an input task, analyze the input task to obtain the task characteristics of the input task, including task domain, task type, and task complexity;
[0108] Formulate an expert selection strategy according to the task characteristics;
[0109] Input the input task into the target expert layer for processing according to the expert selection strategy to obtain a processing result.
[0110] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A task processing method based on multiple expert layers, characterized in that: The following steps are involved: Construct a multi-layer expert network, including a basic expert layer for processing general domain tasks, a domain expert layer for processing several specific domain tasks, and a task expert layer for processing several types of tasks; receiving an input task, analyzing the input task to obtain task characteristics of the input task, wherein the task characteristics include task domain, task type, and task complexity; extracting keywords from the input task, and obtaining the task domain through a keyword matching algorithm; Formulate expert selection strategies based on task characteristics; The number of detection target expert layers; If the number of target expert layers is greater than 1, the pre-trained gating network is used to configure weights for each target expert layer based on task complexity and task domain; Among them, for the i-th expert layer, the weight of the gating network output is Calculated by the following formula, In the formula, is a learned parameter that represents the influence of complexity, domain characteristics, and constants. is an activation function, is the task complexity, For domain characteristics; According to the expert selection strategy, the input task is input into the target expert layer for processing to obtain the processing result, including inputting the input task into each target expert layer to obtain the processing result according to the level and weight of the target expert layer: the levels of the basic expert layer, the domain expert layer, and the task expert layer increase in sequence, and the input task is input into the target expert layer with the lowest level to obtain the current output, and the current output is multiplied by the weight of the target expert layer with the lowest level as input data to the target expert layer of the next level, and so on, until the target expert layer with the highest level is processed to obtain the output, and the output of the target expert layer with the highest level is multiplied by the weight of the target expert layer with the highest level as the processing result.
2. The task processing method based on multiple expert layers according to claim 1 is characterized in that: Analyze the input task to obtain the task characteristics of the input task, including: Process the input task based on the pre-trained classification model to obtain the task type; The input task is processed based on the pre-trained language model to obtain the task complexity.
3. The task processing method based on multiple expert layers according to claim 2 is characterized in that: Formulate an expert selection strategy based on the characteristics of the task, including: Pre-configure expert selection strategy configuration rules; Based on the expert selection strategy configuration rules, an expert selection strategy is formulated according to the task characteristics.
4. The task processing method based on multiple expert layers according to claim 3 is characterized in that: Based on the expert selection strategy configuration rules, an expert selection strategy is formulated according to the task characteristics, including: Detection task area; If the task domain is a general domain, the task complexity is detected, and the basic expert layer is selected as the target expert layer according to the task complexity, or the basic expert layer and the task expert layer are selected as the target expert layer, wherein the task expert layer is a task expert layer adapted to the task type; including if the task complexity is low complexity, only the basic expert layer is selected as the target expert layer to process the input task, and if the task complexity is medium or high complexity, a combination of the basic expert layer and the task expert layer is selected to process the input task; If the task domain is a specific domain, the task complexity is detected, and the domain expert layer is selected as the target expert layer according to the task complexity, or at least one of the basic expert layer, the task expert layer and the domain expert layer are selected as the target expert layer, wherein the domain expert layer is a domain expert layer adapted to the task domain, and the task expert layer is a task expert layer adapted to the task type; including if the task complexity is low complexity, only the domain expert layer is selected, if the task complexity is medium complexity, the basic expert layer and the domain expert layer are selected, and if the task complexity is high complexity, the basic expert layer, the domain expert layer and the task expert layer are selected.
5. The task processing method based on multiple expert layers according to claim 4 is characterized in that: According to the expert selection strategy, the input task is input into the target expert layer for processing to obtain the processing results, including: If there is one target expert layer, the input task is input into the target expert layer for processing, and the output of the target expert layer is the processing result; If the number of target expert layers is greater than one, the input task is input into each target expert layer to obtain the processing result according to the level and weight of the target expert layer.
6. A task processing system based on multiple expert layers, characterized in that: include: A multi-layer expert network building module is used to build a multi-layer expert network, including a basic expert layer for processing general domain tasks, a domain expert layer for processing several specific domain tasks, and a task expert layer for processing several task types; An input task characteristic analysis module is used to receive an input task, analyze the input task to obtain the task characteristics of the input task, and the task characteristics include task domain, task type and task complexity; wherein keywords are extracted from the input task, and the task domain is obtained through a keyword matching algorithm; Expert selection strategy formulation module, used to formulate expert selection strategy according to task characteristics; An input task processing module is used to input the input task into the target expert layer for processing to obtain a processing result according to the expert selection strategy; The input task processing module is also used to detect the number of target expert layers; If the number of target expert layers is greater than 1, the pre-trained gating network is used to configure weights for each target expert layer based on task complexity and task domain; Among them, for the i-th expert layer, the weight of the gating network output is Calculated by the following formula, In the formula, is a learned parameter that represents the influence of complexity, domain characteristics, and constants. is an activation function, is the task complexity, For domain characteristics; The input task processing module is also used to input the input task into each target expert layer to obtain the processing result according to the level and weight of the target expert layer, including the basic expert layer, the domain expert layer, and the task expert layer in increasing order, and input the input task into the target expert layer with the lowest level to obtain the current output, and multiply the current output by the weight of the target expert layer with the lowest level as input data to the target expert layer of the next level, and so on, until the target expert layer with the highest level is processed to obtain the output, and the output of the target expert layer with the highest level is multiplied by the weight of the target expert layer with the highest level as the processing result.
7. A terminal, characterized in that: include: A memory, used for storing a task processing program based on multiple expert layers; A processor is used to implement the steps of the task processing method based on multiple expert layers as claimed in any one of claims 1 to 5 when executing the task processing program based on multiple expert layers.
8. A computer-readable storage medium, characterized in that: The readable storage medium stores a task processing program based on multiple expert layers, and when the task processing program based on multiple expert layers is executed by a processor, the steps of the task processing method based on multiple expert layers as claimed in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Adaptive decision-making method and system based on large language model, and storage medium
CN118503394A
Information interaction method and device, electronic equipment and storage medium
CN119129646A