Target user group-oriented expert model generation method and device, and electronic equipment

CN122819318APending Publication Date: 2026-09-25BEIJING XUEZHITU NETWORK TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611139388.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-29
Publication Date
2026-09-25

AI Technical Summary

Technical Problem

[0007]本申请实施例提供了一种面向目标用户群体的专家模型生成方法、装置及电子设备,以至少解决现有压缩方法对大模型的压缩率较低的技术问题

Benefits of technology

[0012]根据本申请的一个方面,提供了一种计算机程序产品或计算机程序,该计算机程序产品或计算机程序包括计算机指令,该计算机指令存储在计算机可读存储介质中。计算机设备的处理器从计算机可读存储介质读取该计算机指令,处理器执行该计算机指令,使得该计算机设备执行上述方法中任一实施例的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122819318A_ABST
    Figure CN122819318A_ABST
Patent Text Reader

Abstract

The application discloses a target user group-oriented expert model generation method and device and electronic equipment. The method comprises the following steps: acquiring a conditional activation distribution of an expert subnetwork in each layer of a mixed expert model when processing sample data, wherein the sample data is used to represent the demand of a target user group; based on the conditional activation distribution, all expert subnetworks in each layer are divided into multiple types of expert networks by using the statistical usage rate and specificity; a minimum expert subset meeting a preset ability coverage rate constraint is selected from a part of the multiple types of expert networks to construct an expert subnetwork set; and ability reconstruction is performed on the expert subnetwork set in each layer to obtain a target user group-oriented expert model. The application solves the technical problem of low compression rate of a large model by using an existing compression method.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and more specifically, to a method, apparatus, and electronic device for generating expert models for a target user group. Background Technology

[0002] This section is intended to provide background or context for the content set forth in the claims or specification, and the content described herein is not acknowledged as prior art simply because it is included in this section.

[0003] The improvement of large language models has long followed a scaling-up path, which involves increasing the number of parameters, training data, and computational budget while keeping the model family, training methods, and evaluation distribution relatively fixed, in order to reduce the model's loss across a wide range of data distributions. This path is often referred to as scaling up. As model size and inference costs increase, another common path is scaling down, which involves reducing the number of parameters, inference computation, and deployment costs by using methods such as knowledge distillation, pruning, quantization, and sparsification, while maintaining the original coverage capabilities.

[0004] However, in actual deployments, models often only serve a specific department, position, or business scenario, and the composition of their actual requests is far narrower than this broad distribution. This leads to a third path called outward derivation: the target of a single model is conditionalized from a broad distribution to the demand distribution of a specific target user group, allowing the model to relinquish knowledge, tasks, and behavioral capabilities outside this demand distribution in exchange for significantly lower single-instance resource overhead. The fundamental difference between outward derivation and downward compression lies not in the compression technique used, but in whether the original capability coverage is maintained: the former changes the evaluation distribution itself, while the latter only relaxes costs while retaining the distribution.

[0005] At the implementation level, existing model compression methods mainly have the following drawbacks: compressing the model with the constraint of maintaining all capabilities, for deployment scenarios that only serve a specific user group, requires allocating more resources from the limited inference resources of the deployer to deploy the model in order to cope with the very few requests made by this group.

[0006] There is currently no effective solution to the above problems. Summary of the Invention

[0007] This application provides an expert model generation method, apparatus, and electronic device for a target user group, which at least solves the technical problem of low compression rate of existing compression methods for large models.

[0008] According to one aspect of the embodiments of this application, a method for generating an expert model for a target user group is provided, comprising: obtaining the conditional activation distribution of expert subnetworks in each layer of a hybrid expert model when processing sample data, wherein the sample data is used to characterize the needs of the target user group; based on the conditional activation distribution, classifying all expert subnetworks in each layer into multiple categories of expert networks by statistically analyzing usage rate and specificity; selecting the smallest subset of experts that meets a preset capability coverage constraint from the partial categories of expert networks in the multiple categories of expert networks to construct an expert subnetwork set for the target user group, wherein the partial categories of expert networks are superior to other categories of expert networks in the multiple categories of expert networks in at least one of usage rate and specificity; performing capability reconstruction on the expert subnetwork set in each layer, so that the inference performance of the expert subnetworks in each layer on the target user group reaches a preset threshold, thereby obtaining an expert model for the target user group.

[0009] According to another aspect of the embodiments of this application, an expert model generation apparatus for a target user group is also provided, comprising: an acquisition unit, configured to acquire the conditional activation distribution of expert subnetworks in each layer of a hybrid expert model when processing sample data, wherein the sample data is used to characterize the needs of the target user group; a partitioning unit, configured to partition all expert subnetworks in each layer into multiple types of expert networks based on the conditional activation distribution by statistically analyzing usage rate and specificity; a construction unit, configured to select the smallest subset of experts that meets a preset capability coverage constraint from the partial category expert networks of the multiple types of expert networks to construct an expert subnetwork set for the target user group, wherein the partial category expert networks are superior to other categories of expert networks in the multiple types of expert networks in at least one of usage rate and specificity; and a reconstruction unit, configured to perform capability reconstruction on the expert subnetwork set of each layer, so that the inference performance of the expert subnetworks in each layer on the target user group reaches a preset threshold, thereby obtaining an expert model for the target user group.

[0010] According to another aspect of the embodiments of this application, a computer-readable storage medium is also provided, the storage medium including a stored program that executes the above-described method when the program is run.

[0011] According to another aspect of the embodiments of this application, an electronic device is also provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor performs the above-described method through the computer program.

[0012] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the steps of any of the embodiments of the methods described above.

[0013] In this embodiment, a minimum subset selection and capability reconstruction method combining usage rate and specificity dual-threshold classification with coverage constraints is adopted. By obtaining the conditional activation distribution of expert subnetworks in the hybrid expert model on the target user group, the usage rate and specificity parameters of each expert subnetwork are statistically analyzed, and the expert subnetworks are divided into four categories: core, general, rare and exclusive, and low value. The minimum expert subset that meets the preset capability coverage constraint is selected to construct the expert subnetwork set, and capability reconstruction is performed on the expert subnetwork set. This can significantly improve the compression ratio while maintaining the inference performance of the target user group, thereby solving the technical problem of low compression ratio of existing compression methods for large models. It achieves near-base model performance in the target user group to realize real requests with lower single-instance resource overhead. Attached Figure Description

[0014] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 This is a flowchart of an expert model generation method for a target user group according to an embodiment of this application; Figure 2 This is a schematic diagram of an expert model generation scheme for a target user group according to an embodiment of this application; Figure 3 This is a schematic diagram of an optional demand data construction scheme according to an embodiment of this application; Figure 4 This is a schematic diagram comparing the consistency of a subset of data sets with the full set of capabilities according to an embodiment of this application; Figure 5 This is a schematic diagram comparing the consistency of a subset of data sets with the full set of capabilities according to an embodiment of this application; Figure 6 This is a schematic diagram showing the frequency distribution of usage by various capabilities of experts according to embodiments of this application; Figure 7 This is a schematic diagram illustrating the distribution of the specific capabilities of each expert according to an embodiment of this application; Figure 8 This is a schematic diagram of an optional expert preference consistency according to an embodiment of this application; Figure 9This is a schematic diagram of an expert model generation apparatus for a target user group according to an embodiment of this application; Figure 10 This is a structural block diagram of a terminal according to an embodiment of this application. Detailed Implementation

[0015] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort should fall within the scope of protection of the present application.

[0016] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0017] First, some nouns or terms that appear in the description of the embodiments of this application shall be interpreted as follows: Hybrid expert model: A neural network architecture in which several parallel feedforward subnetworks (called experts) are set up in each hybrid expert layer, and a router selects a fixed number of experts for each word to participate in the computation. For example, a hybrid expert model with 60 layers, 512 experts in each layer, and 8 experts selected each time can have a total number of parameters in the hundreds of billions, while the effective computation of a single inference is only about 1 / 64 of the total number of parameters.

[0018] Expert subnetwork: One of the feedforward subnetworks arranged in parallel within each hybrid expert layer in a hybrid expert model. Each expert subnetwork independently transforms its input, and the router determines which expert subnetworks to activate based on the input. For example, in layer l, the input vector is routed to k=8 expert subnetworks, each performing an independent linear transformation and nonlinear activation on the input.

[0019] Router: A component in the mixture-of-experts model responsible for selecting expert sub-networks that participate in computation for each token. The router receives an input vector, calculates the gating weights of each expert sub-network through a linear layer and a softmax function, and selects the top k expert sub-networks with the highest weights to participate in computation. For example, the router calculates gating weights for 512 experts and selects the top 8 experts.

[0020] Token: A basic unit obtained after tokenizing text in natural language processing. A token can be a character, a word or a sub-word. For example, the sentence "contract review process" can be tokenized into ["contract", "review", "process"] etc., and each tokenization result is a token.

[0021] Conditional activation distribution: The frequency distribution that each expert sub-network in each layer of the mixture-of-experts model is selected by the router under the condition of sample data from a specific user group. For example, for the target user group of the legal department, the number of times each expert is selected in the 10th layer for the sample data of this group is counted, so as to obtain the conditional activation distribution of this layer under the condition of legal demand.

[0022] Utilization rate: The ratio of the number of times an expert sub-network is selected by the router in the sample data of a given user group to the total number of tokens in that layer. The utilization rate is divided into a first utilization rate (on the target user group) and a second utilization rate (on the entire user group). For example, in the 20th layer, if an expert is selected 800 times among 10,000 tokens, its utilization rate is 0.08.

[0023] Specificity: The degree of deviation of the utilization rate of an expert sub-network on the target user group relative to the utilization rate on the entire user group, which is quantified as a specificity parameter, specifically the logarithmic value of the ratio of the first utilization rate to the second utilization rate. For example, if the utilization rate of an expert on the legal group is 0.15 and the utilization rate on all users is 0.05, the specificity parameter is log(0.15 / 0.05) which is approximately equal to 1.10, indicating that the expert is obviously biased towards legal demands.

[0024] Demand component: Each type of demand obtained after clustering the demand sample set of the target user group. Each demand component corresponds to a group of demand samples with similar functions, and has corresponding observation weight and failure rate. For example, the demands of the legal department can be clustered into demand components such as "contract review", "legal consultation", "compliance inspection", and "document drafting".

[0025] Calibration set: A data set obtained after dividing samples of various types of demands according to the final weight of demand components, including an adaptation subset and a validation subset. The adaptation subset is used for model fine-tuning, and the validation subset is used for performance monitoring. For example, the calibration set contains 10,000 samples, of which 8,000 are the adaptation subset and 2,000 are the validation subset.

[0026] Fitting subset: The portion of the calibration set used for model fine-tuning. For example, the fitting subset contains 8,000 samples weighted according to the demand components, with 3,200 samples allocated to the "contract review" component and 2,400 samples allocated to the "legal consultation" component, etc.

[0027] Validation subset: The portion of the calibration set used to monitor task metrics separately for each demand component. For example, the validation subset contains 2000 samples, used to check whether the task metrics for each demand component, such as "contract review" and "legal consultation," meet the standards after fine-tuning.

[0028] Expert Group: A set of expert subnetworks extracted through co-activation analysis that are frequently selected together on the same lexical unit. The expert subnetworks within an expert group cooperate functionally to complete a specific type of task. For example, experts #23, #87, and #156 (# represents the number) at layer 15 are frequently selected simultaneously when processing legal reasoning lexical units, thus forming an expert group.

[0029] Subnet set: A collection of expert subnets selected from a certain layer of the hybrid expert model to construct an expert model for a target user group. For example, 48 experts are selected from the 512 experts in layer 20 to form the subnet set of that layer.

[0030] Masking: A technique used to control which expert subnetworks participate in computation during inference. By setting the routing weights of expert subnetworks that do not participate in computation to negative infinity, their weights become zero after softmax, effectively excluding those experts. For example, setting a 0 / 1 mask vector for 512 experts allows only the weights of 48 experts to pass through.

[0031] Empty expert slots: During the capacity rebuilding phase, these are newly added expert positions outside the retained expert subnets. The initial parameters of empty expert slots are obtained through a weighted merging or replication of the pruned experts, rather than random initialization. For example, after retaining 48 experts at a certain layer, 8 new empty expert slots are added, with the initial parameters of each slot being the result of a weighted merging of the 464 pruned experts based on their utilization rates.

[0032] Coverage: The proportion of routing paths covered by the selected expert subnet set on the sample data of the target user group out of all routing paths.

[0033] Demand Association Structure: A data structure representing the relationships between various demands of a target user group, expressed in graph form. Nodes are the domain entities, task types, and capability primitives corresponding to demand components, while edges represent co-occurrence, dependency, and hierarchical relationships between nodes. For example, in the legal demand association structure, there is a dependency relationship between "contract review" and "compliance inspection."

[0034] Knowledge graph: A specific implementation of a requirement association structure, which uses a graph structure to represent the relationships between requirement components. For example, a knowledge graph in the legal field can be constructed using entities such as "contract," "review," "terms," ​​and "risk" as nodes and their co-occurrence relationships in legal sessions as edges.

[0035] Point mutual information: A metric that measures whether the co-occurrence frequency of two events exceeds the expected randomness. In expert group extraction, point mutual information is used to determine whether the co-occurrence of two expert subnetworks exceeds the random level.

[0036] Distillation data: During the capability reconstruction phase, the original hybrid expert model serves as the teacher model, used to train the pruned student model. Distillation data is generated based on the demand association structure, ensuring coverage of various needs of the target user group. For example, based on the weights of each demand component in the legal demand association structure, a distillation dataset containing various types of samples, such as contract review and legal consultation, is generated.

[0037] Routing temperature: A parameter that controls the smoothness of the output distribution of the router's softmax function. A higher routing temperature makes the output distribution smoother, helping new expert slots obtain non-zero gating weights, thus obtaining gradients during backpropagation. For example, increasing the routing temperature from 1.0 to 2.0 allows new slots with previously near-zero weights to obtain larger gradient signals.

[0038] Load balancing: A mechanism used in hybrid expert model training to prevent routers from excessively concentrating on selecting a few experts. By incorporating a load balancing loss term into the training loss, routers are encouraged to distribute terms evenly among the experts. In this application, the load balancing mechanism results in some experts being called frequently across various needs; however, the high frequency of these experts does not necessarily represent a correlation with any specific need.

[0039] In related technologies, outward derivation is described as a paradigm that leverages a core model to derive multiple task-specific interfaces to adapt to diverse tasks and interact with the environment. The Mixture-of-Experts (MoE) architecture provides an operational object for outward derivation. This architecture sets up several parallel feedforward subnetworks (called "experts") at each Mixture-of-Experts layer. A router (also called a gating module) selects a fixed number of experts for each token to participate in computation. Therefore, the total number of parameters can reach hundreds of billions, while the effective computational cost of a single inference is far less than the total number of parameters. If a certain type of requirement only stably calls a portion of the experts during inference, then only these experts can be retained, resulting in a subnetwork much smaller than the original model while still maintaining that type of capability.

[0040] It is evident that when the model only needs to serve a specific user group, the needs of that group can be transformed into a computable basis, based on which experts to undertake these needs can be selected from the hybrid expert base model and the rest can be eliminated, thereby achieving performance close to that of the base model on the real requests of that group with significantly lower single-instance resource overhead.

[0041] To address the above issues, according to one aspect of the embodiments of this application, a method embodiment for generating expert models for a target user group is provided. The input of this scheme includes two parts: first, a pre-trained hybrid expert model as the base model, which contains E experts at each layer, and for each word, the router selects k experts at each layer to participate in the calculation; second, data that can characterize the demand composition of the target user group, referred to as demand source data. The output of the scheme is a deployable model, which retains only a subset of the base model's expert set at each layer and supplements a small number of newly added expert slots. The total number of parameters and memory usage are significantly lower than the base model, and the task performance is close to that of the base model on the real requests of the target group.

[0042] It should be noted that the base model in the above input is not limited to a pre-trained single-layer MoE model, but can also be a multi-layer MoE model with experts and routers set up on multiple MoE layers. For hybrid architecture models with MoE structures set up only in some layers, this solution can perform subnet extraction only on these layers, while leaving the remaining layers unchanged.

[0043] Figure 1 This is a flowchart of an expert model generation method for a target user group according to an embodiment of this application, such as... Figure 1 As shown, the method may include the following steps: Step S102: Obtain the conditional activation distribution of the expert sub-networks in each layer of the hybrid expert model when processing sample data. The sample data is used to characterize the needs of the target user group. Step S104: Based on the conditional activation distribution, all expert subnetworks at each layer are divided into multiple types of expert networks by statistical usage rate and specificity. Step S106: Select the smallest subset of experts that meets the preset capability coverage constraint from a subset of expert networks of multiple types of expert networks to construct an expert subnet set for the target user group; Step S108: Perform capability reconstruction on the expert subnet set of each layer to obtain an expert model oriented towards the target user group.

[0044] In the technical solution of this application, the router of the hybrid expert model is constrained by the load balancing mechanism, and there are a large number of experts who are frequently called by various needs. It is impossible to distinguish the experts who are truly biased towards the target needs and the high-frequency experts caused by load balancing based solely on the usage rate. By introducing a specificity parameter, that is, the relationship between the usage rate of the target user group and the usage rate of the entire user group, the degree of bias of experts towards the distribution of target needs can be quantified. Experts with a specificity parameter greater than zero and whose usage rate ranking and specificity ranking both reach the threshold are classified as core experts. Experts with a specificity not greater than zero but with high usage rate are classified as general experts. Experts with high specificity but low usage rate are classified as rare exclusive experts. This achieves the technical effect of distinguishing the experts who truly undertake the target needs from the high-frequency experts of load balancing, avoiding the retention of experts who are irrelevant to the target needs and the removal of experts with low usage rate but obvious bias towards the target needs. This improves the compression ratio while maintaining the inference performance of the target user group.

[0045] Furthermore, this application uses a preset capacity coverage rate as a constraint, selecting experts sequentially in descending order of usage rate until the cumulative coverage reaches the constraint value. This ensures that the selected expert subset can cover the main routing paths of the target user group. When resource budget allows, experts are supplemented based on specificity and expert groups, maintaining the collaborative structure among experts. Because experts within an expert group cooperate functionally, removing some would disrupt the overall functionality, thus achieving the technical effect of maximizing the compression ratio while ensuring routing coverage. Moreover, empty expert slots are initialized outside the retained expert subnet set. Weighted merging or duplication of the removed experts is used as the initial value of the slots instead of random initialization, giving the new slots a reasonable initial parameter distribution. Through routing temperature adjustment and auxiliary load balancing losses, the new slots obtain gradient updates during fine-tuning, becoming effective experts participating in the computation. Simultaneously, using the original hybrid expert model as the teacher model, distillation data is generated based on the demand association structure for knowledge distillation, thereby achieving the technical effect of recovering the capabilities lost during the pruning process and effectively utilizing the added capacity.

[0046] As an optional embodiment, before obtaining the conditional activation distribution of the expert subnetwork in each layer of the hybrid expert model when processing sample data in step S102, the above method further includes a demand data construction stage. Specifically, a demand sample set of the target user group is obtained; each demand sample in the demand sample set is clustered to obtain the initial demand components, corresponding observation weights, and failure rates of multiple types of demands, and a demand association structure representing the relationship between various types of demands of the target user group is constructed; in the demand association structure, H jumps are extrapolated along the edge of the node corresponding to the demand component to determine the candidate demand component and the corresponding initial weight of the extrapolated node, where H is a positive integer; the weights of the demand components are synthesized and normalized according to the observation weights, initial weights, and failure rates to obtain the final weights of the demand components; the samples of various types of demands are divided according to the final weights of the demand components to obtain a calibration set, which includes an adaptation subset and a validation subset.

[0047] In the implementation process, the sources of the requirement sample set include historical dialogue records of the target user group, real-time requests collected online, and requirement samples generated backward from business documents. Clustering adopts a hierarchical clustering method in semantic embedding space, grouping requirement samples with similar functions into the same component. The observation weight is the proportion of each requirement component in the sample set, and the failure rate is the proportion of the requirement of that component that is not met by the basic model. The purpose of extrapolation H-jump is to discover potential requirements that the target user group may have raised but have not yet appeared in the sample. Requirement source data can also be obtained through requirement migration from similar user groups, that is, using requirement components of similar groups as candidate requirement components.

[0048] By employing the above technical solution, clustering and extrapolation transform the original demand samples into structured demand component representations, covering both known and potential demands of the target user group. This provides accurate demand basis for subsequent route observation and expert selection. By incorporating failure rate into weight synthesis, the weights of demand components already adequately met by the basic model are reduced, allowing more resources to be allocated to demands requiring improvement, thereby enhancing the targeted nature of capability reconstruction.

[0049] In the technical solution provided in step S102, the conditional activation distribution of the expert sub-networks in each layer of the hybrid expert model is obtained when processing sample data, and the sample data is used to characterize the needs of the target user group.

[0050] In the specific implementation process, obtaining the conditional activation distribution includes the following sub-steps: First, a routing record hook is set at each layer of the hybrid expert model. When sample data passes through the model, the hook records the expert sequence number selected by the router for each lexical unit. The routing records of the prompt word stage and the generation stage are uniformly stored as an int16 tensor to save memory. At the same time, background corpus (i.e., representative samples of the entire user group) is collected, and the routing information is recorded in the same way. Then, the number of times each expert is selected is counted for the sample data of the target user group according to the demand component, and the background corpus is counted according to the total amount. The summation is used to obtain the conditional activation distribution. The sample data comes from the calibration set, which is a data set obtained by dividing the various demand samples according to the demand component weight of the target user group, including the adaptation subset and the validation subset. Before obtaining the conditional activation distribution, the process also includes a demand data construction stage: obtaining a demand sample set of the target user group, clustering each demand sample in the demand sample set to obtain initial demand components, corresponding observation weights, and failure rates for multiple types of demands, and constructing a demand association structure; extrapolating H hops along the edges of the corresponding nodes of the demand components in the demand association structure to determine candidate demand components and their corresponding initial weights; synthesizing and normalizing the weights of the demand components based on the observation weights, initial weights, and failure rates to obtain the final weights of the demand components; and dividing the samples of each type of demand according to the final weights to obtain a calibration set.

[0051] The technical advantages of the above solution are as follows: Recording only expert serial numbers instead of gating weights makes the method independent of the specific router implementation, allowing it to be implemented even on frameworks that can only derive expert serial numbers. This expands the method's applicability, enabling various hybrid expert model frameworks to use this method for expert subnet extraction. Unifying the recording of the prompt word stage and the generation stage avoids statistical biases caused by stage division, ensuring that the conditional activation distribution accurately reflects the true routing patterns of the target user group, thereby improving the accuracy of subsequent expert classification. Using int16 tensors for storage reduces the memory footprint of routing records for each term to one-quarter of the original, making observation of large-scale models with hundreds of billions of parameters feasible in engineering, thus supporting expert subnet extraction in resource-constrained environments.

[0052] In the technical solution provided in step S104, based on the conditional activation distribution, all expert subnetworks at each layer are divided into multiple types of expert networks by statistical usage rate and specificity.

[0053] In the specific implementation process, the division of expert subnetworks includes the following sub-steps: First, based on the conditional activation distribution, the number of expert selections under each demand component is weighted and merged according to the observation weights to obtain the first usage rate of each expert subnetwork in the target user group. Simultaneously, based on the statistics of the background corpus, the second usage rate of each expert subnetwork in the general background corpus (or background sample data, general corpus sampling, publicly available comprehensive evaluation sets, or request records from other groups) is obtained. Then, the specificity parameter is calculated, specifically the logarithmic difference between the first and second usage rates. The usage rate ranking of each expert subnetwork is determined based on the first usage rate, and the specificity ranking of each expert subnetwork is determined based on the specificity parameter. Next, experts are classified according to a dual threshold: those with a specificity parameter greater than zero, whose usage rate ranking reaches the usage rate threshold, and whose specificity ranking reaches the specificity threshold are classified as core expert networks; those with a specificity parameter not greater than zero and whose usage rate ranking reaches the usage rate threshold are classified as general expert networks; those with a specificity parameter greater than zero, whose specificity ranking reaches the specificity threshold, but whose usage rate ranking does not reach the usage rate threshold are classified as rare exclusive expert networks; and those who do not meet any of the above conditions are classified as low-value expert networks.

[0054] After classification, inter-component difference verification is required to check the consistency of expert classification under each required component. Furthermore, co-activation structure and expert group extraction are performed: multiple expert sub-networks selected at each layer for each word are obtained; the frequency of any two expert sub-networks being jointly selected on the same word is counted to obtain the co-occurrence frequency; based on the deviation of the co-occurrence frequency from the product of the frequency of occurrence of the two participating expert sub-networks (i.e., point mutual information, PMI), it is determined whether the co-occurrence of the expert sub-network pair exceeds the expected level of random co-occurrence; expert sub-network pairs exceeding the expected level of random co-occurrence are retained, and connection edges are established using the expert sub-networks in the retained expert sub-network pairs as nodes to construct a co-activation graph; expert sub-networks corresponding to closely connected nodes in the co-activation graph are assigned to the same expert group. The PMI de-biasing process involves calculating the logarithm of the product of the co-occurrence frequency and the frequency of occurrence of the two experts, divided by the total number of words; a positive value indicates that the co-occurrence exceeds the expected level of random co-occurrence. It should be noted that the usage rate threshold and the specificity threshold can be adjusted according to actual needs. For example, the usage rate threshold can be set to the top 10%, and the specificity threshold can be set to the top 20%. The division of the expert group can be implemented using community discovery algorithms such as the Louvain algorithm.

[0055] The above technical solution employs a dual-threshold classification based on usage rate and specificity. By simultaneously considering the frequency of expert usage within the target group and the degree of bias towards the target needs, it can distinguish between experts truly biased towards the target needs and high-frequency experts resulting from load balancing. This avoids the problem of retaining irrelevant experts and pruning important low-frequency experts based solely on usage rate, thus maintaining inference performance for the target user group while improving the compression ratio. The PMI-based de-biased co-activation graph construction, by statistically analyzing the deviation of the co-occurrence frequency of two experts from random expectations, can identify functionally complementary expert groups. This ensures the maintenance of collaborative structures among experts in subsequent subnet selection, avoiding disruption of functional units.

[0056] In the technical solution provided in step S106, the smallest subset of experts that meets the preset capability coverage constraint is selected from a subset of expert networks of multiple types of expert networks to construct an expert subnet set for the target user group.

[0057] In the specific implementation process, the selection of the minimum expert subset includes the following sub-steps: First, using a preset capability coverage rate as a constraint, expert subnetworks are selected sequentially in descending order of usage rate at each layer until the cumulative coverage rate of the selected expert subnetworks reaches the preset capability coverage rate, thus obtaining the minimum baseline subnetwork set. Then, under the condition that the resource budget allows, expert networks are supplemented to the minimum baseline subnetwork set based on specificity and expert groups, thus obtaining the minimum expert subset. Supplementation based on specificity includes: first, supplementing the core expert networks that are not selected in the minimum baseline subnetwork set, and then supplementing the rare specialized experts that are not selected according to the auxiliary ranking score from high to low, where the auxiliary ranking score is obtained by weighted summation of the standardized values ​​corresponding to the usage rate ranking and the standardized values ​​corresponding to the specificity ranking. Supplementation based on expert groups includes: if there are target expert groups where some expert subnetworks are selected, then supplementing the unselected expert subnetworks of the target expert group according to the group-level preference, where the group-level preference is the mean of the specificity parameters of all expert subnetworks within the expert group. Next, a masking operation is performed on the selected expert subset, which sets the route weights of the experts not selected to negative infinity, making them zero after softmax, which is equivalent to pruning. Finally, a cross-component check is performed to ensure that the route path for each required component is covered by the selected expert subset.

[0058] The technical advantages of the above-mentioned solution are as follows: The greedy selection under coverage constraints, by selecting experts in descending order of usage rate until the coverage requirement is met, ensures that the selected expert subset covers the main routing paths of the target user group, thus guaranteeing that the model's inference ability on the target user group does not significantly decrease after pruning. The minimum subset selection strategy, through a greedy algorithm, selects the fewest experts while satisfying the coverage constraint, achieving the technical effect of maximizing the compression ratio. Based on specificity and expert group supplementation, by prioritizing the supplementation of highly specific experts and complete expert groups, the collaborative structure among experts and coverage of target needs are maintained, avoiding damage to functional units due to the pruning of some expert group members.

[0059] As an optional embodiment, the above method further includes extracting the expert group in the following manner: obtaining multiple expert sub-networks selected at each layer for each word; counting the frequency of any two expert sub-networks being jointly selected on the same word to obtain the co-occurrence frequency of each expert sub-network pair; taking the frequency of any expert sub-network in each expert sub-network pair being selected individually as the occurrence frequency of that expert sub-network, and judging whether the co-occurrence of the expert sub-network pair exceeds the expected level of random co-occurrence based on the deviation of the co-occurrence frequency from the product of the occurrence frequencies of the two expert sub-networks participating in the co-occurrence; retaining expert sub-network pairs whose co-occurrence frequency exceeds the expected level of random co-occurrence; using the expert sub-networks in the retained expert sub-network pairs as nodes and establishing connection edges for expert sub-network pairs to construct a co-activation graph; and assigning the expert sub-networks corresponding to closely connected nodes in the co-activation graph to the same expert group.

[0060] In the specific implementation process, the degree of deviation is measured using Point Mutual Information (PMI). A positive PMI indicates that co-occurrence exceeds the random expectation. The community division of the co-activation graph uses the Louvain algorithm (the community detection algorithm can also be replaced by spectral clustering, label propagation, or hierarchical clustering).

[0061] By employing the above technical solution and using PMI debiasing, the deviation of statistical co-occurrence from random expectation is used to eliminate false co-occurrence caused by high-frequency use by two experts, thereby accurately identifying expert groups that cooperate functionally.

[0062] In the technical solution provided in step S108, the execution capability of the expert subnet set at each layer is reconstructed so that the inference performance of the expert subnet at each layer on the target user group reaches a preset threshold, thereby obtaining an expert model oriented towards the target user group.

[0063] In the specific implementation process, capacity reconstruction includes the following steps: First, using a hybrid expert model as the teacher model, distillation data is generated based on the demand association structure. Specifically, distillation samples covering various demands are generated according to the weights and relationships of each demand component in the demand association structure. The teacher model generates soft labels for these samples for student model training. Then, multiple empty expert slots are initialized outside the expert subnet set retained in each layer. There are two initialization methods: one is to merge the expert networks removed from this layer into a general fallback expert network as the initial value of the slots by weighting according to their usage rate; the other is to copy the most used experts from the removed expert subnets as the initial value of the slots. Next, the intermediate model, including the expert subnet set and empty expert slots, is fine-tuned using the distillation data. During the fine-tuning process, the routing temperature is adjusted to make the softmax output smoother, and an auxiliary load balancing loss term is added to give the new slots non-zero gating weights so as to obtain gradient updates in backpropagation. On the validation subset, task indicators are monitored separately for each demand component. If the task indicator of any demand component fails to meet the standard, the distillation data quota for that demand component is increased and fine-tuning is repeated. After all task metrics for all required components are met, the coverage rate of each layer in the intermediate model is obtained. If the coverage rate of any layer does not reach the preset capacity coverage, a new expert subnet for that layer is selected; if the coverage rate of all layers meets the target, the intermediate model is used as the expert model. Routing temperature can be adjusted using an annealing strategy, gradually decreasing from a higher temperature. The weight of the auxiliary load balancing loss term can also be dynamically adjusted, giving it a higher weight in the early stages of training to promote the activation of new slots, and gradually decreasing it later to focus on task performance. When weighted merging of pruned experts, the weight is the usage rate of each expert among the target user group.

[0064] The technical advantages of the above-mentioned solution are as follows: Non-random initialization of empty expert slots, using weighted merging or duplication of pruned experts as initial values, ensures new slots have a reasonable initial parameter distribution, enabling effective gradient acquisition and rapid convergence during fine-tuning. Routing temperature adjustment and auxiliary load balancing loss, by smoothing the softmax output and encouraging routers to evenly distribute terms, allow new slots to obtain non-zero gating weights and gradient updates during fine-tuning, thus becoming effective experts in computation. Distillation data is generated based on the demand association structure, reusing the relationships between demand components to generate training samples covering various demands, ensuring the distillation process covers all demands of the target user group, thereby recovering the capabilities lost during pruning. A validation loop monitors indicators on the validation subset according to demand components and dynamically adjusts distillation quotas, ensuring the performance of each demand component meets the standards, thus guaranteeing the overall performance of the final model on the target user group.

[0065] As an optional embodiment, the technical solution of this application is further described in detail below with reference to specific embodiments: The overall process of this solution is as follows: Figure 2 As shown, it includes six stages: demand data construction, route observation and collection, statistical and structural analysis, subnet selection and pruning, capacity reconstruction, deployment and hierarchical update, which are divided into three groups according to the problems they solve.

[0066] The first group is Phase One, which addresses the boundary of the target group's needs.

[0067] The second group, comprising stages two through four, addresses which experts undertook these needs. This group runs the base model on the calibration set and collects route observations layer by layer. Based on this, it determines the strength of the association between each expert and the target need, as well as their collaborative structure. Under the constraint that the layer-by-layer route coverage is no less than a threshold α, it selects the smallest subset of experts. The key here is that the absolute frequency of an expert's selection is insufficient to determine their affiliation; it is necessary to simultaneously examine their preference for relatively general backgrounds, and to use stable, co-activated expert groups rather than individual experts as the selection unit.

[0068] The third group consists of stages five and six, which address "how to maintain usability and evolvability after pruning". Stage five utilizes the demand-related structure distillation training data obtained in stage one, and performs fine-tuning after adding several new expert slots to each layer to restore the necessary capabilities lost by the subnet during pruning. Stage six continuously monitors input and output signals after deployment and triggers updates at the corresponding granularity according to the four-level strategy.

[0069] After fine-tuning in Phase 5, the routing distribution of the model has changed, and the subnet selected in Phase 4 may no longer be optimal. Therefore, this scheme sets up a verification loop after Phase 5, returning to Phase 2: re-collecting routing observations and verifying the layer-by-layer coverage. If it meets the standard, it proceeds to Phase 6 for deployment; otherwise, it returns to Phase 4 to reselect the subnet. This verification is an internal check within a single extraction process and is independent of the hierarchical update loop triggered after deployment in Phase 6 in terms of triggering conditions and execution granularity.

[0070] This approach allows for performance degradation in the resulting model across capabilities outside the target group's needs. This degradation is an expected characteristic, not a defect: by forgoing capabilities corresponding to the few requests made by the target group, limited parameters and inference resources can be concentrated on their high-frequency needs, thus maintaining performance within the target range with significantly lower resource overhead than the base model. Therefore, performance acceptance is based on the demand components defined in Phase 1 and conducted on the validation subset specified in Phase 1, not on a general evaluation set.

[0071] Tables 1 and 2 below list the main terms involved in this embodiment and their meanings: Table 1

[0072] Table 2

[0073] Phase 1 Requirements Data Construction The inputs for Phase 1 include: Basic model: only needed when using offline capability detection methods to determine task completion status; Existing business documents, knowledge bases, or classification systems of the target group: used to build the requirement association structure (optional); Demand source data: See Figure 3 It can be any one or a combination of the following six sources: Historical Request Records: Dialogue records, work order logs, or API call records generated by the target group during actual use; this is the most comprehensive implementation method. Short-Term Online Collection: First, provide services to the target group using the basic model or any general model, collect requests within a certain period, and then transfer them to this solution; suitable for the initial stage of system launch. Business Document Reverse Generation: When there are no request records, a representative request set is generated by reverse engineering using job descriptions, business process documents, internal knowledge bases, API documents, or organizational responsibility descriptions, with the help of a language model. Migration to Similar Groups: Use the demand component division and weights of other groups with similar business attributes to the target group as initial values, and then correct them with a small number of target group samples. Manual Assignment: Personnel familiar with the business of this group directly provide a list of demand components and their weights, and assign or select corresponding public datasets or internal datasets to assemble calibration sets for each component; this is the simplest implementation form of this solution, and it can still be fully implemented in subsequent stages.

[0074] When demand data is derived from actual usage records, a bias needs to be corrected: users will actively abandon the ability to repeatedly fail in actual use, and such requests will decrease or even disappear from the records, making the distribution reflected in the records narrower than the actual demand distribution of this group. If pruning is performed directly based on the records, this bias will be solidified into the model structure and further narrowed in subsequent use. Therefore, this solution includes two correction processes in Phase 1, both of which have multiple implementation methods: Firstly, the determination of task completion status is used to identify needs that have appeared in the records but have been underestimated. This can be achieved through: explicit feedback signals (ratings, dislikes, reports); implicit behavioral signals (retry, rewritten questions, session interruption without acceptance, output not being copied); and offline capability detection—in the absence of any user feedback signals, each request in the records is generated and rated using the basic model, and those with ratings below a threshold are considered incomplete.

[0075] Secondly, bounded extrapolation on the demand association structure is used to estimate components of demand that do not appear in the records but belong to the group. The demand association structure refers to any structure capable of expressing the adjacency relationships between demand components, including but not limited to: knowledge graphs (preferred implementations, whose nodes include domain entities, task types, and capability primitives, and whose edges include co-occurrence, dependency, and hierarchical relationships), existing classification systems or ontology of the business domain, tag trees, and nearest neighbor relationships in the request embedding space. Extrapolation extends outward from the components already appearing in the records along this structure to multiple H hops, the first... The weights of the jump components are determined by the attenuation coefficient. Those below the threshold will be excluded.

[0076] Taking a corporate financial analyst position as a specific example: the components that have appeared in the historical records are "report data verification" and "expense anomaly explanation writing." The former corresponds to two capabilities: table processing and numerical calculation, while the latter corresponds to long text generation. Expanding outward one hop from these two sets of capabilities along the association structure, we can reach capabilities such as multi-step numerical reasoning (which has a dependency edge with numerical calculation, as the causal analysis of the verification results requires it as a prerequisite) and chart reading (which has a co-occurrence edge with table processing). The weights of these two are higher than the threshold, so they are included in the demand components and allocated corresponding quotas in the calibration set. However, code generation, which only has a weak association with the existing capabilities, has a weight lower than the threshold and is not included. Thus, although this position has not made multi-step numerical reasoning requests in the records, those capabilities that are prerequisites for existing requirements and are indeed within its demand scope are retained. At the same time, the expansion is limited to a finite range through hop-by-hop decay and thresholds, preventing the demand scope from expanding to a general distribution.

[0077] The attenuation coefficient γ and the upper limit of the number of hops H together constitute the adjustable parameters of the demand range: if the value is too small, it degenerates into direct pruning based on records, resulting in the aforementioned narrowing; if the value is too large, the demand range tends to be a general distribution, and the compression effect disappears. The method for determining its value will be given in detail in the subsequent Phase 1.

[0078] The processing steps for Phase One are as follows: Step S1, preprocessing: perform anonymization, deduplication, and format normalization on the source data of the requirements.

[0079] De-identification: Using rule-based matching and Named Entity Recognition (NER), personal identification information such as names, contact information, ID numbers, and account numbers is removed, while retaining the semantic content of the request. Specifically, firstly, regular expressions are used to match sensitive information with fixed formats, such as mobile phone numbers, ID card numbers, and bank card numbers, and replace them with placeholders; then, a pre-trained NER model is used to identify entities such as person names, place names, and organization names and perform appropriate processing.

[0080] Deduplication: First, perform precise deduplication (text must be completely identical), then perform approximate deduplication using Locality-Sensitive Hashing (LSH) or cosine similarity of embedded vectors. During deduplication, duplicates are not discarded; instead, they are merged into a single record, and their repetition count (n) is recorded. This count is weighted in step S4. Two types of duplication need to be distinguished: mechanical duplication generated by system templates or scheduled tasks should be eliminated, while duplicate requests genuinely submitted by users at different times should be counted. These can be distinguished by source identifiers, the distribution characteristics of request intervals, or complete text consistency. Specifically, if multiple requests originate from the same system identifier and the request intervals exhibit a periodic distribution, they are considered mechanical duplications; if they originate from different times and the interval distribution is not periodic, they are considered genuine duplications and counted.

[0081] Normalization: Each record is organized into a uniform structure, including the request text, target response (if any), session identifier, round number, timestamp, and repetition count. The session identifier and round number are retained because the implicit signal in step S2 requires the sequence within the session; the timestamp is retained because drift detection in stage six requires a time dimension. The output is a normalized set of request records R.

[0082] Step S2: Determine the task completion status by labeling each record in R as either completed, incomplete, or pending. One or more of the following methods can be used; when multiple methods are used, explicit signals take precedence, and the remaining signals are determined by weighted voting: Explicit feedback signals: User ratings, dislikes, reports, or corrections for this round of output. A rating below a preset threshold (e.g., 1 point in a 3-point scale) is considered incomplete; dislikes or reports are directly considered incomplete; corrections indicate that there are errors in the output and are considered incomplete or partially completed.

[0083] Implicit behavioral signals: If a request with a semantic similarity higher than the threshold (e.g., cosine similarity 0.85) is made again within a set time period (e.g., 30 seconds to 5 minutes) in the same session, it is judged as a retry, and the original request is marked as incomplete; if the session terminates after this round of output and there is no adoption action (e.g., no copying, no clicking of the "helpful" button), it is judged as suspicious incomplete; if the output is not copied, not called by downstream processes, or used after being manually rewritten, it is judged as incomplete.

[0084] Offline Capability Detection: In the absence of any feedback collection mechanism, responses are generated for each request in R using the base model. Scores are then assigned based on the consistency of the responses generated through a scoring model, rule validation, or multiple sampling. Responses falling below a threshold are considered incomplete. This method does not rely on any user feedback collection, allowing this step to be performed on purely offline historical data. Specifically, the scoring model can employ a reward model or a rule-based evaluation function. The consistency of multiple sampling is measured by generating multiple responses with different random seeds and calculating the consistency of the answers. Consistency below a threshold indicates that the model's processing of the request is unstable, and the request is considered incomplete.

[0085] Requests that cannot be determined are marked as pending and processed as completed, but are not included in the failure amplification process of step S5 to avoid misjudgment. The output is a set of requests labeled with their completion status.

[0086] Step S3: Demand association structure construction. A structure G is constructed to express the adjacency relationships between demand components. A preferred implementation is a knowledge graph, and its construction process is as follows: Nodes are divided into three categories: Domain Entity Nodes: extracted from the target group's business documents, knowledge base, and interface documents. Task Type Nodes: extracted from request texts, consisting of combinations of actions and objects, or mapped to a pre-defined intent tagging system. Capability Primitive Nodes: such as numerical calculations, code generation, long text understanding, multi-turn instruction compliance, and table processing, which can be taken from a pre-defined list.

[0087] Edges are categorized into three types: co-occurrence edges: two nodes appear in the same request or session, with weight determined by co-occurrence frequency; dependency edges: tasks corresponding to one node require another node as a prerequisite, such as "report anomaly cause analysis" depending on "report data verification," with weight determined by manual assignment or conditional probability; and hierarchical edges: hierarchical relationships derived from existing ontologies or classification systems, such as "reporting attendance regulations" having "company regulations" as its superior.

[0088] Alternative implementation: G can also adopt an existing classification system or label tree in the target group's business domain, or directly use a k-nearest neighbor graph of the request embedding vector. When using a nearest neighbor graph, the "jump" in step S5 is the order of the nearest neighbor.

[0089] G is reused in stage five for generating the distillation training data; both are the same product and are not constructed repeatedly. The output is the demand-related structure G.

[0090] The construction of the demand association structure G can also incorporate annotation feedback from domain experts: after automatically building the initial graph, the graph is presented to domain experts for review and correction, including deleting unreasonable edges, adding missing nodes and edges, and adjusting edge weights. This approach is particularly suitable for scenarios with complex business logic and extremely high accuracy requirements (such as financial compliance and medical diagnostic assistance).

[0091] Step S4: Demand component partitioning and observation weight estimation.

[0092] Vectorize the requests in R and perform clustering to obtain... K Each component has a specific requirement. Clustering methods can include hierarchical clustering, density clustering, or k-means. The value of k can be determined by the silhouette coefficient, the elbow method, or directly using the community detection results on G. Vectorization can use the output of a pre-trained sentence encoder or the embedding layer of a large language model.

[0093] Each component is mapped to one or more nodes in G, establishing a correspondence between components and structures for extrapolation in step S5. The mapping method is as follows: calculate the cosine similarity between the request text embedding of each component and the embedding described by each node in G, and use the nodes with similarity higher than a preset threshold as the mapping targets for that component.

[0094] Calculate the observation weights of each component. With failure rate : , in The number of repetitions recorded in step S1. Indicates that for the part belonging to the first All records of each demand component that were marked as incomplete in step S2. Number of repetitions Summation, This is represented as the normalized request record set output from step S1. This represents a request record in R, and its index is its sequence number. ∈ Representing records After clustering in step S4, the data is assigned to components. , For step S2, record The completion status label can be one of "completed", "incomplete", or "pending".

[0095] Requirements for component granularity: Component division should not remain at broad categories such as "knowledge," "mathematics," and "reasoning." Experimental results show that different subsets within the same broad category have significantly different preferred expert sets (e.g., ...). Figure 4 , Figure 5 As shown, the overlap between subsets within the same class is between 0.15 and 0.82. If coarse categories are used as components, intra-class differences will be masked, and the selected subnets will not reflect the actual preferences of the target group. The component granularity should be fine enough that the requested task type and domain entity within the same component are basically consistent.

[0096] Its output consists of: K demand components, and the value of each component. and , portion to The mapping.

[0097] Step S5 involves bounded extrapolation and weighted synthesis.

[0098] Failure item: Yes Components exceeding the threshold have underestimated observation weights due to user abandonment and must be amplified. Amplification factor. The determination of the parameters will be explained in the parameter determination section below.

[0099] Sample Size Compensation: Failed components often have too few samples in R to support the routing statistics in Phase 3. Therefore, for such components, a persona is constructed based on the corresponding node in G and the role description of the target group. The language model then generates request variants with the same node and task type to expand the sample size of the component and improve its internal diversity. Specifically, the persona includes information such as the job responsibilities of the target group, a glossary of commonly used terms, and typical workflows. These are used as system prompts for the language model, instructing it to generate request variants that match the characteristics of that role. The generated variants must undergo semantic similarity verification with the original request; those that are too close (e.g., cosine similarity higher than 0.95) are considered duplicates and discarded.

[0100] Extrapolation term: Starting from the nodes mapped by each component in step S4, extend outward along the edges of G to multiple hops H. Each new node reached by the jump constitutes a candidate component, and its weight is: , in, The initial weights of the candidate components decrease exponentially with the number of hops. extrapolation jump number ( =1,2,…,H), representing candidate nodes. The number of hops along the demand-related structure G to the nearest activated component node, where γ∈(0,1) is the hop-by-hop decay coefficient. Let G be the set of activated components that are adjacent to the new node. The number of elements in the set. Indicates the first Components obtained by jump extrapolation Weight increment, component Not by When the extrapolation reaches the target value, this item is set to 0. The weight is below the threshold. The candidate components are not included. The minimum observation weight of the existing components can be set as a percentage (e.g., 10%).

[0101] Weight composition and normalization: , , In the synthesis formula, the first term is the observation weight, and the second term is the failure compensation term. The first term is the weight amplification factor for the failed term, and the third term is the extrapolation term.

[0102] The output of this step is: the expanded component set and its final weights. .

[0103] Step S6: Construct the hierarchical calibration set.

[0104] Set the target sample size M for the calibration set. The value of M must be such that the routing statistics in Phase 3 reach stability. The criterion is: gradually increase M and repeat the usage rate calculation in Phase 3. When the Jensen-Shannon divergence (JSD) between two consecutive usage rate distributions is lower than a preset threshold (e.g., 0.01), the statistics are considered to have reached stability. If this condition is not met, return to this step to expand the sample.

[0105] Each component according to Allocate sample quotas.

[0106] When samples are insufficient, they should be supplemented in the following priority order: other real requests belonging to the component within the same group; request variants generated from role profiles; samples from publicly available datasets consistent with the task type of the component; and synthetic data. Each supplementary source must be recorded for future traceability.

[0107] Each sample must carry a label indicating its component. This label is used in Phase Three to... The route counts of each component are weighted and merged. If all samples are mixed and counted without weighting, the components with high proportions will overwhelm the components with low proportions but necessary ones, causing the experts on which the latter depend to be eliminated in stage four.

[0108] Step S7: Separate the adaptation subset from the validation subset.

[0109] Within each component, the data is randomly divided according to a set ratio (e.g., 8:2) to obtain the fit subset. With verification subset The two subsets do not intersect. Hierarchical partitioning ensures that each component is represented in both subsets.

[0110] The component extrapolated from step S5 must be The quota is occupied in the middle, otherwise it is impossible to verify whether the extrapolated part is valid.

[0111] Cross-contamination check: Perform approximately duplicate checks between two subsets. Zhongyu Nearly duplicate samples are removed. Nearly duplicate samples can be determined using the cosine similarity of the embedding vectors, with a threshold of 0.90.

[0112] It is only used for performance acceptance after Phase 5 and monitoring in Phase 6. It does not participate in the route observation and collection in Phase 2, nor in the fine-tuning in Phase 5.

[0113] The output of Phase 1 includes: a) a) The result of dividing the demand into components, with each component corresponding to a type of request from the target group; b) The weight of each demand component. , representing the proportion of this type of request in the target group's needs; c) a hierarchical calibration set consistent with the above weights, divided into mutually disjoint fitting subsets and validation subsets, where each sample carries a component label.

[0114] Additionally, the demand correlation structure G can be output for use in stage five to generate distillation data, including the failure rate of each component. The source record serves as a baseline for drift determination in Phase Six.

[0115] Regarding parameters Determination: The failure amplification factor λ compensates for undersampling caused by users abandoning their requests. The mechanism is as follows: after repeated failures in a certain component, users gradually stop making requests of that type, thus the observed number of requests is lower than the actual demand of that group. One implementation method is to take: , That is, the reciprocal of the failure rate is used as the first-order compensation, and a fixed maximum value is used. (e.g., 5) Cut off to prevent When the value is close to 1, the value diverges.

[0116] Alternatively, a reference method can be used: if the proportion of this component in similar groups or general scenarios can be obtained, this proportion can be used as the upper bound after compensation. Take the maximum value that ensures the weight after compensation does not exceed the upper bound.

[0117] The upper limit of the extrapolation hop count H is 1 or 2. With each additional hop, the number of candidate components increases according to the average degree of G. When H is greater than or equal to 3, the number of candidate components is usually close to the general distribution, and the extrapolation loses its selectivity. When implementing, H=1 can be taken first. If the size of the subnet obtained in stage four is much smaller than the resource budget, it can be relaxed to H=2.

[0118] The hop-by-hop decay coefficient γ determines the breadth of the demand range and is the core parameter that needs to be adjusted according to the target group in this scheme. Its degradation behavior at both ends is as follows: when γ approaches 0, the scheme degenerates into pruning directly based on historical records, and the model's capabilities will gradually narrow with use; when γ approaches 1, the demand range approaches a general distribution, and the compression effect disappears.

[0119] Determination method (end-to-end criterion): For γ∈{0, 0.2, 0.4, 0.6, 0.8}, execute stages one through five respectively, recording the subnet compression ratio corresponding to each value. Based on the retention rate of the task metric, plot the relationship curve between the two; under the premise of meeting the agreed retention rate, take the value with the largest compression ratio, γ. This curve generally has an inflection point shape: when γ is below the inflection point, the retention rate increases rapidly with the increase of γ while the compression ratio decreases only slightly; when γ is above the inflection point, the gain of retention rate slows down while the compression ratio deteriorates rapidly. The inflection point is the value to be taken.

[0120] Low-cost proxy criterion: When not performing the full procedure, only the number of G nodes covered by the calibration set and the total number of corresponding candidate components under each γ value can be counted, and the inflection point where the growth curve of this number relative to γ ​​becomes steeper is taken. This criterion is used to narrow the scan range, and the final value is still based on the end-to-end criterion.

[0121] It should be noted that in the case of a cold start without historical request records, the observations in steps S2 and S4 cannot be statistically analyzed. In this case, the λ term is zero, and the components and weights are given by the aforementioned alternative scheme. Steps S3 and S5 to S7 are executed as usual.

[0122] Phase Two Route Observation and Collection Its inputs are: the base model and the adapted subset of the output from stage one. (Each sample carries component labels), general background corpus.

[0123] The processing steps are as follows: Step S8, Route Recording and Inference Collection: Deploy recording hooks at the router output of each hybrid expert layer in the base model. Perform forward inference once for each sample, record the routing results without changing the routing algorithm itself.

[0124] The record contains the global index of each word in the k experts selected at each level, organized into a shape (…). An integer tensor R, with elements ranging from 0 to E-1, is stored using int16 to reduce memory overhead. These respectively represent word elements, layer numbers, and experts; Prompt words and all lexical units from the generation phase are recorded uniformly without distinction. The final subnet serves the complete inference process, and the statistical objects must be consistent with the deployment and acceptance objects; moreover, tensors themselves do not carry prompt word length information, and forcibly distinguishing them will introduce meta-information dependency and preprocessing error risks. Only when there are insignificant differences between components or disordered co-activation structures in phase three will separate phase data collection be performed for diagnosis. Using only the selected expert's index, an expert ranking in the top k is recorded as 1, without distinguishing rankings or using gating weights. This allows the solution to be implemented even in reasoning frameworks that can only derive expert indices. If the framework can derive gating weights, it can be used as an alternative implementation for weighted calculations, with the subsequent process remaining unchanged. The tensor of each sample is saved along with its component labels. The output is the routing observation tensor and component labels of each sample.

[0125] In routing records, if the inference framework can derive the gating weights (i.e., the softmax probability values ​​output by the router), a weighted statistical method can be used instead of an equal-weighted one: for each term, the k experts selected at a certain layer are counted according to their gating weights, rather than being counted as 1 with equal weights. This method utilizes more information provided by the router and may improve statistical accuracy. Accordingly, the usage rate calculation in Phase 3 needs to replace the count with a weighted sum.

[0126] Step S9 Background corpus collection: Route observation of general background corpus collection is carried out in the same way.

[0127] The background corpus should represent the broad needs addressed by the base model, rather than the needs of the target group. It can be taken from general corpus sampling, publicly available comprehensive evaluation sets, or request records from other groups, and its size should be no less than the fit subset. Specifically, the background corpus may include mixed sampling of publicly available comprehensive evaluation datasets and random sampling of general dialogue corpora.

[0128] The necessity of this step lies in the fact that the subsequent specificity measure is the ratio of each expert's usage rate in the target group to their usage rate in general needs; if the merged results of each component of the target group are used as background, this indicator will degenerate into relative differences between components, failing to reflect the preferences of the target group relative to the general population. The output is the routing observation tensor of the background corpus.

[0129] Step S10: Vectorized count accumulation, converting route observations into expert-selected counts.

[0130] For a single sample, the same layer kEach selection is flattened and then scattered and accumulated into a counting vector of length E in one operation, completing the counting of all layers in one go without using word-by-word iteration. Specifically, for each layer's k expert indices, a scatter-add operation is used to accumulate them into a zero vector of length E, an operation that has been highly optimized in modern deep learning frameworks.

[0131] Cross-sample processing is performed using streaming: each tensor is read in and its count is added to the count of its component before being released, without being fully loaded into memory.

[0132] According to the accumulation of components and layers, The background corpus is accumulated layer by layer. The output consists of the two sets of counting tensors mentioned above. This indicates the sequence number of the demand component. Indicates the hybrid expert layer sequence number ( =1,…,L), This indicates the expert's serial number ( =0,…,E 1).

[0133] The output of this stage includes: layer-by-layer count sheets for each component. , for components All sample words in the first Layer will experts Cumulative number of times the top k items were selected; background counting tensor , as background corpus lexical units in the first Layer will experts The cumulative number of times the top k experts are selected; the determined values ​​of global parameters E, k, and L: E is the total number of experts in each layer, determined by the maximum value of the routing observation tensor elements plus one or by the known model configuration; k is the number of experts selected per word per layer, determined by the size of the third dimension (topk) of the routing observation tensor or by the known model configuration; L is the number of mixed expert layers, determined by the size of the second dimension of the routing observation tensor or by the known model configuration; stability records for each component.

[0134] Phase Three Statistical and Structural Analysis Phase 3 input: Phase 2 input and Phase 1 component weights .

[0135] The processing steps for this stage are as follows: Step S12: Weighted merging and utilization rate calculation.

[0136] First, normalize the count of each component within the layer. Then, by weighting and summing, we obtain the usage rate distribution of the target group: , Instead of directly adding the counts of each component and then normalizing: the distribution obtained by direct addition is dominated by the component with the largest sample size, which does not conform to the actual needs.

[0137] The background distribution is obtained by directly normalizing the background count within the layer. It cannot be obtained by adding the normalized distributions of each component, as the latter does not constitute a probability distribution.

[0138] Retain each component individually This is used in step S15. The output consists of the three distributions mentioned above.

[0139] Step S13, specificity calculation.

[0140] , Here, ε is a smoothing term used to avoid logarithmic divergence when the probability is zero.

[0141] A value greater than zero indicates that the target group uses the expert more often than the general needs, meaning it is biased towards the target group; a value equal to zero indicates that it is on par with the general needs, meaning it is a general expert used for all types of needs; a value less than zero indicates that it is used relatively less.

[0142] This step cannot be replaced by utilization rate: Actual measurements show that the utilization rate curves for different types of demand almost overlap at the shallow level (e.g., ...). Figure 6 As shown), the specificity curves of each layer show obvious separation (e.g. Figure 7 (As shown). This experimental result directly proves that the method of determining expert attribution based solely on usage rate fails completely at shallow levels, and a specificity dimension must be introduced. The output of this step is a layer-by-layer specificity matrix.

[0143] Step S14, dual-threshold expert classification.

[0144] The experts with counts greater than zero in this layer are selected to form the activation set, and all rankings are calculated within this activation set. Experts with counts of zero are never activated by any sample and do not participate in the classification.

[0145] Rankings are determined by usage rate in descending order within the active set. Ranked in descending order of specificity Setting thresholds , Those whose ranking is not higher than the corresponding threshold are respectively referred to as high usage rate and high specificity.

[0146] Here, This indicates that the experts selected are from the top 20% of those with the highest usage rates. This indicates that the top 10% of experts in terms of specificity are selected. These two threshold values ​​can be adjusted according to the actual situation: for layers with a large number of experts (E), the threshold can be appropriately relaxed; for layers with a small number of experts (E), the threshold can be appropriately tightened.

[0147] Based on this, they are divided into four categories: Core experts: Those with specificity greater than zero and meeting the threshold in both rankings are the most deserving of retention. These experts are frequently used by the target group and show a preference for the target group in terms of general needs, making them the direct bearers of the target group's needs.

[0148] General Experts: Highly used but with no more than zero specificity, they handle basic routing but do not reflect the characteristics of the target group; they are retained to maintain coverage. These experts are frequently called upon for various needs and are a product of the load balancing mechanism in the MoE architecture. Although they do not reflect the unique preferences of the target group, they are indispensable for routing coverage.

[0149] Rare Exclusive Experts: These experts have a specificity greater than zero and meet the specificity ranking threshold, but their usage rate ranking does not. They are prioritized below core experts and will be included during the fourth phase of supplementation. These experts are used infrequently, but when they are used, they are clearly biased towards the target group, belonging to the category of "small but highly specialized" experts.

[0150] Low-value experts: Those who fail to meet both thresholds should be prioritized for elimination. These experts are neither frequently used nor shown a preference for the target group, making them the primary targets for elimination.

[0151] Using both dimensions is necessary because looking only at usage rate would retain a large number of experts who are frequently called upon by various needs due to load balancing, which does not reflect the preferences of the target group; looking only at specificity would retain obscure experts who are rarely used and whose occasional appearances result in artificially high ratios, making their retention meaningless. The dual-threshold design is one of the core technical features of this invention.

[0152] Calculate the auxiliary sorting score: , and For the two normalized values ​​within the activation set, This score is a relative weight. It is only used to rank experts who have passed the threshold, and is not used to determine whether someone has passed the threshold.

[0153] Each layer can use the same default threshold; alternatively, it can be adjusted layer by layer based on the inter-component difference index obtained in step S15, with the larger the difference... The smaller the value (because layers with greater differences require more refined specificity screening).

[0154] The output is a hierarchical four-category division and auxiliary sorting score.

[0155] Step S15, component difference verification, verifies whether different components truly correspond to distinguishable expert distributions, this conclusion is a prerequisite for subnet extraction.

[0156] Header list overlap: For each component, the top N members are selected in descending order of specificity to form a set. The intersection-union ratio (IUR) of these sets for two components at the same layer is calculated. A value close to zero indicates that each component has its own dedicated expert, while a value close to one indicates that the so-called dedicated expert is an illusion. The value of N can be set to 10% of the size of the activation set at that layer or a fixed value (e.g., 50).

[0157] Overall distribution distance: Calculate the Jensen-Shannon divergence (JSD) for the usage distribution of the two components in the same layer. The two components are complementary: the former only considers the head list, is discrete and noise-resistant but loses tail information; the latter examines the complete distribution, is continuous and information-complete but is sensitive to tail noise.

[0158] Grouping Consistency: Calculate the mean of expert-specific rank correlation between different subsets within the same component and the mean of rank correlation between different components. A significantly higher mean of the former indicates that the component partitioning is consistent with the expert distribution structure (e.g., ...). Figure 8 As shown, on the four representative layers, the mean rank correlations among subsets within the same capability are 0.45, 0.45, 0.49, and 0.52, respectively, while the corresponding values ​​across capabilities are -0.10, -0.07, -0.03, and 0.05, respectively.

[0159] Processing of the judgment results: If the indicators show that there are significant differences, proceed to step S16; if the overlap of the two components is higher than the threshold in all layers, it is determined that they are indistinguishable at the routing level, and they are merged into one component and returned to step S12 for recalculation; if there are no significant differences between all components in a certain layer, it is determined that there is no capability division in that layer, and the specificity of that layer is not used as the selection criterion, but only the experts with the highest utilization rate are retained according to the coverage constraint of Phase 4.

[0160] The above judgment results must be recorded and declared in the final conclusion. The output consists of the difference index of each layer and the selection strategy label.

[0161] Step S16: Extract co-activated structures and expert groups, extract stable co-activated expert combinations, and use combinations rather than individual experts as selection units.

[0162] Represent the first k choices of each word as a zero-row vector of length E (the e-th element is 1 if and only if expert e is selected), stack all words into a matrix B, and then use the co-occurrence count matrix. , This represents the number of lexical terms that were simultaneously selected by two experts using the same lexical term.

[0163] Directly using co-occurrence counting can be contaminated by high-frequency general experts, who have high co-occurrence rates with any other expert due to their high frequency rather than any form of collaboration. This bias can be removed using Point Mutual Information (PMI). , in , This indicates the sequence number of two experts within the same layer, where T is the total number of tokens included in the statistics. Let C = B be the number of terms in which expert i and expert j are simultaneously selected in the top k by the same term, i.e., the co-occurrence count matrix. B's ( ) elements, The empirical probability of co-occurrence between two experts; Let i be the empirical probability of expert i appearing. For experts The total number of lemmas selected in the top k, where T is the total number of lemmas. A value greater than zero indicates that the co-occurrence frequency exceeds the expectation when independent, and is judged as true coordination; a value close to zero indicates that it originates only from the high frequencies of each individual lemma.

[0164] Weak edges are removed by thresholding edge weights (e.g., only edges with PMI greater than 0 are retained). A graph is constructed using experts as nodes and point mutual information as edge weights. A community detection algorithm (preferably the Leuven algorithm) is run to obtain the expert group for this layer. The Leuven algorithm discovers the community structure by optimizing the modularity function, does not require a preset number of communities, and is suitable for scenarios with a large number of experts.

[0165] Community discovery is used instead of frequent itemset mining: the former does not require pre-setting the size of the community, and the size of the community is determined by the data; the latter requires enumerating combinations of a specified size, which results in an excessively large combination space, sparse and unstable support when the number of experts reaches hundreds. It can be used as a supplementary implementation method to demonstrate ternary or quaternary combinations as higher-order evidence.

[0166] For each expert group, the mean of the group-specific preferences is taken as the group-level preference. A positive group-level preference indicates that the expert group as a whole is biased towards the target group, while a negative preference indicates that it is biased towards general needs.

[0167] Community discovery can also employ alternative algorithms such as spectral clustering and label propagation; however, sensitivity checks on edge weight thresholds and resolution parameters must be performed and recorded. The output of this step is the division of expert groups at each level and group-level preferences.

[0168] Outputs of this stage: Layer-by-layer usage rate distribution and specificity matrix; Layer-by-layer classification of four types of experts and auxiliary ranking scores; Layer-by-layer component difference indicators and selection strategy markings; Layer-by-layer expert group classification and group-level preferences.

[0169] Stage 4 Subnet Selection and Pruning Inputs for this stage: Layered usage distribution of stage three, four types of experts, auxiliary ranking score, expert group division and selection strategy labeling; resource budget (target memory, target parameter quantity).

[0170] The processing steps for this stage are as follows: Step S17: Coverage definition and layer-by-layer greedy selection.

[0171] For a selected subset of experts at a certain level The coverage of a single term is defined as the percentage of terms that the term falls within the range of k experts to which it is routed at that layer. The proportion, that is ,in The top k experts for this word element in this layer; the coverage of this layer. Take the average of this proportion for all word units.

[0172] Coverage is calculated separately for each component, and then weighted according to the components. The weighted merging is consistent with the approach used in step S12. This weighting method ensures that the less significant but necessary components are not overlooked in the coverage assessment.

[0173] Given threshold Seeking satisfaction The smallest A greedy solution is adopted: experts are added in descending order of usage rate, and the cumulative coverage rate reaches [a certain percentage]. If we stop, we can calculate the required number of experts in one step by taking the prefix sum of the sorted counts.

[0174] scanning Plot the curve showing the relationship between subnet size and coverage for subnet size in the range {0.8, 0.9, 0.95}. The horizontal axis represents subnet size and the vertical axis represents coverage. This curve provides a trade-off between compression ratio and coverage for implementers to choose from.

[0175] Each level makes independent choices, and the final result is... The output consists of the baseline subnet and subnet size-coverage curves for each layer.

[0176] In addition to the greedy solution of sorting by usage rate in descending order, the minimum expert subset selection under coverage constraints can also be achieved using the Integer Linear Programming (ILP) method: set whether each expert is selected as a 0-1 decision variable, take minimizing the number of selected experts as the objective function, take the coverage rate of each component as not less than α as the constraint condition, and use commercial or open source solvers to find the exact solution.

[0177] Step S18, based on specificity and supplemented by the expert panel.

[0178] Step S17 yields a baseline subnet based solely on usage rate, which may lack experts with low usage rates but specific to the target group. Supplementing the baseline subnet follows this order: first, add those core experts identified in step S14 who have not yet been selected; then, add rare, specialized experts, selecting the best from highest to lowest auxiliary ranking score.

[0179] Expert Team Constraints: If an expert is selected but the majority of other members in their expert team are not selected, the expert's collaborative structure is disrupted. The expert should either be replaced by new members or removed from the subnet. Expert teams with a positive team-level preference are preferentially retained as a whole. Specifically, "majority" is defined as follows: when the proportion of selected members in an expert's expert team is less than 50%, an expert team integrity check is triggered; if replacing the team does not exceed the resource budget, the team is retained as a whole; otherwise, the expert is removed.

[0180] Supplementation is subject to resource budget constraints and will continue until the budget limit is reached; when the budget is insufficient, core experts will be given priority over rare exclusive experts.

[0181] For layers marked as having no capability division in step S15, skip this step and directly use the baseline subnet from step S17. The output of this step is the final subnet for each layer. .

[0182] Step S19, masking is implemented.

[0183] After the router completes the selection of the top k experts, a mask is applied to experts not selected for the subnet, preventing them from participating in the calculation and thus not modifying the routing algorithm itself. Specifically, based on the top-k index output by the router, a filtering step is added: if an expert is not... If the output is zeroed out, (optionally, in this case, the remaining gating weights need to be renormalized, and if the number of remaining experts is less than k, the experts should be selected based on the routing scores in the reserved set to ensure that each term is still calculated by k experts), so that the expert does not participate in the subsequent feedforward calculation. This implementation makes the pruning reversible and can run without retraining the router.

[0184] Preferred options include resetting the routing rights of unreserved experts to negative infinity before the top-k selection (the top-k selection is automatically re-selected among the reserved experts, and softmax is automatically normalized).

[0185] Record and persistently save the layer-by-layer counts, specificity matrix, expert group divisions, and the mapping table from the original expert indices to their indices within the subnet of the original complete model. This mapping table is a prerequisite for the local reattachment of experts in Phase Six; if it is not saved, any changes in requirements after deployment will require rerunning the entire process.

[0186] The number of experts retained, the total number of parameters, and the memory usage at each layer after pruning were significantly reduced compared to the basic model.

[0187] The output of this step is the masked model, the mapping table, and the compression ratio.

[0188] It should be noted that the mask is only a reversible verification method and the model size is not reduced. After the verification is passed, the parameters of those who are not retained experts must be physically deleted and the corresponding rows must be deleted from the router output matrix (or permanently excluded from the candidate set). This allows the router to re-execute the top-k selection among the retained experts, resulting in a decrease in the total number of parameters and the actual amount of memory.

[0189] Step S20: Cross-component verification.

[0190] Calculate the coverage rate of each component in the final subnet of this layer. If a subnet has high coverage rates for all components, it indicates that the selected experts are general backbone experts shared by various needs, rather than dedicated experts for the target group. In this case, although the compression ratio can be achieved, the selection basis is invalid. It is necessary to return to step S18 to increase the proportion of specific experts, or return to phase one to refine the component granularity. The verification results must be recorded and stated in the final conclusion. Coverage rate is only a proxy indicator for capability retention and cannot be used to claim a capability retention level. The final conclusion must be based on the actual task indicators after phase five. The output of this step is a cross-component coverage table.

[0191] The output of Phase 4 is: the final subnets of each layer. Subnet size-coverage curve and compression ratio; mapping table from original expert number to subnet number and original model statistics; cross-component coverage table.

[0192] Phase 5 Capacity Reconstruction Inputs for this stage: the mask model and mapping table from stage four; the requirement association structure G and component weights from stage one. Validation subset .

[0193] The processing steps for this stage are as follows: Step S21: Distillation training data generation.

[0194] Using nodes in G that are covered by each component as conditions, requests and responses are generated by the base model, constituting distilled data. The base model here serves as the teacher model, and its output represents the performance before pruning. Specifically, for each node in G, requests matching the characteristics of its assigned component's persona are generated, and then the base model generates high-quality responses as supervision signals.

[0195] The generation quota for each node is based on its component. The allocation of data ensures that the composition of the distillation data aligns with the target requirements.

[0196] For the components included through bounded extrapolation in Phase 1, which have few samples in the original records, they must be generated in full according to their weights in this step; otherwise, the capacity of this part cannot be recovered.

[0197] G is reused here and is the same product as that constructed in step S3 of stage one.

[0198] Distillation data must be consistent with Perform approximate duplicate detection and remove duplicates to ensure the validation set is not contaminated by distilled data. The output of this step is the distillation training dataset.

[0199] Step S22: Empty expert slot settings and router remapping.

[0200] Retained on each floor In addition to the existing expert slots, m new expert slots will be added to accommodate any capabilities that need to be restored or added later.

[0201] Slots are not initialized randomly. Newly added experts, even with random initialization, will not be selected by the router and therefore will not receive gradients, becoming permanently ineffective experts. Initialization can be performed using any of the following methods: Method 1: Merge the pruned experts in the same layer according to their usage rates into a single general fallback expert as the initial value for each slot. This involves taking the parameters of each pruned expert, normalizing them according to their respective usage rates, and then performing a weighted average. This initialization ensures the new slot is within the parameter space of the pruned experts, providing it with basic capabilities. Method 2: Directly copy the most frequently used pruned experts as the initial values ​​for each slot. This method is simpler and more direct, but may result in highly similar initial values ​​for multiple slots.

[0202] The router output dimension was adjusted from E to The mapping table from the original expert serial number to the new serial number is updated and saved accordingly. The determination of m is based on the pruning ratio of the layer, the marginal benefit of the coverage curve obtained in step S17, and the number of components included by extrapolation in stage one. In implementation, a value can be taken first according to a certain proportion (such as 10%) of the number of experts retained in each layer, and then adjusted according to the verification results of step S24.

[0203] The output of this step is the model structure with empty expert slots and the updated mapping table.

[0204] Step S23, fine-tuning.

[0205] The model containing the empty expert slot is fine-tuned using the distillation data from step S21, with the goal of restoring the performance to near that of the base model within the target requirement range.

[0206] To ensure that newly added slots have a gradient, activation measures must be taken during the fine-tuning phase: Measure 1: Increase routing temperature to make the route distribution smoother. Routing temperature is a temperature parameter in the router's softmax operation. The higher the temperature, the smaller the difference in the probability of each expert being selected, so that newly added slots that were not originally selected also have a chance to be selected and participate in the calculation, thereby obtaining the gradient.

[0207] Measure 2: Introduce an auxiliary load balancing loss to encourage a certain proportion of terms to be routed to the new slots. Specifically, in addition to the original load balancing loss, an auxiliary loss term is added for the new slots. This term generates a gradient when the activation ratio of the new slots is lower than the target ratio, prompting the router to route some terms to the new slots. Measure 3: Force some terms to be routed to the new slots within several training steps (forced routing). Specifically, in the first N training steps, the routing results of some terms are forcibly replaced with new slots with a certain probability (e.g., 20%), ensuring that these slots obtain gradient signals in the early stages of training.

[0208] After the fine-tuning is complete, restore the normal routing settings.

[0209] You can choose to train only the expert parameters and router parameters while freezing the rest (such as attention layers, layer normalization layers, etc.) to reduce training overhead.

[0210] During the fine-tuning process The indicators are monitored separately for each component to avoid the recovery of high-weight components at the expense of the degradation of low-weight components. If a component indicator shows a significant decrease during fine-tuning (e.g., a decrease exceeding 5%), the weight of that component should be increased in the loss function. The output of this step is the fine-tuned model.

[0211] Step S24, retest the circuit.

[0212] The routing distribution of the model has changed after fine-tuning. The subnet selected in Phase 4 may no longer be optimal. We can return to Phase 2 to re-collect routing observations for the adapted subset and verify the coverage of each layer and the task indicators of each component.

[0213] Those who meet both coverage and performance indicators will proceed to Phase Six of deployment.

[0214] Those whose coverage does not meet the standard return to step S17 of stage four to reselect a subnet; those whose component index does not meet the standard but whose coverage meets the standard return to step S21 to increase the distillation data quota for that component and then readjust.

[0215] The re-verification is an internal verification of a single extraction process, and it is independent of the hierarchical update after the deployment of Phase 6 in terms of both triggering conditions and execution granularity.

[0216] Acceptance at this stage The subset was isolated from the adaptation data from the beginning of the phase and did not participate in the route observation, collection and fine-tuning.

[0217] The output of this step is a verified deployable model and its component-based metric record.

[0218] Outputs of this phase: Deployable model (including retained experts and newly added slots); updated expert sequence number mapping table; task indicator records by component, serving as the baseline for phase six monitoring; distillation dataset and generated quota records, for reuse in L1 level updates of phase six.

[0219] It should be noted that when the coverage is insufficient and the subnet needs to be expanded, the candidate experts must be restored from the original complete model checkpoints based on the complete model routing statistics, expert ranking and mapping table that were persistently saved before fine-tuning, and then the structure reconstruction and fine-tuning must be performed again; it is not allowed to reselect deleted experts based solely on the observations of the scaled-down routers.

[0220] Phase Six Deployment and Tiered Updates Inputs for this phase: the deployable model for phase five, the baseline metrics by component, distillation data and quota records; the mapping table and original model statistics for phase four; and the component partitioning, weights, failure rate baseline and correlation structure G for phase one.

[0221] The processing steps for this stage are as follows: Step S25, Deployment and Signal Monitoring: After the model goes online, two types of signals are continuously collected.

[0222] Input-side signals: Layer-by-layer routing coverage of new requests on the current subnet; component weights of new requests after phase-one clustering and their Jensen-Shannon divergence relative to the baseline; percentage of requests that do not fall into any known component. Input-side signals are sensitive but indirect and can serve as early warnings.

[0223] Output-side signals: Retention rate of task metrics measured separately for each component on the verification subset and new requests; user retry rate, rewrite rate, and negative feedback rate; changes in the failure rate of each component relative to the baseline, determined according to step S2 in Phase 1. Output-side signals are lagging but reliable.

[0224] Judgments must be based on the task indicators on the output side, and conclusions should not be drawn solely based on coverage.

[0225] The output is a monitoring record organized by time and components.

[0226] Step S26: Graded Judgment and Update Execution. Based on the monitoring results, determine the grade according to the table below and execute the corresponding action. The judgment is checked in the order of L3, L2, L1. If a match is found, the action is executed and lower grades are not checked again. See Table 3 below.

[0227] Table 3

[0228] It should be noted that L2 is valid only if step S19 of phase four has persistently saved the layer-by-layer counts, specificity matrix, expert group division, and expert number mapping table of the original complete model. If any of these are missing, the triggering conditions for L2 should be handled as L3. When the basic model is upgraded, the routing structure is completely different, and all the original statistics become invalid, inevitably resulting in L3. The periodic forced reassessment in L3 is used to prevent the gradual drift from accumulating over a long period without triggering the aforementioned conditions. The set period can be determined according to the business scenario, such as once per quarter or once per half year. After each update, the verification in step S24 of phase five must be re-executed, and the new indicators recorded as the baseline for subsequent monitoring.

[0229] The output of this step is the updated model and the new monitoring baseline.

[0230] Outputs of this phase: a continuously running, deployable model; monitoring and grading records; and traceability records of the level, triggering conditions, and execution results of each update.

[0231] In a more specific implementation, the target user group is the legal department of a certain enterprise. The basic model is a hybrid expert model with the following parameters: L=60 layers, E=512 experts, and k=8 experts selected per word per layer.

[0232] Phase 1: Obtain the department's dialogue records from the past six months, and after anonymization and deduplication, obtain request records; after completion status assessment and clustering, obtain several demand components, among which contract clause parsing accounts for the highest proportion, while foreign-related legal retrieval accounts for a relatively low proportion in the records due to repeated failures. After amplifying the failure items and generating variants of the role profile, its weight increases; extrapolate one hop along the knowledge graph (H=1, γ=0.4) to include several adjacent components that did not appear in the records; construct a hierarchical calibration set according to the composite weight and divide it into an adaptation subset and a validation subset at an 8:2 ratio.

[0233] Phase 2: Collect routing observations for the adapted subset and the general background corpus respectively, and obtain the layer-by-layer counts of each component and the background.

[0234] Phase 3: Weighted pooling to obtain the usage rate distribution and calculate specificity; according to Four categories of experts were identified. Difference verification showed that the distribution of experts in the middle and deep layers differed significantly, while the difference was not significant in the shallow layer. Therefore, the shallow layer was marked as selected only by coverage. The co-activation graph and community analysis revealed the expert groups in each layer.

[0235] Phase 4: Using α=0.9, a greedy algorithm is used to select the baseline subnet layer by layer. Core experts and expert groups with positive preferences are added to obtain the final subnets at each layer. After masking, the compression ratio is calculated, and the mapping table and original statistics are saved. Cross-component verification confirms that the coverage of the selected subnets in other types of needs is significantly lower than that of the current group.

[0236] Phase 5: Distill training data according to component weights using the knowledge graph; add a small number of empty expert slots to each layer and initialize with the weighted merging results of the pruned experts, and adjust the router output dimensions accordingly; increase the router temperature and complete fine-tuning with auxiliary load balancing loss; after verifying the route coverage and each component index, proceed to deployment.

[0237] Phase Six: After going live, monitor coverage, component weight divergence, and task indicators for each component; when a component indicator declines, use L1 to supplement distillation data and only fine-tune empty expert slots; when adding a new international M&A business line, use L2 to partially re-attach relevant experts; when the basic model is updated, use L3 to rerun the entire model.

[0238] Furthermore, this solution applies selection constraints on an expert team basis to maintain the collaborative structure among experts; it uses distilled data generated by reusing the demand association structure, combined with non-random initialization and route activation measures for newly added slots, to recover pruning losses and ensure that the added capacity actually participates in training; and it divides post-deployment updates into four levels, allowing most demand changes to be handled by only fine-tuning the newly added slots or partially reverting experts, without having to rerun the entire process, making the update cost commensurate with the degree of change. The implementation conditions of this solution are also relatively lenient: route observation only requires the selected expert's index and does not depend on gating weights; pruning is implemented by adding a mask to the router output without modifying the routing algorithm itself; and demand source data can be reverse-generated from business documents or directly specified by personnel even when there are no historical records.

[0239] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0240] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0241] According to another aspect of the embodiments of this application, an expert model generation apparatus is also provided for implementing the above-described expert model generation method for a target user group. Figure 9 This is a schematic diagram of an optional expert model generation apparatus for a target user group according to an embodiment of this application, such as... Figure 9 As shown, the device may include: The acquisition unit 91 is used to acquire the conditional activation distribution of the expert sub-networks in each layer of the hybrid expert model when processing sample data, wherein the sample data is used to characterize the needs of the target user group. Partitioning unit 93 is used to divide all expert subnetworks in each layer into multiple types of expert networks based on conditional activation distribution, by statistical usage rate and specificity. The construction unit 95 is used to select the smallest subset of experts that meets the preset capability coverage constraint from some categories of expert networks in a multi-class expert network to construct an expert subnet set for the target user group. Among them, some categories of expert networks are superior to other categories of expert networks in the multi-class expert network in at least one of usage rate and specificity. Reconstruction unit 97 is used to reconstruct the execution capabilities of the expert subnet sets at each layer, so that the inference performance of the expert subnets at each layer on the target user group reaches a preset threshold, thereby obtaining an expert model oriented towards the target user group.

[0242] By employing the above modules and combining minimum subset selection and capability reconstruction with usage rate and specificity dual-threshold classification and coverage constraints, we can solve the technical problems of limited compression ratio and performance degradation after pruning when using hybrid expert models to serve specific user groups, due to the constraint of maintaining all capabilities. This achieves a technical effect that is close to the performance of the base model on the real requests of the target user group with significantly lower single-instance resource overhead.

[0243] It should be noted that the examples and application scenarios implemented by the above modules and corresponding steps are the same, but are not limited to the content disclosed in the above embodiments. It should also be noted that the above modules, as part of the device, can run in a corresponding hardware environment, and can be implemented through software or hardware, wherein the hardware environment includes a network environment.

[0244] According to another aspect of the embodiments of this application, a server or terminal for implementing the above-described expert model generation method for a target user group is also provided.

[0245] Figure 10 This is a structural block diagram of a terminal according to an embodiment of this application, such as... Figure 10 As shown, the terminal may include: one or more (only one is shown in the figure) processors 1001, memory 1003, and transmission devices 1005, such as... Figure 10 As shown, the terminal may also include input / output devices 1007.

[0246] The memory 1003 can be used to store software programs and modules, such as the program instructions / modules corresponding to the expert model generation method and apparatus for a target user group in this embodiment. The processor 1001 executes various functional applications and data processing by running the software programs and modules stored in the memory 1003, thereby realizing the aforementioned expert model generation method for a target user group. The memory 1003 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1003 may further include memory remotely located relative to the processor 1001, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0247] The aforementioned transmission device 1005 is used to receive or send data via a network, and can also be used for data transfer between a processor and memory. Specific examples of the network described above may include wired networks and wireless networks. In one example, the transmission device 1005 includes a Network Interface Controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In another example, the transmission device 1005 is a radio frequency (RF) module, used for wireless communication with the Internet.

[0248] Specifically, memory 1003 is used to store application programs.

[0249] The processor 1001 can invoke the application program stored in the memory 1003 via the transmission device 1005 to perform the following steps: Obtain the conditional activation distribution of expert subnetworks in each layer of the hybrid expert model when processing sample data, where the sample data is used to characterize the needs of the target user group; based on the conditional activation distribution, classify all expert subnetworks in each layer into multiple categories of expert networks by statistical usage rate and specificity; select the smallest subset of experts that meets the preset capability coverage constraint from some categories of expert networks to construct the expert subnetwork set for the target user group; perform capability reconstruction on the expert subnetwork set of each layer to obtain the expert model for the target user group.

[0250] This application provides a scheme for generating expert models for a target user group. By using dual-threshold classification based on usage rate and specificity, minimum subset selection constrained by coverage rate, and capability reconstruction, it achieves the goal of approaching the performance of the base model on real requests from the target user group with significantly lower single-instance resource overhead. This results in a significant improvement in compression ratio while maintaining inference performance for the target user group, thereby solving the technical problems of limited compression ratio and performance degradation after pruning when using hybrid expert models to serve a specific user group due to the constraint of maintaining all capabilities.

[0251] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0252] Those skilled in the art will understand that Figure 10 The structure shown is for illustrative purposes only. The terminal can be a smartphone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile internet device (MID), a PAD, or other terminal devices. Figure 10 This does not limit the structure of the aforementioned electronic device. For example, the terminal may also include components that are more... Figure 10 The more or fewer components shown (such as network interfaces, display devices, etc.), or having the same Figure 10 The different configurations shown.

[0253] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing the hardware related to the terminal device. The program can be stored in a computer-readable storage medium, which may include: flash drive, read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0254] Embodiments of this application also provide a storage medium. Optionally, in this embodiment, the storage medium can be used to execute program code for an expert model generation method targeting a specific user group.

[0255] Optionally, in this embodiment, the storage medium may be located on at least one of the network devices in the network shown in the above embodiment.

[0256] Optionally, in this embodiment, the storage medium is configured to store program code for performing the following steps: Obtain the conditional activation distribution of expert subnetworks in each layer of the hybrid expert model when processing sample data, where the sample data is used to characterize the needs of the target user group; based on the conditional activation distribution, classify all expert subnetworks in each layer into multiple categories of expert networks by statistical usage rate and specificity; select the smallest subset of experts that meets the preset capability coverage constraint from some categories of expert networks to construct the expert subnetwork set for the target user group; perform capability reconstruction on the expert subnetwork set of each layer to obtain the expert model for the target user group.

[0257] Optionally, the storage medium is also configured to store program code for performing the following steps: before obtaining the conditional activation distribution, obtaining a demand sample set of the target user group and performing clustering and calibration set partitioning; after obtaining the expert model, continuously monitoring the input and output signals and performing hierarchical updates.

[0258] Optionally, specific examples in this embodiment can refer to the examples described in the above embodiments, and will not be repeated here.

[0259] Optionally, in this embodiment, the storage medium may include, but is not limited to, various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0260] The sequence numbers of the embodiments in this application are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0261] If the integrated units in the above embodiments are implemented as software functional units and sold or used as independent products, they can be stored in the aforementioned computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause one or more computer devices (which may be personal computers, servers, or network devices, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.

[0262] In the above embodiments of this application, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0263] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection between units or modules, and may be electrical or other forms.

[0264] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0265] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0266] The above description is only a preferred embodiment of this application. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of this application, and these improvements and modifications should also be considered within the scope of protection of this application.

Claims

1. A method for generating expert models for a target user group, characterized in that, include: Obtain the conditional activation distribution of expert subnetworks in each layer of the hybrid expert model when processing sample data, wherein the sample data is used to characterize the needs of the target user group; Based on the conditional activation distribution, all expert subnetworks in each layer are divided into multiple types of expert networks by statistical usage rate and specificity. From the partial category expert networks of the multi-category expert networks, the smallest subset of experts that meets the preset capability coverage constraint is selected to construct an expert subnet set for the target user group, wherein the partial category expert networks are superior to other category expert networks in the multi-category expert networks in at least one of usage rate and specificity. The capability reconstruction of the expert subnet set at each layer is performed so that the inference performance of the expert subnet at each layer on the target user group reaches a preset threshold, thereby obtaining an expert model for the target user group.

2. The method according to claim 1, characterized in that, Before obtaining the conditional activation distributions of the expert subnetworks in each layer of the hybrid expert model when processing sample data, the method further includes: Obtain the demand sample set of the target user group; Cluster the demand samples in the demand sample set to obtain the initial demand components, corresponding observation weights and failure rates of multiple types of demands, and construct a demand association structure that represents the relationship between various demands of the target user group. In the demand association structure, extrapolation H jumps are performed along the edges of the nodes corresponding to the demand components to determine the candidate demand components and their corresponding initial weights corresponding to the extrapolated nodes, where H is a positive integer; The weights of the demand components are synthesized and normalized based on the observation weights, initial weights, and failure rates to obtain the final weights of the demand components, which include initial demand components and candidate demand components. The samples of various demands are divided according to the final weight of the demand components to obtain a calibration set. The calibration set includes an adaptation subset and a validation subset, and includes multiple sample data.

3. The method according to claim 2, characterized in that, Constructing a demand association structure that represents the relationships between various needs of the target user group, including: The domain entities, task types, and capability primitives involved in the aforementioned requirement sample set are used as nodes; Based on the co-occurrence, dependency, and hierarchical relationships of nodes in a session, edges are created between nodes to obtain a knowledge graph representing the structure of the required associations.

4. The method according to claim 1, characterized in that, Based on the conditional activation distribution, all expert subnetworks at each layer are classified into multiple types of expert networks by statistical usage rate and specificity, including: Based on the conditional activation distribution, the first usage rate of each expert subnetwork on the target user group and the second usage rate of each expert subnetwork on the general background corpus are obtained, and the specific parameters of each expert subnetwork are obtained through the first usage rate and the second usage rate. The ranking of each expert subnetwork is determined based on the first usage rate, and the ranking of each expert subnetwork is determined based on the specificity parameter. The expert subnetworks that have a specificity parameter greater than zero, a usage rate ranking that reaches the usage rate threshold, and a specificity ranking that reaches the specificity threshold are classified as core expert networks. The expert subnetworks whose specific parameters are not greater than zero and whose usage ranking reaches the usage threshold are classified as general expert networks. All expert subnetworks whose specificity parameter is greater than zero, whose specificity ranking reaches the specificity threshold, and whose usage rate ranking does not reach the usage rate threshold are classified as rare exclusive expert networks. All expert subnetworks that do not meet any of the above conditions are classified as low-value expert networks.

5. The method according to claim 4, characterized in that, Using a preset capability coverage rate as a constraint, a minimum subset of experts is selected from a subset of expert networks across the multiple expert networks, including: Using the preset capability coverage rate α as a constraint, expert subnetworks are selected sequentially in descending order of usage rate at each layer until the cumulative coverage rate of the selected expert subnetworks reaches the preset capability coverage rate α, thereby obtaining the minimum baseline subnetwork set. Within the constraints of the resource budget, the expert network is supplemented to the minimum baseline subnet set based on specificity and expert groups to obtain the minimum expert subset. Specifically, supplementing the expert network based on specificity includes: first, supplementing the core expert network not selected for the minimum baseline subnet set; then, supplementing rare, specialized experts not selected for the minimum baseline subnet set according to their auxiliary ranking scores from high to low. The auxiliary ranking scores are obtained by weighted summation of the standardized values ​​corresponding to usage ranking and specificity ranking. Supplementing the expert network based on expert groups includes: if some expert subnets are selected for the target expert group, supplementing the unselected expert subnets of the target expert group according to the group-level preference, where the group-level preference is the mean of the specificity parameters of all expert subnets within the expert group.

6. The method according to claim 5, characterized in that, The method also includes extracting the expert panel in the following manner: Obtain multiple expert subnetworks selected for each word element at each layer; The co-occurrence frequency of each expert subnetwork pair is obtained by counting the frequency of two expert subnetworks being co-selected on the same word unit. The frequency at which any expert subnetwork in each expert subnetwork pair is individually selected is taken as the occurrence frequency of that expert subnetwork. Based on the deviation of the co-occurrence frequency from the product of the occurrence frequencies of the two expert subnetworks participating in the co-occurrence, it is determined whether the co-occurrence of the expert subnetwork pair exceeds the expected level of random co-occurrence. The expert subnetwork pairs whose co-occurrence frequency exceeds the expected level of random co-occurrence are retained; By using the expert subnetworks in the preserved expert subnetwork pairs as nodes and establishing connection edges between the expert subnetwork pairs, a co-activation graph can be constructed. The expert subnetworks corresponding to closely connected nodes in the co-activation graph are grouped into the same expert group.

7. The method according to claim 1, characterized in that, The capability reconstruction of the expert subnets at each layer is performed so that the inference performance of each expert subnet on the target user group reaches a preset threshold, thereby obtaining an expert model for the target user group, including: Using a hybrid expert model as the teacher model, distillation data is generated based on the demand-related structure. In addition to the expert subnet set retained in each layer, a number of empty expert slots are initialized. The initialization of multiple empty expert slots includes merging the expert networks removed in this layer into a general fallback expert network as the initial value of the slot by weighting the usage rate, and / or copying the most used experts from the removed expert subnets as the initial value of the slot. The distillation data is used to fine-tune the intermediate model, which includes the expert subnet set and empty expert slots. On the validation subset, task metrics are monitored separately for each demand component. If any demand component fails to meet the task metrics, the distillation data quota for that demand component is increased and then fine-tuned. If the task metrics of all demand components meet the requirements, the coverage of each layer in the intermediate model is obtained. If the coverage of any layer in the intermediate model does not reach the preset capability coverage, then the expert subnet set of that layer is reselected; if the coverage of all layers in the intermediate model reaches the preset capability coverage, then the intermediate model is used as the expert model.

8. The method according to any one of claims 1 to 7, characterized in that, After obtaining the expert model for the target user group, the method further includes: After the expert model is launched, the input and output signals are continuously monitored. The system performs hierarchical updates based on the input signal and the output signal.

9. An expert model generation device for a target user group, characterized in that, include: The acquisition unit is used to acquire the conditional activation distribution of the expert sub-networks in each layer of the hybrid expert model when processing sample data, wherein the sample data is used to characterize the needs of the target user group. A partitioning unit is used to divide all expert subnetworks of each layer into multiple types of expert networks based on the conditional activation distribution and by statistically analyzing usage rate and specificity. The construction unit is used to select the smallest subset of experts that meets the preset capability coverage constraint from the partial category expert networks of the multi-category expert networks to construct an expert subnet set for the target user group, wherein the partial category expert networks are superior to other category expert networks in the multi-category expert networks in at least one of usage rate and specificity. The reconstruction unit is used to reconstruct the performance of the expert subnet set at each layer, so that the inference performance of the expert subnet at each layer on the target user group reaches a preset threshold, thereby obtaining an expert model for the target user group.

10. A computer-readable storage medium, characterized in that, The storage medium includes a stored program, wherein the program executes the method described in any one of claims 1 to 8 when it is run.

11. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the method described in any one of claims 1 to 8 through the computer program.