A dynamic layer pruning system and method for large language models

Through the dynamic layer pruning system and method, the computing resource allocation of large language models is dynamically adjusted, which solves the problems of low reasoning efficiency and resource waste, and achieves efficient computing resource allocation and model performance recovery.

CN120181137BActive Publication Date: 2025-10-03NINGBO ORIENTAL UNIV OF TECH (TEMPORARY NAME)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510652953.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-10-03
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

Existing large-scale language models suffer from low reasoning efficiency, resource waste, and poor adaptability of pruning methods during the reasoning process.

Method used

A dynamic layer pruning system and method for large language models is adopted. The input audio signal is converted into multiple input tokens of equal coding length through a conversion module. Multiple perceptual estimation mappings and self-attention estimation mappings are performed using a routing adjustment device. The allocation of computing resources is dynamically adjusted through a global tag-aware routing algorithm. Combined with the optimization module, a two-stage optimization strategy is used to optimize the parameters of the routing adjustment device.

Benefits of technology

It improves inference efficiency, reduces unnecessary computation, maintains model performance, achieves reasonable allocation of computing resources and model stability, and ensures that the performance of the pruned model recovers or exceeds the original performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120181137B_ABST
    Figure CN120181137B_ABST
Patent Text Reader

Abstract

The present invention relates to a dynamic layer pruning system and method for large-scale language models. By providing a conversion module to convert an input audio signal into embedded representations of multiple input tokens of equal encoding length, and by providing a routing adjustment device to replace the converter model architecture of traditional large-scale language models, a global tag-aware routing algorithm is used to dynamically adjust computing resource allocation, thereby avoiding the inefficiency caused by fixed resource allocation, improving inference efficiency, and reducing unnecessary computation. Furthermore, the routing adjustment device performs multiple perceptual estimation mappings and self-attention estimation mappings on the embedded representation of each input token to obtain the corresponding output token, thereby decoupling the multi-layer perceptron and self-attention layer pruning strategies within the converter model architecture. This makes computing resource allocation more reasonable, effectively avoids the resource waste caused by uniform pruning, and maintains model performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer model technology, and in particular to a dynamic layer pruning system and method for large-scale language models. Background Art

[0002] With the widespread application of large language models (LLMs) in natural language processing, their superior performance and powerful expressive capabilities have achieved remarkable results in various tasks. However, as the scale of large language models continues to increase, the computational resource requirements for these models have also increased dramatically, making computational efficiency during inference a major bottleneck.

[0003] Traditional large-scale language models typically use a deep transformer model architecture. However, the computing resource allocation for each layer in the transformer model architecture is fixed, so even for different input tags, the allocation of computing resources cannot be dynamically adjusted according to actual needs, resulting in wasted computing resources and low inference efficiency.

[0004] To address this issue, existing technologies have disclosed pruning methods to reduce model size and improve inference speed. These methods are categorized into static and dynamic pruning. Static pruning methods specifically include layer-based pruning and importance score-based pruning. Static pruning methods often waste computational resources for some input tags, while some complex input tags fail to receive sufficient computational resources, resulting in severe performance degradation and an inability to dynamically adapt to the complexity of the input tags, leading to poor adaptability.

[0005] Dynamic pruning methods employ a joint training approach, optimizing routing strategies and model parameters simultaneously during the training of the converter model architecture. However, this approach presents potential training instability issues. Furthermore, because the routing strategy in this joint training approach is randomly initialized at the start of training and trained alongside pre-trained model parameters, this can lead to premature pruning decisions, making it difficult for the model to recover its original performance. This ultimately impacts the effectiveness of pruning and the model's inference accuracy, resulting in poor adaptability. Summary of the Invention

[0006] The technical problem to be solved by the present invention is how to overcome the technical shortcomings of existing large-scale language models, such as low inference efficiency, waste of resources, and poor adaptability of pruning methods during inference. To overcome these shortcomings of the existing technology, the present invention provides a dynamic layer pruning system and method for large-scale language models, specifically comprising a dynamic layer pruning system and a dynamic layer pruning method for large-scale language models.

[0007] The present invention provides a dynamic layer pruning system for large language models, comprising:

[0008] a conversion module configured to convert an input audio signal into an embedded representation of a plurality of input tokens of equal encoding length;

[0009] a routing adjustment device, electrically connected to the conversion module, configured to perform multiple perceptual estimation mappings and self-attention estimation mappings on the embedded representation of each segment of the input token to obtain corresponding output tokens; wherein, after each perceptual estimation mapping and each self-attention estimation mapping, a global label-aware routing algorithm is used to perform resource allocation for the perceptual estimation mapping and the self-attention estimation mapping based on the mapping results of the current mapping;

[0010] An optimization module is electrically connected to the routing adjustment device and is configured to perform parameter optimization on the routing adjustment device through a two-stage optimization strategy, and to perform parameter optimization on the global label-aware routing algorithm through a routing optimization loss function in a first-stage optimization strategy, and to perform parameter optimization on the multiple perception estimation mappings and the self-attention estimation mappings through a low-rank adaptation function in a second-stage optimization strategy.

[0011] The disclosed dynamic layer pruning system for large-scale language models utilizes a conversion module to convert input audio signals into embedded representations of multiple input tokens of equal encoding length, thereby enabling the acquisition of audio input tokens. By replacing the traditional converter model architecture of large-scale language models with a routing adjustment device, a global tag-aware routing algorithm is employed to dynamically adjust computational resource allocation, avoiding the inefficiencies associated with fixed resource allocation, improving inference efficiency, and reducing unnecessary computation. Furthermore, the routing adjustment device performs multiple perceptual estimation and self-attention estimation mappings on the embedded representation of each input token to obtain the corresponding output tokens, thereby decoupling the pruning strategies of the multi-layer perceptron and self-attention layers within the converter model architecture. This optimizes computational resource allocation, effectively avoiding the resource waste associated with uniform pruning, and maintaining model performance. Furthermore, an optimization module is employed to optimize the parameters of the routing adjustment device using a two-stage optimization training strategy, ensuring independent optimization of the routing strategy and model parameters, avoiding training instability, and ensuring that the performance of the pruned large-scale language model recovers or exceeds the original performance, resulting in improved environmental adaptability.

[0012] In one possible implementation, the routing adjustment device is a network structure formed by connecting routing-converter networks in parallel, the number of which is equal to the number of embedded representations contained in the embedded representation of each segment of the input token, and each routing-converter network is a network structure formed by connecting multiple routing-converter modules in series.

[0013] In each of the router-converter networks, the last router-converter module is a network structure formed by sequentially connecting a router 1, a normalization unit, and a self-attention unit in series. All of the router-converter modules except the last router-converter module are network structures formed by sequentially connecting a router 1, a self-attention unit, a router 2, and a feedforward network unit in series. Two adjacent router 1s are electrically connected, two adjacent router 2s are electrically connected, and the normalization unit in the last router-converter module is electrically connected to the router 2 in the previous router-converter module.

[0014] The router 1 and the router 2 are both used to execute the global label-aware routing algorithm, the feedforward network unit is used to execute the perceptual estimation mapping, and the self-attention unit is used to execute the self-attention estimation mapping.

[0015] The router-converter module with the above structure not only performs multiple perceptual estimation mappings on the embedded representation of each input token segment through the multiple feedforward network units, but also performs multiple self-attention estimation mappings on the embedded representation of each input token segment through the multiple self-attention units. Finally, the normalization unit and self-attention unit of the router-converter module at the end perform normalization and mapping, respectively, to ultimately obtain an output token corresponding to the embedded representation of the input token segment. Furthermore, since each router-converter module except the router-converter module at the end includes a router 1 and a router 2, a global label-aware routing algorithm can be used to perform resource allocation for the perceptual estimation mapping and self-attention estimation mapping based on the mapping results after each perceptual estimation mapping and each self-attention estimation mapping. This further avoids the inefficiency caused by fixed resource allocation, improves inference efficiency, and further reduces unnecessary computation.

[0016] In one possible implementation, the calculation formula of the global label-aware routing algorithm is as follows:

[0017] ,

[0018] ,

[0019] ,

[0020] Where,

[0021] Representing the input of the global label-aware routing algorithm;

[0022] the encoding length of a single embedding representation representing the embedding representation of each segment of the input tokens;

[0023] a parameter matrix representing the global label-aware routing algorithm;

[0024] represents the noise sampled from the Gumbel distribution;

[0025] represents the temperature parameter used to control the smoothness of the distribution;

[0026] represents the output of the global label-aware routing algorithm;

[0027] For vector , The function definition is:

[0028] .

[0029] The global label-aware routing algorithm, which uses the above computational formula, dynamically allocates computing resources based on the complexity of the embedded representation of each input label, flexibly adjusting the allocation of computing resources to meet the computational requirements of different labels. Specifically, more computing resources are allocated to complex global label-aware routing inputs, while the amount of computation is reduced for simple global label-aware routing inputs, significantly improving inference efficiency and avoiding the inefficiency caused by the fixed allocation of computing resources in static pruning methods.

[0030] In a possible implementation, the calculation formula of the perception estimation mapping is as follows:

[0031] ,

[0032] Where,

[0033] an input representing the perceptual estimation map;

[0034] A result obtained by performing multi-layer perceptron mapping on the input of the perceptual estimation mapping;

[0035] represents a discrete vector obtained by a straight-through Gumbel estimator provided on the feedforward network unit;

[0036] representing an output of the perceptual estimation map;

[0037] This can achieve the goal of pruning and optimizing the multi-layer perceptron in the converter model architecture of traditional large-scale language models, ensuring the rational allocation of computing resources and ensuring that all feedforward network units can replace the multi-layer perceptron and efficiently process task-specific local information.

[0038] In one possible implementation, the self-attention estimation map is calculated as follows:

[0039] ,

[0040] Where,

[0041] represents the input of the self-attention estimation map;

[0042] represents a discrete vector obtained by a straight-through Gumbel estimator placed on the self-attention unit;

[0043] represents the result obtained by performing self-attention mapping on the input of the self-attention estimation map;

[0044] represents the output of the self-attention estimation map;

[0045] This can achieve the goal of pruning and optimizing the self-attention mapping layer in the converter model architecture of traditional large-scale language models, further ensuring the rational allocation of computing resources, and ensuring that all self-attention units can replace the self-attention mapping layer to realize the collection of global contextual relationships.

[0046] In one possible implementation, the optimization module is configured to perform the following steps:

[0047] A1: Keeping the model parameters of all the feedforward network units and all the self-attention units unchanged, optimizing the parameters of all the routers 1 and all the routers 2 using a routing optimization loss function to obtain a routing optimization result;

[0048] A2: Keep the routing optimization result unchanged, optimize the model parameters of all the feedforward network units and all the self-attention units through a low-rank adaptation function, and obtain an optimized routing adjustment device.

[0049] The optimization module, running according to the above steps, freezes the parameters of the pre-trained large language model through step A1, optimizing only the routing strategy. This avoids potential mismatches between the routing strategy and model parameters, ensuring that the routing strategy can be stably optimized based on the model's pre-training, avoiding premature pruning decisions and guaranteeing model stability. Furthermore, by executing step A2, the model parameters of all feedforward network units and all self-attention units are optimized using a low-rank adaptation function, achieving low-rank adaptive fine-tuning. This only adjusts the low-rank adaptation parameters, avoiding large-scale modifications to the original model parameters. This allows the model's performance to be restored after pruning, while reducing the computational resources and time required for fine-tuning, thereby improving training efficiency.

[0050] In one possible implementation, the routing optimization loss function is calculated as follows:

[0051] ,

[0052] ,

[0053] ,

[0054] Where,

[0055] A function value representing the routing optimization loss function;

[0056] represents the standard language modeling loss function;

[0057] Represents the function value of the sparsity loss function;

[0058] Represents weight;

[0059] Represents the target sparsity set for parameter optimization;

[0060] Represents the first order after all the routers 1 and all the routers 2 are sorted in parameter optimization. The router performs the embedding representation of the input token on the A routing skip flag embedded to indicate the decision;

[0061] represents the encoding length of the embedding representation of each input token in parameter optimization;

[0062] represents the number of the routing-switch modules in each routing-switch network;

[0063] This solution introduces a sparsity regularization term (i.e., target sparsity) to control the sparsity of computation, ensuring that the routing strategy can be stably optimized based on the model's pre-trained parameters without being affected by model parameter adjustments.

[0064] In a possible implementation, the calculation formula of the low-rank adaptation function is as follows:

[0065] ,

[0066] Where,

[0067] A parameter matrix representing the model parameters of all the feedforward network units and all the self-attention units before parameter optimization;

[0068] A parameter matrix representing the model parameters of all the feedforward network units and all the self-attention units after parameter optimization;

[0069] and Represents a low-rank matrix that can be adjusted in parameter optimization;

[0070] This scheme introduces an adjustable low-rank matrix, which not only enables fine-tuning of the low-rank matrix, but also recovers the performance loss caused by layer pruning while maintaining the efficiency and computational savings of the fine-tuning process.

[0071] Another technical solution of the present invention is to provide a dynamic layer pruning method for a large language model, comprising the following steps:

[0072] S1: Optimizing the parameters of the routing adjustment device by using a two-stage optimization strategy through an optimization module to obtain an optimized routing adjustment device;

[0073] S2: A transformation module is used to transform the input audio signal into an embedded representation of multiple input tokens of equal encoding length;

[0074] S3: The optimized routing adjustment device performs multiple perceptual estimation mappings and self-attention estimation mappings on the embedded representation of each segment of the input tag to obtain the corresponding output tag; wherein, after each perceptual estimation mapping and each self-attention estimation mapping, a global label perceptual routing algorithm is used to perform resource allocation for perceptual estimation mapping and self-attention estimation mapping according to the mapping results of this time.

[0075] The method disclosed in this application first optimizes the parameters of a routing control device using a two-stage optimization training strategy through an optimization module. This ensures independent optimization of the routing strategy and model parameters, avoids training instability, and ensures that the performance of a pruned large language model recovers or exceeds the original performance, thus having good environmental adaptability. Subsequently, a conversion module converts the input audio signal into an embedded representation of multiple input tokens of equal encoding length, thereby obtaining the audio input tokens. Finally, the routing control device replaces the traditional large language model converter model architecture and adopts a global label-aware routing algorithm to dynamically adjust computing resource allocation, avoiding the inefficiency caused by fixed resource allocation, improving inference efficiency, and reducing unnecessary computation. Furthermore, the routing control device performs multiple perceptual estimation mappings and self-attention estimation mappings on the embedded representation of each input token segment to obtain the corresponding output token, thereby decoupling the pruned strategy of the multi-layer perceptron and self-attention layer in the converter model architecture. This makes computing resource allocation more reasonable, effectively avoids the resource waste caused by uniform pruned representation, and maintains model performance.

[0076] In a possible implementation, step S1 includes the following steps:

[0077] S11: Optimizing parameters of the global label-aware routing algorithm using a routing optimization loss function through the optimization module to obtain a routing optimization result;

[0078] S12: Keeping the routing optimization result unchanged, optimizing the parameters of the multiple perception estimation mapping and the self-attention estimation mapping by the optimization module using a low-rank adaptation function to obtain an optimized routing adjustment device;

[0079] This can further ensure the independent optimization of routing strategies and model parameters, avoid training instability problems, and ensure that the model performance of large language models after pruning recovers or exceeds the original performance, thereby making the optimized routing adjustment device more environmentally adaptable. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] Figure 1 This is a schematic diagram of the structure of a dynamic layer tailoring system for large language models disclosed in an embodiment of the present application;

[0081] Figure 2 Schematic diagram of the structure of each routing-converter module of the routing adjustment device disclosed in the embodiment of the present application, except for the routing-converter module located at the end;

[0082] Figure 3 This is a schematic structural diagram of a routing-converter module at the end of a routing adjustment device disclosed in an embodiment of the present application;

[0083] Figure 4 A flow chart of the method disclosed in the embodiments of this application;

[0084] Figure 5 This is the processing flow of a certain input signal by the optimized routing adjustment device disclosed in the embodiment of this application. DETAILED DESCRIPTION

[0085] First, those skilled in the art should understand that these embodiments are merely used to explain the technical principles of the embodiments of the present application and are not intended to limit the scope of protection of the embodiments of the present application. Those skilled in the art may adjust them as needed to suit specific application scenarios.

[0086] In the embodiments of the present application, unless otherwise clearly specified and limited, the electrical connection between the first feature and the second feature means that there is transmission of electrical signals between the first feature and the second feature, that is, there is an electrical relationship, and the way to achieve the transmission of electrical signals may be electrical connection of wires, radio connection, electrical connection of electromagnetic media (such as semiconductors), communication achieved by channels, etc.

[0087] In the embodiments of the present application, unless otherwise expressly specified or limited, a first feature being "above" or "below" a second feature may mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, a first feature being "above," "above," and "above" a second feature may mean that the first feature is directly above or obliquely above the second feature, or simply means that the first feature is higher in level than the second feature. A first feature being "below," "below," and "below" a second feature may mean that the first feature is directly below or obliquely below the second feature, or simply means that the first feature is lower in level than the second feature.

[0088] The present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0089] See also Figures 1 to 5 The embodiment of the present application discloses a dynamic layer clipping system for large language models. The structural diagram of the dynamic layer clipping system is as follows: Figure 1 As shown, the dynamic layer clipping system includes a conversion module, a routing adjustment device, and an optimization module. The routing adjustment device is electrically connected to the conversion module, and the optimization module is electrically connected to the routing adjustment device. The conversion module is configured to convert the input audio signal into an embedded representation of multiple input tags of equal encoding length. The conversion method is conventional and will not be elaborated here.

[0090] In the dynamic layer pruning system, the routing adjustment device is configured to perform multiple perceptual estimation mappings and self-attention estimation mappings on the embedded representation of each input tag to obtain the corresponding output tag; wherein, after each perceptual estimation mapping and each self-attention estimation mapping, a global label-aware routing algorithm is used to perform resource allocation for the perceptual estimation mapping and self-attention estimation mapping based on the mapping results of this time.

[0091] See also Figure 2 and Figure 3 In this embodiment, the routing adjustment device comprises a network structure composed of a number of routing-converter networks connected in parallel, equal to the number of embedded representations contained in the embedded representation of each input token. Each routing-converter network is composed of multiple routing-converter modules connected in series. The arrows in the figure represent the direction of operation of each routing-converter network. In each routing-converter network, the last routing-converter module is a network structure composed of a router 1, a normalization unit, and a self-attention unit connected in series. All routing-converter modules except the last routing-converter module are a network structure composed of a router 1, a self-attention unit, a router 2, and a feedforward network unit connected in series. Two adjacent router 1s are electrically connected, two adjacent router 2s are electrically connected, and the normalization unit in the last routing-converter module is electrically connected to the router 2 in the previous routing-converter module. Routers 1 and 2 are both used to execute the global label-aware routing algorithm. The feedforward network unit is used to perform perceptual estimation mapping, and the self-attention unit is used to perform self-attention estimation mapping.

[0092] In this embodiment, the calculation formula of the global label-aware routing algorithm is as follows:

[0093] ,

[0094] ,

[0095] ,

[0096] ,

[0097] Where,

[0098] Represents the input of the global label-aware routing algorithm;

[0099] the encoding length of a single embedding representation representing the embedding representation of each segment of the input tokens;

[0100] a parameter matrix representing the global label-aware routing algorithm;

[0101] represents the noise sampled from the Gumbel distribution;

[0102] represents the temperature parameter used to control the smoothness of the distribution;

[0103] Represents the output of the global label-aware routing algorithm.

[0104] For Router 1, The first item represents the probability of skipping the self-attention unit or normalization unit electrically connected to router 1, and the second term represents the probability of executing the algorithm corresponding to the self-attention unit or normalization unit electrically connected to router 1. For router 2, The first item represents the probability of skipping the feedforward network unit electrically connected to router 2, the second term Represents the probability of executing the algorithm of the feedforward network unit electrically connected to router two.

[0105] For vector , The function definition is:

[0106] ,

[0107] because is a two-dimensional vector, so

[0108] ,

[0109] May as well , ,and then,

[0110] ,

[0111] As can be seen, the output of the global label-aware routing algorithm is a two-dimensional vector. For router one, one component represents the value of skipping the self-attention unit or normalization unit electrically connected to router one, and the other component represents the value of executing the algorithm corresponding to the self-attention unit or normalization unit electrically connected to router one. For router two, one component represents the value of skipping the feedforward network unit electrically connected to router two, and the other component represents the value of executing the algorithm corresponding to the feedforward network unit electrically connected to router two. Whether to skip or execute is determined by the size. For example, for router one, if the component representing the value of skipping the self-attention unit or normalization unit electrically connected to router one is greater than the component representing the value of executing the algorithm corresponding to the self-attention unit or normalization unit electrically connected to router one, then the self-attention unit or normalization unit electrically connected to router one is skipped; otherwise, the algorithm corresponding to the self-attention unit or normalization unit electrically connected to router one is executed. The same applies to router two.

[0112] In this embodiment, the calculation formula of the perception estimation mapping is as follows:

[0113] ,

[0114] Where,

[0115] represents the input of the perceptual estimation map;

[0116] The result obtained by performing multi-layer perceptron mapping on the input of the perceptual estimation map;

[0117] represents the discrete vector obtained by the straight-through Gumbel estimator (ST-GumbelEstimator) set on the feedforward network unit;

[0118] Represents the output of the perceptual estimation map.

[0119] In this embodiment, the calculation formula of the self-attention estimation map is as follows:

[0120] ,

[0121] Where,

[0122] Represents the input of the self-attention estimation map;

[0123] represents the discrete vector obtained by the straight-through Gumbel estimator set on the self-attention unit;

[0124] Represents the result obtained by performing self-attention mapping on the input of the self-attention estimation map;

[0125] Represents the output of the self-attention estimation map.

[0126] In this dynamic layer pruning system, the optimization module is configured to perform parameter optimization on the routing adjustment device through a two-stage optimization strategy, and in the first stage optimization strategy, the routing optimization loss function is used to perform parameter optimization on the global label-aware routing algorithm, and in the second stage optimization strategy, the low-rank adaptation function is used to perform parameter optimization on the multiple perception estimation mapping and the self-attention estimation mapping.

[0127] Specifically, the optimization module is configured to perform the following steps:

[0128] A1: Keep the model parameters of all feedforward network units and all self-attention units unchanged, and use the routing optimization loss function to optimize the parameters of all routers 1 and all routers 2 to obtain the routing optimization results.

[0129] The goal of routing optimization is to control the sparsity of computation, which can be achieved by introducing a sparsity regularization term. Set the target sparsity to , defining sparsity is the ratio of the modules calculated in each forward propagation to the total modules. In this embodiment, the calculation formula of the routing optimization loss function is as follows:

[0130] ,

[0131] ,

[0132] ,

[0133] Where,

[0134] Represents the function value of the routing optimization loss function;

[0135] represents the standard language modeling loss function;

[0136] Represents the function value of the sparsity loss function;

[0137] Represents weight;

[0138] Represents the target sparsity set for parameter optimization;

[0139] Represents the first order after all routers 1 and all routers 2 are sorted in parameter optimization. The router's embedding representation of the input token A routing skip flag embedded to indicate the decision;

[0140] represents the encoding length of the embedding representation of each input token in parameter optimization;

[0141] Represents the number of router-switch modules in each router-switch network.

[0142] Furthermore, the parameter matrix of the global label-aware routing algorithm is used as the unknown quantity to be solved, and the minimum value of the routing optimization loss function is set as the goal. Based on the constraint principle that more computing resources are allocated to complex global label-aware routing inputs, while the computational effort is reduced for simple global label-aware routing inputs, the parameter matrix is ​​optimized to obtain the optimized parameter matrix. The specific optimization algorithm can adopt a traditional optimization model solving algorithm or a particle swarm optimization algorithm.

[0143] A2: Keep the routing optimization result unchanged, optimize the model parameters of all feedforward network units and all self-attention units through a low-rank adaptation function, and obtain an optimized routing adjustment device.

[0144] In this embodiment, the calculation formula of the low-rank adaptation function (LoRA) is as follows:

[0145] ,

[0146] Where,

[0147] Represents the parameter matrix composed of the model parameters of all feedforward network units and all self-attention units before parameter optimization;

[0148] Represents the parameter matrix composed of the model parameters of all feedforward network units and all self-attention units after parameter optimization;

[0149] and Represents the low-rank matrix that can be adjusted in parameter optimization. By optimizing these two low-rank matrices and , which can be efficiently fine-tuned for specific tasks without changing the original model.

[0150] In this embodiment, a low-rank matrix is ​​regarded as a matrix whose rank is less than half of the minimum of the number of rows and columns of the matrix. In order to reduce the amount of calculation, two low-rank matrices can be and Set the rank to 1. For a specific task (such as the task that the LLaMA-2 7B model needs to perform), by and performing one or more fine-tuning operations, we obtain , and then obtain the model parameters of all feed-forward network units and all self-attention units, thereby completing the corresponding parameter optimization. Thus, without changing the original model, efficient fine-tuning can be performed for a specific task.

[0151] See Figure 4 , and the usage method corresponding to the dynamic layer pruning system for large language models in this embodiment will be further disclosed below. This method includes the following steps:

[0152] S1: Use the optimization module to perform parameter optimization on the routing adjustment device by adopting a two-stage optimization strategy to obtain an optimized routing adjustment device. Specifically, step S1 includes the following steps: S11: Use the optimization module to perform parameter optimization on the global token-aware routing algorithm by adopting a routing optimization loss function to obtain a routing optimization result; S12: Keep the routing optimization result unchanged, and use the optimization module to perform parameter optimization on the multiple perception estimation mapping and self-attention estimation mapping by adopting a low-rank adaptation function to obtain an optimized routing adjustment device.

[0153] S2: Use the conversion module to convert the input audio signal into an embedded representation of multiple segments of input tokens with equal encoding lengths.

[0154] S3: Use the optimized routing adjustment device to perform multiple perception estimation mappings and self-attention estimation mappings on the embedded representation of each segment of input tokens to obtain corresponding output tokens; among them, after each perception estimation mapping and each self-attention estimation mapping, the global token-aware routing algorithm is used to perform resource allocation for one perception estimation mapping and self-attention estimation mapping according to the mapping result of this time.

[0155] The technical effects generated by the dynamic layer pruning system for large language models in this embodiment will be described below. The routing adjustment device in this embodiment is a network structure formed by connecting three routing-transformer networks in parallel. See Figure 5 , when the routing adjustment device undergoes parameter optimization, for a voice signal with the text "This is a good thing", under the conversion of the conversion module, an embedded representation of two segments of input tokens with equal encoding lengths is formed, and each segment corresponds to three characters, such as Figure 5 "This is" in Figure 5The figure shows the routing of the text message "this is a," with the data flow of each word's embedding represented by a black arrow. Unlike traditional static structural pruning, this embodiment dynamically prunes layers by considering both horizontal and vertical dynamics. In horizontal dynamics, different content receives different computational allocations. In vertical dynamics, the multi-layer perceptron (MLP) and attention modules of the traditional transformer model architecture are decoupled.

[0156] This example further verifies whether the dynamic layer pruning system can maintain the capabilities of the original model after pruning, and compares it with multiple existing pruning methods. Specifically, on the LLaMA-2 7B model, the dynamic layer pruning system (abbreviated as the present invention in Table 1), static pruning methods (such as ShortGPT and Shortened LLaMA), and other dynamic pruning methods (such as MoD-D and D-LLM) were used to evaluate the task accuracy of the pruned model on seven common language understanding tasks (BoolQ, PIQA, HellaSwag, Winogrande, ARC-E, ARC-C, and OpenBookQA). The results are shown in Table 1:

[0157] Table 1

[0158]

[0159] As shown in Table 1, the dynamic layer pruning system achieves excellent task accuracy after pruning 25% of its parameters. Furthermore, experimental measurements show that the dynamic layer pruning system maintains over 90% of its inference performance after pruning 25% of its parameters, significantly outperforming the static pruning method. After fine-tuning with LoRA, the dynamic layer pruning system fully recovers its original performance, even surpassing the accuracy of the original model for some tasks. The dynamic layer pruning system maintains over 90% of its inference performance after pruning 25% of its parameters, significantly outperforming the static pruning method.

[0160] In summary, the dynamic layer pruning system for large language models disclosed in this embodiment achieves audio input tag acquisition by providing a conversion module to convert the input audio signal into embedded representations of multiple input tags of equal encoding length. By providing a routing adjustment device to replace the traditional large language model converter model architecture, a global tag-aware routing algorithm is used to dynamically adjust computing resource allocation, avoiding the inefficiency caused by fixed resource allocation, improving inference efficiency, and reducing unnecessary computation. Furthermore, the routing adjustment device performs multiple perceptual estimation and self-attention estimation mappings on the embedded representation of each input tag to obtain the corresponding output tag, thereby decoupling the pruning strategy of the multi-layer perceptron and self-attention layers within the converter model architecture. This makes computing resource allocation more rational, effectively avoids the resource waste caused by unified pruning, and maintains model performance. Furthermore, by providing an optimization module to optimize the parameters of the routing adjustment device using a two-stage optimization training strategy, the routing strategy and model parameters are optimized independently, avoiding training instability and ensuring that the performance of the pruned large language model recovers or exceeds the original performance, thus achieving good environmental adaptability.

[0161] In the description of the embodiments of the present application, it should be noted that in the description of the present application, terms such as "inside" and "outside" indicating directions or positional relationships are based on the directions or positional relationships shown in the accompanying drawings. This is only for the convenience of description and does not indicate or imply that the device or component must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation on the present application.

[0162] In the description of the present application, the description with reference to the terms "one embodiment", "some embodiments", "in the present embodiment", "specific example", or "some examples" means that the specific features, mechanisms, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, mechanisms, materials or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, those skilled in the art may combine and combine different embodiments or examples described in this specification and the features of different embodiments or examples, unless they are mutually inconsistent.

[0163] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A dynamic layer pruning system for large language models, characterized by: include: a conversion module configured to convert an input audio signal into an embedded representation of a plurality of input tokens of equal encoding length; a routing adjustment device, electrically connected to the conversion module, configured to perform multiple perceptual estimation mappings and self-attention estimation mappings on the embedded representation of each segment of the input token to obtain corresponding output tokens; wherein, after each perceptual estimation mapping and each self-attention estimation mapping, a global label-aware routing algorithm is used to perform resource allocation for the perceptual estimation mapping and the self-attention estimation mapping based on the mapping results of the current mapping; an optimization module, electrically connected to the routing adjustment device, and configured to optimize parameters of the routing adjustment device through a two-stage optimization strategy, and optimize parameters of the global label-aware routing algorithm through a routing optimization loss function in a first-stage optimization strategy, and optimize parameters of the multiple perception estimation maps and the self-attention estimation maps through a low-rank adaptation function in a second-stage optimization strategy; The routing adjustment device is a network structure formed by connecting routing-converter networks in parallel in a number equal to the number of embedded representations contained in the embedded representation of each segment of the input token, and each routing-converter network is a network structure formed by connecting multiple routing-converter modules in series; In each of the router-converter networks, the last router-converter module is a network structure formed by sequentially connecting a router 1, a normalization unit, and a self-attention unit in series. All of the router-converter modules except the last router-converter module are network structures formed by sequentially connecting a router 1, a self-attention unit, a router 2, and a feedforward network unit in series. Two adjacent router 1s are electrically connected, two adjacent router 2s are electrically connected, and the normalization unit in the last router-converter module is electrically connected to the router 2 in the previous router-converter module. The router 1 and the router 2 are both used to execute the global label-aware routing algorithm, the feedforward network unit is used to execute the perceptual estimation mapping, and the self-attention unit is used to execute the self-attention estimation mapping.

2. The dynamic layer pruning system for large language models according to claim 1, characterized in that: The calculation formula of the global label-aware routing algorithm is as follows: IN θ ∈R d×2 , Where, x l Representing the input of the global label-aware routing algorithm; d represents the encoding length of a single embedding representation of each segment of the input token embedding representation; W θ a parameter matrix representing the global label-aware routing algorithm; w l represents the noise sampled from the Gumbel distribution; τ represents the temperature parameter used to control the smoothness of the distribution; y l represents the output of the global label-aware routing algorithm; For vector The definition of the softmax function is: Here, n represents the dimension of X.

3. The dynamic layer pruning system for large language models according to claim 2, characterized in that: The calculation formula of the perception estimation mapping is as follows: a l+1 =g l [1]·(f l (a l )+a l )+g l [0]·a l , Where, a l an input representing the perceptual estimation map; f l (a l ) performing multi-layer perceptron mapping on the input of the perceptual estimation mapping; represents a discrete vector obtained by a straight-through Gumbel estimator provided on the feedforward network unit; a l+1 represents the output of the perceptual estimation map.

4. The dynamic layer pruning system for large language models according to claim 3, characterized in that: The calculation formula of the self-attention estimation map is as follows: b l+1 =h l [1]·(s l (a l )+b l )+h l [0]·b l , Where, b l represents the input of the self-attention estimation map; represents a discrete vector obtained by a straight-through Gumbel estimator placed on the self-attention unit; s l (a l ) represents the result obtained by performing self-attention mapping on the input of the self-attention estimation map; b l+1 represents the output of the self-attention estimation map.

5. The dynamic layer pruning system for large language models according to claim 4, characterized in that: The optimization module is configured to perform the following steps: A1: Keeping the model parameters of all the feedforward network units and all the self-attention units unchanged, optimizing the parameters of all the routers 1 and all the routers 2 using a routing optimization loss function to obtain a routing optimization result; A2: Keep the routing optimization result unchanged, optimize the model parameters of all the feedforward network units and all the self-attention units through a low-rank adaptation function, and obtain an optimized routing adjustment device.

6. The dynamic layer pruning system for large language models according to claim 5, characterized in that: The calculation formula of the routing optimization loss function is as follows: Where, A function value representing the routing optimization loss function; represents the standard language modeling loss function; Represents the function value of the sparsity loss function; α represents the weight; Represents the target sparsity set for parameter optimization; represents the routing skip flag determined by the t-th embedded representation of the embedded representation of the input token by the l-th router after all the routers 1 and all the routers 2 are sorted in the parameter optimization; S represents the encoding length of the embedded representation of each segment of the input token in the parameter optimization; L represents the number of the router-switch modules in each of the router-switch networks.

7. The dynamic layer pruning system for large language models according to claim 5 or 6, characterized in that: The calculation formula of the low-rank adaptation function is as follows: W adapt =W+A·B, Where, W represents a parameter matrix composed of the model parameters of all the feedforward network units and all the self-attention units before parameter optimization; W adapt A parameter matrix representing the model parameters of all the feedforward network units and all the self-attention units after parameter optimization; A and B represent low-rank matrices that can be adjusted in parameter optimization.

8. A dynamic layer pruning method for large language models, characterized in that: The dynamic layer pruning system for large language models according to any one of claims 1 to 7 comprises the following steps: S1: Optimizing the parameters of the routing adjustment device by using a two-stage optimization strategy through an optimization module to obtain an optimized routing adjustment device; S2: A transformation module is used to transform the input audio signal into an embedded representation of multiple input tokens of equal encoding length; S3: The optimized routing adjustment device performs multiple perceptual estimation mappings and self-attention estimation mappings on the embedded representation of each segment of the input tag to obtain the corresponding output tag; wherein, after each perceptual estimation mapping and each self-attention estimation mapping, a global label perceptual routing algorithm is used to perform resource allocation for perceptual estimation mapping and self-attention estimation mapping according to the mapping results of this time.

9. The dynamic layer pruning method for large language models according to claim 8, characterized in that: The step S1 includes the following steps: S11: Optimizing parameters of the global label-aware routing algorithm using a routing optimization loss function through the optimization module to obtain a routing optimization result; S12: Keeping the routing optimization result unchanged, the optimization module performs parameter optimization on the multiple perception estimation mapping and the self-attention estimation mapping by using a low-rank adaptation function to obtain an optimized routing adjustment device.

Citation Information

Patent Citations

  • Combined model compression method and system for pre-training language model

    CN114742036A

  • Joint Speech and Language Model Using Large Language Models

    US20240386881A1