Foundation model fine tuning method with efficient training and reasoning capability
By extracting and merging target unit data pairs, generating compressed data and inputting into multi-head self-attention modules, the problem of low training and inference efficiency in fine-tuning of basic models in visual tasks is solved, and efficient computing resource utilization and model deployment are achieved.
Patent Information
- Application Number
- CN202510218787.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art is difficult to improve training efficiency and inference efficiency simultaneously in the fine-tuning of the basic model of visual tasks, resulting in high computing resources and low deployment efficiency.
By extracting the target unit data pair from the unit sequence data of the target multi-head self-attention module, generating compressed data, and replacing the original data input module, reducing the computational amount and improving inference efficiency.
It realizes efficient training and inference during the fine-tuning of basic models in visual tasks, reduces computing resource consumption, and improves the inference speed of the model in downstream task deployment.
Smart Images

Figure CN120146131A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of visual model fine-tuning, and specifically relates to a basic model fine-tuning method, device, equipment and storage medium with efficient training and reasoning capabilities. Background Art
[0002] With the development of deep learning technology, the application of pre-trained base models in visual tasks has become more and more widespread. The computational overhead of fine-tuning the base model on downstream tasks is very high. At the same time, deploying the fine-tuned base model on downstream tasks often requires computing resources that exceed the resource limits of the downstream tasks. Therefore, improving the training efficiency of the base model during fine-tuning on downstream tasks and the inference efficiency of the fine-tuned base model has become an urgent problem to be solved.
[0003] Existing technical means mainly focus on parameter-efficient fine-tuning (PEFT) and model compression. PEFT achieves the same effect as full fine-tuning by fine-tuning a small number of parameters. Model compression methods directly act on the basic model structure through pruning, knowledge distillation, and model quantization to compress the inference calculation amount of the basic model while maintaining its performance on benchmark tasks as much as possible.
[0004] However, although PEFT effectively reduces the training cost of fine-tuning, most methods inevitably increase the computational complexity of the model in the inference phase, resulting in low inference efficiency of the base model on downstream tasks. Model compression methods are inefficient in training efficiency because they require a large amount of computing resources for retraining after compression to prevent a significant drop in performance. Simply using existing technologies cannot simultaneously solve the two major problems of high training overhead when fine-tuning the base model and low inference efficiency after the base model is deployed, making it impossible to quickly and efficiently deploy the base model on downstream tasks. Summary of the invention
[0005] The present application aims to provide a basic model fine-tuning method, apparatus, device and storage medium with efficient training and reasoning capabilities, at least to solve the problem that the training overhead and reasoning efficiency in the basic model pre-training process of visual tasks cannot be balanced.
[0006] In a first aspect, the present application discloses a basic model fine-tuning method with efficient training and reasoning capabilities, including: Extracting at least one set of target unit data pairs from the unit sequence data to be input into the target multi-head self-attention module; the target unit data pairs include two unit data in the unit sequence data that meet a preset similarity condition; Generating compressed data of the cell sequence data according to the combination of two pieces of the cell data in each group of the target cell data pairs; Inputting the compressed data into the target multi-head self-attention module instead of the cell sequence data.
[0007] In a second aspect, an embodiment of the present application further discloses a basic model fine-tuning device with efficient training and inference capabilities, including: An extraction module, configured to extract at least one group of target cell data pairs from the cell sequence data to be input into the target multi-head self-attention module; the two pieces of cell data in the target cell data pair satisfy a preset similarity condition in the cell sequence data; A compression module, configured to generate compressed data of the cell sequence data according to the combination of two pieces of the cell data in each group of the target cell data pairs; An input module, configured to input the compressed data into the target multi-head self-attention module instead of the cell sequence data.
[0008] In a third aspect, an embodiment of the present application further discloses an electronic device, including a processor and a memory, where the memory stores a program or instruction that can run on the processor, and when the program or instruction is executed by the processor, the steps of the method described in the first aspect are implemented.
[0009] In a fourth aspect, an embodiment of the present application further discloses a readable storage medium, where a program or instruction is stored on the readable storage medium, and when the program or instruction is executed by a processor, the steps of the method described in the first aspect are implemented.
[0010] In summary, in the embodiment of the present application, by selecting the unit data pairs that meet the preset similarity conditions in the downstream task, effectively focusing on the features with high correlation, it is ensured that in the subsequent feature merging and processing steps, the model performance is optimized in a targeted manner, and the targeted unit data selection is used to reduce unnecessary computational overhead and improve the efficiency of fine-tuning; and then the two unit data in each group of target unit data pairs are merged to generate compressed data, so that the model can reduce the amount of calculation in the processing process, improve the reasoning efficiency, and make the compressed data in the merging process retain key information, ensuring that the model can still maintain high performance and accuracy in a compressed state; finally, the compressed data is input into the target multi-head self-attention module instead of the original unit sequence data, so that the computational complexity of the model in the reasoning stage is significantly reduced, which not only reduces the consumption of computing resources in the fine-tuning process, but also improves the reasoning speed of the model in the downstream task deployment. Thus, based on the method of the embodiment of the present application, by extracting and merging the target unit data pairs, and optimizing the feature processing flow, the performance and efficiency of the model in practical applications are improved, ensuring that while ensuring the training efficiency, the reasoning efficiency of the model is improved, and finally realizing the efficient deployment of the basic model on downstream tasks. This solves the problem of the inability to balance the training overhead and inference efficiency in the pre-training process of the basic model for visual tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In the attached picture: Figure 1 It is a flowchart of the steps of a basic model fine-tuning method with efficient training and reasoning capabilities provided in an embodiment of the present application; Figure 2 It is a flowchart of the steps of another basic model fine-tuning method with efficient training and reasoning capabilities provided in an embodiment of the present application; Figure 3 is a compression ratio-performance comparison diagram between two embodiments of the present application; Figure 4 It is a data flow process under the embodiment of the present application; Figure 5 It is a block diagram of a basic model fine-tuning device with efficient training and reasoning capabilities provided in an embodiment of the present application; Figure 6 is a block diagram of an electronic device according to an embodiment of the present application; Figure 7 It is a block diagram of an electronic device of another embodiment provided by the embodiments of the present application. DETAILED DESCRIPTION
[0012] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0013] The terms "first", "second", etc. in the specification and claims of the present application are used to distinguish similar objects, rather than to describe a specific order or sequence. It should be understood that such data used can be interchanged under appropriate circumstances so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually of the same category, and the number of objects is not limited. For example, the first object can be one or multiple. In addition, "and / or" in the specification and claims means at least one of the connected objects, and the character " / " generally means that the related objects before and after are in an "or" relationship.
[0014] In the present application, we mainly focus on the efficient downstream task transfer of training and inference for the Vision Transformer (ViT). A ViT model contains L identical encoders, and each encoder consists of a multi-head self-attention (MHSA) module and a feed-forward network (FFN).
[0015] Formally, the input image is reshaped and linearly projected into a sequence unit data of dimensions. (For simplicity of expression, the description of the classification unit and the distillation unit is omitted here). Then for the encoder , the input is represented as , and the output is represented as ; for the MHSA, the input units are first processed by three fully connected layers to generate , and matrices, and the output is processed by another fully connected layer after being calculated by to obtain the input of the subsequent feed-forward neural network (FFN); for the FFN, its input units are processed by two fully connected layers to obtain the corresponding output. The method disclosed in the present application mainly focuses on the calculation stage before the input unit is sent into the MHSA module.
[0016] For the above-mentioned ViT basic vision model, when , and add a low-rank matrix structure (Low-Rank Adaptation, LoRA) to the projection matrix. When fine-tuning the ViT model on downstream tasks, the update amount of the above matrix in the model actually lies in the subspace of the original matrix. Therefore, the update amount of these matrices during the fine-tuning process can be directly modeled using low-rank decomposition, and only the branch matrix modeled by the low-rank matrix is trained during the fine-tuning process.
[0017] Specifically, for the dense matrix and the input , the forward propagation method of the updated is: , where , . During the inference process, can be merged with to avoid additional computational overhead in the fine-tuned base model.
[0018] As Figure 1 shown, it is a method for fine-tuning a base model with efficient training and inference capabilities provided in this embodiment. It is mainly used for the ViT vision base model with the LoRA structure added above, and uses the sequence unit merging paradigm for parameter-free network compression. Such methods are orthogonal to the Transformer structure and can flexibly change the compression rate.
[0019] The method may include the following steps: Step 101, extract at least one set of target unit data pairs from the unit sequence data to be input into the target multi-head self-attention module.
[0020] Among them, the target unit data pair contains two unit data in the unit sequence data that meet the preset similarity conditions.
[0021] In some embodiments of the present application, in order to effectively focus on features with higher relevance and optimize the model performance, two unit data that meet the preset similarity conditions will be selected by calculating the similarity between units in the unit sequence data to form a target unit data pair. The similarity conditions may include metrics such as Euclidean distance and cosine similarity to ensure that the selected unit data pair has high similarity in the feature space. After extracting the target unit data pair, unnecessary computational overhead can be reduced and the efficiency of fine-tuning can be improved.
[0022] In a specific example, when dealing with an image classification task, target unit data pairs are extracted from the unit sequence data in the multi-head self-attention module of a pre-trained base model. First, the similarity between each pair of units in the unit sequence data can be calculated, and then the unit data pairs with a similarity higher than 0.8 are selected as the target unit data pairs. In this way, it is possible to effectively focus on the relatively relevant image features and improve the fine-tuning efficiency of the model.
[0023] Step 102: Generate compressed data of the unit sequence data according to the combination of the two unit data in each group of target unit data pairs.
[0024] In some embodiments of the present application, in order to reduce the computational amount during the model processing and improve the inference efficiency, the two unit data in the target unit data pairs are combined, so that the key information can be retained and compressed data can be generated. The compressed data in the combination process not only contains the key information of the original data, but also reduces the data redundancy, enabling the model to perform inference calculations efficiently in subsequent processing steps.
[0025] In a specific example, when dealing with an image classification task, the two unit data in the target unit data pairs extracted from the multi-head self-attention module of the pre-trained base model are combined. For example, the compressed data can be generated by calculating the average value of these unit data. In this way, it is possible to effectively reduce the computational amount when the model processes image features, improve the inference efficiency, and ensure that the model can still maintain a high classification accuracy in the compressed state.
[0026] Step 103: Input the compressed data instead of the unit sequence data into the target multi-head self-attention module.
[0027] In some embodiments of the present application, in order to complete the final fine-tuning training, the previously generated compressed data is input into the target multi-head self-attention module, reducing the computational resources required during the inference process while maintaining the original model performance. Specifically, this replacement operation ensures that the compressed data can effectively participate in the multi-head self-attention mechanism of the model, thereby achieving efficient inference.
[0028] In a specific example, when dealing with an image classification task, the compressed data obtained by the previous combination is input into the multi-head self-attention module of the pre-trained base model instead of the original unit sequence data. In this way, it is possible to significantly reduce the computational complexity in the inference stage, while ensuring that the accuracy and performance of the model in the classification task do not decrease significantly. Finally, while improving the inference speed of the model, the goal of efficiently deploying downstream tasks is achieved.
[0029] In summary, in the embodiment of the present application, by selecting the unit data pairs that meet the preset similarity conditions in the downstream task, effectively focusing on the features with high correlation, it is ensured that in the subsequent feature merging and processing steps, the model performance is optimized in a targeted manner, and the targeted unit data selection is used to reduce unnecessary computational overhead and improve the efficiency of fine-tuning; and then the two unit data in each group of target unit data pairs are merged to generate compressed data, so that the model can reduce the amount of calculation in the processing process, improve the reasoning efficiency, and make the compressed data in the merging process retain key information, ensuring that the model can still maintain high performance and accuracy in a compressed state; finally, the compressed data is input into the target multi-head self-attention module instead of the original unit sequence data, so that the computational complexity of the model in the reasoning stage is significantly reduced, which not only reduces the consumption of computing resources in the fine-tuning process, but also improves the reasoning speed of the model in the downstream task deployment. Thus, based on the method of the embodiment of the present application, by extracting and merging the target unit data pairs, and optimizing the feature processing flow, the performance and efficiency of the model in practical applications are improved, ensuring that while ensuring the training efficiency, the reasoning efficiency of the model is improved, and finally realizing the efficient deployment of the basic model on downstream tasks. This solves the problem of the inability to balance the training overhead and inference efficiency in the pre-training process of the basic model for visual tasks.
[0030] Figure 2 This is another basic model fine-tuning method with efficient training and reasoning capabilities provided in the embodiments of the present application.
[0031] The method may include the following steps: Step 201, determining a plurality of unit data pairs from unit data of unit sequence data.
[0032] Among them, each unit data pair has a corresponding similarity evaluation value.
[0033] In some embodiments of the present application, in order to effectively evaluate the similarity between units so as to subsequently select and process unit data pairs with higher correlation, multiple unit data pairs with similarity evaluation values are determined by calculating the similarity between each pair of unit data. The similarity evaluation value can be calculated by a variety of methods, such as Euclidean distance, cosine similarity, etc. After determining multiple unit data pairs, the data selection and processing efficiency in subsequent steps can be improved.
[0034] In a specific example, when processing an image recognition task, multiple pairs of unit data are determined from the unit sequence data. For example, the cosine similarity between each pair of unit data can be calculated first, and a corresponding similarity evaluation value can be assigned to each pair of unit data. In this way, the similarity between the unit data can be effectively evaluated, and a basis can be provided for data selection and processing in subsequent steps.
[0035] Optionally, step 201 includes the following sub-steps: Sub-step 2011: Divide the unit data in the unit sequence data into reference unit data and control unit data, and calculate the similarity evaluation value between each reference unit data and each control unit data respectively.
[0036] In some embodiments of the present application, in order to make the process of evaluating the similarity between unit data not lose generality, the unit sequence data is divided into two groups, and the similarity evaluation value between each reference unit data and each control unit data is calculated respectively to evaluate the similarity between different unit data pairs. The similarity evaluation value can be calculated by various methods, and this process ensures that each unit data pair has a high similarity in the feature space, thereby providing a basis for data selection and processing in subsequent steps.
[0037] In a specific example, when processing an image classification task, the unit data in the unit sequence data of the pre-trained base model can be randomly divided into two groups: reference unit data and control unit data. First, calculate the cosine similarity between each reference unit data and each control unit data as the corresponding similarity evaluation value. In this way, the similarity between unit data can be effectively evaluated, and a basis for data selection and processing in subsequent steps can be provided.
[0038] Suppose in the input units of the th ViT layer of the Encoder, before MHSA, it is randomly divided into two groups: , , where . And assume that the most similar unit is matched for each unit through cosine similarity. Then there is: .
[0039] Sub-step 2012: Respectively determine the control unit data with the maximum similarity evaluation value to each reference unit data as the paired unit data corresponding to the reference unit data.
[0040] In some embodiments of the present application, in order to accurately match the most relevant unit data pairs and improve the efficiency and accuracy of subsequent processing steps, the similarity evaluation value between each reference unit data and all control unit data is calculated, and then the control unit data with the maximum similarity evaluation value is selected as the paired unit data of the reference unit data. This can ensure that each reference unit data can be matched with the most similar control unit data, thereby improving the effect of data merging in subsequent processing steps.
[0041] In a specific example, when dealing with an image classification task, by calculating the cosine similarity between each reference unit data and all control unit data, the control unit data with the highest similarity evaluation value is selected as the paired unit data for each reference unit data. For example, if the similarity evaluation value between reference unit data A and control unit data B is 0.95, and the similarity evaluation values between A and other control unit data are all lower than 0.95, then control unit data B is determined as the paired unit data for reference unit data A. In this way, it can be ensured that each reference unit data can be matched with the most relevant control unit data, improving the performance of the model in subsequent processing steps.
[0042] Continuing with the example in sub-step 2011, then one can select from the units with the highest similarity to the units in to form unit pairs where , , and . Then the paired unit of .
[0043] In sub-step 2013, the corresponding reference unit data and paired unit data are determined as a unit data pair, and the similarity evaluation value between the reference unit data and the paired unit data in the unit data pair is determined as the similarity evaluation value of the unit data pair.
[0044] In some embodiments of the present application, in order to clarify the similarity of each unit data pair and thus optimize the model performance in subsequent processing steps, the previously determined reference unit data and paired unit data are paired, and their similarity evaluation values are recorded to form unit data pairs. This process ensures that each unit data pair has a clear similarity in the feature space, providing a basis for data processing and merging in subsequent steps.
[0045] In a specific example, when dealing with an image classification task, the previously calculated reference unit data A and paired unit data B are determined as a unit data pair, and the similarity evaluation value 0.95 between A and B is recorded as the similarity evaluation value of this unit data pair. In this way, the experimenter can clarify the similarity of each unit data pair and provide a basis for subsequent feature merging and processing steps.
[0046] Step 202, according to the descending order of the similarity evaluation values of all unit data pairs, determine a preset number of target unit data pairs from the unit data pairs.
[0047] In some embodiments of the present application, in order to screen out a sufficient number of target data pairs from a large number of unit data pairs and further optimize the model performance, the similarity evaluation values of the unit data pairs will be sorted, and several unit data pairs with higher rankings will be selected as the target unit data pairs. This process ensures that the selected target unit data pairs can represent the key features in the unit sequence data, thereby improving the overall performance of the model in subsequent processing steps.
[0048] In a specific example, when processing an image classification task, the similarity of all unit data pairs in the unit sequence data is evaluated and arranged in descending order of the similarity evaluation values. Then, according to a preset sampling quantity, for example, the first 10 unit data pairs, the corresponding number of target unit data pairs is selected. In this way, the most relevant target unit data pairs can be screened out from a large number of unit data pairs, ensuring the improvement of the model performance in subsequent feature merging and processing steps.
[0049] Step 203: Generate the compressed data of the unit sequence data according to the combination of the two unit data in each group of target unit data pairs.
[0050] The method shown in this step has been described in step 102 and will not be elaborated here.
[0051] Optionally, in a simple embodiment, data compression is achieved through average pooling combination. At this time, step 203 includes the following sub-steps: Sub-step 2031: Combine the two unit data in each target unit data pair through average pooling respectively to obtain the combined unit data of each target unit data pair.
[0052] In some embodiments of the present application, in order to reduce data redundancy in the model processing process and improve the inference efficiency, the two unit data in each target unit data pair will be averaged through average pooling, which can retain key information and reduce the data volume. Average pooling is a calculation method that generates a new data point by taking the average of two data. After executing this step, the generated combined unit data helps to optimize the model performance in subsequent processing steps, ensuring high accuracy even in the compressed state.
[0053] In a specific example, when processing an image classification task, the two unit data in the target unit data pair of the pre-trained base model are combined through average pooling. For example, for the target unit data pairs A and B, their average value is calculated to generate the combined unit data C. In this way, data redundancy in the model processing process can be effectively reduced and the inference efficiency can be improved.
[0054] Sub-step 2032: Combine all the merged cell data into a merged cell matrix of cell sequence data, and determine the merged cell matrix as the compressed data of the cell sequence data.
[0055] In some embodiments of the present application, in order to further simplify the data structure and improve the processing efficiency of the model on the basis of the merged cell data, all the merged cell data will be integrated into a matrix to form a merged cell matrix of the new cell sequence data, and this matrix will be used as the compressed data and input into the model. The merged cell matrix is a data structure used to represent the compressed cell data.
[0056] In a specific example, when processing an image classification task, all the merged cell data generated by average pooling before are integrated into a matrix. For example, the merged cell data C, D, and E are integrated into a merged cell matrix M, and this matrix is determined as the compressed data of the cell sequence data. In this way, the data structure can be effectively simplified and the computational complexity of the model during the inference process can be reduced.
[0057] Optionally, as Figure 3 shown, the horizontal axis in the figure is the throughput of the model (unit: samples / sec), and the vertical axis is the inference accuracy index of the model (unit: %). Considering that at a relatively low compression ratio (about 1.7 times), the performance of two ViT backbones (ViT-L / 16 and ViT-B / 16) shows a slight decrease (close to 1%) compared with directly performing parameter-efficient model fine-tuning, such a compression effect is unacceptable for the current state-of-the-art model compression algorithms; while at a high compression ratio (greater than 3.0 times), the performance of the baseline scheme drops rapidly, and its performance is even worse than directly performing parameter-efficient model fine-tuning training on a small-scale ViT with the corresponding throughput.
[0058] Through analysis, it is found that the traditional full-scale fine-tuning method guides the model to dynamically align with the target data distribution of the downstream task by fine-tuning a large number of parameters. However, in the context of efficient training-inference downstream task migration, in order to achieve the goal of efficient training, the model can only fine-tune a small number of parameters far less than the number of parameters of the model itself, which poses a major challenge in accurately capturing the subtle differences in the data distributions of the pre-training task and the downstream task. Although the sequence cell merging paradigm is highly effective in reducing model complexity and improving inference efficiency, it introduces a risk of information loss in the process of the ViT processing information layer by layer. According to the previous analysis, due to the limitation of parameter-efficient model fine-tuning on the understanding of the data distribution, this loss is difficult to correct under the condition of efficient fine-tuning. Therefore, directly combining the sequence cell merging and the PEFT algorithm to achieve efficient training-inference downstream task migration may not produce the optimal result.
[0059] Therefore, the Parallel Yielding Re-Activation (PYRA) method is proposed, which adaptively modulates the features of sequence units and enhances the model's perception of data distribution during the unit merging process. Specifically, a pair of lightweight learnable vectors in each ViT model layer are first used to generate weights for adaptive merging in parallel. Then, these generated weights are applied to the sequence units to be merged through re-activation for modulation. PYRA adaptively calibrates the feature distribution of sequence units with low computational complexity. At this time, step 203 includes the following sub-steps: In sub-step 2033, the reference unit data in all target unit data pairs are formed into a reference unit matrix of the unit sequence data, and the control unit data in all target unit data pairs are formed into a control unit matrix of the unit sequence data.
[0060] In some embodiments of the present application, in order to effectively integrate unit data and improve data processing efficiency and accuracy, the reference unit data in all target unit data pairs are integrated into a matrix to form a reference unit matrix. At the same time, all control unit data are integrated into a matrix to form a control unit matrix. The reference unit matrix and the control unit matrix respectively represent the two data sets before merging. This process ensures the consistency and integrity of the data during the data integration process, thus providing a basis for subsequent modulation and compression operations.
[0061] In a specific example, when processing an image classification task, the reference unit data in all target unit data pairs extracted from the pre-trained base model are integrated into a matrix to form a reference unit matrix. At the same time, all control unit data are integrated into a matrix to form a control unit matrix. For example, the reference unit data in the target unit data pairs form part of the reference unit matrix, while the control unit data in the target unit data pairs form part of the control unit matrix. In this way, unit data can be effectively integrated, and data processing efficiency and accuracy can be improved.
[0062] Formally, for a ViT layer with the sequence units to be merged with dimensions and can be grouped and merged into a sequence unit matrix and where .
[0063] In sub-step 2034, an information matrix of the unit sequence data, as well as the first decoupling weight and the second decoupling weight of the unit sequence data, are determined based on the reference unit matrix and the control unit matrix.
[0064] The first decoupling weight is the decoupling weight of the unit sequence data to the length of the unit data in the unit sequence data, and the second decoupling weight is the decoupling weight of the unit sequence data to the compressed length of the unit sequence data.
[0065] In some embodiments of the present application, in order to improve the model's perception of data distribution through precise weight calculation and feature modulation, the reference unit matrix and the control unit matrix are operated to calculate the information matrix; then, the first decoupling weight and the second decoupling weight of the unit sequence data are calculated according to the information matrix. The information matrix represents the feature information after the reference and control unit data are combined, the first decoupling weight is the weight for modulating the length of the unit data, and the second decoupling weight is the weight for modulating the length of the compressed unit data. This process ensures that when performing data modulation, the data distribution in the downstream task can be fully perceived, thereby improving model performance.
[0066] In a specific example, when processing an image classification task, the information matrix is first calculated by using a reference unit matrix and a control unit matrix. For example, the information matrix can be obtained by a linear operation of and. Then, the first decoupling weight and the second decoupling weight are calculated based on the information matrix. The first decoupling weight represents the weight for modulating the unit data length, while the second decoupling weight represents the weight for modulating the unit data compression length. In this way, the model's perception of data distribution can be improved, providing a basis for subsequent feature modulation.
[0067] Following the example of step 2033, a modulation matrix can be learned , to adaptively modulate the sequence unit at the granularity of each channel. In fact, here we directly learn a fixed There are many redundant parameters and fixed It is not possible to adaptively correct the different features of sequence unit pairs in different images. Therefore, the decoupled weights of feature channels can be learned separately. and the decoupling weights of the feature sequence units and through The modulation matrix is generated in the form of low-rank decomposition. And it is generated in parallel and , for the layer To be merged Sequence units create two learnable vectors as modulation weight generators inside the encoder block of the ViT layer: and To achieve the above process.
[0068] Optionally, sub-step 2034 includes the following sub-steps: Sub-step 20341: Normalize the sum of the reference unit matrix and the control unit matrix according to a preset normalization function to obtain an information matrix.
[0069] In some embodiments of the present application, in order to normalize the input data before feature modulation, ensure the balanced data distribution, and improve the training and inference effects of the model, by normalizing the sum of the reference unit matrix and the control unit matrix, the difference of components in the data can be eliminated and the data distribution can be smoothed. The normalization process is a commonly used data preprocessing method that scales the input data to a specific range to reduce the scale difference between different features. The obtained information matrix will be used for subsequent feature modulation to ensure that the model can better learn and adapt to the data distribution.
[0070] In a specific example, when dealing with an image classification task, normalize the sum of the reference unit matrix and the control unit matrix according to a preset normalization function. For example, use the LayerNorm function to normalize the sum of the reference unit matrix and the control unit matrix to obtain an information matrix. In this way, the difference of data components in the data can be eliminated and the data distribution can be made more balanced.
[0071] Continuing from the example of sub-step 2033, to ensure and adaptive extraction of feature information from two units in the current encoded sequence unit pair, the information matrix of the current sequence unit to be merged can be calculated first: , the distribution of can be normalized by using operation to obtain a smoother gradient during training and .
[0072] Sub-step 20342: Determine the matrix product of the information matrix and the pre-trained first modulation vector as the first decoupling weight, and determine the matrix product of the pre-trained second modulation vector and the information matrix as the second decoupling weight.
[0073] In some embodiments of the present application, in order to achieve precise feature modulation through pre-trained modulation vectors and improve the adaptability of the model to the data distribution, the matrix product of the information matrix and the pre-trained first modulation vector is calculated to obtain the first decoupling weight; at the same time, the matrix product of the pre-trained second modulation vector and the information matrix is calculated to obtain the second decoupling weight. The first decoupling weight is used to modulate the length of the unit data, and the second decoupling weight is used to modulate the compressed length of the unit data. This process ensures that during the feature modulation process, the data distribution can be effectively adjusted and the learning ability of the model can be enhanced.
[0074] In a specific example, when dealing with an image classification task, the information matrix is multiplied by a pre-trained first modulation vector to obtain a first decoupled weight. At the same time, the pre-trained second modulation vector is multiplied by the information matrix to obtain a second decoupled weight. In this way, it can be ensured that the modulated data can accurately reflect the feature distribution of the original data and improve the adaptability of the data.
[0075] Continuing with the example of sub-step 20341, given the sequence unit information matrix, the adaptive weights can be generated simultaneously in parallel. and : , .
[0076] In sub-step 2035, the compressed data of the unit sequence data is determined by modulating the reference unit matrix with the first decoupled weight and the second decoupled weight.
[0077] In some embodiments of the present application, in order to precisely modulate the unit features and improve the adaptability of the data distribution, the reference unit matrix will be modulated by using the first decoupled weight and the second decoupled weight to ensure that the modulated data can accurately reflect the feature distribution of the original data. The decoupled weight is the weight used to adjust the feature distribution during the modulation process. The first decoupled weight and the second decoupled weight are respectively used to modulate the length and the compressed length of the unit data. The compressed data generated in this way will be used for subsequent model inference to improve the processing efficiency and model performance.
[0078] In a specific example, when dealing with an image classification task, the reference unit matrix is modulated using the previously calculated first decoupled weight and second decoupled weight. For example, by adjusting the weight of each unit data in, the modulated combined unit data is generated. In this way, it can be ensured that the modulated data can accurately reflect the feature distribution of the original data and improve the adaptability of the data.
[0079] Continuing with the example of sub-step 2034, simply generating through matrix multiplication and still faces some possible problems. First, there are no measures to ensure that and maintain their values within the normal range. Second, simply decoupling the weights into the sequence unit dimension and the channel dimension will result in the modulation weight matrix having low rank and limited expressive ability, so it may not be able to optimally modulate the sequence units in complex data distributions. To address these problems, a reactivation strategy can be adopted for sequence unit modulation.
[0080] Optionally, sub-step 2035 includes the following sub-steps: Sub-step 20351: By activating the first decoupled weight during the broadcast process, determine the first corrected decoupled weight of the first decoupled weight, and modulate the reference unit matrix with the first corrected decoupled weight to obtain the first modulation matrix of the reference unit matrix.
[0081] In some embodiments of the present application, in order to optimize the use of weights during the feature modulation process and improve the modulation accuracy of data, the first corrected decoupled weight will be obtained by activating the first decoupled weight during the broadcast process; then, the reference unit matrix will be modulated with the corrected weight to generate the first modulation matrix of the reference unit matrix. The activation process maps the weight value to a specified range through a specific function (such as the sigmoid function). This process ensures that the weights can accurately adjust the data feature distribution during feature modulation, thereby improving the performance and adaptability of the model.
[0082] In a specific example, when processing an image classification task, activate the calculated first decoupled weight, for example, use the sigmoid function to activate it to obtain the first corrected decoupled weight. Then, modulate the reference unit matrix with the first corrected decoupled weight to generate the first modulation matrix of the reference unit matrix. In this way, the use of weights can be optimized and the modulation accuracy of data can be improved.
[0083] Continuing with the example of sub-step 2034, we can first broadcast to and perform sigmoid activation on : , and then modulate with the activated : , where represents the Hadamard product.
[0084] Sub-step 20352: By activating the second decoupled weight during the broadcast process, determine the second corrected decoupled weight of the second decoupled weight, and modulate the first modulation matrix with the second corrected decoupled weight to obtain the second modulation matrix of the reference unit matrix.
[0085] In some embodiments of the present application, in order to further optimize the feature modulation process and ensure the accuracy of data distribution, the second corrected decoupled weight will be obtained by activating the second decoupled weight during the broadcast process; then, the first modulation matrix will be modulated with the corrected weight to generate the second modulation matrix of the reference unit matrix. This process ensures that the weights can more accurately adjust the data feature distribution during feature modulation, thereby improving the performance and adaptability of the model.
[0086] In a specific example, when dealing with an image classification task, activation processing is performed on the calculated second decoupled weight. For example, the sigmoid function is used to activate it to obtain the second corrected decoupled weight. Then, the experimenter modulates the first modulation matrix using the second corrected decoupled weight to generate the second modulation matrix of the reference cell matrix. In this way, the use of weights can be further optimized, and the modulation accuracy of data can be improved.
[0087] Sub-step 20353: Correct the reference cell matrix through the second modulation matrix, and determine the corrected reference cell matrix as the compressed data of the cell sequence data.
[0088] In some embodiments of the present application, in order to ensure the accuracy and consistency of data features and improve the processing efficiency of the model, by correcting the reference cell matrix using the second modulation matrix, the data feature distribution can be further optimized, and the quality and accuracy of the compressed data can be ensured. The corrected reference cell matrix represents the adjusted data features and can better reflect the feature distribution of the original data. The compressed data generated by executing this step will be used for subsequent model inference, reducing the computational overhead and improving the processing efficiency.
[0089] In a specific example, when dealing with an image classification task, the previously generated second modulation matrix is used to correct the reference cell matrix. For example, each cell data in the reference cell matrix is adjusted through the second modulation matrix to generate the corrected reference cell matrix. Then, the corrected reference cell matrix is determined as the compressed data of the cell sequence data. In this way, the accuracy and consistency of data features can be ensured, and the processing efficiency of the model can be improved.
[0090] Continuing with the example of sub-step 20351, after being modulated by the modulated weight, the broadcast weight activated by sigmoid can be used again to modulate to obtain the modulated sequence units: , In the above formula, the original sequence unit matrix is used to create a residual connection to maintain the gradient flow during training; then the modulated sequence units are merged with through average pooling. The generator is initialized using a random Gaussian distribution and the generator is initialized with zero values , so the reactivation is equivalent to the identity transformation at the beginning of training. Since the sequence unit modulation is only performed on , the parallelism in training and inference can be guaranteed (this is because during the sequence feature merging process The sequence units are unique, while different may point to the same , that is, different may actually point to the same unit object).
[0091] Step 204, input the compressed data instead of the unit sequence data into the target multi-head self-attention module.
[0092] The method shown in this step has been described in step 102 and will not be elaborated here.
[0093] As Figure 4 shown, it is a data flow process in a round corresponding to an Encoder of this application: Step S0: First, train a decoupling weight generator for decoupling the features between the reference unit matrix and the control unit matrix. The decoupling weight generator generates weights in an adaptive manner for use in subsequent steps; Step S1: Add the elements in the reference unit matrix and the control unit matrix item by item to generate a preliminary comprehensive matrix, which helps to integrate the features of the two groups of data and lay a foundation for subsequent normalization processing; Step S2: Perform normalization processing on the comprehensive matrix to eliminate the dimension difference, smooth the data distribution, ensure the consistency of the data in subsequent processing, and obtain an information matrix; Step S2.1: Use the information matrix to perform matrix multiplication with a pre-trained first modulation vector to generate a first decoupling weight; Step S2.2: Use a pre-trained second modulation vector to perform matrix multiplication with the information matrix to generate a second decoupling weight; Step S3.1: Ensure that it can accurately adjust the data feature distribution during the feature modulation process by performing broadcast activation processing on the first decoupling weight; Step S3.2: Modulate the reference unit matrix with the first corrected decoupling weight to generate a first modulation matrix of the reference unit matrix; Step S4.1: Ensure that it can accurately adjust the data feature distribution during the feature modulation process by performing broadcast activation processing on the second decoupling weight; Step S4.2: Modulate the first modulation matrix with the second corrected decoupling weight to generate a second modulation matrix of the reference unit matrix; Step S5: Correct the reference unit matrix through the second modulation matrix to generate a corrected reference unit matrix, and determine it as the compressed data of the unit sequence data. These compressed data will be combined with the control unit matrix and then can be used for subsequent model inference to improve the processing efficiency and model performance.
[0094] In summary, in the embodiment of the present application, by selecting the unit data pairs that meet the preset similarity conditions in the downstream task, effectively focusing on the features with high correlation, it is ensured that in the subsequent feature merging and processing steps, the model performance is optimized in a targeted manner, and the targeted unit data selection is used to reduce unnecessary computational overhead and improve the efficiency of fine-tuning; and then the two unit data in each group of target unit data pairs are merged to generate compressed data, so that the model can reduce the amount of calculation in the processing process, improve the reasoning efficiency, and make the compressed data in the merging process retain key information, ensuring that the model can still maintain high performance and accuracy in a compressed state; finally, the compressed data is input into the target multi-head self-attention module instead of the original unit sequence data, so that the computational complexity of the model in the reasoning stage is significantly reduced, which not only reduces the consumption of computing resources in the fine-tuning process, but also improves the reasoning speed of the model in the downstream task deployment. Thus, based on the method of the embodiment of the present application, by extracting and merging the target unit data pairs, and optimizing the feature processing flow, the performance and efficiency of the model in practical applications are improved, ensuring that while ensuring the training efficiency, the reasoning efficiency of the model is improved, and finally realizing the efficient deployment of the basic model on downstream tasks. This solves the problem of the inability to balance the training overhead and inference efficiency in the pre-training process of the basic model for visual tasks.
[0095] refer to Figure 5 , which shows a basic model fine-tuning device 30 with efficient training and reasoning capabilities provided by an embodiment of the present application, including: An extraction module 301 is used to extract at least one set of target unit data pairs from the unit sequence data to be input into the target multi-head self-attention module; the target unit data pairs include two unit data in the unit sequence data that meet a preset similarity condition; A compression module 302, configured to generate compressed data of the unit sequence data according to the merging of two unit data in each set of target unit data pairs; The input module 303 is used to input the compressed data instead of the unit sequence data into the target multi-head self-attention module.
[0096] Optionally, the extraction module 301 includes: A data pair submodule, used to determine a plurality of unit data pairs from the unit data of the unit sequence data; each unit data pair has a corresponding similarity evaluation value; The target data pair submodule is used to determine a preset sampling number of target unit data pairs from the unit data pairs according to the descending arrangement order of the similarity evaluation values of all unit data pairs.
[0097] Optionally, the data pair submodule includes: An evaluation unit for dividing the unit data in the unit sequence data into reference unit data and control unit data, and calculating the similarity evaluation values between each reference unit data and each control unit data respectively; An extraction unit for respectively determining the control unit data with the maximum similarity evaluation value to each reference unit data as the paired unit data corresponding to the reference unit data; A pairing unit for determining the corresponding reference unit data and the paired unit data as a unit data pair, and determining the similarity evaluation value between the reference unit data and the paired unit data in the unit data pair as the similarity evaluation value of the unit data pair.
[0098] Optionally, the compression module 302 includes: An average pooling module for respectively merging the two unit data in each target unit data pair through average pooling to obtain the merged unit data of each target unit data pair; A first compression sub-module for forming the merged unit matrix of the unit sequence data from all the merged unit data, and determining the compressed data of the unit sequence data as the merged unit matrix.
[0099] Optionally, each target unit data pair respectively includes reference unit data and control unit data, and the compression module 302 includes: A matrix generation sub-module for forming the reference unit matrix of the unit sequence data from the reference unit data in all the target unit data pairs, and forming the control unit matrix of the unit sequence data from the control unit data in all the target unit data pairs; An intermediate parameter sub-module for determining the information matrix of the unit sequence data, as well as the first decoupling weight and the second decoupling weight of the unit sequence data according to the reference unit matrix and the control unit matrix; the first decoupling weight is the decoupling weight of the unit sequence data with respect to the length of the unit data in the unit sequence data, and the second decoupling weight is the decoupling weight of the unit sequence data with respect to the compressed length of the unit sequence data; A second compression sub-module for determining the compressed data of the unit sequence data through the modulation of the reference unit matrix by the first decoupling weight and the second decoupling weight.
[0100] Optionally, the intermediate parameter sub-module includes: An information matrix unit for normalizing the sum of the reference unit matrix and the control unit matrix according to a preset normalization function to obtain the information matrix; A decoupling weight unit for determining the matrix product of the information matrix and the pre-trained first modulation vector as the first decoupling weight, and determining the matrix product of the pre-trained second modulation vector and the information matrix as the second decoupling weight.
[0101] Optionally, the second compression submodule includes: A first modulation unit, configured to determine a first modified decoupling weight of the first decoupling weight by activating the first decoupling weight during the broadcasting process, and modulate the reference unit matrix by the first modified decoupling weight to obtain a first modulation matrix of the reference unit matrix; A second modulation unit, configured to determine a second modified decoupling weight of the second decoupling weight by activating the second decoupling weight during the broadcasting process, and modulate the first modulation matrix by the second modified decoupling weight to obtain a second modulation matrix of the reference unit matrix; The compression unit is used to correct the reference unit matrix through the second modulation matrix, and determine the corrected reference unit matrix as the compressed data of the unit sequence data.
[0102] In summary, in the embodiment of the present application, by selecting the unit data pairs that meet the preset similarity conditions in the downstream task, effectively focusing on the features with high correlation, it is ensured that in the subsequent feature merging and processing steps, the model performance is optimized in a targeted manner, and the targeted unit data selection is used to reduce unnecessary computational overhead and improve the efficiency of fine-tuning; and then the two unit data in each group of target unit data pairs are merged to generate compressed data, so that the model can reduce the amount of calculation in the processing process, improve the reasoning efficiency, and make the compressed data in the merging process retain key information, ensuring that the model can still maintain high performance and accuracy in a compressed state; finally, the compressed data is input into the target multi-head self-attention module instead of the original unit sequence data, so that the computational complexity of the model in the reasoning stage is significantly reduced, which not only reduces the consumption of computing resources in the fine-tuning process, but also improves the reasoning speed of the model in the downstream task deployment. Thus, based on the method of the embodiment of the present application, by extracting and merging the target unit data pairs, and optimizing the feature processing flow, the performance and efficiency of the model in practical applications are improved, ensuring that while ensuring the training efficiency, the reasoning efficiency of the model is improved, and finally realizing the efficient deployment of the basic model on downstream tasks. This solves the problem of the inability to balance the training overhead and inference efficiency in the pre-training process of the basic model for visual tasks.
[0103] Reference Figure 6 , the electronic device 500 may include one or more of the following components: a processing component 502 , a memory 504 , a power component 506 , a multimedia component 508 , an audio component 510 , an input / output (I / O) interface 512 , a sensor component 514 , and a communication component 516 .
[0104] The processing component 502 generally controls the overall operation of the electronic device 500, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The processing component 502 may include one or more processors 520 to execute instructions to complete all or part of the steps of the above - mentioned methods. In addition, the processing component 502 may include one or more modules to facilitate the interaction between the processing component 502 and other components. For example, the processing component 502 may include a multimedia module to facilitate the interaction between the multimedia component 508 and the processing component 502.
[0105] The memory 504 is used to store various types of data to support the operation of the electronic device 500. Examples of such data include instructions for any application or method operating on the electronic device 500, contact data, phone book data, messages, pictures, multimedia, etc. The memory 504 can be implemented by any type of volatile or non - volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read - only memory (EEPROM), erasable programmable read - only memory (EPROM), programmable read - only memory (PROM), read - only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0106] The power component 506 provides power to various components of the electronic device 500. The power component 506 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 500.
[0107] The multimedia component 508 includes an interface that provides an output interface between the electronic device 500 and the user. In some embodiments, the interface may include a liquid crystal display (LCD) and a touch panel (TP). If the interface includes a touch panel, the interface can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can sense not only the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 508 includes a front - facing camera and / or a rear - facing camera. When the electronic device 500 is in an operating mode, such as a shooting mode or a multimedia mode, the front - facing camera and / or the rear - facing camera can receive external multimedia data. Each front - facing camera and rear - facing camera can be a fixed optical lens system or have focal length and optical zoom capabilities.
[0108] The audio component 510 is used to output and / or input audio signals. For example, the audio component 510 includes a microphone (MIC) that is used to receive external audio signals when the electronic device 500 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signal can be further stored in the memory 504 or transmitted via the communication component 516. In some embodiments, the audio component 510 further includes a speaker for outputting audio signals.
[0109] The input / output I / O interface 512 provides an interface between the processing component 502 and a peripheral interface module, and the peripheral interface module may be a keyboard, a click wheel, buttons, etc. These buttons may include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0110] The sensor component 514 includes one or more sensors for providing an assessment of the state of the electronic device 500 in various aspects. For example, the sensor component 514 can detect the on / off state of the electronic device 500, the relative positioning of components, such as the display and the keypad of the electronic device 500. The sensor component 514 can also detect a change in the position of the electronic device 500 or a component of the electronic device 500, the presence or absence of user contact with the electronic device 500, the orientation or acceleration / deceleration of the electronic device 500, and the temperature change of the electronic device 500. The sensor component 514 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 514 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 514 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0111] The communication component 516 is used to facilitate communication between the electronic device 500 and other devices in a wired or wireless manner. The electronic device 500 can access a wireless network based on a communication standard, such as WiFi, a carrier network (such as 2G, 3G, 4G, or 5G), or a combination thereof. In an exemplary embodiment, the communication component 516 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 516 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0112] In an exemplary embodiment, the electronic device 500 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for implementing the method provided by the embodiments of the present application.
[0113] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 504 including instructions, and the above instructions can be executed by a processor 520 of the electronic device 500 to complete the above method. For example, the non-transitory storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0114] Figure 7 It is a block diagram of an electronic device 600 according to another embodiment of the present invention. For example, the electronic device 600 may be provided as a server.
[0115] Referring to Figure 7 , the electronic device 600 includes a processing component 622, which further includes one or more processors, and memory resources represented by a memory 632 for storing instructions executable by the processing component 622, such as application programs. The application programs stored in the memory 632 may include one or more modules each corresponding to a set of instructions. In addition, the processing component 622 is configured to execute instructions to perform the method provided by the embodiments of the present application.
[0116] The electronic device 600 may further include a power supply component 626 configured to perform power management of the electronic device 600, a wired or wireless network interface 650 configured to connect the electronic device 600 to a network, and an input / output (I / O) interface 658. The electronic device 600 may operate based on an operating system stored in the memory 632, such as WindowsServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, or the like.
[0117] It should be noted that for the method embodiments of the present application, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present application are not limited by the described action sequence, because according to the embodiments of the present application, certain steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of the present application.
[0118] Those skilled in the art will readily conceive of other embodiments of the present application after considering the specification and practicing the application disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure. The specification and examples are only regarded as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0119] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A basic model fine-tuning method with efficient training and reasoning capabilities, characterized in that: include: Extract at least one set of target unit data pairs from the unit sequence data to be input into the target multi-head self-attention module; The target unit data pair includes two unit data in the unit sequence data that meet a preset similarity condition; Generate compressed data of the unit sequence data according to the merging of two unit data in each group of the target unit data pair; The compressed data is input into the target multi-head self-attention module instead of the unit sequence data.
2. The basic model fine-tuning method with efficient training and reasoning capabilities as claimed in claim 1, characterized in that: The step of extracting at least one set of target unit data pairs from the unit sequence data to be input into the target multi-head self-attention module comprises: Determine a plurality of unit data pairs from the unit data of the unit sequence data; each of the unit data pairs has a corresponding similarity evaluation value; According to the descending arrangement order of the similarity evaluation values of all the unit data pairs, a preset sampling number of target unit data pairs are determined from the unit data pairs.
3. The basic model fine-tuning method with efficient training and reasoning capabilities as claimed in claim 2, characterized in that: The determining a plurality of unit data pairs from the unit data of the unit sequence data comprises: Dividing the unit data in the unit sequence data into reference unit data and control unit data, and respectively calculating a similarity evaluation value between each reference unit data and each control unit data; respectively determining the comparison unit data having the greatest similarity evaluation value with each of the reference unit data as the pairing unit data corresponding to the reference unit data; The corresponding reference unit data and paired unit data are determined as a unit data pair, and the similarity evaluation value between the reference unit data and the paired unit data in the unit data pair is determined as the similarity evaluation value of the unit data pair.
4. The basic model fine-tuning method with efficient training and reasoning capabilities as claimed in claim 1, characterized in that: The step of generating compressed data of the unit sequence data based on merging two of the unit data in each group of the target unit data pair comprises: Merging the two unit data in each of the target unit data pairs by average pooling to obtain merged unit data of each of the target unit data pairs; All of the merged unit data are combined into a merged unit matrix of the unit sequence data, and the merged unit matrix is determined as compressed data of the unit sequence data.
5. The basic model fine-tuning method with efficient training and reasoning capabilities as claimed in claim 1, characterized in that: Each of the target unit data pairs includes reference unit data and control unit data, and the generating of compressed data of the unit sequence data based on the merging of the two unit data in each group of the target unit data pairs includes: The reference cell data in all the target cell data pairs are used to form a reference cell matrix of the cell sequence data, and the control cell data in all the target cell data pairs are used to form a control cell matrix of the cell sequence data; Determine the information matrix of the cell sequence data, and the first decoupling weight and the second decoupling weight of the cell sequence data according to the reference cell matrix and the control cell matrix; the first decoupling weight is the decoupling weight of the cell sequence data to the length of the cell data in the cell sequence data, and the second decoupling weight is the decoupling weight of the cell sequence data to the compressed length of the cell sequence data; The compression data of the unit sequence data is determined by modulating the reference unit matrix with the first decoupling weight and the second decoupling weight.
6. The basic model fine-tuning method with efficient training and reasoning capabilities as claimed in claim 5, characterized in that: The determining the information matrix of the unit sequence data and the first decoupling weight and the second decoupling weight of the unit sequence data according to the reference unit matrix and the comparison unit matrix comprises: Normalizing the sum of the reference unit matrix and the control unit matrix according to a preset normalization function to obtain the information matrix; The matrix product of the information matrix and the pre-trained first modulation vector is determined as the first decoupling weight, and the matrix product of the pre-trained second modulation vector and the information matrix is determined as the second decoupling weight.
7. The basic model fine-tuning method with efficient training and reasoning capabilities as claimed in claim 5, characterized in that: The step of determining the compressed data of the unit sequence data by modulating the reference unit matrix by the first decoupling weight and the second decoupling weight comprises: Determining a first modified decoupling weight of the first decoupling weight by activating the first decoupling weight during the broadcasting process, and modulating the reference unit matrix by the first modified decoupling weight to obtain a first modulation matrix of the reference unit matrix; Determine a second modified decoupling weight of the second decoupling weight by activating the second decoupling weight during the broadcasting process, and modulate the first modulation matrix by the second modified decoupling weight to obtain a second modulation matrix of the reference unit matrix; The reference unit matrix is corrected by using the second modulation matrix, and the corrected reference unit matrix is determined as the compressed data of the unit sequence data.
8. A basic model fine-tuning device with efficient training and reasoning capabilities, characterized in that: include: An extraction module, used to extract at least one set of target unit data pairs from the unit sequence data to be input into the target multi-head self-attention module; The target unit data pair includes two unit data in the unit sequence data that meet a preset similarity condition; A compression module, used for generating compressed data of the unit sequence data according to the combination of two unit data in each group of the target unit data pair; An input module is used to input the compressed data into the target multi-head self-attention module instead of the unit sequence data.
9. An electronic device, characterized in that: include: A processor, a memory for storing instructions executable by the processor; The processor is configured to execute the instructions to implement the steps of the basic model fine-tuning method with efficient training and reasoning capabilities as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: When the instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform the steps of the basic model fine-tuning method with efficient training and reasoning capabilities as described in any one of claims 1 to 7.