Fine-tuning Method, Device, Electronic Device and Storage Medium of Pre-trained Model

By generating and correcting the integration of the small-scale second feature processing network and pre-trained model, the problems of large amount of computation and high storage occupancy in the model fine-tuning process are solved, efficient model fine-tuning is achieved, and pre-training knowledge is retained.

CN116883781BActive Publication Date: 2025-07-11BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310791801.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-06-29
Publication Date
2025-07-11
Estimated Expiration
2043-06-29

AI Technical Summary

Technical Problem

现有技术中,模型微调方式容易导致预训练模型遗忘学习到的知识,且计算量大、存储占用高,难以达到最佳效果。

Method used

By generating a smaller-scale second feature processing network, it is fused with the pre-trained model, and corrected based on the training data of the target task until the target model is obtained.

Benefits of technology

It reduces the computational volume and storage space requirements of the model fine-tuning process, improves processing speed, reduces costs, and retains the knowledge of pre-trained models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116883781B_ABST
    Figure CN116883781B_ABST
Patent Text Reader

Abstract

The present disclosure provides a fine-tuning method, device, electronic device, and storage medium for a pre-trained model, which relates to the field of computer technologies, and particularly to artificial intelligence technologies such as deep learning, natural language processing, and vision technologies. The specific implementation solution is as follows: First, obtain first training data associated with a target task and a pre-trained model, generate a second feature processing network based on the pre-trained model, then, based on a preset rule, fuse the second feature processing network with the pre-trained model to obtain a fused model, then input input data into the fused model to obtain a prediction result output by the fused model, and finally, correct the second feature processing network in the fused model according to the difference between the prediction result and the annotation result until a target model is obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technologies, and in particular to artificial intelligence technologies such as deep learning, natural language processing, and vision technologies. Specifically, it relates to a fine-tuning method, apparatus, electronic device, and storage medium for a pre-trained model. Background Art

[0002] Data-driven deep learning typically adopts a model pre-training - model fine-tuning approach for industrial applications. For example, first pre-train on a super-large-scale dataset to obtain a pre-trained model, and then fine-tune the downstream task model according to the specific task data of the actual application scenario. For instance, first train an image recognition pre-trained model based on large-scale image data, and then fine-tune the pre-trained model based on different types of image data to obtain an image classification model. Summary of the Invention

[0003] The present disclosure aims to solve at least one of the technical problems in the related art to some extent.

[0004] A first aspect embodiment of the present disclosure proposes a fine-tuning method for a pre-trained model, including:

[0005] Obtain first training data associated with a target task and a pre-trained model, where the first training data includes input data and annotation results;

[0006] Generate a second feature processing network based on the pre-trained model, where the scale of the second feature processing network is smaller than that of the first feature processing network in the pre-trained model;

[0007] Fuse the second feature processing network and the pre-trained model based on a preset rule to obtain a fused model;

[0008] Input the input data into the fused model to obtain a prediction result output by the fused model;

[0009] Correct the second feature processing network in the fused model according to the difference between the prediction result and the annotation result until a target model is obtained.

[0010] A second aspect embodiment of the present disclosure proposes a fine-tuning apparatus for a pre-trained model, including:

[0011] A first acquisition module, configured to obtain first training data associated with a target task and a pre-trained model, where the first training data includes input data and annotation results;

[0012] A generation module for generating a second feature processing network based on the pre-trained model, where the scale of the second feature processing network is smaller than that of the first feature processing network in the pre-trained model;

[0013] A fusion module for fusing the second feature processing network and the pre-trained model based on a preset rule to obtain a fused model;

[0014] A second acquisition module for inputting the input data into the fused model to obtain a prediction result output by the fused model;

[0015] A correction module for correcting the second feature processing network in the fused model according to the difference between the prediction result and the annotation result until a target model is obtained.

[0016] An embodiment of the third aspect of the present disclosure provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the fine-tuning method of the pre-trained model proposed in the first aspect embodiment of the present disclosure.

[0017] An embodiment of the fourth aspect of the present disclosure provides a computer-readable storage medium storing a computer program, which when executed by a processor, implements the fine-tuning method of the pre-trained model proposed in the first aspect embodiment of the present disclosure.

[0018] An embodiment of the fifth aspect of the present disclosure provides a computer program product including a computer program, which when executed by a processor, implements the fine-tuning method of the pre-trained model proposed in the first aspect embodiment of the present disclosure.

[0019] The fine-tuning method, device, computer device, and storage medium of the pre-trained model provided by the present disclosure have the following beneficial effects:

[0020] In the embodiments of the present disclosure, first, the first training data associated with the target task and the pre-trained model are obtained, the second feature processing network is generated based on the pre-trained model, then based on a preset rule, the second feature processing network and the pre-trained model are fused to obtain a fused model, then the input data is input into the fused model to obtain a prediction result output by the fused model, and finally, according to the difference between the prediction result and the annotation result, the second feature processing network in the fused model is corrected until a target model is obtained. Thus, not only the computational amount in the fine-tuning process of the pre-trained model is reduced, the processing speed is improved, the computing power and storage space are saved, the cost of model fine-tuning is reduced, but also it is ensured that the final target model retains the knowledge learned by the pre-trained model.

[0021] Additional aspects and advantages of the present disclosure will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present disclosure. Description of the Drawings

[0022] The above-mentioned and / or additional aspects and advantages of the present disclosure will become apparent and be readily understood from the following description of embodiments in conjunction with the drawings, where:

[0023] Figure 1 is a schematic flowchart of a method for fine-tuning a pre-trained model provided by an embodiment of the present disclosure;

[0024] Figure 2 is a schematic flowchart of a method for fine-tuning a pre-trained model provided by an embodiment of the present disclosure;

[0025] Figure 3 is a schematic flowchart of a method for fine-tuning a pre-trained model provided by an embodiment of the present disclosure;

[0026] Figure 4 is a schematic structural diagram of a fused model provided by the present disclosure;

[0027] Figure 5 is a schematic structural diagram of a device for fine-tuning a pre-trained model provided by an embodiment of the present disclosure;

[0028] Figure 6 shows a block diagram of an exemplary computer device suitable for implementing the embodiments of the present disclosure. Detailed Embodiments

[0029] Embodiments of the present disclosure will be described in detail below. Examples of the embodiments are shown in the drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present disclosure, and should not be construed as limiting the present disclosure.

[0030] Embodiments of the present disclosure relate to technical fields such as deep learning and natural language processing.

[0031] Deep Learning (DL) is to learn the internal laws and representation levels of sample data. The information obtained during these learning processes is very helpful for the interpretation of data such as text, images, and sounds. Its ultimate goal is to enable machines to have the ability to analyze and learn like humans and be able to recognize data such as text, images, and sounds.

[0032] Natural Language Processing (NLP) is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can enable effective communication between humans and computers using natural language.

[0033] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0034] In related technologies, in the method of model fine-tuning, the commonly used methods are full-scale fine-tuning based on the parameters of the pre-trained model or fine-tuning the parameters of the fully connected layer by freezing the backbone network of the pre-trained model. However, it is difficult to achieve the best effect with the above fine-tuning methods. For example, full-scale fine-tuning of large model parameters is likely to cause the model to forget the knowledge learned in the pre-training stage and is prone to overfitting to downstream tasks, while the method of only fine-tuning the fully connected layer is prone to underfitting problems. In addition, although the method based on parameter increment can better handle the problems of model overfitting and underfitting and has an improvement compared with the commonly used non-reference methods, they bring additional costs. On the one hand, they change the original network structure, and on the other hand, large models are required to participate in gradient calculation during downstream fine-tuning, resulting in very high video memory occupancy, which does not meet the actual application and also increases the cost of model deployment.

[0035] In view of the above problems, the present disclosure proposes a fine-tuning method for a pre-trained model. By generating a second feature processing network on the side of the pre-trained model and correcting and training the second feature processing network based on the first training data associated with the target task until the target model is obtained. Since only the parameters of the second feature processing network need to be updated and there is no need to perform gradient operations on the pre-trained model, the computational amount in the fine-tuning process is greatly reduced, the video memory occupancy in the fine-tuning process is saved, the processing speed is improved, the computing power and storage space are saved, the cost of model fine-tuning is reduced, and it is ensured that the final model retains the knowledge learned by the pre-trained model.

[0036] Next, the fine-tuning method, device, electronic device, and storage medium of the pre-trained model according to the embodiments of the present disclosure will be described with reference to the accompanying drawings.

[0037] Figure 1 It is a schematic flowchart of a fine-tuning method for a pre-trained model provided by an embodiment of the present disclosure.

[0038] As Figure 1 shown, the fine-tuning method of the pre-trained model may include the following steps:

[0039] Step 101, obtain the first training data associated with the target task and the pre-trained model, where the first training data includes input data and annotation results.

[0040] Among them, the target task is a downstream task that applies the fine-tuned pre-trained model. For example, the target task may be a classification task, a translation task, etc., and the present disclosure does not limit this.

[0041] Among them, the first training data is a data set associated with the target task and used for fine-tuning the pre-trained model. It can be a text data set, or it can also be an image data set, or it can also be an audio data set, etc. The type of the first training data is determined by the target task, and the present disclosure does not limit this. Among them, the fine-tuning of the pre-trained model is a process of applying the pre-trained model to a specific data set and adapting the parameters of the pre-trained model to the specific data set.

[0042] Among them, the pre-trained model is a large model that has been trained on a large-scale data set, with a complex network structure and a large number of parameters.

[0043] Among them, the annotation result is the result of manually or automatically annotating the input data based on the target task. For example, when the target task is an image classification task and the input data is an image containing a cat, the corresponding annotation result may be "feline".

[0044] Step 102: Generate a second feature processing network based on the pre-trained model, where the scale of the second feature processing network is smaller than that of the first feature processing network in the pre-trained model.

[0045] Among them, the pre-trained model generally can be divided into an input network, a first feature processing network, and an output network. Among them, the input network is used to perform feature extraction and other processing on the input data, the first feature processing network is used to process the features output by the input network, such as convolution, context information dependence capture, automatic importance learning, etc., and the output network is used to determine the processing result corresponding to the input data based on the features output by the first feature processing network, such as the type of the input data, the data after translation of the input data, etc. The second feature processing network can be a network generated based on the pre-trained model and whose parameters can be adjusted during the fine-tuning process of the pre-trained model.

[0046] In the embodiments of the present disclosure, considering that the network in the pre-trained model that affects the model's feature processing is mainly the first feature processing network, the second feature processing network can be generated only based on the first feature processing network.

[0047] In some possible implementation forms, in order to reduce the computational amount during the model fine-tuning process, the first feature processing network can be compressed to a certain extent to obtain the second feature processing network, and the present disclosure does not limit this.

[0048] Step 103: Based on a preset rule, fuse the second feature processing network with the pre-trained model to obtain the fused model.

[0049] The preset rule can be preset according to the type of the target task or according to historical training data. The present disclosure does not limit this.

[0050] In some possible implementation forms, in order to avoid forgetting the knowledge learned in the pre-training stage and minimize the scale of the fine-tuned parameters during the fine-tuning process, the present disclosure can parallelize the second feature processing network with the network having the same function in the pre-trained model to obtain the fused model. For example, connect the input end of the second feature processing network to the input end of the first feature processing network in the pre-trained model, and connect the output end of the second feature processing network to the output end of the first feature processing network in the pre-trained model. Then, perform forward processing of fine-tuning using the fused model, and only fine-tune the parameters of the second feature processing network.

[0051] Step 104: Input the input data into the fused model to obtain the prediction result output by the fused model.

[0052] The prediction result is the prediction result obtained after the fused model processes the input data. For example, if the target task is a classification task, the prediction result is the predicted type; if the target task is a translation task, the prediction result is the translated data corresponding to the input data.

[0053] In some possible implementation forms, when the input data is an image, since the pre-trained model is a Transformer model and the data processed by the Transformer model is sequence data, the image can be first block-encoded to obtain the encoded sequence corresponding to the image, and then the encoded sequence is input into the fused model, so that the fused model can process the image data, thereby improving the applicable range of the fused model.

[0054] In the present disclosure, the granularity of image block division is determined according to the scale of the encoded sequence that the fused model can process. For example, if the scale of the encoded sequence that the fused model can process is large, the image can be divided into finer-grained blocks; if the scale of the encoded sequence that the fused model can process is small, the image can be divided into coarser-grained blocks.

[0055] It should be noted that since the fused model includes the pre-trained model, the obtained prediction result contains both the knowledge learned in the pre-training stage and the knowledge of the second feature processing network.

[0056] Step 105: According to the difference between the prediction result and the annotation result, correct the second feature processing network in the fused model until the target model is obtained.

[0057] Among them, the target model is a model applicable to the target task. The target model can be, for example, a text / image classification model, a text / image detection model, a response model, a translation model, an image segmentation model, an audio classification model, an audio segmentation model, etc. The present disclosure does not limit this.

[0058] In the present disclosure, according to the prediction result and the annotation result, the cross-entropy loss function can be used to calculate the correction gradient, and then the second feature processing network is corrected. Since the prediction result contains both the knowledge learned in the pre-training stage and the knowledge of the second feature processing network, therefore, in the target model obtained by correcting the second feature processing network based on the difference between the prediction result and the annotation result, because it contains the pre-trained model, the knowledge learned in the pre-training stage is retained. At the same time, the second feature processing network is adjusted based on the first training data associated with the target task, so it is also applicable to the target task. It is also possible to correct the parameters in the preset fusion rule according to the difference between the prediction result and the annotation result, so that the model after fusing the second feature processing network and the pre-trained model is more applicable to the target task.

[0059] In the embodiment of the present disclosure, first, obtain the first training data associated with the target task and the pre-trained model, generate the second feature processing network based on the pre-trained model, then fuse the second feature processing network and the pre-trained model based on the preset rule to obtain the fused model, then input the input data into the fused model to obtain the prediction result output by the fused model, and finally, according to the difference between the prediction result and the annotation result, correct the second feature processing network in the fused model until the target model is obtained. Thereby, not only the computational amount in the fine-tuning process of the pre-trained model is reduced, the processing speed is improved, the computing power and storage space are saved, the cost of model fine-tuning is reduced, but also it is ensured that the final target model retains the knowledge learned by the pre-trained model.

[0060] Next, taking the first training data including image data and the corresponding annotation categories of the image data as an example, the fine-tuning process of the pre-trained model will be further described.

[0061] First, based on the sequence scale that the pre-trained model can handle, the image data in the first training data is block-encoded. For example, based on the sequence scale that the pre-trained model can handle, it is determined that the image data needs to be divided into 3*3 image blocks. Then each image block is encoded, and the encodings corresponding to each image block are spliced to obtain the encoding sequence corresponding to the image data.

[0062] After that, the encoded sequence is input into the fused model to obtain the image category predicted by the fused model. According to the difference between the predicted image category and the labeled category, the corrected gradient is determined. Then, based on the corrected gradient, the second feature processing network is corrected until an image classification model is obtained.

[0063] Since only the second feature processing network with a small scale has its parameters corrected during the process of obtaining the image classification model based on the pre-trained model, the small parameter scale reduces the amount of data processed during the fine-tuning of the pre-trained model, improves the processing speed, saves computing power and storage space, and improves the efficiency of model fine-tuning. At the same time, since the pre-trained model obtained during the pre-training stage is retained in the obtained image classification model, the image classification model retains the knowledge learned during the pre-training process.

[0064] Figure 2 It is a schematic flowchart of a method for fine-tuning a pre-trained model provided by an embodiment of the present disclosure.

[0065] As Figure 2 shown, the method for fine-tuning the pre-trained model may include the following steps:

[0066] Step 201, obtain first training data associated with the target task and the pre-trained model, where the first training data includes input data and labeled results.

[0067] The specific implementation form of step 201 can refer to the detailed descriptions in other embodiments of the present disclosure and will not be specifically elaborated here.

[0068] Step 202, prune the first feature processing network in the pre-trained model according to a preset pruning ratio to obtain an initial feature processing network.

[0069] Pruning is a network compression technology that can reduce the number of parameters and computational amount of the network model without affecting the performance of the network model. The preset pruning ratio can be determined according to the type of the target task. For example, the preset pruning ratio corresponding to a classification model is 50%, and the preset pruning ratio corresponding to a translation model is 40%, etc., so that the pruned model is more suitable for the current target task. Alternatively, the preset pruning ratio can also be a preset fixed value, that is, for any type of model, pruning is performed based on this ratio, and the present disclosure does not limit this.

[0070] In some possible implementation forms, the nodes in each layer of the first feature processing network can be pruned according to the preset pruning ratio in the order of the corresponding network parameters from small to large to obtain the initial feature processing network.

[0071] Among them, the nodes in the first feature processing network are the basic computing units in the first feature processing network. Among them, the connection strength between nodes is represented by weights, and the magnitude of the weights represents the importance of the nodes. The network parameters corresponding to the nodes are the node weights.

[0072] Among them, the initial feature processing network is the feature processing network generated after pruning the first feature processing network.

[0073] In the present disclosure, since the nodes in the first feature processing network are already trained nodes, the node matrix corresponding to the first feature processing network composed of these nodes, and the network parameters corresponding to each node in the node matrix are all known. First, the network parameter values corresponding to each layer of nodes are statistically analyzed, and then, according to the magnitude of the corresponding network parameters, a part of the nodes with relatively larger corresponding network parameters (the remaining part after subtracting a part of the nodes with relatively smaller corresponding network parameters according to a preset pruning ratio) are selected to form the initial feature processing network.

[0074] It can be understood that by pruning the network nodes of each layer in the pre-trained model, an initial feature processing network with the same structure but smaller scale can be obtained compared with the pre-trained model.

[0075] Step 203: Input the second training data into the first feature processing network and the initial feature processing network respectively, and obtain the first feature output by the first feature processing network and the second feature output by the initial feature processing network.

[0076] Among them, the second training data is the data processed by the input network in the pre-trained model. For example, it is a vector obtained after mapping text, or a sequence identifier obtained after block encoding of an image, etc. It can be obtained by processing the data used in the pre-trained model stage through the input network, or by processing the first training data through the input network. The present disclosure does not make any limitations in this regard.

[0077] Among them, the first feature is the feature corresponding to the second training data obtained after the first feature processing network processes the second training data. The second feature is the feature corresponding to the second training data obtained after the initial feature processing network processes the second training data.

[0078] Step 204: Based on the difference between the second feature and the first feature, correct the network parameters in the initial feature processing network until the second feature processing network is obtained.

[0079] In the present disclosure, the network parameters in the initial feature processing network can be corrected by correcting the initial feature processing network a preset number of times, or by ensuring that the difference between the second feature output by the corrected second feature processing network and the first feature is within a specified range. That is to say, it is ensured that the second feature processing network and the first feature processing network have the same or similar performance. The present disclosure does not limit this.

[0080] In some possible implementation forms, in order to make the performance of the initial feature processing network the same as or similar to that of the first feature processing network, after obtaining the first feature and the second feature, the mean square error loss function can be used as the distillation loss function to calculate the correction gradient, and then the network parameters in the initial feature processing network are corrected based on the correction gradient, so that the initial feature processing network can maintain as much knowledge as the first feature processing network.

[0081] Step 205: Based on a preset rule, fuse the second feature processing network with the pre-trained model to obtain the fused model.

[0082] Step 206: Input the input data into the fused model to obtain the prediction result output by the fused model.

[0083] Step 207: According to the difference between the prediction result and the annotation result, correct the second feature processing network in the fused model until the target model is obtained.

[0084] Among them, for the specific implementation forms of steps 205 to 207, reference can be made to the detailed descriptions in other embodiments of the present disclosure, and details will not be elaborated here.

[0085] In the embodiments of the present disclosure, after obtaining the first training data and the pre-trained model associated with the target task, first, the first feature processing network in the pre-trained model is pruned according to a preset pruning ratio to obtain the initial feature processing network. Then, the second training data is input into the first feature processing network and the initial feature processing network respectively to obtain the first feature output by the first feature processing network and the second feature output by the initial feature processing network. Then, based on the difference between the second feature and the first feature, the network parameters in the initial feature processing network are corrected until the second feature processing network is obtained. Then, based on a preset rule, the second feature processing network is fused with the pre-trained model to obtain the fused model. After that, the input data is input into the fused model to obtain the prediction result output by the fused model. Finally, based on the difference between the prediction result and the annotation result, the second feature processing network in the fused model is corrected until the target model is obtained. Thus, by transforming the pre-trained model and then fine-tuning some of the transformed parameters based on the target task, not only the computational amount in the fine-tuning process of the pre-trained model is reduced, the speed of model fine-tuning is improved, the computing power and storage space are saved, and the cost of model fine-tuning is reduced, but also it is ensured that the final target model retains the knowledge learned by the pre-trained model.

[0086] Figure 3 It is a schematic flowchart of a method for fine-tuning a pre-trained model provided by an embodiment of the present disclosure.

[0087] As Figure 3 shown, the method for fine-tuning the pre-trained model may include the following steps:

[0088] Step 301, obtain the first training data and the pre-trained model associated with the target task, where the first training data includes input data and annotation results.

[0089] Step 302, generate a second feature processing network based on the pre-trained model, where the scale of the second feature processing network is smaller than that of the first feature processing network in the pre-trained model.

[0090] Among them, for the specific implementation forms of steps 301 to 302, reference may be made to the detailed descriptions in other embodiments of the present disclosure, and details will not be described here.

[0091] Step 303, use the first preset mapping network to connect the output of the input network in the pre-trained model to the first input layer in the second feature processing network.

[0092] Among them, the first preset mapping network can be in any form and is used to change the feature dimension. For example, the first preset mapping network is a fully connected layer. In the present disclosure, since the second feature processing network is obtained by pruning the first feature processing network and the feature dimension it can process is smaller than that of the first feature processing network, in the present disclosure, the first preset mapping network can be a fully connected layer for reducing the feature dimension output by the input network.

[0093] It should be noted that the degree of dimension reduction of the first preset mapping network for the feature dimension can be determined by the number of parameters of the first feature processing network and the number of parameters of the second feature processing network. For example, the amount of feature reduction dimension corresponding to the first mapping network can be determined according to the ratio of the number of the first network parameters in the first feature network to the number of the second network parameters in the second feature processing network. For example, if the number of the second network parameters is 1 / of the number of the first network parameters, then the first mapping network needs to reduce the dimension of the features output by the input network to 1 / of its original dimension.

[0094] Step 304: Use a preset fusion rule to fuse the features obtained by first mapping the output of the (i - 1)-th layer in the first feature processing network with the output of the (i - 1)-th layer in the second feature processing network, and input the fused features into the i-th layer in the second feature processing network.

[0095] Among them, i is a value greater than or equal to 1 and less than or equal to N, and N is the number of layers included in the second feature processing network.

[0096] Among them, the preset fusion rule is a rule set for fusing the features of the (i - 1)-th layer in the second feature processing network and the features obtained by first mapping the (i - 1)-th layer in the first feature processing network. It can be preset according to the type of the target task, or it can also be preset according to historical training data. The present disclosure does not limit this.

[0097] In the present disclosure, in order to ensure that the performance of the second feature processing network is similar to that of the first feature processing network, a dynamic weighting strategy is adopted. The features output by each feature processing layer in the first feature processing network are dynamically weighted and then input into the corresponding feature processing layer in the second feature processing network, so that each layer of the second feature processing network can obtain the output of the corresponding previous layer in the first feature processing network.

[0098] Step 305: Use the second preset mapping network to connect the output of the N-th layer in the second feature processing network with the output network in the pre-trained model to obtain a model after fusing the second feature processing network and the pre-trained model.

[0099] Among them, the second preset mapping network can be in any form and is used to change the feature dimension, such as the second preset mapping network being a fully connected layer. In the present disclosure, since the second feature processing network is obtained by pruning the first feature processing network, the feature dimension output by it is smaller than that of the first feature processing network. Therefore, in the present disclosure, the second preset mapping network can be a fully connected layer used to increase the feature dimension output by the second feature processing network.

[0100] It should be noted that the degree of dimension increase of the second preset mapping network for the feature dimension can be determined by the number of parameters of the first feature processing network and the number of parameters of the second feature processing network. For example, the amount of feature increase dimension corresponding to the second mapping network can be determined according to the quantity ratio between the first network parameter in the first feature network and the second network parameter in the second feature processing network. For example, if the number of parameters of the second network parameter is 1 / r of the number of parameters of the first network parameter, then the second mapping network needs to increase the feature dimension output by the second feature processing network by r times, so as to ensure that the increased feature dimension can be processed by the output network of the first feature processing network. When the feature reduction dimension amount corresponding to the first mapping network and the feature increase dimension amount corresponding to the second mapping network are reciprocal to each other, it can be considered that the feature dimension of the output network is the same as the feature dimension of the first feature processing network, thereby improving the fusion efficiency of the second feature processing network and the pre-trained model.

[0101] Step 306: Input the input data into the fused model, and obtain the prediction result output by the fused model.

[0102] Step 307: According to the difference between the prediction result and the annotation result, correct the second feature processing network in the fused model until the target model is obtained.

[0103] Among them, the specific implementation forms of steps 306 to 307 can refer to the detailed descriptions in other embodiments of the present disclosure, and will not be specifically elaborated here.

[0104] In the embodiments of the present disclosure, first, after obtaining the first training data and the pre-trained model associated with the target task, a second feature processing network is generated based on the pre-trained model. Then, using the first preset mapping network, the output of the input network in the pre-trained model is connected to the first input layer in the second feature processing network. At the same time, using the preset fusion rule, the features after the first mapping of the output of the (i - 1)-th layer in the first feature processing network are fused with the output of the (i - 1)-th layer in the second feature processing network, and the fused features are input into the i-th layer in the second feature processing network. Then, using the second preset mapping network, the output of the N-th layer in the second feature processing network is connected to the output network in the pre-trained model to obtain the model after fusing the second feature processing network and the pre-trained model. After that, the input data is input into the fused model to obtain the prediction result output by the fused model. Finally, according to the difference between the prediction result and the annotation result, the second feature processing network in the fused model is corrected until the target model is obtained. Thus, during the fine-tuning process of the pre-trained model, the pre-trained model only participates in the forward calculation of the fine-tuning process and does not participate in the backpropagation gradient correction. This not only reduces the computational complexity of the fine-tuning process of the pre-trained model, improves the processing speed, saves computing power and storage space, reduces the cost of model fine-tuning, but also ensures that the final target model retains the knowledge learned by the pre-trained model.

[0105] The following Figure 4 , taking the pre-trained model as a classification model as an example, further illustrates a method for fine-tuning a pre-trained model provided by an embodiment of the present disclosure.

[0106] As Figure 4 shown, Figure 4 is a schematic structural diagram of a fused model provided by the present disclosure. Among them, the label "1" in the figure represents the first mapping network, the label "2" represents the second mapping network, the label "3" represents the fusion network, M is the pre-trained model, Θ i is the parameter of the i-th layer of the pre-trained model. The pre-trained model consists of N layers, then the predicted output of the pre-trained model can be expressed as y = M(x; Θ) ∈ R d , Θ = Θ1·Θ2·…·Θ N , d is the dimension of the image vector. m is the second feature processing network, θ i is the parameter of the i-th layer of the second feature processing network; f i M is the input or output of the i-th layer of the pre-trained model, f i m is the input or output of the i-th layer of the second feature processing network; μ i is the weight of the parameter in the i-th layer of the pre-trained model; the encoder is the input network of the pre-trained model, and the classifier is the output network of the pre-trained model; r is the dimension reduction or dimension increase factor.

[0107] In the present disclosure, after obtaining the first training data associated with the target task and the pre-trained model M, importance sampling θ of the parameters of the pre-trained model M can be adopted i = sample(Θ i ), that is, pruning the pre-trained model M to obtain the second feature processing network m. Since the nodes in the pre-trained model are already trained nodes, the node matrix corresponding to the first feature processing network composed of these nodes, and the network parameters corresponding to each node in the node matrix are all known. First, the weights corresponding to each layer of nodes are statistically analyzed. After subtracting a part of the nodes with relatively small corresponding weights according to the preset pruning ratio, the remaining nodes form the initial feature processing network. That is, 1 / r nodes with relatively large weights in each layer of the pre-trained model M can be selected to form the parameters of the initial feature processing network. Then, the method of distilling the initial feature processing network with the pre-trained model M is adopted. The mean square error loss function is used as the distillation loss function to calculate the corrected gradient. Then, based on the corrected gradient, the parameters in the initial feature processing network are corrected until the second feature processing network m is obtained.

[0108] Then, the parameters in the second feature processing network m are fine-tuned while keeping the parameters of the pre-trained model M fixed. To ensure that the performance of the second feature processing network m is similar to that of the pre-trained model M, a fusion rule based on a dynamic weighting strategy is adopted to fuse the features of the pre-trained model M and the features of the second feature processing network m to obtain the fusion features of the fusion network. That is, the input of the i-th layer of the fusion network can be expressed as

[0109]

[0110] Among them, is the input of the i-th layer of the fusion network, μ i is the weight of the parameters in the i-th layer of the pre-trained model included in the fusion rule, is the first preset mapping network, is the output of the i-th layer of the pre-trained model, is the input of the i-th layer of the second feature processing network, and N is the number of layers of the pre-trained model.

[0111] As Figure 4 shown, the input of the second layer of the fusion network can be expressed as

[0112]

[0113] Among them, is the input of the second layer of the fusion network, μ2 is the weight of the parameters in the second layer of the pre-trained model,

[0114] It is a mapping layer connected to the second feature layer of the first feature processing network. It is the output of the second layer of the pre-trained

[0115] model and is the input to the second layer of the second feature processing network. It is the input to the second layer of the second feature processing network.

[0116] Finally, the output of the Nth layer of the pre-trained model M is mapped and dynamically weighted to the Nth layer of the second feature processing network m. Then, the output of the Nth layer of the second feature processing network m is mapped and connected back to the classifier of the pre-trained model M, and the predicted classification result is output by the classifier. After that, the cross-entropy loss function can be used to calculate the corrected gradient to correct the parameters in the second feature processing network m and the weights in the fusion rule until the final target model is obtained.

[0117] In the present disclosure, during the classification training process, only the gradient is backpropagated from the second feature processing network m, and the pre-trained model M only participates in the forward inference, without performing gradient calculation and parameter update, which greatly reduces the computational amount during the training process, improves the processing speed, saves computing power and storage space, and saves the cost of model fine-tuning.

[0118] To implement the above embodiments, the present disclosure also proposes a fine-tuning device for a pre-trained model.

[0119] Figure 5 It is a schematic structural diagram of the fine-tuning device for the pre-trained model provided by the embodiments of the present disclosure.

[0120] As Figure 5 shown, the fine-tuning device 500 for the pre-trained model may include:

[0121] A first acquisition module 501, configured to acquire first training data associated with a target task and a pre-trained model, where the first training data includes input data and annotation results;

[0122] A generation module 502, configured to generate a second feature processing network based on the pre-trained model, where the scale of the second feature processing network is smaller than that of the first feature processing network in the pre-trained model;

[0123] A fusion module 503, configured to fuse the second feature processing network and the pre-trained model based on a preset rule to obtain a fused model;

[0124] A second acquisition module 504, configured to input the input data into the fused model to obtain a prediction result output by the fused model;

[0125] A correction module 505, configured to correct the second feature processing network in the fused model according to the difference between the prediction result and the annotation result until a target model is obtained.

[0126] Optionally, the fine-tuning device 500 of the pre-trained model further includes:

[0127] A first determination module (not shown in the figure), configured to determine a preset pruning ratio according to the type of the target task.

[0128] A second determination module (not shown in the figure), configured to determine the feature reduction dimension of the first mapping network and the feature increase dimension of the second mapping network according to the quantity ratio between the first network parameters in the first feature processing network and the second network parameters in the second feature processing network.

[0129] Optionally, the above-mentioned generation module 502 is further configured to:

[0130] Prune the first feature processing network in the pre-trained model according to the preset pruning ratio to obtain an initial feature processing network;

[0131] Input the second training data into the first feature processing network and the initial feature processing network respectively to obtain the first feature output by the first feature processing network and the second feature output by the initial feature processing network, where the second training data is the data processed by the input network in the pre-trained model;

[0132] Based on the difference between the second feature and the first feature, correct the network parameters in the initial feature processing network until the second feature processing network is obtained.

[0133] Optionally, the above-mentioned generation module 502 is further configured to:

[0134] Prune the nodes in each layer of the first feature processing network according to the preset pruning ratio in the order of the corresponding network parameters from small to large to obtain an initial feature processing network.

[0135] Optionally, the above-mentioned fusion module 503 is further configured to:

[0136] Use the first preset mapping network to connect the output of the input network in the pre-trained model with the first input layer in the second feature processing network;

[0137] Use the preset fusion rule to fuse the feature after the first mapping of the output of the (i - 1)-th layer in the first feature processing network with the output of the (i - 1)-th layer in the second feature processing network, and input the fused feature into the i-th layer in the second feature processing network, where i is a value greater than or equal to 1 and less than or equal to N, and N is the number of layers included in the second feature processing network;

[0138] Use the second preset mapping network to connect the output of the N-th layer in the second feature processing network with the output network in the pre-trained model.

[0139] Optionally, in the above-mentioned second determination module, the feature reduction dimension amount corresponding to the first mapping network and the feature increase dimension amount corresponding to the second mapping network are reciprocal to each other.

[0140] Optionally, the above-mentioned fusion module 503 is further configured to:

[0141] Modify the parameter values in the preset fusion rule according to the difference between the prediction result and the annotation result.

[0142] Optionally, the above-mentioned second acquisition module 504 is further configured to:

[0143] In the case where the input data is an image, perform block coding on the image to obtain an encoded sequence corresponding to the image;

[0144] Input the encoded sequence into the fused model.

[0145] For the functions and specific implementation principles of the above-mentioned modules in the embodiments of the present disclosure, reference may be made to the above-mentioned method embodiments, and details are not described herein again.

[0146] The fine-tuning device of the pre-trained model in the embodiments of the present disclosure first obtains the first training data associated with the target task and the pre-trained model, generates a second feature processing network based on the pre-trained model, then fuses the second feature processing network with the pre-trained model based on a preset rule to obtain a fused model, and then inputs the input data into the fused model to obtain the prediction result output by the fused model. Finally, according to the difference between the prediction result and the annotation result, the second feature processing network in the fused model is corrected until the target model is obtained. Thereby, not only the computational amount in the fine-tuning process of the pre-trained model is reduced, the processing speed is improved, the computing power and storage space are saved, the cost of model fine-tuning is reduced, but also it is ensured that the final target model retains the knowledge learned by the pre-trained model.

[0147] To implement the above embodiments, the present disclosure also proposes an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the fine-tuning method of the pre-trained model proposed in the foregoing embodiments of the present disclosure.

[0148] To implement the above embodiments, the present disclosure also proposes a computer-readable storage medium storing a computer program, which when executed by a processor, implements the fine-tuning method of the pre-trained model proposed in the foregoing embodiments of the present disclosure.

[0149] To implement the above embodiments, the present disclosure also proposes a computer program product including a computer program, which when executed by a processor, implements the fine-tuning method of the pre-trained model proposed in the foregoing embodiments of the present disclosure.

[0150] Figure 6 A block diagram of an exemplary computer device suitable for implementing embodiments of the present disclosure is shown. Figure 6 The computer device 12 shown is merely an example and should not impose any limitation on the functions and usage scope of the embodiments of the present disclosure.

[0151] As Figure 6 shown, the computer device 12 appears in the form of a general-purpose computing device. The components of the computer device 12 may include, but are not limited to: one or more processors or processing units 16, a system memory 28, and a bus 18 connecting different system components (including the system memory 28 and the processing unit 16).

[0152] The bus 18 represents one or more of several types of bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. By way of example, these architectures include, but are not limited to, Industry Standard Architecture (ISA) bus, Micro Channel Architecture (MAC) bus, Enhanced ISA bus, Video Electronics Standards Association (VESA) local bus, and Peripheral Component Interconnection (PCI) bus.

[0153] The computer device 12 typically includes a variety of computer system-readable media. These media can be any available media accessible by the computer device 12, including volatile and non-volatile media, removable and non-removable media.

[0154] The memory 28 may include computer system-readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. The computer device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. Merely by way of example, the storage system 34 may be used for reading and writing non-removable, non-volatile magnetic media ( Figure 6 not shown, commonly referred to as a "hard disk drive"). Although Figure 6Not shown in the figure, a disk drive for reading and writing a removable non-volatile disk (such as a "floppy disk") and an optical disk drive for reading and writing a removable non-volatile optical disk (such as: Compact Disc Read Only Memory; hereinafter referred to as: CD-ROM), Digital Video Disc Read Only Memory; hereinafter referred to as: DVD-ROM) or other optical media) may be provided. In these cases, each drive may be connected to the bus 18 through one or more data medium interfaces. The memory 28 may include at least one program product having a set (such as at least one) of program modules configured to perform the functions of the embodiments of the present disclosure.

[0155] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in the memory 28. Such program modules 42 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment. The program modules 42 generally execute the functions and / or methods in the embodiments described in the present disclosure.

[0156] The computer device 12 may also communicate with one or more external devices 14 (such as a keyboard, a pointing device, a display 24, etc.), and may also communicate with one or more devices that enable a user to interact with the computer device 12, and / or communicate with any device that enables the computer device 12 to communicate with one or more other computing devices (such as a network card, a modem, etc.). Such communication may be carried out through the input / output (I / O) interface 22. In addition, the computer device 12 may also communicate with one or more networks (such as a Local Area Network; hereinafter referred to as: LAN), a Wide Area Network; hereinafter referred to as: WAN) and / or a public network, such as the Internet) through the network adapter 20. As shown in the figure, the network adapter 20 communicates with other modules of the computer device 12 through the bus 18. It should be understood that although not shown in the figure, other hardware and / or software modules may be used in combination with the computer device 12, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0157] The processing unit 16 executes various functional applications and data processing by running the programs stored in the system memory 28, such as implementing the methods mentioned in the foregoing embodiments.

[0158] In the technical solution of the present disclosure, first, the first training data associated with the target task and the pre-trained model are obtained, the second feature processing network is generated based on the pre-trained model, and then, based on the preset rules, the second feature processing network and the pre-trained model are fused to obtain the fused model. After that, the input data is input into the fused model to obtain the prediction result output by the fused model. Finally, according to the difference between the prediction result and the annotation result, the second feature processing network in the fused model is corrected until the target model is obtained. Thus, not only the computational amount in the fine-tuning process of the pre-trained model is reduced, the cost of model fine-tuning is lowered, but also it is ensured that the final target model retains the knowledge learned by the pre-trained model.

[0159] In the description of this specification, the descriptions referring to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present disclosure. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described may be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art may combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0160] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present disclosure, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0161] Any process or method description in the flowchart or described in other ways herein may be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present disclosure includes additional implementations, where the functions may be executed in a substantially simultaneous manner or in a reverse order according to the involved functions, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present disclosure.

[0162] The logic and / or steps represented in the flowchart or otherwise described herein can, for example, be considered a definitional sequence of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device.

[0163] As used in this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which the program can be printed, as the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0164] It should be understood that various parts of the present disclosure can be implemented in hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented in software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques known in the art can be used: discrete logic circuits having logic gates for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0165] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0166] In addition, in various embodiments of the present disclosure, each functional unit may be integrated into a processing module, may exist physically alone for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0167] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present disclosure have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present disclosure.

Claims

1. A fine-tuning method for a pre-trained model, comprising: Obtaining first training data associated with a target task and a pre-trained model, wherein the first training data includes input data and an annotation result, and the input data is any one of the following data: text, image, audio; Generating a second feature processing network based on the pre-trained model, wherein the scale of the second feature processing network is smaller than that of the first feature processing network in the pre-trained model, and the second feature processing network is obtained by pruning the first feature processing network; Fusing the second feature processing network and the pre-trained model based on a preset rule to obtain a fused model; Inputting the input data into the fused model to obtain a prediction result output by the fused model; Correcting the second feature processing network in the fused model according to the difference between the prediction result and the annotation result until a target model is obtained.

2. The method according to claim 1, wherein The generating a second feature processing network based on the pre-trained model includes: Pruning the first feature processing network in the pre-trained model according to a preset pruning ratio to obtain an initial feature processing network; Inputting second training data into the first feature processing network and the initial feature processing network respectively to obtain a first feature output by the first feature processing network and a second feature output by the initial feature processing network, and the second training data is data processed by an input network in the pre-trained model; Correcting the network parameters in the initial feature processing network based on the difference between the second feature and the first feature until a second feature processing network is obtained.

3. The method according to claim 2, wherein The pruning the first feature processing network according to a preset pruning ratio to obtain an initial feature processing network includes: Pruning the nodes in each layer of the first feature processing network according to a preset pruning ratio in the order of corresponding network parameters from small to large to obtain an initial feature processing network.

4. The method according to claim 3, wherein, It further includes: Determining the preset pruning ratio according to the type of the target task.

5. The method according to claim 1, wherein, The fusing the second feature processing network and the pre-trained model based on a preset rule to obtain a fused model includes: Using a first preset mapping network to connect the output of the input network in the pre-trained model with the first input layer in the second feature processing network; Fusing the feature after the first mapping of the output of the (i - 1)-th layer in the first feature processing network and the output of the (i - 1)-th layer in the second feature processing network using a preset fusion rule, and inputting the fused feature into the i-th layer in the second feature processing network, where i is a value greater than or equal to 1 and less than or equal to N, and N is the number of layers included in the second feature processing network; Using a second preset mapping network to connect the output of the N-th layer in the second feature processing network with the output network in the pre-trained model.

6. The method according to claim 5, wherein, It further includes: Determine the feature reduction dimension of the first preset mapping network and the feature increase dimension of the second preset mapping network according to the quantity ratio between the first network parameters in the first feature processing network and the second network parameters in the second feature processing network.

7. The method according to claim 6, wherein, The feature reduction dimension of the first preset mapping network and the feature increase dimension of the second preset mapping network are reciprocal to each other.

8. The method according to claim 5, wherein, It further includes: Correct the parameter values in the preset fusion rule according to the difference between the prediction result and the annotation result.

9. The method according to any one of claims 1-8, wherein, The inputting the input data into the fused model includes: When the input data is an image, perform block coding on the image to obtain the coding sequence corresponding to the image; Input the coding sequence into the fused model.

10. A fine-tuning device for a pre-trained model, including: A first acquisition module, configured to acquire first training data associated with a target task and a pre-trained model, where the first training data includes input data and an annotation result, and the input data is any one of the following data: text, image, audio; A generation module, configured to generate a second feature processing network based on the pre-trained model, where the scale of the second feature processing network is smaller than that of the first feature processing network in the pre-trained model, and the second feature processing network is obtained by pruning the first feature processing network; A fusion module, configured to fuse the second feature processing network and the pre-trained model based on a preset rule to obtain a fused model; A second acquisition module, configured to input the input data into the fused model to obtain a prediction result output by the fused model; A correction module, configured to correct the second feature processing network in the fused model according to the difference between the prediction result and the annotation result until a target model is obtained.

11. The apparatus according to claim 10, wherein, The generation module is specifically configured to: Prune the first feature processing network in the pre-trained model according to a preset pruning ratio to obtain an initial feature processing network; Input second training data into the first feature processing network and the initial feature processing network respectively, to obtain a first feature output by the first feature processing network and a second feature output by the initial feature processing network, where the second training data is data processed by an input network in the pre-trained model; Based on the difference between the second feature and the first feature, correct the network parameters in the initial feature processing network until a second feature processing network is obtained.

12. The apparatus according to claim 11, wherein, The generation module is specifically configured to: Prune the nodes in each layer of the first feature processing network according to a preset pruning ratio in the order of the corresponding network parameters from small to large to obtain an initial feature processing network.

13. The device according to claim 12, wherein, It further includes: A first determination module, configured to determine the preset pruning ratio according to the type of the target task.

14. The apparatus according to claim 10, wherein, The fusion module is further configured to: Use a first preset mapping network to connect the output of the input network in the pre-trained model to the first input layer in the second feature processing network; Using a preset fusion rule, fuse the feature after the first mapping of the output of the (i - 1)-th layer in the first feature processing network with the output of the (i - 1)-th layer in the second feature processing network, and input the fused feature into the i-th layer in the second feature processing network, where i is a value greater than or equal to 1 and less than or equal to N, and N is the number of layers included in the second feature processing network; Using a second preset mapping network, connect the output of the N-th layer in the second feature processing network with the output network in the pre-trained model.

15. The device according to claim 14, wherein, Further comprising: A second determination module, which determines the feature reduction dimension amount corresponding to the first preset mapping network and the feature increase dimension amount corresponding to the second preset mapping network according to the quantity ratio between the first network parameters in the first feature processing network and the second network parameters in the second feature processing network.

16. The apparatus according to claim 15, wherein, The feature reduction dimension amount corresponding to the first preset mapping network and the feature increase dimension amount corresponding to the second preset mapping network are reciprocal to each other.

17. The apparatus according to claim 14, wherein, The fusion module is further configured to: Correct the parameter value in the preset fusion rule according to the difference between the prediction result and the annotation result.

18. The device according to any one of claims 10-17, wherein, The second acquisition module is specifically configured to: In the case where the input data is an image, perform block coding on the image to obtain the coding sequence corresponding to the image; Input the coding sequence into the fused model.

19. An electronic device, characterized in that, Comprising: At least one processor; And a memory communicatively connected to the at least one processor; Wherein, the memory stores instructions that may be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 - 9.

20. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 - 9.

21. A computer program product, comprising a computer program, where the computer program implements the method according to any one of claims 1 - 9 when executed by a processor.

Citation Information

Patent Citations

  • Prediction model training method and device, equipment and storage medium

    CN113762501A

  • System and method for generating 3D objects from 2d images of garments

    US20230046431A1