Determination Method, Device, Equipment and Storage Medium of Transfer Learning Model
By calculating the transfer rates of sample coded information entropy and category coded information entropy, the transfer learning effect of candidate network layers is quickly evaluated, which solves the problem that the optimal model cannot be quickly determined in transfer learning, and improves the training efficiency of the transfer learning model.
Patent Information
- Application Number
- CN202210757206.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-06-29
AI Technical Summary
In transfer learning, it is impossible to quickly determine the optimal transfer learning model for the target task, especially in the presence of multiple transfer learning models.
By determining multiple candidate network layers from the candidate transfer learning model, calculating the mobility based on the sample coded information entropy and the category coded information entropy, the transfer learning effect of the candidate network layer is evaluated, and the optimal transfer learning model is constructed based on the transfer rate.
It realizes the rapid and accurate evaluation of the transfer learning effect of the candidate network layer before transfer learning, reduces the construction time and resource consumption of the transfer learning model, and improves the training efficiency of the transfer learning model.
Smart Images

Figure CN115115050B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer and Internet technologies, and particularly to a method, apparatus, device, and storage medium for determining a transfer learning model. Background Art
[0002] In the process of model learning, through transfer learning, a model developed for one task can be applied to the model training of other tasks.
[0003] Currently, in transfer learning, according to the correlation between the target task and the source task, one or more transfer learning models applicable to the target task are determined. Then, in the case of multiple transfer learning models, the training sample set of the target task is used to train each transfer learning model respectively to obtain multiple deep learning models applicable to the target task. Then, for these multiple deep learning models, the accuracy of the output results of each deep learning model is determined through testing to determine the transfer learning effect of each deep learning model after this transfer learning, and the deep learning model with the highest accuracy, that is, the best transfer learning effect, is determined as the final training model for the target task.
[0004] However, the transfer learning effect can only be evaluated after obtaining the deep learning model through transfer learning. In the case of multiple transfer learning models, it is impossible to quickly determine the optimal transfer learning model corresponding to the target task. Summary of the Invention
[0005] Embodiments of the present application provide a method, apparatus, device, and storage medium for determining a transfer learning model, providing a way to evaluate the transfer learning effect before transfer learning, and being able to quickly and accurately determine the optimal candidate network layer for the training sample set. The technical solutions are as follows.
[0006] According to one aspect of the embodiments of the present application, a method for determining a transfer learning model is provided. The method includes the following steps:
[0007] Determine a plurality of candidate network layers from at least one candidate transfer learning model, where one candidate transfer learning model corresponds to at least one candidate network layer; wherein, different candidate transfer learning models are models trained based on different training data;
[0008] Process the training sample set based on the candidate network layer to obtain the sample coding information entropy and the class coding information entropies corresponding to multiple classes respectively, where the training sample set includes training samples belonging to different classes; wherein, the sample coding information entropy is used to indicate the amount of information contained in the training samples in the training sample set after coding, and the class coding information entropy corresponding to the class is used to indicate the amount of information contained in the training samples belonging to the class in the training sample set after coding;
[0009] Determine the transfer rate of the candidate network layer according to the sample coding information entropy and the multiple class coding information entropies, where the transfer rate is used to indicate the transfer learning effect of the candidate network layer for the training sample set;
[0010] Based on the transfer rates respectively corresponding to each candidate network layer, construct a transfer learning model for the training sample set according to the candidate network layers whose transfer rates meet the first condition.
[0011] According to one aspect of the embodiments of the present application, there is provided an apparatus for determining a transfer learning model, and the apparatus includes the following modules:
[0012] A network layer determination module, configured to determine multiple candidate network layers from at least one candidate transfer learning model, where one candidate transfer learning model corresponds to at least one candidate network layer; wherein, different candidate transfer learning models are models trained based on different training data;
[0013] A sample processing module, configured to process the training sample set based on the candidate network layer to obtain the sample coding information entropy and the class coding information entropies corresponding to multiple classes respectively, where the training sample set includes training samples belonging to different classes; wherein, the sample coding information entropy is used to indicate the amount of information contained in the training samples in the training sample set after coding, and the class coding information entropy corresponding to the class is used to indicate the amount of information contained in the training samples belonging to the class in the training sample set after coding;
[0014] A transfer rate determination module, configured to determine the transfer rate of the candidate network layer according to the sample coding information entropy and the multiple class coding information entropies, where the transfer rate is used to indicate the transfer learning effect of the candidate network layer for the training sample set;
[0015] A model construction module, configured to construct a transfer learning model for the training sample set based on the transfer rates respectively corresponding to each candidate network layer according to the candidate network layers whose transfer rates meet the first condition.
[0016] According to one aspect of the embodiments of the present application, embodiments of the present application provide a computer device, which includes a processor and a memory. At least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the above-mentioned method for determining a transfer learning model.
[0017] According to one aspect of the embodiments of the present application, embodiments of the present application provide a computer-readable storage medium, in which at least one program is stored, and the at least one program is loaded and executed by a processor to implement the above-mentioned method for determining a transfer learning model.
[0018] According to one aspect of the embodiments of the present application, there is provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for determining a transfer learning model.
[0019] The technical solutions provided by the embodiments of the present application can bring the following beneficial effects:
[0020] By determining the transfer rate of a candidate network layer through the sample coding information entropy and the category coding information entropy, and thereby determining the transfer learning effect of the candidate network layer for the training sample set, a way to evaluate the transfer learning effect before transfer learning is provided. It is possible to evaluate the transfer learning effect of the candidate network layer for the training sample set without transfer learning, so as to quickly and accurately determine the optimal candidate network layer for the training sample set; moreover, according to the transfer rate of the candidate network layer, the transfer learning effect is determined with the network layer as the basic unit, which improves the judgment accuracy of the transfer rate. For the same candidate transfer learning model, the optimal network layer suitable for transfer learning in the candidate transfer learning model can be determined. Further, in the subsequent process, a transfer learning model can be constructed with the network layer as the basic unit, which reduces the difference between the initially constructed transfer learning model and the finally trained transfer learning model from the side and improves the training efficiency of the transfer learning model. Description of the Drawings
[0021] Figure 1 is a schematic diagram of a system for determining a transfer learning model provided by an embodiment of the present application;
[0022] Figure 2 Exemplarily shows a schematic diagram of a system for determining a transfer learning model;
[0023] Figure 3 is a flowchart of a method for determining a transfer learning model provided by an embodiment of the present application;
[0024] Figure 4 A schematic diagram exemplarily showing a method for obtaining the migration rate of a candidate migration network;
[0025] Figure 5 A schematic diagram exemplarily showing a method for constructing a transfer learning model;
[0026] Figure 6 A schematic diagram exemplarily showing another method for constructing a transfer learning model;
[0027] Figure 7 A schematic diagram showing the flow of a method for determining a transfer learning model applied to an image classification task;
[0028] Figure 8 A schematic diagram showing the flow of a method for determining a transfer learning model in a chemical molecular structure classification task;
[0029] Figure 9 It is a block diagram of a device for determining a transfer learning model provided by an embodiment of the present application;
[0030] Figure 10 It is a block diagram of a device for determining a transfer learning model provided by another embodiment of the present application;
[0031] Figure 11 It is a block diagram of the structure of a computer device provided by an embodiment of the present application. Detailed implementation manners
[0032] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail with reference to the accompanying drawings.
[0033] Artificial Intelligence (AI) is to use a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, a theory, method, technology, and application system that perceives the environment, acquires knowledge, and uses knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable the machines to have the functions of perception, reasoning, and decision-making.
[0034] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.
[0035] Machine Learning (ML) is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and learning from demonstration.
[0036] The solution provided by the embodiments of this application involves technologies such as machine learning in artificial intelligence, and will be specifically described through the following embodiments.
[0037] Please refer to Figure 1 , which shows a schematic diagram of a system for determining a transfer learning model provided by an embodiment of this application. The system for determining the transfer learning model may include a terminal device 10 and a server 20.
[0038] The terminal device 10 may be an electronic device such as a mobile phone, a tablet computer, a PC (Personal Computer), a smart voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, etc., and the embodiments of this application do not make any limitations in this regard.
[0039] The server 20 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, as well as big data and artificial intelligence platforms.
[0040] The above terminal 10 and the above server 20 may communicate through a network.
[0041] In some embodiments, the above system for determining the transfer learning model is applied in the transfer learning process. Exemplarily, such as Figure 2As shown, the terminal device 10 provides the server 20 with the training sample set for this transfer learning, as well as at least one candidate transfer learning model for the training sample set. Further, the server 20 determines a plurality of candidate network layers from the at least one candidate transfer learning model. Among them, one candidate transfer learning model corresponds to at least one candidate network layer. After that, for each candidate network layer, the server 20 processes each training sample in each training sample set based on the feature extraction function of the candidate network layer to obtain the feature vector corresponding to each training sample, and then constructs a sample feature matrix based on the feature vectors corresponding to each training sample, and constructs a class feature matrix corresponding to the class based on the feature vectors corresponding to the training samples belonging to a certain class. After that, the server 20 calculates and determines the sample coding information entropy according to the sample feature matrix, calculates and determines the class coding information entropy corresponding to each class according to the class feature matrices corresponding to each class. Further, the server 20 determines the transfer rate of the candidate network layer according to the sample coding information entropy and the class coding information entropy corresponding to each class. After that, the server 20 determines the candidate network layer whose transfer rate meets the first condition as the target network layer, and determines the candidate transfer learning model to which the target network layer belongs as the target transfer learning model. In the target transfer learning model, for the target network layer and the network layers before the target network layer, they remain unchanged, and the parameters of the network layers after the target network layer are randomly initialized to construct a transfer learning model for the training sample set. After that, the training sample set is used to train the transfer learning model.
[0042] It should be noted that the above introduction to the terminal device 10 and the server 20 is only exemplary and explanatory. In the exemplary embodiment, the functions of the terminal device 10 and the server 20 can be flexibly set and adjusted. Exemplarily, during the transfer learning process, when the load capacity of the server 20 permits, without relying on the above terminal device 10, the server 20 simultaneously performs data collection, data processing, and model training; or, during the transfer learning process, the terminal device 10 provides a visual configuration interface for the staff to facilitate the staff to configure relevant training parameters in the visual configuration interface, such as the training sample set, the candidate transfer learning model, etc. After that, the server 20 performs data collection, data processing, and model training based on the relevant training parameters configured by the staff.
[0043] Please refer to Figure 3 , which shows a flowchart of a method for determining a transfer learning model provided by an embodiment of the present application. Each step in this method can be executed by the above Figure 1 terminal device 10 and / or server 20 (hereinafter collectively referred to as "computer device"). This method may include at least one of the following steps (301-304):
[0044] Step 301: Determine multiple candidate network layers from at least one candidate transfer learning model.
[0045] A candidate transfer learning model refers to a deep learning model for transfer learning that can be applied to a training sample set. Among them, different candidate transfer learning models are models trained based on different training data. In some embodiments, the training tasks corresponding to different candidate transfer learning models may be the same or different. In a possible implementation manner, the training tasks corresponding to different candidate transfer learning models are different; that is, different candidate transfer learning models are deep learning models for the same task trained based on different training data. In another possible implementation manner, the training tasks corresponding to different candidate transfer learning models are different; for example, different candidate transfer learning models are deep learning models for similar tasks trained based on different training data. Among them, the above-mentioned similar tasks can also be called associated tasks.
[0046] In the embodiment of the present application, before transfer learning, the computer device obtains at least one candidate transfer learning model, and determines multiple candidate network layers from the at least one candidate transfer learning model. Among them, one candidate transfer learning model corresponds to at least one candidate network layer. In some embodiments, after the computer device obtains the at least one candidate transfer learning model, it samples from each candidate transfer learning model to obtain the candidate network layers.
[0047] Step 302: Process the training sample set based on the candidate network layers to obtain the sample coding information entropy and the class coding information entropy corresponding to each of multiple classes.
[0048] In the embodiment of the present application, after the computer device obtains the candidate network layers, it processes the training sample set based on the candidate network layers to obtain the sample coding information entropy and the class coding information entropy corresponding to each of multiple classes. Among them, the sample coding information entropy is used to indicate the amount of information contained in the training samples in the training sample set after coding, and the class coding information entropy corresponding to a class is used to indicate the amount of information contained in the training samples belonging to the class in the training sample set after coding.
[0049] In some embodiments, the above training sample set includes training samples belonging to different categories. Among them, the training sample set may correspond to one or more categories, which is not limited in the embodiments of the present application; the number of training samples corresponding to each category can be any value, which is not limited in the embodiments of the present application. In the embodiments of the present application, after the computer device obtains the candidate network layer, based on each training sample included in the training sample set, the above sample coding information entropy is obtained; based on the training samples belonging to a certain category in the training sample set, the category coding information entropy corresponding to the category is obtained, and then for each category, this step is repeatedly executed to obtain the category coding information entropy corresponding to each category respectively.
[0050] It should be noted that in the embodiments of the present application, a label represents a category, and one category corresponds to one label. That is to say, in the above training sample set, the training samples belonging to the same category correspond to the same label.
[0051] Step 303, determine the transfer rate of the candidate network layer according to the sample coding information entropy and multiple category coding information entropies.
[0052] In the embodiments of the present application, after the computer device obtains the above sample coding information entropy and the above multiple category coding information entropies, it determines the transfer rate of the candidate network layer according to the sample coding information entropy and multiple category coding information entropies. Among them, the transfer rate is used to indicate the transfer learning effect of the candidate network layer for the training sample set.
[0053] Exemplarily, assume that the training sample set x i refers to the i-th training sample in the training sample set, and y i refers to the label of the i-th training data in the training sample set. Among them, the value of y i is {1, 2... C}, that is, the training sample set includes training samples of C categories. Generally speaking, the cross-entropy is used as the loss function of the model of the training sample set. When the loss function is optimized to the optimal minimum, the value of the cross-entropy is approximately equal to the mutual information MI(Z, Y) between the feature vector z i of x i and the label y i :
[0054] MI(Z, Y) = H(Z) - H(Z|Y);
[0055] Among them, H(Z) refers to the information entropy of the feature vectors of the training sample set, and H(Z|Y) refers to the information entropy of the feature vectors of the training samples belonging to the category corresponding to the Y label in the training sample set.
[0056] However, due to the high dimensionality of the feature vectors, algorithms for estimating the information entropy related to the feature vectors are either inaccurate or computationally complex, requiring a large amount of computing resources. Therefore, in the embodiments of the present application, the coding information entropy is used for approximate estimation of the mutual information.
[0057] Physically, the sample coding information entropy R(Z) is approximately equal to evaluating H(Z). Similarly, the class coding information entropy R(Z c ) is approximately equal to evaluating H(Z|Y = c). Then, for the candidate network layer l of the candidate transfer learning model k, the transfer rate T k l is:
[0058]
[0059] where n c refers to the number of training samples included in the c-th class in the training sample set, and N represents the total number of training samples in the training sample set, that is:
[0060]
[0061] Step 304: Based on the transfer rates respectively corresponding to each candidate network layer, construct a transfer learning model for the training sample set according to the candidate network layer whose transfer rate satisfies the first condition.
[0062] In the embodiments of the present application, after the computer device obtains the transfer rates respectively corresponding to each candidate network layer, taking the transfer rate as a benchmark, it determines the candidate network layer whose transfer rate satisfies the first condition from each candidate network layer, and then constructs a transfer learning model for the training sample set according to this candidate network layer.
[0063] The first condition is the judgment condition for the candidate network layer, which can be flexibly set and adjusted according to the actual situation. The embodiments of the present application do not limit this. Exemplarily, the first condition includes but is not limited to at least one of the following: the highest transfer rate, the transfer rate is greater than the target value, etc. The embodiments of the present application do not limit this. Among them, the above target value can be any value, which can be flexibly set and adjusted according to the actual situation. The embodiments of the present application do not limit this.
[0064] In summary, in the technical solution provided by the embodiments of the present application, the migration rate of the candidate network layer is determined by the sample coding information entropy and the category coding information entropy, so as to determine the transfer learning effect of the candidate network layer for the training sample set, providing a way to evaluate the transfer learning effect before transfer learning, and being able to evaluate the transfer learning effect of the candidate network layer for the training sample set without transfer learning, so as to quickly and accurately determine the optimal candidate network layer for the training sample set; moreover, according to the migration rate of the candidate network layer, the transfer learning effect is determined with the network layer as the basic unit, improving the judgment accuracy of the migration rate. For the same candidate transfer learning model, the optimal network layer suitable for transfer learning in the candidate transfer learning model can be determined. Further, in the subsequent process, a transfer learning model can be constructed with the network layer as the basic unit, reducing the difference between the initially constructed transfer learning model and the finally trained transfer learning model from the side, and improving the training efficiency of the transfer learning model.
[0065] Next, taking a certain candidate network layer as an example, the method for obtaining the sample coding information entropy and the category coding information entropy will be introduced.
[0066] In an exemplary embodiment, step 302 above includes at least one of the following:
[0067] 1. Process the training sample set based on the candidate network layer to obtain a sample feature matrix and category feature matrices corresponding to multiple categories respectively.
[0068] In the embodiments of the present application, after the computer device obtains the above candidate network layer, it processes the training sample set based on the candidate network layer to obtain a sample feature matrix and category feature matrices corresponding to multiple categories respectively. In some embodiments, different candidate network layers correspond to different feature extraction functions, and the computer device processes the training sample set based on the feature extraction function of the candidate network layer to obtain the above sample feature matrix and category feature matrices corresponding to multiple categories respectively.
[0069] In some embodiments, after the computer device obtains the above candidate network layer, it determines the feature extraction function of the candidate network layer. Further, based on the feature extraction function of the candidate network layer, each training sample in the training sample set is processed to obtain a feature vector corresponding to each training sample. Exemplarily, assume that the feature extraction function of the candidate network layer l of the candidate transfer learning model k above is f l k (x), then the feature vector z i of the training sample x i = f l k (x i ).
[0070] In some embodiments, after the computer device obtains the feature vectors corresponding to each training sample, it constructs a sample feature matrix according to the feature vectors corresponding to each training sample. Among them, the data in the first target column of the sample feature matrix is the feature vector of the first target training sample, and the first target column refers to any column in the sample feature matrix, and this application embodiment does not make a limitation on this.
[0071] In some embodiments, after the computer device obtains the feature vectors corresponding to each training sample, it constructs a class feature matrix corresponding to each class according to the class to which the training samples included in the training sample set belong. Taking the target class as an example, for at least one training sample belonging to the target class in the training sample set, a class feature matrix corresponding to the target class is constructed according to the feature vectors corresponding to each training sample belonging to the target class. Among them, the data in the second target column of the class feature matrix corresponding to the target class is the feature vector of the second target training sample belonging to the target class, and the second target column refers to any column in the class feature matrix corresponding to the target class, and this application embodiment does not make a limitation on this.
[0072] It should be noted that in the embodiments of this application, the classification of each training sample in the training sample set can be before obtaining the above-mentioned feature vectors or after obtaining the above-mentioned feature vectors, and this application embodiment does not make a limitation on this.
[0073] In a possible implementation manner, after the computer device obtains the above-mentioned training sample set, taking the label as the benchmark, one label corresponds to one class, and based on the labels corresponding to each training sample, the training sample is classified. Then, after obtaining the feature vectors corresponding to each training sample, the above-mentioned sample feature matrix and the above-mentioned class feature matrix are directly constructed.
[0074] In another possible implementation manner, after the computer device obtains the above-mentioned feature vectors, taking the labels corresponding to each training sample in the training sample set as the benchmark, one label corresponds to one class, and the training sample is classified. Then, the above-mentioned sample feature matrix and the above-mentioned class feature matrix are constructed.
[0075] Of course, in other possible implementation manners, in order to reduce the classification time, when collecting the above-mentioned training sample set, the training samples can be directly collected and stored by class distinction to generate the above-mentioned training sample set.
[0076] 2. Determine the sample coding information entropy according to the sample feature matrix.
[0077] In the embodiments of the present application, after the computer device obtains the above sample feature matrix, it determines the sample coding information entropy according to the sample feature matrix. In some embodiments, the computer device obtains the dimension of the feature vector corresponding to the training sample and the coding precision rate for the training sample; further, according to the dimension of the feature vector corresponding to the training sample and the coding precision rate for the training sample, it determines the coding length required to compress the sample feature matrix into the coding indicated by the coding precision rate, and determines the sample coding information entropy based on the coding length. Among them, the above coding precision rate can be any value preset by the staff, and the embodiments of the present application do not limit this.
[0078] Exemplarily, assuming that the dimension of the feature vector corresponding to the training sample is d and the coding precision rate for the training sample is ε, then the sample coding information entropy R(Z) is:
[0079]
[0080] where I d is the d-dimensional identity matrix, Z is the sample feature matrix, and Z T is the transpose matrix of the sample feature matrix.
[0081] In addition, from the above formula, it can be seen that the calculation result of the sample coding information entropy R(Z) can be understood as the log value of the coding length required to compress Z into the coding with a precision of ε.
[0082] 3. Determine the class coding information entropy corresponding to each class according to the class feature matrix corresponding to each class.
[0083] In the embodiments of the present application, after the computer device obtains the above class feature matrix, it determines the class coding information entropy corresponding to each class according to the class feature matrix corresponding to each class. Taking the target class as an example, the computer device obtains the dimension of the feature vector corresponding to the training sample and the coding precision rate for the training sample; further, according to the dimension of the feature vector corresponding to the training sample and the coding precision rate for the training sample, it determines the coding length required to compress the class feature matrix corresponding to the target class into the coding indicated by the coding precision rate, and determines the class coding information entropy corresponding to the target class based on the coding length.
[0084] Exemplarily, for class c, the class coding information entropy R(Z c ) is:
[0085]
[0086] In addition, from the above formula, it can be seen that the calculation result of the class coding information entropy R(Z c ) can be understood as the coding length required to compress Z cThe log value of the coding length required to compress into a code with precision ε.
[0087] Exemplarily, as Figure 4 described, the training sample set includes training samples belonging to class 1, training samples belonging to class 2, and training samples belonging to class 3. The feature extraction function of the candidate network layer l of the candidate transfer learning model k is f l k (x) Obtain the feature vectors of each training sample, then construct a sample feature matrix based on the feature vectors of each training sample, and construct a class feature matrix corresponding to each class based on the feature vectors corresponding to each class. Then, determine the sample coding information entropy according to the sample feature matrix, determine the class coding information entropy according to the class feature matrix, and further determine the transfer rate of the candidate network layer l for the training sample set.
[0088] In summary, in the technical solution provided by the embodiments of the present application, the transfer rate of the candidate network layer is determined by the sample coding information entropy and the class coding information entropy. The information entropy of the feature vectors of the training sample set is evaluated by the sample coding information entropy, and the information entropy of the feature vectors belonging to the class corresponding to the label in the training sample set is evaluated by the class coding information entropy. Compared with the information entropy, the coding rate has the advantages of simple calculation and high accuracy, and improves the evaluation efficiency of the transfer learning effect.
[0089] It should be noted that the above is an introduction to the acquisition methods of the sample coding information entropy and the class coding information entropy taking a certain candidate network layer as an example. In the present application, the above-described steps need to be executed for each candidate network layer.
[0090] Next, the determination method of the candidate network layer will be introduced.
[0091] In an exemplary embodiment, step 301 above includes at least one of the following:
[0092] 1. Determine a candidate transfer learning model based on the training sample set.
[0093] In the embodiments of the present application, after the computer device obtains the above training sample set, it determines a candidate transfer learning model applicable to the training sample set based on the training sample set.
[0094] In a possible implementation, the computer device determines the above-mentioned candidate transfer learning model based on the training task of the training sample set. In some embodiments, after obtaining the above-mentioned training sample set, the computer device determines the training task of the training sample set, and then based on the training task of the training sample set, determines the training model corresponding to the associated task associated with the training task as the candidate transfer learning model. Among them, the above-mentioned associated task can be called a similar task. Exemplarily, if the training task of the training sample set is to recognize the facial images of users, the corresponding associated tasks can be recognizing human images, recognizing employee images, recognizing human motion images, etc.
[0095] In another possible implementation, the computer device determines the above-mentioned candidate transfer learning model based on the training task of the training sample set and the content included in the training sample set. In some embodiments, after obtaining the above-mentioned training sample set, the computer device determines the training task of the training sample set, and then based on the training task of the training sample set, determines the training model corresponding to the associated task associated with the training task as the candidate transfer learning model to be selected. After that, based on the content included in the training data corresponding to each candidate transfer learning model to be selected, a model similar to the content included in the training sample set is determined as the candidate transfer learning model from the candidate transfer learning models to be selected.
[0096] 2. Sample the network layers included in the candidate transfer learning model to obtain at least one candidate network layer corresponding to the candidate transfer learning model.
[0097] In the embodiments of the present application, after obtaining the above-mentioned candidate transfer learning model, the computer device samples the network layers included in the candidate transfer learning model to obtain at least one candidate network layer corresponding to the candidate transfer learning model. Among them, when sampling, the computer device can randomly sample the network layers included in the candidate transfer learning model, or sample the network layers included in the candidate transfer learning model based on the sampling basis. The embodiments of the present application do not limit this. Exemplarily, the sampling basis for the candidate transfer learning model includes but is not limited to at least one of the following: the position between the network layer and the output layer, the importance of the network layer in the model, the preset number of candidate network layers, the preset sampling interval, etc. The embodiments of the present application do not limit this. In some embodiments, different types of candidate transfer learning models correspond to different sampling bases.
[0098] In a possible implementation, to evaluate the transfer learning effect of the overall model as much as possible, candidate network layers are determined based on the positional distance between the network layer and the output layer. In some embodiments, after obtaining the above candidate transfer learning model, the computer device determines, from the candidate transfer learning model, the network layers whose positional distance from the output layer is less than a threshold as the network layers to be sampled, and then samples from the network layers to be sampled to obtain candidate network layers. Among them, the above threshold can be any value, and can be flexibly set and adjusted according to the actual situation, and the embodiments of the present application do not limit this. In some embodiments, when sampling from the network layers to be sampled to obtain candidate network layers, random sampling can be used, or sampling can be performed based on other bases in the above sampling basis except the position between the network layer and the output layer, and the embodiments of the present application do not limit this.
[0099] In another possible implementation, to make the candidate network layers representative of the candidate transfer learning model, candidate network layers are determined based on the importance of the network layers in the model. In some embodiments, after obtaining the above candidate transfer learning model, the computer device obtains at least one network layer included in the candidate transfer learning model, and the importance of each network layer in the candidate transfer learning model. Further, based on this importance, the network layers whose importance meets the second condition are determined from the candidate transfer learning model as the network layers to be sampled, and then samples are taken from the network layers to be sampled to obtain candidate network layers. Among them, the above importance can be determined during model training, and the above second condition can be flexibly set and adjusted according to the actual situation. Exemplarily, the above second condition is that the importance is greater than a threshold value, and this threshold value can be any value, and the embodiments of the present application do not limit this. In some embodiments, when sampling from the network layers to be sampled to obtain candidate network layers, random sampling can be used, or sampling can be performed based on other bases in the above sampling basis except the importance of the network layer in the model, and the embodiments of the present application do not limit this.
[0100] In yet another possible implementation, candidate network layers are determined based on a preset number of candidate network layers. In some embodiments, after obtaining the above candidate transfer learning model, the computer device obtains the total number of network layers included in the candidate transfer learning model, and then determines the sampling interval based on the preset number of candidate network layers, and then performs average sampling based on this sampling interval to obtain the network layers to be sampled, and samples from the network layers to be sampled to obtain candidate network layers. Among them, the above preset number of candidate network layers can be any value, and can be flexibly set and adjusted according to the actual situation, and the embodiments of the present application do not limit this.
[0101] In some embodiments, when sampling and obtaining candidate network layers from the network layers to be sampled, random sampling can be performed, or sampling can be performed based on other bases in the above sampling basis except the preset number of candidate network layers. The embodiments of the present application do not limit this.
[0102] In other possible implementation manners, candidate network layers are determined based on a preset sampling interval. In some embodiments, after the computer device obtains the above candidate transfer learning model, it determines the candidate network layers.
[0103] In some embodiments, after the computer device obtains the above candidate transfer learning model, it obtains the total number of network layers included in the candidate transfer learning model, and then determines the number of candidate network layers based on a preset sampling interval, and samples and obtains the above candidate network layers from the candidate transfer learning model based on the number of candidate network layers.
[0104] It should be noted that the above is an introduction to the acquisition method of candidate network layers taking a certain candidate transfer learning model as an example. In the present application, the above-introduced steps need to be executed for each candidate transfer learning model.
[0105] In summary, in the technical solution provided by the embodiments of the present application, candidate transfer learning models are determined through training tasks, and candidate network layers are selected from the candidate transfer learning models to evaluate the transfer learning ability of the network layers, improving the judgment accuracy of the transfer learning ability. For the same candidate transfer learning model, the optimal network layer suitable for transfer learning in the candidate transfer learning model can be determined, reducing the difference between the initially constructed transfer learning model and the finally trained transfer learning model from the side, and improving the training efficiency of the transfer learning model.
[0106] Next, the construction method of the transfer learning model will be introduced.
[0107] In an exemplary embodiment, the above step 304 includes at least one of the following steps:
[0108] 1. Based on the transfer rates respectively corresponding to each candidate network layer, determine the candidate network layers whose transfer rates meet the first condition as target network layers.
[0109] In the embodiments of the present application, after the computer device obtains the above transfer rates, based on the transfer rates respectively corresponding to each candidate network layer, it determines the candidate network layers whose transfer rates meet the first condition as target network layers.
[0110] In a possible implementation manner, the above first condition is the highest mobility. In some embodiments, after the computer device obtains the mobilities of each candidate network layer, it sorts the mobilities from high to low, and then determines the candidate network layer with the highest mobility as the target network layer, and constructs a transfer learning model based on the target network layer.
[0111] In another possible implementation manner, the above first condition is that the mobility is greater than a target value. In some embodiments, after the computer device obtains the mobilities of each candidate network layer, it determines the candidate network layers with mobilities greater than the target value as the target network layer. Of course, in other possible implementation manners, the above first condition is that the mobility is greater than the target value and the position distance from the output layer is the smallest. Exemplarily, after the computer device obtains the mobilities of each candidate network layer, it determines the candidate network layers with mobilities greater than the target value as the candidate target network layers, and the number of the candidate target network layers is not one. Then, the computer device determines the candidate transfer learning networks to which each candidate target network layer belongs, and further determines the position distances between each candidate target network layer and the output layer in the corresponding candidate transfer learning network, and determines the candidate target network layer with the smallest position distance as the target network layer.
[0112] 2. Construct a transfer learning model for the training sample set according to the target network layer.
[0113] In the embodiments of the present application, after the computer device determines the above target network layer, it constructs a transfer learning model for the training sample set according to the target network layer.
[0114] In a possible implementation manner, in order to improve the construction efficiency of the transfer learning model, the computer device constructs the above transfer learning model based on the candidate transfer learning model to which the target network layer belongs. In some embodiments, the computer device determines the candidate transfer learning model to which the target network layer belongs as the target transfer learning model for the training sample set; further, based on the position of the target network layer in the target transfer learning model, it determines the target network layer and other network layers before the target network layer as the first network layer, and determines other network layers after the target network layer as the second network layer; then, it initializes the parameters of the second network layer to obtain the third network layer, and constructs a transfer learning model for the training sample set according to the first network layer and the third network layer. Exemplarily, as Figure 5 shown, the computer device obtains the candidate transfer learning model to which the target network layer belongs, directly transfers the target network layer and other network layers before the target network layer to the transfer learning model for the training sample set, and randomly initializes and then transfers other network layers after the target network layer to the transfer learning model for the training sample set.
[0115] In another possible implementation, the computer device constructs the above-mentioned transfer learning model only based on the target network layer. In some embodiments, after determining the above-mentioned target network layer, the computer device determines the candidate transfer learning model to which the target network layer belongs as the target transfer learning model for the training sample set, and determines the target network layer in the target transfer learning model and other networks located before the target network layer as the first network layer. Further, a fourth network layer is constructed, and a transfer learning model for the training sample set is constructed based on the first network layer and the fourth network layer. Exemplarily, as Figure 6 shown, after determining the above-mentioned target network layer, the computer device determines the candidate transfer learning model to which the target network layer belongs as the target transfer learning model for the training sample set, and directly transfers the target network layer in the target transfer learning model and other network layers located before the target network layer to the transfer learning model for the training sample set, and constructs other network layers located after the target network layer in the transfer learning model.
[0116] In summary, in the technical solution provided by the embodiments of the present application, according to the transfer rate of the candidate network layer, the transfer learning effect is determined with the network layer as the basic unit, which improves the judgment accuracy of the transfer rate. Subsequently, a transfer learning model is constructed with the network layer as the basic unit, which improves the construction efficiency of the transfer learning model, reduces the difference between the initially constructed transfer learning model and the finally trained transfer learning model from the side, and improves the training efficiency of the transfer learning model.
[0117] Next, taking the image classification task as an example, the method for determining the transfer learning model in the present application will be introduced.
[0118] Please refer to Figure 7 , which shows a flowchart of a method for determining a transfer learning model provided by another embodiment of the present application. Each step in this method can be executed by the terminal device 10 and / or the server 20 (collectively referred to as "computer device") in the above Figure 1 . This method may include at least one of the following steps (701 to 712):
[0119] Step 701, determining at least one first candidate transfer learning model based on the training task of the first training sample set; wherein, the first training sample set includes training sample images belonging to different categories.
[0120] In some embodiments, the above-mentioned training sample images may be images that can reflect diseases in the medical field, commodity images in the shopping field, or book cover images in the education field, etc. The embodiments of the present application do not limit this.
[0121] Step 702: Sample each network layer included in each first candidate transfer learning model to obtain a plurality of first candidate network layers. One first candidate transfer learning model corresponds to one or more first candidate network layers.
[0122] Step 703: Process each training sample image in the first training sample set based on the feature extraction function of the first candidate network layer to obtain a feature vector corresponding to each training sample image.
[0123] Step 704: Determine a first sample feature matrix according to the feature vectors corresponding to each training sample image.
[0124] Step 705: Determine the first sample coding information entropy according to the first sample feature matrix.
[0125] Step 706: Determine a first category feature matrix corresponding to each category according to the feature vectors of the training sample images corresponding to each category.
[0126] Step 707: Determine the first category coding information entropy corresponding to each category according to the first category feature matrix corresponding to each category.
[0127] Step 708: Determine the transfer rate of the first candidate network layer according to the first sample coding information entropy and the first category coding information entropy corresponding to each category.
[0128] Step 709: Based on the transfer rates corresponding to each first candidate network layer, determine the first candidate network layers whose transfer rates meet the first condition as the first target network layers.
[0129] Step 710: Determine the first candidate transfer learning model to which the first target network layer belongs as the target transfer learning model for the first training sample set.
[0130] Step 711: Based on the position of the first target network layer in the target transfer learning model, determine the first target network layer and other network layers before the first target network layer as the first network layer, and determine other network layers after the first target network layer as the second network layer.
[0131] Step 712: Construct a transfer learning model for the first training sample set according to the first network layer and the third network layer.
[0132] In addition, taking the classification task of chemical molecular structures as an example, the method for determining the transfer learning model in this application is introduced.
[0133] Please refer to Figure 8 , which shows a flowchart of the method for determining the transfer learning model provided by another embodiment of this application. Each step in this method can be performed by the aboveFigure 1 The terminal device 10 and / or the server 20 (hereinafter collectively referred to as "computer device") in
[0134] Step 801: Determine at least one second candidate transfer learning model based on the training tasks of the second training sample set; wherein, the second training sample set includes chemical molecules belonging to different structural categories.
[0135] Step 802: Sample the network layers included in each second candidate transfer learning model to obtain a plurality of second candidate network layers. One second candidate transfer learning model corresponds to one or more second candidate network layers.
[0136] Step 803: Process each training sample image of the second training sample set based on the feature extraction function of the second candidate network layer to obtain a feature vector corresponding to each training sample image.
[0137] Step 804: Determine the second sample feature matrix according to the feature vectors corresponding to each training sample image.
[0138] Step 805: Determine the second sample coding information entropy according to the second sample feature matrix.
[0139] Step 806: Determine the second category feature matrix corresponding to each structural category according to the feature vectors of the training sample images corresponding to each structural category.
[0140] Step 807: Determine the second category coding information entropy corresponding to each structural category according to the second category feature matrix corresponding to each structural category.
[0141] Step 808: Determine the transfer rate of the second candidate network layer according to the second sample coding information entropy and the second category coding information entropy corresponding to each structural category.
[0142] Step 809: Based on the transfer rates corresponding to each second candidate network layer, determine the second candidate network layer with the transfer rate meeting the first condition as the second target network layer.
[0143] Step 810: Determine the second candidate transfer learning model to which the second target network layer belongs as the target transfer learning model for the second training sample set.
[0144] Step 811: Based on the position of the second target network layer in the target transfer learning model, determine the second target network layer and other network layers before the second target network layer as the first network layer, and determine other network layers after the second target network layer as the second network layer.
[0145] Step 812, construct a transfer learning model for the second training sample set according to the first network layer and the third network layer.
[0146] It should be noted that the above takes the image classification task and the classification task of chemical molecular structures as examples to introduce the method for determining the transfer learning model in this application. In the exemplary embodiment, the method for determining the transfer learning model can also be used for text recognition, keyword extraction, etc., and the embodiments of this application do not limit this.
[0147] The following is an embodiment of the apparatus of this application, which can be used to execute the method embodiment of this application. For details not disclosed in the embodiment of the apparatus of this application, please refer to the method embodiment of this application.
[0148] Please refer to Figure 9 , which shows a block diagram of a device for determining a transfer learning model provided by an embodiment of this application. This device has the function of implementing the above method for determining the transfer learning model. This function can be implemented by hardware or by hardware executing corresponding software. This device can be a computer device or can be set in a computer device. The device 900 may include: a network layer determination module 910, a sample processing module 920, a transfer rate determination module 930, and a model construction module 940.
[0149] The network layer determination module 910 is configured to determine a plurality of candidate network layers from at least one candidate transfer learning model, where one candidate transfer learning model corresponds to at least one candidate network layer; among them, different candidate transfer learning models are models trained based on different training data.
[0150] The sample processing module 920 is configured to process the training sample set based on the candidate network layer to obtain the sample coding information entropy and the class coding information entropy corresponding to each of the plurality of classes. The training sample set includes training samples belonging to different classes; among them, the sample coding information entropy is used to indicate the amount of information contained in the training samples in the training sample set after coding, and the class coding information entropy corresponding to the class is used to indicate the amount of information contained in the training samples belonging to the class in the training sample set after coding.
[0151] The transfer rate determination module 930 is configured to determine the transfer rate of the candidate network layer according to the sample coding information entropy and the plurality of class coding information entropies, and the transfer rate is used to indicate the transfer learning effect of the candidate network layer for the training sample set.
[0152] The model construction module 940 is configured to construct a transfer learning model for the training sample set based on the transfer rates corresponding to each of the candidate network layers, according to the candidate network layers whose transfer rates meet the first condition.
[0153] In some embodiments, as Figure 10 shown, the sample processing module 920 includes: a matrix acquisition sub-module 921 and an information entropy determination sub-module 922.
[0154] The matrix acquisition sub-module 921 is configured to process the training sample set based on the candidate network layer to obtain a sample feature matrix and class feature matrices corresponding to multiple classes respectively.
[0155] The information entropy determination sub-module 922 is configured to determine the sample coding information entropy according to the sample feature matrix.
[0156] The information entropy determination sub-module 922 is further configured to determine the class coding information entropy corresponding to each class according to the class feature matrix corresponding to each class.
[0157] In some embodiments, as Figure 10 shown, the matrix acquisition sub-module 921 is configured to:
[0158] Process each training sample in the training sample set based on the feature extraction function of the candidate network layer to obtain a feature vector corresponding to each training sample;
[0159] Construct the sample feature matrix according to the feature vectors corresponding to each training sample; wherein, the data in the first target column of the sample feature matrix is the feature vector of the first target training sample;
[0160] For at least one training sample belonging to the target class in the training sample set, construct the class feature matrix corresponding to the target class according to the feature vectors corresponding to each training sample belonging to the target class; wherein, the data in the second target column of the class feature matrix corresponding to the target class is the feature vector of the second target training sample belonging to the target class.
[0161] In some embodiments, as Figure 10 shown, the information entropy determination sub-module 922 is configured to:
[0162] Obtain the dimension of the feature vector corresponding to the training sample and the coding precision rate for the training sample;
[0163] Determine the coding length required to compress the sample feature matrix into the coding indicated by the coding precision rate according to the dimension of the feature vector corresponding to the training sample and the coding precision rate for the training sample;
[0164] Determine the sample coding information entropy based on the coding length.
[0165] In some embodiments, such as Figure 10 shown, the network layer determination module 910 includes: a model determination sub-module 911 and a network layer acquisition sub-module 912.
[0166] The model determination sub-module 911 is configured to determine, based on the training task of the training sample set, the training model corresponding to the associated task associated with the training task as the candidate transfer learning model;
[0167] The network layer acquisition sub-module 912 is configured to sample the network layers included in the candidate transfer learning model to obtain at least one candidate network layer corresponding to the candidate transfer learning model.
[0168] In some embodiments, such as Figure 10 shown, the network layer acquisition sub-module 912 is configured to:
[0169] Determine, from the candidate transfer learning model, the network layers whose position distance from the output layer is less than a threshold as the network layers to be sampled; sample and obtain the candidate network layers from the network layers to be sampled;
[0170] Or,
[0171] Obtain at least one network layer included in the candidate transfer learning model, and the importance degree of each network layer in the candidate transfer model; based on the importance degree, determine the network layers whose importance degree meets the second condition from the candidate transfer learning model as the candidate network layers.
[0172] In some embodiments, the model construction module 940 is configured to:
[0173] Based on the transfer rate corresponding to each candidate network layer respectively, determine the candidate network layers whose transfer rate meets the first condition as the target network layers;
[0174] Determine the candidate transfer learning model to which the target network layer belongs as the target transfer learning model for the training sample set;
[0175] Based on the position of the target network layer in the target transfer learning model, determine the target network layer and other network layers before the target network layer as the first network layer, and determine other network layers after the target network layer as the second network layer;
[0176] Perform initialization processing on the parameters of the second network layer to obtain a third network layer;
[0177] Construct a transfer learning model for the training sample set according to the first network layer and the third network layer.
[0178] In summary, in the technical solution provided by the embodiments of the present application, the migration rate of the candidate network layer is determined through the sample coding information entropy and the category coding information entropy, so as to determine the transfer learning effect of the candidate network layer for the training sample set, providing a way to evaluate the transfer learning effect before transfer learning. Without transfer learning, it is possible to evaluate the transfer learning effect of the candidate network layer on the training sample set, so as to quickly and accurately determine the optimal candidate network layer for the training sample set. Moreover, according to the migration rate of the candidate network layer, the transfer learning effect is determined with the network layer as the basic unit, improving the judgment accuracy of the migration rate. For the same candidate transfer learning model, the optimal network layer suitable for transfer learning in the candidate transfer learning model can be determined. Further, in the subsequent process, a transfer learning model can be constructed with the network layer as the basic unit, which reduces the difference between the initially constructed transfer learning model and the finally trained transfer learning model from the side and improves the training efficiency of the transfer learning model.
[0179] It should be noted that when the device provided in the above embodiment realizes its functions, only the division of the above functional modules is used for illustration. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the method embodiment belong to the same concept, and the specific implementation process is detailed in the method embodiment and will not be repeated here.
[0180] Please refer to Figure 11 , which shows the structural block diagram of a computer device provided by an embodiment of the present application. This computer device can be used to implement the functions of the above method for determining a transfer learning model. Specifically:
[0181] The computer device 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including a random access memory (RAM) 1102 and a read only memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing unit 1101. The computer device 1100 also includes a basic input / output system (Input / Output, I / O system) 1106 for facilitating the transfer of information between various components within the computer, and a mass storage device 1107 for storing an operating system 1113, application programs 1114, and other program modules 1115.
[0182] The basic input / output system 1106 includes a display 1108 for displaying information and input devices 1109 such as a mouse, keyboard, etc. for user input of information. Both the display 1108 and the input devices 1109 are connected to the central processing unit 1101 through an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may also include an input / output controller 1110 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1110 also provides outputs to a display screen, printer, or other types of output devices.
[0183] The mass storage device 1107 is connected to the central processing unit 1101 through a mass storage controller (not shown) connected to the system bus 1105. The mass storage device 1107 and its associated computer-readable medium provide non-volatile storage for the computer device 1100. That is to say, the mass storage device 1107 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0184] Without loss of generality, computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, EPROM (Erasable Programmable Read Only Memory), EEPROM (Electrically Erasable Programmable Read Only Memory), flash memory, or other solid-state storage devices, CD-ROM, DVD (Digital Video Disc), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art know that computer storage media is not limited to the above several types. The above-mentioned system memory 1104 and mass storage device 1107 may be collectively referred to as memory.
[0185] According to various embodiments of the present application, the computer device 1100 can also run on a remote computer on the network through a network such as the Internet. That is, the computer device 1100 can be connected to the network 1112 through the network interface unit 1111 connected to the system bus 1105. Or rather, the network interface unit 1111 can also be used to connect to other types of networks or remote computer systems (not shown).
[0186] The memory further includes a computer program, which is stored in the memory and is configured to be executed by one or more processors to implement the above-mentioned method for determining the transfer learning model.
[0187] In an exemplary embodiment, there is also provided a computer-readable storage medium storing at least one instruction, at least one segment of program, a set of codes or a set of instructions, which, when executed by a processor, implement the above-mentioned method for determining the transfer learning model.
[0188] Optionally, the computer-readable storage medium may include: ROM (Read Only Memory), RAM (Random Access Memory), SSD (Solid State Drives), or optical discs, etc. Among them, the random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0189] In an exemplary embodiment, there is also provided a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned method for determining the transfer learning model.
[0190] It should be understood that the "plurality" mentioned herein refers to two or more. "And / or" describes the association relationship of associated objects and indicates that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after. In addition, the step numbers described herein only exemplarily show a possible execution sequence between steps. In some other embodiments, the above steps may not be executed in the numbered order. For example, two steps with different numbers can be executed simultaneously, or two steps with different numbers can be executed in the reverse order of the illustration. The embodiments of the present application do not limit this.
[0191] The above are only exemplary embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for determining a transfer learning model, characterized in that, The method is executed by a computer device, and the method includes: Determine a plurality of candidate network layers from at least one candidate transfer learning model, where one candidate transfer learning model corresponds to at least one candidate network layer; among them, different candidate transfer learning models are models trained based on different training data; Process the training sample set based on the candidate network layers to obtain the sample coding information entropy and the class coding information entropy corresponding to each of the plurality of classes. The training sample set includes training sample images belonging to different classes; among them, the sample coding information entropy is used to indicate the amount of information contained in the training sample images in the training sample set after coding, and the class coding information entropy corresponding to the class is used to indicate the amount of information contained in the training sample images belonging to the class in the training sample set after coding; Determine the transfer rate of the candidate network layer according to the sample coding information entropy and the plurality of class coding information entropies. The transfer rate is used to indicate the transfer learning effect of the candidate network layer for the training sample set; Based on the transfer rates corresponding to each of the candidate network layers, construct a transfer learning model for the training sample set according to the candidate network layers whose transfer rates meet the first condition. The training sample images in the training sample set are used to train the transfer learning model so that the transfer learning model performs an image classification task.
2. The method according to claim 1, wherein The processing the training sample set based on the candidate network layers to obtain the sample coding information entropy and the class coding information entropy corresponding to each of the plurality of classes includes: Process the training sample set based on the candidate network layers to obtain a sample feature matrix and class feature matrices corresponding to each of the plurality of classes; Determine the sample coding information entropy according to the sample feature matrix; Determine the class coding information entropy corresponding to each of the classes according to the class feature matrix corresponding to each of the classes.
3. The method according to claim 2, wherein The processing the training sample set based on the candidate network layers to obtain a sample feature matrix and class feature matrices corresponding to each of the plurality of classes includes: Process each training sample image in the training sample set based on the feature extraction function of the candidate network layer to obtain a feature vector corresponding to each training sample image; Construct the sample feature matrix according to the feature vectors corresponding to each training sample image; among them, the data in the first target column of the sample feature matrix is the feature vector of the first target training sample image; For at least one training sample image belonging to the target class in the training sample set, construct the class feature matrix corresponding to the target class according to the feature vectors corresponding to each training sample image belonging to the target class; among them, the data in the second target column of the class feature matrix corresponding to the target class is the feature vector of the second target training sample image belonging to the target class.
4. The method according to claim 2, wherein The determining the sample coding information entropy according to the sample feature matrix includes: Obtain the dimension of the feature vector corresponding to the training sample image and the coding accuracy rate for the training sample image; Determine the coding length required to compress the sample feature matrix into the coding indicated by the coding accuracy according to the dimension of the feature vector corresponding to the training sample image and the coding accuracy for the training sample image; Determine the sample coding information entropy based on the coding length.
5. The method according to claim 1, characterized in that, The determining of multiple candidate network layers from at least one candidate transfer learning model includes: Based on the training task of the training sample set, determine the training model corresponding to the associated task associated with the training task as the candidate transfer learning model; Sample the network layers included in the candidate transfer learning model to obtain at least one candidate network layer corresponding to the candidate transfer learning model.
6. The method according to claim 5, wherein The sampling of the network layers included in the candidate transfer learning model to obtain at least one candidate network layer corresponding to the candidate transfer learning model includes: From the candidate transfer learning model, determine the network layers whose position distance from the output layer is less than a threshold as the network layers to be sampled; sample the candidate network layers from the network layers to be sampled; Or, Obtain at least one network layer included in the candidate transfer learning model and the importance degree of each network layer in the candidate transfer learning model; based on the importance degree, determine the network layers whose importance degree satisfies the second condition from the candidate transfer learning model as the network layers to be sampled; sample the candidate network layers from the network layers to be sampled.
7. The method according to any one of claims 1 to 6, characterized in that, The constructing of a transfer learning model for the training sample set according to the transfer rates respectively corresponding to the candidate network layers, including: Based on the transfer rates respectively corresponding to the candidate network layers, determine the candidate network layers whose transfer rates satisfy the first condition as the target network layers; Determine the candidate transfer learning model to which the target network layer belongs as the target transfer learning model for the training sample set; Based on the position of the target network layer in the target transfer learning model, determine the target network layer and other network layers before the target network layer as the first network layer, and determine the other network layers after the target network layer as the second network layer; Perform initialization processing on the parameters of the second network layer to obtain a third network layer; Construct a transfer learning model for the training sample set according to the first network layer and the third network layer.
8. An apparatus for determining a transfer learning model, characterized in that The device is a computer device or is set in a computer device, and the device includes: A network layer determination module, configured to determine multiple candidate network layers from at least one candidate transfer learning model, where one candidate transfer learning model corresponds to at least one candidate network layer; wherein, different candidate transfer learning models are models trained based on different training data; A sample processing module, configured to process a training sample set based on the candidate network layer to obtain a sample coding information entropy and class coding information entropies corresponding to multiple classes respectively, where the training sample set includes training sample images belonging to different classes; wherein, the sample coding information entropy is used to indicate the amount of information contained in the training sample images in the training sample set after coding, and the class coding information entropy corresponding to the class is used to indicate the amount of information contained in the training sample images belonging to the class in the training sample set after coding; A migration rate determination module, configured to determine the migration rate of the candidate network layer according to the sample coding information entropy and the multiple class coding information entropies, where the migration rate is used to indicate the transfer learning effect of the candidate network layer for the training sample set; A model construction module, configured to construct a transfer learning model for the training sample set based on the migration rates corresponding to the respective candidate network layers, and construct a transfer learning model for the training sample set according to the candidate network layers whose migration rates meet the first condition, where the training sample images in the training sample set are used to train the transfer learning model so that the transfer learning model performs an image classification task.
9. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one program is stored in the memory, and the at least one program is loaded and executed by the processor to implement the method for determining the transfer learning model according to any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that, At least one program is stored in the storage medium, and the at least one program is loaded and executed by a processor to implement the method for determining the transfer learning model according to any one of claims 1 to 7.
11. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, and the processor reads and executes the computer instructions from the computer-readable storage medium to implement the method for determining the transfer learning model according to any one of claims 1 to 7.
Citation Information
Patent Citations
Method of converting gesture recognition to identification recognition on the basis of feature transfer learning
CN108960171A
Image processing method and apparatus, electronic device, and storage medium
WO2020199619A1