Cross-scenario migration method and device for a highway disease classification model
By using SimCLR framework and dynamic knowledge distillation technology in the transportation field, deep learning models are trained and merged, and the problem of poor cross-scene migration effect of the highway apparent disease classification model in the situation of insufficient target domain data and incomplete labeling is solved, and effective model migration in the transportation field is achieved.
Patent Information
- Application Number
- CN202310200906.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-02-22
AI Technical Summary
The prior art cannot effectively realize cross-scene migration of deep learning classification models for road apparent diseases in the field of transportation, especially when the target scene data is small and partially annotated, the model effect is not good.
Using the SimCLR framework, the SimCLR framework uses the source domain data set to train students' network feature extractor and classifier, freezes the feature extractor parameters, dynamically distillates the source domain and target domain data to train the feature extractor, and uses the target domain labeled data to train the classifier, and finally merges the model for cross-scene migration.
With the small amount of target domain data and partial annotation, a good cross-scene migration effect is maintained, solving the problem of poor model effect in the prior art.
Smart Images

Figure CN116204827B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of continuous learning, and in particular, to a cross-scene migration method and device for a highway disease classification model. Background Art
[0002] In the case of sufficient data sets, deep learning has good capabilities for detecting and classifying highway apparent diseases. Combining a trained deep learning model with in-vehicle edge devices can detect and record highway apparent diseases in real time, facilitating subsequent corresponding maintenance work on the highway by relevant staff. However, due to the different apparent conditions of different highways, a deep learning model with satisfactory effects depends on a large amount of high-quality highway disease data and computing resources, and has certain requirements for the professional knowledge level of annotators. Currently, in the transportation field, the cross-scene migration effect of deep learning classification models for highway apparent diseases is still not satisfactory.
[0003] The amount of data in the current target scene is small, and due to human and time constraints, most of it has not been manually annotated. Therefore, it is necessary to find a model cross-scene migration method for the transportation field, which can still maintain good effects when migrating the source scene model to the target scene with less data and partial annotation in the target scene.
[0004] The above problems need to be solved urgently. Summary of the Invention
[0005] The present invention aims to overcome at least one of the above disadvantages of the prior art. On the one hand, it provides a cross-scene migration method for a highway disease classification model. The method includes the steps: S110: Obtain the SimCLR framework; S120: Based on the SimCLR framework, use the source domain data set to train the original student network feature extractor to generate the first student network feature extractor; S130: Freeze the parameters of the first student network feature extractor and use the source domain data set to train the original student network classifier to generate the first student network classifier; S140: Use the source domain data set and the target domain data set to train the first student network feature extractor to generate the second student network feature extractor; S150: Use the labeled data in the target domain data set to train the first student network classifier to generate the second student network classifier; S160: Combine the second student network feature extractor and the second student network classifier into a migrated model; S170: Based on the migrated model, classify diseases of highways in different scenes.
[0006] Optionally, step S120 includes: S1201: generating a student network with a network pre-trained on ImageNet; S1202: respectively generating first source domain data and second source domain data from the source domain dataset through different data augmentations; S1203: the first source domain data and the second source domain data respectively generate first features and second features through the feature extractor of the original student network, and the first features and the second features respectively generate first classification results and second classification results through the classifier of the original student network; S1204: calculating a loss function L using the first classification result and the second classification result FE ; S1205: using the loss function L FE to update the original student network feature extractor to generate a first student network feature extractor.
[0007] Optionally, the data augmentation includes strong data augmentation and / or weak data augmentation; the strong data augmentation includes one or a combination of Gaussian blur and random grayscale transformation; the weak data augmentation includes one or a combination of cropping and horizontal flipping.
[0008] Optionally, step S130 includes: S1301: freezing the parameters of the first student network feature extractor; S1302: inputting the source domain dataset into the frozen first student network feature extractor, and generating a first prediction result after passing through the frozen first student network feature extractor and the original student network classifier; S1303: generating a loss function L based on the first prediction result and the source domain data label BCE-1 ; S1304: updating the original student network classifier based on the loss function L BCE-1 to generate a first student network classifier.
[0009] Optionally, step S140 includes: S1401: dividing the target domain dataset into labeled target domain data and unlabeled target domain data; S1402: copying the student network to generate a teacher network; S1403: hiding the labels of the target domain labeled data and mixing them with the unlabeled target domain data to generate a dataset T; S1404: generating a dataset T after weak data augmentation of the dataset T W , and generating a dataset T after strong data augmentation of the dataset T S ; S1405: replacing the classifiers of the teacher network and the student network with MLP layers; S1406: respectively inputting the dataset T W and the dataset T S into the teacher network and the student network, and denoting the outputs of the first M layers of the teacher network and the student network as feature1 and feature2 respectively; S1407: using the results after feature1 and feature2 respectively passing through the paraphraser P and the translator R as the loss function L HT; S1408: Based on the loss function L HT Update the first M layers of the student network; S1409: Denote the output of the teacher network as Denote the output of the student network as Use And Generate the loss function L DK ; S14010: Based on the loss function L DK Update the first student network feature extractor to generate the second student network feature extractor.
[0010] Optionally, the step S14010 further includes: Based on the loss function L DK While updating the first student network feature extractor to generate the second student network feature extractor, dynamically update the feature extractor of the teacher network.
[0011] Optionally, the step S150 includes: S1501: Use the source domain and target domain data to train the first student network feature extractor by the dynamic knowledge distillation method to generate the second student network feature extractor.
[0012] Optionally, the step S1501 includes: S15011: Initialize the classifier of the first student network, and freeze the parameters of the second student network feature extractor; S15012: Input the labeled data in the target domain into the feature extractor of the frozen second student network, output the second prediction result, and use the second prediction result and the label of the labeled data in the target domain as the loss function L BCE_ To update the initialized classifier of the first student network to generate the second student network classifier.
[0013] Optionally, the loss function is: Where m represents the total number of samples, n represents the total number of categories, P ij Represents whether the i-th sample belongs to the j-th category. If so, it is 1, otherwise it is 0, and q ij Represents the probability that the model predicts that the i-th sample belongs to the j-th category.
[0014] On the other hand, the present invention provides a cross-scenario migration device for a highway disease classification model. The device includes: an acquisition unit for acquiring the SimCLR framework; a first student network feature extractor generation unit for training an original student network feature extractor using a source domain dataset based on the SimCLR framework to generate a first student network feature extractor; a first student network classifier generation unit for freezing the parameters of the first student network feature extractor and training an original student network classifier using the source domain dataset to generate a first student network classifier; a second student network feature extractor generation unit for training the first student network feature extractor using the source domain dataset and a target domain dataset to generate a second student network feature extractor; a second student network classifier generation unit for training the first student network classifier using labeled data of the target domain dataset to generate a second student network classifier; a merging unit for merging the second student network feature extractor and the second student network classifier into a migrated model; and a classification unit for classifying highway diseases in different scenarios based on the migrated model.
[0015] In another aspect, the present invention provides a computer-readable storage medium storing one or more instructions, and the computer instructions are used to cause the computer to execute the above-mentioned cross-scenario migration method of the highway disease classification model.
[0016] In yet another aspect, the present invention provides an electronic device, including: a memory and a processor; at least one program instruction is stored in the memory; and the processor loads and executes the at least one program instruction to implement the above-mentioned cross-scenario migration method of the highway disease classification model.
[0017] The beneficial effects of the present invention are as follows: The present invention provides a cross-scenario migration method for a highway disease classification model. The method includes the steps: S110: Obtain the SimCLR framework; S120: Based on the SimCLR framework, use the source domain dataset to train the original student network feature extractor to generate the first student network feature extractor; S130: Freeze the parameters of the first student network feature extractor and use the source domain dataset to train the original student network classifier to generate the first student network classifier; S140: Use the source domain dataset and the target domain dataset to train the first student network feature extractor to generate the second student network feature extractor; S150: Use the labeled data of the target domain dataset to train the first student network classifier to generate the second student network classifier; S160: Combine the second student network feature extractor and the second student network classifier into the migrated model; S170: Based on the migrated model, classify the diseases of highways in different scenarios. By obtaining the SimCLR framework, using the source domain to train the student network feature extractor, freezing the parameters of the student feature extractor and using the source domain to train the student network classifier, dynamically distilling and training the student network feature extractor using the source domain and the target domain data, training the student network classifier using the labeled data of the target domain and combining the models, the problem in the prior art that the model effect after cross-scenario migration cannot be maintained when the target domain data in the traffic field is small and partially labeled is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The present invention will be further described below with reference to the drawings and embodiments.
[0019] Figure 1 It is a flowchart of a cross-scenario migration method for a highway disease classification model provided by an embodiment of the present invention.
[0020] Figure 2 It is a schematic diagram of the process of generating the first student network feature extractor provided by an embodiment of the present invention.
[0021] Figure 3 It is a schematic diagram of the process of generating the first student network classifier provided by an embodiment of the present invention.
[0022] Figure 4 It is a schematic diagram of the process of generating the second student network feature extractor provided by an embodiment of the present invention.
[0023] Figure 5 It is a schematic diagram of the process of generating the second student network classifier provided by an embodiment of the present invention.
[0024] Figure 6 It is a schematic diagram of a cross-scenario migration device for a highway disease classification model provided by an embodiment of the present invention.
[0025] Figure 7It is a partial block diagram of the electronic device provided by the embodiments of the present invention. Detailed implementation manners
[0026] Before discussing the exemplary embodiments in more detail, it should be noted that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations as sequential processes, many of the operations can be implemented in parallel, concurrently, or simultaneously. In addition, the order of the operations can be rearranged. The process can be terminated when its operations are completed, but it can also have additional steps not included in the drawings. The process can correspond to a method, function, procedure, subroutine, subprogram, etc.
[0027] It should be understood that although terms such as "first" and "second" may be used herein to describe various units, these units should not be limited by these terms. These terms are only used to distinguish one unit from another. For example, without departing from the scope of the exemplary embodiments, the first unit can be called the second unit, and similarly, the second unit can be called the first unit. The term "and / or" used herein includes any and all combinations of one or more of the listed related items.
[0028] For ease of understanding, the following professional terms are explained herein:
[0029] SimCLR framework: Chen et.al et al. proposed a new framework in their research paper "SimCLR: A Simple Framework for Contrastive Learning of Visual Representations". The SimCLR framework is very simple. It takes an image and performs random transformations on it to obtain a pair of two augmented images. Each image in the pair passes through an encoder to obtain a representation. Then, a non-linear fully connected layer is applied to obtain a representation.
[0030] Transfer learning: (Transfer Learning) utilizes existing old knowledge to learn new knowledge. The core idea is to find the similarities between the old knowledge and the new knowledge. It is easier to learn new knowledge based on the old knowledge. Therefore, we generally use the old knowledge to assist in learning new knowledge. Transfer learning finds the similarities between different tasks and applies the knowledge of the old tasks to the new tasks.
[0031] Source domain: (Source Domain) The old knowledge that has been mastered is called the source domain.
[0032] Target domain: (Target Domain) The new knowledge to be learned is called the target domain.
[0033] Dynamic Knowledge Distillation: (Dynamic Distillation) is a research field for knowledge transfer and deep model compression. The core idea is that the knowledge contained in a trained complex model (teacher network) can be "distilled into" another small model. The small model has a simpler network structure than the large model, and at the same time, the distillation effect is similar to that of the large model. Therefore, knowledge distillation is also regarded as a model compression technology. The difference between dynamic knowledge distillation and knowledge distillation is that the teacher network in knowledge distillation is a static network, while the teacher network in dynamic knowledge distillation is a dynamic network.
[0034] For the convenience of subsequent understanding, the overall inventive concept of the present invention is a cross-scenario transfer optimization method for a highway disease classification model based on the SimCLR framework and dynamic distillation. Among them, the method solves the problem that the existing technology cannot maintain the model effect after cross-scenario transfer in the case of less target domain data and partial annotation in the traffic field by obtaining the SimCLR framework to train the student network feature extractor using the source domain, freezing the parameters of the student feature extractor to train the student network classifier using the source domain, dynamically distilling and training the student network feature extractor using the source domain and target domain data, and training the student network classifier using the labeled data in the target domain and merging the models.
[0035] Embodiment 1
[0036] Reference Figure 1 , showing a flowchart of a cross-scenario transfer method for a highway disease classification model. The method includes the steps:
[0037] S110: Obtain the SimCLR framework.
[0038] As an example, the SimCLR framework can be directly obtained and is very mature in the prior art, so it will not be elaborated here.
[0039] S120: Based on the SimCLR framework, use the source domain dataset to train the original student network feature extractor to generate the first student network feature extractor.
[0040] As an example, the step S120 includes: S1201: Generate a student network with a network pre-trained on ImageNet; S1202: Respectively generate the first source domain data and the second source domain data by performing different data augmentations on the source domain dataset; S1203: The first source domain data and the second source domain data pass through the feature extractor of the original student network to generate the first feature and the second feature respectively, and the first feature and the second feature pass through the classifier of the original student network to generate the first classification result and the second classification result respectively; S1204: Calculate the loss function L using the first classification result and the second classification result FE ; S1205: Use the loss function L FEThe original student network feature extractor is updated to generate a first student network feature extractor.
[0041] Optional, such as Figure 2 As shown in the figure, it is a schematic diagram of the process of generating the first student network feature extractor. S represents the source domain data, and S1 and S2 are generated after different data enhancements. S1 and S2 are generated by the feature extractor f of the original student network. student Then output the features separately and Then the classifier g of the original student network outputs the classification results respectively and according to and Calculate the loss function L FE , using L FE Update the feature extractor f of the original student network student . Wherein, the data enhancement includes strong data enhancement and / or weak data enhancement; the strong data enhancement includes one or a combination of Gaussian blur and random grayscale transformation; the weak data enhancement includes one or a combination of cropping and horizontal flipping.
[0042] S130: Freeze the first student network feature extractor parameters and use the source domain data set to train the original student network classifier to generate a first student network classifier.
[0043] As an example, step S130 includes: S1301: freezing the first student network feature extractor parameters; S1302: inputting the source domain data set into the frozen first student network feature extractor, and generating a first prediction result after the frozen first student network feature extractor and the original student network classifier; S1303: generating a loss function L based on the first prediction result and the source domain data label BCE-1 ; S1304: Based on the loss function L BCE-1 The original student network classifier is updated to generate the first student network classifier.
[0044] Optional, such as Figure 3 As shown in the figure, it is a schematic diagram of the process of generating the first student network classifier. S also represents the source domain data, through the frozen first student network feature extractor f′ student The output result Prediction1 and its label label of the original student network classifier g S Assume supervised loss function L BCE_1 The original student network classifier is updated to generate a first student network classifier. The above method can enhance the feature extraction capability of the student model for the source domain.
[0045] S140: Generate a second student network feature extractor by training the first student network feature extractor using the source domain dataset and the target domain dataset.
[0046] As an example, step S140 includes: S1401: Divide the target domain dataset into labeled target domain data and unlabeled target domain data; S1402: Copy the student network to generate a teacher network; S1403: Mix the target domain labeled data with the labels hidden and the unlabeled target domain data to generate dataset T; S1404: Generate dataset T after weakly augmenting dataset T W , and generate dataset T after strongly augmenting dataset T S ; S1405: Replace the classifiers of the teacher network and the student network with an MLP layer; S1406: Feed dataset T W and dataset T S into the teacher network and the student network respectively, and denote the outputs of the first M layers of the teacher network and the student network as feature1 and feature2; S1407: Use the results of feature1 and feature2 after passing through the paraphraser P and the translator R respectively as the loss function L HT ; S1408: Update the first M layers of the student network based on the loss function L HT ; S1409: Denote the output of the teacher network as and the output of the student network as Use and to generate the loss function L DK ; S14010: Update the first student network feature extractor based on the loss function L DK to generate a second student network feature extractor.
[0047] Optionally, as shown in Figure 4 , it is a schematic diagram of the process for generating the second student network feature extractor. T W and T S represent the results of dataset T after weak data augmentation and after strong data augmentation respectively, and the teacher network and the student network have N layers. Denote the first M layers of the teacher network as W hint , and the first M layers of the student network as W guuded . Calculate the loss function L hint between the output features of W guided and W HT , and calculate the loss function L between the output of the student network and the output DK of the teacher network U. More specifically, replace the classifiers of the teacher network and the student network with an MLP layer (multi-layer perceptron). Feed T W and TS Input into the teacher network and the student network. The outputs of the first M layers of the teacher network and the student network are denoted as feature1 and feature2 respectively. Use the results after passing feature1 and feature2 through the paraphraser P and the translator R respectively as the loss function L HT Update the first M layers of the student network, using and the loss function L DK While updating the feature extractor of the student network, dynamically update the feature extractor of the teacher network. The MLP can fully cross different dimensions of the feature vector, improving the ability to obtain non-linear features and combined feature information.
[0048] Optionally, the paraphraser P and the translator R represent the paraphraser and the translator respectively. The internal structure of the paraphraser P is similar to that of an encoder and a decoder. Its role is to abstract the feature output of the teacher model into a teacher factor, and the teacher factor can be regarded as the representative of the feature output. Divide P into two parts: the encoder P-1 and the decoder P-2. The T-factor, which is the output of the data after passing through W hint and then through P-1, is called the teacher factor. The output feature-2 of the T-factor after passing through P-2 is the output of feature-1 after passing through P. If feature-1 and feature-2 are close, then the T-factor can be regarded as the core part of feature-1, that is, the most worthy part to learn. Use the reconstruction loss L REC between feature-1 and feature-2 to train P, where L REC = ||feature1 - P(feature1)|| 2 2.
[0049] If the teacher network generates t feature maps, then adjust the number of channels of the teacher factor to t×k, where k is called the paraphrasing rate. In actual work, only use P-1 in P to generate the teacher factor.
[0050] R is the translator, which abstracts the output of W guided into a student factor. Treat the translator R and the paraphraser P similarly. Use the encoder and the decoder to extract the student factor. Divide R into two parts: the encoder R-1 and the decoder R-2. The student factor S-factor is the core part of the student feature map. In actual work, only use R-1 in R to generate the student factor. The above method simplifies the feature layer of its output through the teacher factor, reduces the learning difficulty of the student network, and the generated student network feature extractor has a better effect.
[0051] S150: Train the first student network classifier using the labeled data in the target domain dataset to generate a second student network classifier.
[0052] As an example, step S150 includes: S1501: Train the first student network feature extractor using source domain and target domain data with a dynamic knowledge distillation method to generate a second student network feature extractor.
[0053] Optionally, step S1501 includes: S15011: Initialize the classifier of the first student network and freeze the parameters of the second student network feature extractor; S15012: Input the labeled data in the target domain into the feature extractor of the frozen second student network, output a second prediction result, and use the second prediction result and the label of the labeled data in the target domain as the loss function L BCE_ to update the initialized classifier of the first student network to generate a second student network classifier.
[0054] Optionally, as Figure 5 shown, it is a schematic diagram of the process for generating the second student network classifier. Initialize the classifier of the first student network as g s , freeze the parameters of the feature extractor f″ of the second student network student , denote the labeled data in the target domain as T L , input T L into the feature extractor f″ of the frozen second student network student , denote the output result as Prediction2, and use Prediction2 and the label of the labeled data in the target domain as the loss function L BCE_ to update g s to generate a second student network classifier.
[0055] S160: Combine the second student network feature extractor and the second student network classifier into a migrated model.
[0056] S170: Classify the diseases of roads in different scenarios based on the migrated model.
[0057] As an example, input the taken road pictures into the model after that to automatically generate the types of road diseases based on that road.
[0058] The above method obtains the student network feature extractor by using the SimCLR framework to train with the source domain, freezes the parameters of the student feature extractor and trains the student network classifier with the source domain, dynamically distills and trains the student network feature extractor using the source domain and target domain data, and trains the student network classifier with the labeled data in the target domain and merges the models, solving the problem that the existing technology cannot maintain the model effect after cross-scenario migration when the target domain data in the traffic field is scarce and partially labeled.
[0059] Embodiment 2
[0060] As Figure 6 shown, a cross-scenario migration device for a highway disease classification model is shown. The device includes:
[0061] An acquisition unit 510, configured to acquire the SimCLR framework.
[0062] A first student network feature extractor generation unit 520, configured to generate a first student network feature extractor by training an original student network feature extractor based on the SimCLR framework using a source domain dataset.
[0063] A first student network classifier generation unit 530, configured to freeze the parameters of the first student network feature extractor and train an original student network classifier using the source domain dataset to generate a first student network classifier.
[0064] A second student network feature extractor generation unit 540, configured to generate a second student network feature extractor by training the first student network feature extractor using a source domain dataset and a target domain dataset.
[0065] A second student network classifier generation unit 550, configured to generate a second student network classifier by training the first student network classifier using labeled data in the target domain dataset.
[0066] A merging unit 560, configured to merge the second student network feature extractor and the second student network classifier into a migrated model.
[0067] A classification unit 570, configured to classify highway diseases in different scenarios based on the migrated model.
[0068] Embodiment 3
[0069] An embodiment of the present invention also proposes a storage medium. A cross-scenario migration method for a highway disease classification model is stored on the storage medium. When the cross-scenario migration program of the highway disease classification model is executed by a processor, the steps of the cross-scenario migration method for the highway disease classification model as described above are implemented. Since this storage medium adopts all the technical solutions of the above all embodiments, it has at least all the beneficial effects brought by the technical solutions of the above embodiments, and will not be elaborated herein one by one.
[0070] Example 4
[0071] Please refer to Figure 7 , the embodiment of the present invention also provides an electronic device, including: a memory and a processor; at least one program instruction is stored in the memory; the processor loads and executes the at least one program instruction to implement the cross-scenario migration method of the highway disease classification model provided in Embodiment 1.
[0072] The memory 602 and the processor 601 are connected in a bus manner. The bus may include any number of interconnected buses and bridges, and the bus connects various circuits of one or more processors 601 and the memory 602 together. The bus can also connect various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art, so they will not be further described herein. The bus interface provides an interface between the bus and the transceiver. The transceiver may be an element or multiple elements, such as multiple receivers and transmitters, and provides a unit for communicating with various other devices on the transmission medium. The data processed by the processor 601 is transmitted on the wireless medium through the antenna. Further, the antenna also receives data and transmits the data to the processor 601.
[0073] The processor 601 is responsible for managing the bus and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. The memory 602 can be used to store the data used by the processor 601 when performing operations.
[0074] The above are only the embodiments of the present invention. Common knowledge such as specific structures and characteristics well known in the art are not described in detail herein. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention belongs before the application date or the priority date, can know all the prior art in this field, and have the ability to apply conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, combine their own abilities to complete and implement this solution. Some typical well-known structures or well-known methods should not be an obstacle for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can be made, which should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicability of the patent. The protection scope required by this application should be subject to the content of its claims, and the specific implementation manners described in the specification can be used to interpret the content of the claims.
[0075] The above are only embodiments of the present invention. Specific structures and common knowledge such as characteristics that are well-known in the art are not described in detail herein. Those of ordinary skill in the art know all the common general technical knowledge in the technical field to which the invention pertains before the filing date or the priority date, can learn all the prior art in this field, and have the ability to apply conventional experimental means before this date. Those of ordinary skill in the art can, under the inspiration given in this application, improve and implement this solution in combination with their own abilities. Some typical well-known structures or well-known methods should not become an obstacle for those of ordinary skill in the art to implement this application. It should be noted that for those skilled in the art, without departing from the structure of the present invention, several deformations and improvements can also be made, and these should also be regarded as the protection scope of the present invention, and these will not affect the implementation effect of the present invention and the practicality of the patent. The protection scope required by this application should be based on the content of its claims, and the specific implementation manners and the like recorded in the specification can be used to interpret the content of the claims.
Claims
1. A cross-scenario migration method for a highway disease classification model, characterized in that, The method includes the steps of: S110: Obtain the SimCLR framework; S120: Based on the SimCLR framework, use the source domain dataset to train the original student network feature extractor to generate the first student network feature extractor; S130: Freeze the parameters of the first student network feature extractor and use the source domain dataset to train the original student network classifier to generate the first student network classifier; S140: Use the source domain dataset and the target domain dataset to train the first student network feature extractor to generate the second student network feature extractor; S150: Use the labeled data in the target domain dataset to train the first student network classifier to generate the second student network classifier; S160: Combine the second student network feature extractor and the second student network classifier into the migrated model; S170: Based on the migrated model, classify the diseases of roads in different scenarios.
2. The cross-scenario migration method for a highway disease classification model according to claim 1, characterized in that, The step S120 includes: S1201: Generate a student network with a network pre-trained on ImageNet; S1202: Respectively generate the first source domain data and the second source domain data by performing different data augmentations on the source domain dataset; S1203: The first source domain data and the second source domain data respectively generate the first feature and the second feature through the feature extractor of the original student network, and the first feature and the second feature respectively generate the first classification result and the second classification result through the classifier of the original student network; S1204: Calculate the loss function L using the first classification result and the second classification result FE ; S1205: Use the loss function L FE Update the original student network feature extractor to generate a first student network feature extractor.
3. The cross-scenario migration method for a highway disease classification model according to claim 2, characterized in that, The data augmentation includes strong data augmentation and / or weak data augmentation; The strong data augmentation includes one or a combination of Gaussian blur and random grayscale transformation; The weak data augmentation includes one or a combination of cropping and horizontal flipping.
4. The cross-scenario migration method for a highway disease classification model according to claim 1, characterized in that, The step S130 includes: S1301: Freeze the parameters of the first student network feature extractor; S1302: Input the source domain dataset into the frozen first student network feature extractor, and generate the first prediction result after passing through the frozen first student network feature extractor and the original student network classifier; S1303: Generate a loss function L based on the first prediction result and the source domain data label BCE-1 ; S1304: Based on the loss function L BCE-1 Update the original student network classifier to generate a first student network classifier.
5. The cross-scenario migration method for a highway disease classification model according to claim 1, characterized in that, The step S140 includes: S1401: Divide the target domain dataset into labeled target domain data and unlabeled target domain data; S1402: Copy the student network to generate the teacher network; S1403: Mix the target domain labeled data with the labels hidden and the target domain unlabeled data to generate the dataset T; S1404: Generate dataset T by performing weak data augmentation on dataset T W , and generate dataset T by performing strong data augmentation on dataset T S ; S1405: Use the MLP layer to replace the classifiers of the teacher network and the student network; S1406: Respectively send the data set T W and the data set T S into the teacher network and the student network. The outputs of the first M layers of the teacher network and the student network are respectively denoted as feature1 and feature2; S1407: Use the results of feature1 and feature2 after passing through the paraphraser P and the translator R respectively as the loss function L HT ; S1408: Based on the loss function L HT Update the first M layers of the student network; S1409: Denote the output of the teacher network as Denote the output of the student network as Use and to generate the loss function L DK ; S14010: Based on the loss function L DK Update the first student network feature extractor to generate a second student network feature extractor.
6. The cross-scenario migration method of the highway disease classification model according to claim 5, wherein, The step S14010 further includes: Based on the loss function L DK While updating the first student network feature extractor to generate the second student network feature extractor, dynamically update the feature extractor of the teacher network.
7. The cross-scenario migration method of the highway disease classification model according to claim 1, wherein, The step S150 includes: S1501: Use the source domain and target domain data to train the first student network feature extractor by the dynamic knowledge distillation method to generate the second student network feature extractor.
8. The cross-scenario migration method of the highway disease classification model according to claim 7, wherein, The step S1501 includes: S15011: Initialize the classifier of the first student network and freeze the parameters of the second student network feature extractor; S15012: Input the labeled data in the target domain into the feature extractor of the frozen second student network to output the second prediction result, and use the second prediction result and the label of the labeled data in the target domain as the loss function L BCE_2 Update the classifier of the initialized first student network to generate the classifier of the second student network.
9. The cross-scenario migration method of the highway disease classification model according to any one of claims 2, 4, 5 and 8, wherein, The loss function is: Among them, m represents the total number of samples, n represents the total number of categories, P ij represents whether the i-th sample belongs to the j-th category. If so, it is 1; otherwise, it is 0. q ij represents the probability that the model predicts that the i-th sample belongs to the j-th category.
10. A cross-scenario migration device for a highway disease classification model, wherein, The device includes: An acquisition unit, used to obtain the SimCLR framework; A unit for generating the first student network feature extractor, used to train the original student network feature extractor based on the SimCLR framework using the source domain dataset to generate the first student network feature extractor; Generate a first student network classifier unit for freezing the parameters of the first student network feature extractor and training the original student network classifier using the source domain dataset to generate the first student network classifier; Generate a second student network feature extractor unit for training the first student network feature extractor using the source domain dataset and the target domain dataset to generate the second student network feature extractor; Generate a second student network classifier unit for training the first student network classifier using the labeled data of the target domain dataset to generate the second student network classifier; A merging unit for merging the second student network feature extractor and the second student network classifier into a migrated model; A classification unit for classifying the diseases of roads in different scenarios based on the migrated model.
Citation Information
Patent Citations
Speech classification network training method and device, computing equipment and storage medium
CN113593611A
Rolling bearing and gear fault diagnosis method based on semi-supervised model contrast migration
CN113723491A