Cross-domain object detection method based on domain adaptation

Through a cross-domain object detection method based on domain adaptation, the combination of CycleGAN and Faster RCNN networks is used to solve the problem of low target detection performance caused by the lack of large-scale annotation data in the target domain, and efficient target recognition and positioning in the target domain are achieved.

CN114626461BActive Publication Date: 2025-05-06XIAN UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210258271.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-16
Publication Date
2025-05-06
Estimated Expiration
2042-03-16

AI Technical Summary

Technical Problem

Existing object detection algorithms are difficult to effectively identify and locate targets in image fields without large-scale labeled data, resulting in low accuracy in prediction of instance categories and bounding box positions in the target domain.

Method used

The cross-domain object detection method based on domain adaptation is adopted to convert data domain domains through the CycleGAN network, convert source domain data with instance-level tags into target domain data, and combine Faster RCNN network for training and retraining of target detectors, and use complexity evaluation and pseudo-label iteration technology to improve detection performance.

Benefits of technology

The classification and prediction performance of target detection is significantly improved in the target domain, and can achieve high recognition accuracy without a large amount of labeled data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114626461B_ABST
    Figure CN114626461B_ABST
Patent Text Reader

Abstract

The present invention discloses a cross-domain object detection method based on domain adaptation, including: Step 1, obtaining an object detection data set including a source domain Ds and a target domain D T , performing data augmentation and data set expansion; Step 2, training a CycleGAN network using the expanded data set and outputting a generated data domain D G ; Step 3, constructing a Faster RCNN network as an object detector, and training the object detector using the source domain Ds and the generated data domain D G as a training set; Step 4, evaluating the complexity of the data set of the target domain D T and retraining the object detector; Step 5, using the object detector trained in Step 4 to perform object detection on the data to be detected, and finally obtaining a detection result. The present invention solves the problems of the influence on the performance of the deep model in object detection and the low prediction accuracy of the instance category and bounding box position after training when there is a source domain with instance-level labels while the target domain does not have instance-level labels.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of transfer learning and relates to a cross-domain target detection method based on domain adaptation. Background Art

[0002] In the field of computer vision, target detection technology has always been an important topic and research direction. The task of target detection is to detect and identify some targets that appear in static images or continuous frames of video image sequences, determine the target location and judge the object category. In recent years, target detection has received extensive attention and research in the academic community, and with the continuous breakthroughs in technology, it has been widely used in the real world, such as video surveillance, human-computer interaction, intelligent transportation, autonomous navigation and robot vision. With the rise of deep neural networks and the powerful computing power of GPUs, target detection continues to flourish.

[0003] At present, deep learning models have been widely used in various fields of computer vision, including target detection. Existing target detection algorithms use deep learning networks as their backbone and detection networks to extract features from input images (or videos) for classification and positioning, respectively. Current target detection algorithms can be roughly divided into two-stage and one-stage methods. The two-stage method first proposes a target candidate bounding box on the input image, and then extracts features from the target candidate box through ROI for subsequent target classification and bounding box regression tasks. It has relatively high target recognition and positioning accuracy, but the algorithm has a slow reasoning speed. On the contrary, the one-stage method directly extracts the prediction box from the input image, which has a higher reasoning speed, but the target recognition accuracy is lower than that of the two-stage method.

[0004] Although most existing target detection algorithms can achieve relatively high recognition accuracy on natural images, the premise of these algorithms is that they all need large-scale labeled data to train the network. In real life, large-scale labeled data may not be available in other image fields, because it is usually difficult and impractical to build large datasets with instance-level labels in many image domains, which include many difficulties such as lack of image sources, copyright issues and labeling costs. Therefore, the existing target detection algorithms have certain limitations. Therefore, we consider using the domain adaptation method in transfer learning to apply the model trained on the source domain data with instance-level labels to the target domain with only sample-level labels, and can obtain a higher target recognition accuracy. The instance-level label consists of a label (i.e., the object class of the instance) and a bounding box (i.e., the location of the instance). The sample-level label only knows the instance category in the image but not the instance location. Summary of the invention

[0005] The purpose of the present invention is to provide a cross-domain target detection method based on domain adaptation, which solves the problem of the impact on the performance of deep models in target detection when there is a source domain with instance-level labels but no instance-level labels in the target domain, and the problem of low accuracy in instance category and bounding box position prediction after training.

[0006] The technical solution adopted by the present invention is:

[0007] The cross-domain target detection method based on domain adaptation includes the following specific steps:

[0008] Step 1: Get the source domain Ds and the target domain D T Target detection dataset, perform data enhancement and dataset expansion;

[0009] Step 2: Build the CycleGAN network, use the expanded data set to train the CycleGAN network and output the generated data domain D G ;

[0010] Step 3: Construct the Faster RCNN network as the target detector, using the source domain Ds and the generated data domain D G Use it as a training set to train the object detector;

[0011] Step 4: For the target domain D T The data set is divided into different levels of data by complexity evaluation, and the target detector is retrained according to the results of complexity evaluation;

[0012] Step 5: Use the target detector trained in step 4 to perform target detection on the data to be detected, and finally obtain the detection result.

[0013] The present invention is also characterized in that:

[0014] In step 2, the CycleGAN network structure consists of two generators with the same structure and two discriminators with the same structure. The generator structure is three convolutional layers, six ResNet modules, two deconvolutional layers and one convolutional layer connected in sequence. Each convolutional layer is followed by a nonlinear activation function. The discriminator structure is five convolutional layers and one fully connected layer connected in sequence. The convolutional layer is followed by a nonlinear activation function, and the fully connected layer is followed by a Softmax function.

[0015] Step 2: The training process of the CycleGAN network is:

[0016] Step 2.1: Extract a subset X from the source domain Ds and extract the subset X from the target domain D T A subset Y is also extracted from the network. Taking X as an example, X is input into the first discriminator D of the CycleGAN network.X ;

[0017] Step 2.2: Input X to the discriminator D from step 2.1 X After that, give the generator G X Input random Gaussian white noise, after the generator generates an image, it is input to the discriminator D X , discriminator D X Judge the input image. If the input image is a generated image, the discriminator D X The output is 0 if the input image is a real image discriminator D X The output is 1;

[0018] Step 2.3: Perform the same operation on Y and input Y into the discriminator D. Y , give the generator G Y Input random Gaussian white noise, the generator generates an image and then inputs it into the discriminator D Y , discriminator D Y Judge the input image. If the input image is a generated image, the discriminator D Y The output is 0 if the input image is a real image discriminator D Y The output is 1.

[0019] The CycleGAN network includes a loss function L for implementing the mapping F:Y→X G (G,D Y ), represents the loss function L when realizing the mapping G: X→Y G (F,D X ), cycle consistency loss function L C (G, F) as shown in formulas (1) to (3):

[0020]

[0021] Where L G (G,D Y ) represents the loss function when implementing the mapping F:Y→X, where Indicates that the real sample y passes through the discriminator D Y The loss function is Indicates that the generated sample G(x) passes through the discriminator D Y The loss function, D Y (y) represents the real sample y through the discriminator D Y Score, D Y (G(x)) represents the generated sample G(x) through the discriminator D Y score;

[0022]

[0023] Where L G (F,D X ) represents the loss function when implementing the mapping G:X→Y, where Indicates that the real sample x passes through the discriminator D X The loss function is Indicates that the generated sample F(y) passes through the discriminator D X The loss function, D X (x) represents the real sample x through the discriminator D X Score, D X (F(y)) represents the generated sample F(y) through the discriminator D X score;

[0024] The cycle consistency loss is:

[0025]

[0026] Where L C (G, F) represents the loss generated when aligning the distribution of generated samples and real samples, F(G(x))-x represents the loss value between the generated sample G(x) and the real sample x, G(F(y))-y represents the loss value between the generated sample F(y) and the real sample y, and ||·||1 is the L1 norm of the vector;

[0027] The final optimization function is:

[0028] L(G,F,D X ,D Y )=L G (G,D Y ,X,Y)+L G (F,D X ,Y,X)+L C (G,F) (4).

[0029] The Faster RCNN network structure in step 3 includes connecting the VGG16 feature extraction network F(·) and the RPN network in sequence. The VGG16 feature extraction network F(·) includes two convolutional layers, a RELU activation function, a maximum pooling layer, two convolutional layers, a maximum pooling layer, three convolutional layers, a RELU activation function, a maximum pooling layer, two convolutional layers, a RELU activation function, and a maximum pooling layer; the input image is passed through the VGG16 feature extraction network F(·) to obtain a feature map and then passes through the RPN network. First, after 512 3×3 convolutions, it is divided into two branches. The first branch uses 18 1×1 convolutions to realize the classification of the foreground or background in the image. The second branch uses 36 1×1 convolutions to realize the bounding box regression of the detected image.

[0030] The training process of the target detector in step 3 is:

[0031] Step 3.1: Samples of the source domain Ds And generate data domain D G Samples similar to the target domain Mix and use the VGG16 network as the feature extractor F(·) to extract a high-dimensional feature vector F(D S ), F(D T );

[0032] Step 3.2 Convert the high-dimensional feature vector F(D S ), F(D T ) is input into the subsequent fully connected network, ReLU nonlinear activation function and fully connected network to obtain a feature map S that stores sufficient feature information. After the feature map S is processed by 3*3 convolution, a high-dimensional feature vector is obtained.

[0033] Step 3.3, after two more 1*1 convolution operations, two feature maps are obtained. According to the output scores of these two feature maps, the candidate region R can be obtained. Then, the feature map S and the candidate region R are pooled with the region of interest P(·) to obtain the feature vector P(S, R) of each region of interest. The feature vector P(S, R) is input into the classifier layer to obtain the category and bounding box of the region of interest, and the training of the object detector is completed iteratively.

[0034] The loss function of the target detector in step 3 is the sum of the classification loss and the regression loss, as shown below:

[0035]

[0036]

[0037]

[0038]

[0039] in is the index of the anchor point in the mini-batch, p i It is an anchor point As the predicted probability of the target, True value, when anchor is positive, is 1, when the anchor is negative, is 0, t i is a vector of four parameterized coordinates of the predicted bounding box, are the coordinates of the ground-truth box associated with the positive anchor box, L C is the classification loss of two categories, L r is the loss of bounding box regression, {p i},{ti} represent the outputs of the classification layer and regression layer respectively.

[0040] Step 4 is as follows:

[0041] Step 4.1: For the target domain D T The validation set is used to evaluate the complexity. First, the pre-trained VGG network is used, and its last layer is removed as a feature extractor to extract the features of the samples. At the same time, the input image is enhanced. Finally, the output high-dimensional feature vector is normalized using the L2 norm. The normalized features are then used to train the ridge regression classifier so that the model can predict the ground-truth difficulty score.

[0042] Step 4.2: According to the evaluation results, the target domain D T The validation set samples are divided into k batches according to the difficulty. The sample difficulty evaluation formula is as follows:

[0043]

[0044] Where I is the input image, B is the bounding box coordinates, and w i 、h i is the width and height in bounding box coordinates, n is the number of samples;

[0045] Step 4.3, after performing complexity evaluation on the validation set samples, the samples can be divided into simple, medium and difficult samples according to the complexity evaluation results. Then, the easy samples are input to the target detector first to obtain the target detector for the target domain D T The prediction results of the samples are then used as pseudo labels for simple samples to train the target detector again. Then, the medium-difficulty samples are input into the target detector and the same operations as the simple samples are performed. Finally, the difficult samples are input into the target detector and the same operations as the simple samples are performed. This completes the iteration of the validation set data and completes the final training of the target detector.

[0046] The beneficial effects of the present invention are:

[0047] The present invention proposes a target detection method based on domain adaptation, which ensures the global domain distribution while not changing the distinguishing information between the data in the source domain and the target domain. After training with domain adversarial loss, cycle consistency loss and target detector regression loss, the target domain validation set data is sorted from easy to difficult and input into the target detector for prediction in this order, so as to give pseudo labels to the target domain validation set samples, and then the target detector is trained again using the validation set samples with the existing pseudo labels, and the target detector is finally trained through cyclic iteration, so that it can show better classification and prediction performance in the target domain test set. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 It is a network structure diagram of the target detection method based on domain adaptation of the present invention;

[0049] Figure 2 Schematic diagram of the network structure of the CycleGAN network in step 2 of the present invention;

[0050] Figure 3 It is a schematic diagram of the network structure of Faster RCNN in step 3 of the present invention. DETAILED DESCRIPTION

[0051] The present invention is described in detail below with reference to the accompanying drawings and specific embodiments.

[0052] The present invention is based on a domain-adaptive cross-domain target detection method, and the specific steps include:

[0053] Step 1: Get the source domain Ds and the target domain D T Target detection dataset, perform data enhancement and dataset expansion;

[0054] Step 2: Build the CycleGAN network, use the expanded data set to train the CycleGAN network and output the generated data domain D G ;

[0055] Step 3: Construct the Faster RCNN network as the target detector, using the source domain Ds and the generated data domain D G Use it as a training set to train the object detector;

[0056] Step 4: For the target domain D T The data set is divided into different levels of data by complexity evaluation, and the target detector is retrained according to the results of complexity evaluation;

[0057] Step 5: Use the target detector trained in step 4 to perform target detection on the data to be detected, and finally obtain the detection result.

[0058] Step 1 is as follows: the source domain dataset Ds follows a certain distribution P s (x), the class label is L s ,Right now Target domain dataset D T Subordinate distribution P T (x), the class label is L T ,Right now The source domain and target domain datasets are input into the random data augmentation network in batches. The random data augmentation network rotates, crops, and adds Gaussian white noise to the original source domain and target domain dataset samples, and then restores them to the original input size to form new samples and add them to the original dataset, thereby achieving the purpose of dataset expansion.

[0059] The CycleGAN network structure in step 2 consists of two generators with the same structure and two discriminators with the same structure. The generator structure is three convolutional layers, six ResNet modules, two deconvolutional layers and one convolutional layer connected in sequence, and each convolutional layer is followed by a nonlinear activation function. The discriminator structure is five convolutional layers and one fully connected layer connected in sequence, and the convolutional layer is followed by a nonlinear activation function, and the fully connected layer is followed by a Softmax function.

[0060] The training process of the CycleGAN network in step 2 is:

[0061] Step 2.1: Extract a subset X from the source domain Ds and extract the subset X from the target domain D T A subset Y is also extracted from the network. Taking X as an example, X is input into the first discriminator D of the CycleGAN network. X ;

[0062] Step 2.2: Input X to the discriminator D from step 2.1 X After that, give the generator G X Input random Gaussian white noise, after the generator generates an image, it is input to the discriminator D X , discriminator D X Judge the input image. If the input image is a generated image, the discriminator D X The output is 0 if the input image is a real image discriminator D X The output is 1;

[0063] Step 2.3: Perform the same operation on Y and input Y into the discriminator D. Y , give the generator G Y Input random Gaussian white noise, the generator generates an image and then inputs it into the discriminator D Y , discriminator D Y Judge the input image. If the input image is a generated image, the discriminator DY The output is 0 if the input image is a real image discriminator D Y The output is 1;

[0064] The main goal of the entire network training process in step 2 is to achieve two mapping functions:

[0065] G: X→Y, F: Y→X

[0066] The generated image is made similar to the target domain image, and its adversarial loss function is:

[0067]

[0068] Where L G (G,D Y ) represents the loss function when implementing the mapping F:Y→X, where Indicates that the real sample y passes through the discriminator D Y The loss function is Indicates that the generated sample G(x) passes through the discriminator D Y The loss function, D Y (y) represents the real sample y through the discriminator D Y Score, D Y (G(x)) represents the generated sample G(x) through the discriminator D Y score.

[0069]

[0070] Where L G (F,D X ) represents the loss function when implementing the mapping G:X→Y, where Indicates that the real sample x passes through the discriminator D X The loss function is Indicates that the generated sample F(y) passes through the discriminator D X The loss function, D X (x) represents the real sample x through the discriminator D X Score, D X (F(y)) represents the generated sample F(y) through the discriminator D X score.

[0071] In addition, further optimization is performed, and its cycle consistency loss is

[0072]

[0073] Where L C(G, F) represents the loss generated when aligning the distribution of generated samples and real samples, F(G(x))-x represents the loss value between the generated sample G(x) and the real sample x, G(F(y))-y represents the loss value between the generated sample F(y) and the real sample y, and ||·||1 is the L1 norm of the vector;

[0074] Therefore, the final optimization function is:

[0075] L(G,F,D X ,D Y )=L G (G,D Y ,X,Y)+L G (F,D X ,Y,X)+L C (G,F) (4)

[0076] The Faster RCNN network structure in step 3 includes the VGG16 feature extraction network F(·) and the RPN network.

[0077] The VGG16 network structure consists of 13 convolutional layers and 3 fully connected layers. Because it is used as a feature extraction network, the fully connected layers are removed. The input image is first convolved twice with 64 3×3 convolution kernels, connected to the ReLU activation function, and then convolved twice with 128 3×3 convolution kernels. Then, it is connected to the ReLU activation function, and then convolved once with a 2×2 maximum pooling kernel. After convolving three times with 256 3×3 convolution kernels, it is connected to the ReLU activation function, and then convolved once with a 2×2 maximum pooling kernel. After repeating the convolution twice with 512 3×3 convolution kernels, it is connected to the ReLU activation function, and then convolved once with a 2×2 maximum pooling kernel to obtain the feature map.

[0078] The input of the RPN network is the feature map obtained after VGG16. It first passes through 512 3×3 convolutions and is divided into two branches. The first branch uses 18 1×1 convolutions to classify the foreground or background in the image, and the second branch uses 36 1×1 convolutions to achieve bounding box regression of the detected image.

[0079] The training process of the target detector in step 3 is:

[0080] Step 3.1: Samples of the source domain Ds And generate data domain D G Samples similar to the target domain Mix and use the VGG16 network as the feature extractor F(·) to extract a high-dimensional feature vector F(D S ), F(D T );

[0081] Step 3.2 Convert the high-dimensional feature vector F(D S ), F(D T ) is input into the subsequent fully connected network, ReLU nonlinear activation function and fully connected network to obtain a feature map S that stores sufficient feature information. After the feature map S is processed by 3*3 convolution, a high-dimensional feature vector is obtained.

[0082] Step 3.3, after two more 1*1 convolution operations, two feature maps are obtained. According to the output scores of these two feature maps, the candidate region R can be obtained. Then, the feature map S and the candidate region R are pooled with the region of interest P(·) to obtain the feature vector P(S, R) of each region of interest. The feature vector P(S, R) is input into the classifier layer to obtain the category and bounding box of the region of interest, and the training of the object detector is completed iteratively.

[0083] The loss function of the target detector in step 3 is the sum of the classification loss and the regression loss, as shown below:

[0084]

[0085]

[0086]

[0087]

[0088] in is the index of the anchor point in the mini-batch, p i It is an anchor point As the predicted probability of the target, True value, when anchor is positive, is 1, when the anchor is negative, is 0, t i is a vector of four parameterized coordinates of the predicted bounding box, are the coordinates of the ground-truth box associated with the positive anchor box, L C is the classification loss of two categories, L r is the loss of bounding box regression, {p i},{t i} represent the outputs of the classification layer and regression layer respectively.

[0089] Step 4 specifically includes:

[0090] Step 4.1: For the target domain D TThe validation set is used to evaluate the complexity. First, the pre-trained VGG network is used, and its last layer is removed as a feature extractor to extract the features of the samples. At the same time, the input image is enhanced. Finally, the output high-dimensional feature vector is normalized using the L2 norm. The normalized features are then used to train the ridge regression classifier so that the model can predict the ground-truth difficulty score.

[0091] Step 4.2: According to the evaluation results, the target domain D T The validation set samples are divided into k batches according to the difficulty. The sample difficulty evaluation formula is as follows:

[0092]

[0093] Where I is the input image, B is the bounding box coordinates, and w i 、h i are the width and height in bounding box coordinates, and n is the number of samples.

[0094] Step 4.3, after performing complexity evaluation on the validation set samples, the samples can be divided into simple, medium and difficult samples according to the complexity evaluation results. Then, the easy samples are input to the target detector first to obtain the target detector for the target domain D T The prediction results of the samples are then used as pseudo labels for simple samples to train the target detector again. Then, the medium-difficulty samples are input into the target detector and the same operations as the simple samples are performed. Finally, the difficult samples are input into the target detector and the same operations as the simple samples are performed. This completes the iteration of the validation set data and completes the final training of the target detector.

Claims

1. A cross-domain target detection method based on domain adaptation, characterized in that: The specific steps include: Step 1: Get the source domain and target domain Target detection dataset, perform data enhancement and dataset expansion; Step 2: Build the CycleGAN network, train the CycleGAN network with the expanded data set, and output the generated data domain ; Step 3: Build the Faster RCNN network as the target detector to transform the source domain and generate data domain Use it as a training set to train the object detector; Step 4: Target domain The data set is divided into different levels of data by complexity evaluation, and the target detector is retrained according to the results of complexity evaluation; Step 5, using the target detector trained in step 4 to perform target detection on the data to be detected, and finally obtaining the detection result; The step 4 is specifically as follows: Step 4.1: Target domain The validation set is used to evaluate the complexity. First, the pre-trained VGG network is used, and its last layer is removed as a feature extractor to extract the features of the samples. At the same time, the input image is enhanced. Finally, the output high-dimensional feature vector is normalized using the L2 norm. The normalized features are then used to train the ridge regression classifier so that the model can predict the ground-truth difficulty score. Step 4.2: According to the evaluation results, the target domain The validation set samples are divided according to the difficulty. The sample difficulty evaluation formula is as follows: in is the input image, are the bounding box coordinates, are the width and height in bounding box coordinates, is the sample size; Step 4.3, after evaluating the complexity of the validation set samples, the samples are divided into simple, medium and difficult samples according to the complexity evaluation results. Then the easy samples are input to the target detector first to obtain the target detector for the target domain. The prediction results of the samples are then used as pseudo labels for simple samples to train the target detector again. Then, the medium-difficulty samples are input into the target detector and the same operations as the simple samples are performed. Finally, the difficult samples are input into the target detector and the same operations as the simple samples are performed. This completes the iteration of the validation set data and completes the final training of the target detector.

2. The cross-domain target detection method based on domain adaptation as claimed in claim 1, characterized in that: The CycleGAN network structure described in step 2 consists of two generators with the same structure and two discriminators with the same structure. The generator structure is three convolutional layers, six ResNet modules, two deconvolutional layers and one convolutional layer connected in sequence. Each convolutional layer is followed by a nonlinear activation function. The discriminator structure is five convolutional layers and one fully connected layer connected in sequence. The convolutional layer is followed by a nonlinear activation function, and the fully connected layer is followed by a Softmax function.

3. The cross-domain target detection method based on domain adaptation as claimed in claim 2, characterized in that: Step 2: The training process of the CycleGAN network is as follows: Step 2.1: From the source domain Extract a subset from , and from the target domain We also extract a subset ,by For example, Input to the first discriminator of the CycleGAN network ; Step 2.2: From step 2.1, Input to the discriminator After that, give the generator Input random Gaussian white noise, after the generator generates an image, it is input to the discriminator , discriminator Judge the input image. If the input image is a generated image, the discriminator The output is 0 if the input image is a real image discriminator The output is 1; Step 2.3: Perform the same operation on Y and input Y into the discriminator. , to the generator Input random Gaussian white noise, the generator generates an image and then inputs it into the discriminator , discriminator Judge the input image. If the input image is a generated image, the discriminator The output is 0 if the input image is a real image discriminator The output is 1.

4. The cross-domain target detection method based on domain adaptation as claimed in claim 2, characterized in that: The CycleGAN network includes a mapping The loss function when , represents the implementation mapping The loss function when , cycle consistency loss function As shown in formulas (1) to (3): in Represents the implementation mapping The loss function when Represents the real sample Through the discriminator The loss function is Indicates the generated sample Through the discriminator The loss function is Represents the real sample Through the discriminator The score, Indicates the generated sample Through the discriminator score; in Represents the implementation mapping The loss function when Represents the real sample Through the discriminator The loss function is Indicates the generated sample Through the discriminator The loss function is Represents the real sample Through the discriminator The score, Indicates the generated sample Through the discriminator score; The cycle consistency loss is: in represents the loss in aligning the distribution of generated samples and real samples, Indicates the generated sample and real samples The loss value between Indicates the generated sample and real samples The loss value between is the L1 norm of the vector; The final optimization function is: 。 5. The cross-domain target detection method based on domain adaptation as claimed in claim 1, characterized in that: The Faster RCNN network structure in step 3 includes connecting the VGG16 feature extraction network in sequence And RPN network, the VGG16 feature extraction network Includes two convolutional layers, one RELU activation function, one maximum pooling layer, two convolutional layers, one maximum pooling layer, three convolutional layers, one RELU activation function, one maximum pooling layer, two convolutional layers, one RELU activation function, one maximum pooling layer; the input image is extracted by the VGG16 feature extraction network The feature map is then passed through the RPN network, first through 512 After the convolution, it is divided into two branches. The first branch uses 18 After the convolution, the foreground or background classification of the image is realized. The second branch uses 36 After the convolution, the bounding box regression of the detected image is realized.

6. The cross-domain target detection method based on domain adaptation as claimed in claim 5, characterized in that: The training process of the target detector in step 3 is: Step 3.1: Set the source domain Sample And generate data domain Samples similar to the target domain Hybrid and use the VGG16 network as a feature extractor Extract high-dimensional feature vectors , ; Step 3.2 Convert the high-dimensional feature vector , Input to the subsequent fully connected network, ReLU nonlinear activation function and fully connected network to obtain a feature map that preserves sufficient feature information , the feature map After 3*3 convolution processing, a high-dimensional feature vector is obtained; Step 3.3: After two more 1*1 convolution operations, two feature maps are obtained. According to the output scores of these two feature maps, the candidate regions are obtained. , and then the feature map and candidate regions Perform region of interest pooling Get the feature vector of each region of interest , the feature vector The input classifier layer obtains the category and bounding box of the region of interest, and the training of the object detector is completed iteratively.

7. The cross-domain target detection method based on domain adaptation as claimed in claim 5, characterized in that: The loss function of the target detector in step 3 is the sum of the classification loss and the regression loss, as shown below: in is the index of the anchor point in the mini-batch, It is an anchor point As the predicted probability of the target, True value, when anchor is positive, is 1, when the anchor is negative, is 0, is a vector of four parameterized coordinates of the predicted bounding box, are the coordinates of the ground-truth box associated with the positive anchor box, is the classification loss for two categories, is the loss for bounding box regression, Represent the outputs of the classification layer and regression layer respectively.

Citation Information

Patent Citations

  • Method for constructing target detection adaptive model based on CycleGAN and pseudo tag

    CN111882055A

  • Domain adaptive unsupervised target detection method based on feature separation and alignment

    CN112488229A