Visual Scene Recognition Method Based on Structured Information Feature Decoupling and Knowledge Transfer
By adopting feature decoupling and knowledge transfer methods based on structured information in visual scene recognition, the problem of feature redundancy interleaving under appearance changes is solved, and more accurate visual scene recognition and visual positioning are achieved.
Patent Information
- Application Number
- CN202111000756.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-08-27
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2041-08-27
AI Technical Summary
The prior art has redundant and interlaced features under appearance changes, insufficient image representation ability, and it is difficult to achieve accurate visual scene recognition.
Using a method of feature decoupling and knowledge transfer based on structured information, structured features are extracted through Canny edge detectors and automatic encoders, combined with appearance teacher models and probabilistic knowledge transfer technology, deeply decoupled feature representations are generated to improve the representation ability of image features.
It improves the image feature representation ability under appearance changes, enhances the robot's visual positioning accuracy in large-scale scenarios, improves the scene re-identification ability, and serves application scenarios such as navigation and positioning.
Smart Images

Figure CN114049541B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and robotics, and particularly relates to a visual scene recognition method based on structured information feature decoupling and knowledge transfer. Background Art
[0002] Accurate scene recognition helps a robot recognize its own state and complete work tasks well. A scene refers to the data at a certain location and at a certain moment recorded by a sensor in the real world, which contains combinations of various different objects. The task of a mobile robot is to repeatedly visit the same scene at different time periods and determine whether the scene has been experienced before. Scene recognition generally focuses on "where am I", and analyzes and judges the current scene by detecting and analyzing the targets in the scene or extracting stable features. For example, in the process of visual SLAM (Simultaneous Localization and Mapping), accurate scene recognition can help the robot determine whether it has entered an environment area visited before, so as to form a closed-loop detection and optimize the map, which is crucial for ensuring the consistency of the map and reducing cumulative errors. "IEEE international conference on robotics and automation (ICRA), 1011–1018, 2018" discloses a convertible generator that can transform conditions such as day and night, seasons of an image. The image transformation generator is designed based on the SURF detector and dense descriptors, and is used to assist feature matching, so as to improve the accuracy of visual scene recognition and metric localization under drastic appearance changes. "IEEE International conference on robotics and automation (ICRA), 4489–4495, 2018" proposes an adversarial, lifelong, incremental domain adaptation method. This method approximates the feature distribution of the source domain by using a generative adversarial network, so that the deployment module can be completely independent of a large amount of source training data. "IEEE International Conference on Robotics and Automation (ICRA), 9271–9277, 2020" proposes a multi-spectral domain invariant framework. By introducing new constraints into the objective function, this framework generates invariant images with semantics and strong discrimination using a non-paired image transformation method, showing competitive performance in multi-spectral scene recognition tasks. Therefore, the key problems of visual scene recognition methods lie in network training under appearance change situations, feature decoupling based on adversarial training, and knowledge transfer based on structured information. Summary of the Invention
[0003] In view of the feature redundancy and interleaving and the insufficient image representation ability of previous scene recognition methods under the condition of appearance changes, the present invention proposes a visual scene recognition method based on feature decoupling and knowledge transfer of structured information. This method uses structural information to learn a deeply decoupled feature representation for scene recognition. By introducing a method of probabilistic knowledge transfer, the transfer of structural information from the Canny edge detector to the structure encoder is realized, and an appearance teacher model is added to help the appearance encoder generate more specific features. In addition, an affine transformation is introduced to generate additional noise into the convolutional autoencoder to solve the problem that the edge is too sensitive to the perspective change. This method can improve the representation ability of image features in the case of appearance changes, so as to ensure that the generated image features can cope with complex environmental changes, improve the scene re-recognition ability of the robot, and serve application scenarios such as navigation and positioning.
[0004] The technical solution of the present invention is realized as follows:
[0005] A visual scene recognition method based on feature decoupling and knowledge transfer of structured information includes the following steps:
[0006] Step 1, use the Canny edge detector to extract the edge representation form X of the image X CE and convert it into a vector X based on the autoencoder CT ;
[0007] Step 2, use the fine-tuned ResNet-34 to extract the appearance feature representation X of the image X AT ;
[0008] Step 3, for the input image X, send it into the feature decoupling network, and then the structured feature vector X SC and the appearance feature vector X A will be generated respectively. Subsequently, X SC is sent to D AA to judge whether the extracted structured feature vector comes from the same domain. In addition, the feature distribution of X SC will be compared with the X CT generated by the content teacher module. As for X A , it will not only be optimized by the triplet loss function, but also its distribution will be compared with the X AT generated by the appearance teacher module.
[0009] Step 4, the decoder D E integrates the input features and reconstructs the original image to encourage the learned content features and appearance features to form a complete representation of the input image. Extract the structured feature vector X SCAs the final scene feature, the cosine distance is used to calculate and optimize the similarity between features, realizing visual scene recognition.
[0010] Further, Step 1: First, in order to achieve a two-dimensional projective transformation, four points in the image need to be found to estimate the homography matrix. Four points are randomly selected within the border at the corners of each frame of the image. The size of the border is set to to ensure a reasonable degree of perspective change. H and W are the width and height of the image respectively.
[0011] The edge representation of the image is
[0012] X CE = Canny(X) (1)
[0013] Canny(·) is the edge extraction operation of the Canny edge detector.
[0014] The vector representation of the edge is:
[0015] X CT = Auto_encoder(X CE ) (2)
[0016] Auto_encoder(·) is the feature encoding operation of the autoencoder.
[0017] Further, Step 2: For the input image X, the fine-tuned ResNet-34 is used to extract the appearance feature representation X AT :
[0018] X AT = ResNet(X) (3)
[0019] ResNet(·) is the operation of extracting the second-to-last layer features of ResNet-34.
[0020] Further, Step 3:
[0021] For the appearance feature, it is extracted through the encoder E A and is represented as:
[0022] X A = E A (X) (4)
[0023] The appearance encoder is trained through the following loss function:
[0024]
[0025] where α controls the separated edges, and y ij ∈{-1,1}. θ AThey are the parameters of the appearance encoder.
[0026] The structured content features are extracted by the encoder E SC and are expressed as:
[0027] X SC = E SC (X) (6)
[0028] To obtain appearance-irrelevant features, a discriminative appearance classification loss function is designed. During the training phase, the content features are fed into the appearance discriminator D AA . The purpose of E SC is to deceive D AA so that it cannot correctly classify the content features.
[0029] It is necessary to train the appearance discriminator D based on the generated E SC and the cross-entropy loss function: AA :
[0030]
[0031] where D AA is considered a binary classifier. θ DAA are the parameters of the appearance discriminator, and can also be expressed as:
[0032]
[0033] where x is the concatenated feature of the input content feature pair {x i , x j}. Note that the gradient of will only backpropagate to the classifier and will not update the other layers of E SC . To implement adversarial training, it is necessary to deceive the appearance discriminator:
[0034]
[0035] where the gradient of will backpropagate to E SC , while the weight parameters of the appearance discriminator should remain unchanged at this time.
[0036] Referring to the approach of probability knowledge transfer, it is first necessary to probabilistically model the data sample sets in the two feature spaces. In this case, how to transfer knowledge (marginal information) from X CT to X SCThe problem is then transformed into minimizing the divergence of the joint probability density distribution between distributions P and Q. Considering that the conditional probability distribution represents the probability of each sample selecting its neighborhood, this can more accurately model the geometric structure of the feature space. Therefore, the conditional probability distribution is used to describe the content teacher model:
[0037]
[0038] Similarly, the probability distribution of the student model X SC is expressed as:
[0039]
[0040] where is a symmetric kernel function with width σ t . a and b are input vectors. The sum of the conditional probabilities is 1 and the range is [0, 1].
[0041] In this teacher-student model, a cosine similarity-based metric is adopted:
[0042]
[0043] The Wasserstein distance is used as the divergence metric:
[0044]
[0045] where P 1 and P 2 represent the probability distributions of the teacher model and the student model respectively. Π(P 1 , P 2 ) is all possible joint probability distributions between P 1 and P 2 . As a distance function, the Wasserstein distance has a nice property that the distance between the centroids of the two distributions is the lower bound. Adopting such a lower bound greatly reduces the computational amount. The final loss function for training the student model (structural content encoder) is defined as:
[0046]
[0047] where N is the size of the mini-batch.
[0048] Similar to the content teacher model, the Wasserstein distance is also used to measure the similarity of the probability distributions of X AT and X A . Therefore, the loss function of the appearance teacher-student model is defined as follows:
[0049]
[0050] Further, Step Four:
[0051] Adopt an encoder-decoder architecture, and the reconstruction loss is defined as:
[0052]
[0053] where and θ SC , θ A , θ DE are the parameters of the encoder and decoder respectively.
[0054] Use the trained network to extract the structured feature vector X SC as the final scene feature, and use the cosine distance to calculate the similarity between the optimized features to achieve visual scene recognition. For visual scene recognition using the generated features, the cosine distance is used for calculating the similarity between images:
[0055]
[0056] Advantages of the present invention: The method of the present invention fully considers visual scene recognition under appearance changes, designs and trains the network structure for feature decoupling and structured information integration, and finally calculates the similarity between images using the optimized structured content features to complete accurate visual scene recognition. It greatly improves the visual positioning accuracy of the robot in large-scale scenes and helps to carry out more intelligent visual navigation and other tasks. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 Schematic diagram of the present invention using projective transformation to simulate perspective changes;
[0058] Figure 2 Schematic diagram of the network structure of the autoencoder in the teacher model of the present invention;
[0059] Figure 3 Experimental results of using different sensitivity thresholds in the Canny edge extractor of the present invention;
[0060] Figure 4 Comparison of appearance prediction performance of using different modules and their combinations in the present invention;
[0061] Figure 5 Schematic diagram of the execution flow of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0062] The following further describes the present invention with reference to the accompanying drawings.
[0063] The visual scene recognition method based on structured information feature decoupling and knowledge transfer of the present invention includes the following steps:
[0064] Step 1: Input the images in the Nordland dataset into the network batch by batch. To achieve a two-dimensional projective transformation, four points in the image need to be found to estimate the homography matrix. As shown in Figure 1 the schematic diagram of simulating perspective change using projective transformation. Four points are randomly selected within the border at the corners of each frame of the image. The size of the border is set to to ensure a reasonable degree of perspective change. H and W are the width and height of the image, respectively. Generally, H = W = 224 is taken.
[0065] The edge representation of the image is
[0066] X CE = Canny(X) (1)
[0067] Canny(·) is the operation of extracting edges by the Canny edge detector.
[0068] The vector representation of the edge is:
[0069] X CT = Auto_encoder(X CE ) (2)
[0070] Auto_encoder(·) is the feature encoding operation of the autoencoder. The length of the generated edge features is set to 2048. The structure of the autoencoder is as shown in Figure 2 shown.
[0071] Step 2: For the input image X, use the fine-tuned ResNet-34 to extract the appearance feature representation X AT :
[0072] X AT = ResNet(X) (3)
[0073] ResNet(·) is the operation of extracting the features of the penultimate layer of ResNet-34. The ResNet-34 network is fine-tuned with a learning rate of 1×10 -4 .
[0074] Step 3: For the appearance features, extract them through the encoder E A and represent them as:
[0075] X A = E A (X) (4)
[0076] The appearance encoder is trained through the following loss function:
[0077]
[0078] where α controls the separated edge, and y ij ∈{-1, 1}. θ A is the parameter of the appearance encoder. Set λ = 0.5 and limit the distance to 1.4. The learning rate of the boundary threshold β is set to 0.0002 and the initial value is 1.0.
[0079] The structured content features are extracted by the encoder E SC and are expressed as:
[0080] X SC = E SC (X) (6)
[0081] To obtain appearance-irrelevant features, a discriminative appearance classification loss function is designed. During the training phase, the content features are fed into the appearance discriminator D AA . The purpose of E SC is to deceive D AA so that it cannot correctly classify the content features.
[0082] It is necessary to train the appearance discriminator D SC based on the generated E AA and the cross-entropy loss function:
[0083]
[0084] where D AA is considered a binary classifier. θ DAA is the parameter of the appearance discriminator, and can also be expressed as:
[0085]
[0086] where x is the concatenated feature of the input content feature pair {x i , x j}}. Note that the gradient of SC will only backpropagate to the classifier and will not update the other layers of E
[0087]
[0088] where the gradient of SC will backpropagate to E
[0089] Referring to the practice of probability knowledge transfer, it is first necessary to probabilistically model the data sample sets in the two feature spaces. In this way, how to transfer knowledge (edge information) from XCT Migration to X SC The problem is then transformed into minimizing the divergence of the joint probability density distribution between distributions P and Q. Considering that the conditional probability distribution represents the probability of each sample selecting its neighborhood, this can more accurately model the geometric structure of the feature space. Therefore, the conditional probability distribution is used to describe the content teacher model:
[0090]
[0091] Similarly, the probability distribution of the student model X SC is expressed as:
[0092]
[0093] where is a symmetric kernel function with width σ t . a and b are input vectors. The sum of the conditional probabilities is 1 and the range is [0, 1].
[0094] In this teacher-student model, a cosine similarity-based metric is adopted:
[0095]
[0096] The Wasserstein distance is used as the divergence metric:
[0097]
[0098] where P 1 and P 2 represent the probability distributions of the teacher model and the student model, respectively. Π(P 1 , P 2 ) is all possible joint probability distributions between P 1 and P 2 . As a distance function, the Wasserstein distance has a nice property that it is bounded below by the distance between the centroids of the two distributions. Adopting such a lower bound greatly reduces the computational amount. The final loss function for training the student model (structural content encoder) is defined as:
[0099]
[0100] where N is the size of the mini-batch.
[0101] Similar to the content teacher model, the Wasserstein distance is also used to measure the similarity of the probability distributions of X AT and X A . Therefore, the loss function of the appearance teacher-student model is defined as follows:
[0102]
[0103] Step 4: Adopt an encoder-decoder architecture, and the reconstruction loss is defined as:
[0104]
[0105] where and θ SC , θ A , θ DE are the parameters of the encoder and decoder, respectively.
[0106] Use the trained network to extract the structured feature vector X SC as the final scene feature, and use the cosine distance to calculate the similarity between the optimized features to achieve visual scene recognition. The feature length is set to 512, the size N of the mini-batch is set to 4, the dropout rate in the encoder is set to 0.5, and that in the discriminator is set to 0.25.
[0107] Use the generated features for visual scene recognition, and the cosine distance is used for calculating the similarity between images:
[0108]
[0109] Use the edge detection algorithm as the teacher model to guide the learning of the content encoder. Therefore, as a feature extractor, the parameters of the edge detection algorithm are also extremely important. Different sensitivity thresholds will result in different noises and accuracies in the generated edge information. We adjusted the threshold t from 0.02 to 0.12 and tested the experimental effects of adding the content teacher model. The PR curve plotted is as Figure 3 shown. We found that it is not the case that the smaller the threshold, the richer the image information. On the contrary, a smaller threshold (t = 0.02) will bring more noise and thus reduce the overall performance. When the threshold is larger, such as 0.12 and 0.10, less edge information is obtained, which will also reduce the performance. Only when the threshold is within an appropriate range, such as t = 0.06, can the best results be obtained.
[0110] The decoupled appearance features can be used to predict the appearance characteristics of each image. Evaluate the prediction accuracy of four different appearances on the Nordland dataset. As Figure 4As shown, before adopting ATM, the appearance features extracted by the original FDNet could only achieve an average accuracy of 70.04%. Thanks to the parameters pre-trained by ResNet-34 and its deeper network, the individual ATM could achieve an accuracy of 91.29% after fine-tuning. The accuracy of FDNet_M is even higher than that of FDNet, which shows the effectiveness of distance-weighted sampling and the edge-based loss function. The introduction of CTM can slightly improve the accuracy of appearance features, while the introduction of ATM significantly improves the classification accuracy of appearance features, which means that this structure can effectively transfer knowledge from ATM to the appearance encoder.
Claims
1. A visual scene recognition method based on structured information feature decoupling and knowledge transfer, characterized in that, the specific steps are as follows: Step 1, use the Canny edge detector to extract the edge representation X of image X CE , and convert it into a vector X based on the autoencoder CT ; Step 2: Use the fine-tuned ResNet-34 to extract the appearance feature representation X of image X AT ; Step 3: For the input image X, send it into the feature decoupling network, and then the structured feature vector X SC and the appearance feature vector X A will be generated respectively; subsequently, X SC is sent to D AA to determine whether the extracted structured feature vectors come from the same domain. In addition, the feature distribution of X SC will be compared with that of X CT generated by the content teacher module. As for X A , it will not only be optimized by the triplet loss function, but its distribution will also be compared with that of X AT generated by the appearance teacher module; Step 4, decoder D E Integrate the input features and reconstruct the original image to encourage the learned content features and appearance features to form a complete representation of the input image; extract the structured feature vector X SC As the final scene feature, and use the cosine distance to calculate and optimize the similarity between features to achieve visual scene recognition.
2. The visual scene recognition method based on structured information feature decoupling and knowledge transfer according to claim 1, characterized in that, the specific process of step one is as follows: First, in order to implement a two-dimensional projective transformation, four points in the image need to be found to estimate the homography matrix. Four points are randomly selected within the border at the corners of each frame of the image, and the size of the border is set to to ensure a reasonable degree of view change. H and W are the width and height of the image respectively; The edge representation form of the image is X CE = Canny(X) (1) Canny(·) is the edge extraction operation of the Canny edge detector; The vector representation of the edge is: X CT = Auto_encoder(X CE ) (2) Auto_encoder(·) is the feature encoding operation of the auto-encoder.
3. The visual scene recognition method based on structured information feature decoupling and knowledge transfer according to claim 1, characterized in that, the specific process of step two is: For the input image X, the fine-tuned ResNet-34 is used to extract the appearance feature representation X AT : X AT = ResNet(X) (3) ResNet(·) is the operation of extracting the penultimate layer features of ResNet-34.
4. The visual scene recognition method based on structured information feature decoupling and knowledge transfer according to claim 1, characterized in that, the specific process of step three is: Appearance features are extracted by encoder E A and are represented as: X A = E A (X) (4) The appearance encoder is trained through the following loss function: where α controls the separated edges, and y ij ∈{-1, 1}, θ A are the parameters of the appearance encoder; The structured content features are extracted by the encoder E SC and are represented as: X SC = E SC (X) (6) To obtain appearance-irrelevant features, a discriminative appearance classification loss function is designed. During the training phase, the content features are fed into the appearance discriminator D AA where E SC aims to deceive D AA so that it cannot correctly classify the content features; It is necessary to generate E-based SC Train the appearance discriminator D using the cross-entropy loss function AA : Among them, D AA is considered a binary classifier, and θ DAA are the parameters of the appearance discriminator, and can also be expressed as: where x is the concatenation feature of the input content feature pair {x i , x j}; to achieve adversarial training, it is necessary to deceive the appearance discriminator: Among them, the gradient will be backpropagated to E SC , while the weight parameters of the appearance discriminator should remain unchanged at this time. First, it is necessary to probabilistically model the data sample sets in the two feature spaces; the problem of transferring knowledge (marginal information) from X CT to X SC is transformed into minimizing the divergence of the joint probability density distribution between distributions P and Q; considering that the conditional probability distribution represents the probability of each sample choosing its neighborhood, which can more accurately model the geometric structure of the feature space, therefore, the conditional probability distribution is used to describe the content teacher model: Similarly, the probability distribution of student model X SC is expressed as: Among them, is a symmetric kernel function with a width of σ t, a and b are input vectors, the sum of the conditional probabilities is 1, and the range is [0, 1]. A cosine similarity-based metric is adopted in this teacher-student model: The Wasserstein distance is used as the divergence metric: where, P 1 and P 2 represent the probability distributions of the teacher model and the student model respectively, and Π(P 1 , P 2 ) is all possible joint probability distributions between P 1 and P 2 ; as a distance function, the Wasserstein distance has a very good property that the distance between the centroids of the two distributions is used as a lower bound. Using such a lower bound greatly reduces the computational amount. The final loss function used to train the student model (structural content encoder) is defined as: where N is the size of the mini-batch; Similar to the content teacher model, the Wasserstein distance is also used to measure the similarity between X AT and X A probability distributions. Therefore, the loss function of the appearance teacher-student model is defined as follows:
5. The visual scene recognition method based on structured information feature decoupling and knowledge transfer according to claim 1, characterized in that, the specific process of step four is as follows: An encoder-decoder architecture is adopted, and the reconstruction loss is defined as: Among them, and θ SC , θ A , θ DE are the parameters of the encoder and the decoder respectively; Extract the structured feature vector X using the trained network SC As the final scene feature, calculate the similarity between the optimized features using the cosine distance to achieve visual scene recognition. Use the generated features for visual scene recognition. The cosine distance is used for calculating the similarity between images:
Citation Information
Patent Citations
Multi-modal human body action recognition method based on knowledge distillation and adversarial learning
CN112364708A
Visual scene recognition method based on learnable feature graph filtering and graph attention network
CN113033669A
Cited By
SYSTEM AND METHOD FOR SUPPORTING SMM UPDATES AND RUNTIME TELEMETRY FOR BRIGHT METAL INSERTS
DE102022116740A1