Visual navigation image processing algorithm based on Riemannian manifold learning
Through the visual navigation image processing algorithm based on Riemann manifold learning, the problems of low accuracy, low efficiency and information degradation in the prior art are solved, and efficient and robust image set classification of mobile robots in unknown environments is realized.
Patent Information
- Application Number
- CN202410003232.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-02
- Publication Date
- 2025-07-04
AI Technical Summary
The existing image set classification method based on Riemann manifold learning has problems such as limited accuracy, low iteration optimization efficiency and Riemann network information degradation, which is difficult to meet the visual navigation needs of mobile robots in unknown unstructured environments.
The visual navigation image processing algorithm based on Riemann manifold learning is adopted, including image preprocessing, local feature extraction and description, multi-manifold joint characterization and multi-core metric learning, lightweight SPD manifold neural network design, SPD manifold depth metric learning, and deep SPD manifold neural network. Through the Riemann kernel function mapping and nonlinear learning mechanism, combined with Riemann pooling layer and core discriminant analysis, the SPD manifold depth metric learning framework is built, optimized network design and improved classification capabilities.
It improves the accuracy and computing efficiency of image set classification, overcomes the problem of information degradation, and realizes robust representation and precise classification in unknown environments.
Smart Images

Figure CN120259591A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of navigation visual image processing, and particularly to a visual navigation image processing algorithm based on Riemannian manifold learning. Background Art
[0002] In the field of navigation visual image processing, improving the visual navigation ability is crucial for promoting its application in unknown unstructured environments. Deep neural networks can distinguish different categories of images after being well-trained and exhibit excellent performance. However, there are the following three deficiencies in the existing image set classification methods based on Riemannian manifold learning: 1) The accuracy of shallow models is limited; 2) The efficiency of iterative optimization is low; 3) Information degradation of Riemannian networks. Therefore, the existing technology needs to study new methods of representation learning on Riemannian manifolds from three aspects: mining complementary statistical information, optimizing the learning mode of features, and exploring effective deep Riemannian networks. Summary of the Invention
[0003] To solve the above problems existing in the background art, the present invention provides a visual navigation image processing algorithm technology based on Riemannian manifold learning, which can improve the classification accuracy, optimize the network design to improve the calculation efficiency, and explore deep models to achieve robust representation, and has important theoretical and practical value for enriching the research of pattern recognition and computer vision.
[0004] To achieve the above object, the present invention adopts the following algorithm: A visual navigation image processing algorithm based on Riemannian manifold learning, characterized in that: the visual navigation image processing algorithm based on Riemannian flow learning includes preprocessing the image data obtained by acquisition systems such as camera probes. Image preprocessing is an operation to preprocess the data collected by image sensors to improve data quality. High-quality image data is the guarantee and basis for the performance of visual navigation technology. The purpose of image preprocessing is to improve the signal-to-noise ratio of the image and correct the distortion of the image. Existing image preprocessing algorithms are difficult to meet the application requirements of mobile robots in terms of real-time performance of data processing. Processing large amounts of data and high-speed real-time image data remains an important challenge for image preprocessing algorithms. Image contrast enhancement is an image operation to further improve image quality to preserve image detail information or meet the requirements of subsequent algorithm processing on the basis of obtaining image data with high signal-to-noise ratio, stability and reliability. Common operations include white balance, power-law transformation, etc. Different operations are adopted for different processing purposes. In the image processing algorithm of visual navigation, multiple enhancement operations are often required to process in order to obtain better image quality, which poses a challenge to the real-time processing of mobile robot visual navigation image data. Another problem existing in visual image processing algorithms is image matching. The extraction and description of image feature points are important contents of image matching. The feature extraction of an image is an important basis for image recognition, classification and matching. Compared with the characteristics of large global feature calculation amount and low processing efficiency, local features are more suitable for visual navigation applications due to their convenient and efficient processing characteristics. By extracting features such as image points, lines, and surfaces, similarity analysis is performed on multiple images, and then three-dimensional environmental information is obtained as the data basis for navigation. How to quickly and robustly obtain image features and descriptors is the difficulty and challenge of high-speed visual navigation technology. Starting from three solutions of mining complementary information to improve classification accuracy, optimizing network design to improve calculation efficiency, and exploring deep models to achieve robust representation, a new method for representing and learning image set data in the category of Riemannian manifolds is explored. The learning methods include: a classification algorithm based on multi-manifold joint representation and multi-kernel metric learning (improving accuracy); a classification algorithm based on a lightweight SPD manifold neural network (improving efficiency); a classification algorithm based on SPD manifold deep metric learning (improving classification ability); a classification algorithm based on a deep SPD manifold neural network (strengthening classification ability).
[0005] Preferably, for the classification algorithm based on multi-manifold joint representation and multi-kernel metric learning provided by the present invention, three Riemannian kernel functions induced on three Riemannian manifolds map the image data to three RKHSs respectively. This mapping process can not only better preserve the original data structure, but also provide an effective implementation path for subsequent data fusion.
[0006] Preferably, for the classification algorithm based on the lightweight SPD manifold neural network provided by the present invention, an SPD matrix correction layer is designed, and a non-linear learning mechanism is introduced for the SPD matrix by using the defined activation function. Further, a Riemannian pooling layer is designed, and it is demonstrated from three technical aspects how to define the Riemannian pooling operation more appropriately under the current design mode. To implement the image set classification task, then the LogEig layer
[32] is used to map the inferred feature manifold with certain effectiveness and discriminability to a Euclidean flat space. Finally, the kernel discriminant analysis (KDA) algorithm is used to learn the subspace mapping to further enhance the intra-class compactness and inter-class separability of the features.
[0007] Preferably, the classification algorithm based on SPD manifold deep metric learning provided by the present invention is a novel SPD manifold deep metric learning (SMDML) framework proposed for the image set classification task, and relatively excellent classification results are obtained in complex classification scenarios, realizing the joint optimization of Riemannian neural network and Riemannian metric learning. This design not only overcomes the problem of information degradation existing in Riemannian neural networks to a certain extent, but also effectively alleviates the influence of the intra-class discreteness of data and the inter-class ambiguity problem on the classification ability of the algorithm.
[0008] Preferably, for the classification algorithm based on SPD manifold deep metric learning provided by the present invention, a Riemannian network with an identity mapping mechanism is constructed. First, SPDNet
[32] is selected as the backbone network of the proposed model, aiming to obtain a low-dimensional, compact and discriminative SPD matrix. Then, a stacked Riemannian autoencoder (SRAE) is constructed at the output end of the backbone network for realizing deep Riemannian representation. Under the supervision of the reconstruction error term, the network mapping mechanisms of SRAE and each Riemannian autoencoder (RAE) will gradually tend to the identity mapping, thus simplifying the training of the deep network while overcoming the model degradation problem. Considering that deep Riemannian representations often have large intra-class discreteness and inter-class similarity, two progressive metric learning modules are then combined with SRAE to alleviate the influence of the above problems on the feature expression ability of the model. Finally, for the classifiers learned in different RAEs, the maximum voting strategy is used to implement the image set classification. BRIEF DESCRIPTION OF THE DRAWINGS Figure 1 It is a schematic diagram of the research on image set classification based on Riemannian manifold learning proposed by the present invention.
Claims
1. A visual navigation image processing algorithm based on Riemannian manifold learning, characterized in that: The visual navigation image processing algorithm based on Riemannian manifold learning includes preprocessing the image data obtained by acquisition systems such as camera probes. Image preprocessing is an operation to preprocess the data collected by image sensors to improve data quality. High-quality image data is the guarantee and foundation of the performance of visual navigation technology. The purpose of image preprocessing is to improve the signal-to-noise ratio of the image and correct the distortion of the image. Existing image preprocessing algorithms are difficult to meet the application requirements of mobile robots in terms of the real-time performance of data processing. Processing large amounts of data and high-speed real-time image data remains an important challenge for image preprocessing algorithms. Image contrast enhancement is an image operation that, based on obtaining high-signal-to-noise ratio, stable and reliable image data, further improves the image quality to preserve the detailed information of the image or meet the requirements of subsequent algorithm processing. Common operations include white balance, power-law transformation, etc. Different operations are adopted for different processing purposes. In the image processing algorithm of visual navigation, multiple enhancement operations are often required to obtain better image quality, which poses a challenge to the real-time processing of mobile robot visual navigation image data. Another problem existing in the visual image processing algorithm is image matching. The extraction and description of image feature points are important contents of image matching. The feature extraction of the image is an important basis for image recognition, classification, and matching. Compared with the characteristics of large global feature calculation amount and low processing efficiency, local features are more suitable for the application of visual navigation due to their convenient and efficient processing characteristics. By extracting features such as image points, lines, and surfaces, similarity analysis is performed on multiple images, and then three-dimensional environmental information is obtained as the data basis for navigation. How to quickly and robustly obtain image features and descriptors is the difficulty and challenge of high-speed visual navigation technology. Starting from three solutions: mining complementary information to improve classification accuracy, optimizing network design to improve calculation efficiency, and exploring deep models to achieve robust representation, new methods for representing learning of image set data in the category of Riemannian manifolds are explored. The learning methods include: classification algorithms based on multi-manifold joint representation and multi-kernel metric learning (to improve accuracy); classification algorithms based on lightweight SPD manifold neural networks (to improve efficiency); classification algorithms based on SPD manifold deep metric learning (to improve classification ability); classification algorithms based on deep SPD manifold neural networks (to strengthen classification ability).
2. The classification algorithm based on multi-manifold joint representation and multi-kernel metric learning (improving accuracy) according to claim 1, characterized in that: For a given image set data, three different Riemannian descriptors (covariance matrix, linear subspace, and Gaussian embedding model) are used simultaneously to model it in order to extract complementary structured features; 2) Since the three Riemannian manifolds obtained by modeling, namely the SPD manifold, the Grassmann manifold, and the Gaussian manifold, have heterogeneous topological structures, and as introduced in Section 2.4, the dimensions of the matrix elements on the three Riemannian manifolds are also inconsistent, they cannot be directly fused. If the strategy of concatenating after pulling into column vectors is adopted, it will not only cause the "curse of dimensionality", but also completely destroy the Riemannian geometric structure of the original data. In order to make the extracted multi-view Riemannian features play a positive guiding and promoting role in the final image set classification, the Riemannian kernel functions induced on the three Riemannian manifolds are used to map them to three RKHSs respectively. This mapping process can not only better preserve the original data structure, but also provide an effective implementation path for subsequent data fusion; 3) Given that RKHS conforms to the Euclidean space property, the designed multi-kernel metric learning (MKML) framework is used to fuse the generated three high-dimensional kernel features into a low-dimensional common subspace for classification. Considering that different kernel features and different local regions in the kernel features have different contributions to the final classification decision, while MKML learns the similarity distance metric to achieve cross-domain data fusion, the useful information is further strengthened and the interference information is weakened by introducing an attention mechanism (adaptive weighting). Therefore, the learned subspace features will have smaller within-class scatter and between-class similarity.
3. The classification algorithm based on the lightweight SPD manifold neural network (improving efficiency) according to claim 1, characterized in that: Considering that the original SPDNet [32] has poor generalization ability on small-scale datasets, and the cross-entropy loss function it adopts cannot explicitly encode and integrate the geometric distribution information of the data into the network training process, resulting in the classification hypersphere inferred may not accurately reflect the distribution relationship between different categories. Since the Principal Component Analysis Network (PCANet) [35] can capture relatively key data distribution information by cascading two-stage principal component analysis learning without using the backpropagation algorithm, and achieves competitive classification results, especially when the data scale is limited. At the same time, the lightweight network architecture of PCANet itself greatly improves the computational efficiency. Inspired by this, and referring to the construction paradigm of the original SPDNet [32], this chapter attempts to design a lightweight SPD manifold network for image set classification with the following two characteristics: 1) Compared with some representative image set classification algorithms, this model can show competitive classification ability; 2) It has high computational efficiency. Considering that the Bidirectional Two-Dimensional Principal Component Analysis ((2D)2PCA) [2] algorithm does not destroy the Riemannian geometric properties while operating on the SPD matrix, a bilinear mapping layer of the SPD matrix based on the (2D)2PCA algorithm is first designed to implement an unsupervised weight optimization strategy. Then, an SPD matrix correction layer is designed, and a defined activation function is used to introduce a non-linear learning mechanism for the SPD matrix. Since traditional pooling operations can effectively reduce network parameters and enhance the discriminative ability of the learned features, a Riemannian pooling layer is further designed, and it is demonstrated from three technical aspects how to define the Riemannian pooling operation more appropriately in the current design mode. To achieve the image set classification task, then the LogEig layer [32] is used to map the inferred feature manifold with certain effectiveness and discriminability to a Euclidean flat space.
4. The classification algorithm based on SPD manifold depth metric learning (improving classification ability) according to claim 1, characterized in that: A novel metric learning regularization term is designed and introduced into the loss function of the network. Since this regularization term realizes the explicit encoding and integration of the within-class and between-class distribution information of the data into the end-to-end training of the network, a more discriminative low-dimensional feature manifold can be inferred for image set classification. The ReCov operation can reduce the within-class scatter of the data to a certain extent and enhance the representation learning ability of the model. As also mentioned in Chapter 4, training a Riemannian network with a single cross-entropy loss function is not sufficient to extract powerful geometric semantic information because it cannot effectively encode and analyze the geometric distribution of the data. Therefore, the designed metric learning regularization term is added to the loss function of the network. By explicitly embedding the geometric distribution information of the data into the training process of the network, not only can more discriminative Riemannian features be inferred, but also an effective classifier can be trained.
Citation Information
Cited By
Data intelligent question-answering method and system based on retrieval enhancement generation
CN120596645A
Data intelligent question answering method and system based on search enhancement generated data
CN120596645B
Image classification method based on multi-manifold joint metric learning
CN121330389A
Cross-domain small sample object image classification method and device for intelligent terminal
CN122116010A