Low-resolution aerial video relocation perception learning method based on semi-supervised Hessian regularization

By applying the semi-supervised Hessian regularization method in low-resolution aerial videos, generating gaze transfer path features and evaluating the importance of video grids using Gaussian hybrid models, the problem of relocation of low-resolution aerial videos is solved, and efficient feature selection and video relocation effects are achieved.

CN119992371APending Publication Date: 2025-05-13HANGZHOU YUANTIAO TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411839455.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-13
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively solve the relocation problem of low-resolution aerial videos, especially in the absence of sufficient labeled data, which makes it difficult to simulate human visual perception, resulting in poor feature selection and video relocation effects.

Method used

Using a semi-supervised Hessian regularization method, the gaze transfer path features are generated through uncertainty sampling method, high-quality features are selected from them using Hessian regularization feature selection, and the importance of each grid in the video is evaluated using Gaussian hybrid model to achieve video redirection.

Benefits of technology

It improves the relevance and accuracy of feature selection of low-resolution aerial videos, reduces dependence on a large number of manual annotations, reduces costs, and enhances the robustness of the model for complex aerial videos, ensuring the accuracy and visual effects of information when the video is displayed on different display devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119992371A_ABST
    Figure CN119992371A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of computer vision and video processing, and discloses a low-resolution aerial video relocation perception learning method based on semi-supervised Hessian regularization, which comprises the following steps of: 1, generating a gaze transfer path feature from a low-resolution aerial video by using an uncertainty sampling method; step 2, selecting high-quality features from the gaze transfer path features by using Hessian regularization feature selection; 3, capturing the distribution characteristics of the selected high-quality features by using a Gaussian mixture model; and 4, evaluating the importance of each grid in the video according to the Gaussian mixture model, and redirecting the low-resolution aerial video. According to the method, the gaze transfer path features are generated from the low-resolution aerial video by using the active learning algorithm, the method can effectively simulate human visual attention, and key and salient regions in the video are preferentially identified and highlighted, so that the correlation and accuracy of feature selection are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer vision and video processing, and in particular to a low-resolution aerial video relocation perception learning method based on semi-supervised Hessian regularization. Background Art

[0002] Recent advances in space science, engineering, and long-distance communications have enabled the launch of a large number of Earth observation satellites. These satellites can be roughly divided into two categories based on their operating altitudes: high-altitude satellites and low-altitude satellites. Compared with low-altitude satellites, high-altitude satellites are able to cover larger areas. In practical application areas, the ability to interpret the semantics of low-resolution (LR) aerial photographs is increasingly becoming a key module for various artificial intelligence applications. In remote sensing applications, the technology of relocalizing LR aerial videos by identifying and emphasizing their semantic areas is crucial. For example, this allows for the optimized display of complex street maps on LR displays, facilitating intelligent navigation. The relocalization algorithm can concisely present the planned route, enhance the driver's navigation experience, and potentially reduce road accidents by helping the driver stay focused. In addition, developing a model for relocalizing LR aerial videos for devices such as iPhones or Apple Watches can significantly help refugees flee disasters more effectively.

[0003] Every day, a large number of aerial photographs are synchronously captured by various Earth observation satellites operating at different altitudes. Typically, satellites located at higher altitudes capture low-resolution (LR) videos covering a wide area, while satellites located at lower altitudes capture high-resolution (HR) videos, focusing on a smaller area. Relocalization involves resizing LR aerial photographs by scaling them down vertically or horizontally to ensure that they can be optimally viewed on screens of various sizes. This process is crucial in remote sensing applications such as intelligent navigation and disaster management. However, relocalization poses huge challenges, mainly due to the complexity of replicating human visual perception and the lack of labeled LR aerial videos required to train effective semantic models. Summary of the invention

[0004] The purpose of the present invention is to provide a low-resolution aerial video relocation perception learning method based on semi-supervised Hessian regularization to solve the problems raised in the above background technology.

[0005] To achieve the above object, the present invention provides the following technical solutions:

[0006] A low-resolution aerial video relocalization perception learning method based on semi-supervised Hessian regularization includes a model training phase and a testing phase. The specific process includes:

[0007] Step 1: Generate gaze transfer path features from low-resolution aerial videos using uncertainty sampling (a strategy in active learning, the core idea of ​​which is to select samples that are the most uncertain or most likely to change the model prediction to request annotation);

[0008] Step 2, using Hessian regularized feature selection to select high-quality features from the gaze transfer path features;

[0009] Step 3, using a Gaussian mixture model to capture the distribution characteristics of the selected high-quality features;

[0010] In step 4, the importance of each grid in the video is evaluated based on the Gaussian mixture model, and the low-resolution aerial video is redirected.

[0011] Furthermore, the step 1 comprises:

[0012] Step 1.1, extract multiple overlapping video blocks from the low-resolution aerial video, each video block represents a part of the video;

[0013] Step 1.2: Use the pre-trained convolutional neural network to extract features from each video block and generate the corresponding visual feature vector. In order to further optimize the representation of the features and reconstruct the relationship between each video block and its spatial neighbors, a specific optimization expression is used to measure the importance of the features. The expression is as follows:

[0014]

[0015] The set {y1,y2,…,y N} represents the visual features of N video blocks, y i is one of the traversal values; A ij represents the importance of each video block in reconstructing its spatial neighbors. A is a matrix, i and j represent the i-th row and j-th column of the matrix. N is the total number of blocks in an aerial video. is the neighbor of each block of the aerial video; Argmin represents the minimum value of the parameter;

[0016] Step 1.3, based on the visual feature vector generated in step 1.2, calculate the saliency score of each video block to evaluate its importance in the video; for this purpose, define a metric θ of the reconstruction target, which takes into account the difference between the reconstructed feature vector and the original feature vector of the video block, as well as the relationship between the video block and its spatial neighbors. The specific expression is as follows:

[0017]

[0018] In the formula, θ is the metric of the reconstruction target, b iis the reconstructed feature vector of the ith video block, μ is used to adjust the saliency of the regularization term, K represents the number of visually attractive patches selected, It consists of selected visually rich video blocks, b si is the reconstructed feature vector of the i-th video block, a si is the original eigenvector;

[0019] Step 1.4, using human gaze data or simulating the mechanism of human visual attention, analyze the relationship between video blocks and generate gaze transfer path features, which describe the gaze order of humans when watching videos; help understand how visual attention is transferred between different video blocks. In order to determine the next video block to be watched, the following optimization problem is used for calculation, and the expression is as follows:

[0020]

[0021] In the formula, s K′+1 is the index of the selected video block, argmin means taking the minimum value, H is the matrix related to the feature, H ii For the diagonal elements of this matrix, H i is the vector corresponding to the i-th row of the submatrix, M is a matrix representing the interaction between the features of the video blocks, s1 to s K′ Represents the index of the video block that has been selected;

[0022] In step 1.5, the generated gaze transfer path features are associated with the corresponding video blocks for subsequent feature selection and video redirection processing.

[0023] Further, step 2 includes:

[0024] Step 2.1, receiving the gaze transfer path features generated from step 1 and the labeled video blocks as labeled samples, the labeled samples contain known category labels; using the labeled samples, further optimizing the feature selection process to improve the accuracy of the Gaussian mixture model's understanding of the video content, and achieving this goal by minimizing an objective function that includes a loss function and a regularization term, the expression is as follows:

[0025]

[0026] Where ε(P) is the pre-set loss function; Indicates using the features in the matrix P to perform minimum optimization; represents the regularizer, Indicates weighting of the regularizer;

[0027] Step 2.2, construct a feature selection model, which uses the Hessian regularization algorithm to ensure that the geometric structure of the sample in the feature space can be maintained when selecting features;

[0028] Step 2.3, use the labeled samples to train the feature selection model, optimize the objective function to minimize the difference between the predicted label and the true label, and introduce a regularization term to control the sparsity of the feature; specifically, define the objective function, the expression is as follows:

[0029]

[0030] Wherein, the matrix V represents the diagonal matrix defined by the decision criterion (the performance criterion used to evaluate and compare different decision schemes). When the i-th sample is labeled, V = ∞ (Le in the implementation) is set 10 ), otherwise set V = 1; L is the known label matrix, which contains the labels of the labeled samples, and K is the graph Laplacian matrix used to express the similarity of the samples;

[0031] β weights the regularizer. For the above formula, the first two terms ensure that in the semi-supervised learning stage, the predicted aerial video label J should be consistent with the ground truth and the association map constructed by the present invention to the greatest extent. P is a feature selection matrix used to select important features, G represents a sample matrix containing all features, and the η(P) regularization function is used to constrain the complexity and sparsity of the matrix P. Similarly, at the same time, design the regularizer To force the feature selection matrix P to be sparse enough; in addition, Penalize label predictions;

[0032] Step 2.4, use the trained feature selection model to predict the unlabeled samples and generate their corresponding labels; in order to achieve this goal, the objective function of step 2.3 is optimized, and the expression is as follows:

[0033]

[0034] tr(J T KJ) is the graph Laplacian regularization term (a graph-based regularization method, commonly used in graph representation learning, semi-supervised learning, and collaborative filtering tasks); K is the graph Laplacian matrix (an important matrix representing a graph, usually used to analyze the properties of a graph), which encodes the similarity between samples; tr(J T KJ) This term encourages the predicted labels J of adjacent samples (in feature space) to be similar, thus preserving the geometric structure of the samples;

[0035] tr((JL) TV(JL)) is the label consistency term; L is the label matrix of samples with known labels, and V is a diagonal matrix used to handle labeled and unlabeled samples; tr((JL) T V(JL)) ensures that the predicted label J of the labeled sample is consistent with the true label L;

[0036] is a feature selection term (using the model's prediction results for unlabeled samples and using them as pseudo labels to enhance the training process); G is the feature matrix, P is the feature selection matrix, and β is a regularization parameter; This term penalizes the feature selection matrix P and encourages the selection of features with greater correlation with the feature matrix G;

[0037] φ∥P∥ 1 / 2,1 / 2 is the Hessian regularization term; φ is a regularization parameter, ∥P∥ 1 / 2,1 / 2 is a matrix norm that encourages sparsity of the feature selection matrix P; φ∥P∥ 1 / 2,1 / 2 This term helps to select a set of high-quality features while maintaining the geometric structure of the feature space;

[0038] In step 2.5, based on the features of labeled samples and unlabeled samples, a high-quality feature set is selected for use in subsequent video redirection and classification tasks.

[0039] Furthermore, the step 3 comprises:

[0040] Step 3.1, initialize the parameters of the Gaussian mixture model, including the mean, covariance and mixing weight of each Gaussian component; these parameters are crucial for the model to accurately capture the distribution of spatial features of video blocks. The following formula is used to express the significance of the i-th component, where Δ represents the condition under a selected geospatial spectrum, and the expression is as follows:

[0041]

[0042] Where p(μ|Δ) represents the significance of the i-th component, Δ represents the value of l under a selected geographic spatial spectrum. i represents the weight of the i-th Gaussian component, k i represents the probability density function of the i-th Gaussian component, μ represents the characteristics of the gaze transfer path, and parameter α i and Σ i are the mean and variance of the Gaussian mixture model, respectively;

[0043] Step 3.2, use the expectation maximization algorithm to train the Gaussian mixture model, which alternately performs the expectation step and the maximization step through an iterative process: in the expectation step, the posterior probability of each data point belonging to each Gaussian component is calculated given the current Gaussian mixture model parameters; in the maximization step, the Gaussian mixture model parameters are updated to maximize the likelihood function of the observed data, and this process continues until the Gaussian mixture model parameters converge or the preset number of iterations is reached;

[0044] In step 3.3, a Gaussian mixture model is obtained that can accurately describe the statistical distribution of the selected features, providing an importance evaluation of each feature for subsequent video redirection.

[0045] Furthermore, the step 4 comprises:

[0046] Step 4.1, divide the low-resolution aerial video into several grids, and extract the gaze transfer path features from each grid; this step is to evaluate the importance of each grid in the video, so as to effectively transfer the focus in video processing and analysis. Use the Gaussian mixture model to evaluate the feature significance of each grid, and the specific expression is as follows:

[0047]

[0048] represents the Gaussian mixture model fine-tuned by the expectation maximization (EM) optimization process, μ represents the gaze transfer path features extracted from the current test LR aerial photography at this stage, and Υ represents the probability under the selected gaze transfer path condition. The values ​​determined by the Gaussian mixture model were evaluated according to the equation;

[0049] Step 4.2, use the trained Gaussian mixture model to evaluate the gaze transfer path characteristics of each grid and obtain the importance score of each grid. After determining the horizontal importance of each grid, implement the normalization step to adjust these values ​​accordingly. We implement the normalization step to adjust these values ​​accordingly. This step ensures that the sum of the importance scores of all grids is 1, which makes subsequent processing and analysis more convenient. The specific expression is as follows:

[0050]

[0051] In this expression, Represents the normalized grid g i The importance score, s h (g i ) represents the grid g i The original importance score of i s h (g i) represents the comprehensive importance scores of all grids;

[0052] Step 4.3, adjusting the size and position of each grid according to the importance score of the grid to generate the redirected video;

[0053] Step 4.4, output the low-resolution aerial video after redirection processing and information optimization.

[0054] Compared with the prior art, the beneficial effects of the present invention are as follows: in the model training stage, the present invention generates a gaze transfer path (GSP) feature from a low-resolution aerial video by using an active learning algorithm. The method can effectively simulate human visual attention, preferentially identify and highlight key and salient areas in the video, thereby improving the relevance and accuracy of feature selection. Moreover, the active learning strategy allows the system to learn effectively with limited annotated data, reducing the reliance on a large number of manual annotations and reducing costs. The present invention also adopts a Hessian regularized feature selection (HRFS) strategy to effectively reduce the impact of noise and unreliable labels in low-resolution aerial videos on model performance. Through the calculation of HRFS, the method can retain the key features in the video to the greatest extent, while flexibly encoding the topological structure of each salient area, enhancing the robustness of the model in the face of complex and changeable aerial videos. In addition, the present invention evaluates the importance of each grid in the video through a Gaussian mixture model (GMM), optimizes the video redirection process, and ensures the accuracy and visual effect of video information when displayed on different display devices. This comprehensive approach not only improves the retargeting effect of low-resolution aerial videos, but also balances computational efficiency and energy consumption, providing users with a more reliable and efficient video processing solution for a variety of remote sensing applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 The model adjustment phase flow chart of the method of the present invention;

[0056] Figure 2 This is a flow chart of the redirection phase of the method of the present invention. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] See also Figure 1 and Figure 2,A low-resolution aerial video relocalization perceptual learning method based on semi-supervised Hessian regularization, including a model adjustment stage and a retargeting stage.

[0059] 1. Model adjustment stage

[0060] like Figure 1 As shown in the figure, the specific process of the model adjustment stage is as follows:

[0061] 1.1. Generating Gaze Path (GSP) Features from Low-resolution Aerial Videos

[0062] In the context of low-resolution (LR) aerial photography, many video segments do not have enough information to accurately determine their semantic categories. These include background elements that usually do not attract people's attention. In order to enhance the classification of LR aerial videos, the present invention utilizes active learning algorithms to identify and select video segments that are more representative in aerial videos.

[0063] The goal of the present invention is to deploy a machine learning method that can accurately identify the potential distribution of samples. Considering that video blocks that are close to each other in space often have semantic similarities, it is feasible to model each block by linearly combining it with adjacent blocks. Therefore, the goal of the present invention is to calculate reconstruction parameters that can achieve such a representation, and the method of the present invention is based on the principle that spatially adjacent blocks provide a basis for understanding the semantic structure of the video. In order to take advantage of the semantic similarity of video blocks that are close to each other in space, the reconstruction parameter calculation method of the following formula is adopted:

[0064]

[0065] The set {y1,y2,…,y N} represents the visual features of N video blocks, y i is one of the traversal values, where A ij (A is a matrix, i and j represent the i-th row and j-th column of the matrix) represents the importance of each video block in reconstructing its spatial neighbors. Where N is the total number of blocks in an aerial video. is the neighbor of each block of the aerial video. Argmin represents the minimum value of the parameter.

[0066] Suppose there is a small aerial video containing 4 video blocks, representing different areas: buildings, roads, parks and sky. The goal of this invention is to reconstruct the features of these video blocks. In this example, N = 4, yi represents the feature vector of the ith video block. A is a matrix, where A ij represents the weight of the jth video block when reconstructing the ith video block. If the ith video block and the jth video block are neighbors in space, then A ijwill be a non-zero value if , otherwise it is zero.

[0067] In order to evaluate the information content of the selected video block, by specifying {a i As the reconstructed visual block, the present invention achieves this through an objective function that guides the reconstruction process, which is expressed as follows:

[0068]

[0069] Where θ is the metric of the reconstruction objective, μ is used to adjust the saliency of the regularization term, and K represents the number of visually attractive patches selected. It consists of selected video blocks with rich visual information (a certain frame of the video is divided into different block areas), {s1,s2,…,s K} represents the indices of these selected video blocks.

[0070] Assuming that the present invention selects the 2 most visually appealing video blocks (e.g., buildings and roads), then K=2. i is the reconstructed feature vector of the i-th video block, a i is the original eigenvector. measures the reconstruction error of the selected video block, is a regularization term that ensures that the reconstructed video patch maintains the same geometric structure as the original input sample.

[0071] Define matrices A, B and Ω to optimize the above formula. The expression is as follows:

[0072] A = [a1; a2; a3; a4] is the set of basic image blocks

[0073] B = [b1; b2; b3; b4] is the set of reconstructed image blocks

[0074]

[0075] Define the matrix C = (IA) T (IA). Assume A is a matrix where each row represents the features of a video block and each column represents a spatial neighbor. For simplicity, the present invention assumes that each video block is only associated with its direct neighbors. Then IA will be:

[0076]

[0077] a ij (represents the element in the i-th row and j-th column). After the above definition, the objective function to be optimized in the present invention becomes:

[0078] ∈(B)=tr((BA) T(BA))+μtr(B T CB)

[0079] The present invention needs to minimize this function by setting ∈(B) to zero, where ∈(B) is the reconstruction objective function to be optimized, and the gradient of ∈(B) with respect to B is set to zero, and the following can be obtained:

[0080]

[0081] D is a unit diagonal matrix, and the video blocks reconstructed by the present invention are used to reconstruct the optimization error. The present invention determines the reconstruction error by the following method:

[0082]

[0083] in, represents the Frobenius norm (square root of the sum of squares of matrix elements) applied to the matrix. The present invention adopts a step-by-step strategy. Within this framework, the present invention considers selecting a subset of video blocks from the aerial video, denoted as The present invention uses n to represent an n×N diagonal matrix tailored for this situation, and i represents an n×N matrix with 1s on the diagonal and 0s elsewhere. Therefore, the sth matrix is ​​determined by optimizing K′+1 Video patch selection, s K′+1 is the index of the selected video block in the above text, which means Information speculation

[0084] Observing the sparsity of C, the present invention uses the Sherman-Morrison formula (which is a widely recognized method for simplifying matrix inversion calculations). This method allows the present invention to obtain the following results:

[0085]

[0086] Where J = C -1 , J * iJ * is the outer product of the i-th column and i-th row of J. i * and H i * are the corresponding columns and rows of H, n is the unit diagonal matrix, i is the matrix of all 1s, m is the adjustable regularization factor used to adjust the function, H is the matrix related to the feature, H ii For the diagonal elements of this matrix, the objective function becomes:

[0087]

[0088] Definition M = CAA T C, the objective function is simplified to:

[0089]

[0090] The present invention is able to iteratively determine K visually appealing patches in each aerial video. These patches are then sequentially connected to form a gaze transfer path (GSP) that illustrates how humans navigate in different aerial videos. In general, each GSP obtains a 128K-dimensional feature vector, which effectively captures the essence of human visual perception.

[0091] 1.2. Select high-quality features from GSP features using Hessian regularized feature selection (HRFS)

[0092] Definition part: Let G = [g1; ···; g N ] belongs to R N×T As a matrix containing the entire depth representation calculated from GSP, for simplicity, the first L samples are represented as labeled and the remaining samples are recorded as unlabeled. L ] is represented as a matrix containing the labels from the L labeled samples. T×W It is represented as a matrix reflecting the selected features, where W represents the aerial video category.

[0093] The objective function optimized by the present invention is:

[0094]

[0095] Wherein ε(P) is the loss function preset in the present invention; Indicates using the features in the matrix P to perform minimum optimization; represents the regularizer, The present invention uses a mathematical method to represent the affinity graph (a visual tool for organizing and classifying large amounts of information, opinions, data, etc.) as W, where the i-th entity W ij The similarity between samples is quantified. If the i-th sample and the j-th sample are spatially adjacent in the feature space, the present invention sets W i j =1, otherwise W i j = 0. The present invention defines Y as Y ii =W ii The diagonal matrix of is calculated on this basis to represent the Laplace form of the graph.

[0096] If labeled and unlabeled samples are found, the present invention defines the predicted LR aerial video label as Among them, j iRepresents the calculated label of the i-th training sample.

[0097] The optimal J is highly compatible with the real aerial video label L and the association graph constructed by the present invention, and the expression is:

[0098]

[0099] The matrix V represents a diagonal matrix defined using a decision criterion (a performance criterion for evaluating and comparing different decision schemes). When the i-th sample is marked, the present invention sets V = ∞ (Le 10 ), otherwise set V = 1. L is a known label matrix, which contains the labels of labeled samples, and K is a graph Laplacian matrix used to express the similarity of samples. In theory, the decision criterion can maximize the consistency between the calculated labels and the true values. Consider minimizing the label prediction error by discovering unlabeled aerial videos. Finally, the formula for feature selection using the association graph constructed in semi-supervised mode is as follows:

[0100]

[0101] Here, β weights the regularizer. For the above formula, the first two terms ensure that in the semi-supervised learning stage, the predicted aerial video label J should be consistent with the ground truth and the association map constructed by the present invention to the greatest extent possible. P is a feature selection matrix used to select important features. G represents a sample matrix containing all features, and the η(P) regularization function is used to constrain the complexity and sparsity of the matrix P. Similarly, at the same time, design the regularizer To force the feature selection matrix P to be sparse enough. In addition, Penalize label predictions. You can jointly compute the linear classifier (a method that combines multiple linear classification models for classification, i.e., P) and the predicted label (i.e., J).

[0102] Theoretically, when p=1 / 2, the present invention calls the regularizer a Hessian regularizer. Therefore, the above objective function is reorganized as:

[0103]

[0104] tr(J T KJ) is the graph Laplacian regularization term (a graph-based regularization method, commonly used in graph representation learning, semi-supervised learning, and collaborative filtering tasks). K is the graph Laplacian matrix (an important matrix representing a graph, usually used to analyze the properties of a graph), which encodes the similarity between samples. tr(J T KJ) This term encourages the predicted labels J of neighboring samples (in feature space) to be similar, thus preserving the geometric structure of the samples.

[0105] tr((JL) T V(JL)) is the label consistency term. L is the label matrix of samples with known labels, and V is a diagonal matrix used to handle labeled and unlabeled samples. tr((JL) T V(JL)) ensures that the predicted label J of the labeled sample is consistent with the true label L

[0106] is a feature selection term (using the model's prediction results for unlabeled samples and using them as pseudo labels to enhance the training process.). G is the feature matrix, P is the feature selection matrix, and β is a regularization parameter. This term penalizes the feature selection matrix P and encourages the selection of features that are highly correlated with the feature matrix G.

[0107] φ∥P∥ 1 / 2,1 / 2 is the Hessian regularization term. φ is a regularization parameter, ∥P∥ 1 / 2,1 / 2 is a matrix norm that encourages the sparsity of the feature selection matrix P. v∥P∥ 1 / 2,1 / 2 This term helps select a set of high-quality features while preserving the geometric structure of the feature space.

[0108] 1.3. Train a Gaussian mixture model (GMM) to capture the distribution characteristics of the selected features

[0109] There are often subjective differences in the interpretation of low-resolution (LR) aerial videos by different viewers, as different people may perceive the same LR aerial photos differently. In order to address this variability and enhance the relocalization process of LR aerial videos, a Gaussian mixture model (a probabilistic model used to represent complex data distributions composed of multiple Gaussian distributions and normal distributions) is used to capture selected geospatial spectral (GSP) features during the training phase, expressed as:

[0110] p(μ|Δ)=Σ i l i *k i (μ|α i ,Σ i ),

[0111] Where p(μ|Δ) represents the significance of the i-th component, Δ represents the value of l under a selected geographic spatial spectrum. i represents the weight of the i-th Gaussian component, k i represents the probability density function of the i-th Gaussian component, μ represents the GSP feature, and parameter α i and Σ i are respectively the mean and variance of the Gaussian mixture model (GMM) trained by the present invention.

[0112] Assume that there are 3 Gaussian components to describe the aerial video features of the present invention. Each video block (building, road, park and sky) will be modeled as a mixture of these Gaussian components. The present invention randomly initializes the mean α of each component i and the covariance matrix Σ i For each video block, the present invention calculates the probability that it belongs to each Gaussian component. For example, a building video block may belong more to a Gaussian component representing a building, while a sky video block may belong more to a Gaussian component representing the sky.

[0113] In order to measure the similarity between the selected GPS feature pairs, the present invention uses the Euclidean distance (which measures the straight-line distance between two points in the Euclidean space, ie, the shortest distance between them) in the calculation.

[0114] The way humans understand scenery is similar to their experience with a large number of training videos. For an unfamiliar LR aerial photo, the process first determines its GSP and then selects eligible features. The importance of each mesh in the photo is evaluated. In the process of compressing LR aerial videos, a grid-based strategy is adopted to avoid possible distortion caused by the orientation of the triangle mesh.

[0115] Specifically, the test LR aerial photo is divided into a series of equally sized grids. The importance of the horizontal grid g is then determined as follows:

[0116]

[0117] Assume that the present invention needs to reduce the video. The present invention calculates the importance of each grid, which is based on the probability adjusted by GMM. For example, if a grid contains important features (such as the edge of a building), it will be given a higher weight and keep its size during the reduction process.

[0118] represents the Gaussian mixture model fine-tuned by the expectation maximization (EM) optimization process, μ represents the GSP features extracted from the current test LR aerial photography at this stage, and Υ represents the probability under the selected GSP condition. The values ​​determined by GMM are evaluated according to the equation. After determining the horizontal importance of each grid, a normalization step is implemented to adjust these values ​​accordingly, the expression is:

[0119]

[0120] The present invention assumes that the size of the relocated low-resolution (LR) aerial photograph is X×Y. The width of the i-th grid is compressed to Where [·] means rounding the real number to the nearest integer. Normalized vertical significance (the significant features or mechanisms associated in the vertical direction) is calculated in a similar way to the horizontal significance (the significant features or mechanisms associated in the horizontal direction) described above.

[0121] 2. Redirection process

[0122] The redirection process is as follows Figure 2 As shown, the specific steps are as follows:

[0123] 2.1. Divide the low-resolution aerial video into several grids and extract GPS features from each grid

[0124] In this step, the present invention first receives a low-resolution aerial video and divides it into multiple small grids. The purpose of this segmentation operation is to be able to analyze each area in the video separately, so as to more accurately retain important visual information in the subsequent relocalization process. Each grid represents a small block in the video, and the present invention can extract the Gaze Shifting Path (GSP) feature from each grid. The GSP feature is a deep learning feature that can simulate human visual attention and capture the most attractive parts of the video block. This feature extraction process provides the basis for the subsequent grid importance evaluation.

[0125] 2.2. Use the trained GMM model to evaluate the GSP features of each grid and obtain the importance score of each grid

[0126] After extracting the GSP features, the present invention uses a pre-trained Gaussian mixture model (GMM) to evaluate the features of each grid. The GMM model has learned how to identify the key visual elements in the video from the GSP features during the training phase. The present invention inputs the GSP features of each grid into the GMM, and the model outputs the importance score of each grid. This score reflects the visual importance of the grid in the overall video. Grids with high scores should retain more details in the redirected video to ensure that the key information of the video is not lost.

[0127] 2.3. According to the importance score of the grid, adjust the size and position of each grid to generate the redirected video

[0128] After obtaining the importance score of each grid, the present invention adjusts the size and position of the grid in the redirected video according to these scores. This step is content-aware, meaning that the adjustment operation not only considers the size and position of the grid, but also the content within the grid and the relationship between them. For grids with high importance, the present invention maintains or even increases their size to ensure that key features are still clearly visible in the redirected video. For grids with low importance, the present invention can appropriately reduce their size or even remove them if necessary to optimize the spatial distribution and visual focus of the video.

[0129] 2.4. Output low-resolution aerial video after redirection and information optimization

[0130] Finally, the present invention recombines all the adjusted grids to form a new, redirected low-resolution aerial video. This video optimizes the efficiency of visual performance and information transmission while maintaining the original information. The redirected video can be used in various applications, such as intelligent navigation, disaster management, etc., in which the clarity of the video and the accuracy of the information are crucial. Through this process, the present invention can effectively redirect low-resolution aerial videos while retaining key visual information in the video, improving the usability and effect of the video in different application scenarios.

[0131] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A low-resolution aerial video relocalization perception learning method based on semi-supervised Hessian regularization, characterized in that: include: Step 1: Generate gaze transfer path features from low-resolution aerial videos using uncertainty sampling method; Step 2, using Hessian regularized feature selection to select high-quality features from the gaze transfer path features; Step 3, using a Gaussian mixture model to capture the distribution characteristics of the selected high-quality features; In step 4, the importance of each grid in the video is evaluated based on the Gaussian mixture model, and the low-resolution aerial video is redirected.

2. According to claim 1, a low-resolution aerial video relocalization perception learning method based on semi-supervised Hessian regularization is characterized in that: The step 1 comprises: Step 1.1, extract multiple overlapping video blocks from the low-resolution aerial video, each video block represents a part of the video; Step 1.2: Use the pre-trained convolutional neural network to extract features from each video block and generate the corresponding visual feature vector. In order to further optimize the representation of the features and reconstruct the relationship between each video block and its spatial neighbors, a specific optimization expression is used to measure the importance of the features. The expression is as follows: The set {y1,y2,…,y N } represents the visual features of N video blocks, y i is one of the traversal values; A ij represents the importance of each video block in reconstructing its spatial neighbors. A is a matrix, i and j represent the i-th row and j-th column of the matrix. N is the total number of blocks in an aerial video. is the neighbor of each block of the aerial video; Argmin represents the minimum value of the parameter; Step 1.3, based on the visual feature vector generated in step 1.2, calculate the saliency score of each video block to evaluate its importance in the video; for this purpose, define a metric θ of the reconstruction target, which takes into account the difference between the reconstructed feature vector and the original feature vector of the video block, as well as the relationship between the video block and its spatial neighbors. The specific expression is as follows: In the formula, θ is the metric of the reconstruction target, b i is the reconstructed feature vector of the i-th video block, μ is used to adjust the saliency of the regularization term, K represents the number of visually attractive patches selected, It consists of selected video blocks that are rich in visual information, b si is the reconstructed feature vector of the i-th video block, a si is the original eigenvector; Step 1.4, using human gaze data or a mechanism that simulates human visual attention, analyze the relationship between video blocks and generate gaze transfer path features, which describe the gaze sequence of humans when watching videos; in order to determine the next video block to be watched, the following optimization problem is used for calculation, and the expression is as follows: In the formula, s K′+1 is the index of the selected video block, argmin means taking the minimum value, H is the matrix related to the feature, H ii For the diagonal elements of this matrix, H i is the vector corresponding to the i-th row of the submatrix, M is a matrix representing the interaction between the features of the video blocks, s1 to s K′ Represents the index of the video block that has been selected; In step 1.5, the generated gaze transfer path features are associated with the corresponding video blocks for subsequent feature selection and video redirection processing.

3. The low-resolution aerial video relocation perception learning method based on semi-supervised Hessian regularization according to claim 1 is characterized in that: Step 2 includes: Step 2.1, receiving the gaze transfer path features generated from step 1 and the labeled video blocks as labeled samples, the labeled samples contain known category labels; using the labeled samples, further optimize the feature selection process, and achieve this goal by minimizing an objective function containing a loss function and a regularization term, the expression is as follows: Where ε(P) is the pre-set loss function; Indicates using the features in the matrix P to perform minimum optimization; represents the regularizer, Indicates weighting of the regularizer; Step 2.2, construct a feature selection model, which uses the Hessian regularization algorithm to ensure that the geometric structure of the sample in the feature space can be maintained when selecting features; Step 2.3, use the labeled samples to train the feature selection model, optimize the objective function to minimize the difference between the predicted label and the true label, and introduce a regularization term to control the sparsity of the feature; specifically, define the objective function, the expression is as follows: Wherein, the matrix V represents the diagonal matrix defined by the decision criterion (used to evaluate and compare the performance criteria of different decision schemes). When the i-th sample is labeled, V = ∞, otherwise V = 1; L is the known label matrix, which contains the labels of the labeled samples, and K is the graph Laplacian matrix used to express the similarity of samples; β weights the regularizer, P is the feature selection matrix used to select important features, G represents the sample matrix containing all features, and η(P) regularization function is used to constrain the complexity and sparsity of matrix P. Similarly, at the same time, design the regularizer To force the feature selection matrix P to be sparse enough; in addition, Penalize label predictions; Step 2.4, use the trained feature selection model to predict the unlabeled samples and generate their corresponding labels; in order to achieve this goal, the objective function of step 2.3 is optimized, and the expression is as follows: tr(J T KJ) is the graph Laplacian regularization term; K is the graph Laplacian matrix, which encodes the similarity between samples; tr(J T KJ) This term encourages the predicted labels J of adjacent samples to be similar, thus maintaining the geometric structure of the samples; tr((JL) T V(JL)) is the label consistency term; L is the label matrix of samples with known labels, and V is a diagonal matrix used to handle labeled and unlabeled samples; tr((JL) T V(JL)) ensures that the predicted label J of the labeled sample is consistent with the true label L; is a feature selection item; G is a feature matrix, P is a feature selection matrix, and β is a regularization parameter; This term penalizes the feature selection matrix P and encourages the selection of features with greater correlation with the feature matrix G; φ∥P∥ 1 / 2,1 / 2 is the Hessian regularization term; φ is a regularization parameter, ∥P∥ 1 / 2,1 / 2 is a matrix norm that encourages sparsity of the feature selection matrix P; φ∥P∥ 1 / 2,1 / 2 This term helps to select a set of high-quality features while maintaining the geometric structure of the feature space; In step 2.5, based on the features of labeled samples and unlabeled samples, a high-quality feature set is selected for use in subsequent video redirection and classification tasks.

4. The low-resolution aerial video relocation perception learning method based on semi-supervised Hessian regularization according to claim 1 is characterized in that: The step 3 comprises: Step 3.1, initialize the parameters of the Gaussian mixture model, including the mean, covariance and mixing weight of each Gaussian component; use the following formula to express the significance of the i-th component, where Δ represents the condition of a selected geographic space spectrum, and the expression is as follows: Where p(μ|Δ) represents the significance of the i-th component, Δ represents the value of l under a selected geographic spatial spectrum. i represents the weight of the i-th Gaussian component, k i represents the probability density function of the i-th Gaussian component, μ represents the characteristics of the gaze transfer path, and parameter α i and Σ i are the mean and variance of the Gaussian mixture model, respectively; Step 3.2, use the expectation maximization algorithm to train the Gaussian mixture model, which alternately performs the expectation step and the maximization step through an iterative process: in the expectation step, the posterior probability of each data point belonging to each Gaussian component is calculated given the current Gaussian mixture model parameters; in the maximization step, the Gaussian mixture model parameters are updated to maximize the likelihood function of the observed data, and this process continues until the Gaussian mixture model parameters converge or the preset number of iterations is reached; In step 3.3, a Gaussian mixture model is obtained that can accurately describe the statistical distribution of the selected features, providing an importance evaluation of each feature for subsequent video redirection.

5. The low-resolution aerial video relocation perception learning method based on semi-supervised Hessian regularization according to claim 1 is characterized in that: The step 4 comprises: Step 4.1, divide the low-resolution aerial video into several grids, and extract the gaze transfer path features from each grid; use the Gaussian mixture model to evaluate the feature significance of each grid, the specific expression is as follows: represents the Gaussian mixture model fine-tuned by the expectation maximization optimization process, μ represents the gaze transfer path features extracted from the current test LR aerial photography at this stage, and Υ represents the probability under the selected gaze transfer path condition. The values ​​determined by the Gaussian mixture model were evaluated according to the equation; Step 4.2: Use the trained Gaussian mixture model to evaluate the gaze transfer path characteristics of each grid and obtain the importance score of each grid. The specific expression is as follows: In this expression, Represents the normalized grid g i The importance score, s h (g i ) represents the grid g i The original importance score of i s h (g i ) represents the comprehensive importance scores of all grids; Step 4.3, adjusting the size and position of each grid according to the importance score of the grid to generate the redirected video; Step 4.4, output the low-resolution aerial video after redirection processing and information optimization.

Citation Information

Patent Citations

  • Structured multi-view Hessian regularized sparse feature selection method

    CN109389127A

  • Underwater moving target tracking method based on sonar image

    CN113052872A

  • Driving behavior collaborative recognition method and device fusing eye movement and scene information

    CN118898830A

  • Method and Apparatus for Preconditioned Predictive Control

    US20190243320A1

  • Unsupervised Latent Low-Rank Projection Learning Method for Feature Extraction of Hyperspectral Images

    US20230114877A1