Image Feature Extraction Method Based on Multidimensional Decomposition
By constructing a full-dimensional sampling cube, energy hierarchical projection, and phase alignment, combined with sparse low-rank decomposition and morphological spectrum consistency verification, the feature drift problem caused by sudden changes in illumination and rapid motion is solved, improving the robustness and recognition accuracy of image feature extraction. It is suitable for high-reliability scenarios such as intelligent monitoring and autonomous driving.
Patent Information
- Application Number
- CN202511218399.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-28
- Publication Date
- 2026-06-30
- Estimated Expiration
- 2045-08-28
AI Technical Summary
In existing technologies, when extracting image features in fast-moving scenarios, sudden changes in illumination cause a shift in the distribution of high-order tensor dimension features, resulting in the recognition model being unable to correctly parse the input features, which affects the reliability and safety of high-risk scenarios such as intelligent monitoring and autonomous driving.
By constructing a full-dimensional sampling cube, performing energy layer projection and phase alignment, and combining sparse low-rank decomposition and morphological spectrum consistency test, stable and discriminative high-order tensor dimensions are selected. Fourier-feature joint energy spectrum is introduced to suppress dynamic energy levels and prevent subspace collapse.
It improves the robustness and discriminative ability of image feature extraction, is suitable for high-reliability visual scenarios, and significantly improves the recognition accuracy in complex dynamic environments.
Smart Images

Figure CN121053475B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and more specifically to an image feature extraction method based on multi-dimensional decomposition. Background Technology
[0002] "Image feature extraction based on multi-dimensional decomposition" refers to a process in image processing and computer vision that moves beyond relying solely on single-dimensional information (such as grayscale, color, or edges) to extract features. Instead, it involves structurally decomposing and analyzing images across multiple dimensions, such as spatial, frequency, color channel, temporal (for dynamic images), and higher-order tensor dimensions. Through decomposition, the complex and coupled multi-source information in the original image can be transformed into several independent or complementary subspace components, thus more clearly characterizing the image's fine-grained structure, texture patterns, and global semantic features. This method improves feature discriminability and stability, reduces the interference of noise and redundant information on modeling, and makes subsequent visual tasks such as classification, recognition, detection, and segmentation more accurate and robust. It is particularly suitable for complex scene analysis requiring multi-level, multi-scale information fusion.
[0003] The existing technology has the following shortcomings:
[0004] In existing technologies, feature extraction from images in fast-moving scenes typically relies on high-order tensor modeling based on multi-dimensional decomposition to enhance feature representation capabilities. However, in practical applications, when sudden changes in lighting occur, such as transitioning from a bright area to a dark area or from a shadow area to a highlight area, the dimensions of the decomposed high-order tensors are prone to overall feature distribution shift. This shift directly leads to a loss of correspondence between the decomposed subspace and the original image structure, causing misalignment of the feature basis originally used to characterize the target shape and dynamic trajectory. Consequently, subsequent recognition models are unable to correctly parse the input features within a short period. As a result, the downstream recognition process completely loses its discriminative ability, leading to errors in the identification of key targets or even the interruption of the overall recognition function, severely impacting reliability and safety in high-risk scenarios such as intelligent monitoring, autonomous driving, and motion analysis.
[0005] The information disclosed in the background section is only intended to enhance the understanding of the background of this disclosure, and therefore may include information that does not constitute prior art known to those skilled in the art. Summary of the Invention
[0006] The purpose of this invention is to provide an image feature extraction method based on multi-dimensional decomposition. By constructing a full-dimensional sampling cube, energy hierarchical projection, phase alignment, sparse low-rank decomposition, and morphological spectrum consistency test, a stable extraction of multi-dimensional image features is achieved. Furthermore, a joint energy spectrum map and dynamic energy level suppression mechanism are introduced to effectively cope with feature drift caused by sudden changes in illumination and rapid motion, prevent subspace structure collapse, improve feature robustness and discriminative ability, and make it suitable for high-reliability visual scenarios, thereby solving the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: an image feature extraction method based on multi-dimensional decomposition, comprising the following steps:
[0008] Construct a full-dimensional sampling cube to simultaneously sample the image in the spatial, color, frequency, and temporal dimensions, generating the original multidimensional tensor basis, which will be used as the data foundation for subsequent processing;
[0009] Using the original multidimensional tensor basis as input, energy hierarchical projection is performed based on the adaptive reference anchor to obtain a multi-scale energy spectrum map, and the dominant frequency band is marked in the energy spectrum map to establish an energy baseline.
[0010] Based on the energy baseline of the dominant frequency band, a structural motion coupled phase alignment map is constructed to correct the phase drift in the multi-scale energy spectrum map, thereby obtaining a spatiotemporal-spectral tensor that maintains consistency in time sequence.
[0011] The spatiotemporal-spectral tensor is sparsely decomposed into low-rank co-decomposition to separate steady-state components and transient perturbations, which are then aggregated to obtain candidate dimension clusters, thereby constructing a dimension pool for subsequent screening.
[0012] A morphological spectrum consistency test is performed on the candidate dimension clusters. By combining geometric relationships to backtrack the spatiotemporal-spectral tensor, stable and discriminative high-order tensor dimensions are selected to form a robust feature subspace.
[0013] Based on the selected higher-order tensor dimensions, a Fourier-feature joint energy spectrum is constructed to track energy level trajectories in real time. When a shift in the higher-order tensor dimension is detected and manifested as instantaneous energy concentration, a spectral energy level suppressor is triggered. The absorption weights are adjusted through a time-domain sliding window to suppress subspace collapse and maintain feature integrity.
[0014] Preferably, the steps for generating the original multidimensional tensor basis are as follows:
[0015] For the input dynamic image source, perform frame-by-frame sampling processing. In each frame, extract the spatial position according to the fixed image block division method, and establish a multi-scale spatial sampling grid within each image block to collect pixel gray value and neighborhood gradient magnitude.
[0016] The image is converted to the CI ELAB color space. In each image block, the channel values of the L*, a, and b channels are extracted, the contrast is calculated, and the color difference is statistically analyzed. The space and color information are then stitched together to form a space-color sampling matrix.
[0017] Two-dimensional Fourier transform and wavelet packet decomposition are performed on the L*, a, and b channels of each image block to extract the frequency domain dominant frequency position, energy distribution, and texture direction index, and generate a frequency response matrix.
[0018] Extract consecutive frames of data from an image sequence at fixed time intervals, perform inter-frame differencing and dynamic consistency evaluation on time series features at the same spatial location, and construct a fourth-order real-valued tensor basis that is uniformly aligned across the four dimensions of space, color, frequency, and time.
[0019] Preferably, the steps for establishing the energy baseline are as follows:
[0020] For each image patch in the original multidimensional tensor basis, an energy description vector is constructed, and the mean brightness, chromaticity variance and mutual information entropy of the color channels, the main peak intensity, main frequency position and bandwidth, the inter-frame pixel variation rate and the spectrum shape change rate are extracted to form a complete set of energy description features.
[0021] Based on the energy description feature set, Euclidean distance calculation and K-means clustering are performed, and the image patch with the minimum information entropy and balanced three-dimensional distribution of color, frequency and time is selected as the adaptive benchmark anchor.
[0022] For each non-anchor image patch, vector projection is performed on the reference anchor in three dimensions: color direction, frequency channel, and time series. The color direction projection value, frequency response vector, and time change weight are extracted to construct an energy space mapping map.
[0023] Extract the dominant frequency bands on the frequency channels from the energy space map and construct a parameter set of the dominant frequency bands that includes the center frequency, bandwidth and energy mean.
[0024] Frequency dimension sampling curves are constructed based on the dominant frequency band parameter set and Savitzky-Golay smoothing is performed to generate a two-dimensional energy baseline matrix, which serves as the standard benchmark for subsequent feature alignment and perturbation detection.
[0025] Preferably, the spatiotemporal-spectral tensor generation steps are as follows:
[0026] Based on the multi-scale energy spectrum map obtained after energy layer projection, the phase response values of all time frames are extracted for each image patch at each frequency point in the dominant frequency band, a frequency-time matrix is constructed, and the phase change region is identified by fitting a third-order polynomial.
[0027] Edge extraction and structure detection are performed on each image frame. A structure-motion coupling vector set is constructed by combining the motion trajectories between image blocks. The image blocks with the highest matching score are selected as phase anchor point regions, and their phase response at all frequency points is recorded as alignment benchmarks.
[0028] The phase offset vector between the non-anchor image patch and the anchor image patch is calculated. The phase alignment is completed by using a Lagrange interpolation algorithm based on structural consistency score weighting. Temporal smoothing is used to control the phase change rate in areas with severe drift.
[0029] The corrected frequency-phase-time response is fused with the spatial and color dimensions of the original multidimensional tensor. The image patch intensity map is then restored using inverse Fourier transform, and the tensor units are updated to generate a temporally consistent spatiotemporal-spectral tensor.
[0030] Preferably, the steps for obtaining candidate dimension clusters are as follows:
[0031] After completing frequency-phase alignment, the spatiotemporal-spectral tensor is sliced according to the spatial location of the image patch, preserving the integrity of the color channel, frequency channel, and time frame dimensions. Min-max normalization is performed on all tensor elements, and mean response matrix and standard deviation matrix in the frequency and time dimensions are calculated to characterize energy fluctuation features.
[0032] Normalized tensors are input into the collaborative decomposition process. Kernel norm minimization is used to constrain low-rank tensors to preserve stable structures, while L1 norm minimization is used to constrain sparse tensors to highlight short-term perturbations. Alternating direction multiplier method is used for iterative optimization. The residual convergence threshold is set to 10 to the power of -6. Lagrange multipliers are continuously updated during iteration to enhance convergence.
[0033] The output consists of two tensors with the same structure as the input: a low-rank tensor representing stable structural features in the image sequence and a sparse tensor set representing transient perturbations and non-target responses. These are used to construct candidate dimension clusters and generate a dimension pool for subsequent filtering.
[0034] Preferably, the step of constructing the candidate dimension pool further includes:
[0035] Based on the low-rank tensor and sparse tensor obtained by sparse low-rank collaborative decomposition, slicing operations are performed along the frequency dimension and time frame dimension, respectively. The average energy response and standard deviation of each frequency point in all image blocks and all time frames are calculated. Frequency points with average values higher than the global mean and standard deviations lower than the preset threshold are identified and defined as stable frequency dimension sets.
[0036] A time frame sequence whose frequency response rate of change is lower than a set rate threshold is defined as a stable time dimension set.
[0037] The combination of the frequency points with the highest non-zero response density in a sparse tensor with the time frame is defined as the set of highly sensitive perturbation dimensions.
[0038] The obtained dimension sets are cross-combined to generate candidate dimension clusters of a four-dimensional index structure. They are then sorted from high to low according to their discriminative scores, and dimension entries with scores not lower than a preset discriminative threshold are selected to form a candidate dimension pool.
[0039] The preferred steps for selecting higher-order tensor dimensions are as follows:
[0040] For each combination of frequency dimension and time dimension in the candidate dimension cluster, perform frequency response time trajectory extraction operation, calculate the maximum value, minimum value, mean, standard deviation and response slope change, calculate morphological spectrum consistency score based on response stability, directional continuity and extreme value distribution balance, and filter dimensions with scores higher than the set threshold.
[0041] Image geometric relationship backtracking analysis is performed on dimensions that pass the frequency response consistency test. By calculating the edge gradient direction difference, frequency response direction difference, and edge continuity score between image patches, image patch dimension combinations with spatial structural consistency and temporal edge stability are identified.
[0042] All dimensions obtained through morphological spectrum consistency testing and image geometric consistency backtracking analysis are remapped back into the spatiotemporal-spectral tensor. The corresponding tensor units are extracted to form a sub-tensor set, and a robust feature subspace containing color channel numbers, frequency channel numbers, time frame numbers, and image patch numbers is constructed.
[0043] Preferably, when a shift in the higher-order tensor dimension is detected and manifested as an instantaneous energy concentration, the spectral level suppressor is triggered. The absorption weights are adjusted through a time-domain sliding window to suppress subspace collapse and maintain feature integrity. The steps are as follows:
[0044] The corresponding response sequences of the selected higher-order tensor dimensions are extracted from the original spatiotemporal-spectral tensor, and after performing Fourier transform, they are paired with the structural response information point by point in the frequency channels to construct the Fourier-characteristic joint energy spectrum.
[0045] Based on the joint energy spectrum, the energy evolution value of each dimension in continuous image frames is extracted, the energy level trajectory is constructed, and the energy level rise rate, abrupt increase magnitude and slope change rate are calculated to determine whether there is a serious shift trend.
[0046] For the higher-order tensor dimensions that are determined to be offset, the spectral energy level suppression factor is calculated, and the frequency attenuation coefficient is constructed based on the magnitude of energy surge and the frequency fluctuation range to complete the suppression of the response value of the high-interference dimension.
[0047] A time-domain sliding window is introduced to analyze the average energy and trend direction of neighboring time frames with the current time frame as the center, and dynamically adjust the absorption weight to control the gradual decay process of high-energy dimensions.
[0048] The joint energy spectrum data after spectral level suppression and glide window adjustment is mapped back to the robust feature subspace tensor unit to complete the feature energy update, ensuring that the feature subspace maintains the integrity of expression and trajectory continuity under strong perturbation.
[0049] The technical effects and advantages provided by the present invention in the above technical solution are as follows:
[0050] This invention achieves unified encoding of multi-source image information by constructing a full-dimensional sampling cube covering four dimensions: space, color, frequency, and time. Energy layering projection and dominant frequency band calibration based on an adaptive reference anchor ensure a clear energy baseline for feature extraction. Phase alignment mapping coupled with structural motion effectively corrects temporal distortions caused by phase drift. Sparse low-rank collaborative decomposition separates steady-state information from transient perturbations, ensuring clear and discernible feature structures. Morphological spectrum consistency checks and geometric relationship backtracking screen out truly stable and discriminative high-order tensor dimensions. Finally, a Fourier-feature joint energy spectrum and dynamic energy level suppression mechanism are introduced to respond and adjust in real time to abnormal energy fluctuations, preventing feature subspace structure collapse. In summary, this method possesses high robustness, high discriminativity, and strong adaptability, significantly improving the stability and recognition accuracy of image feature extraction in complex dynamic environments, making it particularly suitable for high-reliability scenarios such as intelligent monitoring, autonomous driving, and behavior analysis. Attached Figure Description
[0051] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this invention. For those skilled in the art, other drawings can be obtained based on these drawings.
[0052] Figure 1 This is a flowchart of the image feature extraction method based on multi-dimensional decomposition according to the present invention. Detailed Implementation
[0053] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, they are provided so that the description of this disclosure will be more complete and fully convey the concept of the exemplary embodiments to those skilled in the art.
[0054] This invention provides, for example Figure 1The image feature extraction method based on multi-dimensional decomposition shown includes the following steps:
[0055] Construct a full-dimensional sampling cube to simultaneously sample the image in the spatial, color, frequency, and temporal dimensions, generating the original multidimensional tensor basis, which will be used as the data foundation for subsequent processing;
[0056] To ensure the accuracy, robustness, and multi-level expressive power of subsequent feature extraction processes, a four-dimensional sampling cube containing spatial, color, frequency, and temporal dimensions needs to be constructed during the image input stage to form the original multi-dimensional tensor basis. The construction process of this sampling cube is implemented through the following steps:
[0057] The input dynamic image source is processed frame by frame, and joint sampling operations are performed on the spatial and color dimensions of each frame. During spatial sampling, the image is divided into several non-overlapping fixed-size image blocks, with a block size of 32×32 pixels. A multi-scale spatial sampling grid is further established within each block. This grid includes center, edge, and corner positions, and pixel grayscale values and neighborhood gradient values are extracted at each sampling point. To enhance multi-scale information representation, a four-level spatial resolution is constructed using a Laplacian pyramid, and the same grid sampling is performed on images at different resolutions. Finally, a spatial sampling vector containing grayscale responses and local gradient magnitudes at each spatial location is obtained. For color sampling, the original RGB image is first converted to the CIELAB color space, which consists of a luminance component L*, a chromaticity component a (green to red), and a chromaticity component b (blue to yellow), better aligning with human visual perception. In this space, the three channels L*, a*, and b* are read pixel by pixel, and their local contrast, inter-channel difference values, and color difference distribution histograms within image blocks are calculated. In each image block, a spatial-color sampling vector is formed, consisting of spatial grayscale information and color channel responses. The sampling vectors of all image blocks are concatenated to form a complete two-dimensional spatial-color sampling matrix for one frame of the image.
[0058] Based on the obtained space-color sampling matrix, the frequency response is further analyzed to extract the distribution features in the frequency dimension. First, a two-dimensional Fast Fourier Transform (FFT) is applied to the L*, a*, and b* channels of each image patch to obtain the amplitude and phase spectra of each channel at the horizontal and vertical frequencies. Energy normalization is performed on the spectral results of each channel, and the center of the dominant frequency distribution, the peak position of the energy density, and the attenuation slope at the edge frequencies are extracted. Then, a three-level subband decomposition is performed on each image patch using symmetric wavelet packet decomposition (e.g., Daubechies-4 wavelet) to obtain the low-frequency approximate subband, horizontal detail subband, vertical detail subband, and diagonal detail subband. The mean, standard deviation, maximum value, and texture direction index of each subband are calculated. Finally, the Fourier spectral features and the wavelet packet frequency subband features are concatenated to construct the complete response description vector of the image patch in the frequency dimension. This step not only captures the periodic distribution characteristics of the image in the global frequency dimension but also preserves information about local texture changes. The frequency response vectors of all image blocks are combined in spatial order to form a frequency response matrix consistent with the aforementioned space-color sampling matrix structure.
[0059] Based on the obtained spatial, color, and frequency static feature sampling results, temporal sampling is performed to complete the structural modeling of the dynamic image in terms of temporal continuity. In the specific implementation, a fixed time interval is set (e.g., sampling one frame every 100 milliseconds), and N frames of images are continuously acquired from the original video sequence, for example, N=16. For each frame, the above two steps are repeated to generate its corresponding spatial-color sampling matrix and frequency response matrix. Then, a time-series feature set is constructed in the inter-frame dimension. To ensure inter-frame consistency, sampling vectors from different times at the same spatial location are concatenated into a time-series vector, and the first and second differences between frames are calculated to evaluate their stability and trend changes. Image patches exhibiting drastic color jumps or abrupt changes in frequency peaks within three consecutive frames are considered as illumination disturbance regions and are smoothed using a bidirectional moving average filter to buffer the impact of sudden changes. Simultaneously, dynamic consistency evaluation is performed on the inter-frame sampling data, assigning weight decay values to image patches with high dynamic blur to reduce their influence weight in subsequent tensor construction. This step ensures that the spatial, color, and frequency features within each time frame are accurately aligned with the time dimension, establishing a stable temporal evolution path and enhancing the ability to withstand disturbances in complex dynamic scenes.
[0060] After completing feature sampling across the four dimensions mentioned above, based on the spatial location index of the image blocks, the spatial sampling vector, color sampling vector, frequency response vector, and corresponding time-series sampling vector of each image block are concatenated in a unified dimensional order to form a fourth-order real-valued tensor containing spatial, color, frequency, and time dimensions. Each tensor unit is a multidimensional coupled feature vector of a set of image blocks, with tensor dimensions being the number of image blocks × channel dimension × frequency channel dimension × number of time frames. For example, for 128 image blocks, each containing 9 spatial point locations × 3 color channels × 16 frequency features × 16 time frames, the final constructed original multidimensional tensor basis dimension is 128 × 9 × 16 × 16. This tensor structure possesses complete spatial structure preservation capabilities, color contrast expression capabilities, frequency feature decoupling capabilities, and temporal evolution visualization capabilities, serving as a unified data carrier for subsequent energy mapping, phase correction, and dimension filtering processing steps.
[0061] The purpose of constructing a full-dimensional sampling cube is to provide a unified, complete, and structured data foundation for the image feature extraction process, ensuring comparability, consistency, and accuracy in subsequent multi-dimensional decomposition processing. In image processing tasks, single-dimensional sampling often fails to effectively express complex features such as spatial structure, color differences, texture frequency, and dynamic changes in images. This step, through simultaneous sampling of the spatial, color, frequency, and temporal dimensions, integrates the multi-level information of the original image into a four-dimensional tensor structure. This allows the feature representation of each image patch to have multi-dimensional linkages, depicting both static appearance features and preserving dynamic evolution paths. This sampling method not only improves the completeness of image data representation but also enhances the adaptability of features to abnormal environmental changes (such as sudden changes in illumination or rapid movement). It provides accurate and stable data support for subsequent energy layer projection, phase correction, and high-order feature selection, fundamentally enhancing the robustness and practicality of the entire image feature extraction method.
[0062] Using the original multidimensional tensor basis as input, energy hierarchical projection is performed based on the adaptive reference anchor to obtain a multi-scale energy spectrum map, and the dominant frequency band is marked in the energy spectrum map to establish an energy baseline.
[0063] To enhance the stability and separability of the multidimensional tensor basis in subsequent feature alignment and discrimination, an energy-layered projection operation based on an adaptive reference anchor is required after the original multidimensional tensor is constructed. The purpose of this operation is to project and transform the coupled high-dimensional image features into a multi-scale energy spectrum space, while simultaneously identifying the frequency regions where the image energy is mainly concentrated, thereby constructing the energy baseline required for alignment and control. To achieve this objective, this step includes the following processing flow:
[0064] Energy structure analysis is performed on the high-order features of each image patch in the original multidimensional tensor basis to determine representative adaptive reference anchors. Each tensor unit of the original tensor basis consists of spatial location, color channel, frequency channel, and time frame, forming a structurally stable and semantically continuous fourth-order real-valued tensor. Based on this tensor, for each spatial image patch, the mean brightness, chromaticity variance, and inter-channel mutual information entropy of the three channels (L*, a*, b*) are extracted in the color dimension; the spectral peak intensity, dominant frequency position, bandwidth, and spectral energy density distribution after two-dimensional fast Fourier transform are extracted in the frequency dimension; and the pixel value variation rate, color drift rate, and spectral shape change rate between consecutive frames are extracted in the time dimension. The features of the above dimensions are used to construct a 48-dimensional energy description vector. Using the feature set composed of the energy description vectors of all image patches in the image frame as input, the Euclidean distance matrix is calculated, and the image patch with the smallest distance is selected as the candidate set of reference anchors for the current image frame based on energy centrality and relative distribution density. K-means clustering is then performed on the candidate set to select image patches with the closest cluster centers, the smallest entropy of the description vector, and a balanced distribution in the three dimensions of color, frequency, and time, which are then used as the adaptive baseline anchor for the current image frame.
[0065] Using a selected adaptive reference anchor as a reference point, an energy projection space with consistent orientation is constructed. For each non-anchor image patch, its energy description vector is projected onto the reference anchor direction in the color, frequency, and time dimensions, respectively. The color dimension projection calculates the unit vector directions of the three channels (L*, a*, b*), calculates the cosine projection amplitude based on the color vector angle, and preserves the energy distribution along the dominant color direction. The frequency dimension projection uses a spectral shape similarity metric (using normalized cross-correlation) to align the spectral peak position and extract the spectral energy mapping along the dominant frequency direction. In the time dimension, a five-frame time sliding window is used to calculate the energy change slope curve between consecutive frames, and the slope change trend of the current image patch is aligned to the slope trajectory of the reference anchor, thereby obtaining the temporal response similarity interval. Finally, each image patch corresponds to a color direction projection value, a frequency channel response vector, and a time change weight; these three constitute its three-dimensional mapping in the energy space. Based on the mapping values of all image patches, they are reorganized into a three-dimensional array structure according to their spatial index positions, thus constructing the completed multi-scale energy spectrum mapping map.
[0066] The dominant frequency band extraction operation is performed on the frequency channel response of each image patch in the multi-scale energy spectrum map to identify the frequency region with the most concentrated information, strongest stability, and highest discriminative power in the image. The specific method is as follows: First, the statistical distribution of the dominant frequency position of all image patches in the frequency dimension is calculated, and an energy peak heatmap of the frequency channels is plotted to identify the frequency region with the highest response. Second, within the identified frequency interval, the symmetry of the spectrum, spectral energy gradient, and energy accumulation density of the interval are further extracted to screen frequency bands with stable high energy aggregation effects. Then, the inter-frame change rate coupled with the frequency band in the time dimension is curve-fitted, and regions with change rates below a set threshold are retained as frequency bands with strong temporal consistency. Next, the stability of the channels corresponding to the dominant frequency region in the color dimension is tested, and their color difference variance and channel universality score in the global image are calculated to screen color combinations with good channel consistency. After cross-validation of the three-dimensional indicators, the final dominant frequency band range is output, typically consisting of 3 to 5 consecutive frequency points, and their center frequency, bandwidth, mean energy, and frequency response stability are recorded as a complete parameter set for the dominant frequency band.
[0067] An overall energy baseline for the image is constructed as a reference for subsequent feature alignment, interference detection, and subspace suppression operations. The energy baseline is constructed as follows: The average energy at the center frequency of the dominant frequency band parameter set is used as the baseline starting point. Five frequency points are sampled in both the low-frequency and high-frequency directions, spanning the bandwidth, resulting in a total of 11 frequency references. At each reference frequency, the response values of all image patches within the entire image range are taken, the average energy response is calculated, and a one-dimensional energy trend curve is constructed. This curve is then smoothed using Savitzky-Golay filtering to eliminate the noise effects of occasional energy fluctuations. Subsequently, a two-dimensional energy baseline matrix is generated based on the smoothed curve. The columns of the matrix are frequency reference points, the rows are spatial image patch indices, and the cell value is the spatial average energy response at that frequency point. This matrix reflects the correspondence between spatial location and frequency energy, serving as a standard reference surface for the global energy distribution of the image in the frequency domain. In subsequent processing, any region deviating from this baseline matrix by more than a set threshold (e.g., 20%) is identified as an abnormal perturbation region and requires appropriate alignment or suppression processing.
[0068] Savitzky-Golay is a commonly used smoothing filtering algorithm in signal processing. Its core idea is to perform a polynomial fit on the data points within a sliding window, and then replace the original data points with the fitted curve, thus achieving smoothing while preserving the local trends and shape characteristics of the data. In image spectral analysis, the advantage of the Savitzky-Golay filter lies in its ability to better maintain the peak positions, slope variations, and local extremum structures of the data while reducing high-frequency noise interference, unlike ordinary moving average filters which can cause signal dulling or peak "collapse."
[0069] In the embodiments of this invention, Savitzky-Golay smoothing filtering is applied to the processing of frequency energy trend curves. Its main function is to smoothly fit the frequency response curve near the dominant frequency band, thereby eliminating "peak" interference caused by sudden increases or decreases in energy due to individual image blocks or local texture abrupt changes. Through this algorithm, a smooth and continuous energy distribution curve with true statistical significance can be obtained without changing the overall structure of the energy peaks. This makes the constructed energy baseline more stable, continuous, and interpretable, providing a more reliable basis for subsequent feature alignment and energy deviation detection.
[0070] The core function of the step of energy-level projection based on an adaptive reference anchor, using the original multidimensional tensor basis as input, is to perform directional and scale-ordered energy mapping processing on the complexly coupled spatial, color, frequency, and temporal features in high-dimensional image data, thereby extracting energy distribution patterns that truly reflect the image structure and dynamic state at multiple scales. This step establishes a unified energy projection reference direction by selecting representative and stable image patches as reference anchors, avoiding feature misalignment and information aliasing caused by fixed filtering or uniform scales in traditional methods. By performing directional projection in each dimension, a multi-scale energy spectrum map reflecting color changes, frequency response, and temporal evolution is constructed, which helps to comprehensively characterize the energy structure features of the image. Based on this, by statistically analyzing the dominant frequency bands in the frequency distribution and constructing a global energy baseline, subsequent phase alignment and anomaly identification have a clear reference standard. Overall, this step not only improves the separability and stability of image features in the energy domain but also provides the necessary energy foundation and directional guidance for subsequent structure restoration and robust feature extraction.
[0071] Based on the energy baseline of the dominant frequency band, a structural motion coupled phase alignment map is constructed to correct the phase drift in the multi-scale energy spectrum map, thereby obtaining a spatiotemporal-spectral tensor that maintains consistency in time sequence.
[0072] After completing the initial multidimensional tensor construction and energy hierarchical projection based on adaptive reference anchors in the aforementioned steps, a multi-scale energy spectrum map and the energy baseline of the dominant frequency band have been obtained. However, for dynamic image sequences, issues such as actual object motion, scene illumination changes, edge structure perturbations, and viewpoint transformations can cause phase shifts in the frequency response of the same image region in the energy spectrum between different time frames. This shift disrupts the continuity of temporal information and affects the comparability of tensor data in the time dimension. This invention proposes a phase alignment method based on joint modeling of structural information and motion trajectory. By constructing a structural motion coupling mapping and performing frequency phase correction, a spatiotemporal-spectral tensor with phase consistency in the time dimension is finally obtained. The specific steps include:
[0073] In the multi-scale energy spectrum map, the phase response of each image patch is extracted at all frequency points covered by the dominant frequency band. The position of each image patch in the original multidimensional tensor basis is fixed, and its corresponding frequency response data has been generated in the previous stage. For each frequency point in the dominant frequency band, the phase value of the image patch at that frequency is calculated and recorded as the phase angle in radians within a real number interval. This process is repeated to obtain the phase response sequence of each image patch across multiple time frames. A frequency-time matrix is constructed, where the columns represent the dominant frequency points, the rows represent the frame sequence, and the matrix elements are the phase values of the corresponding frames and frequency points. Time series analysis is performed on each column, and the phase change trend curve at each frequency point is calculated using a third-order polynomial fitting method. Time periods where the rate of change exceeds a preset threshold are marked. These abrupt change segments reflect the discontinuous response of the image structure at that frequency point and are the target areas for subsequent correction.
[0074] By combining the edge structure distribution in image frames with the motion trends between image blocks, a mapping relationship between spatial structure and temporal displacement is established to construct a phase alignment model with structure-motion coupling. The specific steps are as follows: For each image frame, the Sobel operator is applied to extract the edge map, and the Canny edge detection method is used for boundary enhancement to form a structural edge map. For image blocks between two consecutive frames, the changes in gray-level centroid position and edge direction gradients are calculated, and motion trajectory description vectors between image blocks are constructed using Euclidean distance and the angle between directions. The edge map and motion vectors are combined to form a spatial-temporal coupling vector set, and all image blocks are matched and scored. Image block pairs with high matching scores are defined as structurally continuous and motion-stable regions. Based on this, a structure-motion consistency reference surface is established. Within this surface, a set of image blocks with stable distribution and the highest score is selected as the phase anchor point region, and their phase response at all dominant frequency points is recorded as the reference benchmark for the entire phase alignment.
[0075] Using the phase anchor point region as an alignment reference, the phase values of other image blocks in the frequency-time matrix are corrected by difference. The correction process consists of two steps. The first step is local phase difference calculation: for each non-anchor image block, the difference between its phase value at each dominant frequency point and the corresponding time frame phase value of the anchor image block is calculated to form a phase offset vector. The second step is phase alignment interpolation: for the above phase offset vector, a weighted Lagrange interpolation method is used to perform interpolation correction in the frequency-time coordinate system to obtain the corrected phase value. The weights are calculated based on the image block's score in the structural consistency score; the higher the score, the stronger the influence during interpolation, thus enhancing the dependence on phase-stable regions. After performing this correction process on all image blocks, the phase matrix of each frame image at all frequency points in the dominant frequency band is regenerated to ensure that its phase trajectory maintains a consistent trend with the anchor point trajectory. For edge regions or regions with structural jumps, a temporal smoothing method is used to control the phase change rate to not exceed a set average rate threshold, thereby suppressing drastic drift.
[0076] The corrected frequency-phase-time data is fused with the spatial and color dimensions of the original tensor basis to generate a final spatiotemporal-spectral tensor with phase consistency characteristics. The fusion process is as follows: using the coordinate index of each image patch's spatial location as a fixed index, the corresponding L*, a*, and b* channel response values in the color channels are used as the color dimension input. Simultaneously, the corrected phase value at each frequency point in the dominant frequency band is combined with the original amplitude response value to form a complex frequency vector. An inverse Fourier transform is used to restore the complex frequency vector to its image domain representation, resulting in an image patch intensity map with phase alignment characteristics. This map is then embedded into the corresponding spatial location, color channel, frequency index, and time frame in the original tensor structure to complete the tensor update. The final constructed fourth-order tensor has dimensions of the total number of image patches, the number of color channels, the number of dominant frequency points, and the number of time frames, respectively, and each frequency point has a unified phase evolution trajectory in the time dimension. This tensor not only maintains the consistency of image structure in space, preserves the discriminative characteristics between channels in the color dimension, and enhances the dominant information in the frequency dimension, but also achieves complete phase alignment in the time dimension, greatly improving the stability and accuracy of feature representation.
[0077] The core function of this step, "constructing a structure-motion coupled phase alignment map based on the energy baseline of the dominant frequency band, correcting phase drift in the multi-scale energy spectrum map, and obtaining a spatiotemporal-spectral tensor that maintains consistency in time sequence," is to solve the problem of frequency and phase drift in images caused by object motion, sudden changes in illumination, or structural perturbations in time series. Traditional methods often rely on fixed window averaging or static filtering, which is difficult to handle structural displacement and spectral misalignment of image content in consecutive frames, easily causing mismatch in the time dimension of high-dimensional tensors. This step extracts image structural edges and motion trajectories, constructs a structure-motion joint vector field, and performs phase difference correction at multiple frequency points based on energy anchor points in the dominant frequency band, thereby achieving phase consistency of dynamic content in frequency space. This method not only restores the spectral synchronization trajectory of the same target in different time frames, but also enhances the continuity and stability of high-order tensors in the time dimension, ensuring that subsequent decomposition and discrimination operations can be performed based on the physically real feature evolution trend, significantly improving the temporal robustness and discrimination accuracy of the entire image feature extraction process in complex motion scenes.
[0078] The spatiotemporal-spectral tensor is sparsely decomposed into low-rank co-decomposition to separate steady-state components and transient perturbations, which are then aggregated to obtain candidate dimension clusters, thereby constructing a dimension pool for subsequent screening.
[0079] After constructing the spatiotemporal-spectral tensor and ensuring frequency-phase alignment, the structural, color, frequency, and temporal information of the original image sequence are embedded in a unified tensor data structure. Although this tensor has high information representation capabilities, it still contains a large amount of interference data from image background, short-term perturbations, and illumination changes. These factors can obscure the feature responses of key targets, affecting subsequent recognition accuracy and model generalization ability. Therefore, a sparse low-rank collaborative decomposition method is proposed to explicitly separate the steady-state components and transient perturbations in the spatiotemporal-spectral tensor. Based on this separation result, a set of candidate dimension clusters is constructed for subsequent high-order screening, forming a dimension pool. The entire processing flow includes the following steps:
[0080] The spatiotemporal-spectral tensors that have completed phase alignment are formatted and their dimensions integrated to establish a unified decomposition input structure. Specifically, the tensor is sliced according to the spatial location of image patches, while preserving the integrity of the color channel, frequency channel, and time frame dimensions corresponding to each image patch. In the color channel dimension, three channels (L*, a*, and b*) are retained. In the frequency dimension, all frequency points in the dominant frequency band are selected as the observation range, for example, 12 discrete frequency channels. In the time dimension, 16 consecutively sampled image frames are selected to form a time series. Therefore, each tensor unit has four distinct dimensions: spatial index, color channel index, frequency channel index, and time frame index, forming a fourth-order tensor of size M×3×12×16, where M is the total number of image patches. Min-max normalization is performed on all elements of this tensor to unify all amplitude response values to the closed interval [0,1]. To facilitate energy distribution analysis, the mean response matrix and standard deviation matrix of the tensor are calculated in the frequency and time dimensions, respectively, to evaluate the fluctuation amplitude of each frequency point in different time frames and provide basic indicators for subsequent disturbance identification.
[0081] Based on the normalized tensor, a sparse low-rank collaborative decomposition is performed, representing the input tensor as the sum of two structurally identical but complementary tensors: one a low-rank tensor and the other a sparse tensor. The decomposition process aims to optimize tensor reconstruction accuracy, employing a joint optimization model of kernel norm minimization and L1 norm minimization. Kernel norm minimization constraints limit the redundancy of the low-rank tensor across color channels, frequency channels, and time frames, ensuring that this part retains only the overall trend and stable structure. L1 norm minimization constraints enhance the non-zero response of the sparse tensor in transient perturbation regions, highlighting abrupt changes while suppressing background regions. The algorithm uses an alternating direction multiplier method, setting the maximum number of iterations to 200. The initial low-rank part is filled with the mean of the entire tensor, while the initial value of the sparse part is set to a zero tensor. In each iteration, the sparse part is fixed first to solve the low-rank part, then the low-rank part is fixed again to solve the sparse part, while simultaneously updating the Lagrange multipliers based on the residuals. The iteration terminates when the residual decreases below 10^(-6), and the final decomposition result is output. During this process, all tensor operations are performed within a four-dimensional structure to ensure consistency in the structural dimension of the decomposition result. Ultimately, two tensors with dimensions identical to the input tensors are obtained: the low-rank tensor represents stable structural features in the image sequence, while the sparse tensor represents short-term perturbations, structural edge variations, and non-target burst responses.
[0082] Here, "Lagrange" refers to the "multiplier" in the Lagrange multiplier method, used to solve optimization problems under constraints. In this invention, the sparse low-rank collaborative decomposition employs an iterative solution process based on the Alternating Direction Multiplier Method (ADMM), where the "Lagrange multiplier" plays a role in coordinating the consistency of constraints among different subproblems. During the decomposition process, the original spatiotemporal tensor is represented as the sum of the low-rank and sparse parts, while requiring that the sum of the two parts approximate the original tensor as closely as possible in each iteration. To satisfy both the optimization objective function (i.e., minimizing the nuclear norm and L1 norm) and the constraint that the sum of the two tensors approximates the original tensor in each iteration, the algorithm introduces the Lagrange multiplier as a "penalty term" for deviations of the current solution from the constraints, and updates the multiplier in each iteration based on the residual (i.e., the difference between the current estimate and the true value). By continuously updating the Lagrange multiplier, the optimization process gradually tends towards a consistent solution that satisfies the constraints, thereby achieving stable and high-precision tensor decomposition. This mechanism is the key mathematical foundation for ensuring the convergence and effectiveness of sparse low-rank cooperative decomposition.
[0083] Based on the two decomposed tensors, candidate dimension clusters are extracted and a complete dimension pool is constructed to prepare for the next stage of feature selection. Specifically, the low-rank tensor is sliced along the frequency dimension, and the average energy response and stability score of each frequency point across all image patches and time frames are calculated. Frequency points with high average response intensity and a response standard deviation less than a set threshold (e.g., 0.03) are selected to form a "stable frequency dimension set." Then, the low-rank tensor is sliced along the time dimension, and time frame sequences with frequency response change rates lower than the average change rate over continuous time periods are extracted to form a "stable time dimension set." Simultaneously, non-zero element density statistics are performed on the sparse tensor, and the combination of the frequency points with the highest non-zero element frequency and the time frames is extracted and defined as the "highly sensitive perturbation dimension set." These three sets are then cross-combined to generate dimension clusters containing a complete four-element index (image patch number, color channel number, frequency point number, and time frame number). For each item in the dimension cluster, its energy contribution ratio in the low-rank and sparse tensors is calculated, and the clusters are sorted according to their discriminant scores. Entries with scores below a set threshold are removed, such as dimension items with a discriminative score below 0.15, ultimately forming a candidate dimension pool. This dimension pool retains complete mapping information for the four dimensions of space, color, frequency, and time. Each dimension combination included has been verified through tensor and decomposition structures, possessing the basic conditions for further processing by subsequent filtering algorithms.
[0084] The step of sparse low-rank co-decomposition of the spatiotemporal-spectral tensor primarily serves to explicitly distinguish stable structural information from transient perturbation information in high-dimensional image data, providing a clear and structured foundational feature representation for subsequent feature selection and discriminative modeling. Since the original tensor contains complex coupled data across four dimensions—space, color, frequency, and time—direct use would lead to the masking or perturbation of key features by factors such as background interference, instantaneous illumination changes, and rapid target movement, severely impacting the discriminativeness and robustness of the features. This step decomposes the tensor using a joint optimization model, extracting the low-rank component to preserve long-term stable core structural responses in the image, such as static backgrounds, persistent textures, and stable frequency responses. Simultaneously, it utilizes the sparse component to capture local anomalies occurring within a short period, such as sudden illumination changes, edge deformations, or atypical motion trajectories. This decomposition not only improves the interpretability of features across different dimensions but also, by constructing candidate dimension clusters, filters out representative and discriminative feature regions, laying a solid foundation for the next step of morphological spectrum consistency verification and significantly improving the stability, robustness, and accuracy of subsequent image understanding and recognition tasks.
[0085] A morphological spectrum consistency test is performed on the candidate dimension clusters. By combining geometric relationships to backtrack the spatiotemporal-spectral tensor, stable and discriminative high-order tensor dimensions are selected to form a robust feature subspace.
[0086] After completing the sparse low-rank collaborative decomposition and extracting candidate dimension clusters, although these clusters exhibit local saliency in frequency and time, some dimensions remain sensitive only to short-term perturbations or local spurious responses, lacking global stability or structural discriminative ability. To further improve the usability and expressive power of feature dimensions, a more rigorous screening process is needed for the candidate dimension clusters. A method combining morphological spectrum consistency testing and image geometric relationship backtracking is proposed to screen high-order tensor dimensions with temporal consistency, frequency structural stability, and spatial structural integrity from the candidate dimension clusters, ultimately constructing a robust feature subspace. The specific steps of this screening process are as follows:
[0087] A morphological spectrum consistency test is performed to identify whether candidate dimensions have a continuous and stable response structure in the frequency and time dimensions. Specifically, this involves extracting the image patch position, color channel number, frequency channel number, and time frame number corresponding to each dimension from the candidate dimension cluster. In the low-rank tensor, the response values of that dimension over the complete time frame sequence within the dominant frequency range are extracted to construct a frequency response time trajectory. Statistical analysis is performed on the trajectory data, including calculating the maximum, minimum, mean, standard deviation, and inter-frame response slope changes. Further, the number of inflection points and local extreme value distribution of the response curves are counted to determine whether they maintain a certain trend continuity and periodic consistency throughout the time series. A morphological spectrum consistency scoring index is set, with response stability, response direction continuity, and extreme value distribution balance as the main evaluation criteria. Dimensions with scores higher than a preset threshold (e.g., 0.85) are considered stable and reliable in spectral morphology and are retained; dimensions with scores lower than the threshold are considered to have severe morphological drift or are severely affected by transient interference and are discarded. This step ensures that the retained dimensions have good temporal continuity and spectral structure consistency, providing a solid foundation for subsequent structural analysis.
[0088] Based on the dimensions identified through morphological spectrum testing, a retrospective analysis based on image geometric structure relationships is performed to filter out dimensions that are spatially isolated, structurally incoherent, or semantically weakly related. Specifically, the spatial positions of all dimensions that passed the previous step are marked in the original image frame, and these positions are mapped to a point matrix on a two-dimensional image coordinate system. The edge gradient direction, edge intensity value, and relative structural distribution of the image patch containing each dimension are calculated. A structural continuity score is constructed by comparing the brightness gradient angle difference, frequency response direction difference, and edge continuity degree between the current image patch and its four adjacent image patches (up, down, left, and right). Image patch groups with high structural continuity scores are considered structural units with good geometric consistency. Furthermore, response trajectories are constructed on time frames for image patches that pass the structural test, and the edge morphological change trends between adjacent frames are evaluated to eliminate candidate dimensions with abrupt boundary directions or broken structures. All dimensions that pass both structural and temporal stability tests are marked as geometrically stable dimensions. This step ensures that the final retained candidate dimensions not only exhibit consistency in spectral response but also possess geometric characteristics such as edge continuity, consistent orientation, and spatial clustering at the image structure level.
[0089] All candidate dimensions that pass both frequency response consistency and spatial geometric consistency tests are remapped back to the original spatiotemporal-spectral tensor to construct a robust feature subspace. Specifically, for each selected dimensional quadruple (image patch number, color channel number, frequency channel number, time frame number), its complete tensor unit is extracted from the corresponding position in the original spatiotemporal-spectral tensor. These tensor units are then aggregated into a subtensor set to construct a higher-order tensor quantum space structure. In this subspace, all dimensional units exhibit strong discriminative responses in the color channel dimension, focus on frequency bands with stable energy structures in the dominant frequency band in the frequency channel dimension, form continuous trajectories without abrupt changes in the time dimension, and cluster in image structure regions with consistent edge directions and good regional connectivity in the spatial dimension. To ensure the overall consistency of feature expression within the subspace, the response variance is calculated for all tensor units constituting the subspace. If a tensor unit with an extreme deviation from the average response exists at a certain frequency point or color channel, interpolation smoothing is performed using local mean regression to correct this. The final robust feature subspace consists of a high-confidence dimension and has clear spatial location identification, time frame index, frequency channel number and color channel mapping. It can be used as a direct input data structure for subsequent Fourier-feature joint energy spectrum tracking and energy level suppression processing.
[0090] The main function of this step, "performing morphological spectrum consistency checks on candidate dimension clusters and combining geometric relationships to backtrack the spatiotemporal-spectral tensor to screen out stable and discriminative high-order tensor dimensions, forming a robust feature subspace," is to significantly improve the discriminability, stability, and physical consistency of the final feature dimensions through a refined screening mechanism during the construction of multidimensional tensor features. Although energy concentration regions have been extracted in previous steps through sparse low-rank decomposition, they may still contain transient interference or structurally isolated pseudo-response dimensions. This step introduces frequency response curve analysis to perform consistency checks on the energy morphology of candidate dimensions over time, removing dimensions with drastic fluctuations or trend breaks. Simultaneously, image geometric relationships are introduced to couple the frequency dimension with the image structure, performing edge direction and connectivity judgments in the spatial structure to further eliminate interfering dimensions outside the target contour or in structurally broken regions. The robust feature subspace that is finally constructed not only has the ability to continuously express in multi-frame sequences, but also has a clear spatial semantic classification, ensuring that subsequent Fourier-feature modeling and energy level change tracking are based on a physically interpretable and mathematically stable set of feature dimensions, thereby comprehensively improving the discrimination accuracy and robustness of subsequent visual tasks.
[0091] Based on the selected higher-order tensor dimensions, a Fourier-feature joint energy spectrum is constructed to track the energy level trajectory in real time. When a shift in the higher-order tensor dimension is detected and manifested as instantaneous energy concentration, a spectral energy level suppressor is triggered. The absorption weight is adjusted through a time-domain sliding window to suppress subspace collapse and maintain feature integrity.
[0092] After completing the screening of high-order tensor dimensions and constructing a robust feature subspace, it is still necessary to monitor and dynamically control these high-confidence feature dimensions in real time to address the problem of abnormal concentration of feature energy caused by sudden perturbations, abrupt changes in illumination, or sudden increases in target motion in image sequences. To this end, a joint method based on Fourier analysis, combined with energy trajectory modeling and spectral energy level suppression mechanisms, is proposed to ensure timely adjustment of feature responses in the event of high-energy perturbations, thus preventing the collapse of the feature subspace structure. This method includes the following steps:
[0093] For each of the higher-order tensor dimensions selected in the previous stage, the image patch location, color channel number, frequency channel number, and time frame number corresponding to that dimension are extracted, and their response sequences in the original spatiotemporal-spectral tensor are obtained. This response sequence represents the temporal variation of a specific region of the image under specific color and frequency channels. Subsequently, a Fourier transform is performed on this response sequence to convert it from the time domain to the frequency domain, obtaining the amplitude and phase information of the corresponding frequency components. While constructing a spectrum in the frequency domain, the structural response information corresponding to that dimension in the original tensor is preserved. The two are then paired point-by-point through the frequency channels to construct a Fourier-feature joint energy spectrum. Each coordinate point in the spectrum represents the product of the structural response intensity of that tensor dimension at a certain time frame and the amplitude of the corresponding frequency component. This joint energy spectrum preserves three aspects of information: temporal variation trend, frequency response intensity, and original image structural features.
[0094] Based on the joint energy spectrum, an energy level trajectory is constructed for each higher-order tensor dimension. This trajectory consists of a sequence of response values for that dimension in the joint energy spectrum along the time frame direction, reflecting its energy evolution trend in consecutive image frames. Local smoothing is performed at each time point in the energy level trajectory, for example, by using a five-point weighted moving average to eliminate high-frequency jitter caused by sampling errors. Furthermore, trend features of the trajectory are extracted, including: the rate of energy level increase between consecutive frames, the sudden increase in the current frame energy level relative to the historical average level, and the rate of change of the trajectory's slope within the sliding window. These quantitative indicators collectively constitute the conditions for determining whether a feature shift has occurred. When the energy level growth rate of a trajectory exceeds twice the historical average slope within a short period (e.g., within 3 frames), and its current frame energy level value is more than twice the historical average energy level value, it can be determined that there is a serious shift trend in that dimension.
[0095] Based on the anomaly detection results of the energy level trajectory, a spectral energy level suppression operation is initiated. Specifically, in higher-order tensor dimensions identified as having shifted, their energy level changes are marked as anomalous, and their spectral energy level suppression factors are calculated. The suppression factor is calculated based on a weighted combination of factors such as the magnitude, duration, and frequency response variation range of that dimension. A set of thresholds is set; if more than three consecutive frames in a certain dimension show significant energy value jumps, and the amplitude fluctuation of the corresponding frequency component exceeds a preset proportion (e.g., ±30%), it is considered a high-interference dimension. For this dimension, a frequency attenuation coefficient is constructed, and the response value of the current frame is multiplied by this coefficient, thereby reducing its actual intensity in the joint energy spectrum.
[0096] A time-domain sliding window is introduced to dynamically adjust the energy level suppression process. The sliding window aims to smoothly control energy response changes and prevent information gaps caused by one-time suppression. A five-frame sliding window is formed, extending two time frames forward and two frames backward from the current time frame. Within this window, the local energy mean and trend direction are calculated for the high-energy dimension in each frame. Based on this trend data, the absorption weights are dynamically adjusted: if the high-energy dimension energy continues to rise within the window, the absorption weight is increased, gradually decaying its energy in subsequent frames; if the energy has begun to decline, the original absorption weight is maintained to avoid over-suppression. The absorption operation decays energy exponentially, effectively suppressing structural perturbations caused by abnormal peaks while ensuring response continuity.
[0097] The joint energy spectrum data, after spectral level suppression and glide window adjustment, is mapped back to the tensor cells at the corresponding locations in the robust feature subspace, replacing the original high-energy response values. This replacement preserves the original structure, channels, frequencies, and temporal information, updating only its energy representation. The resulting feature subspace maintains spatial structural integrity, spectral response stability, and temporal trajectory continuity even under strong external perturbations. This step enables the entire image processing chain to adaptively adjust to anomalies.
[0098] The main function of this step is to dynamically monitor the energy evolution process of features after selecting high-order tensor dimensions with structural stability and discriminative capabilities, and to intervene and control potential anomalies in a timely manner by constructing a Fourier-feature joint energy spectrum, ensuring the robustness and integrity of the feature subspace under complex environments. In image sequences, environmental factors such as sudden strong light, rapid target entry, and occlusion changes can easily cause a sudden surge in the energy of local feature dimensions, leading to the collapse of the overall feature structure and affecting the accuracy of subsequent recognition and modeling. To address such perturbations, this step transforms the time-series response into a frequency domain representation through Fourier analysis, constructs an energy spectrum based on structural features, and detects energy mutation trends in the spectral trajectory. By setting an energy level mutation threshold, the dimension with energy shift is identified, and a time-domain sliding window is used to adjust the energy absorption weights, gradually suppressing abnormal responses and avoiding information breakage. Ultimately, the updated feature subspace possesses strong adaptability and dynamic response capabilities, ensuring consistent performance in terms of temporal continuity, spectral stability, and structural representation integrity, providing a solid foundation for subsequent classification, tracking, and recognition tasks.
[0099] This invention achieves unified encoding of multi-source image information by constructing a full-dimensional sampling cube covering four dimensions: space, color, frequency, and time. Energy layering projection and dominant frequency band calibration based on an adaptive reference anchor ensure a clear energy baseline for feature extraction. Phase alignment mapping coupled with structural motion effectively corrects temporal distortions caused by phase drift. Sparse low-rank collaborative decomposition separates steady-state information from transient perturbations, ensuring clear and discernible feature structures. Morphological spectrum consistency checks and geometric relationship backtracking screen out truly stable and discriminative high-order tensor dimensions. Finally, a Fourier-feature joint energy spectrum and dynamic energy level suppression mechanism are introduced to respond and adjust in real time to abnormal energy fluctuations, preventing feature subspace structure collapse. In summary, this method possesses high robustness, high discriminativity, and strong adaptability, significantly improving the stability and recognition accuracy of image feature extraction in complex dynamic environments, making it particularly suitable for high-reliability scenarios such as intelligent monitoring, autonomous driving, and behavior analysis.
[0100] The foregoing has only described certain exemplary embodiments of the present invention by way of illustration. Undoubtedly, those skilled in the art can modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the foregoing drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.
Claims
1. An image feature extraction method based on multi-dimensional decomposition, characterized in that, Includes the following steps: Construct a full-dimensional sampling cube to simultaneously sample the image in the spatial, color, frequency, and temporal dimensions, generating the original multidimensional tensor basis, which will be used as the data foundation for subsequent processing; Using the original multidimensional tensor basis as input, energy hierarchical projection is performed based on the adaptive reference anchor to obtain a multi-scale energy spectrum map, and the dominant frequency band is marked in the energy spectrum map to establish an energy baseline. The steps for establishing an energy baseline are as follows: For each image patch in the original multidimensional tensor basis, an energy description vector is constructed, and the mean brightness, chromaticity variance and mutual information entropy of the color channels, the main peak intensity, main frequency position and bandwidth, the inter-frame pixel variation rate and the spectrum shape change rate are extracted to form a complete set of energy description features. Based on the energy description feature set, Euclidean distance calculation and K-means clustering are performed, and the image patch with the minimum information entropy and balanced three-dimensional distribution of color, frequency and time is selected as the adaptive benchmark anchor. For each non-anchor image patch, vector projection is performed on the reference anchor in three dimensions: color direction, frequency channel, and time series. The color direction projection value, frequency response vector, and time change weight are extracted to construct an energy space mapping map. Extract the dominant frequency bands on the frequency channels from the energy space map and construct a parameter set of the dominant frequency bands that includes the center frequency, bandwidth and energy mean. Based on the dominant frequency band parameter set, a frequency dimension sampling curve is constructed and Savitzky-Golay smoothing is performed to generate a two-dimensional energy baseline matrix, which serves as the standard benchmark for subsequent feature alignment and perturbation detection. Based on the energy baseline of the dominant frequency band, a structural motion coupled phase alignment map is constructed to correct the phase drift in the multi-scale energy spectrum map, thereby obtaining a spatiotemporal-spectral tensor that maintains consistency in time sequence. The spatiotemporal-spectral tensor is sparsely decomposed into low-rank co-decomposition to separate steady-state components and transient perturbations, which are then aggregated to obtain candidate dimension clusters, thereby constructing a dimension pool for subsequent screening. A morphological spectrum consistency test is performed on the candidate dimension clusters. By combining geometric relationships to backtrack the spatiotemporal-spectral tensor, stable and discriminative high-order tensor dimensions are selected to form a robust feature subspace. Based on the selected higher-order tensor dimensions, a Fourier-feature joint energy spectrum is constructed to track energy level trajectories in real time. When a shift in the higher-order tensor dimension is detected and manifested as instantaneous energy concentration, a spectral energy level suppressor is triggered. The absorption weights are adjusted through a time-domain sliding window to suppress subspace collapse and maintain feature integrity.
2. The image feature extraction method based on multi-dimensional decomposition according to claim 1, characterized in that, The steps for generating the original multidimensional tensor basis are as follows: For the input dynamic image source, perform frame-by-frame sampling processing. In each frame, extract the spatial position according to the fixed image block division method, and establish a multi-scale spatial sampling grid within each image block to collect pixel gray value and neighborhood gradient magnitude. Convert the image to the CIELAB color space, and adjust the L color space in each image block. Channel values are extracted, contrast is calculated, and color difference is statistically analyzed for the three channels a, b. The splicing space and color information are then combined to form a space-color sampling matrix. L for each image patch Two-dimensional Fourier transform and wavelet packet decomposition are performed on channels a, b, and c respectively to extract the frequency domain dominant frequency position, energy distribution, and texture direction index, and generate a frequency response matrix. Extract consecutive frames of data from an image sequence at fixed time intervals, perform inter-frame differencing and dynamic consistency evaluation on time series features at the same spatial location, and construct a fourth-order real-valued tensor basis that is uniformly aligned across the four dimensions of space, color, frequency, and time.
3. The image feature extraction method based on multi-dimensional decomposition according to claim 1, characterized in that, The steps for generating the spacetime-spectral tensor are as follows: Based on the multi-scale energy spectrum map obtained after energy layer projection, the phase response values of all time frames are extracted for each image patch at each frequency point in the dominant frequency band, a frequency-time matrix is constructed, and the phase change region is identified by fitting a third-order polynomial. Edge extraction and structure detection are performed on each image frame. A structure-motion coupling vector set is constructed by combining the motion trajectories between image blocks. The image blocks with the highest matching score are selected as phase anchor point regions, and their phase response at all frequency points is recorded as alignment benchmarks. The phase offset vector between the non-anchor image patch and the anchor image patch is calculated. The phase alignment is completed by using a Lagrange interpolation algorithm based on structural consistency score weighting. Temporal smoothing is used to control the phase change rate in areas with severe drift. The corrected frequency-phase-time response is fused with the spatial and color dimensions of the original multidimensional tensor. The image patch intensity map is then restored using inverse Fourier transform, and the tensor units are updated to generate a temporally consistent spatiotemporal-spectral tensor.
4. The image feature extraction method based on multi-dimensional decomposition according to claim 1, characterized in that, The steps for obtaining candidate dimension clusters are as follows: After completing frequency-phase alignment, the spatiotemporal-spectral tensor is sliced according to the spatial location of the image patch, preserving the integrity of the color channel, frequency channel, and time frame dimensions. Min-max normalization is performed on all tensor elements, and mean response matrix and standard deviation matrix in the frequency and time dimensions are calculated to characterize energy fluctuation features. Normalized tensors are input into the collaborative decomposition process. Kernel norm minimization is used to constrain low-rank tensors to preserve stable structures, while L1 norm minimization is used to constrain sparse tensors to highlight short-term perturbations. Alternating direction multiplier method is used for iterative optimization. The residual convergence threshold is set to 10 to the power of -6. Lagrange multipliers are continuously updated during iteration to enhance convergence. The output consists of two tensors with the same structure as the input: a low-rank tensor representing stable structural features in the image sequence and a sparse tensor set representing transient perturbations and non-target responses. These are used to construct candidate dimension clusters and generate a dimension pool for subsequent filtering.
5. The image feature extraction method based on multi-dimensional decomposition according to claim 4, characterized in that, The steps for constructing the candidate dimension pool further include: Based on the low-rank tensor and sparse tensor obtained by sparse low-rank collaborative decomposition, slicing operations are performed along the frequency dimension and time frame dimension, respectively. The average energy response and standard deviation of each frequency point in all image blocks and all time frames are calculated. Frequency points with average values higher than the global mean and standard deviations lower than the preset threshold are identified and defined as stable frequency dimension sets. A time frame sequence whose frequency response rate of change is lower than a set rate threshold is defined as a stable time dimension set. The combination of the frequency points with the highest non-zero response density in a sparse tensor with the time frame is defined as the set of highly sensitive perturbation dimensions. The obtained dimension sets are cross-combined to generate candidate dimension clusters of a four-dimensional index structure. They are then sorted from high to low according to their discriminative scores, and dimension entries with scores not lower than a preset discriminative threshold are selected to form a candidate dimension pool.
6. The image feature extraction method based on multi-dimensional decomposition according to claim 1, characterized in that, The steps for selecting higher-order tensor dimensions are as follows: For each combination of frequency dimension and time dimension in the candidate dimension cluster, perform frequency response time trajectory extraction operation, calculate the maximum value, minimum value, mean, standard deviation and response slope change, calculate morphological spectrum consistency score based on response stability, directional continuity and extreme value distribution balance, and filter dimensions with scores higher than the set threshold. Image geometric relationship backtracking analysis is performed on dimensions that pass the frequency response consistency test. By calculating the edge gradient direction difference, frequency response direction difference, and edge continuity score between image patches, image patch dimension combinations with spatial structural consistency and temporal edge stability are identified. All dimensions obtained through morphological spectrum consistency testing and image geometric consistency backtracking analysis are remapped back into the spatiotemporal-spectral tensor. The corresponding tensor units are extracted to form a sub-tensor set, and a robust feature subspace containing color channel numbers, frequency channel numbers, time frame numbers, and image patch numbers is constructed.
7. The image feature extraction method based on multi-dimensional decomposition according to claim 1, characterized in that, When a shift in the higher-order tensor dimension is detected and manifested as an instantaneous energy concentration, the spectral level suppressor is triggered. The absorption weights are adjusted through a time-domain sliding window to suppress subspace collapse and maintain feature integrity. The steps are as follows: The corresponding response sequences of the selected higher-order tensor dimensions are extracted from the original spatiotemporal-spectral tensor, and after performing Fourier transform, they are paired with the structural response information point by point in the frequency channels to construct the Fourier-characteristic joint energy spectrum. Based on the joint energy spectrum, the energy evolution value of each dimension in continuous image frames is extracted, the energy level trajectory is constructed, and the energy level rise rate, abrupt increase magnitude and slope change rate are calculated to determine whether there is a serious shift trend. For the higher-order tensor dimensions that are determined to be offset, the spectral energy level suppression factor is calculated, and the frequency attenuation coefficient is constructed based on the magnitude of energy surge and the frequency fluctuation range to complete the suppression of the response value of the high-interference dimension. A time-domain sliding window is introduced to analyze the average energy and trend direction of neighboring time frames with the current time frame as the center, and dynamically adjust the absorption weight to control the gradual decay process of high-energy dimensions. The joint energy spectrum data after spectral level suppression and glide window adjustment is mapped back to the robust feature subspace tensor unit to complete the feature energy update, ensuring that the feature subspace maintains the integrity of expression and trajectory continuity under strong perturbation.
Citation Information
Patent Citations
Unsupervised hyperspectral remote sensing image spatial-spectral feature extraction method
CN111368691A
System and method for optimized processing of information on quantum systems
US20220092035A1