A Smart Image Signal Processing Method and System Based on Multimodal Fusion

By decomposing visible light imaging data and non-visible light spectral data into different functional layers and constructing a multimodal joint optimization model, the contribution ratio is dynamically adjusted, solving the problems of poor adaptability to dynamic scenes and lack of correlation between material reflection information and motion trajectory in multimodal image processing, and achieving high-precision image fusion effect.

CN120318603BActive Publication Date: 2025-10-31BEIJING ZHAOKE HENGXING SCI & TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510797128.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-10-31
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

Existing image processing methods suffer from poor adaptability to dynamic scenes, insufficient collaborative optimization of cross-modal data physical properties, and lack of correlation between material reflection information and motion trajectory in multimodal image processing, resulting in low image fusion accuracy, loss of information in occluded areas, and significant motion artifacts.

Method used

By decomposing visible light imaging data and non-visible light spectral data into different functional layers, and based on the physical property correlation between multimodal data, a multimodal joint optimization model is constructed to dynamically adjust the data contribution ratio and generate fused image signals.

Benefits of technology

It significantly improves the ability to restore information in occluded areas and the continuity of motion trajectories, and solves the problems of detail loss and poor dynamic adaptability caused by fixed weights or single-modal limitations in traditional methods. It is particularly suitable for complex lighting, dynamic target tracking and material recognition scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318603B_ABST
    Figure CN120318603B_ABST
Patent Text Reader

Abstract

This application provides an intelligent image signal processing method and system based on multimodal fusion. Specifically, this application receives multimodal input signals and divides them into a spatial distribution feature extraction region and a dynamic trajectory capture region. These are further decomposed into a penetration feature layer and a material reflection characteristic layer. Missing pixel information in the dynamic trajectory capture region is correlated and mapped with the energy distribution of the penetration feature layer to generate enhanced dynamic trajectory data. A specific reflection mode is identified in the material reflection characteristic layer, and the corresponding frequency band response is frequency-domain superimposed with the spatial distribution feature extraction region to generate composite spatial features. A multimodal joint optimization model is then constructed, and the contribution ratio of the two data points is adjusted to generate a fused image signal. The technical solution provided by this application not only solves the problems of detail loss, artifact generation, and poor dynamic adaptability caused by the limitations of single-modality processing, but also improves the spatial resolution and tracking accuracy of the image signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an intelligent image signal processing method and system based on multimodal fusion. Background Technology

[0002] With rapid social development and progress, image signal processing technology has been widely used in fields such as intelligent security, autonomous driving, medical imaging, and remote sensing monitoring.

[0003] Currently, common image processing methods mainly include weighted averaging strategies based on pixel-level fusion, frequency domain decomposition techniques based on multi-scale transformation (such as wavelet transform and Laplacian pyramid decomposition), and feature-level fusion models based on deep learning. Weighted averaging methods achieve data fusion through fixed or empirical weight allocation, but they are difficult to adapt to the nonlinear correlation of modal characteristics in dynamic scenes. Multi-scale transformation techniques fuse high-frequency and low-frequency components of images by separating them, but they lack targeted modeling of physical differences (such as penetration and reflection characteristics) across modal data. Deep learning-based methods automatically learn fusion rules through end-to-end networks, but they rely on a large amount of labeled data and have poor model interpretability, limiting their application in resource-constrained scenarios.

[0004] However, existing methods have significant shortcomings in multimodal image processing. First, fixed-weight fusion strategies cannot dynamically adapt to the differences in physical characteristics of different modal data, leading to the loss of information in occluded areas or fusion artifacts. Second, multi-scale decomposition methods have poor compatibility with the frequency domain characteristics of non-visible light data (such as millimeter waves and terahertz waves), and imbalances in frequency band energy distribution can easily cause a break in the correlation between material reflection information and motion features. In addition, data-driven deep learning models, due to their black-box nature, have difficulty in achieving cross-modal temporal synchronization and synergistic optimization of physical mechanisms, and are prone to insufficient model generalization ability in low signal-to-noise ratio or small sample scenarios. Summary of the Invention

[0005] This application provides an intelligent image signal processing method and system based on multimodal fusion to solve the technical problems in the prior art, such as poor adaptability to dynamic scenes, insufficient collaborative optimization of cross-modal data physical characteristics, and lack of correlation between material reflection information and motion trajectory, resulting in low image fusion accuracy, loss of occluded area information, and significant motion artifacts.

[0006] In a first aspect, this application provides an intelligent image signal processing method based on multimodal fusion, comprising:

[0007] Receive multimodal input signals, wherein the multimodal input signals include visible light imaging data and non-visible light spectral data;

[0008] The visible light imaging data is divided into a spatial distribution feature extraction region and a dynamic change trajectory capture region, while the non-visible light spectral data is decomposed into a penetration feature layer and a material reflection characteristic layer.

[0009] The missing pixel information in the dynamically changing trajectory capture area is correlated and mapped with the energy distribution of the penetrating feature layer to generate enhanced dynamic trajectory data;

[0010] The reflection pattern matching a specific material in the target scene is identified in the material reflection characteristic layer, and the frequency band response corresponding to the reflection pattern is superimposed with the spatial distribution feature extraction area in the frequency domain to generate a composite spatial feature;

[0011] Based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, a multimodal joint optimization model is constructed. The contribution ratio of the visible light imaging data and the non-visible light spectral data is dynamically adjusted using the multimodal joint optimization model to generate a fused image signal.

[0012] Optionally, a reflection pattern matching a specific material in the target scene is identified in the material reflection characteristic layer, and the frequency band response corresponding to the reflection pattern is superimposed with the spatial distribution feature extraction area in the frequency domain to generate a composite spatial feature, including:

[0013] Based on a preset material reflection characteristic library, the material reflection characteristic layer is segmented by frequency band, and the frequency band range of the reflection mode corresponding to a specific material in the target scene is extracted.

[0014] The energy distribution within the frequency band of the reflection mode is dynamically extracted by bandpass filtering, and the filtered reflection mode frequency band response containing the target material density parameter is generated step by step based on the extracted energy.

[0015] The frequency band response of the filtered reflection mode is multiplied by the frequency domain data corresponding to the spatial distribution feature extraction region to obtain the modulated spatial feature frequency domain distribution.

[0016] The frequency domain distribution of the modulated spatial features is normalized based on the target material density parameter to eliminate the energy scale differences between different modes.

[0017] The normalized frequency domain distribution is weighted and superimposed with the original spatial features of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier.

[0018] Optionally, the normalized frequency domain distribution is weighted and superimposed with the original spatial features of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier, including:

[0019] Based on the target material density parameter and the energy distribution of the corresponding region in the penetrating feature layer, the dynamic fusion weight is calculated;

[0020] The normalized frequency domain distribution is converted into a complex domain expression, and the original spatial features are decomposed into real and imaginary components.

[0021] The magnitude of the complex domain representation is nonlinearly scaled according to the dynamic fusion weights to obtain the scaled magnitude distribution, while retaining the phase information of the original spatial features;

[0022] The scaled amplitude distribution is orthogonally projected onto the imaginary component of the original spatial features to generate initial fused frequency domain data.

[0023] Based on the energy gradient changes in adjacent regions of the penetrating feature layer, energy equalization compensation is performed on the initial fused frequency domain data to eliminate local distortions caused by differences in multimodal data coverage.

[0024] The compensated frequency domain data is combined with the real components of the original spatial features to perform complex number reconstruction, generating a composite spatial feature of the target material identifier.

[0025] Optionally, the missing pixel information in the dynamically changing trajectory capture area is correlated and mapped with the energy distribution of the penetrating feature layer to generate enhanced dynamic trajectory data, including:

[0026] Detect pixel-deficient areas caused by dynamic blur in the dynamic change trajectory capture area, and mark the pixel-deficient areas as areas to be compensated;

[0027] Extract the energy distribution pattern in the penetrating feature layer that corresponds to the spatial location of the region to be compensated;

[0028] Based on the energy distribution pattern, energy attenuation directionality analysis is performed on the peripheral adjacent areas of the area to be compensated. By calculating the energy change rate of each pixel in the penetrating feature layer along the motion trajectory, an energy attenuation gradient matrix is ​​generated.

[0029] Based on the distribution characteristics of the energy attenuation gradient matrix and the local energy peak of the energy distribution pattern, dynamic trajectory compensation parameters are generated, which include a compensation intensity coefficient and a direction correction factor.

[0030] Based on the compensation intensity coefficient, multi-scale interpolation is performed on the pixel missing portion of the region to be compensated, and the interpolation path is constrained by the direction correction factor to generate initial compensation trajectory data.

[0031] The initial compensation trajectory data is dynamically corrected using the energy attenuation gradient matrix. The corrected compensation trajectory data is then spatially aligned and energy-fused with the original trajectory data of the dynamically changing trajectory capture area to generate enhanced dynamic trajectory data containing cross-modal compensation information.

[0032] Optionally, multi-scale interpolation is performed on the pixel-missing portion of the region to be compensated based on the compensation intensity coefficient, and the interpolation path is constrained by the orientation correction factor to generate initial compensation trajectory data, including:

[0033] The energy decay gradient matrix is ​​decomposed into multi-scale values ​​according to a preset Gaussian pyramid hierarchy to generate a gradient distribution hierarchy set that matches the resolution of the region to be compensated. Each level in the gradient distribution hierarchy set corresponds to the energy decay change characteristics at different scales.

[0034] Based on the energy attenuation directionality of each level in the gradient distribution hierarchy set, a multi-scale interpolation template is constructed.

[0035] Based on the direction correction factor, the main motion directions in the gradient distribution hierarchy set are filtered to generate hierarchical constraint path parameters;

[0036] In each level, the weight distribution of the multi-scale interpolation template and the directional tolerance range of the hierarchical constraint path parameters are used to perform adaptive bilinear interpolation on the pixel missing part of the region to be compensated, thereby generating hierarchical compensation trajectory data.

[0037] The compensation trajectory data of each level are fused across scales according to the Gaussian pyramid reconstruction rules to generate initial compensation trajectory data that is consistent with the multi-scale energy attenuation characteristics of the penetrating feature layer.

[0038] Optionally, based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, a multimodal joint optimization model is constructed, including:

[0039] The timestamp sequence of the enhanced dynamic trajectory data is phase-matched with the energy change period of the composite spatial feature to generate synchronization time alignment parameters;

[0040] Based on the synchronization timing alignment parameters, a sliding window mutual information calculation is performed on the motion vector amplitude of consecutive frames in the enhanced dynamic trajectory data and the frequency domain energy peak distribution of the composite spatial features to generate modal contribution weight coefficients.

[0041] Extract the spatial distribution boundary of the frequency domain energy peak in the composite spatial features, and map the spatial distribution boundary to the motion vector coverage area corresponding to the enhanced dynamic trajectory data. Calculate the coverage area overlap rate between the frequency domain energy peak distribution boundary and the motion vector coverage area, and generate a spatial coupling difference matrix based on the coverage area overlap rate.

[0042] Based on the gradient distribution direction of the spatial coupling difference matrix, the modal contribution weight coefficient is multiplied with the target material density parameter in the material reflection characteristic layer to generate a dynamic fusion constraint factor.

[0043] Based on the dynamic fusion constraint factor, a trajectory energy conservation equation is constructed in the motion vector space of the enhanced dynamic trajectory data, and a material reflection energy transfer equation is constructed in the frequency domain space of the composite spatial features. By iteratively solving the minimum cross-entropy of the trajectory energy conservation equation and the material reflection energy transfer equation, a joint optimization objective function is generated.

[0044] The energy attenuation gradient matrix of the penetrating feature layer is embedded as a regularization term into the joint optimization objective function. By constraining the frequency band energy allocation ratio of the visible light imaging data and the non-visible light spectral data, a multimodal joint optimization model is constructed to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectral data.

[0045] Optionally, the contribution ratio of the visible light imaging data and the non-visible light spectral data is dynamically adjusted using the multimodal joint optimization model to generate a fused image signal, including:

[0046] Based on the dynamic fusion constraint factor and the energy attenuation gradient matrix of the penetrating feature layer, calculate the frequency band energy allocation ratio parameters of visible light imaging data and non-visible light spectral data;

[0047] Based on the frequency band energy allocation ratio parameters, a visible light frequency band energy allocation weight map and a non-visible light penetrability frequency band energy allocation weight map are generated.

[0048] The visible light imaging data is converted to the frequency domain to generate visible light frequency domain data. The visible light frequency domain data is dynamically truncated according to the visible light frequency band energy allocation weight mapping, and the frequency band components that match the frequency domain energy peak of the composite spatial features are retained to generate weighted visible light frequency domain data.

[0049] The penetrability feature layer of the non-visible light spectrum data is decomposed into frequency bands to obtain the energy of each frequency band after decomposition. The energy of each frequency band after decomposition is dynamically weighted according to the weight mapping of the non-visible light penetrability frequency band energy to generate penetration-compensated frequency domain data.

[0050] The weighted visible light frequency domain data and the penetration compensation frequency domain data are superimposed in frequency bands to obtain a fused frequency domain distribution;

[0051] The fused frequency domain distribution is subjected to inverse transformation to generate an initial fused spatial signal, and the initial fused spatial signal is spatially consistent based on the motion vector coverage area of ​​the enhanced dynamic trajectory data.

[0052] The corrected fused spatial signal is modulated by dot product with the material identification information of the composite spatial features to generate a fused image signal that includes multimodal physical property associations.

[0053] Secondly, this application provides an intelligent image signal processing system based on multimodal fusion, comprising:

[0054] A receiving module is used to receive multimodal input signals, wherein the multimodal input signals include visible light imaging data and non-visible light spectral data;

[0055] The decomposition module is used to divide the visible light imaging data into a spatial distribution feature extraction region and a dynamic change trajectory capture region, and at the same time decompose the non-visible light spectral data into a penetration feature layer and a material reflection characteristic layer.

[0056] The mapping module is used to associate and map the missing pixel information in the dynamic trajectory capture area with the energy distribution of the penetrating feature layer to generate enhanced dynamic trajectory data.

[0057] The identification module is used to identify the reflection pattern that matches a specific material in the target scene in the material reflection characteristic layer, and to superimpose the frequency band response corresponding to the reflection pattern with the spatial distribution feature extraction area in the frequency domain to generate composite spatial features;

[0058] The construction module is used to construct a multimodal joint optimization model based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, and to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectral data using the multimodal joint optimization model to generate a fused image signal.

[0059] Thirdly, embodiments of this application provide a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are to be invoked and executed by the processing component to implement an intelligent image signal processing method based on multimodal fusion as described in the first aspect above.

[0060] Fourthly, embodiments of this application provide a computer storage medium storing a computer program, which, when executed by a computer, implements an intelligent image signal processing method based on multimodal fusion as described in the first aspect.

[0061] In this embodiment, visible light imaging data and non-visible light spectral data are decomposed into different functional layers (spatial distribution feature extraction region, dynamic trajectory capture region, penetrability feature layer, and material reflection characteristic layer). Based on the physical property correlation between multimodal data (such as penetrability energy distribution compensating for dynamic trajectory loss and reflection mode frequency band superposition optimizing spatial features), adaptive collaborative optimization of cross-modal data in dynamic scenes is achieved. By constructing a multimodal joint optimization model to dynamically adjust the contribution ratio, the information restoration capability and motion trajectory continuity of the fused image in the occluded area are significantly improved. This solves the problems of detail loss, artifact generation, and poor dynamic adaptability caused by fixed weights or single-modal limitations in traditional methods, and is especially suitable for complex lighting, dynamic target tracking, and material recognition scenarios.

[0062] Furthermore, by segmenting the frequency bands of the material reflectivity library and dynamically truncating the bands using bandpass filtering, the frequency band responses of the reflectivity modes of specific materials in the target scene are accurately extracted. Combined with dot multiplication and normalization, the energy scale differences between visible and non-visible light modes are effectively eliminated. By weighted superposition of the modulated frequency domain distribution and the original spatial features, a composite spatial feature with material identification is generated, which strengthens the correlation between material reflectivity and spatial distribution and avoids the material misjudgment or edge blurring problems caused by energy imbalance in traditional frequency domain superposition methods. This mechanism significantly improves the accuracy of target material identification in complex scenes and provides feature inputs with consistent physical properties for subsequent multimodal joint optimization, ensuring the synergistic optimization effect of the fused image in terms of material details and motion trajectory compensation.

[0063] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description

[0064] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0065] Figure 1 A flowchart of an intelligent image signal processing method based on multimodal fusion provided in this application is shown;

[0066] Figure 2A schematic diagram of the structure of an intelligent image signal processing system based on multimodal fusion provided in this application is shown;

[0067] Figure 3 A schematic diagram of the structure of a computing device provided in this application is shown. Detailed Implementation

[0068] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0069] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.

[0070] Researchers have found that existing image fusion techniques struggle to balance occlusion information compensation and motion trajectory continuity in dynamic and complex scenes. Furthermore, the physical differences between visible and non-visible light modal data lead to an imbalance in frequency band energy distribution, affecting the synergistic optimization of material recognition and motion tracking. To address this, a multimodal fusion-based intelligent image signal processing method is proposed. This method decomposes the functional layers of visible and non-visible light data, constructs a cross-modal correlation mapping mechanism, and establishes a multimodal joint optimization model. This enables adaptive synergy between material reflection characteristics and motion trajectory compensation in dynamic scenes, effectively improving the spatial resolution of the fused image and the accuracy of dynamic target tracking.

[0071] The technical solutions of this application are applicable to scenarios requiring simultaneous processing of occlusion compensation, material analysis, and motion reconstruction, such as complex road condition perception in autonomous driving, dynamic target recognition in security monitoring, and multimodal lesion analysis in medical imaging. The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this application without creative effort are within the scope of protection of this application.

[0072] Figure 1A flowchart of an intelligent image signal processing method based on multimodal fusion is provided as an embodiment of this application, such as... Figure 1 As shown, the method includes:

[0073] Step 101: Receive a multimodal input signal, wherein the multimodal input signal includes visible light imaging data and non-visible light spectral data;

[0074] In this step, receiving multimodal input signals refers to simultaneously acquiring data sets with different physical properties through multi-source sensing devices. Visible light imaging data consists of image information from the portion of the electromagnetic spectrum perceptible to human vision, acquired by optical imaging devices, including texture details, color distribution, and dynamic change characteristics of the scene. Non-visible light spectrum data consists of electromagnetic wave signals beyond the visible light band acquired by special sensors, such as millimeter-wave signals with penetrating capabilities, infrared radiation data reflecting thermodynamic properties, or terahertz spectra characterizing the resonance features of matter molecules, which can reveal the outline and material properties of obscured objects.

[0075] In this embodiment, a multi-sensor collaborative acquisition mechanism is first used to acquire visible light image sequences and non-visible light spectral data streams in parallel under time synchronization constraints. Dynamic stability processing is applied to the visible light imaging data, employing a motion estimation-based inter-frame alignment algorithm to eliminate jitter artifacts during acquisition. Environmental interference suppression is implemented for the non-visible light spectral data, using frequency domain filtering techniques to remove background noise. Subsequently, a cross-modal spatial mapping model is constructed. By extracting structured feature points from the visible light images and performing geometric matching with high-confidence reflection regions in the non-visible light data, a multi-modal spatially aligned coordinate transformation relationship is generated. Finally, the spatiotemporally aligned visible light and non-visible light data are integrated into a unified format multi-modal input signal, providing a foundation for subsequent feature decomposition.

[0076] For example, in intelligent security scenarios, visible light cameras and millimeter-wave radars deployed indoors operate synchronously. The visible light cameras capture high-definition images of people's activities, while the millimeter-wave radar penetrates walls to detect reflected signals from moving objects in obscured areas. The system ensures that the acquisition times of the two types of data are aligned through hardware synchronization signals and performs electronic image stabilization on the visible light video to eliminate image blur caused by slight camera movement. The millimeter-wave data undergoes adaptive filtering to remove fixed echo interference from wall reflections. By extracting the edge features of doors and windows from the visible light images and the spatial coordinates of moving targets in the millimeter-wave point cloud, a cross-modal spatial mapping relationship is established, ultimately outputting a fused sensing signal of the trajectory of people's activities and objects hidden behind walls.

[0077] 102, the visible light imaging data is divided into a spatial distribution feature extraction region and a dynamic change trajectory capture region, and the non-visible light spectral data is decomposed into a penetration feature layer and a material reflection characteristic layer.

[0078] In this step, the spatial distribution feature extraction region refers to the static image partition in the visible light imaging data that characterizes the geometric structure and texture details of the target object; the dynamic trajectory capture region refers to the temporally continuous image sequence partition containing the displacement features of the moving target. The penetration feature layer refers to the energy distribution layer in the non-visible light spectrum data that reflects the ability of electromagnetic waves to penetrate obstacles, and can characterize the spatial position of the obscured object; the material reflection characteristic layer refers to the feature layer that reveals the material properties of the target through the differences in reflection intensity at different frequency bands.

[0079] In this embodiment, moving target segmentation is first performed on the visible light imaging data, using a background subtraction algorithm to separate static scenes from dynamic regions. Static image partitions containing stable textures are designated as spatial distribution feature extraction regions, while time-varying image sequences containing moving targets are designated as dynamic trajectory capture regions. Simultaneously, frequency domain energy analysis is performed on the non-visible light spectral data. Low-frequency penetrating echoes are extracted using a band-stop filter to form a penetrating feature layer, and a spectral clustering algorithm is used to divide high-frequency reflection signals according to material type, generating a material reflection characteristic layer. Finally, the energy distribution of the penetrating feature layer is mapped to the visible light imaging coordinate system using a spatial projection matrix, ensuring cross-modal data spatial alignment.

[0080] Continuing with intelligent security scenarios, for example, visible light video streams are processed using a three-frame differential method to detect areas of human activity. Ten consecutive frames of images of this area are designated as a dynamic trajectory capture region, while fixed backgrounds such as walls and furniture are designated as spatial distribution feature extraction regions. Millimeter-wave radar data is separated through time-frequency transformation to extract low-frequency signals that penetrate walls, forming a penetration feature layer. High-frequency reflection signals are matched using a material property database to identify the reflection differences between metal instruments and human tissue, constructing a material reflection property layer. The coordinate information of objects hidden behind walls in the penetration feature layer is mapped to the visible light image space through polar coordinates to pixel coordinates, providing a cross-modal data foundation for subsequent dynamic trajectory compensation.

[0081] Step 103: Associate and map the missing pixel information in the dynamic trajectory capture area with the energy distribution of the penetrating feature layer to generate enhanced dynamic trajectory data;

[0082] In this step, missing pixel information refers to image data gaps caused by motion blur, occlusion, or sensor limitations within the dynamic trajectory capture area; the energy distribution of the penetrating feature layer refers to the energy intensity and spatial gradient characteristics formed after electromagnetic waves penetrate obstacles in the non-visible light spectrum data; correlation mapping refers to the compensation mechanism for establishing visible light missing areas and non-visible light energy distribution through cross-modal spatial relationships; and enhanced dynamic trajectory data refers to the continuous motion trajectory dataset generated by fusing visible light dynamic information and non-visible light penetrating features.

[0083] In this embodiment, firstly, motion optical flow analysis is used to detect the missing pixel locations in the dynamically changing trajectory region, and edge contour detection is used to determine the geometric boundaries of the missing region. Then, the energy distribution pattern corresponding to the spatial coordinates in the penetrating feature layer is extracted, and the attenuation trend of the penetrating energy is analyzed using gradient direction histograms. Next, a motion trajectory prediction model based on the energy attenuation gradient is constructed, and an adaptive interpolation algorithm is used to convert the energy intensity of the penetrating feature layer into pixel compensation values. Finally, a multi-scale fusion strategy is employed to spatiotemporally weight and superimpose the non-visible light compensation data with the original visible light trajectory, generating an enhanced dynamic trajectory that eliminates occlusion artifacts. This process overcomes the perception limitations of visible light sensors by leveraging the physical characteristics of the penetrating feature layer, achieving complementary enhancement of cross-modal motion information.

[0084] Continuing with intelligent security scenarios, for example, a trajectory break occurs when a person passes through an area obstructed by a pillar in a visible light surveillance image. Millimeter-wave data from the penetrating feature layer detects a continuously moving energy hotspot behind the pillar. Spatial coordinate mapping determines the location of the corresponding visible light gap area. Based on millimeter-wave energy gradient analysis, the direction of movement is inferred to be from left to right. A bilinear interpolation algorithm with directional constraints is used to generate transition pixels to fill the trajectory gaps. The final output is an enhanced trajectory containing the complete movement path, preserving the high-precision contour features of the visible light while incorporating millimeter-wave penetrating data to reconstruct the motion state during the obstructed phase.

[0085] Step 104: Identify the reflection mode that matches a specific material in the target scene in the material reflection characteristic layer, and superimpose the frequency band response corresponding to the reflection mode with the spatial distribution feature extraction area in the frequency domain to generate composite spatial features;

[0086] In this step, the material reflection characteristic layer refers to the set of features in the non-visible light spectrum data that characterize the electromagnetic wave reflection law of different materials; the reflection mode refers to the reflection intensity distribution characteristics of a specific material in a specific frequency band; the frequency band response refers to the energy distribution pattern of the target material in the characteristic frequency band; and the composite spatial characteristics refer to the multi-dimensional feature expression that integrates the spatial distribution characteristics of visible light and the reflection characteristics of non-visible light materials.

[0087] In this embodiment, firstly, based on a pre-built database of material reflectivity characteristics, a spectral matching algorithm is used to scan the frequency bands of the material reflectivity layer to identify the reflectivity characteristic frequency bands matching the target material. The target frequency band response is extracted using a dynamic bandwidth filter and combined with material density parameters to generate a frequency domain modulation template. Subsequently, a Fourier transform is performed on the visible light spatial distribution feature extraction region, and the frequency domain features are multiplied with the material frequency band response to achieve spectral modulation. Energy density-based normalization eliminates cross-modal energy scale differences, and finally, a phase-preserving inverse transform is used to reconstruct the spatial features, forming a composite feature that simultaneously includes optical texture and material properties. This process achieves deep coupling of material information and spatial features through a physically driven frequency domain fusion mechanism.

[0088] Continuing with intelligent security scenarios, for example, identifying the unique high-frequency reflection patterns of metal tools within millimeter-wave material reflection characteristic layers. Adaptive bandpass filtering is used to extract the characteristic frequency bands of this material, generating a frequency domain mask containing metal density parameters. The visible light spatial distribution feature extraction region, after frequency domain transformation, is multiplied and modulated with the metal frequency band response to enhance the edge sharpness of metal objects in the image. After inverse transformation and reconstruction, metal knives that were previously confused with plastic items in visible light images reveal their unique material identification features, while preserving the spatial details of the original image, forming a composite spatial feature with material discrimination capabilities.

[0089] Step 105: Based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, a multimodal joint optimization model is constructed. The contribution ratio of the visible light imaging data and the non-visible light spectral data is dynamically adjusted using the multimodal joint optimization model to generate a fused image signal.

[0090] In this step, temporal synchronization refers to the alignment of the time series of enhanced dynamic trajectory data with the energy change period of composite spatial features on the time axis; multimodal joint optimization model refers to the fusion framework that coordinates the differences in physical characteristics between visible and non-visible light data through mathematical modeling; contribution ratio refers to the weight coefficients dynamically allocated to different modal data according to scene characteristics during the fusion process; fused image signal refers to the image output with high resolution and physical feature correlation generated by integrating the advantages of multimodal data.

[0091] In this embodiment, a timestamp sequence is first constructed for the enhanced dynamic trajectory data. This sequence is then aligned with the frequency domain energy fluctuation period of the composite spatial features using a phase-matching algorithm to generate frame-level synchronization parameters. The mutual information between the trajectory amplitude and the frequency domain energy peak is calculated using a sliding window, generating weighting coefficients that reflect the modal correlation strength. Subsequently, the spatial boundary distribution of the composite features is extracted, and its overlap rate with the area covered by the trajectory is calculated, constructing a spatial coupling difference matrix. Gradient direction analysis is used to fuse the weighting coefficients with the material density parameters, forming a dynamic fusion constraint factor. Finally, a joint optimization function containing the energy conservation equation and the reflection transfer equation is constructed, embedding the gradient regularization term of the penetration feature. The optimal frequency band energy allocation ratio is iteratively solved to achieve adaptive fusion of multimodal data.

[0092] Continuing with intelligent security scenarios, for example, when a person enters an area obstructed by a pillar, the energy changes in the frequency band of metal reflection in the composite spatial features are temporally synchronized. Through mutual information analysis, the fusion weights are dynamically adjusted to significantly enhance the fusion contribution of millimeter-wave material features in obstructed areas, while appropriately reducing the weight of visible light dynamic trajectories. A joint optimization model, combined with penetration gradient features, ensures that the final fused image highlights the details of the metal contours compensated for by millimeter waves in obstructed areas, while retaining high-definition texture features of visible light imaging in unobstructed areas. Through dynamic weight allocation across modal data, the behavioral trajectory of a person carrying a weapon through an obstructed area is fully reconstructed, while ensuring the spatial resolution and material identification accuracy of the image.

[0093] Researchers have found that existing technologies often misidentify materials in material reflectivity identification due to coarse frequency band segmentation or differences in energy scale. Furthermore, the frequency domain superposition process lacks adaptation to the physical properties of non-visible light spectra, affecting the accuracy of the association between spatial features and material identification. Therefore, in some embodiments, according to step 104, a reflection pattern matching a specific material in the target scene is identified in the material reflectivity layer, and the frequency band response corresponding to the reflection pattern is frequency-domain superimposed with the spatial distribution feature extraction area to generate composite spatial features, including:

[0094] Step 201: Based on a preset material reflection characteristic library, perform frequency band segmentation on the material reflection characteristic layer and extract the frequency band range of the reflection mode corresponding to a specific material in the target scene.

[0095] In this step, the material reflection characteristic library refers to a database of pre-stored reflection intensity and waveform characteristics of different materials under specific electromagnetic bands, including standard reflection spectra of materials such as metals, plastics, and fabrics; the reflection mode frequency band range refers to the characteristic frequency range in which the target material has significant distinguishability in the material reflection characteristic layer, which is manifested as an energy peak area or a waveform abrupt change area.

[0096] In this embodiment, a pre-built material reflectivity library is first loaded, and a sliding window scan is performed on the full-band spectrum of the material's reflectivity layer. A spectral clustering algorithm is then used to match the scanned local spectrum with the standard reflectivity spectrum in the material library, filtering out candidate frequency bands with similarity exceeding a threshold. Next, energy continuity analysis is used to merge adjacent candidate frequency bands, ultimately segmenting a frequency range that completely covers the reflectivity characteristics of the target material.

[0097] Step 202: Dynamically extract the energy distribution within the frequency band of the reflection mode by bandpass filtering, and generate the filtered reflection mode frequency band response containing the target material density parameter step by step based on the extracted energy.

[0098] In this step, bandpass filtering refers to the technique of selectively extracting signals in a specific frequency band by setting a frequency passband range; dynamic truncation refers to the adaptive process of adjusting the filtering bandwidth according to the real-time energy distribution; the target material density parameter refers to the quantitative index derived by relating the reflected energy intensity to the physical density of the material; and the filtered reflection mode frequency band response refers to the energy distribution matrix that retains the characteristic frequency bands of the target material and carries density information.

[0099] In this embodiment, a bandpass filter bank with adaptive bandwidth adjustment capability is first constructed based on the generated frequency band boundary markers. By monitoring the energy distribution gradient within the target frequency band in real time, the cutoff frequency range of the filter is dynamically narrowed or expanded to ensure complete capture of the material's reflection characteristics. Spatial integration is performed on the captured frequency band energy, and the energy intensity is converted into a density parameter by combining the density mapping function in the material reflection characteristic library. Finally, a frequency band response matrix containing the spatial density distribution is generated, where the amplitude represents the distribution density of the target material, and the phase retains the original spectral characteristics.

[0100] Step 203: Perform a dot product operation between the filtered reflection mode frequency band response and the frequency domain data corresponding to the spatial distribution feature extraction region to obtain the modulated spatial feature frequency domain distribution.

[0101] In this step, dot multiplication refers to the mathematical operation of multiplying corresponding elements of two frequency domain matrices; modulated spatial feature frequency domain distribution refers to enhancing the feature expression of the target material's relevant frequency components through frequency domain operations; and target material density parameter refers to the quantized feature parameter generated through the mapping relationship between reflected energy and physical density.

[0102] In this embodiment, a Fast Fourier Transform (FFT) is first performed on the visible light spatial distribution feature extraction region to generate a frequency domain matrix containing spatial texture details. Then, the obtained filtered reflection mode frequency band response matrix is ​​multiplied element-wise with this matrix, enhancing the amplitude of frequency components related to the target material in the visible light frequency domain while suppressing irrelevant components. Finally, the modulated frequency domain amplitude is nonlinearly scaled using the target material density parameter, preserving the original phase information to generate a frequency domain distribution that simultaneously contains optical texture and material density features.

[0103] Step 204: Normalize the frequency domain distribution of the modulated spatial features based on the target material density parameter to eliminate the energy scale differences between different modes;

[0104] In this step, the target material density parameter refers to the quantitative characteristic value generated by mapping the reflected energy intensity to the physical density, which is used to characterize the material's distribution density in space; normalization refers to the process of scaling the energy intensity of different modal data to achieve comparable dimensions; energy scale difference refers to the difference in signal intensity magnitude between visible light and non-visible light modes due to different sensing principles.

[0105] In this embodiment, a normalized coefficient matrix is ​​first constructed based on the target material density parameter. This coefficient is generated by a preset density-energy conversion function in the material reflectivity library. The modulated spatial characteristic frequency domain distribution matrix is ​​then multiplied element-wise with the normalized coefficients, compressing the high-energy region of the non-visible light mode proportionally to the density parameter and moderately enhancing the low-density region. Subsequently, a frequency domain energy histogram matching algorithm is used to adjust the high-frequency component energy distribution of the visible light mode, making it approximate the shape of the normalized non-visible light energy histogram, ultimately achieving scale alignment of cross-modal frequency domain energy.

[0106] Step 205: The normalized frequency domain distribution is superimposed with the original spatial features of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier.

[0107] The normalized frequency domain distribution refers to the frequency domain feature matrix after energy scale alignment, which retains the reflectivity of the target material and eliminates cross-modal dimensional differences; the original spatial features refer to the texture, edge and other detailed information obtained from the spatial distribution feature extraction area in the visible light imaging data; the composite spatial features of the target material identifier refer to the multi-dimensional feature expression that integrates visible light spatial details and non-visible light material properties, and has the dual representation capability of physical properties and visual information.

[0108] First, a dynamic fusion weight matrix is ​​calculated based on the target material density parameter and the energy gradient distribution of the penetrating feature layer. The normalized frequency domain distribution is converted into a complex domain expression, separating the amplitude and phase components. Two-dimensional wavelet decomposition is performed on the original spatial features to extract low-frequency contour and high-frequency detail components. The frequency domain amplitude component is nonlinearly scaled according to the dynamic fusion weights, while preserving the phase information of the original spatial features. The scaled amplitude and high-frequency detail components are spatially fused using an orthogonal projection algorithm, and then residually connected with the low-frequency contour component. Finally, local energy equalization compensation is performed on the fusion result based on the energy gradient of the penetrating feature layer to eliminate edge distortion caused by modal coverage differences, generating a composite feature that simultaneously contains material identification and spatial details.

[0109] To address the problem that traditional frequency domain weighted superposition methods in the prior art struggle to simultaneously preserve phase information and maintain energy balance, leading to distortion or information loss in the fused features at material boundaries, some embodiments, according to step 205, involve weighted superposition of the normalized frequency domain distribution with the original spatial features of the extracted spatial distribution feature region to generate a composite spatial feature containing the target material identifier, including:

[0110] Step 301: Calculate the dynamic fusion weight based on the target material density parameter and the energy distribution of the corresponding region in the penetrating feature layer;

[0111] In this step, the target material density parameter refers to a quantitative index generated by mapping the reflection energy intensity to the physical density, used to characterize the material's spatial distribution concentration; the penetrability feature layer energy distribution refers to the spatial gradient change of low-frequency energy in the non-visible light spectrum data that characterizes the electromagnetic wave penetration capability; and the dynamic fusion weight refers to the fusion ratio coefficient of visible and non-visible light data dynamically adjusted according to the material density and penetrability energy, used to optimize the complementarity of multimodal information.

[0112] In this embodiment, spatial gradient calculation is first performed on the penetrating feature layer to generate a gradient matrix characterizing the direction and intensity of energy attenuation. The target material density parameter matrix is ​​then element-wise multiplied with the gradient matrix (Hadamard product operation) to obtain a preliminary weight distribution. The preliminary weights are nonlinearly mapped using the Sigmoid function, and combined with spatial continuity constraints (such as smoothing based on bilateral filtering) to eliminate the influence of isolated noise points. Finally, a dynamically fused weight matrix is ​​generated, whose numerical distribution satisfies the following: higher weights are assigned to non-visible light modes in regions with high material density and significant penetrating energy attenuation, while visible light contributions are enhanced in low-density and energy-stable regions.

[0113] Step 302: Convert the normalized frequency domain distribution into a complex domain expression and decompose the original spatial features into real and imaginary components;

[0114] In this step, the normalized frequency domain distribution refers to the frequency domain feature matrix after energy scale alignment. Its amplitude represents the energy intensity of the target material, and its phase retains spatial structure information. The complex domain expression refers to the mathematical representation of decomposing the frequency domain data into amplitude and phase components. The real component corresponds to the frequency domain cosine coefficient, which represents the low-frequency profile of the space. The imaginary component corresponds to the frequency domain sine coefficient, which represents the high-frequency details of the space.

[0115] In this embodiment, the normalized frequency domain distribution is first reconstructed using complex number conversion: the frequency domain amplitude is used as the modulus, and the angle parameters are initialized based on the phase information of the original visible light spatial features to generate a complex matrix containing material energy intensity and spatial phase features. Then, a Hilbert transform is performed on the original visible light spatial features, decomposing them into a real part (low-frequency components reflecting the overall contour) and an imaginary part (high-frequency detail components containing edge textures). An orthogonal projection algorithm couples the amplitude and imaginary components of the complex matrix, retaining the real components as the basic spatial structure.

[0116] Step 303: The magnitude of the complex domain representation is nonlinearly scaled according to the dynamic fusion weights to obtain the scaled magnitude distribution, while retaining the phase information of the original spatial features;

[0117] In this step, dynamic fusion weight refers to the modal contribution coefficient that is dynamically adjusted based on material density and penetrating energy; the amplitude in complex domain form refers to the modulus component in the frequency domain feature matrix that characterizes energy intensity; phase information refers to the angular component in the frequency domain data that reflects the spatial structure relationship; nonlinear scaling refers to non-uniform intensity adjustment of the amplitude according to the weight coefficient, while keeping the phase unchanged to maintain the integrity of the spatial structure.

[0118] In this embodiment, the dynamically fused weight matrix is ​​first input into the adaptive gain controller to generate scaling coefficients that are positively correlated with the amplitude intensity. An exponential function is used to perform a nonlinear mapping on the complex domain amplitude: the amplitude is increased exponentially in the weight-dominant region and linearly scaled in the low-weight region. The original phase angle is strictly preserved during scaling, and the adjusted amplitude and original phase are recombine into a complex form through polar coordinate transformation. Finally, DC component correction is performed on the scaled frequency domain distribution to eliminate the baseline offset introduced by the nonlinear operation.

[0119] Step 304: Perform orthogonal projection operation between the scaled amplitude distribution and the imaginary component of the original spatial features to generate initial fused frequency domain data;

[0120] In this step, the scaled amplitude distribution refers to the frequency domain energy intensity matrix after dynamic weight nonlinear adjustment; the imaginary component of the original spatial features refers to the sinusoidal component representing high-frequency details after the visible light image is decomposed in the frequency domain; the orthogonal projection operation refers to the mathematical operation of superimposing the components of two vector spaces in the vertical direction, which is used to preserve the principal component features and eliminate redundant information.

[0121] In this embodiment, the scaled amplitude distribution is first converted into vector form to construct a feature vector space based on energy intensity. Gram-Schmidt orthogonalization is then performed on the imaginary components of the original spatial features to generate an orthogonal basis vector set. The scaled amplitude vectors are then projected onto the orthogonal basis space of the imaginary components using a projection matrix, preserving amplitude components that are linearly independent of the imaginary components. Finally, the linear combination of the projection residuals and the orthogonal basis components is reconstructed into frequency domain data, forming an initial fused frequency domain matrix that retains high-frequency details while enhancing key features.

[0122] Step 305: Based on the energy gradient changes in adjacent regions in the penetrating feature layer, perform energy equalization compensation on the initial fused frequency domain data to eliminate local distortions caused by differences in multimodal data coverage;

[0123] In this step, the energy gradient change in adjacent regions of the penetrability feature layer refers to the spatial distribution difference of penetrability energy in the non-visible light spectrum data, which is obtained by calculating the energy intensity difference between adjacent pixels; energy balance compensation refers to adjusting the energy distribution of the fused data according to the gradient change to eliminate brightness abrupt changes or texture distortion caused by different modal coverage ranges; local distortion refers to image artifacts caused by data registration errors or energy distribution imbalances during multimodal fusion.

[0124] In this embodiment, the penetrating feature layer is first subjected to Gaussian filtering to generate a smoothed energy distribution map. The Sobel gradient of this energy map is calculated to identify regions of abrupt energy changes. A compensation coefficient matrix is ​​constructed based on the gradient magnitude, where high gradient regions correspond to low compensation coefficients to suppress energy overshoot, and low gradient regions are given high compensation coefficients to enhance details. The compensation coefficient matrix is ​​then multiplied element-wise with the initial fused frequency domain data to achieve adaptive adjustment of the frequency domain energy. Finally, a deconvolution operation is used to restore the edge sharpness lost due to the compensation operation, generating equalized fused frequency domain data.

[0125] Step 306: Perform complex number reconstruction on the compensated frequency domain data and the real component of the original spatial feature to generate a composite spatial feature of the target material identifier;

[0126] In this step, the compensated frequency domain data refers to the optimized frequency domain matrix after energy equalization, which eliminates multimodal fusion artifacts and retains key features; the real component of the original spatial features refers to the cosine coefficient component of the visible light image after frequency domain decomposition, which represents the low-frequency contour information of the space; complex reconstruction refers to the operation of recombinating the real and imaginary components of the frequency domain into a complex form and performing an inverse transformation; the composite spatial features of the target material identifier refer to the final output features that fuse visible light spatial details and non-visible light material properties, possessing a dual expression of physical properties and visual information.

[0127] In this embodiment, the compensated frequency domain data (including the optimized imaginary component) is first complexly combined with the real component of the original spatial features to generate a complex matrix containing complete frequency domain information. The complex matrix is ​​then transformed to the spatial domain using an inverse Fourier transform to obtain a preliminary fused image. Subsequently, based on the material density distribution of the penetrating feature layer, a spatial domain masking operation is performed on the fused image: a semi-transparent color marker is superimposed on high-density material areas while preserving the natural transition of low-frequency contours. The final generated composite spatial features, while retaining visible light texture details, enhance the visual recognizability of the target material through hue and saturation adjustments.

[0128] Since pixel loss due to dynamic blur or occlusion often relies on visible light data interpolation for compensation in traditional methods, the trajectory is prone to breakage due to lack of penetration information. Therefore, in some embodiments, as described in step 103, the missing pixel information in the dynamically changing trajectory capture area is correlated and mapped with the energy distribution of the penetration feature layer to generate enhanced dynamic trajectory data, including:

[0129] Step 401: Detect pixel missing areas caused by dynamic blur in the dynamic change trajectory capture area, and mark the pixel missing areas as areas to be compensated;

[0130] In this step, the dynamic trajectory capture area refers to the spatial range in a video sequence or image frame used to track the target's motion trajectory; dynamic blur refers to the image ghosting phenomenon caused by the target's rapid movement or the camera's excessive exposure time; pixel missing area refers to the part of the dynamic blur area where details are lost due to the discontinuity of the motion trajectory; and the area to be compensated refers to the set of spatial coordinates that are marked and need to be compensated for across modal information.

[0131] In this embodiment, the motion vector field between adjacent frames is first calculated using optical flow to identify trajectory regions where the motion speed exceeds a preset threshold (e.g., 15 pixels per second). Laplacian edge detection is then performed on these high-speed regions, locating edge breakage areas caused by blurring by comparing the gradient difference between the static background template and the current frame. Next, an adaptive threshold segmentation algorithm is used to mark pixel clusters with gradient differences below the dynamic blur determination threshold as regions to be compensated, generating a binary mask label matrix. Finally, morphological closing operations are used to fill holes in the mask, ensuring the spatial continuity of the regions to be compensated.

[0132] Step 402: Extract the energy distribution pattern in the penetrability feature layer that corresponds to the spatial location of the region to be compensated;

[0133] In this step, the penetrability feature layer refers to the low-frequency energy distribution layer in the non-visible light spectrum data that characterizes the penetration capability of electromagnetic waves; spatial location correspondence refers to the coordinate mapping relationship established through cross-modal spatial registration; and energy distribution mode refers to the multi-dimensional feature set composed of the intensity, gradient direction, and frequency band of penetrability energy within a specific spatial region.

[0134] In this embodiment, firstly, based on the mask matrix of the region to be compensated generated in step 401, a spatial coordinate transformation is performed on the penetrating feature layer: through pre-calibrated affine transformation parameters, the region to be compensated in the visible light image coordinate system is mapped to the polar coordinate system of the non-visible light data. A bilinear interpolation algorithm is used to resample the penetrating feature layer, extracting the energy distribution of the sub-region that precisely corresponds to the region to be compensated. Then, principal component analysis (PCA) is used to extract the first three principal components of this sub-region, constructing a three-dimensional feature vector containing energy intensity, gradient direction, and frequency band correlation, forming an energy distribution pattern descriptor.

[0135] Step 403: Based on the energy distribution pattern, perform energy attenuation directionality analysis on the peripheral adjacent areas of the area to be compensated, and generate an energy attenuation gradient matrix by calculating the energy change rate of each pixel in the penetrating feature layer along the motion trajectory direction.

[0136] In this step, the energy attenuation directionality analysis refers to quantifying the energy attenuation law along the motion trajectory direction based on the physical characteristics of the penetrating feature layer; the motion trajectory direction refers to the main direction of target motion calculated by optical flow method or motion vector field; the energy attenuation gradient matrix refers to the matrix composed of the first derivative of each pixel along the motion direction, and its value reflects the rate of energy change per unit displacement.

[0137] In this embodiment, the main motion direction of the target object is first determined based on the energy distribution features extracted in step 402. Gradient analysis is performed on the penetrability feature layer around the peripheral pixels of the area to be compensated: by calculating the energy change intensity of each pixel in the horizontal and vertical directions, the horizontal gradient value and the vertical gradient value are obtained respectively. Then, the gradient values ​​in these two directions are synthesized according to the proportion of the main motion angle. For example, in the main motion direction, the horizontal gradient is superimposed with a high weight, and the vertical gradient is merged according to a corresponding proportion, generating a comprehensive gradient value along the motion trajectory. To eliminate dimensional bias caused by differences in the reflective properties of different materials, the gradient values ​​are normalized to make the data in each region comparable. Next, with each pixel as the center, the energy difference between it and its adjacent pixels is calculated along the motion direction, and converted into the energy change rate per unit distance based on the actual physical distance—if the energy of the subsequent pixel is lower than that of the current pixel, it is marked as a negative value (energy attenuation), and vice versa (energy enhancement). Finally, the energy change rate of all pixels is integrated into a matrix, in which continuous negative value regions accurately map the energy attenuation band of the motion path, while positive value regions show energy enhancement characteristics, providing a quantitative basis driven by physical characteristics for subsequent compensation.

[0138] Step 404: Based on the distribution characteristics of the energy attenuation gradient matrix and the local energy peak of the energy distribution pattern, generate dynamic trajectory compensation parameters, which include compensation intensity coefficient and direction correction factor.

[0139] In this step, the local energy peak refers to the extreme point in the energy distribution pattern where the intensity is significantly higher than that of the surrounding area; the dynamic trajectory compensation parameters include the compensation intensity coefficient (which controls the intensity level of pixel compensation) and the orientation correction factor (which adjusts the offset angle of the compensation path), which are used to guide the repair of cross-modal missing information.

[0140] In this embodiment, a local maximum detection algorithm is first used in the energy distribution pattern to identify energy peak points as anchor points for trajectory compensation. Based on the distribution characteristics of the energy decay gradient matrix, the angle difference between the gradient descent direction and the original motion direction is calculated with the peak point as the center, generating a direction correction factor. Simultaneously, a compensation intensity coefficient is calculated based on the product of the gradient decay amplitude and the peak intensity; regions with more severe decay and higher peak values ​​are assigned higher compensation intensities. Finally, a Gaussian mixture model is used to spatially smooth the compensation parameters, eliminating the influence of isolated noise points and generating a dynamic compensation parameter set that matches the physical characteristics of the motion trajectory.

[0141] Step 405: Perform multi-scale interpolation on the pixel missing part of the region to be compensated according to the compensation intensity coefficient, and constrain the interpolation path in combination with the direction correction factor to generate initial compensation trajectory data.

[0142] In this step, the compensation intensity coefficient refers to the interpolation weight parameter that is dynamically adjusted according to the material attenuation characteristics, controlling the intensity level of pixel compensation; the orientation correction factor refers to the path offset angle parameter generated based on the gradient orientation deviation, constraining the interpolation direction; multi-scale interpolation refers to adaptive compensation of missing regions at multiple resolution levels, taking into account both the overall trajectory continuity and the restoration of local details; the initial compensation trajectory data refers to the preliminary repair results generated after fusing cross-modal information.

[0143] In this embodiment, firstly, based on the compensation intensity coefficient and direction correction factor generated in step 404, the area to be compensated is decomposed into multiple resolution levels using a Gaussian pyramid. A large-scale interpolation template is constructed in the high-level, low-resolution image, and the interpolation window size is dynamically adjusted according to the compensation intensity coefficient—a large-scale bilinear interpolation is used to fill the main body of the motion trajectory in high-coefficient areas, while a small window is used to preserve details in low-coefficient areas. Simultaneously, the direction correction factor constrains the angle of the interpolation path, prioritizing the selection of pixels with the same correction direction among adjacent pixels in the area to be compensated for weighted calculation. Subsequently, in the low-level, high-resolution image, the interpolation results are optimized in detail by combining the energy distribution characteristics of the penetrating feature layer: the edge sharpness of the metal material is enhanced in areas of concentrated millimeter-wave energy, while the natural transition of visible light textures is maintained in areas of energy attenuation. Finally, the interpolation results from each level are superimposed level by level using a pyramid inverse fusion algorithm. The high-level results ensure the macroscopic continuity of the motion trajectory, while the low-level corrections enhance microscopic details such as blade teeth, generating initial compensated trajectory data that conforms to the physical characteristics of non-visible light while preserving spatial details of visible light.

[0144] Step 406: The initial compensation trajectory data is dynamically corrected using the energy attenuation gradient matrix. The corrected compensation trajectory data is then spatially aligned and energy-fused with the original trajectory data of the dynamically changing trajectory capture area to generate enhanced dynamic trajectory data containing cross-modal compensation information.

[0145] In this step, dynamic correction refers to iterative optimization of the compensation data based on physical attenuation characteristics; spatial alignment refers to coordinate system matching between the cross-modal compensation trajectory and the original visible light trajectory; energy fusion refers to the integration of information combining visible light texture and non-visible light physical characteristics; and enhanced dynamic trajectory data refers to the final output of a high-precision spatiotemporal continuous trajectory.

[0146] In this embodiment, the initial compensation trajectory data is first weighted and fused with the energy attenuation gradient matrix to enhance the compensation intensity in regions with significant gradient attenuation. The matching degree between the compensation trajectory and the gradient distribution is optimized using an iterative least squares method. Then, an affine transformation model is used to map the corrected compensation trajectory to the visible light dynamic trajectory capture region coordinate system, achieving sub-pixel-level spatial alignment based on feature point matching (such as SIFT keypoints). Finally, a wavelet domain energy fusion strategy is employed to preserve the macroscopic continuity of the visible light trajectory in the low-frequency band and inject material features of the non-visible light compensation trajectory in the high-frequency band, generating enhanced trajectory data with unified spatiotemporal dimensions and complete physical properties.

[0147] To address the issue that existing single-scale interpolation methods struggle to adapt to the multi-scale feature differences in occluded regions, leading to significant deviations between the compensation trajectory and the actual motion direction, some embodiments, as described in step 405, perform multi-scale interpolation on the pixel-missing portion of the region to be compensated based on the compensation intensity coefficient, and constrain the interpolation path using the direction correction factor to generate initial compensation trajectory data, including:

[0148] Step 501: Decompose the energy attenuation gradient matrix into a multi-scale structure according to a preset Gaussian pyramid hierarchy to generate a gradient distribution hierarchy set that matches the resolution of the region to be compensated. Each level in the gradient distribution hierarchy set corresponds to the energy attenuation change characteristics at different scales.

[0149] In this step, Gaussian pyramid hierarchy refers to the multi-resolution hierarchical structure of the image generated by progressive downsampling and Gaussian filtering; multi-scale decomposition refers to decomposing the original data into feature sets with different spatial resolutions to separate macroscopic trends from microscopic details; gradient distribution hierarchy set refers to the representation of the energy decay gradient matrix at different resolution levels, with higher levels representing large-scale energy decay trends and lower levels preserving local detail changes.

[0150] In this embodiment, a Gaussian pyramid is first constructed on the energy decay gradient matrix: by recursively applying Gaussian filtering and interval downsampling, a hierarchical sequence with progressively decreasing resolution is generated. Gradient features at the corresponding scale are preserved in each level—higher levels (lower resolution) highlight the overall decay trend of the motion trajectory, while lower levels (higher resolution) maintain fine gradient changes in the edge regions. Subsequently, an interpolation algorithm is used to align the gradient data of each level to the original resolution space of the region to be compensated, forming a gradient distribution hierarchy set whose scale characteristics match the compensation requirements.

[0151] Step 502: Based on the energy attenuation directionality of each level in the gradient distribution hierarchy set, construct a multi-scale interpolation template;

[0152] In this step, the multi-scale interpolation template refers to an adaptive compensation operator constructed based on the gradient direction characteristics of different resolution levels. Its shape and scale change with the level to match the repair needs of trajectory features of different granularities.

[0153] In this embodiment, firstly, a direction consistency analysis is performed on each level in the gradient distribution hierarchy set: the dominant gradient direction of each level is detected by Hough transform and used as the principal axis direction of the interpolation template. The template size is adjusted according to the level resolution—a large-scale elliptical template is used for higher levels (lower resolution), and a compact rhombic template is used for lower levels (higher resolution). The template weight distribution is dynamically adjusted by the gradient magnitude: high gradient decay regions are assigned high weights at the center and low weights at the edges; low gradient regions are assigned uniform weights. Finally, an interpolation template set matching the physical characteristics of each level is generated.

[0154] Step 503: Filter the main motion directions in the gradient distribution hierarchy set according to the direction correction factor to generate hierarchical constraint path parameters;

[0155] In this step, the orientation correction factor refers to the path angle adjustment parameter generated by energy attenuation orientation deviation calibration; the hierarchical constraint path parameter refers to the cross-level motion path descriptor generated by fusing multi-scale gradient orientation features and orientation correction factor, which is used to unify the consistency of compensation path orientation at different resolution levels.

[0156] In this embodiment, the main motion direction is first extracted for each level of the gradient distribution hierarchy set, and the detected main direction is calibrated using a direction correction factor. In high-level (low-resolution) data, valid directions are filtered based on the angle threshold of the direction correction factor, and noisy directions with large deviations are eliminated. In low-level (high-resolution) data, the correction factor is used to fine-tune the local direction to eliminate direction jitter at the detail level. Subsequently, the calibrated multi-scale direction features are fused according to the hierarchical weights to generate a hierarchical constraint path parameter set. The macroscopic hierarchical parameters define the overall direction of the motion trajectory, and the microscopic hierarchical parameters refine the precise angle of the local path, ultimately ensuring the direction consistency of the cross-scale compensation path.

[0157] Step 504: In each level, adaptive bilinear interpolation is performed on the pixel missing part of the region to be compensated using the weight distribution of the multi-scale interpolation template and the directional tolerance range of the hierarchical constraint path parameters to generate hierarchical compensation trajectory data.

[0158] In this step, hierarchical compensation trajectory data refers to the local repair results generated by multi-scale interpolation that match the characteristics of each resolution level, which preserves both the macroscopic motion trend and repairs the microscopic details; adaptive bilinear interpolation refers to the compensation method that dynamically adjusts the interpolation coefficients according to the template weights and orientation constraints to achieve physical characteristic alignment of cross-modal data.

[0159] In this embodiment, based on the weight distribution of the multi-scale interpolation template and the directional tolerance range of the hierarchical constraint path parameters, the area to be compensated is repaired in layers: In the high-level (low-resolution) layer, a large-scale bilinear interpolation is used to fill the trajectory breakage area along the main motion direction, and the template weight distribution enhances the compensation intensity in areas with significant gradient attenuation (such as the blade of a metal instrument); In the low-level (high-resolution) layer, the interpolation neighborhood is filtered in combination with the directional tolerance range, and only pixels consistent with the correction direction are selected for weighted calculation. At the same time, the energy distribution characteristics of the penetrating feature layer (such as millimeter wave reflection intensity) are injected to optimize the detail texture. Through layer-by-layer interpolation compensation, the high-level results ensure the continuity of the motion trajectory, while the low-level repair restores fine structures such as jagged edges, ultimately generating hierarchical compensation trajectory data that matches the physical attenuation characteristics and visual features.

[0160] Step 505: Perform cross-scale fusion of the compensation trajectory data at each level according to the Gaussian pyramid reconstruction rule to generate initial compensation trajectory data that is consistent with the multi-scale energy attenuation characteristics of the penetrating feature layer.

[0161] In this step, the Gaussian pyramid reconstruction rule refers to the method of multi-scale data fusion through stepwise upsampling and weighted superposition; the initial compensation trajectory data refers to the compensation output that fuses multi-level repair results and retains cross-scale physical characteristics, and needs to be coupled with the energy attenuation characteristics of the penetrating feature layer in the frequency domain and spatial domain.

[0162] In this embodiment, the compensation trajectory data generated in step 504 is upsampled level by level from low resolution to high resolution according to the pyramid hierarchy. After each upsampling, it is weighted and superimposed with the adjacent higher-level data: the higher-level data provides the macroscopic energy attenuation framework of the motion trajectory, while the lower-level data injects the local attenuation features of the detailed texture. During the superposition process, the fusion weights are dynamically adjusted based on the energy distribution ratio of each level of the penetration feature layer—the weight of the detail layer is increased in the high-frequency energy concentration area of ​​millimeter waves, and the contribution of the macroscopic layer is enhanced in the low-frequency energy area, so that the fusion result is completely matched with the multi-scale attenuation of the penetration feature layer in the frequency domain. The final generated initial compensation trajectory data has both the spatial continuity of the visible light trajectory and the physical attenuation characteristics of the non-visible light modes.

[0163] Because existing multimodal fusion models lack temporal synchronization constraints and physical characteristic-driven optimization mechanisms, they suffer from lag in contribution ratio adjustment or energy distribution imbalance in dynamic scenarios. To address this issue, in some embodiments, according to step 105, a multimodal joint optimization model is constructed based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, including:

[0164] Step 601: Phase-match the timestamp sequence of the enhanced dynamic trajectory data with the energy change period of the composite spatial feature to generate synchronization timing alignment parameters;

[0165] In this step, enhanced dynamic trajectory data refers to the spatiotemporally continuous motion trajectory after fusing cross-modal compensation, which includes the temporal information of the target motion; the energy change period of composite spatial features refers to the regular pattern of fluctuations in visible and non-visible light modes (such as material reflection and penetrating energy) over time; phase matching refers to aligning the periodic fluctuations of different modal data on the time axis by adjusting the temporal offset; and synchronous temporal alignment parameters refer to the key parameters for eliminating cross-modal temporal differences, including delay compensation amount, phase scaling factor, etc.

[0166] In this embodiment, the main period of energy change in the composite spatial features is first extracted, while the motion period corresponding to the timestamp sequence of the enhanced trajectory data is analyzed. The phase difference between the two is calculated using a dynamic time warping algorithm to generate time-series compensation parameters: delay compensation is applied to the trajectory data to align with the energy peak moment, and the motion period length is adjusted by a scaling factor to make it consistent with the energy fluctuation period of non-visible light. Finally, cross-correlation verification using a sliding window ensures that the cross-modal data is completely synchronized at key event moments such as the start and end points of the waving trajectory and the velocity peak.

[0167] Step 602: Based on the synchronization timing alignment parameters, perform sliding window mutual information calculation on the motion vector amplitude of consecutive frames in the enhanced dynamic trajectory data and the frequency domain energy peak distribution of the composite spatial features to generate modal contribution weight coefficients.

[0168] In this step, sliding window mutual information calculation refers to measuring the statistical correlation between the distributions of two types of data within a fixed time window, quantifying the degree of information sharing between modalities; modal contribution weight coefficient refers to the fusion weight dynamically allocated based on mutual information values, used to balance the contribution ratio of different modalities in feature representation.

[0169] In this embodiment, based on synchronization time alignment parameters, the motion vector amplitude of consecutive frames in enhanced dynamic trajectory data is spatiotemporally aligned with the frequency domain energy peak distribution of composite spatial features. A sliding window mechanism is employed to calculate the joint probability density of the motion vector amplitude histogram and the frequency domain energy distribution within the window, and the correlation between the two is quantified using a mutual information formula. A higher mutual information value indicates a greater contribution of non-visible light modes (such as millimeter-wave energy distribution) to the trajectory feature expression during that time period, and vice versa, visible light modes dominate. Finally, a modal contribution weighting coefficient that dynamically changes over time is generated to guide the weighted fusion of multimodal features.

[0170] Step 603: Extract the spatial distribution boundary of the frequency domain energy peak in the composite spatial features, and map the spatial distribution boundary to the motion vector coverage area corresponding to the enhanced dynamic trajectory data. Calculate the coverage area overlap rate between the frequency domain energy peak distribution boundary and the motion vector coverage area, and generate a spatial coupling difference matrix based on the coverage area overlap rate.

[0171] In this step, the spatial distribution boundary of the frequency domain energy peak refers to the outline of the non-visible light energy concentration area determined by threshold segmentation; the motion vector coverage area refers to the spatial projection range of motion vectors of consecutive frames in the enhanced dynamic trajectory data; the coverage area overlap rate refers to the proportion of the intersection area of ​​two types of regions to their union area, which is used to quantify spatial consistency; the spatial coupling difference matrix refers to a two-dimensional matrix generated based on the overlap rate that reflects the cross-modal spatial matching degree, and the lower the value, the better the spatial consistency.

[0172] In this embodiment, morphological closing operations are first performed on the frequency domain energy distribution of the composite spatial features to eliminate noise interference. Then, closed boundaries of the energy peak regions are extracted through adaptive threshold segmentation. The boundary coordinates are mapped to the motion vector space of the enhanced dynamic trajectory data. The area ratio of the intersection and union regions of these two regions is calculated using polygon Boolean operations to obtain the pixel-by-pixel coverage overlap rate. A sliding window statistical method is used to calculate the mean local overlap rate for each block of the full-frame image. This is combined with the motion vector amplitude to generate a spatial coupling difference matrix. Low-value regions in the matrix represent high-level matching of cross-modal spatial features, while high-value regions indicate spatial misalignment locations requiring further correction.

[0173] Step 604: Based on the gradient distribution direction of the spatial coupling difference matrix, multiply the modal contribution weight coefficient with the target material density parameter in the material reflection characteristic layer to generate a dynamic fusion constraint factor.

[0174] In this step, the dynamic fusion constraint factor refers to the optimization coefficient generated by combining spatial coupling difference, modal weight and material density, which is used to constrain the strength and direction of multimodal data fusion; the product operation refers to the weighted fusion of modal weight coefficient and material density parameter according to the gradient direction of spatial coupling difference, so as to realize the dynamic adaptation of physical properties and motion characteristics.

[0175] In this embodiment, the gradient direction is first extracted from the spatial coupling difference matrix to identify the main direction of cross-modal spatial misalignment. The modal contribution weight coefficients and the target material density parameters are then directionally multiplied according to the gradient direction: in regions where the gradient direction coincides with the main motion direction, the product operation enhances the fusion weights; in regions where the gradient deviates, the product operation suppresses the weights. The product result is spatially smoothed using a non-uniform sampling algorithm to eliminate local abrupt noise and generate a dynamic fusion constraint factor matrix that strictly matches the physical characteristics of the motion trajectory.

[0176] Step 605: Based on the dynamic fusion constraint factor, construct the trajectory energy conservation equation in the motion vector space of the enhanced dynamic trajectory data, and construct the material reflection energy transfer equation in the frequency domain space of the composite spatial features. By alternately iterating to solve the minimum cross-entropy of the trajectory energy conservation equation and the material reflection energy transfer equation, a joint optimization objective function is generated.

[0177] In this step, the trajectory energy conservation equation refers to the mathematical constraint that ensures the continuity of kinetic and potential energy conversion within the motion vector space; the material reflection energy transfer equation refers to the physical model that describes the interaction between material properties and frequency domain reflection energy; the cross-entropy minimization refers to achieving the optimization objective by minimizing the difference in probability distributions between the two types of equations; and the joint optimization objective function refers to the unified optimization framework that integrates kinematic and physical property constraints.

[0178] In this embodiment, based on a dynamic fusion constraint factor, a trajectory energy conservation equation is constructed in the motion vector space of the enhanced trajectory data: through Lagrange mechanics modeling, the rate of change of motion vector amplitude is correlated with energy attenuation to ensure trajectory continuity. Simultaneously, a material reflection energy transfer equation is constructed in the frequency domain space of the composite spatial features: based on Fresnel's law of reflection, a mapping relationship is established between the material density parameter and the frequency domain energy reflection coefficient. The cross-entropy of the two equations is iteratively solved using the Alternating Direction Multiplier Method (ADMM), i.e., minimizing the difference in energy distribution probability between the two equations, ultimately generating a joint optimization objective function that simultaneously satisfies motion continuity and physical reflection characteristics.

[0179] Step 606: The energy attenuation gradient matrix of the penetrating feature layer is embedded as a regularization term into the joint optimization objective function. By constraining the frequency band energy allocation ratio of the visible light imaging data and the non-visible light spectral data, a multimodal joint optimization model is constructed to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectral data.

[0180] In this step, the regularization term refers to the constraint introduced to prevent the model from overfitting; the frequency band energy allocation ratio refers to the contribution weight of visible light and non-visible light data in different frequency bands; and the multimodal joint optimization model refers to the fusion framework that balances the contributions of the two types of data through mathematical constraints.

[0181] In this embodiment, the energy attenuation gradient matrix of the penetrating feature layer is embedded as an L2 regularization term into the joint optimization objective function to constrain the energy distribution ratio of visible and non-visible light data in the frequency domain: increasing the weight of non-visible light in the millimeter-wave high-frequency attenuation region and enhancing the contribution of visible light in the low-frequency stable region. The constrained optimization problem is solved by the projected gradient descent method, dynamically adjusting the fusion ratio of the two types of data to generate a robust multimodal joint optimization model that can adapt to scene changes.

[0182] Furthermore, to address the problem of fragmentation in frequency band energy allocation and spatial consistency correction in traditional fusion methods, which leads to artifacts or material identification distortion in the fused image, some embodiments, according to step 105, include:

[0183] Step 701: Based on the dynamic fusion constraint factor and the energy attenuation gradient matrix of the penetrating feature layer, calculate the frequency band energy allocation ratio parameter of visible light imaging data and non-visible light spectral data;

[0184] In this step, the frequency band energy allocation ratio parameter refers to the contribution ratio of the generated visible and non-visible light data in different frequency bands based on the dynamic fusion constraint factor and the energy attenuation characteristics of the penetrating feature layer, and is used to quantify the fusion weight distribution of multimodal data.

[0185] In this embodiment, the dynamic fusion constraint factor matrix and the energy attenuation gradient matrix of the penetrating feature layer are first normalized, and an initial ratio parameter is generated using a weighted average algorithm. For high-energy attenuation regions, the allocation ratio of non-visible light spectral data is increased by combining the intensity value of the dynamic fusion constraint factor; in low-attenuation regions, the proportion of non-visible light is reduced and the weight of visible light is increased based on the constraint factor. Finally, through spatial adaptive threshold segmentation, a frequency band energy allocation ratio parameter matrix matching the material density and motion characteristics is generated.

[0186] Step 702: Based on the frequency band energy allocation ratio parameters, generate a visible light frequency band energy allocation weight map and a non-visible light penetrability frequency band energy allocation weight map;

[0187] In this step, visible light frequency band energy allocation weight mapping and non-visible light penetration frequency band energy allocation weight mapping refer to the spatial weight distribution map generated according to the frequency band energy allocation ratio parameter, which respectively controls the fusion intensity of visible light and non-visible light data in different regions. The former enhances the preservation of visual details, while the latter enhances the expression of physical properties.

[0188] In this embodiment, based on the frequency band energy allocation ratio parameter, independent weight mappings are generated for visible light and non-visible light data respectively: the visible light ratio parameter is converted into a weight value, and a spatial continuous distribution map is generated by bilinear interpolation. The high ratio region is assigned a weight close to the full value, while the weight in the low ratio region approaches zero. At the same time, the non-visible light ratio parameter is subjected to a Hadamard product operation with the energy attenuation gradient of the penetrating feature layer to strengthen the weight contribution of the high attenuation region, and Gaussian filtering is used to eliminate abrupt noise and generate a smooth weight distribution.

[0189] Step 703: Convert the visible light imaging data to the frequency domain to generate visible light frequency domain data. Dynamically extract the visible light frequency domain data according to the visible light frequency band energy allocation weight mapping, retain the frequency band components that match the frequency domain energy peak of the composite spatial features, and generate weighted visible light frequency domain data.

[0190] In this step, visible light frequency domain data refers to the feature representation of visible light imaging data converted to the frequency domain through Fourier transform; dynamic truncation refers to selectively retaining specific frequency band components according to weight mapping; weighted visible light frequency domain data refers to the optimized frequency domain data generated after retaining the components that match the energy peak of the composite spatial feature frequency domain.

[0191] In this embodiment, a Fast Fourier Transform (FFT) is first performed on the visible light imaging data to generate frequency domain data containing amplitude and phase information. Based on the visible light frequency band energy allocation weight mapping, the components of each frequency band are dynamically truncated in the frequency domain space: the full amplitude is retained in high-weight frequency bands (such as those corresponding to high-frequency textures), while the amplitude of low-weight frequency bands (such as those corresponding to low-frequency noise) is attenuated according to the mapping ratio. Through frequency domain bandpass filtering technology, only the components that match the frequency domain energy peak distribution of the composite spatial features are retained, ultimately generating weighted visible light frequency domain data that retains key features while suppressing noise.

[0192] Step 704: The non-visible light penetration frequency band energy allocation weight mapping dynamically weights the energy of each frequency band after decomposition to generate penetration compensation frequency domain data.

[0193] In this step, the penetration compensation frequency domain data refers to the optimized data generated by dynamically weighting the energy of each decomposed frequency band through the non-visible light frequency band energy allocation weight mapping, which enhances the contribution of frequency bands with significant physical characteristics and suppresses interference frequency bands.

[0194] In this embodiment, non-visible light spectral data (such as millimeter-wave penetration characteristics) is decomposed into frequency bands to extract the energy distribution of each frequency band. Based on the energy allocation weight mapping of the penetration frequency bands, high weight coefficients are assigned to high-frequency attenuation regions to enhance their energy amplitude; while the weight coefficients are reduced for low-frequency penetration regions. The weighted frequency band energy is reconstructed into penetration-compensated frequency domain data through inverse frequency domain transformation, and its high-frequency components are precisely matched with the high-frequency edges of the visible light weighted frequency domain data in spatial dimension.

[0195] Step 705: The weighted visible light frequency domain data and the penetration compensation frequency domain data are superimposed to obtain a fused frequency domain distribution;

[0196] In this step, the fusion frequency domain distribution refers to the joint frequency domain expression generated by superimposing weighted visible light frequency domain data and penetration compensation frequency domain data through frequency bands. It integrates visible light texture details and non-visible light physical properties to form cross-modal complementary frequency domain features.

[0197] In this embodiment, frequency band superposition is performed on weighted visible light frequency domain data (preserving high-frequency edges and low-frequency contours) and penetration compensation frequency domain data (enhancing high-frequency reflection and low-frequency penetration characteristics): in the low-frequency band, a weighted average fusion dominated by visible light data is used to ensure spatial continuity; in the high-frequency band, a maximum value fusion prioritizing non-visible light data is implemented to enhance the sharpness of material edges. Through smooth transition processing of the frequency band overlap region, frequency domain jump noise is eliminated, generating a fused frequency domain distribution that combines visual fidelity and physical properties.

[0198] Step 706: Perform inverse transformation on the fused frequency domain distribution to generate an initial fused spatial signal, and perform spatial consistency correction on the initial fused spatial signal based on the motion vector coverage area of ​​the enhanced dynamic trajectory data;

[0199] In this step, the initial fused spatial signal refers to the preliminary spatial domain fusion result generated by performing an inverse Fourier transform on the fused frequency domain distribution; spatial consistency correction refers to adjusting the spatial coordinates of the fused signal through affine transformation based on the motion vector coverage area of ​​the enhanced dynamic trajectory data to eliminate cross-modal registration errors.

[0200] In this embodiment, an inverse Fourier transform is first performed on the fused frequency domain distribution to generate an initial fused spatial signal. Based on the motion vector coverage area of ​​the enhanced dynamic trajectory data, key feature points are extracted, and their spatial offset from the original visible light imaging data is calculated. A non-rigid transformation model is constructed using thin-plate spline interpolation (TPS) to perform local deformation correction on the initial fused signal, ensuring spatial consistency in key areas while preserving natural deformation.

[0201] Step 707: The corrected fused spatial signal is multiplied and modulated with the material identification information of the composite spatial features to generate a fused image signal containing multimodal physical property associations.

[0202] In this step, the corrected fused spatial signal refers to the cross-modal image data that has undergone frequency domain fusion and spatial consistency correction; the material identification information refers to the label matrix representing the target material properties in the composite spatial features (such as the physical feature encoding of high-density material regions and low-frequency penetration regions); dot product modulation refers to the mathematical operation of multiplying the two types of data pixel by pixel to enhance specific material features; the fused image signal with multimodal physical properties refers to the final output image that simultaneously retains visible light texture details and non-visible light material properties, possessing physical interpretability for cross-modal perception.

[0203] In this embodiment, the corrected fused spatial signal and material identification information are first multiplied pixel by pixel: in high-density material regions, millimeter-wave reflection features are amplified using high weighting coefficients, while in low-frequency penetration regions, non-visible light noise is suppressed using low weighting coefficients. Next, an adaptive gamma correction algorithm is used to optimize the dynamic range of the modulated image, increasing contrast in high-density regions to enhance edge sharpness, and maintaining a natural grayscale distribution in low-frequency penetration regions. Finally, spatial filtering eliminates high-frequency noise introduced during modulation, generating a fused image signal that retains visible light visual realism while enhancing non-visible light material characteristics.

[0204] Figure 2 This application provides a schematic diagram of the structure of an intelligent image signal processing system based on multimodal fusion, as shown in the embodiment. Figure 2 As shown, the system includes:

[0205] The receiving module 21 is used to receive multimodal input signals, wherein the multimodal input signals include visible light imaging data and non-visible light spectral data;

[0206] The decomposition module 22 is used to divide the visible light imaging data into a spatial distribution feature extraction region and a dynamic change trajectory capture region, and at the same time decompose the non-visible light spectral data into a penetration feature layer and a material reflection characteristic layer.

[0207] Mapping module 23 is used to associate and map the missing pixel information in the dynamic trajectory capture area with the energy distribution of the penetrating feature layer to generate enhanced dynamic trajectory data;

[0208] The identification module 24 is used to identify the reflection mode that matches a specific material in the target scene in the material reflection characteristic layer, and to superimpose the frequency band response corresponding to the reflection mode with the spatial distribution feature extraction area in the frequency domain to generate a composite spatial feature.

[0209] The construction module 25 is used to construct a multimodal joint optimization model based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, and to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectral data using the multimodal joint optimization model to generate a fused image signal.

[0210] Figure 2 The aforementioned intelligent image signal processing system based on multimodal fusion can perform... Figure 1The implementation principle and technical effects of the intelligent image signal processing method based on multimodal fusion described in the illustrated embodiment will not be repeated here. The specific methods by which each module and unit performs operations in the intelligent image signal processing system based on multimodal fusion described in the above embodiments have been described in detail in the embodiments related to this method, and will not be elaborated upon here.

[0211] In one possible design, Figure 2 The intelligent image signal processing system based on multimodal fusion shown in the embodiment can be implemented as a computing device, such as... Figure 3 As shown, the computing device may include a storage component 31 and a processing component 32;

[0212] The storage component 31 stores one or more computer instructions, wherein the one or more computer instructions are invoked and executed by the processing component 32.

[0213] The processing component 32 is used for the above Figure 1 The embodiment describes an intelligent image signal processing method based on multimodal fusion.

[0214] The processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above-described method. Alternatively, the processing component may be implemented as one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the above-described method.

[0215] Storage component 31 is configured to store various types of data to support operations at the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.

[0216] Of course, computing devices may also include other components, such as input / output interfaces, display components, communication components, etc.

[0217] Input / output interfaces provide interfaces between processing components and peripheral interface modules, which can be output devices, input devices, etc.

[0218] The communication components are configured to facilitate wired or wireless communication between computing devices and other devices.

[0219] The computing device can be a physical device or an elastic computing host provided by a cloud computing platform. In this case, the computing device can refer to a cloud server, and the aforementioned processing components, storage components, etc., can be basic server resources rented or purchased from the cloud computing platform.

[0220] This application also provides a computer storage medium storing a computer program, which, when executed by a computer, can perform the above-described functions. Figure 1 The embodiment shown is an intelligent image signal processing method based on multimodal fusion.

[0221] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0222] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0223] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0224] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A smart image signal processing method based on multimodal fusion, characterized in that, include: Receive multimodal input signals, wherein the multimodal input signals include visible light imaging data and non-visible light spectral data; The visible light imaging data is divided into a spatial distribution feature extraction region and a dynamic change trajectory capture region, while the non-visible light spectral data is decomposed into a penetration feature layer and a material reflection characteristic layer. The missing pixel information in the dynamically changing trajectory capture area is correlated and mapped with the energy distribution of the penetrating feature layer to generate enhanced dynamic trajectory data; The reflection pattern matching a specific material in the target scene is identified in the material reflection characteristic layer, and the frequency band response corresponding to the reflection pattern is superimposed with the spatial distribution feature extraction area in the frequency domain to generate a composite spatial feature; Based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, a multimodal joint optimization model is constructed. The contribution ratio of the visible light imaging data and the non-visible light spectral data is dynamically adjusted using the multimodal joint optimization model to generate a fused image signal.

2. The method according to claim 1, characterized in that, The reflection patterns matching specific materials in the target scene are identified in the material reflection characteristic layer, and the frequency band response corresponding to the reflection pattern is superimposed with the spatial distribution feature extraction area in the frequency domain to generate composite spatial features, including: Based on a preset material reflection characteristic library, the material reflection characteristic layer is segmented by frequency band, and the frequency band range of the reflection mode corresponding to a specific material in the target scene is extracted. The energy distribution within the frequency band of the reflection mode is dynamically extracted by bandpass filtering, and the filtered reflection mode frequency band response containing the target material density parameter is generated step by step based on the extracted energy. The frequency band response of the filtered reflection mode is multiplied by the frequency domain data corresponding to the spatial distribution feature extraction region to obtain the modulated spatial feature frequency domain distribution. The frequency domain distribution of the modulated spatial features is normalized based on the target material density parameter to eliminate the energy scale differences between different modes. The normalized frequency domain distribution is weighted and superimposed with the original spatial features of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier.

3. The method according to claim 2, characterized in that, The normalized frequency domain distribution is weighted and superimposed with the original spatial features of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier, including: Based on the target material density parameter and the energy distribution of the corresponding region in the penetrating feature layer, the dynamic fusion weight is calculated; The normalized frequency domain distribution is converted into a complex domain expression, and the original spatial features are decomposed into real and imaginary components. The magnitude of the complex domain representation is nonlinearly scaled according to the dynamic fusion weights to obtain the scaled magnitude distribution, while retaining the phase information of the original spatial features; The scaled amplitude distribution is orthogonally projected onto the imaginary component of the original spatial features to generate initial fused frequency domain data. Based on the energy gradient changes in adjacent regions of the penetrating feature layer, energy equalization compensation is performed on the initial fused frequency domain data to eliminate local distortions caused by differences in multimodal data coverage. The compensated frequency domain data is combined with the real components of the original spatial features to perform complex number reconstruction, generating a composite spatial feature of the target material identifier.

4. The method according to claim 1, characterized in that, The missing pixel information in the dynamically changing trajectory capture area is correlated and mapped with the energy distribution of the penetrating feature layer to generate enhanced dynamic trajectory data, including: Detect pixel-deficient areas caused by dynamic blur in the dynamic change trajectory capture area, and mark the pixel-deficient areas as areas to be compensated; Extract the energy distribution pattern in the penetrating feature layer that corresponds to the spatial location of the region to be compensated; Based on the energy distribution pattern, energy attenuation directionality analysis is performed on the peripheral adjacent areas of the area to be compensated. By calculating the energy change rate of each pixel in the penetrating feature layer along the motion trajectory, an energy attenuation gradient matrix is ​​generated. Based on the distribution characteristics of the energy attenuation gradient matrix and the local energy peak of the energy distribution pattern, dynamic trajectory compensation parameters are generated, which include a compensation intensity coefficient and a direction correction factor. Based on the compensation intensity coefficient, multi-scale interpolation is performed on the pixel missing portion of the region to be compensated, and the interpolation path is constrained by the direction correction factor to generate initial compensation trajectory data. The initial compensation trajectory data is dynamically corrected using the energy attenuation gradient matrix. The corrected compensation trajectory data is then spatially aligned and energy-fused with the original trajectory data of the dynamically changing trajectory capture area to generate enhanced dynamic trajectory data containing cross-modal compensation information.

5. The method according to claim 4, characterized in that, Multi-scale interpolation is performed on the pixel-missing portion of the region to be compensated based on the compensation intensity coefficient, and the interpolation path is constrained by the orientation correction factor to generate initial compensation trajectory data, including: The energy decay gradient matrix is ​​decomposed into multi-scale values ​​according to a preset Gaussian pyramid hierarchy to generate a gradient distribution hierarchy set that matches the resolution of the region to be compensated. Each level in the gradient distribution hierarchy set corresponds to the energy decay change characteristics at different scales. Based on the energy attenuation directionality of each level in the gradient distribution hierarchy set, a multi-scale interpolation template is constructed. Based on the direction correction factor, the main motion directions in the gradient distribution hierarchy set are filtered to generate hierarchical constraint path parameters; In each level, the weight distribution of the multi-scale interpolation template and the directional tolerance range of the hierarchical constraint path parameters are used to perform adaptive bilinear interpolation on the pixel missing part of the region to be compensated, thereby generating hierarchical compensation trajectory data. The compensation trajectory data of each level are fused across scales according to the Gaussian pyramid reconstruction rules to generate initial compensation trajectory data that is consistent with the multi-scale energy attenuation characteristics of the penetrating feature layer.

6. The method according to claim 1, characterized in that, Based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, a multimodal joint optimization model is constructed, including: The timestamp sequence of the enhanced dynamic trajectory data is phase-matched with the energy change period of the composite spatial feature to generate synchronization time alignment parameters; Based on the synchronization timing alignment parameters, a sliding window mutual information calculation is performed on the motion vector amplitude of consecutive frames in the enhanced dynamic trajectory data and the frequency domain energy peak distribution of the composite spatial features to generate modal contribution weight coefficients. Extract the spatial distribution boundary of the frequency domain energy peak in the composite spatial features, and map the spatial distribution boundary to the motion vector coverage area corresponding to the enhanced dynamic trajectory data. Calculate the coverage area overlap rate between the frequency domain energy peak distribution boundary and the motion vector coverage area, and generate a spatial coupling difference matrix based on the coverage area overlap rate. Based on the gradient distribution direction of the spatial coupling difference matrix, the modal contribution weight coefficient is multiplied with the target material density parameter in the material reflection characteristic layer to generate a dynamic fusion constraint factor. Based on the dynamic fusion constraint factor, a trajectory energy conservation equation is constructed in the motion vector space of the enhanced dynamic trajectory data, and a material reflection energy transfer equation is constructed in the frequency domain space of the composite spatial features. By iteratively solving the minimum cross-entropy of the trajectory energy conservation equation and the material reflection energy transfer equation, a joint optimization objective function is generated. The energy attenuation gradient matrix of the penetrating feature layer is embedded as a regularization term into the joint optimization objective function. By constraining the frequency band energy allocation ratio of the visible light imaging data and the non-visible light spectral data, a multimodal joint optimization model is constructed to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectral data.

7. The method according to claim 6, characterized in that, The multimodal joint optimization model is used to dynamically adjust the contribution ratio of visible light imaging data and non-visible light spectral data to generate a fused image signal, including: Based on the dynamic fusion constraint factor and the energy attenuation gradient matrix of the penetrating feature layer, calculate the frequency band energy allocation ratio parameters of visible light imaging data and non-visible light spectral data; Based on the frequency band energy allocation ratio parameters, a visible light frequency band energy allocation weight map and a non-visible light penetrability frequency band energy allocation weight map are generated. The visible light imaging data is converted to the frequency domain to generate visible light frequency domain data. The visible light frequency domain data is dynamically truncated according to the visible light frequency band energy allocation weight mapping, and the frequency band components that match the frequency domain energy peak of the composite spatial features are retained to generate weighted visible light frequency domain data. The penetrability feature layer of the non-visible light spectrum data is decomposed into frequency bands to obtain the energy of each frequency band after decomposition. The energy of each frequency band after decomposition is dynamically weighted according to the weight mapping of the non-visible light penetrability frequency band energy to generate penetration-compensated frequency domain data. The weighted visible light frequency domain data and the penetration compensation frequency domain data are superimposed in frequency bands to obtain a fused frequency domain distribution; The fused frequency domain distribution is subjected to inverse transformation to generate an initial fused spatial signal, and the initial fused spatial signal is spatially consistent based on the motion vector coverage area of ​​the enhanced dynamic trajectory data. The corrected fused spatial signal is modulated by dot product with the material identification information of the composite spatial features to generate a fused image signal that includes multimodal physical property associations.

8. An intelligent image signal processing system based on multimodal fusion, characterized in that, include: A receiving module is used to receive multimodal input signals, wherein the multimodal input signals include visible light imaging data and non-visible light spectral data; The decomposition module is used to divide the visible light imaging data into a spatial distribution feature extraction region and a dynamic change trajectory capture region, and at the same time decompose the non-visible light spectral data into a penetration feature layer and a material reflection characteristic layer. The mapping module is used to associate and map the missing pixel information in the dynamic trajectory capture area with the energy distribution of the penetrating feature layer to generate enhanced dynamic trajectory data. The identification module is used to identify the reflection pattern that matches a specific material in the target scene in the material reflection characteristic layer, and to superimpose the frequency band response corresponding to the reflection pattern with the spatial distribution feature extraction area in the frequency domain to generate composite spatial features; The construction module is used to construct a multimodal joint optimization model based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, and to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectral data using the multimodal joint optimization model to generate a fused image signal.

9. A computing device, characterized in that, It includes a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are invoked and executed by the processing component to implement the intelligent image signal processing method based on multimodal fusion as described in any one of claims 1 to 7.

10. A computer storage medium, characterized in that, The device contains a computer program that, when executed by a computer, implements an intelligent image signal processing method based on multimodal fusion as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Satellite-borne multi-band synthetic aperture radar signal processing optimization method and system

    CN119471684A

  • Highway abnormal event detection and alarm method based on deep learning

    CN119917970A