Intelligent image signal processing method and system based on multi-modal fusion

The method addresses dynamic scene adaptability and cross-modal data optimization by decomposing visible and non-visible light data into spatial and material features, enhancing image fusion precision and material tracking accuracy.

CN120318603AActive Publication Date: 2025-07-15BEIJING ZHAOKE HENGXING SCI & TECH CO LTD

Patent Information

Application Number
CN202510797128.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-15
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

In the multimodal image fusion, existing image processing methods have poor adaptability to dynamic scenes, insufficient coordinated optimization of physical characteristics across modal data, and lack of correlation between material reflection information and motion trajectory, resulting in low image fusion accuracy, loss of occlusion area information and significant motion artifacts.

Method used

By decomposing the visible light imaging data and non-visible light spectrum data into different functional layers, and based on the physical characteristic correlation between multimodal data, a multimodal joint optimization model is constructed, the data contribution ratio is dynamically adjusted, and the fused image signal is generated.

Benefits of technology

It significantly improves the ability to restore information in the occlusion area and the continuity of motion trajectory, and solves the problems of loss of details and poor dynamic adaptability caused by fixed weights or single-modal limitations in traditional methods. It is especially suitable for complex lighting, dynamic target tracking and material recognition scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318603A_ABST
    Figure CN120318603A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent image signal processing method and system based on multi-modal fusion. According to the invention, a multi-mode input signal is received, is divided into a spatial distribution feature extraction region and a dynamic change trajectory capture region, and is decomposed into a penetrability feature layer and a substance reflection feature layer; carrying out association mapping on missing pixel information in the dynamic change trajectory capture region and energy distribution of the penetrability feature layer to generate enhanced dynamic trajectory data, identifying a determined reflection mode in the substance reflection feature layer, and carrying out frequency domain superposition on corresponding frequency band response and the spatial distribution feature extraction region; generating a composite spatial feature, then constructing a multi-modal joint optimization model, then adjusting the contribution ratio of the two data, and generating a fused image signal; the technical scheme provided by the invention not only solves the problems of detail loss, artifact generation and poor dynamic adaptability caused by single-mode limitation, but also improves the spatial resolution and tracking precision of the image signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and particularly to an intelligent image signal processing method and system based on multimodal fusion. Background Art

[0002] With the rapid progress of social development, image signal processing technology has been widely applied in fields such as intelligent security, autonomous driving, medical imaging, and remote sensing monitoring.

[0003] Currently, common image processing methods mainly include the weighted average strategy based on pixel-level fusion, the frequency-domain decomposition technology based on multi-scale transformation (such as wavelet transform, Laplacian pyramid decomposition), and the feature-level fusion model based on deep learning; the weighted average method realizes data fusion through fixed or empirical weight assignment, but it is difficult to adapt to the non-linear correlation of modal characteristics in dynamic scenes; the multi-scale transformation technology separates the high-frequency and low-frequency components of the image for fusion, but lacks targeted modeling of the physical differences of cross-modal data (such as penetrability, reflection characteristics); the deep learning-based method automatically learns the fusion rules through an end-to-end network, but it depends on a large amount of labeled data and the interpretability of the model is poor, and it is limited in application in resource-constrained scenarios.

[0004] However, existing methods have significant deficiencies in multimodal image processing. First, the fusion strategy with fixed weights cannot dynamically adapt to the physical property differences of different modal data, resulting in the loss of information in occluded areas or fusion artifacts; second, the multi-scale decomposition method has poor compatibility with the frequency-domain characteristics of non-visible light data (such as millimeter waves, terahertz waves), and the unbalanced energy distribution in frequency bands is likely to cause the disconnection of the correlation between material reflection information and motion characteristics; in addition, due to the black-box characteristics of data-driven deep learning models, it is difficult to achieve the collaborative optimization of cross-modal temporal synchronization and physical mechanisms, and there is a problem of insufficient model generalization ability in low signal-to-noise ratio or small sample scenarios. Summary of the Invention

[0005] This application provides an intelligent image signal processing method and system based on multimodal fusion to solve the technical problems of low image fusion accuracy, loss of information in occluded areas, and significant motion artifacts caused by poor adaptability to dynamic scenes, insufficient collaborative optimization of cross-modal data physical properties, and lack of correlation between material reflection information and motion trajectories in the prior art.

[0006] In a first aspect, this application provides an intelligent image signal processing method based on multimodal fusion, including: Receiving a multimodal input signal, where the multimodal input signal includes visible light imaging data and non-visible light spectrum data; Divide the visible light imaging data into a spatial distribution feature extraction region and a dynamic change trajectory capture region, and at the same time decompose the non-visible light spectrum data into a penetration feature layer and a material reflection characteristic layer; Associate and map the missing pixel information in the dynamic change trajectory capture region with the energy distribution of the penetration feature layer to generate enhanced dynamic trajectory data; Identify the reflection patterns matching specific materials in the target scene in the material reflection characteristic layer, and perform frequency domain superposition of the frequency band responses corresponding to the reflection patterns with the spatial distribution feature extraction region to generate composite spatial features; Based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, construct a multi-modal joint optimization model, and use the multi-modal joint optimization model to dynamically adjust the contribution ratios of the visible light imaging data and the non-visible light spectrum data to generate a fused image signal.

[0007] Optionally, identifying the reflection patterns matching specific materials in the target scene in the material reflection characteristic layer, and performing frequency domain superposition of the frequency band responses corresponding to the reflection patterns with the spatial distribution feature extraction region to generate composite spatial features, includes: Based on a preset material reflection characteristic library, perform frequency band segmentation on the material reflection characteristic layer to extract the reflection pattern frequency band range corresponding to specific materials in the target scene; Dynamically intercept the energy distribution within the reflection pattern frequency band range through band-pass filtering, and generate a filtered reflection pattern frequency band response containing the target material density parameter according to the intercepted energy distribution; Perform a dot product operation on the filtered reflection pattern frequency band response and the frequency domain data corresponding to the spatial distribution feature extraction region to obtain the modulated spatial feature frequency domain distribution; Based on the target material density parameter, perform normalization processing on the modulated spatial feature frequency domain distribution to eliminate the energy scale difference between different modalities; Perform weighted superposition of the normalized frequency domain distribution and the original spatial features of the spatial distribution feature extraction region to generate composite spatial features containing target material identifiers.

[0008] Optionally, performing weighted superposition of the normalized frequency domain distribution and the original spatial features of the spatial distribution feature extraction region to generate composite spatial features containing target material identifiers, includes: Calculate a dynamic fusion weight based on the target material density parameter and the energy distribution in the corresponding region of the penetration feature layer; Convert the normalized frequency domain distribution into a complex domain expression form, and decompose the original spatial features into real and imaginary components; Non-linearly scale the amplitude of the complex-domain expression form according to the dynamic fusion weight to obtain a scaled amplitude distribution, while retaining the phase information of the original spatial features; Perform an orthogonal projection operation on the scaled amplitude distribution and the imaginary component of the original spatial features to generate initial fused frequency-domain data; Based on the energy gradient change in adjacent regions of the penetrative feature layer, perform energy equalization compensation on the initial fused frequency-domain data to eliminate local distortions caused by multi-modal data coverage differences; Perform complex reconstruction on the compensated frequency-domain data and the real component of the original spatial features to generate a composite spatial feature of the target material identification.

[0009] Optionally, associate and map the missing pixel information in the dynamic change trajectory capture area with the energy distribution of the penetrative feature layer to generate enhanced dynamic trajectory data, including: Detect the pixel missing area caused by dynamic blur in the dynamic change trajectory capture area, and mark the pixel missing area as an area to be compensated; Extract the energy distribution pattern corresponding to the spatial position of the area to be compensated in the penetrative feature layer; Based on the energy distribution pattern, perform energy attenuation directionality analysis on the peripheral adjacent areas of the area to be compensated, and generate an energy attenuation gradient matrix by calculating the energy change rate of each pixel point in the penetrative feature layer along the movement trajectory direction; Generate dynamic trajectory compensation parameters according to the distribution characteristics of the energy attenuation gradient matrix and the local energy peak of the energy distribution pattern, where the dynamic trajectory compensation parameters include a compensation intensity coefficient and a direction correction factor; Perform multi-scale interpolation on the pixel missing part of the area to be compensated according to the compensation intensity coefficient, and combine the direction correction factor to constrain the interpolation path to generate initial compensated trajectory data; Dynamically correct the initial compensated trajectory data using the energy attenuation gradient matrix, and perform spatial alignment and energy fusion on the corrected compensated trajectory data and the original trajectory data in the dynamic change trajectory capture area to generate enhanced dynamic trajectory data containing cross-modal compensation information.

[0010] Optionally, perform multi-scale interpolation on the pixel missing part of the area to be compensated according to the compensation intensity coefficient, and combine the direction correction factor to constrain the interpolation path to generate initial compensated trajectory data, including: Perform multi-scale decomposition on the energy attenuation gradient matrix according to a preset Gaussian pyramid level to generate a set of gradient distribution levels that match the resolution of the area to be compensated, where each level in the set of gradient distribution levels corresponds to the energy attenuation change characteristics of different scales; Construct a multi-scale interpolation template based on the energy attenuation directionality of each level in the set of gradient distribution levels; Screen the main motion direction in the set of gradient distribution levels according to the direction correction factor to generate level constraint path parameters; In each level, use the weight distribution of the multi-scale interpolation template and the direction tolerance range of the level constraint path parameters to perform adaptive bilinear interpolation on the missing pixel part of the area to be compensated to generate level compensation trajectory data; Fuse the level compensation trajectory data across scales according to the Gaussian pyramid reconstruction rule to generate initial compensation trajectory data that is consistent with the multi-scale energy attenuation characteristics of the penetrability feature layer.

[0011] Optionally, based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, construct a multi-modal joint optimization model, including: Perform phase matching on the timestamp sequence of the enhanced dynamic trajectory data and the energy change period of the composite spatial features to generate synchronous temporal alignment parameters; Based on the synchronous temporal alignment parameters, perform sliding window mutual information calculation on the magnitude of the motion vectors of consecutive frames in the enhanced dynamic trajectory data and the peak frequency domain energy distribution of the composite spatial features to generate modal contribution weight coefficients; Extract the spatial distribution boundary of the peak frequency domain energy in the composite spatial features, map the spatial distribution boundary to the motion vector coverage area corresponding to the enhanced dynamic trajectory data, calculate the overlap rate of the coverage area between the peak frequency domain energy distribution boundary and the motion vector coverage area, and generate a spatial coupling difference matrix according to the overlap rate of the coverage area; According to the gradient distribution direction of the spatial coupling difference matrix, perform a multiplication operation on the modal contribution weight coefficient and the target material density parameter in the material reflection characteristic layer to generate a dynamic fusion constraint factor; Based on the dynamic fusion constraint factor, construct a trajectory energy conservation equation in the motion vector space of the enhanced dynamic trajectory data and a material reflection energy transfer equation in the frequency domain space of the composite spatial features, and generate a joint optimization objective function by alternately iteratively solving the minimum cross entropy of the trajectory energy conservation equation and the material reflection energy transfer equation; Embed the energy attenuation gradient matrix of the penetrative feature layer as a regularization term into the joint optimization objective function, and construct a multi-modal joint optimization model for dynamically adjusting the contribution ratio of the visible light imaging data and the non-visible light spectrum data by constraining the frequency band energy allocation ratio of the visible light imaging data and the non-visible light spectrum data.

[0012] Optionally, use the multi-modal joint optimization model to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectrum data to generate a fused image signal, including: Calculate the frequency band energy allocation ratio parameters of the visible light imaging data and the non-visible light spectrum data based on the dynamic fusion constraint factor and the energy attenuation gradient matrix of the penetrative feature layer; Generate a visible light band energy allocation weight map and a non-visible light penetrative band energy allocation weight map based on the frequency band energy allocation ratio parameters; Convert the visible light imaging data to the frequency domain to generate visible light frequency domain data, dynamically intercept the visible light frequency domain data according to the visible light band energy allocation weight map, retain the frequency band components matching the frequency domain energy peak of the composite spatial feature, and generate weighted visible light frequency domain data; Perform frequency band decomposition on the penetrative feature layer of the non-visible light spectrum data to obtain the energy of each decomposed frequency band, and dynamically weight the energy of each decomposed frequency band according to the non-visible light penetrative band energy allocation weight map to generate penetrative compensation frequency domain data; Perform frequency band superposition on the weighted visible light frequency domain data and the penetrative compensation frequency domain data to obtain a fused frequency domain distribution; Perform inverse transformation processing on the fused frequency domain distribution to generate an initial fused spatial signal, and perform spatial consistency correction on the initial fused spatial signal based on the motion vector coverage area of the enhanced dynamic trajectory data; Perform dot product modulation on the corrected fused spatial signal and the material identification information of the composite spatial feature to generate a fused image signal containing multi-modal physical property associations.

[0013] In a second aspect, the present application provides an intelligent image signal processing system based on multi-modal fusion, including: A receiving module for receiving a multi-modal input signal, where the multi-modal input signal includes visible light imaging data and non-visible light spectrum data; A decomposition module for dividing the visible light imaging data into a spatial distribution feature extraction area and a dynamic change trajectory capture area, and at the same time decomposing the non-visible light spectrum data into a penetrative feature layer and a material reflection characteristic layer; A mapping module, configured to perform an association mapping on the missing pixel information in the dynamically changing trajectory capture area and the energy distribution of the penetrability feature layer, so as to generate enhanced dynamic trajectory data; An identification module, configured to identify a reflection pattern matching a specific material in the target scene in the material reflection characteristic layer, and perform a frequency-domain superposition of the frequency-band response corresponding to the reflection pattern and the spatial distribution feature extraction area, so as to generate a composite spatial feature; A construction module, configured to construct a multi-modal joint optimization model based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial feature, and use the multi-modal joint optimization model to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectrum data, so as to generate a fused image signal.

[0014] In a third aspect, an embodiment of the present application provides a computing device, including a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a multi-modal fusion-based intelligent image signal processing method as described in the first aspect above.

[0015] In a fourth aspect, an embodiment of the present application provides a computer storage medium, storing a computer program, and when the computer program is executed by a computer, it implements a multi-modal fusion-based intelligent image signal processing method as described in the first aspect.

[0016] In the embodiment of the present application, by decomposing the visible light imaging data and the non-visible light spectrum data into different functional layers (spatial distribution feature extraction area, dynamically changing trajectory capture area, penetrability feature layer, and material reflection characteristic layer), and based on the physical property correlation between multi-modal data (such as compensating for missing dynamic trajectories with penetrability energy distribution, and optimizing spatial features by frequency-band superposition of reflection patterns), the adaptive collaborative optimization of cross-modal data in a dynamic scene is realized; by constructing a multi-modal joint optimization model to dynamically adjust the contribution ratio, the information restoration ability of the fused image in the occlusion area and the continuity of the motion trajectory are significantly improved, and the problems of detail loss, artifact generation, and poor dynamic adaptability caused by fixed weights or single-modal limitations in traditional methods are solved, which is particularly suitable for complex illumination, dynamic target tracking, and material identification scenarios.

[0017] Furthermore, through the frequency band segmentation and band - pass filtering dynamic interception of the material reflection characteristic library, the frequency band response of the reflection mode of specific materials in the target scene is accurately extracted. Combining dot - product operation and normalization processing effectively eliminates the energy scale difference between visible light and non - visible light modalities. By weighted superposition of the modulated frequency - domain distribution and the original spatial features, a composite spatial feature with material identification is generated, strengthening the correlation between material reflection characteristics and spatial distribution, and avoiding problems such as material misjudgment or edge blurring caused by energy imbalance in traditional frequency - domain superposition methods. This mechanism significantly improves the accuracy of target material recognition in complex scenes, and at the same time provides a feature input with consistent physical characteristics for subsequent multi - modal joint optimization, ensuring the collaborative optimization effect of the fused image in terms of material details and motion trajectory compensation.

[0018] These aspects or other aspects of the present application will be more clearly understood in the following description of the embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the technical solutions in the embodiments of the present application or in the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0020] Figure 1 It shows a flowchart of an intelligent image signal processing method based on multi - modal fusion provided by the present application; Figure 2 It shows a schematic structural diagram of an intelligent image signal processing system based on multi - modal fusion provided by the present application; Figure 3 It shows a schematic structural diagram of a computing device provided by the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] In order to enable those skilled in the art to better understand the solutions of the present application, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application.

[0022] In some of the processes described in the specification, claims, and the above-mentioned drawings of the present application, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish the different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0023] Researchers have found that existing image fusion technologies are difficult to balance the compensation of occluded area information and the continuity of motion trajectories in dynamic and complex scenarios, and the energy distribution of visible light and non-visible light modal data is unbalanced due to physical property differences, affecting the collaborative optimization effect of material recognition and motion tracking. Based on this, an intelligent image signal processing method based on multimodal fusion is provided. This method realizes the adaptive collaboration of material reflection characteristics and motion trajectory compensation in dynamic scenarios by decomposing the functional layers of visible light and non-visible light data, constructing a cross-modal association mapping mechanism, and a multimodal joint optimization model, effectively improving the spatial resolution of the fused image and the tracking accuracy of dynamic targets.

[0024] The technical solution of the present application is applicable to scenarios that require simultaneous processing of occlusion compensation, material analysis, and motion reconstruction, such as complex road condition perception in autonomous driving, dynamic target recognition under security monitoring, and multimodal lesion analysis of medical images. The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present application without creative efforts shall fall within the scope of protection of the present application.

[0025] Figure 1 The flowchart of an intelligent image signal processing method based on multimodal fusion is provided for the embodiments of the present application, as Figure 1 shown. This method includes: Step 101, receiving a multimodal input signal, where the multimodal input signal includes visible light imaging data and non-visible light spectrum data; In this step, receiving the multi-modal input signal means synchronously acquiring a set of data with different physical characteristics through multi-source sensing devices. Visible light imaging data is the image information in the part of the electromagnetic spectrum that can be perceived by the human vision collected by an optical imaging device, including the texture details, color distribution, and dynamic change characteristics of the scene; non-visible light spectrum data is the electromagnetic wave signal beyond the visible light band obtained through special sensors, such as millimeter wave signals with penetration ability, infrared radiation data reflecting thermodynamic characteristics, or terahertz spectra characterizing the resonance characteristics of matter molecules, which can reveal the contour and material characteristics of occluded objects.

[0026] In this embodiment, first, through the multi-sensor collaborative acquisition mechanism, the visible light image sequence and the non-visible light spectrum data stream are captured in parallel under the time synchronization constraint. Dynamic stability processing is performed on the visible light imaging data, and the inter-frame alignment algorithm based on motion estimation is used to eliminate the jitter artifacts in the acquisition process; environmental interference suppression is implemented on the non-visible light spectrum data, and frequency domain filtering technology is used to remove background noise. Subsequently, a cross-modal space mapping model is constructed, and geometric matching is performed by extracting the structured feature points in the visible light image and the high-confidence reflection regions in the non-visible light data to generate the coordinate transformation relationship of multi-modal space alignment. Finally, the visible light data and the non-visible light data after spatio-temporal alignment are integrated into a multi-modal input signal in a unified format, providing a basis for subsequent feature decomposition.

[0027] For example, in an intelligent security scenario, a visible light camera and a millimeter wave radar deployed indoors operate synchronously. The visible light camera captures high-definition images of personnel activities, and the millimeter wave radar penetrates the wall to detect the reflected signals of moving objects in the occluded area. The system ensures the alignment of the acquisition moments of the two types of data through hardware synchronization signals, and performs electronic image stabilization processing on the visible light video to eliminate the blurred image caused by the slight shaking of the camera. The millimeter wave data removes the fixed echo interference reflected by the wall through adaptive filtering. By extracting the edge features of doors and windows in the visible light image and the spatial coordinates of moving targets in the millimeter wave point cloud, a cross-modal space mapping relationship is established, and finally, a fusion perception signal of the personnel activity trajectory and the hidden objects behind the wall is output.

[0028] 102, divide the visible light imaging data into a spatial distribution feature extraction area and a dynamic change trajectory capture area, and at the same time decompose the non-visible light spectrum data into a penetration feature layer and a material reflection characteristic layer; In this step, the spatial distribution feature extraction region refers to the static image partition in the visible light imaging data that represents the geometric structure and texture details of the target object; the dynamic change trajectory capture region refers to the time-continuous image sequence partition that contains the displacement characteristics of the moving target. The penetrability feature layer refers to the energy distribution layer in the non-visible light spectrum data that reflects the ability of electromagnetic waves to penetrate obstacles, and can characterize the spatial position of the occluded object; the material reflection characteristic layer refers to the feature layer that reveals the material properties of the target through the difference in reflection intensity in different frequency bands.

[0029] In this embodiment, first, motion target segmentation is performed on the visible light imaging data, and the background difference algorithm is used to separate the static scene and the dynamic region. The static image partition containing stable texture is designated as the spatial distribution feature extraction region, and the time-varying image sequence where the moving target is located is designated as the dynamic change trajectory capture region. At the same time, frequency domain energy analysis is performed on the non-visible light spectrum data, and the low-frequency penetrability echo is extracted through a band-stop filter to form the penetrability feature layer, and the spectral clustering algorithm is used to classify the high-frequency reflection signals according to the material categories to generate the material reflection characteristic layer. Finally, the energy distribution of the penetrability feature layer is mapped to the visible light imaging coordinate system through a spatial projection matrix to ensure cross-modal data space alignment.

[0030] Continuing with the intelligent security scenario, for example, the visible light video stream detects the personnel activity area through the three-frame difference method, and the continuous ten-frame images in this area are designated as the dynamic change trajectory capture region, while the fixed backgrounds such as walls and furniture are designated as the spatial distribution feature extraction region. The millimeter-wave radar data separates the low-frequency signal that penetrates the wall through time-frequency transformation to form the penetrability feature layer. After the high-frequency reflection signals are matched with the material property library, the reflection differences between metal instruments and human tissues are identified to construct the material reflection characteristic layer. The coordinate information of the objects hidden behind the wall in the penetrability feature layer is mapped to the visible light image space through the conversion from polar coordinates to pixel coordinates, providing a cross-modal data basis for subsequent dynamic trajectory compensation.

[0031] Step 103, associate and map the missing pixel information in the dynamic change trajectory capture region with the energy distribution of the penetrability feature layer to generate enhanced dynamic trajectory data; In this step, the missing pixel information refers to the image data holes in the dynamic change trajectory capture region caused by motion blur, occlusion, or sensor limitations; the energy distribution of the penetrability feature layer refers to the energy intensity and spatial gradient characteristics formed after electromagnetic waves penetrate obstacles in the non-visible light spectrum data; the association mapping refers to establishing a compensation mechanism for the visible light missing region and the non-visible light energy distribution through cross-modal spatial relationships; the enhanced dynamic trajectory data refers to the continuous motion trajectory data set generated by fusing the visible light dynamic information and the non-visible light penetrability characteristics.

[0032] In this embodiment, first, the motion optical flow analysis method is used to detect the pixel missing positions in the dynamically changing trajectory region, and the geometric boundary of the missing region is determined through edge contour detection. Subsequently, the energy distribution pattern corresponding to the spatial coordinates in the penetrative feature layer is extracted, and the attenuation trend of the penetrative energy is analyzed using the histogram of oriented gradients. Then, a motion trajectory prediction model based on the energy attenuation gradient is constructed, and the energy intensity of the penetrative feature layer is converted into a pixel compensation value through an adaptive interpolation algorithm. Finally, a multi-scale fusion strategy is adopted to perform spatio-temporal weighted superposition of the non-visible light compensation data and the visible light original trajectory to generate an enhanced dynamic trajectory that eliminates occlusion artifacts. This process breaks through the perception limitations of visible light sensors through the physical characteristics of the penetrative feature layer, realizing complementary enhancement of cross-modal motion information.

[0033] Continuing with the intelligent security scenario, for example, when a person passes through the occlusion area of a pillar in a visible light surveillance video, a trajectory break occurs. The millimeter-wave data of the penetrative feature layer detects a continuously moving energy hotspot behind the pillar, and the position of the hotspot corresponding to the visible light missing area is determined through spatial coordinate mapping. Based on the millimeter-wave energy gradient analysis, it is inferred that the motion direction is from left to right, and the bilinear interpolation algorithm with direction constraints is used to generate transitional pixels to fill the trajectory holes. Finally, an enhanced trajectory containing the complete movement path is output, which not only retains the high-precision contour features of visible light but also fuses the millimeter-wave penetrative data to restore the motion state in the occluded stage.

[0034] Step 104, identify the reflection pattern in the material reflection characteristic layer that matches a specific material in the target scene, and perform frequency-domain superposition of the frequency-band response corresponding to the reflection pattern and the spatial distribution feature extraction region to generate a composite spatial feature; In this step, the material reflection characteristic layer refers to the set of features that characterize the electromagnetic wave reflection laws of different materials in the non-visible light spectrum data; the reflection pattern refers to the reflection intensity distribution characteristics shown by a specific material in a specific frequency band; the frequency-band response refers to the energy distribution form of the target material within the characteristic frequency band; and the composite spatial feature refers to the multi-dimensional feature expression that fuses the visible light spatial distribution feature and the non-visible light material reflection characteristic.

[0035] In this embodiment, first, based on a pre-built material reflection characteristic database, a spectral matching algorithm is used to perform frequency band scanning on the material reflection characteristic layer to identify the reflection characteristic frequency band that matches the target material. The target frequency band response is extracted through a dynamic bandwidth filter, and a frequency domain modulation template is generated in combination with the material density parameter. Subsequently, a Fourier transform is performed on the visible light spatial distribution feature extraction region, and the frequency domain feature is multiplied by the material frequency band response to achieve spectral modulation. The cross-modal energy scale difference is eliminated through normalization processing based on energy density, and finally, an inverse transform with phase preservation is used to reconstruct the spatial feature, forming a composite feature that simultaneously includes optical texture and material attributes. This process realizes the deep coupling of material information and spatial features through a frequency domain fusion mechanism driven by physical characteristics.

[0036] Continuing with the intelligent security scenario, for example, a unique high-frequency reflection pattern of metal instruments is identified in the millimeter-wave material reflection characteristic layer. The material feature frequency band is extracted through adaptive band-pass filtering, and a frequency domain mask including the metal density parameter is generated. After the visible light spatial distribution feature extraction region is transformed into the frequency domain, it is multiplied and modulated with the metal frequency band response to enhance the edge sharpness of metal objects in the image. After inverse transform reconstruction, the metal knife that was originally confused with plastic items in the visible light image shows unique material identification characteristics, while retaining the spatial details of the original image, forming a composite spatial feature with material discrimination ability.

[0037] Step 105: Based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial feature, construct a multi-modal joint optimization model, and use the multi-modal joint optimization model to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectrum data to generate a fused image signal; In this step, the temporal synchronization refers to the alignment relationship between the time series of the enhanced dynamic trajectory data and the energy change period of the composite spatial feature on the time axis; the multi-modal joint optimization model refers to a fusion framework that coordinates the physical characteristic differences between visible light and non-visible light data through mathematical modeling; the contribution ratio refers to the weight coefficient dynamically allocated by different modal data according to the scene characteristics during the fusion process; the fused image signal refers to an image output with high resolution and physical feature correlation generated by integrating the advantages of multiple modalities.

[0038] In this embodiment, first, a timestamp marking sequence is constructed for the enhanced dynamic trajectory data. The sequence is aligned with the frequency-domain energy fluctuation period of the composite spatial features through a phase matching algorithm to generate frame-level synchronization parameters. The mutual information between the amplitude of the motion trajectory and the peak value of the frequency-domain energy is calculated based on a sliding window to generate a weight coefficient reflecting the modal correlation strength. Subsequently, the spatial boundary distribution of the composite features is extracted, and the overlap rate with the area covered by the motion trajectory is calculated to construct a spatial coupling difference matrix. The weight coefficient and the material density parameter are fused through gradient direction analysis to form a dynamic fusion constraint factor. Finally, a joint optimization function containing an energy conservation equation and a reflection transfer equation is constructed, and a gradient regularization term of the penetration feature is embedded. The optimal energy distribution ratio of the frequency band is iteratively solved to achieve the adaptive fusion of multimodal data.

[0039] Continuing with the intelligent security scenario, for example, when a person enters the area blocked by the column, the energy change in the metal reflection frequency band of the enhanced trajectory data and the composite spatial features shows temporal synchronization. The fusion weight is dynamically adjusted through mutual information analysis to significantly enhance the fusion contribution of the millimeter-wave material features in the blocked area, while appropriately reducing the weight ratio of the visible light dynamic trajectory. The joint optimization model combines the penetration gradient feature, making the final fused image highlight the metal contour details compensated by millimeter waves in the blocked area and retain the high-definition texture features of the visible light imaging in the non-blocked area. Through the dynamic weight allocation of cross-modal data, the behavior trajectory of a person carrying a weapon passing through the blocked area is completely restored, while ensuring the spatial resolution and material identification accuracy of the image.

[0040] The researchers found that in the prior art, misjudgment of materials often occurs in the identification of material reflection characteristics due to rough frequency band segmentation or energy scale differences, and the lack of adaptation to the physical characteristics of non-visible light spectra in the frequency domain superposition process affects the correlation accuracy between spatial features and material identification; based on this, in some embodiments, according to step 104, the reflection mode matching the specific material in the target scene is identified in the material reflection characteristic layer, and the frequency domain superposition of the frequency band response corresponding to the reflection mode and the spatial distribution feature extraction area is performed to generate composite spatial features, including: Step 201: Based on a preset material reflection characteristic library, segment the frequency bands of the material reflection characteristic layer to extract the frequency band range of the reflection mode corresponding to the specific material in the target scene; In this step, the material reflection characteristic library refers to a database that pre-stores the reflection intensity and waveform characteristics of different materials in specific electromagnetic frequency bands, including the standard reflection spectra of materials such as metals, plastics, and fabrics; the frequency band range of the reflection mode refers to the characteristic frequency interval with significant distinguishability of the target material in the material reflection characteristic layer, which is manifested as an energy peak area or a waveform mutation area.

[0041] In this embodiment, first, a pre-built material reflection characteristic library is loaded, and a sliding window scan is performed on the full-frequency band spectrum of the material reflection characteristic layer. The spectral clustering algorithm is used to match the similarity between the local spectrum obtained by the scan and the standard reflection spectrum in the material library, and the candidate frequency bands with similarity exceeding the threshold are screened out. Then, adjacent candidate frequency bands are merged through energy continuity analysis, and finally, the frequency band range that completely covers the reflection characteristics of the target material is segmented.

[0042] Step 202: Dynamically intercept the energy distribution within the frequency band range of the reflection mode through band-pass filtering, and generate a filtered reflection mode frequency band response containing the density parameters of the target material according to the intercepted energy step by step. In this step, band-pass filtering refers to a technique for selectively extracting signals in a specific frequency band by setting the frequency passband range; dynamic interception refers to an adaptive process of adjusting the filter bandwidth according to the real-time energy distribution; the density parameter of the target material refers to a quantitative index derived by correlating the reflection energy intensity with the physical density of the material; the filtered reflection mode frequency band response refers to an energy distribution matrix that retains the characteristic frequency band of the target material and carries density information.

[0043] In this embodiment, first, a band-pass filter bank with the ability to adjust the bandwidth adaptively is constructed based on the generated frequency band boundary markers. By real-time monitoring of the energy distribution gradient within the target frequency band, the cut-off frequency range of the filter is dynamically shrunk or expanded to ensure complete interception of the material reflection characteristics. The spatial integration calculation is performed on the energy of the intercepted frequency band, and combined with the density mapping function in the material reflection characteristic library, the energy intensity is converted into density parameters. Finally, a frequency band response matrix containing the spatial density distribution is generated, whose amplitude represents the distribution density of the target material, and the phase retains the original spectral characteristics.

[0044] Step 203: Perform an element-wise multiplication operation on the filtered reflection mode frequency band response and the frequency domain data corresponding to the spatial distribution feature extraction region to obtain the modulated spatial feature frequency domain distribution. In this step, the element-wise multiplication operation refers to a mathematical operation of multiplying the corresponding elements of two frequency domain matrices; the modulated spatial feature frequency domain distribution refers to enhancing the feature expression of the frequency components related to the target material through frequency domain operations; the density parameter of the target material refers to a quantitative characteristic parameter generated through the mapping relationship between the reflection energy and the physical density.

[0045] In this embodiment, first, a fast Fourier transform is performed on the visible light spatial distribution feature extraction region to generate a frequency domain matrix containing spatial texture details. Subsequently, an element-wise multiplication operation is performed on the obtained filtered reflection mode frequency band response matrix, so that the amplitude of the frequency components related to the target material in the visible light frequency domain is enhanced, and the irrelevant components are suppressed. Finally, the modulated frequency domain amplitude is non-linearly scaled in combination with the density parameter of the target material, and the original phase information is retained to generate a frequency domain distribution that simultaneously contains optical texture and material density characteristics.

[0046] Step 204: Normalize the modulated spatial feature frequency domain distribution based on the target material density parameter to eliminate the energy scale differences between different modalities; In this step, the target material density parameter refers to a quantization eigenvalue generated through the mapping relationship between reflection energy intensity and physical density, and is used to characterize the distribution density of the material in space; normalization processing refers to the process of making the energy intensities of data in different modalities reach a comparable dimension through proportional scaling; the energy scale difference refers to the signal intensity magnitude difference caused by different sensing principles between the visible light and non-visible light modalities.

[0047] In this embodiment, first, a normalization coefficient matrix is constructed based on the target material density parameter, and this coefficient is generated by a preset density-energy conversion function in the material reflection characteristic library. Multiply the modulated spatial feature frequency domain distribution matrix element by element with the normalization coefficient, so that the high-energy regions of the non-visible light modality are proportionally compressed according to the density parameter, and the low-density regions are moderately enhanced. Subsequently, through the frequency domain energy histogram matching algorithm, adjust the energy distribution of the high-frequency components of the visible light modality to make its morphology approach that of the normalized non-visible light energy histogram, and finally achieve the scale alignment of cross-modal frequency domain energy.

[0048] Step 205: Weightedly superimpose the normalized frequency domain distribution and the original spatial features of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier; The normalized frequency domain distribution refers to a frequency domain feature matrix that has undergone energy scale alignment processing, which retains the reflection characteristics of the target material and eliminates cross-modal dimension differences; the original spatial features refer to the texture, edges and other detailed information obtained through the spatial distribution feature extraction region in the visible light imaging data; the composite spatial feature of the target material identifier refers to a multi-dimensional feature expression that fuses visible light spatial details and non-visible light material attributes, and has the dual representation ability of physical characteristics and visual information.

[0049] First, calculate the dynamic fusion weight matrix based on the target material density parameter and the energy gradient distribution of the penetration feature layer. Convert the normalized frequency domain distribution into a complex domain expression form, and separate the amplitude and phase components. Perform two-dimensional wavelet decomposition on the original spatial features to extract the low-frequency contour and high-frequency detail components. Nonlinearly scale the amplitude component of the frequency domain according to the dynamic fusion weight, while retaining the phase information of the original spatial features. Fuse the scaled amplitude and high-frequency detail components in the spatial domain through the orthogonal projection algorithm, and then perform a residual connection with the low-frequency contour component. Finally, perform local energy equalization compensation on the fusion result based on the energy gradient of the penetration feature layer to eliminate the edge distortion caused by the modality coverage difference, and generate a composite feature that simultaneously contains the material identifier and spatial details.

[0050] To address the problem in the prior art that the traditional frequency-domain weighted superposition method has difficulty in balancing the retention of phase information and energy equilibrium, resulting in distortion or information loss of the fused features at the material boundaries; in some embodiments, as described in step 205, the normalized frequency-domain distribution is weighted and superimposed with the original spatial features of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier, including: Step 301, calculate a dynamic fusion weight based on the target material density parameter and the energy distribution in the corresponding region of the penetrability feature layer; In this step, the target material density parameter refers to a quantization index generated through the mapping relationship between the reflection energy intensity and the physical density, which is used to characterize the distribution density of the material in space; the energy distribution of the penetrability feature layer refers to the gradient change in space of the low-frequency energy representing the electromagnetic wave penetration ability in the non-visible light spectrum data; the dynamic fusion weight refers to the fusion ratio coefficient of visible light and non-visible light data dynamically adjusted according to the material density and penetrability energy, which is used to optimize the complementarity of multi-modal information.

[0051] In this embodiment, first, a spatial gradient calculation is performed on the penetrability feature layer to generate a gradient matrix representing the energy attenuation direction and intensity. The target material density parameter matrix is multiplied element by element (Hadamard product operation) with the gradient matrix to obtain a preliminary weight distribution. The preliminary weight is non-linearly mapped through the Sigmoid function, and combined with the spatial continuity constraint (such as smoothing processing based on bilateral filtering), the influence of isolated noise points is eliminated. Finally, a dynamic fusion weight matrix is generated, and its numerical distribution satisfies: a higher weight is assigned to the non-visible light modality in the regions with high material density and significant attenuation of penetrability energy, and the contribution of visible light is enhanced in the regions with low density and stable energy.

[0052] Step 302, convert the normalized frequency-domain distribution into a complex-domain expression form, and decompose the original spatial feature into a real component and an imaginary component; In this step, the normalized frequency-domain distribution refers to the frequency-domain feature matrix after energy scale alignment processing, whose amplitude represents the energy intensity of the target material, and the phase retains the spatial structure information; the complex-domain expression form refers to the mathematical representation of decomposing the frequency-domain data into amplitude and phase components; the real component corresponds to the frequency-domain cosine coefficient, representing the spatial low-frequency contour; the imaginary component corresponds to the frequency-domain sine coefficient, representing the spatial high-frequency details.

[0053] In this embodiment, first, a complex reconstruction is performed on the normalized frequency-domain distribution: taking the frequency-domain amplitude as the modulus length, initializing the angle parameter based on the phase information of the original visible light spatial features, and generating a complex matrix containing the material energy intensity and spatial phase features. Subsequently, a Hilbert transform is performed on the original visible light spatial features to decompose them into a real part (low-frequency component reflecting the overall contour) and an imaginary part (high-frequency detail component containing edge textures). The amplitude of the complex matrix is coupled with the imaginary part component through an orthogonal projection algorithm, and the real part component is retained as the basic spatial structure.

[0054] Step 303: Nonlinearly scale the amplitude of the complex-domain representation form according to the dynamic fusion weight to obtain the scaled amplitude distribution, while retaining the phase information of the original spatial features. In this step, the dynamic fusion weight refers to the modal contribution coefficient dynamically adjusted based on the material density and penetration performance energy; the amplitude of the complex-domain representation form refers to the modulus length component representing the energy intensity in the frequency-domain feature matrix; the phase information refers to the angle component reflecting the spatial structure relationship in the frequency-domain data; the nonlinear scaling refers to non-uniform intensity adjustment of the amplitude according to the weight coefficient while keeping the phase unchanged to maintain the integrity of the spatial structure.

[0055] In this embodiment, first, the dynamic fusion weight matrix is input into an adaptive gain controller to generate a scaling coefficient positively correlated with the amplitude intensity. A nonlinear mapping is performed on the complex-domain amplitude using an exponential function: enhancing the amplitude according to the exponential law in the weight-dominated region and maintaining linear scaling in the low-weight region. During the scaling process, the original phase angle is strictly retained, and the adjusted amplitude and the original phase are recombined into a complex form through polar coordinate transformation. Finally, a DC component correction is performed on the scaled frequency-domain distribution to eliminate the baseline shift introduced by the nonlinear operation.

[0056] Step 304: Perform an orthogonal projection operation on the scaled amplitude distribution and the imaginary part component of the original spatial features to generate initial fusion frequency-domain data. In this step, the scaled amplitude distribution refers to the frequency-domain energy intensity matrix nonlinearly adjusted by the dynamic weight; the imaginary part component of the original spatial features refers to the sine component representing the high-frequency details after the frequency-domain decomposition of the visible light image; the orthogonal projection operation refers to the mathematical operation of superimposing the components of two vector spaces in the vertical direction, which is used to retain the principal component features and eliminate redundant information.

[0057] In this embodiment, first, the scaled amplitude distribution is converted into a vector form, and a feature vector space based on energy intensity is constructed. The imaginary part components of the original spatial features are subjected to Gram - Schmidt orthogonalization processing to generate an orthogonal basis vector group. The scaled amplitude vector is projected onto the orthogonal basis space of the imaginary part components through the projection matrix calculation, and the amplitude components that are linearly independent of the imaginary part components are retained. Finally, the projection residuals and the linear combination of the orthogonal basis components are reconstructed into frequency - domain data to form an initial fusion frequency - domain matrix that not only retains high - frequency details but also enhances key features.

[0058] Step 305: Based on the energy gradient change in adjacent regions of the penetrability feature layer, perform energy equalization compensation on the initial fusion frequency - domain data to eliminate local distortion caused by multi - modal data coverage differences. In this step, the energy gradient change in adjacent regions of the penetrability feature layer refers to the spatial distribution difference of the energy representing penetrability in non - visible light spectral data, which is obtained by calculating the energy intensity difference between adjacent pixels; energy equalization compensation refers to adjusting the energy distribution of the fusion data according to the gradient change to eliminate sudden brightness changes or texture distortions caused by different modal coverage ranges; local distortion refers to image artifacts generated during multi - modal fusion due to data registration errors or energy distribution imbalances.

[0059] In this embodiment, first, perform Gaussian filtering on the penetrability feature layer to generate a smoothed energy distribution map. Calculate the Sobel gradient of this energy map to identify energy mutation regions. Based on the gradient amplitude, construct a compensation coefficient matrix, where high - gradient regions correspond to low compensation coefficients to suppress energy overshoot, and low - gradient regions are given high compensation coefficients to enhance details. Multiply the compensation coefficient matrix element - by - element with the initial fusion frequency - domain data to achieve adaptive adjustment of the frequency - domain energy. Finally, restore the edge sharpness lost due to the compensation operation through deconvolution operation to generate the equalized fusion frequency - domain data.

[0060] Step 306: Perform complex reconstruction on the compensated frequency - domain data and the real - part components of the original spatial features to generate composite spatial features of the target material identification. In this step, the compensated frequency - domain data refers to the optimized frequency - domain matrix after energy equalization processing, which eliminates multi - modal fusion artifacts and retains key features; the real - part components of the original spatial features refer to the cosine coefficient components after frequency - domain decomposition of the visible - light image, representing spatial low - frequency contour information; complex reconstruction refers to the operation of recombining the real and imaginary part components in the frequency domain into a complex form and performing an inverse transformation; the composite spatial features of the target material identification refer to the final output features that fuse visible - light spatial details and non - visible - light material attributes, possessing dual expressions of physical characteristics and visual information.

[0061] In this embodiment, first, the compensated frequency-domain data (including the optimized imaginary component) is combined with the real component of the original spatial feature to generate a complex matrix containing complete frequency-domain information. The complex matrix is transformed into the spatial domain through inverse Fourier transform to obtain a preliminary fused image. Subsequently, based on the material density distribution of the penetrability feature layer, a spatial-domain masking operation is performed on the fused image: a semi-transparent color mark is superimposed on the high-density material area, while the natural transition of the low-frequency contour is retained. The finally generated composite spatial feature enhances the visual recognition of the target material through hue and saturation adjustment while retaining the visible light texture details.

[0062] In traditional methods, the pixel missing areas caused by dynamic blur or occlusion mostly rely on visible light data interpolation compensation, which is prone to trajectory breakage due to the lack of penetrability information. Based on this, in some embodiments, according to step 103, the missing pixel information in the dynamically changing trajectory capture area is associated and mapped with the energy distribution of the penetrability feature layer to generate enhanced dynamic trajectory data, including: Step 401, detecting the pixel missing areas caused by dynamic blur in the dynamically changing trajectory capture area, and marking the pixel missing areas as areas to be compensated; In this step, the dynamically changing trajectory capture area refers to the spatial range in a video sequence or image frame for tracking the target motion trajectory; dynamic blur refers to the image blurring phenomenon caused by the rapid movement of the target or the excessive camera exposure time; the pixel missing area refers to the part of the details lost due to the discontinuous motion trajectory in the dynamic blur area; the area to be compensated refers to the set of spatial coordinates marked for cross-modal information compensation.

[0063] In this embodiment, first, the motion vector field between adjacent frames is calculated by the optical flow method to identify the trajectory area where the motion speed exceeds a preset threshold (such as 15 pixels per second). Laplacian edge detection is performed on the high-motion-speed area, and the edge breakage area caused by blur is located by comparing the gradient difference between the static background template and the current frame. Then, an adaptive threshold segmentation algorithm is used to mark the pixel clusters with gradient difference values lower than the dynamic blur determination threshold as areas to be compensated, and a binary mask marking matrix is generated. Finally, the holes in the mask are filled by morphological closing operation to ensure the spatial continuity of the areas to be compensated.

[0064] Step 402, extracting the energy distribution pattern corresponding to the spatial position of the area to be compensated in the penetrability feature layer; In this step, the penetrability feature layer refers to the low-frequency energy distribution layer representing the electromagnetic wave penetration ability in the non-visible light spectrum data; spatial position correspondence refers to the coordinate mapping relationship established through cross-modal spatial registration; the energy distribution pattern refers to the multi-dimensional feature set composed of the intensity, gradient direction, and frequency band of the penetration energy in a specific spatial area.

[0065] In this embodiment, first, based on the mask matrix of the area to be compensated generated in step 401, a spatial coordinate transformation is performed on the penetrative feature layer: the area to be compensated in the visible light image coordinate system is mapped to the polar coordinate system of the non-visible light data through pre-calibrated affine transformation parameters. The bilinear interpolation algorithm is used to resample the penetrative feature layer, and the energy distribution of the sub-region corresponding exactly to the area to be compensated is extracted. Then, the first three principal components of this sub-region are extracted through principal component analysis (PCA), and a three-dimensional feature vector including energy intensity, gradient direction, and frequency band correlation is constructed to form an energy distribution pattern descriptor.

[0066] Step 403: Based on the energy distribution pattern, perform an energy attenuation directivity analysis on the adjacent area around the area to be compensated, and generate an energy attenuation gradient matrix by calculating the energy change rate of each pixel point in the penetrative feature layer along the motion trajectory direction; In this step, the energy attenuation directivity analysis refers to quantifying the attenuation law of energy along the motion trajectory direction based on the physical characteristics of the penetrative feature layer; the motion trajectory direction refers to the main direction of target motion calculated by the optical flow method or the motion vector field; the energy attenuation gradient matrix refers to a matrix composed of the first-order derivatives of each pixel point along the motion direction, and its value reflects the energy change rate under unit displacement.

[0067] In this embodiment, first, based on the energy distribution characteristics extracted in step 402, the main motion direction of the target object is determined. Perform a gradient analysis on the penetrative feature layer around the adjacent pixels of the area to be compensated: by calculating the energy change intensity of each pixel point in the horizontal and vertical directions, the horizontal gradient value and the vertical gradient value are obtained respectively. Subsequently, the gradient values in these two directions are synthesized according to the ratio of the main motion angle. For example, in the main motion direction, the horizontal gradient is superimposed with a high proportion of weight, and the vertical gradient is fused according to the corresponding ratio to generate a comprehensive gradient value along the motion trajectory direction. To eliminate the dimensional deviation caused by the difference in reflection characteristics of different materials, the gradient values are normalized to make the data of each area comparable. Then, with each pixel as the center, calculate the energy difference between it and the adjacent pixel along the motion direction, and convert it into the energy change rate per unit distance in combination with the actual physical distance - if the energy of the latter pixel is lower than that of the current pixel, it is marked as a negative value (energy attenuation), otherwise it is a positive value (energy enhancement). Finally, the energy change rates of all pixels are integrated into a matrix, where the continuous negative value area accurately maps the energy attenuation band of the motion path, and the positive value area shows the energy enhancement characteristics, providing a quantitative basis driven by physical characteristics for subsequent compensation.

[0068] Step 404: Generate dynamic trajectory compensation parameters according to the distribution characteristics of the energy attenuation gradient matrix and the local energy peak of the energy distribution pattern. The dynamic trajectory compensation parameters include a compensation intensity coefficient and a direction correction factor; In this step, the local energy peak refers to the extreme point in the energy distribution pattern with a significantly higher intensity than the surrounding area; the dynamic trajectory compensation parameters include a compensation intensity coefficient (controlling the intensity level of pixel compensation) and a direction correction factor (adjusting the offset angle of the compensation path), which are used to guide the repair of cross-modal missing information.

[0069] In this embodiment, first, a local maximum detection algorithm is used in the energy distribution pattern to identify the energy peak point as the anchor point for trajectory compensation. Based on the distribution characteristics of the energy attenuation gradient matrix, the angle difference between the gradient descent direction centered on the peak point and the original motion direction is calculated to generate the direction correction factor. At the same time, the compensation intensity coefficient is calculated according to the product of the gradient attenuation amplitude and the peak intensity. Higher compensation intensities are given to regions with more severe attenuation and higher peaks. Finally, the Gaussian mixture model is used to perform spatial smoothing on the compensation parameters to eliminate the influence of isolated noise points and generate a set of dynamic compensation parameters that match the physical characteristics of the motion trajectory.

[0070] Step 405: Perform multi-scale interpolation on the missing pixel part of the area to be compensated according to the compensation intensity coefficient, and combine the direction correction factor to constrain the interpolation path to generate initial compensation trajectory data; In this step, the compensation intensity coefficient refers to the interpolation weight parameter dynamically adjusted according to the material attenuation characteristics, which controls the intensity level of pixel compensation; the direction correction factor refers to the path offset angle parameter generated based on the gradient direction deviation, which constrains the interpolation direction; multi-scale interpolation refers to adaptively compensating the missing area at multiple resolution levels, taking into account both the overall trajectory continuity and the local detail restoration; the initial compensation trajectory data refers to the preliminary repair result generated after fusing cross-modal information.

[0071] In the embodiment of the present application, first, based on the compensation intensity coefficient and the direction correction factor generated in step 404, the area to be compensated is decomposed into multiple resolution levels according to the Gaussian pyramid. A large-scale interpolation template is constructed in the high-level low-resolution image, and the interpolation window size is dynamically adjusted according to the compensation intensity coefficient - a large-range bilinear interpolation is used to fill the main body of the motion trajectory in the high-coefficient area, and a small window is used to retain details in the low-coefficient area; at the same time, the interpolation path is angle-constrained by the direction correction factor, and pixel points consistent with the correction direction are preferentially selected from adjacent pixels in the area to be compensated for weighted calculation. Subsequently, in the low-level high-resolution image, combined with the energy distribution characteristics of the penetrative feature layer, the interpolation result is optimized for details: the edge sharpness of the metal material is enhanced in the area where the millimeter-wave energy is concentrated, and the natural transition of the visible light texture is maintained in the energy attenuation area. Finally, the interpolation results of each level are stacked step by step through the pyramid reverse fusion algorithm. The high-level results ensure the macroscopic continuity of the motion trajectory, and the low-level corrections improve the microscopic details such as the blade tooth pattern, generating initial compensation trajectory data that not only conforms to the physical characteristics of non-visible light but also retains the spatial details of visible light.

[0072] Step 406: Dynamically correct the initial compensation trajectory data using the energy attenuation gradient matrix, spatially align and perform energy fusion on the corrected compensation trajectory data and the original trajectory data in the dynamically changing trajectory capture region to generate enhanced dynamic trajectory data containing cross-modal compensation information; In this step, dynamic correction refers to iteratively optimizing the compensation data based on physical attenuation characteristics; spatial alignment refers to matching the coordinate systems of the cross-modal compensation trajectory and the original visible light trajectory; energy fusion refers to integrating the information of visible light texture and non-visible light physical characteristics; enhanced dynamic trajectory data refers to the finally output high-precision spatio-temporally continuous trajectory.

[0073] In this embodiment, first, the initial compensation trajectory data and the energy attenuation gradient matrix are weighted and fused to enhance the compensation intensity in the regions with significant gradient attenuation, and the matching degree between the compensation trajectory and the gradient distribution is optimized by iterative least squares method. Subsequently, the corrected compensation trajectory is mapped to the coordinate system of the visible light dynamic trajectory capture region using an affine transformation model, and sub-pixel level spatial alignment is achieved based on feature point matching (such as SIFT key points). Finally, through a wavelet domain energy fusion strategy, the macroscopic continuity of the visible light trajectory is retained in the low frequency band, and the material characteristics of the non-visible light compensation trajectory are injected in the high frequency band to generate enhanced trajectory data with unified spatio-temporal dimensions and complete physical characteristics.

[0074] To explain the problem in the prior art that the single-scale interpolation method is difficult to adapt to the multi-scale feature differences in the occlusion region, resulting in a large deviation between the compensation trajectory and the true motion direction; in some embodiments, according to step 405, multi-scale interpolation is performed on the pixel missing part of the region to be compensated according to the compensation intensity coefficient, and the interpolation path is constrained by combining the direction correction factor to generate the initial compensation trajectory data, including: Step 501: Perform multi-scale decomposition on the energy attenuation gradient matrix according to the preset Gaussian pyramid levels to generate a set of gradient distribution levels matching the resolution of the region to be compensated, where each level in the set of gradient distribution levels corresponds to different scale energy attenuation change characteristics; In this step, the Gaussian pyramid level refers to the multi-resolution level structure of the image generated by successive downsampling and Gaussian filtering; multi-scale decomposition refers to decomposing the original data into a set of features with different spatial resolutions to separate the macroscopic trend and microscopic details; the set of gradient distribution levels refers to the manifestation of the energy attenuation gradient matrix at different resolution levels, where the high level represents the large-range energy attenuation trend and the low level retains the local detail changes.

[0075] In this embodiment, first, a Gaussian pyramid is constructed for the energy attenuation gradient matrix: by recursively applying Gaussian filtering and downsampling at every other point, a hierarchical sequence with gradually decreasing resolution is generated. The gradient features at the corresponding scale are retained at each level - the high level (low resolution) highlights the overall attenuation trend of the motion trajectory, and the low level (high resolution) maintains the fine gradient changes in the edge region. Subsequently, the gradient data at each level is aligned to the original resolution space of the region to be compensated through an interpolation algorithm, forming a set of hierarchical gradient distributions that match the scale characteristics and compensation requirements.

[0076] Step 502: Based on the energy attenuation directionality of each level in the set of hierarchical gradient distributions, construct a multi-scale interpolation template. In this step, the multi-scale interpolation template refers to an adaptive compensation operator constructed based on the gradient direction characteristics at different resolution levels. Its shape and scale vary with the level and are used to match the repair requirements of trajectory features at different granularities.

[0077] In this embodiment, first, a direction consistency analysis is performed on each level in the set of hierarchical gradient distributions: the dominant gradient direction of each level is detected through the Hough transform and used as the main axis direction of the interpolation template. The template size is adjusted according to the level resolution - a large-scale elliptical template is used for the high level (low resolution), and a compact rhombus template is used for the low level (high resolution). The template weight distribution is dynamically adjusted by the gradient amplitude: a high weight is given to the center and a low weight to the edge in the high gradient attenuation region; a uniform weight is used in the low gradient region. Finally, a set of interpolation templates that match the physical characteristics of each level is generated.

[0078] Step 503: Screen the main motion direction in the set of hierarchical gradient distributions according to the direction correction factor to generate hierarchical constraint path parameters. In this step, the direction correction factor refers to a path angle adjustment parameter generated by calibrating the energy attenuation direction deviation; the hierarchical constraint path parameter refers to a cross-level motion path descriptor generated by fusing multi-scale gradient direction features and the direction correction factor, which is used to unify the direction consistency of the compensation paths at different resolution levels.

[0079] In this embodiment, first, the main motion direction is extracted from each level in the set of hierarchical gradient distributions, and the detected main direction is calibrated by the direction correction factor. In the high-level (low-resolution) data, effective directions are screened based on the angle threshold of the direction correction factor, and noise directions with large deviations are removed; in the low-level (high-resolution) data, the local direction is fine-tuned using the correction factor to eliminate the direction jitter at the detail level. Subsequently, the calibrated multi-scale direction features are fused according to the level weights to generate a set of hierarchical constraint path parameters. Its macroscopic level parameters define the overall trend of the motion trajectory, and the microscopic level parameters refine the precise angle of the local path, ultimately ensuring the direction consistency of the cross-scale compensation path.

[0080] Step 504: In each level, using the weight distribution of the multi-scale interpolation template and the direction tolerance range of the hierarchical constraint path parameters, perform adaptive bilinear interpolation on the missing pixel parts of the area to be compensated to generate hierarchical compensation trajectory data; In this step, the hierarchical compensation trajectory data refers to the local repair results generated by multi-scale interpolation that match the characteristics of each resolution level, which not only retains the macroscopic motion trend but also repairs the microscopic detail features; adaptive bilinear interpolation refers to a compensation method that dynamically adjusts the interpolation coefficients according to the template weights and direction constraints to achieve the alignment of the physical characteristics of cross-modal data.

[0081] In this embodiment, based on the weight distribution of the multi-scale interpolation template and the direction tolerance range of the hierarchical constraint path parameters, the area to be compensated is repaired layer by layer: in the high level (low resolution), a large-range bilinear interpolation is used to fill the trajectory break area along the main motion direction, and the template weight distribution enhances the compensation intensity in the areas where the gradient decays significantly (such as the edge of the metal instrument); in the low level (high resolution), the interpolation neighborhood is screened in combination with the direction tolerance range, and only the pixel points consistent with the correction direction are selected for weighted calculation, and at the same time, the energy distribution characteristics of the penetrating feature layer (such as the millimeter wave reflection intensity) are injected to optimize the detailed texture. Through layer-by-layer interpolation compensation, the high-level result ensures the continuity of the motion trajectory, and the low-level repair restores fine structures such as serrated edges, and finally generates hierarchical compensation trajectory data that matches the physical attenuation characteristics and visual features.

[0082] Step 505: Cross-scale fuse the hierarchical compensation trajectory data according to the Gaussian pyramid reconstruction rule to generate initial compensation trajectory data that is consistent with the multi-scale energy attenuation characteristics of the penetrating feature layer; In this step, the Gaussian pyramid reconstruction rule refers to a method for multi-scale data fusion achieved by successive upsampling and weighted superposition; the initial compensation trajectory data refers to the compensation output that fuses the repair results of multiple levels and retains the cross-scale physical characteristics, and needs to be coupled with the energy attenuation characteristics of the penetrating feature layer in the frequency domain and the spatial domain.

[0083] In this embodiment, the hierarchical compensation trajectory data generated in step 504 is upsampled level by level from low resolution to high resolution according to the pyramid levels. After each level of upsampling, it is weighted and superimposed with the adjacent high-level data: the high-level data provides the macroscopic energy attenuation framework of the motion trajectory, and the low-level data injects the local attenuation characteristics of the detailed texture. During the superimposition process, the fusion weight is dynamically adjusted based on the energy distribution ratio of each level of the penetrative feature layer - the weight of the detail layer is increased in the area where the millimeter-wave high-frequency energy is concentrated, and the contribution of the macroscopic layer is enhanced in the low-frequency energy area, so that the fusion result completely matches the multi-scale attenuation of the penetrative feature layer in terms of frequency domain characteristics. The finally generated initial compensation trajectory data not only has the spatial continuity of the visible light trajectory but also integrates the physical attenuation characteristics of non-visible light modalities.

[0084] Since the existing multi-modal fusion models lack the time synchronization constraint and the optimization mechanism driven by physical characteristics, the adjustment of the contribution ratio lags behind or the energy distribution is unbalanced in dynamic scenarios; to solve the above problems, in some embodiments, according to step 105, based on the time synchronization between the enhanced dynamic trajectory data and the composite spatial features, a multi-modal joint optimization model is constructed, including: Step 601, perform phase matching on the timestamp sequence of the enhanced dynamic trajectory data and the energy change period of the composite spatial features to generate synchronous timing alignment parameters; In this step, the enhanced dynamic trajectory data refers to the spatio-temporally continuous motion trajectory after fusing cross-modal compensation, which contains the timing information of the target motion; the energy change period of the composite spatial features refers to the regular pattern of the visible light and non-visible light modalities (such as material reflection and penetrative energy) fluctuating over time; phase matching refers to aligning the periodic fluctuations of different modal data on the time axis by adjusting the timing offset; the synchronous timing alignment parameters refer to the key parameters for eliminating cross-modal timing differences, including delay compensation amount, phase scaling factor, etc.

[0085] In this embodiment, first, the main energy change period of the composite spatial features is extracted, and at the same time, the motion period corresponding to the timestamp sequence of the enhanced trajectory data is analyzed. The dynamic time warping algorithm is used to calculate the phase difference between the two to generate timing compensation parameters: a delay compensation is applied to the trajectory data to align the energy peak moments, and the motion period length is adjusted through a scaling factor to make it consistent with the non-visible light energy fluctuation period. Finally, through the sliding window cross-correlation verification, it is ensured that the cross-modal data is completely synchronized at the key event moments such as the start and end points of the waving trajectory and the speed peak.

[0086] Step 602, based on the synchronous timing alignment parameters, perform sliding window mutual information calculation on the motion vector amplitude of consecutive frames in the enhanced dynamic trajectory data and the frequency domain energy peak distribution of the composite spatial features to generate modal contribution weight coefficients; In this step, the sliding window mutual information calculation refers to measuring the statistical correlation of the distributions of two types of data within a fixed time window, and quantifying the degree of information sharing between modalities; the modal contribution weight coefficient refers to the fusion weight dynamically allocated based on the mutual information value, which is used to balance the contribution ratios of different modalities in feature representation.

[0087] In this embodiment, based on the synchronous timing alignment parameters, spatio-temporal alignment is performed on the amplitude of the motion vectors of consecutive frames of the enhanced dynamic trajectory data and the peak distribution of the frequency domain energy of the composite spatial features. A sliding window mechanism is adopted to calculate the joint probability density of the motion vector amplitude histogram and the frequency domain energy distribution within the window, and the correlation between the two is quantified through the mutual information formula. The higher the mutual information value, the greater the contribution of the non-visible light modality (such as millimeter wave energy distribution) to the trajectory feature representation during this period, and vice versa, the visible light modality dominates. Finally, a modal contribution weight coefficient that changes dynamically with time is generated to guide the weighted fusion of multi-modal features.

[0088] Step 603: Extract the spatial distribution boundary of the peak of the frequency domain energy in the composite spatial feature, map the spatial distribution boundary to the motion vector coverage area corresponding to the enhanced dynamic trajectory data, calculate the overlap rate of the coverage area between the peak distribution boundary of the frequency domain energy and the motion vector coverage area, and generate a spatial coupling difference matrix according to the overlap rate of the coverage area; In this step, the spatial distribution boundary of the peak of the frequency domain energy refers to the contour of the non-visible light energy concentration area determined by threshold segmentation; the motion vector coverage area refers to the spatial projection range of the consecutive frame motion vectors in the enhanced dynamic trajectory data; the overlap rate of the coverage area refers to the ratio of the intersection area of the two types of areas to their union area, which is used to quantify the spatial consistency; the spatial coupling difference matrix refers to a two-dimensional matrix reflecting the cross-modal spatial matching degree generated based on the overlap rate, and the lower the value, the better the spatial consistency.

[0089] In this embodiment, first, a morphological closing operation is performed on the frequency domain energy distribution of the composite spatial feature. After eliminating noise interference, the closed boundary of the energy peak area is extracted through adaptive threshold segmentation. The boundary coordinates are mapped to the motion vector space of the enhanced dynamic trajectory data, and the area ratio of the intersection area and the union area of the two is calculated through polygon Boolean operations to obtain the pixel-by-pixel coverage area overlap rate. The sliding window statistical method is used to calculate the local overlap rate mean for each block of the full-frame image, and a spatial coupling difference matrix is generated in combination with the motion vector amplitude. The low-value areas in the matrix indicate a high degree of cross-modal spatial feature matching, and the high-value areas indicate the spatial misalignment positions that need to be further corrected.

[0090] Step 604: According to the gradient distribution direction of the spatial coupling difference matrix, perform a multiplication operation on the modal contribution weight coefficient and the target material density parameter in the material reflection characteristic layer to generate a dynamic fusion constraint factor; In this step, the dynamic fusion constraint factor refers to an optimization coefficient comprehensively generated by spatial coupling difference, modal weight, and material density, which is used to constrain the intensity and direction of multi-modal data fusion; the product operation refers to the weighted fusion of the modal weight coefficient and the material density parameter in the gradient direction of the spatial coupling difference to achieve the dynamic adaptation of physical characteristics and motion features.

[0091] In this embodiment, first, the gradient direction of the spatial coupling difference matrix is extracted to identify the main direction of cross-modal spatial misalignment. The modal contribution weight coefficient and the target material density parameter are subjected to a directional product in the gradient direction: in the region where the gradient direction is consistent with the main motion direction, the product operation enhances the fusion weight; in the region where the gradient deviates, the product operation suppresses the weight. The product result is spatially smoothed by a non-uniform sampling algorithm to eliminate local mutation noise, and a dynamic fusion constraint factor matrix that strictly matches the physical characteristics of the motion trajectory is generated.

[0092] Step 605: Based on the dynamic fusion constraint factor, construct a trajectory energy conservation equation in the motion vector space of the enhanced dynamic trajectory data, and construct a material reflection energy transfer equation in the frequency domain space of the composite space feature. By alternately iteratively solving the cross-entropy minimum of the trajectory energy conservation equation and the material reflection energy transfer equation, a joint optimization objective function is generated. In this step, the trajectory energy conservation equation refers to a mathematical constraint that ensures the continuity of the conversion between kinetic energy and potential energy in the motion vector space; the material reflection energy transfer equation refers to a physical model that describes the interaction relationship between material characteristics and frequency domain reflection energy; the cross-entropy minimum refers to achieving the optimization objective by minimizing the probability distribution difference between the two types of equations; the joint optimization objective function refers to a unified optimization framework that fuses kinematic and physical characteristic constraints.

[0093] In this embodiment, based on the dynamic fusion constraint factor, a trajectory energy conservation equation is constructed in the motion vector space of the enhanced trajectory data: through Lagrangian mechanics modeling, the change rate of the motion vector amplitude is associated with the energy decay to ensure trajectory continuity. At the same time, a material reflection energy transfer equation is constructed in the frequency domain space of the composite space feature: based on the Fresnel reflection law, a mapping relationship between the material density parameter and the frequency domain energy reflection coefficient is established. The cross-entropy of the two equations is iteratively solved by the alternating direction method of multipliers (ADMM), that is, the difference in their energy distribution probabilities is minimized, and finally a joint optimization objective function that simultaneously satisfies motion continuity and physical reflection characteristics is generated.

[0094] Step 606: Embed the energy attenuation gradient matrix of the penetrative feature layer as a regularization term into the joint optimization objective function, and construct a multi-modal joint optimization model for dynamically adjusting the contribution ratio of the visible light imaging data and the non-visible light spectrum data by constraining the band energy allocation ratio of the visible light imaging data and the non-visible light spectrum data. In this step, the regularization term refers to the constraint condition introduced to prevent model overfitting; the band energy allocation ratio refers to the contribution weights of visible light and non-visible light data in different frequency bands; the multi-modal joint optimization model refers to a fusion framework that balances the contributions of the two types of data through mathematical constraints.

[0095] In this embodiment, the energy attenuation gradient matrix of the penetrative feature layer is embedded as an L2 regularization term into the joint optimization objective function to constrain the band energy allocation ratio of visible light and non-visible light data: increase the weight of non-visible light in the significantly attenuated high-frequency region of millimeter waves, and enhance the contribution of visible light in the stable low-frequency region. Solve the constrained optimization problem through the projected gradient descent method, dynamically adjust the fusion ratio of the two types of data, and generate a robust multi-modal joint optimization model that can adapt to scene changes.

[0096] Furthermore, to solve the problem that traditional fusion methods are disjointed in band energy allocation and spatial consistency correction, resulting in artifacts in the fused image or distortion of material identification, in some embodiments, according to step 105, it includes: Step 701: Calculate the band energy allocation ratio parameters of the visible light imaging data and the non-visible light spectrum data based on the dynamic fusion constraint factor and the energy attenuation gradient matrix of the penetrative feature layer. In this step, the band energy allocation ratio parameter refers to the contribution ratio values of visible light and non-visible light data in different frequency bands generated according to the dynamic fusion constraint factor and the energy attenuation characteristics of the penetrative feature layer, and is used to quantify the fusion weight distribution of multi-modal data.

[0097] In this embodiment, first normalize the dynamic fusion constraint factor matrix and the energy attenuation gradient matrix of the penetrative feature layer, and generate initial ratio parameters through a weighted average algorithm. For the high energy attenuation region, increase the allocation ratio of the non-visible light spectrum data in combination with the intensity value of the dynamic fusion constraint factor; in the low attenuation region, reduce the non-visible light ratio and enhance the visible light weight according to the constraint factor. Finally, generate a band energy allocation ratio parameter matrix that matches the material density and motion characteristics through spatial adaptive threshold segmentation.

[0098] Step 702: Generate a visible light band energy allocation weight map and a non-visible light penetrative band energy allocation weight map based on the band energy allocation ratio parameters. In this step, the visible light band energy distribution weight map and the non-visible light penetrability band energy distribution weight map refer to the spatial weight distribution maps generated according to the band energy distribution ratio parameters, which respectively control the fusion intensity of visible light and non-visible light data in different regions. The former strengthens the retention of visual details, and the latter enhances the expression of physical characteristics.

[0099] In this embodiment, based on the band energy distribution ratio parameters, independent weight maps are generated for visible light and non-visible light data respectively: the visible light proportion parameter is converted into a weight value, and a spatially continuous distribution map is generated through bilinear interpolation. A weight close to the full value is assigned to the high proportion region, and the weight in the low proportion region approaches zero; at the same time, the non-visible light proportion parameter and the energy attenuation gradient of the penetrability feature layer are subjected to a Hadamard product operation to strengthen the weight contribution of the high attenuation region, and Gaussian filtering is used to eliminate the mutation noise to generate a smooth weight distribution.

[0100] Step 703: Convert the visible light imaging data into the frequency domain to generate visible light frequency domain data, and dynamically intercept the visible light frequency domain data according to the visible light band energy distribution weight map, retain the band components matching the frequency domain energy peak of the composite spatial feature, and generate weighted visible light frequency domain data; In this step, the visible light frequency domain data refers to the feature expression obtained by converting the visible light imaging data into the frequency domain through Fourier transform; dynamic interception means selectively retaining specific band components according to the weight map; the weighted visible light frequency domain data refers to the optimized frequency domain data generated after retaining the components matching the frequency domain energy peak of the composite spatial feature.

[0101] In this embodiment, first, a fast Fourier transform is performed on the visible light imaging data to generate frequency domain data containing amplitude and phase information. Based on the visible light band energy distribution weight map, dynamic interception is performed on each band component in the frequency domain space: the full amplitude is retained in the high weight band (such as the band corresponding to high-frequency texture), and the amplitude is attenuated according to the mapping ratio in the low weight band (such as the band corresponding to low-frequency noise). Through the frequency domain band-pass filtering technology, only the components matching the frequency domain energy peak distribution of the composite spatial feature are retained, and finally, weighted visible light frequency domain data that retains key features and suppresses noise is generated.

[0102] Step 704: Dynamically weight the energies of each decomposed band by the non-visible light penetrability band energy distribution weight map to generate penetrability compensation frequency domain data; In this step, the penetrability compensation frequency domain data refers to the optimized data generated by dynamically weighting the energies of each decomposed band by the non-visible light band energy distribution weight map, which strengthens the contribution of the bands with significant physical characteristics and suppresses the interference bands.

[0103] In this embodiment, the non-visible light spectrum data (such as millimeter wave penetration characteristics) is decomposed into frequency bands, and the energy distribution of each frequency band is extracted. Based on the penetration frequency band energy allocation weight mapping, a high weight coefficient is assigned to the region with significant high-frequency attenuation to enhance its energy amplitude; the weight coefficient is reduced for the low-frequency penetration region. The weighted frequency band energy is reconstructed into penetration compensation frequency domain data through inverse frequency domain transformation, and its high-frequency components are precisely matched with the high-frequency edges of the visible light weighted frequency domain data in the spatial dimension.

[0104] Step 705: Perform frequency band superposition on the weighted visible light frequency domain data and the penetration compensation frequency domain data to obtain a fused frequency domain distribution. In this step, the fused frequency domain distribution refers to the joint frequency domain expression generated by performing frequency band superposition on the weighted visible light frequency domain data and the penetration compensation frequency domain data. It integrates the visible light texture details and the non-visible light physical characteristics to form cross-modal complementary frequency domain features.

[0105] In this embodiment, frequency band superposition is performed on the weighted visible light frequency domain data (retaining the high-frequency edges and low-frequency contours) and the penetration compensation frequency domain data (strengthening the high-frequency reflection and low-frequency penetration characteristics): weighted average fusion dominated by visible light data is used in the low-frequency band to ensure spatial continuity; maximum value fusion with priority given to non-visible light data is implemented in the high-frequency band to enhance the sharpness of the material edges. Through smooth transition processing of the frequency band overlapping region, frequency domain jump noise is eliminated, and a fused frequency domain distribution with both visual fidelity and physical characteristics is generated.

[0106] Step 706: Perform inverse transformation processing on the fused frequency domain distribution to generate an initial fused spatial signal, and perform spatial consistency correction on the initial fused spatial signal based on the motion vector coverage area of the enhanced dynamic trajectory data. In this step, the initial fused spatial signal refers to the preliminary spatial domain fusion result generated by performing inverse Fourier transform on the fused frequency domain distribution; spatial consistency correction refers to adjusting the spatial coordinates of the fused signal through affine transformation based on the motion vector coverage area of the enhanced dynamic trajectory data to eliminate cross-modal registration errors.

[0107] In this embodiment, first, inverse Fourier transform is performed on the fused frequency domain distribution to generate an initial fused spatial signal. Based on the motion vector coverage area of the enhanced dynamic trajectory data, key feature points are extracted, and the spatial offset between them and the original visible light imaging data is calculated. A non-rigid transformation model is constructed through thin plate spline interpolation (TPS) to perform local deformation correction on the initial fused signal, ensuring spatial consistency in the key areas while retaining natural deformations.

[0108] Step 707: Perform dot product modulation on the corrected fused spatial signal and the material identification information of the composite spatial feature to generate a fused image signal containing multi-modal physical property associations. In this step, the corrected fused space signal refers to the cross-modal image data that has undergone frequency-domain fusion and spatial consistency correction; the material identification information refers to the label matrix that characterizes the target material attributes in the composite spatial features (such as the physical feature encoding of the high-density material area and the low-frequency penetration area); dot product modulation refers to the mathematical operation of multiplying two types of data pixel by pixel to enhance specific material features; the fused image signal associated with multi-modal physical characteristics refers to the final output image that simultaneously retains the visible light texture details and non-visible light material attributes, and has physical interpretability of cross-modal perception.

[0109] In this embodiment, first, perform a pixel-by-pixel dot product operation on the corrected fused space signal and the material identification information: amplify the millimeter-wave reflection characteristics through a high weight coefficient in the high-density material area, and suppress the non-visible light noise through a low weight coefficient in the low-frequency penetration area. Then, use the adaptive gamma correction algorithm to optimize the dynamic range of the modulated image, enhance the contrast in the high-density area to strengthen the edge sharpness, and maintain the natural gray distribution in the low-frequency penetration area. Finally, eliminate the high-frequency noise introduced during the modulation process through spatial filtering to generate a fused image signal that not only retains the visual authenticity of visible light but also enhances the non-visible light material characteristics.

[0110] Figure 2 The following is a schematic structural diagram of an intelligent image signal processing system based on multi-modal fusion provided by an embodiment of the present application, as Figure 2 shown. The system includes: A receiving module 21, configured to receive multi-modal input signals, where the multi-modal input signals include visible light imaging data and non-visible light spectrum data; A decomposition module 22, configured to divide the visible light imaging data into a spatial distribution feature extraction area and a dynamic change trajectory capture area, and at the same time decompose the non-visible light spectrum data into a penetration feature layer and a material reflection characteristic layer; A mapping module 23, configured to perform an association mapping between the missing pixel information in the dynamic change trajectory capture area and the energy distribution of the penetration feature layer to generate enhanced dynamic trajectory data; An identification module 24, configured to identify the reflection pattern that matches a specific material in the target scene in the material reflection characteristic layer, and perform frequency-domain superposition of the frequency band response corresponding to the reflection pattern and the spatial distribution feature extraction area to generate composite spatial features; A construction module 25, configured to construct a multi-modal joint optimization model based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial features, and use the multi-modal joint optimization model to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectrum data to generate a fused image signal.

[0111] Figure 2The described intelligent image signal processing system based on multi-modal fusion can execute Figure 1 the intelligent image signal processing method based on multi-modal fusion described in the illustrated embodiment. Its implementation principle and technical effects will not be elaborated further. For the intelligent image signal processing system based on multi-modal fusion in the above embodiment, the specific ways in which each module and unit perform operations have been described in detail in the embodiment related to this method, and will not be elaborated here.

[0112] In a possible design, Figure 2 the intelligent image signal processing system based on multi-modal fusion in the illustrated embodiment can be implemented as a computing device, such as Figure 3 as shown, this computing device may include a storage component 31 and a processing component 32; The storage component 31 stores one or more computer instructions, where the one or more computer instructions are called and executed by the processing component 32.

[0113] The processing component 32 is used for the above Figure 1 intelligent image signal processing method based on multi-modal fusion in the illustrated embodiment.

[0114] Among them, the processing component 32 may include one or more processors to execute computer instructions to complete all or part of the steps in the above method. Of course, the processing component may also be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors or other electronic components for executing the above method.

[0115] The storage component 31 is configured to store various types of data to support operations on the terminal. The storage component can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0116] Of course, the computing device will necessarily also include other components, such as input / output interfaces, display components, communication components, etc.

[0117] The input / output interface provides an interface between the processing component and the peripheral interface module, and the above peripheral interface module can be an output device, an input device, etc.

[0118] The communication component is configured to facilitate communication between the computing device and other devices in a wired or wireless manner, etc.

[0119] Among them, the computing device can be a physical device or an elastic computing host provided by a cloud computing platform, etc. At this time, the computing device can refer to a cloud server. The above processing component, storage component, etc. can be basic server resources rented or purchased from a cloud computing platform.

[0120] The embodiment of the present application also provides a computer storage medium storing a computer program, and when the computer program is executed by a computer, it can implement the above Figure 1 a method for intelligent image signal processing based on multimodal fusion shown in the embodiment.

[0121] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be repeated here.

[0122] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative labor.

[0123] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solution, or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0124] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. And these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. An intelligent image signal processing method based on multimodal fusion, characterized in that, Including: Receiving a multi-modal input signal, wherein the multi-modal input signal includes visible light imaging data and non-visible light spectrum data; Dividing the visible light imaging data into a spatial distribution feature extraction region and a dynamic change trajectory capture region, and at the same time decomposing the non-visible light spectrum data into a penetration feature layer and a material reflection characteristic layer; Associating and mapping the missing pixel information in the dynamic change trajectory capture region with the energy distribution of the penetration feature layer to generate enhanced dynamic trajectory data; Identifying a reflection pattern matching a specific material in the target scene in the material reflection characteristic layer, and performing frequency domain superposition of the frequency band response corresponding to the reflection pattern with the spatial distribution feature extraction region to generate a composite spatial feature; Based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial feature, constructing a multi-modal joint optimization model, and using the multi-modal joint optimization model to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectrum data to generate a fused image signal.

2. The method according to claim 1, wherein Identifying a reflection pattern matching a specific material in the target scene in the material reflection characteristic layer, and performing frequency domain superposition of the frequency band response corresponding to the reflection pattern with the spatial distribution feature extraction region to generate a composite spatial feature, including: Based on a preset material reflection characteristic library, performing frequency band segmentation on the material reflection characteristic layer to extract the frequency band range of the reflection pattern corresponding to a specific material in the target scene; Dynamically intercepting the energy distribution within the frequency band range of the reflection pattern through band-pass filtering, and generating a filtered reflection pattern frequency band response containing the target material density parameter according to the intercepted energy distribution; Performing a dot product operation on the filtered reflection pattern frequency band response and the frequency domain data corresponding to the spatial distribution feature extraction region to obtain a modulated spatial feature frequency domain distribution; Normalizing the modulated spatial feature frequency domain distribution based on the target material density parameter to eliminate the energy scale difference between different modalities; Performing weighted superposition of the normalized frequency domain distribution and the original spatial feature of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier.

3. The method according to claim 2, wherein Performing weighted superposition of the normalized frequency domain distribution and the original spatial feature of the spatial distribution feature extraction region to generate a composite spatial feature containing the target material identifier, including: Calculating a dynamic fusion weight based on the target material density parameter and the energy distribution in the corresponding region of the penetration feature layer; Converting the normalized frequency domain distribution into a complex domain expression form, and decomposing the original spatial feature into a real component and an imaginary component; Non-linearly scaling the amplitude of the complex domain expression form according to the dynamic fusion weight to obtain a scaled amplitude distribution, while retaining the phase information of the original spatial feature; Performing an orthogonal projection operation on the scaled amplitude distribution and the imaginary component of the original spatial feature to generate initial fusion frequency domain data; Based on the energy gradient change in adjacent regions of the penetrative feature layer, perform energy equalization compensation on the initial fused frequency-domain data to eliminate local distortions caused by multi-modal data coverage differences; Perform complex reconstruction on the compensated frequency-domain data and the real part component of the original spatial feature to generate a composite spatial feature of the target material identification.

4. The method according to claim 1, wherein Associate and map the missing pixel information in the dynamic change trajectory capture region with the energy distribution of the penetrative feature layer to generate enhanced dynamic trajectory data, including: Detect the pixel missing regions caused by dynamic blur in the dynamic change trajectory capture region, and mark the pixel missing regions as regions to be compensated; Extract the energy distribution pattern corresponding to the spatial position of the region to be compensated in the penetrative feature layer; Based on the energy distribution pattern, perform energy attenuation directivity analysis on the peripheral adjacent regions of the region to be compensated. By calculating the energy change rate of each pixel point in the penetrative feature layer along the movement trajectory direction, generate an energy attenuation gradient matrix; According to the distribution characteristics of the energy attenuation gradient matrix and the local energy peak value of the energy distribution pattern, generate dynamic trajectory compensation parameters, where the dynamic trajectory compensation parameters include a compensation intensity coefficient and a direction correction factor; Perform multi-scale interpolation on the pixel missing part of the region to be compensated according to the compensation intensity coefficient, and combine the direction correction factor to constrain the interpolation path to generate initial compensated trajectory data; Use the energy attenuation gradient matrix to dynamically correct the initial compensated trajectory data, spatially align and energy fuse the corrected compensated trajectory data with the original trajectory data in the dynamic change trajectory capture region to generate enhanced dynamic trajectory data containing cross-modal compensation information.

5. The method according to claim 4, wherein Perform multi-scale interpolation on the pixel missing part of the region to be compensated according to the compensation intensity coefficient, and combine the direction correction factor to constrain the interpolation path to generate initial compensated trajectory data, including: Perform multi-scale decomposition on the energy attenuation gradient matrix according to the preset Gaussian pyramid levels to generate a set of gradient distribution levels matching the resolution of the region to be compensated, where each level in the set of gradient distribution levels corresponds to different scale energy attenuation change characteristics; Based on the energy attenuation directivity of each level in the set of gradient distribution levels, construct a multi-scale interpolation template; According to the direction correction factor, screen the main movement direction in the set of gradient distribution levels to generate level constraint path parameters; In each level, use the weight distribution of the multi-scale interpolation template and the direction tolerance range of the level constraint path parameters to perform adaptive bilinear interpolation on the pixel missing part of the region to be compensated to generate level compensated trajectory data; Fuse the level compensated trajectory data across scales according to the Gaussian pyramid reconstruction rule to generate initial compensated trajectory data consistent with the multi-scale energy attenuation characteristics of the penetrative feature layer.

6. The method according to claim 1, characterized in that, Based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial feature, construct a multi-modal joint optimization model, including: Phase-match the timestamp sequence of the enhanced dynamic trajectory data with the energy change period of the composite spatial features to generate synchronous timing alignment parameters; Based on the synchronous timing alignment parameters, perform sliding window mutual information calculation on the motion vector amplitudes of consecutive frames in the enhanced dynamic trajectory data and the frequency domain energy peak distribution of the composite spatial features to generate modal contribution weight coefficients; Extract the spatial distribution boundary of the frequency domain energy peaks in the composite spatial features, map the spatial distribution boundary to the motion vector coverage area corresponding to the enhanced dynamic trajectory data, calculate the coverage area overlap rate between the frequency domain energy peak distribution boundary and the motion vector coverage area, and generate a spatial coupling difference matrix according to the coverage area overlap rate; According to the gradient distribution direction of the spatial coupling difference matrix, perform a multiplication operation on the modal contribution weight coefficients and the target material density parameters in the material reflection characteristic layer to generate dynamic fusion constraint factors; Based on the dynamic fusion constraint factors, construct a trajectory energy conservation equation in the motion vector space of the enhanced dynamic trajectory data and a material reflection energy transfer equation in the frequency domain space of the composite spatial features, and generate a joint optimization objective function by alternately iteratively solving the cross-entropy minimum of the trajectory energy conservation equation and the material reflection energy transfer equation; Embed the energy attenuation gradient matrix of the penetrability feature layer as a regularization term into the joint optimization objective function, and construct a multi-modal joint optimization model for dynamically adjusting the contribution ratio of the visible light imaging data and the non-visible light spectrum data by constraining the frequency band energy allocation ratio of the visible light imaging data and the non-visible light spectrum data; 7. The method according to claim 1, wherein Use the multi-modal joint optimization model to dynamically adjust the contribution ratio of the visible light imaging data and the non-visible light spectrum data to generate a fused image signal, including: Based on the dynamic fusion constraint factors and the energy attenuation gradient matrix of the penetrability feature layer, calculate the frequency band energy allocation ratio parameters of the visible light imaging data and the non-visible light spectrum data; Based on the frequency band energy allocation ratio parameters, generate a visible light frequency band energy allocation weight map and a non-visible light penetrability frequency band energy allocation weight map; Convert the visible light imaging data to the frequency domain to generate visible light frequency domain data, dynamically intercept the visible light frequency domain data according to the visible light frequency band energy allocation weight map, and retain the frequency band components matching the frequency domain energy peaks of the composite spatial features to generate weighted visible light frequency domain data; Perform frequency band decomposition on the penetrability feature layer of the non-visible light spectrum data to obtain the energies of the decomposed frequency bands, and dynamically weight the energies of the decomposed frequency bands according to the non-visible light penetrability frequency band energy allocation weight map to generate penetrability compensation frequency domain data; Perform frequency band superposition on the weighted visible light frequency domain data and the penetrability compensation frequency domain data to obtain a fused frequency domain distribution; Perform an inverse transformation on the fused frequency-domain distribution to generate an initial fused spatial signal, and perform spatial consistency correction on the initial fused spatial signal based on the motion vector coverage area of the enhanced dynamic trajectory data; Dot-multiply modulate the corrected fused spatial signal with the material identification information of the composite spatial feature to generate a fused image signal containing multi-modal physical property associations.

8. An intelligent image signal processing system based on multimodal fusion, characterized in that, Comprising: A receiving module for receiving multi-modal input signals, where the multi-modal input signals include visible light imaging data and non-visible light spectrum data; A decomposition module for dividing the visible light imaging data into a spatial distribution feature extraction area and a dynamic change trajectory capture area, and at the same time decomposing the non-visible light spectrum data into a penetrability feature layer and a material reflection property layer; A mapping module for associatively mapping the missing pixel information in the dynamic change trajectory capture area with the energy distribution of the penetrability feature layer to generate enhanced dynamic trajectory data; An identification module for identifying a reflection pattern matching a specific material in the target scene in the material reflection property layer, and performing frequency-domain superposition of the frequency band response corresponding to the reflection pattern with the spatial distribution feature extraction area to generate a composite spatial feature; A construction module for constructing a multi-modal joint optimization model based on the temporal synchronization between the enhanced dynamic trajectory data and the composite spatial feature, and dynamically adjusting the contribution ratios of the visible light imaging data and the non-visible light spectrum data by using the multi-modal joint optimization model to generate a fused image signal.

9. A computing device, characterized in that, Comprising a processing component and a storage component; the storage component stores one or more computer instructions; the one or more computer instructions are used to be called and executed by the processing component to implement a multi-modal fusion-based intelligent image signal processing method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that, Stored with a computer program, when the computer program is executed by a computer, it implements a multi-modal fusion-based intelligent image signal processing method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Satellite-borne multi-band synthetic aperture radar signal processing optimization method and system

    CN119471684A

  • Highway abnormal event detection and alarm method based on deep learning

    CN119917970A

  • Unmanned aerial vehicle video stream analysis method based on high-frequency digital information analysis

    CN120071198A

  • Three-dimensional choroidal vessel imaging and quantitative analysis method and apparatus based on optical coherence tomography system

    WO2022007352A1

Cited By

  • Electronic detonator bridge wire welding quality online detection method and system

    CN120516261A

  • Video fusion method and system based on large model

    CN120568159A

  • Multi-modal brain image intelligent feature extraction method and system based on deep learning

    CN120747705A

  • Binocular three-dimensional scanning system and method with texture mapping function

    CN120807827A

  • Image generation and editing method and system based on large model

    CN120876701A