A high-temporal-spatial sampling fusion method, system, terminal and storage medium

By acquiring the event stream of the event camera and the grayscale image of the visible light camera in the same time dimension and performing feature matching and affine transformation, the problems of high-resolution cameras being unable to analyze high-frequency vibrations and event cameras having difficulty in obtaining complete information of static scenes are solved, thus achieving the simultaneous monitoring of high-frequency vibration capture and submillimeter deformation analysis.

CN120259829BActive Publication Date: 2025-09-09SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510733552.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-09
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In existing technologies, high-resolution cameras cannot analyze high-frequency vibrations, and event cameras find it difficult to directly obtain complete information of static scenes, resulting in poor structural monitoring effects.

Method used

By acquiring the event stream of the event camera and the grayscale image of the visible light camera in the same time dimension, feature matching and affine transformation are performed to obtain the registered grayscale image, and the resolution is improved through event upsampling to achieve high spatiotemporal sampling fusion.

Benefits of technology

It achieves the synchronization of high-frequency vibration capture and sub-millimeter deformation analysis, improves the monitoring effect, and breaks through the technical barriers of traditional single-modal sensors in time-space-dynamic range.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120259829B_ABST
    Figure CN120259829B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of structural health monitoring, and discloses a high spatiotemporal sampling fusion method, system, terminal and storage medium, wherein the high spatiotemporal sampling fusion method comprises: obtaining an event stream, a grayscale image and a visible light image of a monitoring target in the same time dimension, wherein the event stream and the grayscale image are homologous data, and the resolution of the visible light image is higher than the resolution of the event stream; performing feature matching and affine transformation on the grayscale image and the visible light image to obtain a registered grayscale image; performing event upsampling on the event stream and the visible light image to obtain a matching event stream, wherein the matching event stream has the same resolution as the matching grayscale image; and generating monitoring data of the monitoring target based on the matching grayscale image and the matching event stream. The present application can simultaneously realize high-frequency vibration capture and compression millimeter-level deformation analysis of the monitoring target, thereby improving the monitoring effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of structural health monitoring technology, and in particular to a high-temporal-spatial sampling fusion method, system, terminal and storage medium. Background Art

[0002] With the rapid development of structural health monitoring (SHM), industrial vibration analysis, and intelligent operation and maintenance technologies, the demand for coordinated sensing of high temporal and spatial resolution in engineering structures (such as bridges, wind turbine blades, and precision instruments) is becoming increasingly urgent. Under complex operating conditions, high-frequency vibrations (such as fatigue crack initiation in composite materials and micro-resonance in bearings) require microsecond temporal resolution (≥1MHz sampling rate) to capture transient events, while micron-level deformations (such as bolt loosening displacement <0.1mm) require sub-pixel spatial resolution (≥4K pixel density) for precise measurement. However, existing technologies face a fundamental contradiction in balancing temporal and spatial resolution while accommodating dynamic range.

[0003] While traditional contact monitoring devices (such as accelerometers and laser rangefinders) can provide highly accurate data, they must be in contact with the object being measured or installed in close proximity, leading to complex deployment, high maintenance costs, and potential interference with the dynamic response of the structure. Especially in long-term, unattended monitoring, contact sensors are susceptible to failure due to environmental corrosion or mechanical vibration, significantly reducing data reliability. Video monitoring technology, with its advantages of non-contact and full-field coverage, has gradually become a research hotspot, but existing technologies still face core challenges: high-resolution cameras are limited by their frame rate and cannot resolve high-frequency vibrations; while event cameras have microsecond temporal resolution, their spatial resolution is low and their output is a sparse event stream, requiring complex reconstruction algorithms to obtain complete information directly from static scenes.

[0004] Therefore, the existing technology still needs to be improved and developed. Summary of the Invention

[0005] The main purpose of this application is to provide a high-temporal-spatial sampling fusion method, system, terminal and storage medium, aiming to solve the problem in the prior art that high-resolution cameras used in structural monitoring cannot analyze high-frequency vibrations or event cameras have difficulty in directly obtaining complete information of static scenes, resulting in poor monitoring effects.

[0006] A first aspect of an embodiment of the present application provides a high spatiotemporal sampling fusion method, which includes the following steps: obtaining an event stream, a grayscale image, and a visible light image of a monitored target in the same time dimension, wherein the event stream and the grayscale image are homologous data, and the resolution of the visible light image is higher than the resolution of the event stream; performing feature matching and affine transformation on the grayscale image and the visible light image to obtain a registered grayscale image; performing event upsampling on the event stream and the visible light image to obtain a matching event stream, wherein the matching event stream has the same resolution as the matching grayscale image; and generating monitoring data of the monitored target based on the registered grayscale image and the matching event stream.

[0007] Optionally, in one embodiment of the present application, the process of obtaining the event stream, grayscale image, and visible light image of the monitored target in the same time dimension further includes: obtaining the event timestamp of the event camera after aligning the global shutter exposure period of the visible light camera with the time window for event sampling of the event camera; and compensating for the row exposure transmission delay of the event timestamp so that the visible light camera and the event camera are synchronized in the data acquisition time dimension.

[0008] Optionally, in one embodiment of the present application, the feature matching and affine transformation are performed on the grayscale image and the visible light image to obtain a registered grayscale image, which specifically includes: extracting and matching features of the grayscale image and the visible light image to obtain feature matching point pairs; performing affine transformation on the feature matching point pairs to obtain a mapping relationship between the coordinate system of the visible light image and the coordinate system of the grayscale image; and mapping the coordinates of the visible light image to the coordinate system of the grayscale image according to the mapping relationship to obtain a registered grayscale image.

[0009] Optionally, in one embodiment of the present application, performing an affine transformation on the feature matching point pairs to obtain a mapping relationship between the coordinate system of the visible light image and the coordinate system of the grayscale image specifically includes:

[0010] Calculating a homography matrix according to the feature matching points;

[0011] The homography matrix is ​​expressed as:

[0012] ;

[0013] in, is the homography matrix, is the rotation parameter of the first row of the homography matrix, is the scaling parameter of the first row of the homography matrix, is the translation parameter of the first row element of the homography matrix, is the rotation parameter of the second row element of the homography matrix, is the scaling parameter of the second row of the homography matrix, is the translation parameter of the second row element of the homography matrix;

[0014] Minimizing a reprojection error according to the homography matrix to obtain a mapping relationship between the coordinate system of the visible light image and the coordinate system of the grayscale image;

[0015] The mapping relationship is expressed as:

[0016] ;

[0017] in, is the horizontal coordinate of the pixel point of the grayscale image. is the vertical coordinate of the pixel point of the grayscale image. is the horizontal coordinate of the pixel point of the visible light image, is the vertical coordinate of the pixel point of the visible light image.

[0018] Optionally, in one embodiment of the present application, event upsampling is performed based on the event stream and the visible light image to obtain a matching event stream, specifically including: dividing the event stream into time windows according to the period of the visible light image to obtain event frames; calculating the weight distribution of each event in the spatiotemporal domain based on the event frames to obtain three-dimensional Gaussian weights corresponding to multiple events in the event frames; reconstructing the event frames based on the three-dimensional Gaussian weights corresponding to multiple events in the event frames to obtain a matching event stream, wherein the matching event stream is consistent with the resolution of the visible light image.

[0019] Optionally, in one embodiment of the present application, the time window segmentation formula is: ;

[0020] in, For the The set of events within a time window, is a single event in the time window, 、 For the The spatial horizontal and vertical position coordinates of an event, For the The time in an event, For the The polarity of an event, is the trigger cycle;

[0021] The formula for the weight distribution is:

[0022] ;

[0023] in, is the three-dimensional Gaussian weight, 、 are the horizontal and vertical coordinates of the spatial domain, is the time domain coordinate, is the standard deviation of the spatial Gaussian kernel, is the standard deviation of the temporal Gaussian kernel.

[0024] Optionally, in one embodiment of the present application, the event upsampling is performed based on the event stream and the visible light image to generate a matching event stream corresponding to the resolution of the registered grayscale image, and then further includes: generating monitoring data of the monitoring target based on the registered grayscale image and the matching event stream.

[0025] A second aspect of the embodiments of the present application further provides a high-speed spatiotemporal sampling fusion system, wherein the high-speed spatiotemporal sampling fusion system includes:

[0026] An image acquisition module, configured to acquire an event stream, a grayscale image, and a visible light image of the monitored target in the same time dimension, wherein the event stream and the grayscale image are homologous data, and the resolution of the visible light image is higher than that of the event stream;

[0027] a spatial alignment module, configured to perform feature matching and affine transformation on the grayscale image and the visible light image to obtain a registered grayscale image;

[0028] a resolution synchronization module, configured to perform event upsampling based on the event stream and the visible light image to obtain a matching event stream, wherein the matching event stream has the same resolution as the registered grayscale image;

[0029] A data fusion module is used to generate monitoring data of the monitoring target according to the registered grayscale image and the matching event stream.

[0030] The third aspect of an embodiment of the present application also provides a terminal, wherein the terminal includes: a memory, a processor, and a high spatiotemporal sampling fusion program stored on the memory and runnable on the processor, wherein the high spatiotemporal sampling fusion program, when executed by the processor, implements the steps of the high spatiotemporal sampling fusion method described above.

[0031] The fourth aspect of an embodiment of the present application further provides a computer-readable storage medium, wherein the computer-readable storage medium stores a high spatiotemporal sampling fusion program, and when the high spatiotemporal sampling fusion program is executed by a processor, the steps of the high spatiotemporal sampling fusion method described above are implemented.

[0032] Beneficial effects: The present application provides a high-temporal-spatial sampling fusion method, system, terminal and storage medium. The present application obtains the event stream and grayscale image captured by the event camera and the visible light image captured by the visible light camera in the same time dimension, replaces the grayscale of the visible light image into the grayscale image to complete spatial alignment, and then improves the resolution of the event stream to be consistent with the resolution of the visible light image, thereby obtaining a registered grayscale image and a matching event stream with consistent resolution, that is, through the complementary fusion of the submillimeter static deformation data of the visible light image and the microsecond dynamic vibration signal of the event stream, thereby realizing the simultaneous high-frequency vibration capture and compressive millimeter deformation analysis of the monitored target, thereby improving the monitoring effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0034] Figure 1 This is a structural principle diagram of a preferred embodiment of the high spatiotemporal sampling and fusion device of this application;

[0035] Figure 2 This is a structural diagram of a preferred embodiment of the high spatiotemporal sampling and fusion device of the present application;

[0036] Figure 3 This is a flow chart of a preferred embodiment of the high spatiotemporal sampling fusion method of the present application;

[0037] Figure 4 This is a flow chart of event stream upsampling in a preferred embodiment of the high spatiotemporal sampling fusion method of the present application;

[0038] Figure 5 This is a structural diagram of a preferred embodiment of the high spatiotemporal sampling fusion system of the present application;

[0039] Figure 6 This is a structural diagram of a preferred embodiment of the terminal of this application.

[0040] Description of reference numerals:

[0041] 11. Translation stage; 12. Visible light camera; 13. Beam splitter prism; 14. Event camera;

[0042] 100. Image acquisition module; 200. Spatial alignment module; 300. Resolution synchronization module; 400. Data fusion module. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solutions and effects of this application clearer and more specific, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the drawings in the embodiments of this application. The described embodiments are only possible technical implementations of this application and are not all possible implementations. Based on the embodiments in this application, those skilled in the art can fully combine the embodiments of this application to obtain other embodiments without creative work, and these embodiments are also within the scope of protection of this application.

[0044] First, the nouns involved in the embodiments of this application are introduced:

[0045] CMOS (Complementary Metal-Oxide-Semiconductor), complementary metal oxide semiconductor, refers to the sensor technology type;

[0046] DR (Dynamic Range) refers to the range of signal strength that the sensor can process;

[0047] DAVIS (Dynamic and Active-pixel Vision Sensor) is a type of event camera and refers to a dynamic vision sensor.

[0048] RGB (Red, Green, Blue), the three primary colors of red, green and blue, refers to the three-channel color model of color images;

[0049] 4K: refers to the resolution of an image or display, usually 3840×2160 pixels;

[0050] 1000fps@4K refers to the 4K resolution video capture speed of 1000 frames per second.

[0051] In related technologies, a single camera is used to monitor structures, or multiple visible light cameras are used to fuse images for structural monitoring. The following is an introduction to visible light cameras, event cameras, and high-speed cameras in related technologies:

[0052] Visible light cameras work by capturing light intensity information across three channels (red, green, and blue) (or a single channel for grayscale information) at a fixed frame rate using a global or rolling shutter CMOS sensor. This generates high-resolution still images. Their core advantage lies in their ability to resolve spatial details, making them suitable for texture feature extraction (such as morphological analysis of surface cracks). Technical advantages include high spatial resolution, supporting sub-pixel detail resolution (e.g., pixel size ≤ 1.4μm at 4K resolution), and cost-effectiveness and ease of use. Industrial-grade visible light cameras are cost-effective and require no complex calibration, making them suitable for large-scale deployment.

[0053] However, visible light cameras have low temporal resolution and are limited by the sensor readout speed and shutter mechanism, making them unable to capture high-frequency dynamic events (such as transient crack extension under impact loads). Visible light cameras also have a limited dynamic range, with the dynamic range (DR) typically ≤70dB in global shutter mode, which can easily lead to overexposure or underexposure in areas with strong light reflections (such as metal structure surfaces) and weak light shadows.

[0054] How an event camera works: Based on an asynchronous pixel-level light intensity change detection mechanism, it only records pixels with brightness changes and their timestamps. Its temporal resolution can reach microseconds, but at the expense of spatial resolution. Technical advantages: Ultra-high temporal resolution enables the capture of high-speed dynamic events (such as bullet penetration and structural impact); low power consumption and high dynamic range: Each pixel responds independently, achieving a dynamic range (DR) exceeding 120dB.

[0055] However, event cameras have sparse data output, that is, only a discrete event stream of brightness changes is output, which lacks complete information of static scenes; motion blur and artifacts, that is, fast-moving objects easily lead to event accumulation (such as deformation areas under structural vibration), resulting in distorted reconstructed images.

[0056] Therefore, although visible light cameras support sub-pixel spatial resolution analysis, they cannot capture high-frequency vibrations. Although event cameras have microsecond temporal resolution, their spatial resolution is low, resulting in poor structural monitoring effects using the same type of camera. In addition, related technologies cannot be used in combination due to contradictions in the coordination of temporal and spatial resolution and dynamic range adaptation between visible light cameras and event cameras.

[0057] High-speed cameras operate by reducing resolution or compressing pixel depth to achieve ultra-high frame rates, often used to capture transient phenomena. Technical advantages: Ultra-high frame rates enable recording of millisecond-level dynamic processes; trigger synchronization capabilities: Support for external signal triggering, allowing precise alignment with physical events.

[0058] However, high-speed cameras have limited continuous recording time. Even with terabyte-level storage, continuous recording at 1000fps at 4K mode lasts only about one to two minutes, making it difficult to meet long-term monitoring needs. Hardware costs are high, requiring dedicated storage and processing units, and system integration is complex. Post-processing consumes significant resources: a single hour of high-speed video can contain millions of frames, making manual labeling or traditional processing algorithms inefficient and requiring AI acceleration. In other words, while high-speed cameras can capture transient events and offer sufficient resolution, they are expensive, require powerful acquisition equipment, and are not portable.

[0059] To address the problem that high-resolution cameras used in structural monitoring cannot resolve high-frequency vibrations or event cameras have difficulty in directly obtaining complete information of static scenes, resulting in poor monitoring effects, the present application obtains the event stream and grayscale image captured by the event camera and the visible light image captured by the visible light camera in the same time dimension, replaces the grayscale of the visible light image into the grayscale image to complete spatial alignment, and then improves the resolution of the event stream to be consistent with the resolution of the visible light image, thereby obtaining a registered grayscale image and a matching event stream with consistent resolution, that is, through the complementary fusion of the submillimeter-level static deformation data of the visible light image and the microsecond-level dynamic vibration signal of the event stream, thereby achieving simultaneous capture of high-frequency vibrations and analysis of compressive millimeter-level deformations of the monitored target, thereby improving the monitoring effect.

[0060] The following specific embodiments are used to describe the technical solution of the present application in detail. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments.

[0061] like Figure 1 and Figure 2 As shown, an embodiment of the present application provides a high spatiotemporal sampling fusion device, which is applied to a high spatiotemporal sampling fusion method. The high spatiotemporal sampling fusion device includes a visible light camera 12, a spectroscopic prism 13, an event camera 14 and a controller. The visible light camera 12 and the event camera 14 are coaxially arranged, and the visible light camera 12 and the event camera 14 are respectively arranged opposite to the spectroscopic prism 13; the controller includes a time registration module and a space registration module, and the time registration module and the space registration module are respectively connected to the visible light camera 12 and the event camera 14.

[0062] Specifically, the visible light camera 12, the beam splitter prism 13, and the event camera 14 are mounted on the translation stage 11 to form an optical coupling module. In the optical coupling module, a 5:5 beam splitter prism 13 is used to achieve optical path coupling. The visible light camera 12 and the DAVIS event camera 14 are coaxially mounted on the precision translation stage 11. The beam splitter prism 13 is used to achieve optical path coupling between the visible light camera 12 and the DAVIS event camera 14. The beam splitter prism 13 evenly distributes the incident light to the two cameras, ensuring consistency in their observation fields. Due to differences in physical size and imaging characteristics between the high-resolution photosensitive element of the visible light camera 12 and the asynchronous event sensor of the DAVIS event camera 14, the precision translation stage 11 is used to fine-tune the camera position to compensate for field of view offset. The coaxial mounting design combined with the mechanical adjustment mechanism allows for initial spatial alignment of the visible light image and the event stream, providing a foundation for subsequent processing.

[0063] To address the core shortcomings of existing technologies in structural health monitoring and vibration analysis, such as conflicting spatiotemporal resolution (low frame rate of high-resolution sensors and insufficient spatial resolution of high-speed sensors), limited dynamic range (poor signal-to-noise ratio in bright and low-light environments), and high system complexity (delays in time-sharing multiplexing and low mechanical turntable reliability), this application proposes a high-spatial-temporal sampling fusion method. This method utilizes heterogeneous visual fusion and a phase-adaptive algorithm to implement a non-contact monitoring strategy. Its core concept is to simultaneously capture high-frequency vibration and analyze submillimeter deformation through multimodal sensing collaboration and dynamic signal enhancement, thus overcoming the technical barriers of traditional single-modal sensors in the dimensions of time, space, and dynamic range. With heterogeneous visual fusion at its core, this application uses a 5k visible light camera to capture high-resolution static details, DAVIS to capture high-dynamic micro-vibration signals, and a dynamic range extension module to optimize full-scene imaging quality. This application eliminates the need for contact sensors or complex mechanical structures, directly enabling non-contact, all-weather dynamic sensing of structural health through visual data streams and embedded edge computing. The core value of this application lies in solving the industry problem of collaborative perception of "high-frequency vibration-submillimeter deformation", providing a non-contact, high-precision, low-cost real-time health monitoring solution for large-scale infrastructure such as bridges and wind turbine blades. It can also be expanded to precision manufacturing, aerospace and other fields, and has both technological advancement and large-scale application potential.

[0064] Compared with related technologies, the specific advantages of this application are as follows:

[0065] Breakthrough heterogeneous fusion architecture: Hardware-level optical coupling is achieved by coupling a high-spatial-resolution, low-temporal-resolution visible light camera (such as a 25-megapixel RGB or grayscale camera, 10 Hz) with a high-temporal-resolution, low-spatial-resolution event camera (such as a DAVIS camera with less than 1 megapixel, 10 MHz). Through a coaxial beam splitter prism design and the integration of cross-scale spatiotemporal registration methods, this architecture enables physical-level data homogeneity, enabling simultaneous perception of 10μm sub-pixel displacement measurement and 1ms vibration capture, solving the industry problem of "high frame rate necessarily reduces resolution, and high resolution necessarily reduces frame rate" in traditional methods.

[0066] The measurement range is wide and adaptable: in the spatial dimension, it supports cross-scale perception from cm-level overall deformation (such as bridge deflection) to micron-level local damage (such as structural microcracks); in the temporal dimension, it is compatible with quasi-static deformation (below 0.1Hz) to high-frequency vibration (above 5kHz); and it can still maintain a signal-to-noise ratio of >35dB in strong light reflections (such as metal surfaces) and weak light shadow areas.

[0067] Cost controllable: Compared with traditional high-speed cameras, the present invention uses standard industrial camera equipment and DAVIS, which reduces the cost of the system and simplifies the acquisition process and the information density of the processing.

[0068] The high spatiotemporal sampling fusion method described in the preferred embodiment of this application is as follows: Figure 3 As shown, the high spatiotemporal sampling fusion method includes the following steps:

[0069] In step S101, an event stream, a grayscale image, and a visible light image of a monitoring target in the same time dimension are obtained, wherein the event stream and the grayscale image are homologous data, and the resolution of the visible light image is higher than that of the event stream.

[0070] It's important to note that the event stream and grayscale image captured by the event camera are derived from the same source. The event stream provides information about dynamic changes in the scene, while the grayscale image (generated by the Dynamic Vision Sensor (DAVIS) in the DAVIS event camera) provides static background information about the scene. It's understandable that an event stream is a temporal sequence consisting of a series of discrete, non-uniformly distributed dynamic events, represented as a collection of events. Each dynamic event is a pixel-level brightness change captured by the DAVIS event camera. When the logarithmic brightness change at a point in the scene exceeds a set threshold, DAVIS records the event. Each dynamic event includes spatial coordinates, a timestamp, and polarity.

[0071] In one possible implementation, before the visible light camera and the event camera acquire an image, after aligning the global shutter exposure period of the visible light camera with the time window for event sampling of the event camera, an event timestamp of the event camera is obtained; and the row exposure transmission delay of the event timestamp is compensated to synchronize the visible light camera and the event camera in the data acquisition time dimension.

[0072] It should be noted that the purpose of time calibration is to eliminate cross-sensor time deviation and ensure that the event stream is strictly synchronized with the visible light image.

[0073] Specifically, the temporal registration module ensures data temporal consistency through the collaboration of hardware signals and software. At the hardware level, an external trigger box generates a synchronization pulse signal that strictly aligns the visible light camera's global shutter exposure with the DAVIS event sampling time window, eliminating the timing deviations associated with traditional time-sharing acquisition. At the software level, a timestamp-based dynamic calibration algorithm further compensates for inherent sensor delays (such as exposure transmission time), achieving microsecond-level time synchronization accuracy and ensuring a consistent time base across all modal data.

[0074] Furthermore, in the hardware triggered synchronization model, the trigger signal period is T , the exposure time of the visible light camera is T , the event sampling window of the DAVIS event camera is , DAVIS Event Camera The start time of the subsampling. Hardware synchronization ensures: , ;in, is the clock jitter, is the clock jitter standard deviation, <1 μs .

[0075] During software timestamp compensation, DAVIS event timestamps , compensates for the line exposure transmission delay : ;in, is the original timestamp of the event, is the compensated event timestamp, is the row exposure transmission delay, is the total number of DAVIS sensor rows, The row coordinates where the event occurred.

[0076] In step S102, feature matching and affine transformation are performed based on the grayscale image and the visible light image to obtain a registered grayscale image.

[0077] In one possible implementation, feature extraction and matching are performed on the grayscale image and the visible light image to obtain feature matching point pairs; an affine transformation is performed based on the feature matching point pairs to obtain a mapping relationship between the coordinate system of the visible light image and the coordinate system of the grayscale image; based on the mapping relationship, the coordinates of the visible light image are mapped to the coordinate system of the grayscale image to obtain a registered grayscale image.

[0078] It should be noted that the spatial registration module solves the problem of spatial alignment and temporal synchronization of cross-modal data. The purpose of spatial registration is to establish a spatial mapping relationship between the visible light image and the DAVIS event stream (including grayscale images and event streams), compensating for differences in sensor resolution and field of view.

[0079] In a possible implementation, a homography matrix is ​​calculated based on the feature matching points; and a reprojection error is minimized based on the homography matrix to obtain a mapping relationship between the coordinate system of the visible light image and the coordinate system of the grayscale image.

[0080] Specifically, the spatial registration process begins with feature extraction and matching. Local feature points (such as corners and edges) are extracted from the visible light image and the DAVIS grayscale image using a feature matching method. A feature scanner matching strategy is then used to calculate the similarity between the feature points. A nearest neighbor matching search is performed to filter out inliers and remove outliers to eliminate mismatches and ensure matching accuracy. It is important to note that the grayscale of the visible light image is registered with the grayscale image. Because the fields of view of the visible light image and the grayscale image do not match (they have different fields of view), image cropping is performed to ensure proper registration of the grayscale of the visible light image with the grayscale image. Next, a 3×3 homography matrix (affine transformation matrix) is calculated based on the matching point pairs. The mapping relationship is obtained by optimizing the objective by minimizing the reprojection error. The matrix parameters include rotation, scaling, translation, and shearing transformations to compensate for sensor installation errors. The high-resolution visible light image coordinates are then mapped to the low-resolution DAVIS coordinate system using the homography matrix to achieve spatial alignment. In other words, the grayscale of the visible light image is replaced with the grayscale of the DAVIS grayscale image to complete the grayscale registration.

[0081] Furthermore, in the physical model establishment, the visible light camera imaging model is assumed to be:

[0082] ;

[0083] in, The pixel coordinates of the visible light camera image, representing a three-dimensional space point projection in visible light images; is the intrinsic parameter matrix of the visible light camera, is the external parameter matrix of the visible light camera, which represents the rotation between the camera coordinate system and the world coordinate system ( ) and translation ( )relation; is the coordinate of a point in three-dimensional space.

[0084] The DAVIS event camera model is:

[0085] ;

[0086] ;

[0087] in, is the pixel coordinate of the DAVIS event camera, representing a three-dimensional space point Projection in the DAVIS event stream; is the intrinsic parameter matrix of the DAVIS event camera, is the internal parameter matrix; is the external parameter matrix of the DAVIS event camera, which represents the rotation between the event camera coordinate system and the world coordinate system ( ) and translation ( )relation; is the event noise term, which represents sensor noise or model error; 、 is the horizontal and vertical focal length, indicating the distance from the camera lens to the imaging plane; 、 are the horizontal and vertical principal point coordinates, indicating the projection position of the optical axis on the imaging plane.

[0088] Solve the affine transformation and perform image registration using feature matching to obtain the homography matrix between the two sensors , the homography matrix is ​​expressed as:

[0089] ;

[0090] in, is the homography matrix, is the rotation parameter of the first row of the homography matrix, is the scaling parameter of the first row of the homography matrix, is the translation parameter of the first row element of the homography matrix, is the rotation parameter of the second row element of the homography matrix, is the scaling parameter of the second row of the homography matrix, is the translation parameter of the second row of the homography matrix.

[0091] Then the mapping relationship can be obtained, which is expressed as:

[0092] ;

[0093] in, is the horizontal coordinate of the pixel point of the grayscale image. is the vertical coordinate of the pixel point of the grayscale image. is the horizontal coordinate of the pixel point of the visible light image, is the vertical coordinate of the pixel point of the visible light image.

[0094] In existing technologies, sensors with high spatial resolution (visible light cameras) and high temporal resolution (DAVIS) typically operate independently, limited by the conflict between spatial and temporal resolution. High-frame-rate sensors (DAVIS) suffer from insufficient spatial resolution due to pixel size limitations, while high-resolution sensors (5K cameras) struggle to achieve a higher frame rate (30fps) due to bandwidth constraints. High-speed cameras, on the other hand, are prohibitively expensive and lack portability. This application utilizes heterogeneous data cross-scale fusion. Through a heterogeneous data fusion architecture, the submillimeter static deformation data from a visible light camera is complementarily fused with the microsecond dynamic vibration signals from a DAVIS event camera, surpassing the performance limits of traditional single-modal sensors in both spatial and temporal dimensions. This approach enables cross-scale perception: global high-resolution deformation monitoring and local high-frequency vibration capture are simultaneously achieved. Furthermore, the dynamic range is enhanced: by fusing the wide dynamic range of visible light with the high-frequency response of DAVIS, it covers the full range of monitoring scenarios, from static deformation to transient impacts.

[0095] In step S103, event upsampling is performed according to the event stream and the visible light image to obtain a matching event stream, wherein the matching event stream has the same resolution as the registered grayscale image.

[0096] In one possible implementation, the event stream is divided into time windows according to the period of the visible light image to obtain event frames; the weight distribution of each event in the spatiotemporal domain is calculated based on the event frames to obtain three-dimensional Gaussian weights corresponding to multiple events in the event frames; and the event frames are reconstructed based on the three-dimensional Gaussian weights corresponding to multiple events in the event frames to obtain a matching event stream, wherein the matching event stream is consistent with the resolution of the visible light image.

[0097] It should be noted that the purpose of event upsampling is to improve the spatial resolution of the DAVIS event stream to match the level of detail of the visible light image.

[0098] Specifically, such as Figure 4 As shown in the figure, event upsampling exploits the homology between the DAVIS event stream and the grayscale image. First, the event stream is time-windowed to generate event frames. Specifically, the continuous event stream is segmented into discrete time windows at a fixed period, with each window containing multiple events, forming event frames. Upsampling is then performed, and high-precision registration is achieved using 3D Gaussian kernel modeling. Finally, at each pixel location, the Gaussian weights of all events within the window are accumulated to output a high-resolution event frame with the same spatial resolution as the visible light image. In other words, upsampling the event stream raises its resolution to match that of the visible light camera.

[0099] Furthermore, let the trigger period of DAVIS be , split the DAVIS event stream into the same time window. The time window splitting formula is: ;

[0100] in, For the The set of events in a time window, each time window Include events, is a single event in the time window, 、 For the The spatial horizontal and vertical position coordinates of an event, For the The time in an event, For the The polarity of an event, is the trigger cycle;

[0101] Three-dimensional Gaussian kernel modeling is to model each event ∈ , calculate its weight distribution in the spatiotemporal domain (without considering polarity), the formula for the weight distribution is:

[0102] ;

[0103] in, is the three-dimensional Gaussian weight; 、 are the horizontal and vertical coordinates of the spatial domain; is the time domain coordinate; is the standard deviation of the spatial Gaussian kernel, which controls the spread of events in space; is the standard deviation of the temporal Gaussian kernel, which controls the spread of events in time.

[0104] Specifically, during the time calibration process, the hardware trigger signal and event timestamp are combined to build a time offset compensation model to ensure that the data stream is strictly aligned in the time dimension.

[0105] In existing technologies, multi-sensor registration usually relies on single-modality feature matching (such as grayscale or edge-based registration), which makes it difficult to solve the spatiotemporal alignment problem across modalities (such as visible light images and event streams) with significant differences in resolution. In this application, a spatiotemporal synchronous registration scheme, namely a cross-modal three-level registration scheme, is adopted: first, preliminary spatial coarse registration: based on the feature matching of the visible light image and the grayscale image of the DAVIS camera, an initial spatial mapping relationship is established. Since the grayscale pixels and event pixels of the DAVIS camera are homologous, the visible light image space is mapped to the grayscale image space of the DAVIS camera to ensure the rough registration of the event stream and the visible light image; second, dynamic spatial fine registration: a dense 3D point cloud is generated by spatiotemporal upsampling of the event stream, and spatiotemporal Gaussian kernel interpolation is used to optimize the sub-pixel alignment of the visible light image and DAVIS; third, timing synchronization calibration: a hardware-level trigger synchronization architecture is adopted, and the microsecond-level time base alignment of the visible light camera and the DAVIS event camera is achieved through the precise trigger signal generated by the trigger circuit. The exposure delay is eliminated in combination with the timestamp calibration circuit to build a unified time base for the entire system. This application can generate a dense motion field that matches the resolution of the visible light camera through event stream spatiotemporal upsampling technology for a resolution difference of up to 15 times between visible light images (5496×3672) and DAVIS event streams (346×240), achieving high-precision alignment and breaking through the technical bottleneck of traditional methods that are difficult to align when the resolution difference is greater than 3 times.

[0106] In step S104, after obtaining the registered grayscale image and the matching event stream, data fusion is performed to generate monitoring data of the monitoring target based on the registered grayscale image and the matching event stream.

[0107] Specifically, the controller also includes a data fusion module, which integrates the static texture details of visible light and the dynamic response characteristics of the event stream, and realizes cross-modal feature fusion through a dynamic weight allocation algorithm.

[0108] In one possible implementation, a static fusion weight and a dynamic fusion weight are obtained; a fusion value of each pixel position is calculated based on the static fusion weight, the dynamic fusion weight, the registered grayscale image and the matching event stream; and monitoring data of the monitoring target is obtained based on the fusion values ​​of all pixel positions.

[0109] It can be understood that fusion weights are generated based on feature importance (such as deformation significance and vibration energy); pixel-level fusion calculates the fusion value for each pixel position; spatiotemporal consistency optimization is combined with the time offset compensation model output by the time calibration module to ensure the continuity of the fused data stream in the spatiotemporal domain.

[0110] Furthermore, static feature enhancement aims to leverage the high-resolution background information of visible light images to locate and analyze structural deformations. During this process, visible light images are preprocessed to identify regions of structural deformation, and static feature mapping is performed to encode this structural deformation information into feature vectors. Dynamic feature extraction aims to capture high-frequency vibration signals through event streams and enhance micron-level displacement details. During dynamic feature extraction, the event stream captures high-frequency vibration signals and enhances micro-displacement details through time-domain amplification techniques. During the fusion output process, an attention mechanism is incorporated to dynamically adjust the contribution weights of different modalities, generating a composite data stream (i.e., monitoring data) with both high spatial and temporal resolution, enabling precise analysis from overall vibration patterns to localized micron-level defects.

[0111] In another embodiment of the present application, the optical components are adjusted, and the beam splitter prism can be replaced with a dichroic mirror or a polarization beam splitter to adapt to specific bands (such as near-infrared) or polarization-sensitive scenarios; a low-cost version can use a fixed bracket instead of a precision translation stage, and compensate for the field of view offset through software algorithms.

[0112] In another embodiment of the present application, the synchronization mechanism is optimized, and the hardware trigger signal can be expanded to multi-node cascade to support nanosecond-level synchronization of distributed monitoring networks; the software level supports NTP / PTP protocol (NTP is the Network Time Protocol, PTP is the Precision Time Protocol) supplementary calibration to improve robustness in complex environments.

[0113] In another embodiment of the present application, an algorithm variation is performed, and dense optical flow calculation is used instead of sparse feature matching and linear interpolation to increase the computational load and improve partial accuracy; an end-to-end neural network is constructed to directly input raw data and output fusion results to improve the level of automation.

[0114] In another embodiment of the present application, cross-modal expansion is performed, and an infrared camera or lidar is integrated to realize multimodal fusion monitoring of thermal-mechanical coupling or three-dimensional deformation; combined with a wireless transmission module, it supports drone-mounted and mobile inspection scenarios.

[0115] The technical advantages of this application are reflected in the following aspects:

[0116] High spatiotemporal resolution synergy: The submillimeter spatial resolution of the visible light camera (static deformation analysis) and the kilohertz temporal resolution of DAVIS (dynamic vibration capture) enable high-precision monitoring of the entire area.

[0117] Non-contact measurement: No need to touch or mark the object being measured, avoiding interference with the structure caused by traditional sensor installation. It is particularly suitable for sensitive scenes such as bridges and aerospace vehicles.

[0118] Dynamic adaptability: Maintaining stable data collection and processing capabilities under complex lighting (strong light / weak light) and fast-moving conditions;

[0119] Real-time: The system can collect and process data in real time, quickly providing feedback on vibration and displacement information. It is suitable for scenarios requiring real-time monitoring, such as vibration monitoring of bridges, buildings, and equipment.

[0120] Multi-scenario expansion: Supports full-scale monitoring from large-scale infrastructure (such as wind turbine blades) to precision equipment (such as chip packaging machines), covering various working conditions such as static deformation, high-frequency vibration, and transient impact.

[0121] Next, the high spatiotemporal sampling fusion system proposed according to the embodiments of the present application will be described with reference to the accompanying drawings.

[0122] Figure 5 It is a structural diagram of the high spatiotemporal sampling fusion system of an embodiment of the present application.

[0123] like Figure 5 As shown, the high spatiotemporal sampling fusion system includes: an image acquisition module 100 , a spatial alignment module 200 , a resolution synchronization module 300 and a data fusion module 400 .

[0124] Specifically, the image acquisition module 100 is used to acquire an event stream, a grayscale image, and a visible light image of the monitored target in the same time dimension, wherein the event stream and the grayscale image are homologous data, and the resolution of the visible light image is higher than that of the event stream;

[0125] A spatial alignment module 200 is configured to perform feature matching and affine transformation on the grayscale image and the visible light image to obtain a registered grayscale image;

[0126] a resolution synchronization module 300 for performing event upsampling based on the event stream and the visible light image to obtain a matching event stream, wherein the matching event stream has the same resolution as the registered grayscale image;

[0127] The data fusion module 400 is configured to generate monitoring data of the monitoring target according to the registered grayscale image and the matching event stream.

[0128] Figure 6 This is a diagram of the structure of a terminal provided in an embodiment of the present application. The terminal may include:

[0129] Memory 501 , processor 502 , and computer programs stored in the memory 501 and executable on the processor 502 .

[0130] When the processor 502 executes the program, the high spatiotemporal sampling fusion method provided in the above embodiment is implemented.

[0131] Furthermore, the terminal further includes:

[0132] The communication interface 503 is used for communication between the memory 501 and the processor 502 .

[0133] The memory 501 is used to store computer programs that can be run on the processor 502 .

[0134] The memory 501 may include a high-speed RAM memory, and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0135] If the memory 501, processor 502, and communication interface 503 are implemented independently, the communication interface 503, memory 501, and processor 502 can be connected to each other via a bus and communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EIS) bus. Buses can be divided into address buses, data buses, control buses, etc. For ease of representation, Figure 6 Only one thick line is used in the diagram, but this does not mean that there is only one bus or one type of bus.

[0136] Optionally, in a specific implementation, if the memory 501, the processor 502 and the communication interface 503 are integrated on a chip, the memory 501, the processor 502 and the communication interface 503 can communicate with each other through an internal interface.

[0137] The processor 502 may be a central processing unit (CPU), an application specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of the present application.

[0138] This embodiment also provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the high spatiotemporal sampling fusion method as described above is implemented.

[0139] One embodiment of the present application provides a computer program product, including a computer program, which, when executed by a processor, implements the Figure 3 The high spatiotemporal sampling fusion method provided in any embodiment of the corresponding embodiment.

[0140] In the description of this specification, the description with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials or characteristics described in conjunction with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, those skilled in the art can combine and combine different embodiments or examples described in this specification and features of different embodiments or examples without contradiction.

[0141] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be understood to indicate or imply relative importance or implicitly specify the number of technical features indicated. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one such feature. In the description of this application, "N" means at least two, for example, two, three, etc., unless otherwise specifically defined.

[0142] Any process or method description in a flowchart or otherwise described herein may be understood to represent a module, fragment or portion of code comprising one or N executable instructions for implementing a custom logical function or process step, and the scope of the preferred embodiments of the present application includes alternative implementations in which functions may be performed in a different order than shown or discussed, including performing functions in a substantially simultaneous manner or in a reverse order depending on the functions involved, which should be understood by those skilled in the art to which the embodiments of the present application pertain.

[0143] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable storage medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable storage medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (not exhaustive) of computer-readable storage media include the following: an electrical connection with one or more wires (electronic devices), a portable computer disk cartridge (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and programmable read-only memory (EPROM or flash memory), fiber optic devices, and a portable compact disc read-only memory (CDROM). In addition, the computer-readable storage medium may even be paper or other suitable medium on which the program is printed, since the program can be obtained electronically by optically scanning the paper or other medium and then editing, interpreting or processing it in other suitable ways as necessary, and then storing it in a computer memory.

[0144] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiment, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented using hardware, as in another embodiment, any one of the following technologies known in the art or a combination thereof can be used: a discrete logic circuit having logic gate circuits for implementing logical functions on data signals, an application-specific integrated circuit having suitable combinational logic gate circuits, a programmable gate array (PGA), a field-programmable gate array (FPGA), etc.

[0145] Those skilled in the art will understand that all or part of the steps in the method of the above embodiment can be completed by instructing related hardware through a program, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiment.

[0146] In addition, the functional units in the various embodiments of the present application may be integrated into a processing module, or each unit may exist physically separately, or two or more units may be integrated into a module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.

[0147] The storage medium mentioned above may be a read-only memory, a magnetic disk, or an optical disk, etc. Although the embodiments of the present application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present application. Persons skilled in the art may make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

[0148] It should be understood that the application of this application is not limited to the above examples. For ordinary technicians in this field, they can make improvements or changes based on the above description. All these improvements and changes should fall within the scope of protection of the claims attached to this application.

[0149] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some or all of the technical features therein. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A high spatiotemporal sampling fusion method, characterized in that: The high spatiotemporal sampling fusion method includes: Acquire an event stream, a grayscale image, and a visible light image of the monitored target in the same time dimension, wherein the event stream and the grayscale image are homologous data, and the resolution of the visible light image is higher than that of the event stream; Performing feature matching and affine transformation on the grayscale image and the visible light image to obtain a registered grayscale image; Performing event upsampling on the event stream and the visible light image to obtain a matching event stream, wherein the matching event stream has the same resolution as the registered grayscale image; generating monitoring data of the monitoring target according to the registered grayscale image and the matching event stream; The process of acquiring the event stream, grayscale image, and visible light image of the monitoring target in the same time dimension also includes: After aligning the global shutter exposure period of the visible light camera with the time window for event sampling of the event camera, obtaining the event timestamp of the event camera; Compensating for a row exposure transmission delay of the event timestamp so that the visible light camera and the event camera are synchronized in a data acquisition time dimension; The performing event upsampling according to the event stream and the visible light image to obtain a matching event stream specifically includes: Dividing the event stream into time windows according to the period of the visible light image to obtain event frames; Calculating the weight distribution of each event in the spatiotemporal domain according to the event frame to obtain three-dimensional Gaussian weights corresponding to the multiple events in the event frame; reconstructing the event frame according to three-dimensional Gaussian weights corresponding to a plurality of events in the event frame to obtain a matching event stream, wherein the matching event stream has a resolution consistent with that of the visible light image; The formula for splitting the time window is: ; in, For the The set of events within a time window, is a single event in the time window, 、 For the The spatial horizontal and vertical position coordinates of an event, For the The time in an event, For the The polarity of an event, is the trigger cycle; The formula for the weight distribution is: ; in, is the three-dimensional Gaussian weight, 、 are the horizontal and vertical coordinates of the spatial domain, is the time domain coordinate, is the standard deviation of the spatial Gaussian kernel, is the standard deviation of the temporal Gaussian kernel.

2. The high spatiotemporal sampling fusion method according to claim 1, characterized in that: The performing feature matching and affine transformation on the grayscale image and the visible light image to obtain a registered grayscale image specifically includes: Performing feature extraction and matching on the grayscale image and the visible light image to obtain feature matching point pairs; Performing an affine transformation on the feature matching point pairs to obtain a mapping relationship between the coordinate system of the visible light image and the coordinate system of the grayscale image; According to the mapping relationship, the coordinates of the visible light image are mapped to the coordinate system of the grayscale image to obtain a registered grayscale image.

3. The high spatiotemporal sampling fusion method according to claim 2, characterized in that: The performing affine transformation on the feature matching point pairs to obtain a mapping relationship between the coordinate system of the visible light image and the coordinate system of the grayscale image specifically includes: Calculating a homography matrix according to the feature matching points; The homography matrix is ​​expressed as: ; in, is the homography matrix, is the rotation parameter of the first row of the homography matrix, is the scaling parameter of the first row of the homography matrix, is the translation parameter of the first row element of the homography matrix, is the rotation parameter of the second row element of the homography matrix, is the scaling parameter of the second row of the homography matrix, is the translation parameter of the second row element of the homography matrix; Minimizing a reprojection error according to the homography matrix to obtain a mapping relationship between the coordinate system of the visible light image and the coordinate system of the grayscale image; The mapping relationship is expressed as: ; in, is the horizontal coordinate of the pixel point of the grayscale image. is the vertical coordinate of the pixel point of the grayscale image. is the horizontal coordinate of the pixel point of the visible light image, is the vertical coordinate of the pixel point of the visible light image.

4. The high spatiotemporal sampling fusion method according to any one of claims 1 to 3, characterized in that: Generating the monitoring data of the monitoring target according to the registered grayscale image and the matching event stream specifically includes: Obtain static fusion weight and dynamic fusion weight; Calculating a fusion value for each pixel position according to the static fusion weight, the dynamic fusion weight, the registered grayscale image, and the matching event stream; The monitoring data of the monitoring target is obtained according to the fusion values ​​of all pixel positions.

5. A high spatiotemporal sampling fusion system, characterized in that: The high-temporal-spatial sampling fusion system is applied to the high-temporal-spatial sampling fusion method according to any one of claims 1 to 4; The high spatiotemporal sampling fusion system includes: An image acquisition module, configured to acquire an event stream, a grayscale image, and a visible light image of the monitored target in the same time dimension, wherein the event stream and the grayscale image are homologous data, and the resolution of the visible light image is higher than that of the event stream; a spatial alignment module, configured to perform feature matching and affine transformation on the grayscale image and the visible light image to obtain a registered grayscale image; a resolution synchronization module, configured to perform event upsampling based on the event stream and the visible light image to obtain a matching event stream, wherein the matching event stream has the same resolution as the registered grayscale image; A data fusion module is used to generate monitoring data of the monitoring target according to the registered grayscale image and the matching event stream.

6. A terminal, characterized in that: The terminal includes: a memory, a processor, and a high spatiotemporal sampling fusion program stored in the memory and runnable on the processor. When the high spatiotemporal sampling fusion program is executed by the processor, the steps of the high spatiotemporal sampling fusion method according to any one of claims 1 to 4 are implemented.

7. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a high spatiotemporal sampling fusion program, which, when executed by a processor, implements the steps of the high spatiotemporal sampling fusion method according to any one of claims 1 to 4.

Citation Information

Patent Citations

  • Dual-camera space-time fusion method based on embedded terminal

    CN111898625A

  • Optical flow estimation method and system based on asynchronous event flow and grayscale image fusion

    CN113269699A