A multimodal large model interpretation method for event streams

By constructing a conversion process from event camera data to visible light images, the target perception and interpretation problems of large multimodal models in complex scenes are solved, efficient target perception and accurate interpretation are achieved, and the input modality of large multimodal models is expanded.

CN120563563BActive Publication Date: 2025-10-03PLA PEOPLES LIBERATION ARMY OF CHINA STRATEGIC SUPPORT FORCE AEROSPACE ENG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510771386.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-10-03
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

Existing large multimodal models are unable to effectively process data captured by event cameras, resulting in poor target perception and interpretation in high dynamic range and fast motion scenarios.

Method used

Build a complete process from event camera data to event description, convert event camera data into visible light images through technologies such as Lucas-Kanade optical flow algorithm, Fourier transform, complex controllable pyramid, S transform, and high- and low-pass filters, and interpret them through a large multimodal model.

Benefits of technology

It improves target perception performance under complex lighting and motion conditions, achieves accurate interpretation of scenes that are difficult to capture with traditional visible light images, reduces computing resources and costs, and provides new visual task solutions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120563563B_ABST
    Figure CN120563563B_ABST
Patent Text Reader

Abstract

The present invention discloses a multimodal large model interpretation method for event streams, which solves the problem that multimodal large models have difficulty in perceiving and interpreting event streams, and belongs to the field of computer vision. The method comprises the following steps: reconstructing polarity event stream data into a grayscale image sequence; determining whether motion amplification is required by a Lucas-Kanade optical flow algorithm, and if so, performing Fourier transform, complex controllable pyramid, and S transform, performing denoising by a bandpass filter, and performing enhancement by a bidirectional difference method combined with amplification parameters; fusing the enhanced phase component feature subband with the amplitude component feature subband to restore the spatial resolution and obtain a reconstructed frequency domain signal; obtaining a motion-amplified grayscale image sequence by an inverse Fourier transform, and dynamically adjusting the amplification parameters by calculating normalized cross-correlation parameters; inputting the motion-amplified grayscale image sequence into a multimodal large model to output a description of the interpreted event. The present invention realizes effective perception and accurate interpretation of event streams by the multimodal large model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision technology and relates to a multimodal large model interpretation method for event streams. Background Art

[0002] Large multimodal models in the field of artificial intelligence can simultaneously process and understand multiple data types such as text, images, audio and video through pre-training combined with fine-tuning, achieve cross-modal semantic understanding and generation, as well as information fusion and reasoning, and enhance the generalization capabilities of AI systems. They have made significant progress in image and video information extraction and have shown great potential in environmental monitoring, urban planning, disaster management and other fields.

[0003] Current large multimodal models mainly rely on visible light images; however, visible light images are less sensitive to small motion changes and are prone to overexposure in high-exposure scenes; large multimodal models will suffer from motion blur and loss of important information when capturing fast-moving objects in high dynamic range scenes.

[0004] For example, the application number is: 202411826999.5, the publication number is: CN 119809925 A, and the invention name is: A method, device and medium for enhancing the visual perception ability of a multimodal model. Although this method can significantly improve the visual perception effect of a large multimodal model by integrating the advantages of different visual encoders, its image interpretation effect is poor in challenging scenarios such as strong exposure and low light.

[0005] As a new type of sensor, event cameras offer the advantages of high dynamic range, high temporal resolution, and high frame rate, making them suitable for object perception in high-speed motion and high-dynamic range scenarios. However, existing large multimodal models are unable to process the data captured by event cameras, preventing the effective utilization of this valuable information. Therefore, how to combine event camera data with large multimodal models to achieve effective perception and interpretation of complex scenes has become an urgent problem. Summary of the Invention

[0006] To address the technical challenges of existing large multimodal models in effectively perceiving and accurately interpreting scenes captured by traditional visible light imagery, as well as the technical challenges of integrating event camera data with large multimodal models to achieve effective perception and interpretation of complex scenes, this paper proposes a large multimodal model interpretation method for event streams. This method constructs a complete process from event input to event description output, expanding the input modalities of large multimodal models to enable them to process data captured by event cameras. This improves object perception performance under complex lighting conditions such as strong exposure and low light, and in complex motion scenarios such as rapid and subtle motion. A motion estimation and amplification method for event streams based on an improved PVMM is proposed. By implementing effective noise suppression and feature enhancement mechanisms, as well as real-time amplification effect evaluation and parameter adjustment mechanisms, it can accurately extract subtle motion information. By constructing a dataset containing challenging scenes such as strong exposure and rapid motion, and conducting extensive experimental validation, the results demonstrate that this method can obtain interpretation descriptions that cannot be provided by traditional visible light imagery. This method provides a new solution for vision tasks and has broad application prospects in areas such as environmental monitoring, autonomous driving, and security surveillance. This method effectively improves the performance of multimodal tasks without requiring additional training of large multimodal models, while also reducing the computing resources and costs required for model deployment.

[0007] The purpose of the present invention is specifically achieved through the following technical solutions:

[0008] The present invention discloses a multimodal large model interpretation method for event streams, the method comprising:

[0009] Step 1: Convert the polar event stream data of the target in the task scene captured by the event camera into a grayscale image sequence; calculate the optical flow velocity of the pixel intensity in the grayscale image sequence using the Lucas-Kanade optical flow algorithm, and obtain the root mean square velocity of the optical flow field from the optical flow velocity; if the root mean square velocity is less than or equal to the motion amplification threshold, execute steps 2 to 7; otherwise, execute step 7 directly;

[0010] Step 2: Perform Fourier transform on the grayscale image sequence to obtain a frequency domain complex signal that is separated into a set of amplitude components and a set of phase components; decompose the frequency domain complex signal into a low-frequency component and multi-scale and multi-directional feature subbands using a complex controllable pyramid, wherein the feature subbands are composed of amplitude component feature subbands and phase component feature subbands;

[0011] Step 3: S-transform the phase component characteristic subband to construct a two-dimensional time-frequency distribution matrix, and retain the phase component characteristic subband in the time-frequency pass domain of the two-dimensional time-frequency distribution matrix through a bandpass filter to obtain the denoised phase component characteristic subband;

[0012] Step 4: Use the differential method combined with the amplification parameter to achieve motion enhancement of the denoised phase component feature subband to obtain the enhanced phase component feature subband;

[0013] Step 5: The enhanced phase component feature subband is fused with the amplitude component feature subband one by one, and the spatial resolution is restored to the same level as the grayscale image sequence through high-pass and low-pass filters to obtain the reconstructed frequency domain signal.

[0014] Step 6: Convert the reconstructed frequency domain signal to the time domain through inverse Fourier transform to generate a motion-amplified grayscale image sequence; calculate the normalized cross-correlation parameter in real time based on the grayscale image sequence and the motion-amplified grayscale image sequence; if the normalized cross-correlation parameter is lower than the adjustment threshold, adjust the amplification parameter;

[0015] Step 7: Input the grayscale image sequence or the motion-amplified grayscale image sequence into the target multimodal large model and output the interpreted event description.

[0016] In step 1, the optical flow velocity of the pixel intensity in the grayscale image sequence is calculated using the Lucas-Kanade optical flow algorithm. The method for obtaining the root mean square velocity of the optical flow field from the optical flow velocity is:

[0017] ;

[0018] ;

[0019] Where, is the root mean square velocity, is the width of the grayscale image, is the height of the grayscale image, is the spatial coordinate of the grayscale image, is the optical flow velocity component of the optical flow vector on the horizontal x-axis, is the optical flow velocity component of the optical flow vector on the vertical y-axis; is the optical flow vector, is the time variable, i.e. the frame index of the traversal grayscale image, , is the maximum value of the time variable.

[0020] In step 2, the grayscale image sequence is subjected to Fourier transform to obtain a frequency domain complex signal that is separated into a set of amplitude components and a set of phase components:

[0021] ;

[0022] ;

[0023] Where, is the set of amplitude components, is the phase component set; is a complex signal in the frequency domain, is the frequency domain coordinate, Indicates the current moment of the target, represents the real part of the complex signal in the frequency domain, Represents the imaginary part of the complex signal in the frequency domain.

[0024] In step 2, the method of decomposing the frequency domain complex signal into low-frequency components and multi-scale and multi-directional characteristic subbands through the complex controllable pyramid is as follows:

[0025] ;

[0026] Where, is a multi-scale and multi-directional feature subband, s represents the number of scale layers, where , is the maximum number of scale layers; represents the direction angle, ; is the low-frequency component, is a complex steerable pyramid operator.

[0027] In step 3, the method of implementing S transform on the phase component characteristic subband to construct a two-dimensional time-frequency distribution matrix is:

[0028] ;

[0029] Where, For the scale layers, At the direction angle, Time and frequency The two-dimensional time-frequency distribution matrix formed; Represents the phase component characteristic subband, which is the scale layers, Frequency domain coordinates under direction angle Department, The phase component value at time , is the Gaussian window function, It is the time delay between the historical frame and the current frame; is a natural constant, is a plural unit;

[0030] The method of obtaining the denoised phase component characteristic subband by retaining the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix through a bandpass filter is as follows:

[0031] ;

[0032] Where, Represents the denoised phase component characteristic subband, which is scale layers, At the direction angle, Pass through the bandpass filter After filtering, the characteristic subband of the phase component in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. It is the inverse transform of S transform.

[0033] In step 4, the enhanced phase component characteristic sub-band includes the first enhanced phase component characteristic sub-band, the second enhanced phase component characteristic sub-band and / or the third enhanced phase component characteristic sub-band, and the calculation method includes:

[0034] S1, non-first frame And not the last frame When , the bidirectional difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the first enhanced phase component characteristic subband is obtained; wherein the calculation method of the first enhanced phase component characteristic subband is:

[0035] ;

[0036] S2, first frame When , the backward difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the second enhanced phase component characteristic subband is obtained; wherein the calculation method of the second enhanced phase component characteristic subband is:

[0037] ;

[0038] S3, last frame When , the forward difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the third enhanced phase component characteristic subband is obtained; wherein the calculation method of the third enhanced phase component characteristic subband is:

[0039] ;

[0040] Where, Represents the first enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. is the amplification parameter used to control the amplitude of motion enhancement;

[0041] Represents the second enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained;

[0042] Represents the third enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained.

[0043] In step 5, the enhanced phase component feature subband is fused with the amplitude component feature subband one by one, and the spatial resolution is restored to the same level as the grayscale image sequence through high-pass and low-pass filters. The method for reconstructing the frequency domain signal is as follows:

[0044] ;

[0045] Where, To reconstruct the frequency domain signal, is the original high-frequency component of the grayscale image sequence, is the characteristic subband of the amplitude component, Low frequency component The lowest level low-frequency component, / represents or.

[0046] In step 6, the reconstructed frequency domain signal is converted to the time domain by inverse Fourier transform to generate a motion-amplified grayscale image sequence as follows:

[0047] ;

[0048] Where, is a motion-amplified grayscale image sequence, is the inverse Fourier transform.

[0049] In step 6, the normalized cross-correlation parameter is calculated in real time based on the grayscale image sequence and the motion-magnified grayscale image sequence. If the normalized cross-correlation parameter is lower than the adjustment threshold, the method for adjusting the magnification parameter is as follows:

[0050] ;

[0051] Where, is the normalized cross-correlation parameter. The closer the normalized cross-correlation parameter value is to 1, the higher the image similarity between the grayscale image sequence and the motion-amplified grayscale image sequence is. If it is lower than the adjustment threshold, the amplification parameters are adjusted according to the preset rules; A grayscale image sequence The pixel mean, Grayscale image sequence magnified for motion The pixel mean.

[0052] In step 7, the grayscale image sequence or the motion-amplified grayscale image sequence is input into the target multimodal large model, and the method for outputting the interpreted event description is as follows:

[0053] ;

[0054] Where, To interpret the event description, is the cross-modal fusion function of the target multimodal large model, Query data for text.

[0055] The beneficial effects of the present invention are:

[0056] 1. The present invention constructs a complete process from the input of the target's polar event stream data to the output of the interpreted event description, expands the input mode of the multimodal large model, enables it to process data captured by the event camera, improves the perception performance of the target under complex lighting conditions such as strong exposure and low light, and complex motion conditions such as fast motion and small motion, and at the same time improves the accuracy of interpretation in complex scenes, realizing effective perception and precise interpretation of scenes that are difficult to capture with traditional visible light images.

[0057] 2. By constructing an effective noise suppression and feature enhancement mechanism, the denoised phase component characteristic subband and the enhanced phase component characteristic subband are obtained. At the same time, real-time evaluation and parameter adjustment of the amplification effect are performed, which can accurately extract micro-motion information and provide richer data support for subsequent analysis and decision-making.

[0058] 3. By acquiring polar event stream data of targets in the task scene captured by an event camera, a large number of experimental verifications are conducted on the obtained dataset containing challenging scenes such as strong target exposure and rapid movement. The results show that the present invention can obtain interpreted event descriptions that traditional visible light images cannot provide, providing a new solution for visual tasks and has broad application prospects in environmental monitoring, autonomous driving, security monitoring and other fields.

[0059] 4. The present invention can effectively improve the performance of multimodal tasks without the need for additional training of large multimodal models; at the same time, it reduces the computing resources and cost overhead required for model deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0060] The present invention will be described in further detail below with reference to the accompanying drawings and examples.

[0061] Figure 1 Schematic diagram of a visible light image in a strong exposure scenario provided by an example of the present invention.

[0062] Figure 2 3 is a schematic diagram of a motion-amplified grayscale image in a strong exposure scenario provided by an example of the present invention.

[0063] Figure 3 Schematic diagram of a visible light image in a low-light scenario provided by an example of the present invention.

[0064] Figure 4 Schematic diagram of a motion-amplified grayscale image in a low-light scenario provided by an example of the present invention.

[0065] Figure 5 This is a schematic diagram of a visible light image in a high dynamic range scenario provided by an example of the present invention.

[0066] Figure 6 Schematic diagram of a motion-amplified grayscale image in a high dynamic range scenario provided by an example of the present invention.

[0067] Figure 7 3 is a schematic diagram of a grayscale image of a first motion amplification in a fast motion and slight motion scene provided by an example of the present invention.

[0068] Figure 8 3 is a schematic diagram of a grayscale image of a second motion amplification in a fast motion and slight motion scene provided by an example of the present invention.

[0069] Figure 9 3 is a schematic diagram of a grayscale image of a third motion magnification in a fast motion and slight motion scene provided by an example of the present invention.

[0070] Figure 10 3 is a schematic diagram of a grayscale image of the fourth motion magnification in a fast motion and slight motion scene provided by an example of the present invention.

[0071] Figure 11 This is a schematic diagram of the description results of the interpreted events in the strong exposure scenario provided by an example of the present invention.

[0072] Figure 12 This is a schematic diagram of the description results of the interpreted events in a low-light scenario provided by an example of the present invention.

[0073] Figure 13 It is a schematic diagram of the description results of the interpreted events in fast motion and slight motion scenes provided by an example of the present invention. DETAILED DESCRIPTION

[0074] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0075] An embodiment of the present invention provides a multimodal large model interpretation method for event streams, the method comprising:

[0076] Step 1: Convert the polar event stream data of the target in the task scene captured by the event camera into a grayscale image sequence; calculate the optical flow velocity of the pixel intensity in the grayscale image sequence using the Lucas-Kanade optical flow algorithm, and obtain the root mean square velocity of the optical flow field from the optical flow velocity; if the root mean square velocity is less than or equal to the motion amplification threshold, execute steps 2 to 7; otherwise, execute step 7 directly;

[0077] Mission scenarios include challenging scenarios with a high dynamic range, such as high exposure and low light conditions, as well as high-speed and subtle motion. The event camera operates based on a photocurrent differential mode and consists of a large number of independent pixel units, each of which can independently and asynchronously sense changes in light intensity. When the light intensity change at a pixel exceeds a preset threshold, a polarity event is generated that records the time, location, and intensity change. This information captures the mission scenario, including the target's trajectory, subtle motion changes, and its state under complex lighting conditions.

[0078] A method for converting polar event stream data of targets in an event camera mission scene into a grayscale image sequence, preferably using the E2VID image reconstruction algorithm (IEEE Transactions on Pattern Analysis and Machine Intelligence, 2019, "High Speed ​​and High Dynamic Range Video with an Event Camera") to convert the polar event stream data into multiple visible light image frames, outputting a grayscale image sequence. This network structure utilizes a U-Net encoder-decoder architecture, introducing a recurrent structure of a convolutional long short-term memory model in the encoder to address temporal correlations, maintain continuity between frames, and provide a clear and reliable image foundation for subsequent image interpretation.

[0079] Step 2: Perform Fourier transform on the grayscale image sequence to obtain a frequency domain complex signal that is separated into a set of amplitude components and a set of phase components; decompose the frequency domain complex signal into a low-frequency component and multi-scale and multi-directional feature subbands using a complex controllable pyramid, wherein the feature subbands are composed of amplitude component feature subbands and phase component feature subbands;

[0080] Step 3: S-transform the phase component characteristic subband to construct a two-dimensional time-frequency distribution matrix, and retain the phase component characteristic subband in the time-frequency pass domain of the two-dimensional time-frequency distribution matrix through a bandpass filter to obtain the denoised phase component characteristic subband;

[0081] Step 4: Use the differential method combined with the amplification parameter to achieve motion enhancement of the denoised phase component feature subband to obtain the enhanced phase component feature subband;

[0082] Step 5: The enhanced phase component feature subband is fused with the amplitude component feature subband one by one, and the spatial resolution is restored to the same level as the grayscale image sequence through high-pass and low-pass filters to obtain the reconstructed frequency domain signal.

[0083] Step 6: Convert the reconstructed frequency domain signal to the time domain through inverse Fourier transform to generate a motion-amplified grayscale image sequence; calculate the normalized cross-correlation parameter in real time based on the grayscale image sequence and the motion-amplified grayscale image sequence; if the normalized cross-correlation parameter is lower than the adjustment threshold, adjust the amplification parameter;

[0084] Step 7: Input the grayscale image sequence or the motion-amplified grayscale image sequence into the target multimodal large model and output the interpreted event description.

[0085] In step 1, the optical flow velocity of the pixel intensity in the grayscale image sequence is calculated using the Lucas-Kanade optical flow algorithm. The method for obtaining the root mean square velocity of the optical flow field from the optical flow velocity is:

[0086] ;

[0087] ;

[0088] Where, is the root mean square velocity, is the width of the grayscale image, is the height of the grayscale image, is the spatial coordinate of the grayscale image, is the optical flow velocity component of the optical flow vector on the horizontal x-axis, is the optical flow velocity component of the optical flow vector on the vertical y-axis; is the optical flow vector, is the time variable, i.e. the frame index of the traversal grayscale image, , is the maximum value of the time variable.

[0089] In step 2, the grayscale image sequence is subjected to Fourier transform to obtain a frequency domain complex signal that is separated into a set of amplitude components and a set of phase components:

[0090] ;

[0091] ;

[0092] Where, is the set of amplitude components, is the phase component set; is a complex signal in the frequency domain, is the frequency domain coordinate, Indicates the current moment of the target, corresponding to the current frame index of the target in the grayscale image sequence, represents the real part of the complex signal in the frequency domain, Represents the imaginary part of the complex signal in the frequency domain.

[0093] In step 2, the method of decomposing the frequency domain complex signal into low-frequency components and multi-scale and multi-directional characteristic subbands through the complex controllable pyramid is as follows:

[0094] ;

[0095] Where, is a multi-scale and multi-directional feature subband, s represents the number of scale layers, where , is the maximum number of scale layers; represents the direction angle, ; is a low-frequency component, which can be further iteratively decomposed to obtain coarser-scale features. It is a complex controllable pyramid operator that realizes multi-scale and multi-directional decomposition.

[0096] In step 3, the method of implementing S transform on the phase component characteristic subband to construct a two-dimensional time-frequency distribution matrix is:

[0097] ;

[0098] Where, For the scale layers, At the direction angle, Time and frequency The two-dimensional time-frequency distribution matrix formed; Represents the phase component characteristic subband, which is the scale layers, Frequency domain coordinates under direction angle Department, The phase component value at time , is the Gaussian window function, It is the time delay between the historical frame and the current frame; is a natural constant, is a plural unit;

[0099] The method of obtaining the denoised phase component characteristic subband by retaining the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix through a bandpass filter is as follows:

[0100] ;

[0101] Where, Represents the denoised phase component characteristic subband, which is scale layers, At the direction angle, Pass through the bandpass filter After filtering, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained to suppress noise. It is the inverse transform of S transform.

[0102] In step 4, the enhanced phase component characteristic sub-band includes the first enhanced phase component characteristic sub-band, the second enhanced phase component characteristic sub-band and / or the third enhanced phase component characteristic sub-band, and the calculation method includes:

[0103] S1, non-first frame And not the last frame When , the bidirectional difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the first enhanced phase component characteristic subband is obtained; wherein the calculation method of the first enhanced phase component characteristic subband is:

[0104] ;

[0105] S2, first frame When , the backward difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the second enhanced phase component characteristic subband is obtained; wherein the calculation method of the second enhanced phase component characteristic subband is:

[0106] ;

[0107] S3, last frame When , the forward difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the third enhanced phase component characteristic subband is obtained; wherein the calculation method of the third enhanced phase component characteristic subband is:

[0108] ;

[0109] Where, Represents the first enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. is the amplification parameter used to control the amplitude of motion enhancement;

[0110] Represents the second enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained;

[0111] Represents the third enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained.

[0112] In step 5, the enhanced phase component feature subband is fused with the amplitude component feature subband one by one, and the spatial resolution is restored to the same level as the grayscale image sequence through high-pass and low-pass filters. The method for reconstructing the frequency domain signal is as follows:

[0113] ;

[0114] Where, To reconstruct the frequency domain signal, is the original high-frequency component of the grayscale image sequence, is the characteristic subband of the amplitude component, Low frequency component The lowest level low-frequency component, / represents or.

[0115] In step 6, the reconstructed frequency domain signal is converted to the time domain by inverse Fourier transform to generate a motion-amplified grayscale image sequence as follows:

[0116] ;

[0117] Where, is a motion-amplified grayscale image sequence, is the inverse Fourier transform.

[0118] In step 6, the normalized cross-correlation parameter is calculated in real time based on the grayscale image sequence and the motion-magnified grayscale image sequence. If the normalized cross-correlation parameter is lower than the adjustment threshold, the method for adjusting the magnification parameter is as follows:

[0119] ;

[0120] Where, is the normalized cross-correlation parameter. The closer the normalized cross-correlation parameter value is to 1, the higher the image similarity between the grayscale image sequence and the motion-amplified grayscale image sequence is. If it is lower than the adjustment threshold, the amplification parameters are adjusted according to the preset rules; A grayscale image sequence The pixel mean, Grayscale image sequence magnified for motion The pixel mean.

[0121] In step 7, the grayscale image sequence or the motion-amplified grayscale image sequence is input into the target multimodal large model, and the method for outputting the interpreted event description is as follows:

[0122] ;

[0123] Where, To interpret the event description, is the cross-modal fusion function of the target multimodal large model, is text query data; wherein, the target multimodal large model is preferably: a Transformer-based multimodal large model, such as: ViT-G / 14.

[0124] This step fuses text query data through the cross-modal fusion function of the target multimodal large model Generate interpreted event descriptions of mission scenarios , accurately mine key information in grayscale images, output interpreted event descriptions, and achieve semantic analysis under complex lighting conditions such as strong exposure and low light, and complex motion conditions such as fast motion and small motion.

[0125] In order to explain the technical solution of the present invention in detail, a specific example is now provided:

[0126] This example conducts event interpretation based on a multimodal large-scale model interpretation method for event streams. It aims to deeply explore the target perception capabilities of event cameras in challenging scenarios such as high dynamic range (including high exposure and low light) and fast and small motions. The multimodal large-scale model is used to describe the scene for visible light images and event camera reconstructed images respectively.

[0127] 1. Event Dataset: This example uses two datasets. One set of data, sourced from publicly available datasets, covers scenarios such as a "plaster dwarf" and a "ceramic mug" being shot with a rifle (with a muzzle velocity of approximately 376 m / s) and a "vehicle exiting a tunnel." These scenarios, which incorporate both rapid and subtle motion and intense exposure, provide highly targeted test samples for the example and effectively validate the performance of the research method in complex dynamic scenarios. The shooting scenario combines the rapid nanosecond motion of a bullet penetrating a target (speed > 300 m / s) with micro-damage to the target surface (millimeter-scale fractures and cracks). In addition to existing data, this example utilizes a DAVIS346 event camera to autonomously collect scene data in real-world environments, encompassing complex scenarios such as intense exposure, low light, rapid and subtle motion, and more. This aims to closely simulate the various extreme conditions likely encountered in the real world.

[0128] 2. Image reconstruction and motion amplification: Obtain polar event stream data of the target in the event camera shooting mission scene, reconstruct the polar event stream data into a grayscale image sequence using the E2VID image reconstruction algorithm, calculate the optical flow velocity of the pixel intensity in the grayscale image sequence using the Lucas-Kanade optical flow algorithm, and obtain the root mean square velocity of the optical flow field from the optical flow velocity. If the root mean square velocity is less than or equal to the motion amplification threshold, it is determined that motion amplification is required. Through steps 2 to 6 of the present invention, a motion-amplified grayscale image sequence is generated for accurate extraction of motion information.

[0129] exist Figure 1 In the strong exposure scene, the bright part of the visible light image is severely overexposed and a lot of information is lost; Figure 2 In strong exposure scenes, the grayscale image after motion amplification can clearly show the bright area, effectively avoiding the overexposure problem.

[0130] exist Figure 3 In low-light scenes, the visible light image is almost completely black due to the extremely weak light, and it is difficult to provide scene information; in sharp contrast, Figure 4 The target information is successfully captured in the motion-enlarged grayscale image.

[0131] exist Figure 5 In high dynamic range scenes, the difference in light intensity inside and outside the tunnel is huge, up to several orders of magnitude. Traditional visible light imaging systems find it difficult to capture details of both bright and dark areas in the same frame. Usually, the strong light outside the tunnel is overexposed, and key information such as road signs and vehicle details are lost. Figure 6The motion-magnified grayscale image shows that the event camera exhibits excellent adaptability. It can respond to rapidly changing lighting conditions in the scene in real time, whether it is the reflection of the road surface under strong light at the tunnel entrance or the weak light changes in the dim environment inside the tunnel, it can be recorded with high temporal resolution and high dynamic range.

[0132] Figures 7 to 10 This is a grayscale image sequence of a "plaster dwarf" being shot by a rifle, processed with motion amplification. Figure 7 shows the static state before the bullet hits the target. The dwarf statue is structurally intact and the background is clear. Figure 8 It shows the moment the bullet contacts the target and cracks appear on the plaster dwarf; Figure 9 The plaster dwarf is shown in the process of breaking apart, with the head separated from the body and a large amount of debris flying to the left; Figure 10 The image shows the process of fragments flying out at the end of a fragmentation splash. As can be seen, visible light images are prone to motion blur, resulting in blurred edges at the moment the dwarf statue shatters, making it impossible to discern the trajectory of the fragments. However, the event camera, with its high temporal resolution and high frame rate, accurately captures and amplifies the details of the dwarf's shattering moment, clearly depicting the dynamic process. Therefore, it is intuitively evident that the event camera has powerful target perception capabilities in high-exposure, low-light, fast-motion, and subtle motion scenarios.

[0133] 3. Model Input and Scene Description: To evaluate the ability of event camera-based scene reconstruction to express information in image interpretation, this example conducts a comparative experiment. A sequence of motion-amplified grayscale images is fed into a large multimodal model. This model, with its robust semantic understanding and interpretation capabilities, can describe scenes in natural language. Simultaneously, a visible light image of the same scene is fed into the model for description. By comparing the resulting descriptions, we explore whether image interpretation based on event camera-based image reconstruction can capture and present more valuable information.

[0134] The interpretation event description result after interpretation is as follows Figure 11 、 Figure 12 and Figure 13 shown. Figure 11 These are the interpretation results from a strongly exposed scene. In this scene, the visible light image is blurry and has unusual colors, severely limiting discernible information. The scene appears to be indoors, with a wooden door and the faint outline of furniture resembling a chair visible next to it. However, the event-reconstructed image, after motion amplification, clearly shows a person holding a water cup. These results clearly demonstrate the significant limitations of traditional visible light imaging in information acquisition under strongly exposed conditions. Figure 12 The interpretation results of low-light scenes are shown: in low-light environments, the visible light image is extremely dim and can only be roughly judged as an indoor environment, with details difficult to distinguish; the event reconstructed image after motion amplification shows a palm in the field of view, as if indicating the number 5. Figure 13 The interpretation results obtained under the scenes of rapid motion and micro-motion of bullet shooting are displayed. In this scene, a group of pictures depicts a small dwarf-shaped statue, which shattered after being hit by an external force, with fragments flying and dust flying. When trying to judge the direction of the external impact based on visible light images, it is difficult to draw a definite conclusion due to lack of information; however, the grayscale image with motion amplification clearly shows the fragments flying backwards, raising dust and leaving impact marks on the background cloth, vividly showing the dynamic process of the ornament being shattered by external force. Further observation found that the fragments of the ornament mainly flew to the left and back of the picture, and there were obvious impact marks on the left side of the background cloth. From this, it is inferred that the direction of the external force impact was most likely from right to left.

[0135] The beneficial effects of the embodiments of the present invention are:

[0136] 1. The present invention constructs a complete process from the input of the target's polar event stream data to the output of the interpreted event description, expands the input mode of the multimodal large model, enables it to process data captured by the event camera, improves the perception performance of the target under complex lighting conditions such as strong exposure and low light, and complex motion conditions such as fast motion and small motion, and at the same time improves the accuracy of interpretation in complex scenes, realizing effective perception and precise interpretation of scenes that are difficult to capture with traditional visible light images.

[0137] 2. By constructing an effective noise suppression and feature enhancement mechanism, the denoised phase component characteristic subband and the enhanced phase component characteristic subband are obtained. At the same time, real-time evaluation and parameter adjustment of the amplification effect are performed, which can accurately extract micro-motion information and provide richer data support for subsequent analysis and decision-making.

[0138] 3. By acquiring polar event stream data of targets in the task scene captured by an event camera, a large number of experimental verifications are conducted on the obtained dataset containing challenging scenes such as strong target exposure and rapid movement. The results show that the present invention can obtain interpreted event descriptions that traditional visible light images cannot provide, providing a new solution for visual tasks and has broad application prospects in environmental monitoring, autonomous driving, security monitoring and other fields.

[0139] 4. The present invention can effectively improve the performance of multimodal tasks without the need for additional training of large multimodal models; at the same time, it reduces the computing resources and cost overhead required for model deployment.

[0140] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any modifications or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in the present invention should be included in the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be based on the scope of protection of the claims.

Claims

1. A multimodal large model interpretation method for event streams, characterized in that: The method includes: Step 1: Convert the polar event stream data of the target in the task scene captured by the event camera into a grayscale image sequence; calculate the optical flow velocity of the pixel intensity in the grayscale image sequence using the Lucas-Kanade optical flow algorithm, and obtain the root mean square velocity of the optical flow field from the optical flow velocity; if the root mean square velocity is less than or equal to the motion amplification threshold, execute steps 2 to 7; otherwise, execute step 7 directly; Step 2: Perform Fourier transform on the grayscale image sequence to obtain a frequency domain complex signal that is separated into a set of amplitude components and a set of phase components; decompose the frequency domain complex signal into a low-frequency component and multi-scale and multi-directional feature subbands using a complex controllable pyramid, wherein the feature subbands are composed of amplitude component feature subbands and phase component feature subbands; Step 3: S-transform the phase component characteristic subband to construct a two-dimensional time-frequency distribution matrix, and retain the phase component characteristic subband in the time-frequency pass domain of the two-dimensional time-frequency distribution matrix through a bandpass filter to obtain the denoised phase component characteristic subband; Step 4: Use the differential method combined with the amplification parameter to achieve motion enhancement of the denoised phase component feature subband to obtain the enhanced phase component feature subband; Step 5: The enhanced phase component feature subband is fused with the amplitude component feature subband one by one, and the spatial resolution is restored to the same level as the grayscale image sequence through high-pass and low-pass filters to obtain the reconstructed frequency domain signal. Step 6: Convert the reconstructed frequency domain signal to the time domain through inverse Fourier transform to generate a motion-amplified grayscale image sequence; calculate the normalized cross-correlation parameter in real time based on the grayscale image sequence and the motion-amplified grayscale image sequence; if the normalized cross-correlation parameter is lower than the adjustment threshold, adjust the amplification parameter; Step 7: Input the grayscale image sequence or the motion-amplified grayscale image sequence into the target multimodal large model and output the interpreted event description.

2. The method according to claim 1, wherein In step 1, the optical flow velocity of the pixel intensity in the grayscale image sequence is calculated using the Lucas-Kanade optical flow algorithm. The method for obtaining the root mean square velocity of the optical flow field from the optical flow velocity is: ; ; Where, is the root mean square velocity, is the width of the grayscale image, is the height of the grayscale image, is the spatial coordinate of the grayscale image, is the optical flow velocity component of the optical flow vector on the horizontal x-axis, is the optical flow velocity component of the optical flow vector on the vertical y-axis; is the optical flow vector, is the time variable, i.e. the frame index of the traversal grayscale image, , is the maximum value of the time variable.

3. The method according to claim 2, wherein In step 2, the grayscale image sequence is subjected to Fourier transform to obtain a frequency domain complex signal that is separated into a set of amplitude components and a set of phase components: ; ; Where, is the set of amplitude components, is the phase component set; is a complex signal in the frequency domain, is the frequency domain coordinate, Indicates the current moment of the target, represents the real part of the complex signal in the frequency domain, Represents the imaginary part of the complex signal in the frequency domain.

4. The method according to claim 3, wherein In step 2, the method of decomposing the frequency domain complex signal into low-frequency components and multi-scale and multi-directional characteristic subbands through the complex controllable pyramid is as follows: ; Where, is a multi-scale and multi-directional feature subband, s represents the number of scale layers, where , is the maximum number of scale layers; represents the direction angle, ; is the low-frequency component, is a complex steerable pyramid operator.

5. The method according to claim 4, wherein In step 3, the method of implementing S transform on the phase component characteristic subband to construct a two-dimensional time-frequency distribution matrix is: ; Where, For the scale layers, At the direction angle, Time and frequency The two-dimensional time-frequency distribution matrix formed; Represents the phase component characteristic subband, which is the scale layers, Frequency domain coordinates under direction angle Department, The phase component value at time , is the Gaussian window function, It is the time delay between the historical frame and the current frame; is a natural constant, is a plural unit; The method of obtaining the denoised phase component characteristic subband by retaining the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix through a bandpass filter is as follows: ; Where, Represents the denoised phase component characteristic subband, which is scale layers, At the direction angle, Pass through the bandpass filter After filtering, the characteristic subband of the phase component in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. It is the inverse transform of S transform.

6. The method according to claim 5, wherein In step 4, the enhanced phase component characteristic sub-band includes the first enhanced phase component characteristic sub-band, the second enhanced phase component characteristic sub-band and / or the third enhanced phase component characteristic sub-band, and the calculation method includes: S1, non-first frame And not the last frame When , the bidirectional difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the first enhanced phase component characteristic subband is obtained; wherein the calculation method of the first enhanced phase component characteristic subband is: ; S2, first frame When , the backward difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the second enhanced phase component characteristic subband is obtained; wherein the calculation method of the second enhanced phase component characteristic subband is: ; S3, last frame When , the forward difference method is combined with the amplification parameter to realize the motion enhancement of the denoised phase component characteristic subband, and the third enhanced phase component characteristic subband is obtained; wherein the calculation method of the third enhanced phase component characteristic subband is: ; Where, Represents the first enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. is the amplification parameter used to control the amplitude of motion enhancement; Represents the second enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained; Represents the third enhanced phase component characteristic subband, which is scale layers, At the direction angle, The enhanced phase component characteristic subband at time , For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained. For the scale layers, At the direction angle, After filtering by a bandpass filter, the phase component characteristic subband in the time-frequency domain of the two-dimensional time-frequency distribution matrix is ​​retained.

7. The method according to claim 6, wherein In step 5, the enhanced phase component feature subband is fused with the amplitude component feature subband one by one, and the spatial resolution is restored to the same level as the grayscale image sequence through high-pass and low-pass filters. The method for reconstructing the frequency domain signal is as follows: ; Where, To reconstruct the frequency domain signal, is the original high-frequency component of the grayscale image sequence, is the characteristic subband of the amplitude component, Low frequency component The lowest level low-frequency component.

8. The method according to claim 7, wherein In step 6, the reconstructed frequency domain signal is converted to the time domain by inverse Fourier transform to generate a motion-amplified grayscale image sequence as follows: ; Where, is a motion-amplified grayscale image sequence, is the inverse Fourier transform.

9. The method according to claim 8, wherein In step 6, the normalized cross-correlation parameter is calculated in real time based on the grayscale image sequence and the motion-magnified grayscale image sequence. If the normalized cross-correlation parameter is lower than the adjustment threshold, the method for adjusting the magnification parameter is as follows: ; Where, is the normalized cross-correlation parameter. The closer the normalized cross-correlation parameter value is to 1, the higher the image similarity between the grayscale image sequence and the motion-amplified grayscale image sequence is. If it is lower than the adjustment threshold, the amplification parameters are adjusted according to the preset rules; A grayscale image sequence The pixel mean, Grayscale image sequence magnified for motion The pixel mean.

10. The method according to claim 9, wherein In step 7, the grayscale image sequence or the motion-amplified grayscale image sequence is input into the target multimodal large model, and the method for outputting the interpreted event description is as follows: ; Where, To interpret the event description, is the cross-modal fusion function of the target multimodal large model, Query data for text.

Citation Information

Patent Citations

  • Multi-modal model visual perception ability enhancement method and device, and medium

    CN119809925A

  • Structural vibration video measurement method and system based on derivative phase optical flow method

    CN115841504A

  • Digital signal real-time processing system based on multi-core distributed processing

    CN119316532A