An event stream cross-modal point process alignment method and system based on a pulse network
Patent Information
- Application Number
- CN202611040145.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2046-07-14
AI Technical Summary
[0008]针对现有技术的以上缺陷或改进需求,本发明提供了一种基于脉冲网络的事件流跨模态点过程对齐方法及系统,由此解决将事件流强行连续化后嵌入连续向量空间进行度量,导致事件流的高时间分辨率信息和天然稀疏性被大量浪费,以致精度和效率受到限制的技术问题
通过将事件流建模为离散时空点过程,通过脉冲神经网络提取事件流特征后,不将其压缩为稠密向量,而是在稀疏观测空间上以点过程能量作为跨模态对齐度量,避免了事件流的强制连续化所导致的时序坍缩问题,保留了异步脉冲的精细时序结构;同时,以点过程能量替代余弦相似度作为度量标准,使度量空间与事件流作为离散点过程的物理本质相匹配,从根本上解决了现有方法度量对象错位的问题,从而能在不牺牲事件稀疏时序特性的前提下,实现与连续语义模态的有效对齐;同时,通过稀疏编码获得各时刻的事件流稀疏观测表示,充分利用了事件流自身稀疏、关键信息集中于少数高激活区域的特点,提升了计算效率,也确保了跨模态对齐的稳定性和可靠性。
Smart Images

Figure CN122595063B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of cross-modal retrieval and alignment, and more specifically, relates to a method and system for aligning cross-modal point processes of event streams based on pulse networks. Background Technology
[0002] An event camera is a bio-inspired neuromorphic visual sensor that operates on a completely different principle than a traditional camera: traditional cameras synchronously acquire global brightness information at a fixed frame rate, while event cameras asynchronously respond to brightness changes at each pixel, outputting an "event stream" consisting of microsecond-level timestamps, pixel coordinates, and polarity. The event stream is essentially a discrete spatiotemporal point process, naturally possessing the characteristics of sparseness, asynchronicity, and high temporal resolution.
[0003] With the rapid development of artificial intelligence technology, the need for cross-modal alignment of event streams with natural language (text) is becoming increasingly urgent. For example, in scenarios such as drone reconnaissance, autonomous driving, and intelligent surveillance, users want to directly retrieve relevant event fragments from the event stream database using natural language queries (such as "a red car drove past from left to right"). The prerequisite for achieving this goal is to establish an alignment model that can accurately measure the correspondence between text semantics and event stream content.
[0004] Currently, most cross-modal alignment methods for event streams and text follow a "continuous + contrastive learning" approach: first, asynchronous event streams are forcibly converted into two-dimensional image frames or three-dimensional voxel grids through voxelization, time slicing, frame aggregation, etc.; then, features are extracted using pre-trained visual-language models (such as CLIP); and finally, cosine similarity or InfoNCE loss is used for measurement and optimization in a continuous vector space. This approach has the following fundamental flaws: First, there's the issue of temporal collapse. The core value of event streams lies in their microsecond-level asynchronous temporal information. However, regardless of whether voxelization, frame aggregation, or time pooling is used, asynchronous pulses are forcibly squeezed into a uniform discrete time window, causing the fine temporal structure between events to be flattened. This "continuous preprocessing" is physically equivalent to using continuous geometry to constrain discrete-point processes, resulting in a fundamental misalignment of the measurement object.
[0005] Second, there is a mismatch in metric spaces. Text, after being encoded by a language model, forms a continuous semantic manifold, whose natural metric is Euclidean distance or cosine similarity; while an event stream is essentially a spatiotemporal point set, whose natural object is an intensity function, and whose natural metric is the negative log-likelihood of a point process. Its core information lies in the spatiotemporal distribution pattern of events. Forcibly embedding a discrete spatiotemporal distribution into a continuous vector space and calculating vector distance inherently results in a metric mismatch in mathematical terms.
[0006] Third, the sparsity mechanism is not fully utilized. The key semantic information in an event stream is often carried by a few highly activated regions, but existing methods uniformly compress the sparse event stream into dense features during the continuation process, thus losing the computational efficiency improvement brought by sparse representation.
[0007] In general, existing methods forcibly convert event streams into continuous data and embed them into a continuous vector space for measurement. This results in a significant waste of the high temporal resolution information and natural sparsity of the event streams, thus limiting accuracy and efficiency. Furthermore, the measurement method is mathematically misaligned with the discrete spatiotemporal distribution characteristics of the data. How to achieve effective alignment with continuous semantic modalities without sacrificing the high temporal resolution and sparsity of events is a pressing problem that needs to be solved. Summary of the Invention
[0008] To address the above-mentioned deficiencies or improvement needs of existing technologies, this invention provides a method and system for aligning event streams across modal points based on pulse networks. This solves the technical problem that forcibly making event streams continuous and embedding them into a continuous vector space for measurement results in a significant waste of the high temporal resolution information and natural sparsity of the event streams, thus limiting accuracy and efficiency.
[0009] To achieve the above objectives, according to a first aspect of the present invention, a method for aligning event streams across modal points based on pulse networks is provided, comprising: After voxelization preprocessing, the candidate event stream output by the event camera is split into positive event stream mesh and negative event stream mesh according to the polarity channel; The positive event flow grid and the negative event flow grid are encoded using two spiking neural networks respectively, and the positive event flow features and negative event flow features obtained at each time step are concatenated to obtain the event flow features at each time step. Sparse coding is performed on the positive event stream features and negative event stream features at each time step, and the sparse coding of the positive event stream and negative event stream at each time step is concatenated to obtain the sparse observation representation of the event stream at each time step; Semantic vectors are extracted from text queries to obtain semantic vectors to be aligned. These semantic vectors are then interacted with event flow features at each time step. The semantic modulation components obtained from the interaction at each time step are combined with the basic intensity components obtained from the event flow features at the same time step to obtain the logarithm of the conditional intensity function at each time step. After flattening the logarithm of the conditional intensity function at each time point and the sparse observation representation of the event stream according to the time and channel dimensions to obtain a one-dimensional logarithmic vector and a one-dimensional sparse vector, the index of the element with a value of 1 in the one-dimensional sparse vector is determined as the activation position. The reward energy is calculated based on the elements of the activation positions in the one-dimensional logarithmic vector, and the penalty energy is calculated based on the one-dimensional logarithmic vector. The difference between the reward energy and the penalty energy is calculated as the point process energy. The larger the element of the activation position in the one-dimensional logarithmic vector, the larger the reward energy; the larger the average value of the elements in the one-dimensional logarithmic vector, the larger the penalty energy. The degree of matching between the candidate event stream and the text query is determined based on the point process energy. If the point process energy is obtained by subtracting the penalty energy from the reward energy, then the greater the point process energy, the greater the degree of matching between the candidate event stream and the text query; otherwise, the smaller the point process energy, the greater the degree of matching between the candidate event stream and the text query.
[0010] Based on the above-mentioned event flow cross-modal point process alignment method based on pulse networks, sparse encoding is performed on the positive and negative event flow features at time t, specifically including: The positive event flow features and negative event flow features at time t are scored using a score selection network to obtain the scores of each channel in the positive and negative event flow features at time t. The scores of each channel in the positive event stream feature are sorted in descending order, and a first threshold is selected from the scores of each channel after sorting based on the sparsity ratio, so as to perform sparse encoding on the positive event stream feature based on the first threshold. The scores of each channel in the negative event stream feature are sorted in descending order, and a second threshold is selected from the scores of each channel after sorting based on the sparsity ratio, so as to perform sparse encoding on the negative event stream feature based on the second threshold.
[0011] According to the above-mentioned event flow cross-modal point process alignment method based on pulse networks, the positive event flow features are sparsely encoded based on the first threshold, specifically including: Set the feature value corresponding to the channel with a score greater than or equal to the first threshold in the positive event stream feature to 1, and set the feature value corresponding to the other channels to 0 to obtain the hard mask of the positive event stream feature; The soft mask for the positive event stream features is generated based on the following formula: ; in, For the soft mask, , This is a score vector composed of the scores of each channel in the positive event stream feature. The first threshold, Temperature coefficient; The sparse encoding of the positive event stream at time t is obtained based on the following formula: ; in, Sparse encoding of the positive event stream. The hard mask; This indicates that the gradient operator is stopped during forward propagation. .
[0012] Based on the above-described event flow cross-modal point process alignment method based on pulse networks, the logarithm of the conditional intensity function at each moment is: ; in, Let be the logarithm of the conditional intensity function at time t. The fundamental intensity component at time t, Let be the semantic modulation component at time t. It is a learnable modulation intensity factor.
[0013] According to the above-mentioned event flow cross-modal point process alignment method based on pulse networks, interaction is performed based on the semantic vector to be aligned and the event flow features at any time t, specifically including: The semantic vector to be aligned is mapped into a query vector via a query projection network; The event flow features at time t are mapped to the event interaction vector at time t via an event projection network; The query vector and the event interaction vector at time t are multiplied element-wise and then the semantic modulation component at time t is output through a semantic modulation conversion network; the semantic modulation conversion network consists of layer normalization and multilayer perceptron.
[0014] Based on the above-described event flow cross-modal point process alignment method based on pulse networks, the fundamental intensity component at any time t is: ; in, The event stream features at time t are represented by base_head, which is a base intensity transformation network consisting of layer normalization and a multilayer perceptron.
[0015] According to the above-mentioned event flow cross-modal point process alignment method based on pulse networks, the step of extracting semantic vectors from text queries to obtain semantic vectors to be aligned specifically includes: The text query is input into a pre-trained language model to obtain a text semantic vector.
[0016] The text semantic vector is concatenated with the learnable missing modality compensation vector and then input into the semantic fusion network to obtain the semantic vector to be aligned.
[0017] According to the above-described event flow cross-modal point process alignment method based on pulse networks, the reward energy is: ; in, The reward energy is represented by i, which corresponds to the index of each element. Let i be the element at index i in a one-dimensional logarithmic vector. Let i be the element at index i in a one-dimensional sparse vector.
[0018] According to the above-described event flow cross-modal point process alignment method based on pulse networks, the penalty energy is: ; in, Let i be the penalty energy, and i be the index of each element. Let i be the element at index i in a one-dimensional logarithmic vector. The dimension of time. is the dimension of the channel dimension.
[0019] According to a second aspect of the present invention, an event flow cross-modal point process alignment system based on a pulse network is provided, comprising: The voxelization preprocessing module is used to preprocess the candidate event stream output by the event camera by voxelization and then split it into positive event stream meshes and negative event stream meshes according to the polarity channel. The spiking neural network encoding module is used to encode the positive event flow grid and the negative event flow grid based on two spiking neural networks respectively, and to concatenate the positive event flow features and negative event flow features obtained at each time step to obtain the event flow features at each time step; The sparse observation generation module is used to sparsely encode the positive event stream features and negative event stream features at each time step, and then concatenate the sparse codes of the positive and negative event streams at each time step to obtain the sparse observation representation of the event stream at each time step. The conditional intensity function calculation module is used to extract semantic vectors from text queries to obtain semantic vectors to be aligned. Based on the semantic vectors to be aligned, it interacts with the event flow features at each time point. The semantic modulation components obtained at each time point are combined with the basic intensity components transformed based on the event flow features at the same time point to obtain the logarithm of the conditional intensity function at each time point. The point process energy calculation module is used to flatten the logarithm of the conditional intensity function at each time step and the sparse observation representation of the event stream according to the time dimension and the channel dimension to obtain a one-dimensional logarithmic vector and a one-dimensional sparse vector. Then, it determines the index of the element with a value of 1 in the one-dimensional sparse vector as the activation position, calculates the reward energy based on the elements at the activation positions in the one-dimensional logarithmic vector, calculates the penalty energy based on the one-dimensional logarithmic vector, and calculates the difference between the reward energy and the penalty energy as the point process energy. Specifically, the larger the element at the activation position in the one-dimensional logarithmic vector, the larger the reward energy; the larger the average value of the elements in the one-dimensional logarithmic vector, the larger the penalty energy. The matching degree determination module is used to determine the matching degree between the candidate event stream and the text query based on the point process energy; if the point process energy is obtained by subtracting the penalty energy from the reward energy, then the larger the point process energy, the larger the matching degree between the candidate event stream and the text query; otherwise, the smaller the point process energy, the larger the matching degree between the candidate event stream and the text query.
[0020] According to a third aspect of the present invention, an electronic device is provided, comprising: a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in the first aspect.
[0021] According to a fourth aspect of the invention, a computer-readable storage medium is provided, the computer-readable storage medium storing computer instructions for causing a processor to perform the method as described in the first aspect.
[0022] According to a fifth aspect of the invention, a computer program product is provided, comprising a computer program or instructions that, when executed by a processor, implement the method as described in the first aspect.
[0023] In summary, compared with the prior art, the above-described technical solutions conceived by this invention can achieve the following beneficial effects: By modeling event flows as discrete spatiotemporal point processes and extracting event flow features through spiking neural networks, instead of compressing them into dense vectors, point process energy is used as a cross-modal alignment metric in the sparse observation space. This avoids the temporal collapse problem caused by the forced continuous nature of event flows and preserves the fine temporal structure of asynchronous pulses. Simultaneously, using point process energy instead of cosine similarity as the metric standard ensures that the metric space matches the physical nature of event flows as discrete point processes, fundamentally solving the problem of misaligned metric objects in existing methods. This allows for effective alignment with continuous semantic modalities without sacrificing the sparse temporal characteristics of events. Furthermore, sparse coding is used to obtain sparse observation representations of the event flow at each time step, fully utilizing the sparsity of the event flow itself and the concentration of key information in a few highly activated regions, improving computational efficiency and ensuring the stability and reliability of cross-modal alignment. Attached Figure Description
[0024] Figure 1 This is a flowchart illustrating the event flow cross-modal point process alignment method based on pulse networks provided in an embodiment of the present invention.
[0025] Figure 2 This is a schematic diagram of the process for obtaining sparse observation representations of event streams provided in an embodiment of the present invention.
[0026] Figure 3 This is a schematic flowchart for point process energy calculation provided in an embodiment of the present invention.
[0027] Figure 4 This is an overall flowchart of cross-modal alignment provided in an embodiment of the present invention. Detailed Implementation
[0028] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention. Furthermore, the technical features involved in the various embodiments of this invention described below can be combined with each other as long as they do not conflict with each other.
[0029] This invention provides a method for aligning event stream processes across modal points based on pulse networks, such as... Figure 1 As shown, it includes: Step 110: After voxelization preprocessing, the candidate event stream output by the event camera is split into positive event stream mesh and negative event stream mesh according to the polarity channel; Step 120: Encode the positive event flow grid and the negative event flow grid based on two spiking neural networks respectively, and concatenate the positive event flow features and negative event flow features obtained at each time step to obtain the event flow features at each time step; Step 130: Sparsely encode the positive event stream features and negative event stream features at each time point, and concatenate the sparse codes of the positive event stream and negative event stream at each time point to obtain the sparse observation representation of the event stream at each time point; Step 140: Extract semantic vectors from the text query to obtain semantic vectors to be aligned. Interact with the event flow features at each time step based on the semantic vectors to be aligned. Combine the semantic modulation components at each time step obtained from the interaction with the basic intensity components transformed based on the event flow features at the same time step to obtain the logarithm of the conditional intensity function at each time step. Step 150: After flattening the logarithm of the conditional intensity function at each time point and the sparse observation representation of the event stream according to the time dimension and the channel dimension to obtain a one-dimensional logarithmic vector and a one-dimensional sparse vector, the index of the element with a value of 1 in the one-dimensional sparse vector is determined as the activation position. The reward energy is calculated based on the elements of the activation positions in the one-dimensional logarithmic vector, and the penalty energy is calculated based on the one-dimensional logarithmic vector. The difference between the reward energy and the penalty energy is calculated as the point process energy. Wherein, the larger the element of the activation position in the one-dimensional logarithmic vector, the larger the reward energy; the larger the average value of the elements in the one-dimensional logarithmic vector, the larger the penalty energy. Step 160: Determine the matching degree between the candidate event stream and the text query based on the point process energy; if the point process energy is obtained by subtracting the penalty energy from the reward energy, then the larger the point process energy, the greater the matching degree between the candidate event stream and the text query; otherwise, the smaller the point process energy, the greater the matching degree between the candidate event stream and the text query.
[0030] Here, the overall idea of this invention is to rigorously model the event stream as a discrete spatiotemporal point process, extract latent features and generate sparse observation representations through a polarity-separated spiking neural network encoder, and use the negative log-likelihood energy of the point process instead of the traditional cosine similarity as the alignment metric in the sparse observation space. This transforms text semantics into conditional modulation of the event stream intensity function, thereby achieving efficient and effective alignment with the text semantic modality without sacrificing the sparse temporal characteristics of the events. This scheme can be applied to any cross-modal retrieval scenario from text to event streams, such as text-based video clip retrieval, UAV reconnaissance image retrieval, and autonomous driving scene query. Without loss of generality, the following embodiments use the cross-modal retrieval task from text to event streams on the N-Caltech101 dataset as an example. The N-Caltech101 dataset contains 101 categories and approximately 8709 event stream samples. Each sample is equipped with a corresponding RGB image and category label. Text queries are constructed using the category labels to achieve cross-modal retrieval and alignment from text to event streams.
[0031] It should be noted that the cross-modal alignment of event streams and text queries is applied to each candidate event stream output by the event camera. The processing flow for each candidate event stream is the same; this embodiment of the invention only describes one candidate event stream as an example. For any candidate event stream, it undergoes voxelization preprocessing, transforming it into a fixed-size voxel mesh. Its dimensions are The specific parameter can be set to: number of time steps. Number of polar channels (Channel 0 corresponds to negative polarity events, and channel 1 corresponds to positive polarity events), spatial resolution , Voxel mesh Each element Indicates time step ,Location polarity The cumulative event count. (This refers to the voxel mesh.) Divide into positive event flow grids according to polarity channels. and negative event flow grid Both have the same dimension. .
[0032] Subsequently, the positive event flow grid and the negative event flow grid are respectively input into two structurally identical but parameter-independent spiking neural networks for encoding, to obtain the positive event flow features and negative event flow features at each time step. In some embodiments, each spiking neural network consists of two spiking convolutional modules, each including a convolutional layer, a batch normalization layer, and a LIF neuron layer.
[0033] The dynamic equation for each LIF neuron in the LIF neuron layer is as follows: ; in, and The membrane potentials at the current and historical moments are respectively. The input current from the batch normalization layer output at the current moment. This is the membrane time constant. When the membrane potential... Exceeding the distribution threshold At this time, the LIF neuron fires a pulse and resets the membrane potential to Pulse output Represented as: ; in It is a step function.
[0034] In some embodiments, the first pulse convolution module has a convolutional layer with 1 input channel, 32 output channels, a 3×3 kernel size, 1% padding, and no bias; followed by a batch normalization layer and a LIF neuron layer. The LIF neuron membrane time constant... Issuance threshold Reset potential The second pulse convolution module has a convolutional layer with 32 input channels, 96 output channels, a kernel size of 3×3, a stride of 2, padding of 1, and no bias; followed by a batch normalization layer and a LIF neuron layer, with parameters set for the first pulse convolution module.
[0035] Taking the positive event flow grid as an example, The pulses are fed into the first pulse convolution module step by step to obtain the intermediate pulse sequence. After a second pulse convolutional module and global average pooling, the spatial size is compressed to 1×1. Finally, a linear projection layer maps the pooled features to the positive event stream features at that moment. ,in The dimension of the total latent features. Similarly, the negative event flow grid is derived from... The negative event flow characteristics are obtained step-by-step. , and All dimensions are , It can be 256.
[0036] By concatenating the outputs of the two spiking neural networks at each time step along the channel dimension, the event flow features at each time step are obtained: ; in , This indicates vector concatenation.
[0037] To obtain sparse and interpretable event observation representations, sparse encoding is performed on the positive and negative event flow features at each time step. The sparse encodings of the positive and negative event flows at each time step are then concatenated to obtain the sparse event flow observation representations at each time step.
[0038] In some embodiments, when sparsely encoding the positive and negative event flow features at time t, a learnable score selection network can be used to score the positive and negative event flow features at time t, respectively, to obtain the score of each channel in the positive and negative event flow features at time t. In some embodiments, the score selection network can be composed of layer normalization, a first fully connected layer, a GELU nonlinear activation function, and a second fully connected layer cascaded together. For positive event flow features, the output score vector is... ,Depend on The fractional composition of each channel; similarly, the negative event stream characteristics yield a fractional vector. Each element in the score vector represents the importance of the corresponding channel at the current time step. Subsequently, the scores of each channel in the positive event stream features are sorted in descending order, and a first threshold is selected from the sorted channel scores based on the sparsity ratio to perform sparse encoding of the positive event stream features based on the first threshold. The number of positive event stream features to be retained is set. ,in For the preset sparsity ratio, the fractional vector Sort in descending order, and denote the th... The largest score is the first threshold. Channels with scores less than the first threshold are encoded as 0, and channels with scores greater than or equal to the first threshold are encoded as 1. Similarly, the scores of each channel in the negative event stream feature are sorted in descending order, and a second threshold is selected from the scores of each sorted channel based on the sparsity ratio, so as to perform sparse encoding on the negative event stream feature based on the second threshold.
[0039] In other embodiments, the sparse encoding of positive and negative event stream features based on a threshold is performed in the same way; here, only the processing method for positive event stream features is further explained. Specifically, the positive event stream features at time t are selected based on a score greater than or equal to a first threshold. The feature value corresponding to the positive event stream is set to 1, and the feature values corresponding to the other channels are set to 0, thus obtaining a hard mask for the positive event stream features. ; Subsequently, a soft mask for positive event stream features is generated based on the following formula. : ; in, , This is a score vector composed of the scores of each channel in the positive event stream features. Temperature coefficient; Then, obtain the sparse encoding of the positive event stream at time t based on the following formula. : ; in, This indicates that the gradient operator is stopped.
[0040] Here, the gradient stopping operator works as follows: (1) During forward propagation, ,therefore The output is a precise 0 / 1 hard mask, which satisfies the discrete requirements of sparse observations.
[0041] (2) During back propagation, It is not differentiable (the gradient is zero). The gradient is also cut off (the gradient is zero), therefore The gradient is equal to The gradient is then passed back to the score selection network via the differentiable path of the sigmoid function, enabling end-to-end differentiable training.
[0042] In short, by introducing a stopping gradient operator, a hard mask is used to ensure sparsity during forward propagation, and the gradient of a soft mask is used to update the network parameters during back propagation, thus solving the problem of non-differentiability of discrete selection operations.
[0043] Similarly, the sparse encoding of the negative event stream at time t can be obtained. Then, the sparse encodings of the positive and negative event streams at each time are concatenated to obtain the sparse observation representation of the event stream at each time. Its dimensions are Obtain sparse observation representations of event streams. The overall process is as follows Figure 2 As shown, during forward propagation, the sparse encoding of positive and negative event streams serves as the corresponding hard mask, thereby significantly improving the efficiency of point process energy calculation during inference.
[0044] For text queries in cross-modal alignment tasks, semantic vectors are extracted from the text query to obtain the semantic vector to be aligned. In some embodiments, the text query can be input into a pre-trained language model (e.g., CLIP text encoder (ViT-B / 32)) to obtain the text semantic vector. This text semantic vector is then concatenated with a learnable missing modality compensation vector and input into a semantic fusion network to obtain the semantic vector to be aligned. The semantic fusion network can consist of alternating cascaded fully connected layers and GELU activation functions, mapping the concatenated vector of the text semantic vector and the learnable missing modality compensation vector to the semantic vector to be aligned. Its dimensions are Specifically, the introduction of missing modality compensation vectors to fill in unused RGB modalities ensures that the training and testing objectives remain consistent in inference scenarios using only text queries, effectively avoiding performance loss caused by training-test inconsistency. Furthermore, the design of the missing modality compensation vectors provides an interface for future expansion to text + RGB multimodal joint queries, demonstrating good scalability.
[0045] Subsequently, the semantic vector to be aligned is interacted with the event flow features at each time step, and the semantic modulation components at each time step obtained from the interaction are combined with the basic intensity components transformed from the event flow features at the same time step to obtain the logarithm of the conditional intensity function at each time step.
[0046] In some embodiments, the candidate event stream can be timed. Logarithm of conditional strength function The model is the sum of the fundamental intensity component and the semantic modulation component at that moment: ; in, Let be the logarithm of the conditional intensity function at time t. The fundamental intensity component at time t, Let be the semantic modulation component at time t. This is a learnable modulation intensity factor, initialized to 10.0 during training. The reason for initializing it to a large positive value is: to make The variance of the component was significantly greater than that of the baseline strength component in the early stages of training. The variance of the model guides it to rely on textual semantic information for cross-modal alignment decisions from the initial training stage, avoiding the dominance of basic strength components in ranking and thus weakening semantic alignment capabilities. After training converges, The value remains around 9.7. The function restricts the semantic modulation components to Within the range, it prevents semantic modulation from diverging unbounded, while ensuring that the conditional strength function is always non-negative.
[0047] In other embodiments, the fundamental strength component at any time t is: ; in, For the event stream features at time t, the base_head is a base intensity transformation network, consisting of layer normalization and a multilayer perceptron. Specifically, the structure of the base_head can be: layer normalization (dimension 256) → fully connected layer (256→256) → GELU activation function → fully connected layer (256→256). It can be seen that... It relies solely on event flow characteristics to characterize the underlying activity intensity of the event flow.
[0048] In other embodiments, the semantic modulation component at time t The methods of obtaining it include: The semantic vector c to be aligned is passed through a query projection network. Mapped to query vector q, the event stream features at time t via event projection network Mapped to the event interaction vector at time t Query vector q and event interaction vector All are D-dimensional vectors. In some embodiments, the projection network is queried. and event projection network Both are fully connected layers from D-dimensional to D-dimensional. The query vector q and the event interaction vector at time t are connected. After element-wise multiplication, the semantic modulation component at time t is output by the semantic modulation conversion network. The semantic modulation and transformation network consists of layer normalization and a multilayer perceptron. Specifically, the structure of the semantic modulation and transformation network is: layer normalization (dimension 256) → fully connected layer (256→256) → GELU activation function → fully connected layer (256→256). This design allows the text semantics to play a decisive role in the direction and amplitude of event intensity modulation.
[0049] Next, the logarithm of the conditional strength function at each time step is... and sparse observation representation of event flow Flattening by time and channel dimensions yields a one-dimensional logarithmic vector. and one-dimensional sparse vectors Specifically, all Each time step and Flattening along the time and channel dimensions yields two dimensions. vector and Determine a one-dimensional sparse vector. The index of an element with a value of 1 is the activation position, based on a one-dimensional logarithmic vector. Calculate reward energy for elements in the active position. Based on one-dimensional logarithmic vectors Calculate penalty energy And calculate reward energy. With punishment energy The difference between them is the point process energy. Among them, the one-dimensional logarithmic vector The larger the element in the activated position, the greater the reward energy. Larger; one-dimensional logarithmic vector The higher the average value of the elements in the mixture, the greater the penalty energy. The larger the reward energy. The average logarithm of the intensity at activation locations in sparse observations; a higher logarithm indicates a more complete semantic explanation of the event at that location; penalty energy. The average of the intensity logarithms across all locations is used to suppress excessively high overall intensity logarithms.
[0050] In some embodiments, reward energy for: ; Where i corresponds to the index of each element. Let i be the element at index i in a one-dimensional logarithmic vector. Let i be the element at index i in a one-dimensional sparse vector.
[0051] Punishment energy for: ; in, The dimension of time. is the dimension of the channel dimension.
[0052] The calculation process for the energy of the entire point process is as follows: Figure 3 As shown. The point process energy obtained based on the above method. It is possible to determine the degree of matching (or alignment) between the candidate event stream and the text query. Among these, if the point process energy... Reward energy Subtract penalty energy Get (i.e.) Then the energy of the point process The larger the value, the greater the match between the candidate event stream and the text query; otherwise (i.e.) Point process energy The smaller the value, the greater the match between the candidate event stream and the text query.
[0053] During the inference phase, for a given text query, it can be paired with each candidate event stream one by one in the manner described above, and the point process energy of each candidate event stream can be calculated. Subsequently, the energy of the process at each point. Sort the candidate event streams in ascending / descending order. The candidate event streams with the lowest / highest energy are those that best align with the semantics of the text query and can be output. The candidate event stream is used as the cross-modal retrieval result. The entire process is as follows: Figure 4 As shown.
[0054] During the training phase, for those containing For each batch of matched samples, calculate the energy matrix between each pair of text queries and candidate event streams. ,in Indicates the first The first text query and the first The point process energy between event flows, where the th event... The first text query and the first There are matching relationships between the event streams. The following contrastive loss function can be constructed: ; in The temperature hyperparameter is used. Minimizing this loss is equivalent to maximizing the likelihood of matching text-event pairs while suppressing the likelihood of non-matching pairs. End-to-end parameter optimization is performed on the entire model (including two spiking neural networks, a score selection network, a query projection network, an event projection network, a semantic modulation transformation network, a base intensity transformation network, and a semantic fusion network) using backpropagation and gradient descent algorithms. The entire training process consistently uses unidirectional text-to-event retrieval as the primary training objective. The AdamW optimizer with stochastic gradient descent can be used, with an initial learning rate of... The weight decays to The total number of training rounds is 100.
[0055] In summary, the method provided by this invention models the event flow as a discrete spatiotemporal point process. After extracting event flow features through a spiking neural network, it does not compress them into dense vectors. Instead, it uses point process energy as a cross-modal alignment metric in a sparse observation space. This avoids the temporal collapse problem caused by the forced continuous nature of the event flow and preserves the fine temporal structure of asynchronous pulses. At the same time, by using point process energy instead of cosine similarity as the metric standard, the metric space matches the physical nature of the event flow as a discrete point process, fundamentally solving the problem of misaligned metric objects in existing methods. This allows for effective alignment with continuous semantic modalities without sacrificing the sparse temporal characteristics of the event. Furthermore, by obtaining sparse observation representations of the event flow at each time step through sparse coding, it fully utilizes the sparsity of the event flow itself and the concentration of key information in a few highly activated regions, improving computational efficiency and ensuring the stability and reliability of cross-modal alignment.
[0056] The following describes an event flow cross-modal point process alignment system based on pulse networks provided by the present invention. The event flow cross-modal point process alignment system based on pulse networks described below can be referred to in correspondence with the event flow cross-modal point process alignment method based on pulse networks described above.
[0057] This invention provides an event flow cross-modal point process alignment system based on pulse networks, comprising: The voxelization preprocessing module is used to preprocess the candidate event stream output by the event camera by voxelization and then split it into positive event stream meshes and negative event stream meshes according to the polarity channel. The spiking neural network encoding module is used to encode the positive event flow grid and the negative event flow grid based on two spiking neural networks respectively, and to concatenate the positive event flow features and negative event flow features obtained at each time step to obtain the event flow features at each time step; The sparse observation generation module is used to sparsely encode the positive event stream features and negative event stream features at each time step, and then concatenate the sparse codes of the positive and negative event streams at each time step to obtain the sparse observation representation of the event stream at each time step. The conditional intensity function calculation module is used to extract semantic vectors from text queries to obtain semantic vectors to be aligned. Based on the semantic vectors to be aligned, it interacts with the event flow features at each time point. The semantic modulation components obtained at each time point are combined with the basic intensity components transformed based on the event flow features at the same time point to obtain the logarithm of the conditional intensity function at each time point. The point process energy calculation module is used to flatten the logarithm of the conditional intensity function at each time step and the sparse observation representation of the event stream according to the time dimension and the channel dimension to obtain a one-dimensional logarithmic vector and a one-dimensional sparse vector. Then, it determines the index of the element with a value of 1 in the one-dimensional sparse vector as the activation position, calculates the reward energy based on the elements at the activation positions in the one-dimensional logarithmic vector, calculates the penalty energy based on the one-dimensional logarithmic vector, and calculates the difference between the reward energy and the penalty energy as the point process energy. Specifically, the larger the element at the activation position in the one-dimensional logarithmic vector, the larger the reward energy; the larger the average value of the elements in the one-dimensional logarithmic vector, the larger the penalty energy. The matching degree determination module is used to determine the matching degree between the candidate event stream and the text query based on the point process energy; if the point process energy is obtained by subtracting the penalty energy from the reward energy, then the larger the point process energy, the larger the matching degree between the candidate event stream and the text query; otherwise, the smaller the point process energy, the larger the matching degree between the candidate event stream and the text query.
[0058] This invention provides an electronic device, including: a computer-readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is configured to read executable instructions stored in the computer-readable storage medium and execute the method as described in any of the above embodiments.
[0059] This invention provides a computer-readable storage medium storing computer instructions that cause a processor to perform the method described in any of the above embodiments.
[0060] This invention provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the method described in any of the above embodiments.
[0061] Those skilled in the art will readily understand that the above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for aligning event stream processes across modal points based on pulse networks, characterized in that, include: After voxelization preprocessing, the candidate event stream output by the event camera is split into positive event stream mesh and negative event stream mesh according to the polarity channel; The positive event flow grid and the negative event flow grid are encoded using two spiking neural networks respectively, and the positive event flow features and negative event flow features obtained at each time step are concatenated to obtain the event flow features at each time step. Sparse coding is performed on the positive event stream features and negative event stream features at each time step, and the sparse coding of the positive event stream and negative event stream at each time step is concatenated to obtain the sparse observation representation of the event stream at each time step; Semantic vectors are extracted from text queries to obtain semantic vectors to be aligned. These semantic vectors are then interacted with event flow features at each time step. The semantic modulation components obtained from the interaction at each time step are combined with the basic intensity components obtained from the event flow features at the same time step to obtain the logarithm of the conditional intensity function at each time step. After flattening the logarithm of the conditional intensity function at each time point and the sparse observation representation of the event stream according to the time and channel dimensions to obtain a one-dimensional logarithmic vector and a one-dimensional sparse vector, the index of the element with a value of 1 in the one-dimensional sparse vector is determined as the activation position. The reward energy is calculated based on the elements at the activation positions in the one-dimensional logarithmic vector, and the penalty energy is calculated based on the one-dimensional logarithmic vector. The difference between the reward energy and the penalty energy is determined as the point process energy, or vice versa. Specifically, the larger the element at the activation position in the one-dimensional logarithmic vector, the larger the reward energy; the larger the average value of the elements in the one-dimensional logarithmic vector, the larger the penalty energy. The degree of matching between the candidate event stream and the text query is determined based on the point process energy. If the point process energy is obtained by subtracting the penalty energy from the reward energy, then the greater the point process energy, the greater the degree of matching between the candidate event stream and the text query; otherwise, the smaller the point process energy, the greater the degree of matching between the candidate event stream and the text query.
2. The event flow cross-modal point process alignment method based on pulse networks as described in claim 1, characterized in that, Sparse encoding is performed on the positive and negative event stream features at time t, specifically including: The positive event flow features and negative event flow features at time t are scored using a score selection network to obtain the scores of each channel in the positive and negative event flow features at time t. The scores of each channel in the positive event stream feature are sorted in descending order, and a first threshold is selected from the scores of each channel after sorting based on the sparsity ratio, so as to perform sparse encoding on the positive event stream feature based on the first threshold. The scores of each channel in the negative event stream feature are sorted in descending order, and a second threshold is selected from the scores of each channel after sorting based on the sparsity ratio, so as to perform sparse encoding on the negative event stream feature based on the second threshold.
3. The event flow cross-modal point process alignment method based on pulse networks as described in claim 2, characterized in that, Sparse encoding of the positive event stream features based on the first threshold specifically includes: Set the feature value corresponding to the channel with a score greater than or equal to the first threshold in the positive event stream feature to 1, and set the feature value corresponding to the other channels to 0 to obtain the hard mask of the positive event stream feature; The soft mask for the positive event stream features is generated based on the following formula: ; in, For the soft mask, , This is a score vector composed of the scores of each channel in the positive event stream feature. For the first threshold, Temperature coefficient; The sparse encoding of the positive event stream at time t is obtained based on the following formula: ; in, Sparse encoding of the positive event stream. The hard mask; This indicates that the gradient operator is stopped during forward propagation. .
4. The event flow cross-modal point process alignment method based on pulse networks as described in claim 1, characterized in that, The logarithms of the conditional intensity function at each time point are: ; in, Let be the logarithm of the conditional intensity function at time t. The fundamental intensity component at time t, Let be the semantic modulation component at time t. It is a learnable modulation intensity factor.
5. The event flow cross-modal point process alignment method based on pulse networks as described in claim 4, characterized in that, Interaction is performed based on the semantic vector to be aligned and the event stream features at any time t, specifically including: The semantic vector to be aligned is mapped into a query vector via a query projection network; The event flow features at time t are mapped to the event interaction vector at time t via an event projection network; The query vector and the event interaction vector at time t are multiplied element-wise and then the semantic modulation component at time t is output through a semantic modulation conversion network; the semantic modulation conversion network consists of layer normalization and multilayer perceptron.
6. The event flow cross-modal point process alignment method based on pulse networks as described in claim 4, characterized in that, The foundation strength components at any time t are: ; in, The event stream features at time t are represented by base_head, which is a base intensity transformation network consisting of layer normalization and a multilayer perceptron.
7. The event flow cross-modal point process alignment method based on pulse networks as described in claim 1, characterized in that, The step of extracting semantic vectors from the text query to obtain the semantic vector to be aligned specifically includes: The text query is input into a pre-trained language model to obtain a text semantic vector. The text semantic vector is concatenated with the learnable missing modality compensation vector and then input into the semantic fusion network to obtain the semantic vector to be aligned.
8. The method for aligning event streams across modal points based on pulse networks as described in claim 1, characterized in that, The reward energy is: ; in, The reward energy is represented by i, which corresponds to the index of each element. Let i be the element at index i in a one-dimensional logarithmic vector. Let i be the element at index i in a one-dimensional sparse vector.
9. The method for aligning event streams across modal points based on pulse networks as described in claim 1, characterized in that, The penalty energy is: ; in, Let i be the penalty energy, and i be the index of each element. Let i be the element at index i in a one-dimensional logarithmic vector. The dimension of time. is the dimension of the channel dimension.
10. A cross-modal point process alignment system for event streams based on pulse networks, characterized in that, include: The voxelization preprocessing module is used to preprocess the candidate event stream output by the event camera by voxelization and then split it into positive event stream meshes and negative event stream meshes according to the polarity channel. The spiking neural network encoding module is used to encode the positive event flow grid and the negative event flow grid based on two spiking neural networks respectively, and to concatenate the positive event flow features and negative event flow features obtained at each time step to obtain the event flow features at each time step; The sparse observation generation module is used to sparsely encode the positive event stream features and negative event stream features at each time step, and then concatenate the sparse codes of the positive and negative event streams at each time step to obtain the sparse observation representation of the event stream at each time step. The conditional intensity function calculation module is used to extract semantic vectors from text queries to obtain semantic vectors to be aligned. Based on the semantic vectors to be aligned, it interacts with the event flow features at each time point. The semantic modulation components obtained at each time point are combined with the basic intensity components transformed based on the event flow features at the same time point to obtain the logarithm of the conditional intensity function at each time point. The point process energy calculation module is used to flatten the logarithm of the conditional intensity function at each time step and the sparse observation representation of the event stream according to the time dimension and the channel dimension to obtain a one-dimensional logarithmic vector and a one-dimensional sparse vector. Then, it determines the index of the element with a value of 1 in the one-dimensional sparse vector as the activation position. Based on the elements at the activation positions in the one-dimensional logarithmic vector, it calculates the reward energy and the penalty energy. It then determines the difference between the reward energy and the penalty energy as the point process energy, or vice versa. Specifically, the larger the element at the activation position in the one-dimensional logarithmic vector, the larger the reward energy; the larger the average value of the elements in the one-dimensional logarithmic vector, the larger the penalty energy. The matching degree determination module is used to determine the matching degree between the candidate event stream and the text query based on the point process energy; if the point process energy is obtained by subtracting the penalty energy from the reward energy, then the larger the point process energy, the larger the matching degree between the candidate event stream and the text query; otherwise, the smaller the point process energy, the larger the matching degree between the candidate event stream and the text query.
Citation Information
Patent Citations
Target recognition method, device and system and computer readable storage medium
CN111275742A
Sensing, storing and computing integrated signal processing method and system based on SNN and ANN hybrid architecture
CN121807773A