Isolated sign language word recognition method based on event camera space-time multi-scale feature fusion
By using the spatiotemporal multi-scale feature fusion method of event cameras in isolated sign word recognition, multi-scale features are extracted using adaptive attenuation factors and feature pyramid networks, the problems of insufficient modeling of spatiotemporal characteristics of event camera data and data redundancy in traditional methods in the prior art are solved, and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202510204355.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-03
AI Technical Summary
The prior art is difficult to effectively model the complex spatiotemporal characteristics of event camera data in isolated sign word recognition, and traditional frame-based methods have limitations such as high data redundancy and motion blur, which affects the recognition accuracy.
The space-time multi-scale feature fusion method based on event camera is adopted to generate event frame sequences through logarithmic perception mechanism, optimize temporal multi-scale features using adaptive attenuation factors, and extract spatial multi-scale features in combination with feature pyramid networks, and finally perform isolated sign word recognition through a full connection layer.
It improves the modeling ability and feature expression of complex spatiotemporal characteristics of event camera data, enhances the accuracy of isolated sign language word recognition, and overcomes the data redundancy and motion fuzzy problems of traditional methods.
Smart Images

Figure CN120088860A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of isolated sign language word recognition, and in particular relates to an isolated sign language word recognition method based on spatio-temporal multi-scale feature fusion of an event camera. Background Art
[0002] Sign language recognition technology aims to convert sign language actions into text or speech through computer vision and machine learning methods, so as to help deaf people communicate with hearing people without barriers. Traditional visual sign language recognition methods mainly rely on frame-based visual sensors (such as RGB cameras) and deep learning models (such as convolutional neural networks). Through training with a large amount of sign language data, the model extracts feature information such as hand key points and gesture shapes, and then recognizes sign language actions. According to different recognition tasks, sign language recognition can be divided into isolated sign language word recognition and continuous sign language recognition. Among them, isolated sign language word recognition mainly classifies the vocabulary corresponding to a single gesture, which is the basic task of sign language recognition and has wide applications in scenarios such as sign language dictionary query and human-computer interaction. However, since traditional visual methods use frame-level information processing, the data redundancy is relatively high, which not only leads to waste of computing resources, but also reduces the overall processing efficiency. Moreover, when the gesture actions of some sign language words are relatively fast, motion blur problems are likely to occur, resulting in the loss or distortion of gesture information in key frames, thereby affecting the recognition accuracy. In addition, in high dynamic range scenarios (such as strong light or weak light conditions), traditional RGB cameras have poor adaptability to light changes and complex backgrounds, further limiting their performance and promotion in practical applications.
[0003] In recent years, event cameras have been introduced into the field of sign language recognition. Based on the principle of biological inspiration, event cameras construct a bionic vision system by simulating the structure and working mechanism of the biological visual cortex, enabling the machine to possess efficient perception capabilities similar to those of humans. This sensor abandons the traditional recording method based on absolute light intensity and instead encodes visual information by capturing the change in light intensity, converting the dynamic changes in the external scene into an asynchronous and sparse spatio-temporal pulse event stream, and using the address-event representation protocol to express such data. This unique working mechanism enables the event camera to generate data only when the scene changes, and its output has the characteristics of spatial sparsity and temporal density. This design not only minimizes data redundancy but also ensures high response capabilities in a rapidly changing environment. Compared with traditional frame-based sensors, event cameras have advantages such as low latency, high dynamic range, low power consumption, and low redundant data, making them an ideal choice for dynamic vision tasks such as sign language recognition.
[0004] However, the non-uniform and discontinuous characteristics of event data pose significant challenges to extracting meaningful spatio-temporal features. How to learn and process the new data paradigm is a daunting task. Early methods mostly relied on manually designed rules and tailored some simple feature extraction methods for the sparsity and asynchronous characteristics of event streams, such as counting the number of events, event time distribution, etc. Although these methods can capture some features, they are limited by prior knowledge, difficult to effectively represent the complex spatio-temporal relationships in event streams, highly dependent on specific datasets or tasks, and lack generalization ability. To address these limitations, recent research has drawn on the successful experience of neural networks in traditional visual data and explored their potential in mining the complex spatio-temporal characteristics of event streams. Some methods introduce graph neural networks (GNNs) to capture local correlations by taking events as graph nodes and constructing spatio-temporal relationship edges. Although the data sparsity is retained, there are challenges in edge relationship definition and time dynamics modeling. Other research uses LSTM to capture time dynamics. Although the effect is good, it faces problems of high computational complexity and large number of parameters. Spiking neural networks (SNNs) show potential due to their matching with the asynchronous characteristics of event streams, but are restricted by the bottlenecks of high event data noise and difficult training. Recent research mainly focuses on converting asynchronous events into frame-based synchronous representations and using deep learning pre-trained models, such as Resnet, Googlenet, and VGG, to extract features. Although this method effectively balances computational complexity and the richness of spatio-temporal information, these frameworks usually extract features at specific time or space scales, resulting in incomplete capture of complex spatio-temporal dependencies. Summary of the Invention
[0005] Aiming at the above deficiencies in the prior art, an isolated sign language word recognition method based on spatio-temporal multi-scale feature fusion of an event camera provided by the present invention solves the problems of insufficient modeling of the complex spatio-temporal characteristics of event camera data and weak generalization ability in the prior art. At the same time, it overcomes the limitations of high data redundancy and motion blur in traditional frame-based methods, and improves the accuracy of isolated sign language word recognition.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is: an isolated sign language word recognition method based on spatio-temporal multi-scale feature fusion of an event camera, comprising the following steps: S1. Use an event camera to collect asynchronous event stream data of sign language actions; S2. Based on the collected asynchronous event stream data of sign language actions, use a logarithmic perception mechanism to generate an event frame sequence; S3. Based on the event frame sequence, use an adaptive decay factor to optimize the time multi-scale features; S4. Based on the optimized time multi-scale features, use a feature pyramid network to extract spatial multi-scale features; S5. Based on the extracted spatial multi-scale features, use the fully connected layer to obtain the recognition results of isolated sign language words.
[0007] Advantages of the present invention: The present invention combines the time multi-scale encoding and the spatial multi-scale fusion strategy to achieve the efficient extraction of complex spatio-temporal features of event camera data. By using the adaptive decay factor to capture the dynamic changes of the initial event frames after logarithmic perception encoding at different time scales, accurate time multi-scale feature representations are generated; and combined with the feature pyramid structure, the feature information at different spatial scales is effectively integrated to comprehensively represent complex spatio-temporal dependence relationships. By connecting the extracted spatio-temporal feature representations to the fully connected layer, the target classification results are obtained. That is, the present invention constructs a spatio-temporal multi-scale feature extraction and classification method based on a neuromorphic vision sensor. This spatio-temporal multi-scale modeling method combining logarithmic perception encoding, adaptive decay factor and feature pyramid structure can comprehensively capture the feature information of event data at different time and spatial scales, which helps to improve the modeling ability of the complex spatio-temporal characteristics of neuromorphic data and the integrity of feature expression, and enhance the accuracy of its target classification.
[0008] Further, the expression of the asynchronous event stream data is as follows: ; where represents the asynchronous event stream data for collecting sign language actions, represents the i th event, represents the total number of event streams, represents the pixel coordinate information, represents the time when the event occurs, represents the polarity information. When belongs to a positive event, it represents an increase in pixel light intensity, belongs to a negative event, representing a decrease in pixel light intensity.
[0009] Advantages of the above further solution: By accurately recording the input event data (including time, spatial coordinates and polarity information), the spatio-temporal changes in sign language actions can be efficiently captured, providing high-quality raw data, which provides a reliable basis for subsequent feature extraction and classification.
[0010] Still further, the specific content of S2 is as follows: Based on the asynchronous event stream data of the collected sign language actions, divide the original continuous event stream into multiple windows according to time, and calculate the logarithmic time features of each event within each time window; According to the logarithmic time features of each event, aggregate the event features at each pixel position to generate an event frame; Perform downsampling processing on the generated event frames to obtain an event frame sequence.
[0011] The beneficial effects of the above further solution are as follows: By means of the logarithmic perception mechanism, a non-linear weighting method effectively enhances the sensitivity to recent events. At the same time, through time window division and event aggregation, the noise impact is reduced, and the accuracy and stability of feature extraction are improved.
[0012] Furthermore, the expression of the event frame is as follows: ; ; where represents the eigenvalue at the pixel position within the time window. represents the pixel position coordinates. represents the n th time window. represents the total number of events at the coordinate position within the time window. represents the i th event. represents all events within the time window, classified by polarity (positive and negative). represents the step function, which is 1 when the coordinate of the event is the same as the target pixel position, and 0 otherwise. represents the i th event's coordinate information. represents the logarithmic time feature of each event. represents the termination time of the current time window. represents the time stamps of the events that occurred in the past within the time window.
[0013] Furthermore, the specific content of S3 is as follows: Project the event frame sequence into a high-dimensional feature space; Based on the projection result, by introducing a learnable decay factor , fuse the high-dimensional feature of the current time window at each time step with the high-dimensional feature of the previous time step to form a time multi-scale multi-feature representation, and complete the optimization of time multi-scale features.
[0014] The beneficial effects of the above further solution are as follows: By using the adaptive decay factor to capture the dynamic changes of the initial event frame encoded by the logarithmic perception at different time scales, an accurate time multi-scale time series feature representation is generated. This method effectively improves the modeling ability of the dynamic changes of events at different time scales, thereby enhancing the extraction and expression of time series features.
[0015] Furthermore, the expression of the time multi-scale multi-feature representation is as follows: ; where represents the time multi-scale multi-feature representation.
[0016] Furthermore, S4 is specifically as follows: Extract multi-scale spatial features according to the optimized time multi-scale features; Fuse the multi-scale spatial features using upsampling and addition operations, and splice them to obtain the spatial multi-scale feature representation, completing the extraction of the spatial multi-scale features.
[0017] The beneficial effect of the above further solution is: effectively integrating information of different spatial scales, enhancing the ability to express complex spatio-temporal relationships, and improving the richness of the overall feature representation.
[0018] Furthermore, the expression of the spatial multi-scale feature representation is as follows: ; ; ; where, represents the spatial multi-scale feature representation, represents the splicing operation, represents the upsampling operation, represents the multi-scale spatial features after the fusion of the feature maps of the layer, represents the convolution operation with a convolution kernel size of 1, and respectively represent the output feature maps after the residual block processing of the layer and the layer, represents the element-wise addition operation of the feature maps, represents the layer's residual convolution operation, represents the n time multi-scale feature representation of the
[0019] Furthermore, S5 is specifically as follows: Perform global average pooling on the extracted spatial multi-scale features; Flatten the result of the average pooling process; According to the flattened feature vector, use a fully connected layer for classification to obtain the isolated sign language word recognition result.
[0020] The beneficial effect of the above further solution is: reducing the feature dimension and retaining key information through global average pooling, and combining with a fully connected layer for classification, improving the accuracy of isolated sign language word recognition.
[0021] Furthermore, the expression of the isolated sign language word recognition result is as follows: ; ; ; where, Indicates the recognition result of isolated sign language words, Indicates the Softmax activation function, Indicates the weight matrix of the second fully connected layer, Indicates the ReLU function, Indicates the weight matrix of the first fully connected layer, Indicates the flattened feature vector, and Respectively indicate the bias terms of the first fully connected layer and the second fully connected layer, Indicates the flattening operation, which converts the high-dimensional feature map into a one-dimensional vector, Indicates the pooled feature map, Indicates the c Pooling result of the Indicates the height of the feature map, Indicates the width of the feature map, Indicates the spatial coordinates in the feature map, Indicates the Channel of the spatial multi-scale feature map, Indicates the channel index in the feature map, Indicates the number of channels of the feature map. Brief Description of the Drawings
[0022] Figure 1 Is the flowchart of the method of the present invention.
[0023] Figure 2 Is the system schematic diagram of the present invention. Detailed Embodiment
[0024] The following describes the detailed embodiment of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the detailed embodiment. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.
[0025] Embodiment As Figure 1 and Figure 2 shown, the present invention provides an isolated sign language word recognition method based on spatio-temporal multi-scale feature fusion of an event camera, and its implementation method is as follows: S1. Use an event camera to collect asynchronous event stream data of sign language actions; In this embodiment, a Prophesee EVK4 event camera is used for data collection. The collected event stream can be expressed as: ; where Represents the asynchronous event stream data collected for sign language actions, Indicates the i th event, Indicates the total number of the event stream, Represents the pixel coordinate information, Indicates the time when the event occurs, Represents the polarity information. When belongs to a positive event, it represents an increase in pixel light intensity, belongs to a negative event, it represents a decrease in pixel light intensity. Only the pixels that trigger the event will transmit and process information, greatly reducing power consumption and the accumulation of redundant information.
[0026] To improve the diversity and generalization ability of the data, this embodiment invited multiple participants of different ages, genders, and gesture habits to collect data under different lighting environments. All participants demonstrated gestures according to a unified sign language standard, and each gesture action was repeated multiple times to ensure the stability and coverage of the data. During data collection, each participant recorded 22 isolated sign language words respectively, and the specific duration depends on the complexity of the sign language action.
[0027] S2. Based on the collected asynchronous event stream data of sign language actions, use the logarithmic perception mechanism to generate an event frame sequence, specifically: Based on the collected asynchronous event stream data of sign language actions, divide it into multiple windows according to time, and within each time window, calculate the logarithmic time feature of each event; According to the logarithmic time feature of each event, aggregate the event features at each pixel position to generate an event frame; Perform downsampling processing on the generated event frames to obtain an event frame sequence.
[0028] In this embodiment, since the event camera generates asynchronous and sparse event stream data, which cannot be directly used for deep learning processing, it is necessary to convert it into an event frame sequence. This embodiment uses the logarithmic perception mechanism for data conversion to enhance the system's sensitivity to recent events while retaining the information of historical events.
[0029] First, to extract high-quality time features, divide the continuous event stream into 3 windows according to time , and within each time window, calculate the logarithmic time feature of each event: ; Through the application of logarithmic perception, the influence of historical events with large time intervals can be compressed, the sensitivity to recent events can be enhanced, and the dynamic change features within a short time interval can be retained. As Figure 2 shown, the change amount of logarithmic perception encoding per unit time It shows that events closer to the current moment have a more significant impact on feature representation. After calculating the temporal features, the event features at each pixel position are aggregated to generate an event frame: ; where represents in the time window, the feature value at pixel position , represents the pixel position coordinates, represents the n th time window, represents the total number of events at the coordinate position within the time window, represents the i th event, represents all events within the time window, classified by polarity, represents the step function, which is 1 when the coordinates of the event are the same as the target pixel position, otherwise 0, represents the i th event's coordinate information, represents the logarithmic temporal feature of each event, represents the end time of the current time window, represents the timestamps of events that occurred in the past within the time window. This formula represents the feature value of pixel position in the th time window , and its value is obtained by aggregating all events within the time window. These events are processed separately according to the polarity of the event (positive or negative polarity), represents the total number of events within the time window, used as a normalization factor to average the feature contributions of each event, The function represents 1 when the position of the event is the same as the target pixel position , otherwise 0.
[0030] Since the resolution of the Prophesee EVK4 is 1280×720, directly processing high-resolution data will result in excessive computational complexity, affecting the training efficiency and real-time performance of the model. Therefore, we use bilinear interpolation to downsample the event frame and convert it to a standard size of 224×224 to adapt to the input of the deep learning network.
[0031] S3. Based on the event frame sequence, optimize the temporal multi-scale features using an adaptive decay factor, specifically: Project the event frame sequence into a high-dimensional feature space; Based on the projection result, by introducing a learnable decay factor , the high-dimensional features of the current time window at each time step with the high-dimensional features of the previous time step are fused to form a time multi-scale multi-feature representation, completing the optimization of time multi-scale features.
[0032] In this embodiment, after obtaining the event frame sequence, it is necessary to further optimize the time features to better capture short-term and long-term sign language change patterns. In this embodiment, an adaptive decay factor is designed to dynamically adjust the influence of historical information on the current time features, thereby optimizing the time multi-scale features.
[0033] To process each time series, it is projected into a latent feature space through a convolution operation, thereby avoiding problems such as excessive sparsity or local feature loss that may occur when directly processing the original event data. First, the event frame sequence is projected into a high-dimensional feature space to fully mine the information in the time dimension: .
[0034] To capture the time multi-scale features between events, not only the information within the current time is considered, but also the decaying influence of the previous time is introduced, so that historical information can be gradually transmitted to subsequent time series. For this purpose, an adaptive decay factor is designed, which can dynamically adjust the influence of past event information on the current features, thereby adaptively learning the time relationship between event sequences.
[0035] Aiming at the temporal dependence characteristics of different sign language actions, an adaptive decay factor (Adaptive DecayFactor) is designed to dynamically adjust the weight of historical information, so that the model can balance short-term dynamic changes and long-term temporal dependence relationships. Specifically, at each time step, the features of the current time window need to be fused with the features of the previous time step to form a time multi-scale feature representation. A learnable decay factor is introduced to enable historical features to adaptively affect the feature calculation of the current time step: ; where represents the time multi-scale multi-feature representation of the n th time window.
[0036] By combining the decayed historical features and the convolutional results of the current input, this feature representation can simultaneously reflect the time evolution of event data and the current information. This method enables the model to dynamically adjust its attention to recent events while retaining relevant historical context, thereby improving the capture of time dependence and enhancing the overall feature extraction ability. Through this mechanism, the model can integrate information at different time scales, thereby enhancing the modeling ability of time dependence and enabling the system to more accurately capture sign language actions at different rates.
[0037] S4. Based on the optimized time multi-scale features, use the Feature Pyramid Network to extract spatial multi-scale features, specifically as follows: Extract multi-scale spatial features according to the optimized time multi-scale features; Use upsampling and addition operations to fuse the multi-scale spatial features, and splice them to obtain the spatial multi-scale feature representation, completing the extraction of the spatial multi-scale features.
[0038] In this embodiment, in order to further enhance the spatial representation of features, a feature pyramid structure is introduced, as Figure 2 shown. The feature pyramid can extract and fuse features at different spatial scales, thereby effectively capturing multi-level spatial information. This method enables the gradual extraction of spatial details from coarse to fine and combines these details with temporal features, thus generating a more comprehensive feature representation. On the basis of optimizing the temporal features, spatial features are further extracted to better identify the structural information of sign language actions. For this purpose, this embodiment uses the Feature Pyramid Network (FPN) for spatial multi-scale feature extraction.
[0039] First, input the optimized time feature sequence into ResNet-34, and extract multi-scale spatial features layer by layer. That is, the spatial multi-scale feature extraction process is: input the feature map obtained through the time multi-scale encoding process into the pyramid structure. In this module, first splice the event sequences spanning different time scales, and then extract features layer by layer through the residual block structure to generate the corresponding feature map for each layer: .
[0040] Subsequently, in order to fuse spatial information at different scales, upsampling and addition operations are used for multi-scale feature fusion. That is, first use a convolution operation with size to align the channel dimensions of the feature maps at each scale, ensuring that features from different layers can be effectively integrated. Subsequently, for the low-resolution feature maps, they are enlarged to the same spatial scale as the high-resolution feature maps through upsampling operations, and addition operations are used for feature fusion. This process not only retains the global information in the low-resolution features but also combines it with the local details in the high-resolution features. Through layer-by-layer feature fusion, the model can simultaneously capture the fine-grained details and global context features in the event stream, thereby further enhancing the understanding of complex spatial structures and dynamic changes, .
[0041] Finally, the spatial multi-scale feature representation obtained by splicing is: ; where represents the spatial multi-scale feature representation, Indicates a splicing operation, Indicates an upsampling operation, Indicates the Multi-scale spatial features after the fusion of feature maps of the Indicates the hierarchical index in multi-scale feature extraction, Indicates a convolution operation with a convolution kernel size of 1, And Respectively indicate the Output feature maps after the processing of the residual blocks of the layer and the Indicates the element-wise addition operation of the feature maps, Indicates the Residual convolution operation of the Indicates the n Time multi-scale feature representation of the
[0042] The present invention adopts four main stages ( ) of ResNet-34 to extract feature maps. Each stage represents a unique feature extraction level, where the spatial resolution gradually decreases while the semantic information becomes more and more abstract. These four stages provide multi-scale feature maps that can capture both fine-grained details and reflect high-level context information, thereby achieving effective feature integration between different levels and scales, obtaining spatio-temporal multi-scale features of the event stream. This way can capture both local detail information and overall spatial structure of sign language gestures, which helps to improve the recognition accuracy of the system.
[0043] S5. According to the extracted spatial multi-scale features, use a fully connected layer to obtain the recognition result of isolated sign language words, specifically: Perform global average pooling on the extracted spatial multi-scale features; Flatten the result after average pooling; According to the flattened feature vector, use a fully connected layer for classification to obtain the recognition result of isolated sign language words.
[0044] In this embodiment, in order to use it for classification tasks, it is necessary to further process the feature maps. After completing spatio-temporal feature extraction, it is necessary to input it into a classification network to finally recognize isolated sign language words. First, perform global average pooling on the feature maps : .
[0045] The result obtained after the pooling operation is a one-dimensional vector with a length of , flatten the pooled features and keep them in the form of a one-dimensional vector, that is, when flattening the feature maps output by the pyramid , Before entering the fully connected layer, a global average pooling operation is first performed on it. Global average pooling extracts global information by taking the average of all spatial positions of each feature map, thereby compressing the spatial dimension of the feature map and obtaining the global features of each channel: .
[0046] The flattened feature vector is fed into the fully connected layer for classification, that is, the flattened feature vector is input into the fully connected layer. The fully connected layer maps the features to the classification space and completes the non-linear transformation through the activation function. Finally, the Softmax activation function is used to convert the output into the probability distribution of each category to complete the classification task: ; where represents the recognition result of isolated sign language words, represents the Softmax activation function, represents the weight matrix of the second fully connected layer, represents the ReLU function, represents the weight matrix of the first fully connected layer, represents the flattened feature vector, and represent the bias terms of the first fully connected layer and the second fully connected layer respectively, represents the flattening operation, which converts the high-dimensional feature map into a one-dimensional vector, represents the pooled feature map, represents the c pooling result of the th channel, represents the height of the feature map, represents the spatial coordinates in the feature map, represents the th channel of the spatial multi-scale feature map, represents the channel index in the feature map, represents the number of channels of the feature map.
[0047] The Softmax layer is used to output the probability distribution of 22 sign language words and finally determine the sign language category.
[0048] In summary, through the time dynamic modeling of event stream data and spatial multi-scale feature extraction, the present invention significantly improves the representation ability of complex spatio-temporal relationships, solves the problems of insufficient modeling of complex spatio-temporal characteristics of event camera data and weak generalization ability in the prior art, and at the same time overcomes the limitations of traditional frame-based methods such as high data redundancy and motion blur, and improves the accuracy of isolated sign language word recognition.
Claims
1. A method for isolated sign language word recognition based on event camera spatiotemporal multi-scale feature fusion, characterized in that: The following steps are involved: S1, using event camera to collect asynchronous event stream data of sign language actions; S2, based on the collected asynchronous event stream data of sign language movements, generate event frame sequences using a logarithmic perception mechanism; S3, based on the event frame sequence, the time multi-scale features are optimized using adaptive attenuation factors; S4. Based on the optimized temporal multi-scale features, the feature pyramid network is used to extract spatial multi-scale features; S5. Based on the extracted spatial multi-scale features, the fully connected layer is used to obtain the isolated sign language word recognition results.
2. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 1 is characterized in that: The expression of the asynchronous event stream data is as follows: in, Indicates the asynchronous event stream data of collecting sign language actions. Indicates i events, represents the total number of event streams, Represents pixel coordinate information, Indicates the time when the event occurred. Indicates polarity information. It is a positive event, indicating an increase in pixel light intensity. It is a negative event, indicating a decrease in pixel intensity.
3. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 1 is characterized in that: The S2 is specifically: Based on the collected asynchronous event stream data of sign language movements, the data is divided into multiple windows according to time, and the logarithmic time feature of each event is calculated in each time window; According to the logarithmic time feature of each event, the event feature of each pixel position is aggregated to generate an event frame; The generated event frames are downsampled to obtain an event frame sequence.
4. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 3 is characterized in that: The expression of the event frame is as follows: in, Indicated in Time window, located at pixel position The characteristic value of Represents the pixel position coordinates, Indicates time window, represents the total number of events at coordinate locations within the time window, Indicates i events, represents all events in the time window, classified by polarity, Represents a step function, which is 1 when the coordinates of the event coincide with the target pixel position, otherwise it is 0. Indicates i The coordinate information of each event, represents the logarithmic time characteristics of each event, Indicates the end time of the current time window. A timestamp representing an event that occurred in the past within the time window.
5. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 1 is characterized in that: The S3 is as follows: Project the event frame sequence into a high-dimensional feature space; Based on the projection results, a learnable attenuation factor is introduced , the high-dimensional features of the current time window at each time step High-dimensional features with the previous time step The fusion is performed to form a temporal multi-scale and multi-feature representation and complete the optimization of temporal multi-scale features.
6. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 5 is characterized in that: The expression of the time multi-scale multi-feature representation is as follows: in, Indicates n Temporal multi-scale feature representation of time windows.
7. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 1 is characterized in that: The S4 is specifically as follows: Extract multi-scale spatial features based on the optimized time multi-scale features; The multi-scale spatial features are fused by upsampling and addition operations, and then concatenated to obtain the spatial multi-scale feature representation, thus completing the extraction of spatial multi-scale features.
8. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 7 is characterized in that: The expression of the spatial multi-scale feature representation is as follows: in, Representation space multi-scale feature representation, Represents a splicing operation, represents the upsampling operation, Indicates The multi-scale spatial features after the fusion of layer feature maps, Represents the hierarchical index in multi-scale feature extraction, represents a convolution operation with a convolution kernel size of 1. and Respectively represent Layer and The output feature map after the residual block of the layer is processed. represents the element-wise addition operation of the feature map, Indicates The residual convolution operation of the layer, Indicates n Temporal multi-scale feature representation of time windows.
9. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 1, characterized in that: The S5 is specifically as follows: Perform global average pooling on the extracted spatial multi-scale features; Flatten the result of average pooling; According to the flattened feature vector, the fully connected layer is used for classification to obtain the isolated sign language word recognition result.
10. The isolated sign language word recognition method based on event camera spatiotemporal multi-scale feature fusion according to claim 9, characterized in that: The expression of the isolated sign language word recognition result is as follows: in, represents the isolated sign language word recognition result, represents the Softmax activation function, represents the weight matrix of the second fully connected layer, represents the ReLU function, represents the weight matrix of the first fully connected layer, represents the flattened eigenvector, and Represent the bias terms of the first fully connected layer and the second fully connected layer respectively, Represents the flattening operation, which converts the high-dimensional feature map into a one-dimensional vector. represents the feature map after pooling, Indicates c The pooling result of channels is represents the height of the feature map, represents the width of the feature map, represents the spatial coordinates in the feature map, Represents the spatial multi-scale feature map channels, represents the channel index in the feature map, Indicates the number of channels of the feature map.
Citation Information
Cited By
Visual multi-mode non-contact gesture unlocking method
CN120526487A