Fingertip detection method and device based on event camera, equipment and storage medium

By using an event camera-based method in fingertip detection, the interference problems of complex backgrounds, light changes and dynamic scenes on fingertip detection are solved, achieving higher detection accuracy and tracking capabilities for fast-changing gestures.

CN119942588APending Publication Date: 2025-05-06ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411996628.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is susceptible to interference from factors such as complex backgrounds, ambient light changes, and motion blur in dynamic scenes in fingertip detection, resulting in a decrease in detection accuracy.

Method used

Using the fingertip detection method based on the event camera, the target event image of the hand is collected through the event camera, and the feature extraction module, Mamba Mamba module and detection head module in the fingertip detection model are used to extract the image, adjust the state space parameter and recognize the fingertip to achieve high-accurate fingertip detection.

Benefits of technology

By replacing the RGB camera, avoiding interference with the fingertip recognition image by lighting changes, complex backgrounds and dynamic scenes, significantly improving the accuracy of fingertip detection and improving the tracking ability of fast-changing gestures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942588A_ABST
    Figure CN119942588A_ABST
Patent Text Reader

Abstract

The invention discloses a fingertip detection method and device based on an event camera, computer equipment and a computer readable storage medium, according to the fingertip detection method, an RGB camera is replaced by the event camera, interference of factors such as illumination change, a complex background and a dynamic scene on a fingertip recognition image is avoided, the fingertip detection difficulty is reduced, and the fingertip detection efficiency is improved. And the fingertip detection accuracy is improved. According to the application, fingertip detection is further carried out on a target event image collected by an event camera according to a fingertip detection model of a Mama module, and SSM layer parameters of the fingertip detection model are dynamically adjusted by using a selection mechanism of the Mama module and feature data corresponding to the target event image, so that feature extraction is carried out through an adjusted SSM layer; according to the method, fine detection of fingertip small targets is realized, and the SSM layer is utilized to extract the time sequence association information of the target event image, so that the tracking capability of rapid change gestures is effectively improved, and the success rate and accuracy of fingertip recognition are further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a fingertip detection method, device, computer equipment and computer-readable storage medium based on an event camera. Background Art

[0002] With the development of science and technology, fingertip detection technology has been applied in more and more fields, such as virtual reality (VR), augmented reality (AR) and sign language recognition. The main purpose of fingertip detection is to accurately locate the key points of the hand, especially to calibrate the end points of each finger. By locating the fingertip position, the hand posture, finger direction and overall shape of the hand can be further inferred, thereby providing data support for applications such as gesture recognition, touch control, and virtual object operation. Fingertip detection usually combines target detection and key point positioning technology to accurately describe the state of the hand in different application scenarios.

[0003] However, the images currently captured by cameras are easily disturbed by factors such as complex backgrounds, changes in ambient light, and motion blur in dynamic scenes. Therefore, how to improve the accuracy of fingertip detection has become an urgent problem to be solved. Summary of the invention

[0004] The present application provides a fingertip detection method, apparatus, computer device and computer-readable storage medium based on an event camera to improve the accuracy of fingertip detection.

[0005] In a first aspect, the present application provides a fingertip detection method based on an event camera, the method comprising:

[0006] Capturing a target event image of a hand through an event camera, and performing feature extraction on the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map;

[0007] Determine a target parameter of a state space SSM layer in the Mamba module based on the convolution layer of the Mamba module of the fingertip detection model and the first feature map, and generate a second feature map based on the target parameter and the SSM layer;

[0008] Based on the detection head module of the fingertip detection model, fingertip recognition is performed on the second feature map to complete the fingertip detection of the target event image.

[0009] In a second aspect, the present application further provides a fingertip detection device based on an event camera, the device comprising:

[0010] A first feature extraction module is used to collect a target event image of a hand through an event camera, and extract features from the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map;

[0011] A second feature extraction module, configured to determine a target parameter of a state space SSM layer in the Mamba module based on a convolution layer of the Mamba module of the fingertip detection model and the first feature map, and generate a second feature map based on the target parameter and the SSM layer;

[0012] The image fingertip detection module is used to perform fingertip recognition on the second feature map based on the detection head module of the fingertip detection model to complete the fingertip detection of the target event image.

[0013] In a third aspect, the present application also provides a computer device, comprising a memory and a processor; the memory is used to store a computer program; the processor is used to execute the computer program and implement the above-mentioned event camera-based fingertip detection method when executing the computer program.

[0014] In a fourth aspect, the present application further provides a computer-readable storage medium storing a computer program, which, when executed by a processor, enables the processor to implement the above-mentioned event camera-based fingertip detection method.

[0015] The present application discloses a fingertip detection method based on an event camera, the fingertip detection method comprising: collecting a target event image of the hand through an event camera, and extracting features from the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map; determining the target parameters of the state space SSM layer in the Mamba module based on the convolution layer of the Mamba module of the fingertip detection model and the first feature map, and generating a second feature map based on the target parameters and the SSM layer; performing fingertip recognition on the second feature map based on a detection head module of the fingertip detection model to complete fingertip detection of the target event image. The present application replaces the RGB camera with an event camera to avoid interference of factors such as lighting changes, complex backgrounds, and dynamic scenes on the fingertip recognition image, thereby reducing the difficulty of fingertip detection and improving the accuracy of fingertip detection. The present application further performs fingertip detection on the target event image captured by the event camera based on the fingertip detection model of the Mamba module. Compared with the fixed model parameters in the traditional model, the present application utilizes the selection mechanism of the Mamba module and the feature data corresponding to the target event image to determine the SSM layer parameters in the Mamba module, realizes dynamic adjustment of the SSM layer parameters, and performs feature extraction through the adjusted SSM layer to realize fine detection of small targets such as fingertips, and uses the SSM layer to extract the temporal correlation information of the target event image, which effectively improves the tracking capability of rapidly changing gestures and further improves the success rate and accuracy of fingertip recognition. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0017] Figure 1 is a schematic flow chart of a fingertip detection method based on an event camera provided in the first embodiment of the present application;

[0018] Figure 2 is a schematic flow chart of a fingertip detection method based on an event camera provided in the second embodiment of the present application;

[0019] Figure 3 A schematic diagram of a processing flow based on a fingertip detection model provided in an embodiment of the present application;

[0020] Figure 4 A schematic diagram of feature fusion of a multi-scale fusion submodule provided in an embodiment of the present application;

[0021] Figure 5A schematic block diagram of a fingertip detection device based on an event camera provided in an embodiment of the present application;

[0022] Figure 6 A schematic block diagram of the structure of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0023] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0024] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0025] It should be understood that the terms used in this application specification are only for the purpose of describing specific embodiments and are not intended to limit the application. As used in this application specification and the appended claims, unless the context clearly indicates otherwise, the singular forms "a", "an" and "the" are intended to include plural forms.

[0026] It should be further understood that the term “and / or” used in the specification and appended claims refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0027] In conjunction with the accompanying drawings, some embodiments of the present application are described in detail below. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0028] See also Figure 1 , Figure 1 It is a schematic flow chart of a fingertip detection method based on an event camera provided in an embodiment of the present application.

[0029] The images currently captured by cameras are easily disturbed by factors such as complex backgrounds, changes in ambient light, and motion blur in dynamic scenes. Specifically:

[0030] In strong light, backlight or low light conditions, the quality of hand images captured by the RGB camera will be greatly reduced, thus affecting the accuracy of fingertip recognition results. For example, in an environment with uneven lighting, the hand image may be overexposed or underexposed, resulting in missing information during the detection process, making it difficult to accurately identify the fingertip position.

[0031] In a complex background environment, the hand and the background are likely to have similar colors or textures, making it difficult to distinguish the hand area from the background, increasing the difficulty of fingertip detection. Especially in outdoor scenes where there are a lot of dynamic elements in the background, it not only increases the difficulty of fingertip detection, but also reduces the accuracy of fingertip detection.

[0032] In the case of fast movement, RGB images are prone to motion blur. Therefore, when the hand moves quickly, the RGB image captured by the camera will appear blurred, making it impossible for the model to effectively extract hand features, thus affecting the recognition results of the fingertip position and reducing the accuracy of fingertip detection.

[0033] In order to solve the above problems, the present application embodiment proposes a fingertip detection method based on an event camera. The fingertip detection method based on an event camera provided in this embodiment can be applied to smart wearable devices such as VR devices or AR devices, and can also be applied to sign language recognition devices. In this embodiment, the event camera replaces the RGB camera to avoid the interference of factors such as lighting changes, complex backgrounds, and dynamic scenes on the fingertip recognition image, thereby reducing the difficulty of fingertip detection and improving the accuracy of fingertip detection.

[0034] like Figure 1 As shown, the fingertip detection method based on event camera specifically includes steps S101 to S103.

[0035] S101, collecting a target event image of a hand through an event camera, and performing feature extraction on the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map;

[0036] In one embodiment, an event camera is used to capture target event images. The event camera, which has low power consumption, can effectively filter interference in the scene, and has high sensitivity and rapid response to dynamic changes, replaces the traditional RGB camera for capturing images. This solves the fingertip detection problem of the traditional RGB camera in lighting changes, complex backgrounds and dynamic scenes.

[0037] An event camera is a sensor that can capture fast dynamic scenes and generate event image data by detecting pixel-level brightness changes. When the brightness change of a pixel in the scene reaches a preset threshold, an event is immediately triggered, recording the location of the pixel, timestamp, and polarity of the brightness change (brightening or darkening).

[0038] After the target event image is collected, fingertip recognition is further performed through a fingertip detection model including a Mamba module. First, a feature extraction module of the fingertip detection model is used to extract features of the target event image to obtain a first feature map.

[0039] Specifically, the target event image can be first convolved through the depth-wise separable convolution layer (DW-Conv) in the feature extraction module, and then the multi-scale feature fusion of the convolved feature map is performed through the multi-scale module (MS-block) in the feature extraction module to obtain the first feature map.

[0040] S102, based on the convolution layer of the Mamba module of the fingertip detection model and the first feature map, determine the target parameters of the state space SSM layer in the Mamba module, and generate a second feature map based on the target parameters and the SSM layer;

[0041] In this embodiment, after the first feature map is extracted, the key area features in the first feature map are further enhanced through the Mamba module in the fingertip detection model, which can effectively focus on the key area and achieve fine detection of small targets such as fingertips. At the same time, the timing correlation information is fully utilized to improve the tracking ability of rapidly changing gestures. Compared with the traditional YOLO detection model, the Mamba module has lower computational complexity and power consumption, and is more suitable for deployment in embedded devices and end-side devices, thereby effectively improving the accuracy and efficiency of fingertip detection.

[0042] Traditional detection models use unchanged model parameters when processing sequences. Traditional detection models are suitable for static patterns with small changes, but are not suitable for processing complex changing sequence data such as rapidly changing gestures. This embodiment introduces a selection mechanism through the Mamba module, so that the model parameters that affect sequence interaction in the fingertip detection model are no longer fixed, and can be dynamically adjusted according to the sequence corresponding to the input first feature map, so that the fingertip detection model based on the Mamba module can adaptively adjust the corresponding SSM layer parameters according to different input feature sequences. Through dynamic parameter adjustment, the fingertip detection model can achieve time variability when processing sequence data, that is, the model output can change with the time step, so that it is more accurate and flexible when processing dynamically changing sequence data.

[0043] Specifically, the Mamba module includes an SSM (State-Space Model Layer) layer, a convolutional layer, and a gated multilayer perceptron (MLP) submodule, in which the SSM layer, convolutional layer, and gated MLP submodule are arranged in sequence. Each module is connected through normalization and residual connections to improve the stability and optimization effect of the Mamba module.

[0044] The first feature map is converted into a state sequence suitable for SSM layer processing through the convolution layer, and the target parameters of the SSM layer are determined according to the convolution result. The target parameter states include state transfer matrix, control matrix, and observation matrix. The SSM layer convolves the input vector (i.e., the state sequence after the first feature map of the convolution layer is converted) according to the target parameters to obtain the output vector (i.e., the second feature map).

[0045] By adding an SSM layer to the Mamba module, that is, by adding a state space model, a multi-level state tracking capability is built inside the Mamba module, so that the Mamba module can effectively manage the context of long sequence data corresponding to the moving fingertips.

[0046] After determining the target parameters of the SSM layer, the SSM layer corresponding to the target parameters is used to perform state tracking and recursive prediction on the feature sequence data processed by the convolution layer, so that the fingertip detection model can continuously store the preceding information of the fingertip movement and obtain the context information of the fingertip movement. Recursive prediction is performed based on the feature sequence data of the target event image and the context of the fingertip movement to obtain the second feature map.

[0047] Furthermore, after obtaining the second feature map, the second feature map can be input into the gated MLP submodule for feature semantic enrichment and effective information screening. The gated MLP submodule consists of a multi-layer perceptron (MLP) structure and a gating mechanism. The second feature map output by the SSM layer is mapped to a high-dimensional feature space through the MLP structure to enrich the semantic information of the second feature map, and the feature capture capability of the Mamba module is deepened through multiple fully connected layers. The gating structure is introduced through the gating mechanism, so that the Mamba module can selectively pass or filter out specific information in the input feature sequence data, so that the fingertip detection model can effectively focus on the key area when processing the long sequence data corresponding to the fingertip changes.

[0048] S103: Based on the detection head module of the fingertip detection model, perform fingertip recognition on the second feature map to complete the fingertip detection of the target event image.

[0049] In this embodiment, after multiple feature extraction and fusion of the target event image, the final second feature map is transmitted to the detection head module. The detection head module determines whether the prior box on the feature point contains the target object (i.e., the fingertip), thereby determining whether there is a fingertip at the position, and completing the fingertip detection of the target event image.

[0050] Furthermore, the detection head module based on the fingertip detection model performs fingertip recognition on the second feature map to complete the fingertip detection of the target event image, specifically comprising:

[0051] Based on the detection head module, fingertip detection is performed on the second feature map, and the position coordinates of the fingertip in the target event image are determined, and a fingertip recognition result is generated based on the position coordinates.

[0052] In this embodiment, through the detection head module, based on the features corresponding to the pre-stored fingertip image, the fingertip position is determined in the second feature map, thereby determining the predicted information of the fingertip position in the target event image, and using the fingertip position as the fingertip recognition result; if the fingertip does not exist in the second feature map, the reminder information that there is no fingertip in the target event image is output as the fingertip recognition result.

[0053] The above embodiment provides a fingertip detection method based on an event camera, the fingertip detection method comprising: collecting a target event image of the hand through an event camera, and extracting features from the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map; determining the target parameters of the state space SSM layer in the Mamba module based on the convolution layer of the Mamba module of the fingertip detection model and the first feature map, and generating a second feature map based on the target parameters and the SSM layer; performing fingertip recognition on the second feature map based on a detection head module of the fingertip detection model to complete fingertip detection of the target event image. The present application replaces the RGB camera with an event camera to avoid interference of factors such as lighting changes, complex backgrounds, and dynamic scenes on the fingertip recognition image, thereby reducing the difficulty of fingertip detection and improving the accuracy of fingertip detection. The present application further performs fingertip detection on the target event image captured by the event camera based on the fingertip detection model of the Mamba module. Compared with the fixed model parameters in the traditional model, the present application utilizes the selection mechanism of the Mamba module and the feature data corresponding to the target event image to determine the SSM layer parameters in the Mamba module, realizes dynamic adjustment of the SSM layer parameters, and performs feature extraction through the adjusted SSM layer to realize fine detection of small targets such as fingertips, and uses the SSM layer to extract the temporal correlation information of the target event image, which effectively improves the tracking capability of rapidly changing gestures and further improves the success rate and accuracy of fingertip recognition.

[0054] See also Figure 2 , Figure 2 It is a schematic flow chart of a fingertip detection method based on an event camera provided in an embodiment of the present application.

[0055] like Figure 2 As shown, the fingertip detection method based on event camera also includes steps S104 to S106.

[0056] S104, performing point-by-point convolution on the depth convolution feature map based on the depth separable convolution submodule to obtain the initial feature map;

[0057] S105, based on the data set construction module of the fingertip detection model, calculating the pixel brightness change value corresponding to each pixel of the front and back frame RGB images, and generating the simulation event image corresponding to the front and back frame RGB images based on the pixel point set whose brightness change value exceeds a preset threshold;

[0058] S106. Based on the data set construction module, generate a model training set corresponding to the simulation event image set, and based on the model training set, train and generate the fingertip detection model.

[0059] Since event images are not convenient for generating image labels or adding image annotations, it is difficult for the fingertip detection model to obtain sufficient data samples. In order to solve the above problem, in this embodiment, the RGB image is converted into a simulated event image to supplement the training data set.

[0060] Specifically, a series of RGB image sets corresponding to the moving fingertips are first obtained through an RGB camera. The brightness changes between two consecutive frames in the RGB image set are analyzed to calculate the offset of the pixel brightness. Based on the brightness offset, simulated event data similar to that of a real event camera is generated.

[0061] Among them, simulated event images are used to convert ordinary RGB images into a form similar to the output of an event camera, so as to supplement the training dataset in the absence of real event image data.

[0062] Specifically, after the RGB image set is acquired, any two consecutive preceding and following frame RGB images in the RGB image set are sequentially acquired for image conversion until all RGB images in the RGB image set are converted into corresponding event graphics.

[0063] Based on the dataset construction module, the pixel brightness values ​​in the grayscale images corresponding to the two consecutive RGB images are compared to calculate the brightness change of each pixel over time. Through this pixel-by-pixel difference calculation, the increase or decrease of the color value at each RGB image position is captured, thereby effectively simulating the response characteristics of the event camera to the brightness change during the object movement.

[0064] Then, the brightness change value of each pixel point is compared with a preset threshold value, which is a preset brightness change threshold value. Based on the preset brightness change threshold value, a dynamic pixel point set with a large brightness change is screened out from the RGB image, and a simulated event image set is generated based on the dynamic pixel point set. A model training set is generated based on the simulated event image set.

[0065] Through the above method, this embodiment uses RGB data to generate a large number of simulated event images. After the RGB images are annotated, they can be converted into corresponding event images, thereby solving the problem of difficulty in event image annotation, and providing a flexible and low-cost event image data acquisition method, which can quickly generate rich simulated event data for the training of fingertip detection models. It also improves the diversity of training data and significantly enhances the generalization ability of fingertip detection models in diverse scenarios. The fingertip detection model based on the Mamba module in this embodiment is trained using the data set constructed by the above method, and through multi-source data fusion, accurate detection of hand fingertips is achieved. The fingertip detection method based on the event camera provided in this embodiment is not only low in energy consumption and high in precision, but also fast in response, and is easy to deploy on smart wearable devices with limited computing resources such as AR glasses and smart helmets.

[0066] Furthermore, the S105 specifically includes:

[0067] Based on the three-channel grayscale values ​​of the front and back frame RGB images, the front and back frame RGB images are respectively converted into corresponding grayscale images;

[0068] Based on the grayscale images corresponding to the front and back frame RGB images, a pixel brightness change value of each pixel point in the front and back frame RGB images is calculated;

[0069] The pixel brightness change value of each pixel point is compared with a preset threshold value respectively, and based on the comparison result of the pixel brightness change value of each pixel point, the pixel point set whose pixel brightness change value exceeds the preset threshold value is determined in the grayscale image corresponding to the previous and next frame RGB images, and based on the pixel point set whose pixel brightness change value exceeds the preset threshold value, the simulated event image is generated.

[0070] In this embodiment, first, the front and back frame RGB images are converted into grayscale images based on the three-channel grayscale values ​​of the front and back frame RGB images. The RGB image contains three color channels, and the grayscale value of each pixel is calculated based on the weighted sum of the three channels. The grayscale value calculation formula is as follows:

[0071] Grayscale (i,x,y) =0.299×R+0.587×G+0.114×B

[0072] Grayscale (i,x,y) is the grayscale pixel value of the pixel at position (x, y) of the i-th frame, R is the R channel value of the pixel at position (x, y) of the i-th frame, G is the G channel value of the pixel at position (x, y) of the i-th frame, and B is the B channel value of the pixel at position (x, y) of the i-th frame.

[0073] Then, by calculating the pixel brightness change value of each pixel in the previous and next frame RGB images, the brightness change of each pixel over time is calculated. The calculation formula of the pixel brightness change value is as follows:

[0074] diff(x,y)=Grayscale (i,x,y) -Grayscale (i-1,x,y)

[0075] Grayscale (i-1,x,y) is the grayscale pixel value of the pixel at position (x, y) in the previous frame of the i-th frame, and diff(x, y) is the brightness change value of the pixel at position (x, y).

[0076] Through the above pixel-by-pixel difference calculation, the increase or decrease of the color value at each pixel position is captured, thereby effectively simulating the response characteristics of the event camera to the brightness change during the movement of the object.

[0077] Then, based on the preset brightness change threshold, a simulated event image is generated. For each pixel, if the brightness change exceeds the set threshold, the pixel is marked as increasing or decreasing in brightness, and a conversion event is triggered to generate a simulated event image. The marking process is as follows:

[0078] event(x,y)=[0,255,0],ifdiff(x,y)>threshold

[0079] event(x,y)=[0,0,255],ifdiff(x,y)<-threshold

[0080] Among them, event(x,y) is a marked event, threshold is the preset threshold, diff(x,y)>threshold indicates a positive event, diff(x,y)<-threshold indicates a negative event, [0,255,0] indicates green, that is, the pixels corresponding to positive events are marked as green, and [0,0,255] indicates red, that is, the pixels corresponding to negative events are marked as red.

[0081] The pixel points whose brightness changes exceed the threshold are marked in red or green, and the other pixel points whose brightness changes do not exceed the threshold are marked in black, thereby obtaining a simulated event image.

[0082] It is understandable that the pixel points whose brightness changes exceed the threshold value can also be marked with other colors, which can be set by the user as needed.

[0083] In the above way, the simulated event images obtained through conversion and the real data of the event camera are used to jointly construct a diverse data set, providing richer training samples for the fingertip detection model. This not only solves the problem of insufficient event data collection, but also further improves the generalization ability of the model.

[0084] Furthermore, the feature extraction module includes a depth-separable convolution submodule and a multi-scale fusion submodule, and the feature extraction of the target event image based on the feature extraction module to obtain the first feature map specifically includes:

[0085] Extracting features of the target event image based on the depth-separable convolution submodule to obtain an initial feature map;

[0086] Based on the multi-scale fusion submodule, the initial feature map is divided into regions to obtain at least one feature region map;

[0087] Based on the multi-scale fusion submodule, convolution operations are performed on each feature region map respectively, and each feature region map after convolution is fused to obtain the first feature map.

[0088] In this embodiment, the target event image is convolved by depthwise separable convolution (DW-Conv) to reduce the computational complexity, and the multi-scale feature fusion is performed on the above convolution features by multi-scale module (MS-block). The key areas in the fused features are then enhanced by Mamba module, thereby effectively improving the accuracy and efficiency of fingertip detection and the accuracy of small target detection.

[0089] Specifically, after the target event image is input into the backbone feature extraction network (Backbone) for processing, Figure 3As shown in , the target event image is first convolved (Conv), and then the convolved image is further convolved through a depth-wise separable convolution to achieve preliminary feature extraction and obtain an initial feature map. Figure 4 As shown in the figure, the initial feature map is divided into four sub-regions (x1, x2, x3, x4) through the multi-scale fusion sub-module MS-block sub-module, and the feature map corresponding to each sub-region is independently convolved, and then the features of the above sub-regions are gradually spliced ​​and fused to extract multi-scale features with richer semantic information, thereby completing the extraction of the first feature map.

[0090] It can be understood that in order to improve the accuracy of feature extraction, the features of two sub-areas can be first convolved and then spliced, and the features of the other two sub-areas can be first spliced ​​and then convolved, and the spliced ​​features can continue to be spliced ​​and convolved, and the first feature map can be updated until the accuracy of the feature map reaches the preset accuracy threshold, and the fusion is stopped to output the first feature map.

[0091] Therefore, through the above-mentioned multi-scale fusion method, the ability of the fingertip detection model to capture target details is improved, and it is particularly suitable for detecting small targets such as fingertips.

[0092] Furthermore, the step of extracting features from the target event image based on the depth-separable convolution submodule to obtain an initial feature map specifically includes:

[0093] Performing depth convolution on the target event image based on the depth separable convolution submodule to obtain a depth convolution feature map;

[0094] Based on the depthwise separable convolution submodule, the depthwise convolution feature map is convolved point by point to obtain the initial feature map.

[0095] In this embodiment, the depth separable convolution submodule decomposes the standard convolution into depth convolution and point-by-point convolution. First, the target event image is subjected to depth convolution to obtain a depth convolution feature map, and then the depth convolution feature map is subjected to point-by-point convolution to obtain the initial feature map. Among them, the depth convolution is to perform a separate convolution operation on each input channel, while the point-by-point convolution is to perform a dot product operation on the result of the depth convolution.

[0096] Through the above method, this embodiment further effectively reduces the model calculation complexity and parameter quantity, makes the detection network more lightweight, and improves the detection efficiency.

[0097] Furthermore, the fingertip detection model further includes an image upsampling module, and before determining the target parameters of the state space SSM layer in the Mamba module based on the convolution layer in the Mamba module and the first feature map, it further includes:

[0098] Based on the image upsampling module of the fingertip detection model, the first feature map is enlarged, and the enlarged first feature map is fused and spliced ​​with the first feature map to obtain a fused feature map;

[0099] Determining a target parameter corresponding to the fused feature map based on the convolution layer, and converting the fused feature map to obtain a target feature sequence;

[0100] Based on the SSM layer corresponding to the target parameter, the target feature sequence is convolved to generate the second feature map.

[0101] In this embodiment, Figure 3 As shown, after the temporal image is convolved and subjected to depth-separable convolution and multi-scale feature fusion, the fused feature map can be output as the first feature map on the one hand, and the fused feature map can be further subjected to iterative processing of depth-separable convolution and multi-scale fusion submodules until the accuracy of the feature map reaches a preset threshold, and the iterative processing is stopped. Then, the iterative first feature map is convolved to integrate information, and the image upsampling module is used to upsample and amplify the first feature map after the convolution to fuse features at different resolutions in the first feature map after the convolution, make up for information loss, and thus capture high-level semantic information and enrich the semantic information in the first feature map. Finally, the upsampled first feature map is concatenated with the output fused first feature map on the other hand to achieve feature integration.

[0102] After determining the target parameters of the SSM layer, the SSM layer corresponding to the target parameters performs state tracking and recursive prediction on the feature sequence data processed by the convolution layer, so that the fingertip detection model can continuously store the preceding information of the fingertip movement and obtain the context information of the fingertip movement. Recursive prediction is performed based on the feature sequence data of the target event image (i.e., the target feature sequence) and the context of the fingertip movement to obtain the second feature map.

[0103] Through the above method, this embodiment splices and fuses the feature maps with different resolutions obtained after processing in different ways, compensates for the missing information, thereby capturing high-level semantic information and enriching the semantic information in the first feature map.

[0104] See also Figure 5 , Figure 5The embodiment of the present application provides a schematic block diagram of a fingertip detection device based on an event camera, and the fingertip detection device based on an event camera is used to execute the aforementioned fingertip detection method based on an event camera. The fingertip detection device based on an event camera can be configured on a server.

[0105] like Figure 5 As shown, the fingertip detection device 500 based on the event camera includes:

[0106] A first feature extraction module 510 is used to collect a target event image of a hand through an event camera, and perform feature extraction on the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map;

[0107] A second feature extraction module 520 is used to determine a target parameter of a state space SSM layer in the Mamba module based on the convolution layer of the Mamba module of the fingertip detection model and the first feature map, and generate a second feature map based on the target parameter and the SSM layer;

[0108] The image fingertip detection module 530 is used to perform fingertip recognition on the second feature map based on the detection head module of the fingertip detection model to complete the fingertip detection of the target event image.

[0109] Furthermore, the event camera-based fingertip detection device 500 further includes a data set generation module, wherein the data set generation module is used to:

[0110] Collecting an RGB image set of the fingertip during movement through an RGB camera, and acquiring any two consecutive RGB images of the preceding and following frames in the RGB image set based on the data set construction module;

[0111] Based on the data set construction module, the pixel brightness change value corresponding to each pixel of the front and back frame RGB images is calculated, and based on the pixel point set whose brightness change value exceeds a preset threshold, the simulated event image corresponding to the front and back frame RGB images is generated;

[0112] Based on the data set construction module, a model training set corresponding to the simulation event image set is generated, and based on the model training set, the fingertip detection model is trained and generated.

[0113] Furthermore, the data set generation module is also used for:

[0114] Based on the three-channel grayscale values ​​of the front and back frame RGB images, the front and back frame RGB images are respectively converted into corresponding grayscale images;

[0115] Based on the grayscale images corresponding to the front and back frame RGB images, a pixel brightness change value of each pixel point in the front and back frame RGB images is calculated;

[0116] Comparing the pixel brightness change value of each pixel point with a preset threshold value respectively, and based on the comparison result of the pixel brightness change value of each pixel point, determining a pixel point set whose pixel brightness change value exceeds the preset threshold value in the grayscale images corresponding to the previous and next frame RGB images;

[0117] The simulated event image is generated based on a set of pixel points whose pixel brightness change values ​​exceed a preset threshold.

[0118] Furthermore, the feature extraction module includes a depth-separable convolution submodule and a multi-scale fusion submodule, and the first feature extraction module 510 specifically includes:

[0119] An initial feature map extraction unit, used to extract features from the target event image based on the depth-separable convolution submodule to obtain an initial feature map;

[0120] A feature sub-region division unit, configured to divide the initial feature map into regions based on the multi-scale fusion sub-module to obtain at least one feature region map;

[0121] The feature sub-region fusion unit is used to perform convolution operations on each feature region map based on the multi-scale fusion sub-module, and fuse the convolved feature region maps to obtain the first feature map.

[0122] Furthermore, the initial feature map extraction unit is also used for:

[0123] Performing depth convolution on the target event image based on the depth separable convolution submodule to obtain a depth convolution feature map;

[0124] Based on the depthwise separable convolution submodule, the depthwise convolution feature map is convolved point by point to obtain the initial feature map.

[0125] Furthermore, the fingertip detection model further includes an image upsampling module. The event camera-based fingertip detection device 500 further includes an image upsampling module for:

[0126] Based on the image upsampling module of the fingertip detection model, the first feature map is enlarged, and the enlarged first feature map is fused and spliced ​​with the first feature map to obtain a fused feature map;

[0127] Determining a target parameter corresponding to the fused feature map based on the convolution layer, and converting the fused feature map to obtain a target feature sequence;

[0128] Based on the SSM layer corresponding to the target parameter, the target feature sequence is recursively predicted to generate the second feature map.

[0129] Furthermore, the image fingertip detection module 530 specifically includes:

[0130] A fingertip detection unit is used to perform fingertip detection on the second feature map based on the detection head module, determine the position coordinates of the fingertip in the target event image based on the detection result, and generate a fingertip recognition result based on the position coordinates.

[0131] It should be noted that those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the above-described device and each module can refer to the corresponding process in the aforementioned method embodiment, and will not be repeated here.

[0132] The above-mentioned device can be implemented in the form of a computer program. Figure 4 Runs on the computer device shown.

[0133] See also Figure 6 , Figure 6 1 is a schematic block diagram of a computer device provided in an embodiment of the present application. The computer device may be a server.

[0134] See also Figure 6 The computer device includes a processor, a memory and a network interface connected through a system bus, wherein the memory may include a non-volatile storage medium and an internal memory.

[0135] The non-volatile storage medium can store an operating system and a computer program. The computer program includes program instructions, and when the program instructions are executed, the processor can execute any one of the event camera-based fingertip detection methods.

[0136] The processor is used to provide computing and control capabilities and support the operation of the entire computer equipment.

[0137] The internal memory provides an environment for the operation of the computer program in the non-volatile storage medium. When the computer program is executed by the processor, the processor can execute any one of the fingertip detection methods based on the event camera.

[0138] The network interface is used for network communication, such as sending assigned tasks, etc. Those skilled in the art will appreciate that Figure 6 The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0139] It should be understood that the processor may be a central processing unit (CPU), and the processor may also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0140] In one embodiment, the processor is used to run a computer program stored in the memory to implement the following steps:

[0141] Capturing a target event image of a hand through an event camera, and performing feature extraction on the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map;

[0142] Determine a target parameter of a state space SSM layer in the Mamba module based on the convolution layer of the Mamba module of the fingertip detection model and the first feature map, and generate a second feature map based on the target parameter and the SSM layer;

[0143] Based on the detection head module of the fingertip detection model, fingertip recognition is performed on the second feature map to complete the fingertip detection of the target event image.

[0144] In one embodiment, the fingertip detection model further includes a data set construction module, and the processor is further configured to implement:

[0145] Collecting an RGB image set of the fingertip during movement through an RGB camera, and acquiring any two consecutive RGB images of the preceding and following frames in the RGB image set based on the data set construction module;

[0146] Based on the data set construction module, the pixel brightness change value corresponding to each pixel of the front and back frame RGB images is calculated, and based on the pixel point set whose brightness change value exceeds a preset threshold, the simulated event image corresponding to the front and back frame RGB images is generated;

[0147] Based on the data set construction module, a model training set corresponding to the simulation event image set is generated, and based on the model training set, the fingertip detection model is trained and generated.

[0148] In one embodiment, when the processor implements the calculation of the pixel brightness change value corresponding to each pixel of the front and back frame RGB images, and generates the simulated event image corresponding to the front and back frame RGB images based on the pixel point set whose brightness change value exceeds the preset threshold, it is used to implement:

[0149] Based on the three-channel grayscale values ​​of the front and back frame RGB images, the front and back frame RGB images are respectively converted into corresponding grayscale images;

[0150] Based on the grayscale images corresponding to the front and back frame RGB images, a pixel brightness change value of each pixel point in the front and back frame RGB images is calculated;

[0151] The pixel brightness change value of each pixel point is compared with a preset threshold value respectively, and based on the comparison result of the pixel brightness change value of each pixel point, the pixel point set whose pixel brightness change value exceeds the preset threshold value is determined in the grayscale image corresponding to the previous and next frame RGB images, and based on the pixel point set whose pixel brightness change value exceeds the preset threshold value, the simulated event image is generated.

[0152] In one embodiment, the feature extraction module includes a depth-separable convolution submodule and a multi-scale fusion submodule. When the processor implements the feature extraction of the target event image based on the feature extraction module to obtain the first feature map, it is used to implement:

[0153] Extracting features of the target event image based on the depth-separable convolution submodule to obtain an initial feature map;

[0154] Based on the multi-scale fusion submodule, the initial feature map is divided into regions to obtain at least one feature region map;

[0155] Based on the multi-scale fusion submodule, convolution operations are performed on each feature region map respectively, and each feature region map after convolution is fused to obtain the first feature map.

[0156] In one embodiment, when the processor extracts features from the target event image based on the depth-separable convolution submodule to obtain an initial feature map, the processor is used to implement:

[0157] Performing depth convolution on the target event image based on the depth separable convolution submodule to obtain a depth convolution feature map;

[0158] Based on the depthwise separable convolution submodule, the depthwise convolution feature map is convolved point by point to obtain the initial feature map.

[0159] In one embodiment, the fingertip detection model further includes an image upsampling module, and the processor implements the convolution layer of the Mamba module based on the fingertip detection model and the first feature map, determines the target parameters of the state space SSM layer in the Mamba module, and generates a second feature map based on the target parameters and the SSM layer, so as to implement:

[0160] Based on the image upsampling module of the fingertip detection model, the first feature map is enlarged, and the enlarged first feature map is fused and spliced ​​with the first feature map to obtain a fused feature map;

[0161] Determining a target parameter corresponding to the fused feature map based on the convolution layer, and converting the fused feature map to obtain a target feature sequence;

[0162] Based on the SSM layer corresponding to the target parameter, the target feature sequence is recursively predicted to generate the second feature map.

[0163] In one embodiment, when the processor implements the detection head module based on the fingertip detection model to perform fingertip recognition on the second feature map to complete the fingertip detection of the target event image, it is used to implement:

[0164] Based on the detection head module, fingertip detection is performed on the second feature map, the position coordinates of the fingertip in the target event image are determined based on the detection result, and a fingertip recognition result is generated based on the position coordinates.

[0165] A computer-readable storage medium is also provided in an embodiment of the present application, wherein the computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and the processor executes the program instructions to implement any one of the event camera-based fingertip detection methods provided in the embodiments of the present application.

[0166] The computer-readable storage medium may be an internal storage unit of the computer device described in the foregoing embodiment, such as a hard disk or memory of the computer device. The computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a smart memory card (SmartMedia Card, SMC), a secure digital (Secure Digital, SD) card, a flash card (Flash Card), etc., equipped on the computer device.

[0167] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in the present application, and these modifications or replacements should be included in the protection scope of the present application. Therefore, the protection scope of the present application shall be based on the protection scope of the claims.

Claims

1. A fingertip detection method based on event camera, characterized in that: The fingertip detection method comprises: Capturing a target event image of a hand through an event camera, and performing feature extraction on the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map; Determine a target parameter of a state space SSM layer in the Mamba module based on the convolution layer of the Mamba module of the fingertip detection model and the first feature map, and generate a second feature map based on the target parameter and the SSM layer; Based on the detection head module of the fingertip detection model, fingertip recognition is performed on the second feature map to complete the fingertip detection of the target event image.

2. The fingertip detection method based on event camera according to claim 1, characterized in that: The fingertip detection method further comprises: Collecting an RGB image set of the fingertip during movement through an RGB camera, and acquiring any two consecutive RGB images of the preceding and following frames in the RGB image set based on the data set construction module; Based on the data set construction module, the pixel brightness change value corresponding to each pixel of the front and back frame RGB images is calculated, and based on the pixel point set whose brightness change value exceeds a preset threshold, the simulated event image corresponding to the front and back frame RGB images is generated; Based on the data set construction module, a model training set corresponding to the simulation event image set is generated, and based on the model training set, the fingertip detection model is trained and generated.

3. The fingertip detection method based on event camera according to claim 1, characterized in that: The step of calculating the pixel brightness change value corresponding to each pixel of the front and back frame RGB images, and generating the simulated event image corresponding to the front and back frame RGB images based on a pixel point set whose brightness change value exceeds a preset threshold comprises: Based on the three-channel grayscale values ​​of the front and back frame RGB images, the front and back frame RGB images are respectively converted into corresponding grayscale images; Based on the grayscale images corresponding to the front and back frame RGB images, a pixel brightness change value of each pixel point in the front and back frame RGB images is calculated; Comparing the pixel brightness change value of each pixel point with a preset threshold value respectively, and based on the comparison result of the pixel brightness change value of each pixel point, determining a pixel point set whose pixel brightness change value exceeds the preset threshold value in the grayscale images corresponding to the previous and next frame RGB images; The simulated event image is generated based on a set of pixel points whose pixel brightness change values ​​exceed a preset threshold.

4. The fingertip detection method based on event camera according to claim 1, characterized in that: The feature extraction module includes a depth-separable convolution submodule and a multi-scale fusion submodule. Based on the feature extraction module, the target event image is subjected to feature extraction to obtain a first feature map, which includes: Extracting features of the target event image based on the depth-separable convolution submodule to obtain an initial feature map; Based on the multi-scale fusion submodule, the initial feature map is divided into regions to obtain at least one feature region map; Based on the multi-scale fusion submodule, convolution operations are performed on each feature region map respectively, and each feature region map after convolution is fused to obtain the first feature map.

5. The fingertip detection method based on event camera according to claim 4, characterized in that: The step of extracting features from the target event image based on the depth-separable convolution submodule to obtain an initial feature map includes: Performing depth convolution on the target event image based on the depth separable convolution submodule to obtain a depth convolution feature map; Based on the depthwise separable convolution submodule, the depthwise convolution feature map is convolved point by point to obtain the initial feature map.

6. The fingertip detection method based on event camera according to claim 1, characterized in that: The convolution layer of the Mamba module of the fingertip detection model and the first feature map, determining the target parameters of the state space SSM layer in the Mamba module, and generating the second feature map based on the target parameters and the SSM layer includes: Based on the image upsampling module of the fingertip detection model, the first feature map is enlarged, and the enlarged first feature map is fused and spliced ​​with the first feature map to obtain a fused feature map; Determining a target parameter corresponding to the fused feature map based on the convolution layer, and converting the fused feature map to obtain a target feature sequence; Based on the SSM layer corresponding to the target parameter, the target feature sequence is recursively predicted to generate the second feature map.

7. The fingertip detection method based on event camera according to any one of claims 1 to 6, characterized in that: The detection head module based on the fingertip detection model performs fingertip recognition on the second feature map to complete the fingertip detection of the target event image, including: Based on the detection head module, fingertip detection is performed on the second feature map, the position coordinates of the fingertip in the target event image are determined based on the detection result, and a fingertip recognition result is generated based on the position coordinates.

8. A fingertip detection device based on an event camera, characterized in that: include: A first feature extraction module is used to collect a target event image of a hand through an event camera, and extract features from the target event image based on a feature extraction module of a fingertip detection model to obtain a first feature map; A second feature extraction module, configured to determine a target parameter of a state space SSM layer in the Mamba module based on a convolution layer of the Mamba module of the fingertip detection model and the first feature map, and generate a second feature map based on the target parameter and the SSM layer; The image fingertip detection module is used to perform fingertip recognition on the second feature map based on the detection head module of the fingertip detection model to complete the fingertip detection of the target event image.

9. A computer device, characterized in that: The computer device includes a memory and a processor; The memory is used to store computer programs; The processor is used to execute the computer program and implement the fingertip detection method based on an event camera as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, which, when executed by a processor, enables the processor to implement the fingertip detection method based on an event camera according to any one of claims 1 to 7.

Citation Information

Cited By

  • Target tracking method based on fusion of RGB data and event data of visual Mama

    CN120125617A