Palm center input method, device and equipment of intelligent wearable equipment and storage medium

By combining event cameras and RGB cameras, using event cameras for fingertip detection and projecting palm area recognition, the problem that RGB cameras are difficult to accurately capture user gestures in high dynamic scenarios is solved, and the interaction accuracy and fluency of smart wearable devices are improved.

CN120010655APending Publication Date: 2025-05-16ZHUHAI MOJIE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411933577.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-25
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

RGB cameras are difficult to accurately capture user gestures or clicking actions in high dynamic scenarios, resulting in low interactive accuracy and fluency of smart wearable devices.

Method used

Combining the event camera and RGB camera for palm input, acquiring event images through the event camera for fingertip detection, and combining the images acquired by the RGB camera for projection palm area recognition and projection of the virtual input interface.

Benefits of technology

It effectively reduces the visual smear or afterimage problems caused by camera lag, improves the accuracy of operating fingertip position capture in high dynamic scenes, and enhances the interaction accuracy and fluency of smart wearable devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120010655A_ABST
    Figure CN120010655A_ABST
Patent Text Reader

Abstract

The invention provides a palm center input method and device of intelligent wearable equipment, the intelligent wearable equipment and a computer readable storage medium. The palm center input method of the intelligent wearable device comprises the steps that a scene image is collected through an RGB camera of the intelligent wearable device for projection palm area recognition; after virtual input interface projection is carried out in the projection palm area, a first event image of a virtual input interface is acquired through an event camera of the intelligent wearable device; performing fingertip detection based on an event image set including the first event image through a trained fingertip detection model to obtain an operation fingertip position of the virtual input interface; and performing input processing based on the operation fingertip position and the virtual input interface. According to the method and the device, the problem of visual smear or ghosting caused by camera lag can be reduced, so that the operation fingertip position of a user in a high-dynamic scene can be accurately captured, and the interaction accuracy and fluency of the intelligent wearable equipment are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer vision technology, and in particular to a palm input method and device for a smart wearable device, a smart wearable device, and a computer-readable storage medium. Background Art

[0002] With the development of smart wearable technology, augmented reality devices and virtual reality devices have entered the lives of users. At present, there are many main interaction methods for smart wearable devices such as augmented reality devices and virtual reality devices. Among them, palm input is a new human-computer interaction method. The operation interface is projected onto the palm of the user through smart wearable devices (such as VR glasses). The user can directly touch or gesture the virtual interface on the palm to achieve input and control.

[0003] In the related art, RGB cameras are mainly used to capture the projected palm and operating fingertips of the hand image recognition operation interface. However, during the actual research and development process, the inventors of the embodiments of the present application found that the frame rate of RGB cameras is usually low. When the user moves his hands quickly to perform input operations, the image may have ghosting or afterimages. In high-dynamic scenes, the user's gestures or click actions are difficult to be accurately captured, resulting in low interaction accuracy and smoothness of smart wearable devices. Summary of the invention

[0004] The present application provides a palm input method and device for a smart wearable device, a smart wearable device, and a computer-readable storage medium, which can reduce the visual ghosting or afterimage problems caused by camera lag, so that the user's operating fingertip position in high-dynamic scenes can be accurately captured, thereby improving the interaction accuracy and fluency of the smart wearable device.

[0005] In a first aspect, the present application provides a palm input method for a smart wearable device, the method comprising:

[0006] The RGB camera of the smart wearable device is used to collect scene images and identify the projected palm area;

[0007] After the virtual input interface is projected on the projection palm area, a first event image of the virtual input interface is acquired through an event camera of the smart wearable device;

[0008] Performing fingertip detection based on an event image set including the first event image by using a trained fingertip detection model to obtain an operating fingertip position of the virtual input interface;

[0009] Input processing is performed based on the operating fingertip position and the virtual input interface.

[0010] In a second aspect, the present application provides a palm input device for a smart wearable device, the palm input device for the smart wearable device comprising:

[0011] A projection unit, used to collect scene images through the RGB camera of the smart wearable device to identify the projected palm area;

[0012] an acquisition unit, configured to acquire a first event image of the virtual input interface through an event camera of the smart wearable device after the virtual input interface is projected on the projection palm area;

[0013] a detection unit, configured to perform fingertip detection based on an event image set including the first event image by using a trained fingertip detection model, and obtain a fingertip position for operating the virtual input interface;

[0014] An input unit is used to perform input processing based on the operating fingertip position and the virtual input interface.

[0015] In a third aspect, the present application also provides a smart wearable device, which includes a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, it executes any palm input method of the smart wearable device provided in the present application.

[0016] In a fourth aspect, the present application also provides a computer-readable storage medium on which a computer program is stored, and the computer program is loaded by a processor to execute the palm input method of the smart wearable device.

[0017] The present application combines an event camera and an RGB camera for palm input. On the one hand, the position of the operating fingertip is identified by using the first event image captured by the RGB camera. Since the event camera can respond to brightness changes in microseconds, the visual ghosting or afterimage problems caused by camera lag are greatly reduced. In high-dynamic scenes, the position of the user's operating fingertip can be accurately captured, thereby improving the interaction accuracy and smoothness of the smart wearable device, so that the user can still obtain clear and continuous visual feedback when performing input operations quickly; on the second hand, since the projected palm area is in a relatively static state, the projection of the virtual input interface is projected by capturing the scene image through the RGB camera, which can make up for the deficiency of the event camera in being unable to detect static objects and ensure accurate and stable detection of the projected palm area. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0019] Figure 1 It is a schematic block diagram of the structure of a smart wearable device provided in an embodiment of the present application;

[0020] Figure 2 This is a flow chart of a palm input method for a smart wearable device provided in an embodiment of the present application;

[0021] Figure 3 It is a schematic diagram of the principle structure of a trained semantic segmentation model provided in an embodiment of the present application;

[0022] Figure 4 It is a schematic diagram of the principle structure of the trained fingertip detection model provided in the embodiments of the present application;

[0023] Figure 5 This is a schematic diagram illustrating an embodiment of a palm input process of a smart wearable device in an embodiment of the present application;

[0024] Figure 6 It is a schematic diagram of the structure of an embodiment of a palm input device for a smart wearable device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0026] The flowcharts shown in the accompanying drawings are only examples and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps may also be decomposed, combined or partially merged, so the actual execution order may change according to actual conditions.

[0027] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and cannot be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, the features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, the meaning of "multiple" is two or more, unless otherwise clearly and specifically defined.

[0028] In order to enable any person skilled in the art to implement and use the present application, the following description is provided. In the following description, details are listed for the purpose of explanation. It should be understood that those of ordinary skill in the art can recognize that the present application can also be implemented without using these specific details. In other examples, the known process will not be elaborated in detail to avoid unnecessary details that make the description of the present application embodiment obscure. Therefore, the present application is not intended to be limited to the embodiments shown, but is consistent with the widest range of principles and features disclosed in accordance with the embodiments of the present application.

[0029] The embodiments of the present application provide a palm input method and device for a smart wearable device, a smart wearable device, and a computer-readable storage medium. The palm input device of the smart wearable device can be integrated in the smart wearable device. The smart wearable device can be smart glasses, smart helmets, etc. The smart glasses can be AR (augmented reality) glasses, VR (Virtual Reality) glasses, MR (Mixed Reality) glasses, XR (eXtended Reality) glasses, etc. The smart helmet can be an AR helmet, etc.

[0030] The executor of the palm input method of the smart wearable device in the embodiment of the present application can be the palm input device of the smart wearable device provided in the embodiment of the present application, or the smart wearable device integrated with the palm input device of the smart wearable device, wherein the palm input device of the smart wearable device can be implemented in hardware or software.

[0031] Some embodiments of the present application are described in detail below in conjunction with the accompanying drawings. In the absence of conflict, the following embodiments and features in the embodiments can be combined with each other.

[0032] Figure 1 It is a schematic block diagram of the structure of a smart wearable device provided in an embodiment of the present application.

[0033] like Figure 1As shown, the smart wearable device 100 includes a processor 101 and a memory 102, and the processor 101 and the memory 102 are connected via a bus 103, such as an I2C (Inter-integrated Circuit) bus.

[0034] Specifically, the processor 101 is used to provide computing and control capabilities to support the operation of the entire smart wearable device 100. The processor 101 can be a central processing unit (CPU), and the processor 101 can also be other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, etc. Among them, the general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0035] Specifically, the memory 102 may be a Flash chip, a read-only memory (ROM) disk, an optical disk, a USB flash drive, or a mobile hard disk.

[0036] Those skilled in the art will understand that Figure 1 The structure shown in the figure is only a block diagram of a partial structure related to the embodiment scheme of the present application, and does not constitute a limitation on the smart wearable device to which the embodiment scheme of the present application is applied. The specific smart wearable device may include more or fewer components than those shown in the figure, or combine certain components, or have a different arrangement of components.

[0037] The processor 101 is used to run the computer program stored in the memory 102, and implements any one of the palm input methods for the smart wearable device provided in the embodiments of the present application when executing the computer program. For example, the processor 101 is used to run the computer program stored in the memory 102, and can implement the following steps when executing the computer program:

[0038] The scene image is collected by the RGB camera of the smart wearable device to identify the projected palm area; after the virtual input interface is projected in the projected palm area, the first event image of the virtual input interface is obtained by the event camera of the smart wearable device; through the trained fingertip detection model, fingertip detection is performed based on the event image set including the first event image to obtain the operating fingertip position of the virtual input interface; input processing is performed based on the operating fingertip position and the virtual input interface.

[0039] It should be noted that technicians in the relevant field can clearly understand that for the convenience and brevity of description, the specific working process of the smart wearable device described above can refer to the corresponding process in the palm input method embodiment of the smart wearable device below, and will not be repeated here.

[0040] In the following, Figure 1 Taking the smart wearable device shown in as an example as the execution subject of the palm input method of the smart wearable device, the palm input method of the smart wearable device provided in the embodiment of the present application is described in detail. For the sake of simplicity and convenience of description, the execution subject will be omitted in the subsequent method embodiments.

[0041] See also Figure 2 , Figure 2 201 is a flowchart of a palm input method for a smart wearable device provided in an embodiment of the present application. The palm input method for a smart wearable device includes steps 201 to 204, wherein:

[0042] 201. The scene image is collected by the RGB camera of the smart wearable device to identify the projected palm area.

[0043] The projected palm area refers to the palm area used to project the virtual input interface, such as the left palm area, the right palm area, etc.

[0044] In some embodiments, in order to accurately detect the projected palm area to improve the projection accuracy of the virtual input interface, the projected palm area may be detected by referring to the following steps A1 to A5:

[0045] A1. Capture scene images through the RGB camera.

[0046] The scene image is an image captured for the scene in which the smart wearable device is located. For example, it may include an image of the palm of the wearer of the smart wearable device.

[0047] In order to identify the palm area and project the virtual input interface onto the palm area, the scene image is collected by the RGB camera of the smart wearable device.

[0048] A2. Perform a multi-scale convolution operation based on the scene image by training the encoder of the semantic segmentation model to obtain a multi-scale feature map of the scene image.

[0049] In order to better understand this embodiment, the trained semantic segmentation model in this embodiment is first introduced below. Figure 3 As shown, Figure 3 It is a schematic diagram of the principle structure of a trained semantic segmentation model provided in an embodiment of the present application. The trained semantic segmentation model may include an encoder and a decoder. The trained semantic segmentation model can be trained based on a preset semantic segmentation model. Semantic segmentation is a computer vision task whose goal is to divide each pixel in an image into a specific category. In the palm input method, semantic segmentation is used to accurately detect the hand and segment the palm area therefrom, so that the palm surface can be identified as an interactive interface, thereby supporting the user to perform touch or gesture operations on the palm. The working principles of the encoder and decoder are as follows:

[0050] 1. Encoder: It is used to perform feature encoding processing such as multi-scale convolution, splicing, channel attention enhancement, etc. on the scene image to obtain the channel-enhanced feature map of the scene image.

[0051] There are many ways to implement the structure of the encoder, exemplarily including the following: <1> and <2> :

[0052] <1> In some implementations, such as Figure 3 As shown, the encoder may include a first convolution layer, a multi-scale convolution layer, a concatenation layer, a channel attention layer, and a second convolution layer. Exemplarily, at this time, the working process of the encoder may be as follows:

[0053] First, in the first convolution layer: the input is the scene image, and the output is the preliminary scene feature map. The scene image passes through the first convolution layer to extract preliminary features and appropriately reduce the number of channels, thereby providing effective support for subsequent multi-scale convolution operations. For example, Figure 3 As shown, the first convolutional layer can be a 1*1 convolutional layer structure.

[0054] Then, in the multi-scale convolution layer: the input is the preliminary scene feature map, and multiple feature maps of different scales are output in parallel. The multi-scale convolution layer specifically includes multiple parallel convolution layers of different scales, for example, Figure 3 As shown, the multi-scale convolution layer can include four parallel convolution layers of different scales (such as: 1*1, 3*3, 5*5 and 7*7 convolution layers). The multi-scale convolution layer structure set by the encoder can capture features of different scales, so that the semantic segmentation model can adapt to objects and details of different sizes in the image, thereby improving the detection and segmentation accuracy of the projected palm area by the semantic segmentation model.

[0055] Next, in the concatenation layer, the input is a feature map of multiple scales, and the output is a fused multi-scale feature map. The concatenation layer set by the encoder can concatenate the output features of the multi-scale convolution layer to form a feature map that fuses information of multiple scales (i.e., a multi-scale feature map), thereby providing rich feature expressions for subsequent processing.

[0056] Next, in the channel attention layer, the input is a multi-scale feature map and the output is a channel-enhanced feature map. Figure 3 As shown in the figure, the multi-scale feature map is further processed by the channel attention layer of the encoder to obtain the channel-enhanced feature map, which can enhance key features while suppressing irrelevant or redundant information, so that the semantic segmentation model can focus on the target area more accurately in a complex background. The channel attention layer is used to perform selective enhancement in the multi-scale feature map to further improve the feature expression effect.

[0057] Finally, the channel enhanced feature map and the preliminary scene feature map are input into the decoder part for subsequent processing. In some embodiments, the channel enhanced feature map can be directly input into the decoder part. In this case, the encoder can include a first convolution layer, a multi-scale convolution layer, a splicing layer, and a channel attention layer. In some embodiments, a second convolution layer (for example, a 1*1 convolution layer) can be set after the channel attention layer in the encoder to integrate the channel enhanced feature map, and then the integrated channel enhanced feature map is input into the decoder part to improve the input feature quality of the decoder.

[0058] <2> In some implementations, the encoder may include a first convolution layer, a multi-scale convolution layer, a concatenation layer, and a channel attention layer. At this time, the channel enhanced feature map output by the channel attention layer and the preliminary scene feature map output by the first convolution layer are directly input to the decoder part.

[0059] 2. Decoder: It is used to perform upsampling, convolution and other operations based on the channel-enhanced feature map of the scene image, gradually restore the spatial resolution of the image, and output the category probability corresponding to each pixel to determine the target pixel probability map of the image. Among them, the pixel value of each pixel in the pixel probability map is the category probability corresponding to each pixel, and the category probability corresponding to each pixel i is used to indicate the probability that the pixel i is the pixel in the palm area.

[0060] There are many ways to implement the structure of the decoder, illustratively, including the following cases ① and ②:

[0061] ① In some embodiments, Figure 3As shown, the decoder may include a third convolutional layer, a first upsampling layer, a fusion layer, a self-attention layer, a prediction layer, and a second upsampling layer. Exemplarily, at this time, the working process of the decoder may be as follows:

[0062] First, in the third convolutional layer: the input is the preliminary scene feature map and the output is the integrated preliminary scene feature map.

[0063] In the first upsampling layer (upsample), the input is the channel-enhanced feature map, and the output is the sampled channel-enhanced feature map. In this way, the channel-enhanced feature map is upsampled by the first upsampling layer of the decoder to restore the spatial resolution of the feature map.

[0064] Then, in the fusion layer (concat): the input is the sampled channel enhanced feature map and the integrated preliminary scene feature map, and the output is the spliced ​​feature map.

[0065] Next, in the self-attention layer, the input is the concatenated feature map, and the output is the self-attention feature map. The self-attention feature map is obtained by the self-attention layer of the decoder. In this way, the self-attention layer can capture the global dependency between pixels in the feature map after channel enhancement, further improving the semantic segmentation model's ability to focus on key areas.

[0066] Next, in the prediction layer: the input is the self-attention feature map, and the output is the preliminary pixel probability map. The prediction layer of the decoder can be a convolutional structure, for example Figure 3 As shown, the prediction layer can be a 3*3 convolution layer, through which the corresponding category probability (such as the probability that each pixel is the pixel in the palm area) of each pixel in the self-attention feature map can be predicted, thereby obtaining a preliminary pixel probability map.

[0067] Finally, in the second upsampling layer (upsample), the input is the preliminary pixel probability map, and the output is the target pixel probability map. The second upsampling layer of the decoder performs upsampling again, gradually restores to the resolution of the input image (i.e., the scene image), and finally outputs the target pixel probability map. The scene image is divided and processed based on the target pixel probability map, thereby obtaining a segmentation result image of the projected palm area and other backgrounds, so that the projected palm area can be segmented from the scene image. Among them, the preliminary pixel probability map and the target pixel map are both feature maps used to indicate the probability that each pixel point is the pixel point where the palm area is located. The difference is that the size of the preliminary pixel probability map and the target pixel probability map is different. In order to improve the accuracy of the category probability prediction corresponding to each pixel point, the feature map obtained after deep encoding by the encoder is used to predict the category probability to obtain a preliminary pixel probability map. However, this preliminary pixel probability map will be compressed into a size smaller than the original input scene image due to the feature encoding process. Therefore, it is necessary to restore the preliminary pixel probability map to obtain a target pixel probability map with the same size as the original input scene image.

[0068] ② In some embodiments, the decoder may include a third convolutional layer, a first upsampling layer, a fusion layer, a prediction layer, and a second upsampling layer. Exemplarily, at this time, the working process of the decoder may be as follows:

[0069] First, in the third convolution layer: the input is the preliminary scene feature map, and the output is the integrated preliminary scene feature map. In the first upsampling layer: the input is the channel enhanced feature map, and the output is the sampled channel enhanced feature map. Then, in the fusion layer: the input is the sampled channel enhanced feature map and the integrated preliminary scene feature map, and the output is the spliced ​​feature map. Next, in the prediction layer: the input is the spliced ​​feature map, and the output is the preliminary pixel probability map. Finally, in the second upsampling layer: the input is the preliminary pixel probability map, and the output is the target pixel probability map. Some implementation details are similar to situation ①. Please refer to the relevant instructions in the previous article for details, which will not be repeated here.

[0070] Exemplarily, step A2 may specifically include: first, performing preliminary feature extraction on the scene image through the first convolution layer of the encoder to obtain a preliminary scene feature map; then, performing multi-scale convolution operations on the preliminary scene feature map in parallel through the multi-scale convolution layer of the encoder to obtain feature maps of multiple scales; then, through the splicing layer of the encoder, splicing the multiple feature maps of different scales to obtain a multi-scale feature map of the scene image.

[0071] A3. Perform channel enhancement processing on the multi-scale feature map through the encoder to obtain a channel enhanced feature map.

[0072] Exemplarily, in step A3, the channel enhancement processing is performed on the multi-scale feature map through the channel attention layer of the encoder to obtain a channel-enhanced feature map. In some embodiments, the channel-enhanced feature map can be directly input into the decoder for subsequent decoding process; in some embodiments, the channel-enhanced feature map can also be integrated through the second convolutional layer and then input into the decoder for subsequent decoding process.

[0073] A4. Predicting based on the channel-enhanced feature map using the decoder of the trained semantic segmentation model to obtain a target pixel probability map of the scene image.

[0074] The target pixel probability map is used to indicate the probability that each pixel point in the scene image is a pixel point where the palm area is located.

[0075] In some embodiments, the decoder includes a third convolutional layer, a first upsampling layer, a fusion layer, a prediction layer, and a second upsampling layer. Step A4 may specifically include: first, through the first convolutional layer of the encoder, a preliminary feature extraction is performed on the scene image to obtain a preliminary scene feature map; through the first upsampling layer of the decoder, the channel enhanced feature map is upsampled to restore the spatial resolution of the feature map to obtain the sampled channel enhanced feature map. Then, through the fusion layer of the decoder, the sampled channel enhanced feature map and the integrated preliminary scene feature map are spliced ​​and fused to obtain the spliced ​​feature map. Then, through the prediction layer of the decoder, the corresponding category probability (such as the probability that each pixel is a pixel in the palm area) is predicted for each pixel of the spliced ​​feature map, thereby obtaining a preliminary pixel probability map. Finally, through the second upsampling layer of the decoder, the preliminary pixel probability map is upsampled again, gradually restored to the resolution of the input image (i.e., the scene image), and finally the target pixel probability map is output.

[0076] In some embodiments, the decoder includes a third convolutional layer, a first upsampling layer, a fusion layer, a self-attention layer, a prediction layer, and a second upsampling layer. Step A4 may specifically include: performing preliminary feature extraction on the scene image through the encoder to obtain a preliminary scene feature map; performing splicing processing based on the channel enhanced feature map and the preliminary scene feature map through the decoder to obtain a spliced ​​feature map; performing self-attention feature extraction on the spliced ​​feature map to obtain a self-attention feature map; performing prediction processing based on the self-attention feature map to obtain a preliminary pixel probability map of the scene image; performing an upsampling operation on the preliminary pixel probability map to obtain the target pixel probability map. Specifically, first, performing preliminary feature extraction on the scene image through the first convolutional layer of the encoder to obtain a preliminary scene feature map; performing an upsampling operation on the channel enhanced feature map through the first upsampling layer of the decoder to restore the spatial resolution of the feature map to obtain a sampled channel enhanced feature map. Then, through the fusion layer of the decoder, the sampled channel enhanced feature map and the integrated preliminary scene feature map are spliced ​​and fused to obtain a spliced ​​feature map. Next, the decoder's self-attention layer is used to capture the global dependencies of the concatenated feature map to obtain a self-attention feature map. The decoder's prediction layer predicts the corresponding category probability of each pixel in the self-attention feature map (such as the probability that each pixel is a pixel in the palm area), thereby obtaining a preliminary pixel probability map. Finally, the second upsampling layer of the decoder is used to upsample the preliminary pixel probability map again, gradually restoring it to the resolution of the input image (i.e., the scene image), and finally outputting the target pixel probability map.

[0077] A5. Divide the scene image based on the target pixel probability map to obtain the projected palm area.

[0078] Among them, the target pixel probability map is a probability map of the same size as the scene image, and the pixel value of each pixel in the target pixel probability map represents the probability that the pixel is the pixel where the palm area is located. Exemplarily, based on the pixel value of each pixel in the target pixel probability map, it is determined whether the pixel at the corresponding position in the scene image is the pixel where the palm area is located, thereby determining all the pixels of the palm area in the scene image; and the scene image is divided and processed according to the pixels of all the palm areas in the scene image to obtain the projected palm area.

[0079] 202. After projecting a virtual input interface on the projection palm area, obtain a first event image of the virtual input interface through an event camera of the smart wearable device.

[0080] The first event image is an event image captured by an event camera.

[0081] An event camera is a sensor that can capture fast dynamic scenes and generates data by detecting pixel-level brightness changes. When the brightness change of a pixel in the scene reaches a preset threshold, the pixel immediately triggers an event, recording the pixel's position, timestamp, and polarity of the brightness change (brightening or darkening). After the virtual input interface is projected on the palm area, the first event image of the virtual input interface is obtained through the event camera of the smart wearable device, which facilitates the subsequent identification of the fingertip position.

[0082] 203. Perform fingertip detection based on the event image set including the first event image through the trained fingertip detection model to obtain the operating fingertip position of the virtual input interface.

[0083] Among them, the operating fingertip position refers to the position of the input operating fingertip (the input operating fingertip refers to the fingertip of the finger used for input operation), which can be set according to the actual business scenario. For example, the projected palm area can be set to the left palm area and the input operating fingertip can be set to the fingertip of the right hand.

[0084] In order to better understand the present embodiment, the trained fingertip detection model in the present embodiment is first introduced below. Figure 4 As shown, Figure 4 : is a schematic diagram of the principle structure of the trained fingertip detection model provided in the embodiment of the present application. The trained fingertip detection model may include a convolution module, a spatiotemporal feature extraction module, a fusion module and a prediction module. Further, the trained fingertip detection model may also include a residual module. The trained fingertip detection model may be obtained by training based on a preset fingertip detection model.

[0085] Taking the trained fingertip detection model as an example, it can include a convolution module, a spatiotemporal feature extraction module, a fusion module, a residual module, and a prediction module. The working principles of each module are as follows:

[0086] 1. A convolution module is used to perform preliminary feature extraction on each event image in the event image set to obtain a preliminary event feature map. For example, Figure 4As shown, the convolution module may include two convolution submodules, each of which may be composed of a convolution layer, a batch normalization layer (Batch Normalization) and an activation layer (such as ReLU). After the event image is subjected to convolution, batch normalization, activation and other operations in the two convolution submodules in turn, a preliminary event feature map is output, and the preliminary event feature map is input into the spatiotemporal feature extraction module. The first convolution submodule first performs convolution, batch normalization, activation and other operations on the event image set, and then outputs the first convolution feature map; and the second convolution submodule performs convolution, batch normalization, activation and other operations on the first convolution feature map again, and then outputs the second convolution feature map as the preliminary event feature map.

[0087] 2. The spatiotemporal feature extraction module is used to extract spatiotemporal features from the preliminary event feature graph to obtain the spatiotemporal feature graph. The spatiotemporal feature extraction module can combine the information of the event image in the spatial and temporal dimensions to capture and focus on the dynamic features in multiple frames of event images in the event image set, effectively improving the ability to focus on the position of the operating fingertip. In this way, the spatiotemporal feature extraction module can analyze the movement trajectory and change trend of the finger, ensure the consistency of the position of the fingertip in multiple frames, and use the spatiotemporal feature extraction model to extract the fingertip. The fingertip detection model can still maintain high-precision detection when processing subtle movements of the fingertip, and can still perform stable recognition even when the fingertip pauses or moves slowly. The Mamba network is an efficient sequence modeling architecture based on the state space model (SSM). It achieves efficient sequence modeling by independently mapping the input vector to a high-dimensional potential state and outputting the result, thereby improving computational efficiency and network performance. For example, the spatiotemporal feature extraction module can use the Mamba network structure. Since the Mamba network can focus on key frame information, even if the user's finger is temporarily stationary, the fingertip detection model using the Mamba network structure as the spatiotemporal feature extraction module can still perform effective recognition based on the previous and next frame information. This overcomes the limitations of event cameras in static scenes, ensures the continuity and accuracy of gesture recognition, realizes seamless recognition of dynamic and static gestures, and effectively improves the user experience of palm input.

[0088] 3. Fusion module, used to splice and fuse the preliminary event feature map and spatiotemporal features output by the convolution module to obtain a fused feature map. In this way, through the connection operation of the fusion module, a richer feature expression is formed, so that the fingertip detection model can contain multi-level information in a feature map, and realize the integration of features of different scales and levels. The design of the fusion module enables the feature map to maintain the detail information while enhancing the global semantic information of the feature map, providing deeper support for subsequent processing, so that the fingertip detection model can more accurately locate the fingertip area in a complex background.

[0089] 4. A residual module is used to perform residual processing on the fused feature map to obtain a processed feature map. For example, the residual module can adopt the residual module (ResBlock) of the Resnet network. The residual module can adopt a short connection structure to directly add the input to the output, thereby alleviating the gradient vanishing problem that may occur in the deep network. In this way, the residual module can be used to maintain the feature map. Figure 1 At the same time, the information highly related to the fingertip position is further proposed, making the final feature map more refined and accurate.

[0090] 5. Prediction module, used to predict the fingertip position according to the processed feature map to obtain the operating fingertip position. Through a series of processing steps including convolution module, spatiotemporal feature extraction module, fusion module, and residual module, the fingertip detection model fully explores and integrates the dynamic features and static details of the hand, laying a solid foundation for fingertip key point detection; illustratively, the prediction module can use a linear module to output the accurate coordinates of the operating fingertip as the operating fingertip position, thereby achieving high-precision detection of the fingertip, providing reliable technical support for application scenarios with high accuracy requirements such as palm input.

[0091] There are many ways to implement step 203, which include, for example:

[0092] (1) In some embodiments, the trained fingertip detection model includes a convolution module, a spatiotemporal feature extraction module, a fusion module and a prediction module. In this case, step 203 may specifically include: performing preliminary feature extraction on each event image in the event image set through the convolution module of the trained fingertip detection model to obtain a preliminary event feature map, wherein the event image set includes the second event image and the first event image; performing feature extraction on the preliminary event feature map through the spatiotemporal feature extraction module of the trained fingertip detection model to obtain a spatiotemporal feature map; performing fusion processing on the spatiotemporal feature map and the preliminary event feature map through the fusion module of the trained fingertip detection model to obtain a fused feature map; performing fingertip position prediction based on the fused feature map through the prediction module of the trained fingertip detection model to obtain the operating fingertip position.

[0093] (2) In some embodiments, the trained fingertip detection model includes a convolution module, a spatiotemporal feature extraction module, a fusion module, a residual module and a prediction module. In this case, step 203 may specifically include: performing preliminary feature extraction on each event image in the event image set through the convolution module of the trained fingertip detection model to obtain a preliminary event feature map, wherein the event image set includes the second event image and the first event image; performing feature extraction on the preliminary event feature map through the spatiotemporal feature extraction module of the trained fingertip detection model to obtain a spatiotemporal feature map; performing fusion processing on the spatiotemporal feature map and the preliminary event feature map through the fusion module of the trained fingertip detection model to obtain a fused feature map; performing residual processing on the fused feature map through the residual module of the trained fingertip detection model to obtain a processed feature map; performing fingertip position prediction based on the processed feature map through the prediction module of the trained fingertip detection model to obtain the operating fingertip position.

[0094] In some embodiments of step 203, the event image set is a set of first event images, and the first event images are used to detect the operating fingertip position. In some embodiments of step 203, the event image set is a set of first event images and second event images, and the first event images and the second event images are used to detect the operating fingertip position. The second event image is obtained by converting the hand image sequence collected by the RGB camera. Exemplarily, the second event image can be obtained by the following steps B1 to B2:

[0095] B1. After the virtual input interface is projected on the projection palm area, a hand image sequence is collected by the RGB camera of the smart wearable device;

[0096] Exemplarily, after the virtual input interface is projected on the projection palm area, a continuous multi-frame depth image can be collected by the RGB camera of the smart wearable device, and the continuous multi-frame depth image forms a hand image sequence.

[0097] B2. Performing a superposition process based on the hand image sequence to obtain a second event image of the virtual input interface.

[0098] The second event image is obtained by converting the hand image sequence captured by the RGB camera through superposition processing.

[0099] Thus, on the first hand, by using the first event image captured by the event camera to detect the position of the operating fingertip, the event camera is a sensor that can capture fast dynamic scenes, and generates data by detecting pixel-level brightness changes. When the brightness change of a pixel in the scene reaches a preset threshold, the pixel immediately triggers an event to record the position, timestamp and polarity of the brightness change (brightening or darkening) of the pixel. Therefore, the first event image can be used to effectively capture the dynamic characteristics of the gesture, avoiding the problem that the RGB camera is difficult to capture the user's fast hand movements, and effectively avoiding the problem that the user's gestures or click actions are difficult to be accurately captured in high-dynamic scenes, thereby avoiding When the user moves his hand quickly to perform input operations, the image may have a trailing image or afterimage, which improves the smoothness and accuracy of input interaction operations and improves the input experience; secondly, when detecting the position of the operating fingertip, the second event image can be used to capture static features, avoid the problem of not being able to obtain valid data when the finger is stationary, and overcome the limitations of the event camera in static scenes; thirdly, the first event image and the second event image are combined to detect the position of the operating fingertip, thereby overcoming the limitations of the event camera in static scenes, and ensuring the continuity and accuracy of gesture recognition, achieving seamless recognition of dynamic and static gestures, and effectively improving the user experience of palm input. In addition, multimodal data refers to data obtained from different perception methods. Each data mode (modality) provides a different perspective on the same thing or scene, and is complementary to each other. By combining the first event image captured by the event camera and the second event image formed by the image captured by the RGB camera, the multimodal data can be combined for fingertip detection to improve the recognition accuracy of the operating fingertip position.

[0100] Exemplarily, step B2 may specifically include the following steps B21 to B24:

[0101] B21. Perform grayscale conversion on each hand image in the hand image sequence to obtain a grayscale image of each hand image.

[0102] Wherein, the hand image sequence includes multiple hand images, and each hand image is an RGB image. In some embodiments, the weighted average of the first channel value, the second channel value, and the third channel value of each pixel point of each hand image is calculated to obtain a channel weighted result image of each hand image as a grayscale image of each image, wherein the first channel value, the second channel value, and the third channel value are the R channel value, the G channel value, and the B channel value, respectively. In this way, the color information of the RGB three channels of each hand image can be compressed into the brightness information of a single channel, which is convenient for subsequent processing. Exemplarily, the weighted formula is shown in the following formula 1:

[0103] Grayscale=0.299×R+0.587×G+0.114×B Formula 1

[0104] In Formula 1, Grayscale represents the grayscale image of the hand image, R represents the R channel value, G represents the G channel value, and B represents the B channel value.

[0105] B22. Based on the grayscale image of each hand image, obtain grayscale difference images of adjacent frames of the hand image sequence.

[0106] Among them, the adjacent frame grayscale difference map refers to the grayscale difference map of two adjacent hand images in the hand image sequence. For example, assuming that the hand image sequence includes M frames of hand images, the grayscale value map of the i-th frame hand image and the i-1-th frame hand image can be calculated as the adjacent frame grayscale difference map corresponding to the i-th frame hand image, where 1<i≤M, so that M-1 adjacent frame grayscale difference maps can be obtained. In this way, the pixel difference between the current frame grayscale image and the previous frame grayscale image can be calculated as the adjacent frame grayscale difference map to reflect the movement changes of objects in the image. The pixel value of any pixel position in the adjacent frame grayscale illustration is the grayscale difference between the corresponding pixel position of the current frame grayscale image i and the previous frame grayscale image i-1. The grayscale difference of the corresponding pixel position is positive or negative, indicating that the brightness of the pixel position increases or decreases. Exemplarily, the calculation formula of the adjacent frame grayscale difference map is shown in Formula 2 below:

[0107] diff(x,y)=Grayscale (i,x,y) -Grayscale (i-1,x,y) Formula 2

[0108] In formula 2, Grayscale (i,x,y) Grayscale represents the grayscale value of the grayscale image of the i-th hand image at the pixel position (x, y). (i-1,x,y) represents the grayscale value of the grayscale image of the i-1th hand image at the pixel position (x, y), and diff(x, y) represents the pixel value of the grayscale difference image of the adjacent frame corresponding to the i-th frame hand image at the pixel position (x, y).

[0109] B23. Generate multiple variation images of the hand image sequence based on the grayscale difference images of each adjacent frame and a preset pixel threshold.

[0110] Exemplarily, a preset pixel threshold may be preset, and each adjacent frame grayscale difference map is processed, and each adjacent frame grayscale difference map generates a corresponding change image, as shown in the following formula 3:

[0111]

[0112] In formula 3, threshold represents the preset pixel threshold, diff(x,y) represents the pixel value of the grayscale difference image of the adjacent frame corresponding to the i-th frame hand image at the pixel position (x,y), and event(x,y) represents the pixel value of the change image generated by the grayscale difference image of the adjacent frame corresponding to the i-th frame hand image at the pixel position (x,y). Among them, [255,0,0] and [0,255,0] represent red and green respectively.

[0113] It can be seen that the generated change image contains green and black-red pixel locations, which represent the areas where the brightness increases and decreases, respectively.

[0114] B24. Perform superposition processing based on the multiple change images to obtain the second event image.

[0115] In some embodiments, the maximum value of the pixel values ​​of the multiple changing images at the pixel point position (x, y) is used as the pixel value of the superimposed image at the pixel point position (x, y), and the multiple changing images are superimposed to obtain the superimposed image as the second event image. In this way, when multiple changing images are superimposed, the pixel values ​​of the same pixel point position are read from multiple consecutive frames of changing images, and the pixel values ​​of these pixel point positions are complemented to enhance the motion characteristics in the image. For example, assuming that there are N consecutive frames of changing images, the pixel value of the pixel point of each frame of the changing image is expressed as I t (x, y), where t = 1, 2, ..., N represents the frame number, (x, y) represents the pixel position; for each pixel position (x, y), the maximum value of the pixel position (x, y) in the N frames of continuous change images is used as the pixel value of the pixel position (x, y) of the superimposed image, so that the pixel position (x, y) in the second event image obtained after superposition can retain the maximum brightness change in the motion, highlighting the characteristics of significant changes in the event, so as to facilitate the subsequent effective capture of fingertip motion. Exemplarily, the second event image obtained after superposition can be expressed as the following formula 4:

[0116]

[0117] In formula 4, I add (x, y) represents the second event image obtained after superposition, I t (x,y) represents the change image of the tth frame.

[0118] Thus, through steps B21 to B24, the hand image sequence captured by the RGB camera can be converted into an event image containing dynamic information, and the RGB frame captured by the depth image can be converted into an event image, such as Figure 5 This ensures that the operating fingertip position can be accurately extracted even when the finger moves at a slow speed or is stationary for a short time.

[0119] 204. Perform input processing based on the operating fingertip position and the virtual input interface.

[0120] Among them, after the virtual input interface is projected, the RGB camera continuously collects scene images at a certain frame rate. The trained semantic segmentation model can be used to update the projected palm area of ​​the scene image collected in real time to update the display position of the virtual input interface, so as to facilitate the detection of the relative position between the button position of the virtual input interface and the operating fingertip position, and determine whether the button click operation is triggered.

[0121] In some embodiments, the relative distance between each key position of the virtual input interface and the operating fingertip position can be obtained; based on the relative distance, it is determined whether to trigger a click operation on the key position to use the keys of the virtual input interface for input.

[0122] For a better understanding of the embodiments of this application, please refer to Figure 5 , the following takes "the projection palm area is the left and right palms, the operating fingertips are the fingertips of the right hand, and the virtual input interface is the virtual keyboard" as an example to illustrate the palm input process of the smart wearable device, as follows:

[0123] 1. First, the scene image is captured by the RGB camera of the smart wearable device, and the projected palm area is identified based on the scene image by training the semantic segmentation model to accurately obtain the palm boundary of the left hand, and a square area is constructed based on the boundary for projecting the virtual keyboard. In this way, the square area provides a stable input interface on the left palm, so that the position and size of the virtual keyboard can match the actual palm area of ​​the user, thereby improving the naturalness and accuracy of the interaction.

[0124] 2. Next, on the one hand, the hand image sequence is collected in real time by the RGB camera for overlay processing to obtain the second event image; on the other hand, the first event image is collected in real time by the event camera.

[0125] 3. Then, the trained fingertip detection model detects the fingertip position of the right hand (i.e., the operating fingertip position) in real time based on the event image set obtained by superimposing the first event image and the second event image; and uses the distance calculation method to measure the relative distance between the current fingertip position of the right hand and each key of the virtual keyboard. The distance is used to determine whether to trigger the click operation to ensure the accuracy of the click action, thereby realizing the palm input function of smart wearable devices (such as VR glasses). The whole process combines semantic segmentation and fingertip detection, allowing users to perform intuitive touch operations on the virtual interface, enhancing the immersion and convenience of the interactive experience.

[0126] In the embodiments of the present application, on the one hand, the following limitations are considered in the use of RGB cameras to capture hand images and recognize fingertip changes: 1> Low frame rate leads to ghosting or afterimages (the frame rate of RGB cameras is relatively low, and when the user quickly moves the fingertips to perform input operations, the image may have ghosting or afterimages, especially in high-dynamic scenes. The user's gestures or click actions are difficult to be accurately captured, resulting in low interaction accuracy and smoothness of the smart wearable device), 2> High power consumption makes it difficult to meet the needs of long-term palm input (the power consumption of RGB cameras is relatively high, which is particularly evident in smart wearable devices, thereby limiting the battery life of smart wearable devices, making users When using palm input for a long time, you will face the problem of frequent charging, which is not conducive to the continuous interactive experience in actual applications), 3> Sensitivity to light leads to low recognition accuracy. (Collecting images through RGB cameras for hand recognition such as finger or palm tips is sensitive to light. In low-light scenes, the RGB camera's ability to capture hand details is greatly reduced, resulting in reduced positioning accuracy of fingertips and palm areas, affecting the stability and accuracy of palm input, and limiting the applicability of palm input in complex environments). By combining event cameras and RGB cameras for palm input, 1> Because the event camera can respond to brightness changes in microseconds, it greatly reduces the camera lag caused by 1> The visual smear or afterimage problem can be solved by using the first event image captured by the RGB camera to identify the position of the operating fingertip, so that the user can still get clear and continuous visual feedback when performing input operations quickly. 2> Since the data acquisition efficiency of the event camera is higher and the power consumption is significantly lower than that of the RGB camera, by using the first event image captured by the RGB camera to identify the position of the operating fingertip, the time spent on collecting images with the RGB camera can be reduced, thereby reducing the time spent on using the RGB camera, which can effectively extend the battery life of the device and improve the overall user experience of palm input. 3> Since the event camera detects pixel-level changes to generate data, the operation fingertip The recognition result of the fingertip position is relatively less affected by lighting, and the operating fingertip position can still be accurately identified even in scenes with weak lighting. 4> Since the projected palm area is in a relatively static state, the RGB camera is used to collect scene images to identify the projected palm area for projection of the virtual input interface, which can make up for the deficiency that the event camera cannot detect static objects and ensure accurate and stable detection of the projected palm area. Therefore, it can effectively solve the problems of smear or afterimage caused by low frame rate when using RGB camera alone for palm input, the problem of difficulty in meeting long-term palm input requirements due to high power consumption, and the problem of low recognition accuracy due to sensitivity to light.

[0127] Secondly, by combining the fingertip detection model of the spatiotemporal feature extraction module and other modules, an event image set containing multiple event images is used as the input of the fingertip detection model to perform fingertip position detection. Firstly, since the second event image in the event image set can capture static features, it avoids the problem of not being able to obtain valid data when the finger is still, overcomes the limitations of the event camera in static scenes, and effectively solves the problem that the event camera can only capture moving objects and cannot obtain valid data when the finger is still; secondly, since the event image set includes multiple continuous frames, the frame superposition processing method of inputting the event image set enables the fingertip detection model to focus on key frame information, even if the user's finger Even when there is a brief static state, the network can still perform effective recognition based on the previous and next frame information; thirdly, the position of the operating fingertip is detected by combining the first event image and the second event image, thereby overcoming the limitations of event cameras in static scenes, ensuring the continuity and accuracy of gesture recognition, and achieving seamless recognition of dynamic and static gestures, effectively improving the user experience of palm input; fourthly, by combining the fingertip detection model with modules such as the spatiotemporal feature extraction module, the problem that traditional CNN cannot focus on key information and is not suitable for long sequence stream data is effectively solved, while overcoming the difficulty of traditional CNN models with high computational complexity, high power consumption, slow inference speed, and difficulty in deployment on smart wearable devices such as AR glasses.

[0128] Thirdly, the lightweight semantic segmentation model provided by the embodiment of the present application combines the multi-scale convolution layer, the channel attention layer and the self-attention layer to achieve efficient feature extraction and accurate projection palm area segmentation effect. The semantic segmentation model takes into account the focusing ability of multi-scale information and key areas, overcomes the shortcomings of traditional models in feature expression and resource requirements, and provides technical support for real-time and efficient palm input on smart wearable devices. First, through the multi-scale convolution layer structure, the semantic segmentation model can capture image features of different scales and adapt to target areas of various sizes and shapes, thereby enhancing the adaptability of projected palm area segmentation; second, through the channel attention layer, important features are further screened and enhanced to ensure that the semantic segmentation model can accurately focus on the projected palm area under complex backgrounds; third, through the self-attention layer, global dependencies are captured in the feature map, so that the semantic segmentation model has good recognition ability for both dynamic and static projected palm areas. The lightweight design of the semantic segmentation model greatly reduces the computational complexity and power consumption, and is particularly suitable for deployment on smart wearable devices such as AR glasses. It overcomes the high demand for computing resources of traditional models, realizes real-time and high-precision self-attention layer detection, and effectively improves the user interaction experience.

[0129] From the above content, it can be seen that by combining the event camera and the RGB camera for palm input, on the one hand, by using the first event image captured by the RGB camera to identify the position of the operating fingertip, since the event camera can respond to brightness changes at the microsecond level, it greatly reduces the visual smear or afterimage problem caused by camera lag, and the user's operating fingertip position can be accurately captured in high dynamic scenes, thereby improving the interaction accuracy and fluency of the smart wearable device, so that the user can still obtain clear and continuous visual feedback when performing input operations quickly; on the other hand, since the projected palm area is in a relatively static state, the RGB camera is used to capture the scene image to identify the projected palm area for projection of the virtual input interface, which can make up for the deficiency that the event camera cannot detect static objects and ensure accurate and stable detection of the projected palm area; on the third hand, since the event camera has higher data acquisition efficiency and significantly lower power consumption than the RGB camera, by using the first event image captured by the RGB camera to identify the position of the operating fingertip, it can reduce the use of the RGB camera to collect images and thus reduce the use time of the RGB camera, which can effectively extend the battery life of the device and improve the overall use experience of palm input.

[0130] In addition, in order to better implement the palm input method of the smart wearable device in the embodiment of the present application, based on the palm input method of the smart wearable device, the embodiment of the present application also provides a palm input device for a smart wearable device, such as Figure 6 FIG. 1 is a schematic diagram of a palm input device for a smart wearable device according to an embodiment of the present application. The palm input device 600 for the smart wearable device includes:

[0131] The projection unit 601 is used to collect scene images through the RGB camera of the smart wearable device to identify the projected palm area;

[0132] An acquisition unit 602 is configured to acquire a first event image of the virtual input interface through an event camera of the smart wearable device after the virtual input interface is projected on the projection palm area;

[0133] A detection unit 603 is used to perform fingertip detection based on the event image set including the first event image by using a trained fingertip detection model to obtain an operation fingertip position of the virtual input interface;

[0134] The input unit 604 is used to perform input processing based on the operating fingertip position and the virtual input interface.

[0135] In some embodiments, the projection unit 601 is used to:

[0136] Acquire scene images by using the RGB camera;

[0137] By training the encoder of the semantic segmentation model, a multi-scale convolution operation is performed based on the scene image to obtain a multi-scale feature map of the scene image;

[0138] Performing channel enhancement processing on the multi-scale feature map through the encoder to obtain a channel-enhanced feature map;

[0139] A target pixel probability map of the scene image is obtained by predicting based on the channel-enhanced feature map through the decoder of the trained semantic segmentation model, wherein the target pixel probability map is used to indicate the probability that each pixel point in the scene image is a pixel point where the palm area is located;

[0140] The scene image is divided based on the target pixel probability map to obtain the projected palm area.

[0141] In some embodiments, the projection unit 601 is used to:

[0142] Performing preliminary feature extraction on the scene image through the encoder to obtain a preliminary scene feature map;

[0143] Through the decoder, a splicing process is performed based on the channel enhanced feature map and the preliminary scene feature map to obtain a spliced ​​feature map;

[0144] Performing self-attention feature extraction on the concatenated feature map to obtain a self-attention feature map;

[0145] Performing prediction processing based on the self-attention feature map to obtain a preliminary pixel probability map of the scene image;

[0146] An upsampling operation is performed on the preliminary pixel probability map to obtain the target pixel probability map.

[0147] In some embodiments, the event image set further includes a second event image, and the acquiring unit 602 is used to:

[0148] After the virtual input interface is projected on the projection palm area, a hand image sequence is collected by an RGB camera of the smart wearable device;

[0149] A superposition process is performed based on the hand image sequence to obtain a second event image of the virtual input interface.

[0150] In some embodiments, the acquisition unit 602 is used to:

[0151] Performing grayscale conversion on each hand image in the hand image sequence to obtain a grayscale image of each hand image;

[0152] Based on the grayscale image of each hand image, obtaining grayscale difference images of adjacent frames of the hand image sequence;

[0153] Based on the grayscale difference images of each adjacent frame and a preset pixel threshold, generating a plurality of change images of the hand image sequence;

[0154] The second event image is obtained by performing a superposition process based on the multiple change images.

[0155] In some embodiments, the acquisition unit 602 is used to:

[0156] The maximum value of the pixel values ​​of the multiple changed images at the pixel point position (x, y) is used as the pixel value of the superimposed image at the pixel point position (x, y), and the multiple changed images are superimposed to obtain the superimposed image as the second event image.

[0157] In some embodiments, the detection unit 603 is used to:

[0158] Performing preliminary feature extraction on each event image in the event image set through the convolution module of the trained fingertip detection model to obtain a preliminary event feature map;

[0159] Extracting features from the preliminary event feature graph through the spatiotemporal feature extraction module of the trained fingertip detection model to obtain a spatiotemporal feature graph;

[0160] The spatiotemporal feature map and the preliminary event feature map are fused by a fusion module of the trained fingertip detection model to obtain a fused feature map;

[0161] The fingertip position is predicted based on the fused feature map to obtain the operating fingertip position.

[0162] In some embodiments, the detection unit 603 is used to:

[0163] Performing residual processing on the fused feature map through the residual module of the trained fingertip detection model to obtain a processed feature map;

[0164] The prediction module of the trained fingertip detection model is used to predict the fingertip position based on the processed feature map to obtain the operating fingertip position.

[0165] In some embodiments, the input unit 604 is used to:

[0166] Acquire the relative distance between each key position of the virtual input interface and the operating fingertip position;

[0167] Based on the relative distance, it is determined whether to trigger a click operation at the button position, so as to use the buttons of the virtual input interface to perform input.

[0168] In specific implementation, the above units can be implemented as independent entities, or can be arbitrarily combined to be implemented as the same or several entities. The specific implementation of the above units can refer to the palm input method embodiment of the smart wearable device above, which will not be repeated here.

[0169] Those skilled in the art will appreciate that all or part of the steps in the palm input method of the above-mentioned smart wearable device can be completed by instructions, or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0170] To this end, an embodiment of the present application provides a computer-readable storage medium, which stores multiple computer programs. The computer programs can be loaded by a processor to execute any palm input method for a smart wearable device provided in an embodiment of the present application.

[0171] The computer-readable storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0172] In the palm input device, computer-readable storage medium, and smart wearable device embodiments of the above-mentioned smart wearable device, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, refer to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working process and beneficial effects of the palm input device, computer-readable storage medium, smart wearable device, and corresponding units of the above-mentioned smart wearable device can refer to the description of the palm input method of the smart wearable device in the above embodiment, and the specific details will not be repeated here.

[0173] The above is a detailed introduction to the palm input method, device, smart wearable device and computer-readable storage medium for a smart wearable device provided in the embodiments of the present application. Specific examples are used in this article to illustrate the principles and implementation methods of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for technical personnel in this field, according to the ideas of the present application, there will be changes in the specific implementation methods and application scope. In summary, the content of this specification should not be understood as a limitation on the present application.

Claims

1. A palm input method for a smart wearable device, characterized in that: The method comprises: The RGB camera of the smart wearable device is used to collect scene images and identify the projected palm area; After the virtual input interface is projected on the projection palm area, a first event image of the virtual input interface is acquired through an event camera of the smart wearable device; By using the trained fingertip detection model, fingertip detection is performed based on the event image set including the first event image to obtain the operating fingertip position of the virtual input interface; Input processing is performed based on the operating fingertip position and the virtual input interface.

2. The palm input method for a smart wearable device according to claim 1, characterized in that: The method of collecting scene images through the RGB camera of the smart wearable device to identify the projected palm area includes: Acquire scene images by using the RGB camera; By training the encoder of the semantic segmentation model, a multi-scale convolution operation is performed based on the scene image to obtain a multi-scale feature map of the scene image; Performing channel enhancement processing on the multi-scale feature map through the encoder to obtain a channel-enhanced feature map; A target pixel probability map of the scene image is obtained by predicting based on the channel-enhanced feature map through the decoder of the trained semantic segmentation model, wherein the target pixel probability map is used to indicate the probability that each pixel point in the scene image is a pixel point where the palm area is located; The scene image is divided based on the target pixel probability map to obtain the projected palm area.

3. The palm input method for a smart wearable device according to claim 2, characterized in that: The method further comprises: Performing preliminary feature extraction on the scene image through the encoder to obtain a preliminary scene feature map; The decoder of the trained semantic segmentation model performs prediction based on the channel enhanced feature map to obtain a target pixel probability map of the scene image, including: Through the decoder, a splicing process is performed based on the channel enhanced feature map and the preliminary scene feature map to obtain a spliced ​​feature map; Performing self-attention feature extraction on the concatenated feature map to obtain a self-attention feature map; Performing prediction processing based on the self-attention feature map to obtain a preliminary pixel probability map of the scene image; An upsampling operation is performed on the preliminary pixel probability map to obtain the target pixel probability map.

4. The palm input method for a smart wearable device according to claim 1, characterized in that: The event image set further includes a second event image, and the method further includes: After the virtual input interface is projected on the projection palm area, a hand image sequence is collected by an RGB camera of the smart wearable device; A superposition process is performed based on the hand image sequence to obtain a second event image of the virtual input interface.

5. The palm input method for a smart wearable device according to claim 4, characterized in that: The performing superposition processing based on the hand image sequence to obtain a second event image of the virtual input interface includes: Performing grayscale conversion on each hand image in the hand image sequence to obtain a grayscale image of each hand image; Based on the grayscale image of each hand image, obtaining grayscale difference images of adjacent frames of the hand image sequence; Based on the grayscale difference images of each adjacent frame and a preset pixel threshold, generating a plurality of change images of the hand image sequence; The second event image is obtained by performing a superposition process based on the multiple change images.

6. The palm input method for a smart wearable device according to claim 5, characterized in that: The performing superposition processing based on the multiple change images to obtain the second event image includes: The maximum value of the pixel values ​​of the multiple changed images at the pixel point position (x, y) is used as the pixel value of the superimposed image at the pixel point position (x, y), and the multiple changed images are superimposed to obtain the superimposed image as the second event image.

7. The palm input method for a smart wearable device according to claim 1, characterized in that: The fingertip detection model is trained to perform fingertip detection based on an event image set including the first event image to obtain the operation fingertip position of the virtual input interface, including: Performing preliminary feature extraction on each event image in the event image set through the convolution module of the trained fingertip detection model to obtain a preliminary event feature map; Extracting features from the preliminary event feature graph through the spatiotemporal feature extraction module of the trained fingertip detection model to obtain a spatiotemporal feature graph; The spatiotemporal feature map and the preliminary event feature map are fused by a fusion module of the trained fingertip detection model to obtain a fused feature map; The fingertip position is predicted based on the fused feature map to obtain the operating fingertip position.

8. The palm input method for a smart wearable device according to claim 7, characterized in that: The predicting of the fingertip position based on the fused feature map to obtain the operating fingertip position includes: Performing residual processing on the fused feature map through the residual module of the trained fingertip detection model to obtain a processed feature map; The prediction module of the trained fingertip detection model is used to predict the fingertip position based on the processed feature map to obtain the operating fingertip position.

9. The palm input method for a smart wearable device according to claim 1, characterized in that: The input processing based on the operating fingertip position and the virtual input interface includes: Acquire the relative distance between each key position of the virtual input interface and the operating fingertip position; Based on the relative distance, it is determined whether to trigger a click operation at the button position, so as to use the buttons of the virtual input interface to perform input.

10. A palm input device for a smart wearable device, characterized in that: The palm input device of the smart wearable device comprises: A projection unit, used to collect scene images through the RGB camera of the smart wearable device to identify the projected palm area; an acquisition unit, configured to acquire a first event image of the virtual input interface through an event camera of the smart wearable device after the virtual input interface is projected on the projection palm area; a detection unit, configured to perform fingertip detection based on an event image set including the first event image by using a trained fingertip detection model, and obtain a fingertip position for operating the virtual input interface; An input unit is used to perform input processing based on the operating fingertip position and the virtual input interface.

11. A smart wearable device, characterized in that: The invention comprises a processor and a memory, wherein a computer program is stored in the memory, and when the processor calls the computer program in the memory, the palm input method of the smart wearable device according to any one of claims 1 to 9 is executed.

12. A computer-readable storage medium, characterized in that: A computer program is stored thereon, and the computer program is loaded by a processor to execute the palm input method for a smart wearable device according to any one of claims 1 to 9.

Citation Information

Cited By

  • Visual multi-mode non-contact gesture unlocking method

    CN120526487A