Password input system, method, device, equipment, storage medium and program product
By processing gesture images through multi-angle shooting and depth-separable convolution, the problems of low accuracy and insufficient security in gesture recognition are solved, enabling high-precision and secure password input.
Patent Information
- Application Number
- CN202511413694.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-01-13
AI Technical Summary
Existing gesture recognition-based password input systems suffer from low password recognition accuracy due to the large changes in the viewing angle of the gesture images, and also pose a security risk of being spied on.
The system employs multi-angle capture of user hand gesture images, combined with a gesture recognition algorithm that incorporates depthwise separable convolution, inverted residual structure, multi-scale feature map detection, and default bounding box mechanism. The system processes multiple frames of gesture images through an information processing device and inputs the recognition results into the business processing terminal through an information transmission device.
It improves the accuracy and security of password recognition, reduces the risk of being spied on due to differences in perspective, and meets the high security requirements of banking and other businesses.
Smart Images

Figure CN121327809A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence, and in particular to a password input system, method, apparatus, device, storage medium, and program product. Background Technology
[0002] With the development of electronic payment and smart terminals, password input, as a key step in identity verification, has increasingly higher requirements for security and convenience.
[0003] Currently, in order to avoid reducing physical contact and lower the risk of leakage, password input systems based on gesture recognition have gradually become a research hotspot. These systems typically acquire images of users' hand gestures through image acquisition devices, and then use feature extraction and pattern recognition algorithms to process the images to obtain the input results.
[0004] However, when users enter their passwords using the above method, the password input system currently suffers from low password recognition accuracy. Summary of the Invention
[0005] This application provides a password input system, method, apparatus, device, storage medium, and program product to improve password recognition accuracy.
[0006] In a first aspect, this application provides a password input system, the system comprising:
[0007] An information acquisition device is used to capture multiple frames of gesture images of the user's hand from n angles during the password input process when the user's hand is inserted into the shielding shell; wherein, the shielding shell is used to form a closed or semi-closed space to shield the user's hand.
[0008] An information processing device is used to process multiple frames of gesture images from each angle using a preset gesture recognition algorithm to obtain gesture recognition results. The gesture recognition algorithm employs depthwise separable convolution, inverted residual structure, multi-scale feature map detection, and a default box mechanism, and the scale of the default box increases linearly as the feature map size decreases.
[0009] An information transmission device is used to receive the gesture recognition result and input the gesture recognition result into a business processing terminal.
[0010] Secondly, this application provides a password input method, the method comprising:
[0011] When the user's hand is inserted into the shielding shell, multiple frames of gesture images are captured from n angles during the user's password input process; wherein, the shielding shell is used to form a closed or semi-closed space to shield the user's hand;
[0012] A preset gesture recognition algorithm is used to process multiple frames of gesture images from each angle to obtain gesture recognition results. The gesture recognition algorithm adopts depthwise separable convolution, inverted residual structure, multi-scale feature map detection and default box mechanism, and the scale of the default box increases linearly as the feature map size decreases.
[0013] The gesture recognition result is input into the business processing terminal.
[0014] Thirdly, this application provides a password input device, the device comprising:
[0015] The acquisition module is used to capture multiple frames of gesture images of the user's hand from n angles during the password input process when the user's hand is inserted into the cover shell; wherein, the cover shell is used to form a closed or semi-closed space to cover the user's hand.
[0016] The processing module is used to process multi-frame gesture images from each angle using a preset gesture recognition algorithm to obtain gesture recognition results. The gesture recognition algorithm uses depthwise separable convolution, inverted residual structure, multi-scale feature map detection, and default box mechanism, and the scale of the default box increases linearly as the feature map size decreases.
[0017] The input module is used to input the gesture recognition result into the business processing terminal.
[0018] Fourthly, this application provides an electronic device, including: at least one processor, and a memory communicatively connected to the processor;
[0019] The memory stores computer-executed instructions;
[0020] The at least one processor executes computer execution instructions stored in the memory to implement the method as described in any of the second aspects.
[0021] Fifthly, this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method as described in any of the second aspects.
[0022] In a sixth aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method described in any of the second aspects.
[0023] The password input system provided in this application enables password input based on user gestures. Specifically, the system includes an information acquisition device, an information processing device, and an information transmission device. The information acquisition device captures multiple frames of gesture images of the user's hand from n angles during the password input process when the user's hand is inserted into the obscuring shell. The information processing device uses a preset gesture recognition algorithm to process the multiple frames of gesture images from each angle to obtain a gesture recognition result. The information transmission device receives the gesture recognition result and transmits it to the business processing terminal to complete the password input. This password input system effectively improves the accuracy of password recognition by capturing multiple frames of gesture images from n angles to comprehensively capture hand features and by employing a gesture recognition algorithm that includes mechanisms such as depthwise separable convolution and inverse residual structures to optimize feature extraction and processing accuracy. Furthermore, the obscuring shell creates a closed or semi-closed space to cover the user's hand, reducing the risk of being spied on during password input and thus effectively improving the security of password input. Attached Figure Description
[0024] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0025] Figure 1 This is a schematic diagram illustrating an application scenario of a password input system provided in an embodiment of this application;
[0026] Figure 2 This is a schematic diagram of the structure of a password input system provided in an embodiment of this application;
[0027] Figure 3A A schematic diagram of the principle of a gesture recognition algorithm provided in this application embodiment. Figure 1 ;
[0028] Figure 3B A schematic diagram of the principle of a gesture recognition algorithm provided in this application embodiment. Figure 2 ;
[0029] Figure 3C Schematic diagram three illustrating the principle of a gesture recognition algorithm provided in this application embodiment;
[0030] Figure 4 A flowchart illustrating a password input method provided in an embodiment of this application;
[0031] Figure 5 This is a schematic diagram of the structure of a password input device provided in an embodiment of this application;
[0032] Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0033] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0034] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0035] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with the relevant laws, regulations, and standards of the relevant countries and regions, have taken necessary confidentiality measures, do not violate public order and good morals, and provide corresponding operation access points for users to choose to authorize or refuse.
[0036] It should be noted that the password input system, method, apparatus, device, storage medium and program product provided in this application can be used in the field of artificial intelligence, or in any field other than artificial intelligence. The application field of the password input system, method, apparatus, device, storage medium and program product in this application is not limited.
[0037] With the widespread adoption of electronic payments and smart terminals, password input, as a crucial step in identity verification, has garnered significant attention for its security and convenience. Traditional password input methods primarily rely on physical keyboards or touchscreens: keyboard-based password input devices typically consist of a multi-grid keyboard and a casing, with an opening only on the user-facing side, while the other sides are closed. Users must insert their hands through the opening and press the corresponding keys to input the password. Touchscreen password input devices, on the other hand, allow users to select numbers on the screen. However, these methods are vulnerable to being spied on or stolen, especially in public places where the possibility of password leakage increases significantly.
[0038] Therefore, to improve password input security, gesture recognition-based password input schemes have emerged. These schemes replace traditional keypad input by recognizing the user's hand gestures, reducing security risks associated with physical contact. Specifically, gesture recognition-based password input systems typically acquire images of the user's hand gestures using image acquisition devices, and then process the images using feature extraction and pattern recognition algorithms to obtain the input result.
[0039] However, the aforementioned gesture recognition-based password input scheme suffers from several drawbacks. Since many gestures are similar, and the viewing angles of the captured gesture images vary greatly, the differences between images of the same gesture from different viewing angles may be greater than the differences between different gestures from the same viewing angle when using traditional feature extraction networks. This ultimately leads to lower accuracy in gesture recognition and an inability to accurately identify the user's input password.
[0040] Therefore, this application provides a password input system, method, apparatus, device, storage medium, and program product, aiming to solve the above-mentioned technical problems of the known art. Specifically, the password input system of this application uses an information acquisition device to capture multiple frames of gesture images of the user's hand from n angles when the user's hand is inserted into the obscuring shell, obtaining the gesture recognition result. The gesture recognition algorithm used employs depthwise separable convolution, inverse residual structure, multi-scale feature map detection, and default box mechanism. Finally, the gesture recognition result is input into the business processing terminal through an information transmission device.
[0041] It should be understood that the password input system of this application can be used in any password input scenario that requires authentication. For example, Figure 1 This is a schematic diagram illustrating an application scenario of a password input system provided in an embodiment of this application. For example... Figure 1 As shown, the system of this application can be used in the password input scenario of banking business processing terminals.
[0042] In this scenario, the banking terminal interacts with the password input system. When using the terminal, the user enters their password through the system. Specifically, the user inserts their hand into the concealed casing and inputs the password by making corresponding gestures. During this process, an information acquisition device captures multiple frames of gesture images from n angles. The information processing device uses a preset gesture recognition algorithm to process these images from each angle to obtain the gesture recognition result. Finally, the information transmission device transmits the gesture recognition result to the banking terminal, enabling it to receive the user's entered password.
[0043] In the above process, by using multi-angle shooting and fusing the recognition results from each angle, the recognition deviation caused by the difference in perspective is reduced, thereby improving accuracy; by forming a closed space by covering the shell, external peeping is avoided, thereby improving security, while meeting the high requirements of banking business for security and accuracy.
[0044] In addition to the password input scenario of the aforementioned banking business processing terminal, the system of this application can also be used in scenarios such as mobile payment terminals, vending machines, and intelligent access control systems, and this application does not limit it to these scenarios.
[0045] Furthermore, the password input system of this application can be integrated into the business processing terminal and share the power or data interface of the business processing terminal, or it can be set up independently and communicate with the business processing terminal via wired or wireless means. This application does not limit this.
[0046] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0047] This application provides a password input system. Figure 2 This is a schematic diagram of the structure of a password input system provided in an embodiment of this application, as shown below. Figure 2 As shown, the system in this embodiment includes an information acquisition device, an information processing device, and an information transmission device. The information acquisition device interacts directly with the user. The output of the information acquisition device is connected to the input of the information processing device. The output of the information processing device is connected to the input of the information transmission device. The output of the information transmission device is connected to the input of the business processing terminal.
[0048] In this embodiment, the information acquisition device is used to capture multiple frames of gesture images of the user's hand during the password input process from n angles when the user's hand is inserted into the shielding shell; wherein, the shielding shell is used to form a closed or semi-closed space to shield the user's hand.
[0049] In this embodiment, the system further includes a shielding shell, which has an opening only on the side facing the user, and the information acquisition device is disposed inside it. Furthermore, a proximity sensor is installed at the entrance of the shielding shell to detect whether the user's hand enters the shell. Specifically, when the proximity sensor detects that the user's hand approaches and enters the opening of the shielding shell, it triggers the information acquisition device to start operating and begin capturing images of the user's hand from n preset angles. If the proximity sensor detects that the user's hand leaves the shielding shell, it sends a stop signal, causing the information acquisition device to pause capturing images, thereby reducing unnecessary resource consumption and ensuring that effective image acquisition only occurs during user operation.
[0050] More specifically, proximity sensors can be infrared proximity sensors, ultrasonic proximity sensors, or capacitive proximity sensors. For example, infrared proximity sensors emit infrared light and receive reflected signals. When a user's hand enters the sensing range, the intensity of the reflected signal changes, thus determining whether the hand has entered. Ultrasonic proximity sensors utilize the emission and echo reception of ultrasonic waves, calculating the time difference of sound wave propagation to sense the distance to the hand and triggering corresponding actions. Capacitive proximity sensors monitor based on the capacitance change caused by a hand approaching, suitable for non-contact sensing scenarios at entrances to obstructed enclosures. The appropriate type of sensor can be selected based on actual application requirements.
[0051] In this embodiment, the information acquisition device specifically includes five high-definition cameras. These cameras are deployed at different locations inside the shielding shell, enabling them to simultaneously capture images of the user's hand extending into the shielding shell from five preset angles (e.g., front, left side, right side, upper diagonally, and lower diagonally), thereby comprehensively capturing the morphological features of the gesture in different dimensions. Each camera has a fast response capability, and can acquire multiple consecutive frames of dynamic gesture images in real time after the proximity sensor triggers the start signal, providing rich and complete raw data support for subsequent gesture recognition algorithms.
[0052] In practical applications, the outer casing can also be completely enclosed, with an opening on one side only when the user needs to enter a password, allowing the user to insert their hand. Furthermore, the information collection device can also include other numbers of high-definition cameras; this embodiment does not limit this, as long as they can capture multiple frames of gesture images of the user's hand from corresponding angles when the user's hand is inserted into the obstructing casing, and can transmit these multiple frames of gesture images to the information processing device. This embodiment does not limit this either.
[0053] As a preferred example, the obscuring shell also incorporates lighting. To ensure image clarity, when a user's hand enters the obscuring shell, the proximity sensor triggers the data acquisition device and simultaneously activates the internal lighting, illuminating the hand area with soft, uniform light. This lighting utilizes a specific wavelength of light, reducing the impact of hand reflections and shadows on image quality. It particularly highlights key gesture features such as finger joints and hand contours, providing stable illumination for the data acquisition device to capture clear gesture images from multiple angles, further ensuring the accuracy of subsequent gesture recognition algorithms.
[0054] In this embodiment, the information processing device is used to process multiple frames of gesture images from each angle using a preset gesture recognition algorithm to obtain gesture recognition results. The gesture recognition algorithm employs depthwise separable convolution, inverse residual structure, multi-scale feature map detection, and a default bounding box mechanism, where the scale of the default bounding box increases linearly as the feature map size decreases. In this embodiment, the information processing device is specifically a microprocessor or a programmable logic device.
[0055] In this embodiment, the gesture recognition algorithm is based on MobileNetV2 and the SSD object detection network. This algorithm uses MobileNetV2 as the basic feature extraction network, leveraging its depthwise separable convolution characteristics to achieve efficient feature extraction. Simultaneously, it integrates the multi-scale feature map detection and default box mechanism of the SSD object detection network, forming a lightweight network structure suitable for gesture recognition. This algorithm retains the efficiency of MobileNetV2 in mobile deployments while improving the robustness of recognition for gestures of different angles and sizes through the multi-scale detection capabilities of SSD, making it particularly suitable for the multi-angle gesture image processing requirements of this system.
[0056] Specifically, in this embodiment, the information processing device uses a gesture recognition algorithm to process multiple frames of gesture images from each angle to obtain corresponding sub-recognition results. Then, the sub-recognition results for each angle are averaged to obtain the average probability of each preset gesture being recognized, and the gesture recognition result is obtained based on the average probability. The sub-recognition results indicate the probability of each preset gesture being recognized.
[0057] It should be understood that sub-recognition results are intermediate outputs of single-angle gesture recognition. For example, when an information processing device processes 10 frames of gesture images from a frontal angle, it uses an algorithm to calculate the probability of each preset gesture, such as "number 1" and "number 2," at that angle (e.g., the probability of "number 3" is 92%), forming the sub-recognition result for the frontal angle. Similarly, sub-recognition results for other angles can be obtained. The gesture recognition result is a comprehensive judgment of all angle sub-recognition results: by calculating the average probability of each preset gesture at different angles (e.g., the average probability of "number 3" at 5 angles is 89%), the preset gesture with the highest average probability is finally determined as the final recognition result, thereby reducing the impact of single-angle recognition bias on the overall result.
[0058] By integrating the recognition results from multiple angles, the system can effectively offset recognition errors caused by limitations in the shooting angle, partial obstruction, or lighting interference from a single angle. This results in a final gesture recognition result that more closely matches the user's actual input intent. For example, from one angle, a finger might obstruct the view and misjudge "5" as "3," but the recognition results from other angles can correct this deviation. By calculating the average probability, the correct result is ultimately selected, significantly improving the accuracy and stability of password recognition and reducing input errors caused by single-viewpoint errors.
[0059] More specifically, in this embodiment, for multi-frame gesture images at each angle, the information processing device uses depthwise separable convolution to apply a single convolutional filter to each input channel of each frame of the gesture image for lightweight filtering, and applies 1x1 pointwise convolution to linearly combine the features of each input channel after depthwise convolution processing to construct new fusion features; it uses an inverse residual structure to process the fusion features, and extracts gesture features of different size feature layers through multi-scale feature map detection; it uses a default box mechanism to generate gesture bounding boxes and category confidence to obtain the sub-recognition result corresponding to the current angle.
[0060] It should be understood that depthwise separable convolution reduces computation while preserving the detailed features of gestures (such as finger texture and joint shape) by splitting standard convolution into depthwise convolution and pointwise convolution; inverse residual structure enhances the non-linear expressive power of features by first increasing the dimensionality to expand the number of feature channels and then reducing the dimensionality to compress them, thus capturing subtle differences in gestures more accurately; multi-scale feature map detection extracts information from feature layers of different resolutions, which can recognize both the overall gesture outline (such as the width of the palm opening) and capture local details (such as fingertip posture); the default box mechanism configures default boxes with increasing scales for different feature layers to adapt to the different sizes of gestures in the image, and the final generated category confidence directly reflects the recognition probability of each preset gesture, thus constituting the sub-recognition result for that angle.
[0061] As can be seen from the above, the gesture recognition algorithm used by the information processing device achieves efficient feature extraction by leveraging the depthwise separable convolution of MobileNetV2. It combines multi-scale feature detection of SSD with the default box mechanism to adapt to gesture changes of different angles and sizes. At the same time, it enhances the feature representation ability through the inverse residual structure to distinguish subtle differences. Furthermore, it reduces the single-viewpoint bias through the averaging fusion strategy of multi-angle sub-recognition results. This not only achieves efficient operation on lightweight devices to meet real-time requirements, but also significantly improves the accuracy and robustness of gesture recognition. It effectively solves the problems of traditional algorithms being sensitive to viewpoints and having difficulty distinguishing similar gestures. It works in synergy with the system's multi-angle acquisition scheme to provide reliable recognition assurance for password input.
[0062] As a further explanation of the gesture recognition algorithm, this algorithm is a gesture recognition algorithm based on MobileNetV2 and SSD object detection network. MobileNetV2 is used as the basic feature extraction network, and the multi-scale feature map detection and default box mechanism of SSD object detection network are integrated to form a lightweight network structure. The overall structure can be divided into feature extraction layer, enhancement processing layer, multi-scale detection layer and result output layer from bottom to top.
[0063] The feature extraction layer employs depthwise separable convolution, the basic idea of which is to divide the complete convolution into two independent decomposed versions. The first layer is called depthwise convolution, which performs lightweight filtering by applying a single convolutional filter to each input channel; the second layer is a 1x1 convolution, called pointwise convolution, which constructs new features by calculating a linear combination of the input channels, preserving the detailed features of gestures (such as finger texture and joint shape) while reducing computational cost.
[0064] Specifically, Figure 3A A schematic diagram of the principle of a gesture recognition algorithm provided in this application embodiment. Figure 1 ,like Figure 3A As shown, the feature extraction layer of the algorithm is presented in the form of multiple connected modules. Each module realizes feature extraction through two key operation units: first, a depthwise convolution unit that filters each channel of the input feature map individually; second, a pointwise convolution unit that integrates channel features through 1x1 convolution. Together, they form a depthwise separable convolution structure, ensuring that lightweight feature extraction is achieved while maintaining the feature map size (C×H×W) basically stable.
[0065] Figure 3B A schematic diagram of the principle of a gesture recognition algorithm provided in this application embodiment. Figure 2 The enhanced processing layer is deployed with features such as Figure 3BThe inverted residual structure shown here suffers from severe information loss due to ReLU in feature maps with a small number of channels. Therefore, LinearBottlenecks and InvertedResidual are introduced to enhance feature representation through "dimensional expansion-dimensional compression" and adapt to the capture of subtle differences in gestures.
[0066] Specifically, such as Figure 3B As shown, the structure begins with a 1x1 convolutional layer (Conv2d) with a stride of 1, using the ReLU6 activation function for dimensionality upscaling. Next, a depthwise separable convolutional layer (Conv2d dw) with a 3x3 kernel size and a stride of 2, using the ReLU6 activation function, is applied for feature extraction. Finally, a 1x1 convolutional layer (Conv2d) with a stride of 1, using the Linear activation function, is applied for dimensionality reduction. This dimensionality upscaling-feature extraction-dimensionality reduction operation effectively enhances feature representation and avoids information loss. The structure includes parameters such as the expansion factor t, the number of output feature channels c, the number of repetitions n of the inverse residual structure, and the stride s.
[0067] The multi-scale detection layer adds convolutional feature layers with gradually decreasing sizes to the truncated end of MobileNetV2, allowing for multi-scale prediction. Unlike Overfeat and YOLO, which operate on a single-scale feature map, it can achieve multi-dimensional feature extraction of both the overall contour and local details.
[0068] Figure 3C The third schematic diagram illustrates the principle of a gesture recognition algorithm provided in this application embodiment. The multi-scale detection layer specifically employs the following... Figure 3C The diagram shows a complex neural network module. From left to right, the module includes: a 1x1 convolutional layer (Conv2d) with a stride of 1 and using the ReLU6 activation function for initial feature processing; a depthwise separable convolutional layer (Conv2d dw) with a 3x3 kernel size and a stride of 1, using the ReLU6 activation function for further feature extraction; and then another 1x1 convolutional layer (Conv2d) with a stride of 1 and using the Linear activation function for feature adjustment. Furthermore, the module contains additional branches such as average pooling (AP), global average pooling (GAP), 1x1 convolution (Conv1d), and the Sigmoid activation function. These operations may be used for feature fusion or attention mechanisms to ensure effective extraction of features at different scales, adapting to the morphological changes of gestures at different angles and sizes.
[0069] The output layer consists of a fixed set of default boxes associated with each feature unit of multiple feature maps at the top of the network. These default boxes are convolved onto the feature maps, and each location predicts the offset relative to the default boxes and the confidence score for each class. The scale of the prior boxes follows a linear increasing rule; as the size of the feature maps decreases, the scale of the prior boxes increases linearly. The final output is the bounding box offset and the class confidence score.
[0070] More specifically, in this embodiment, because the input feature channels and output feature channels are not equal, the feature map addition operation of short connections cannot be performed. Therefore, when the inverse residual structure processes the fused features, the first layer has no short connection branches, the second layer has a step size of 1, and the number of input feature matrix channels is equal to the number of output feature matrix channels of the previous layer, and the number of output feature channels is equal to the number of input channels. The output features are added to the input features through short connections, thereby enhancing the non-linear expressive power of the features and improving the distinguishability of gestures with similar shapes.
[0071] More specifically, in this embodiment, multi-scale feature map detection is achieved by adding convolutional feature layers with gradually decreasing sizes to the tail of the truncated base network. The default box mechanism associates each feature unit of multiple feature maps at the top of the network with a fixed set of default boxes. Each position predicts the offset relative to the default box and the confidence score of each category. The scale of the default box increases linearly as the size of the feature map decreases, thereby adapting to the size changes of the gesture at different angles and distances, and ensuring the comprehensiveness of feature extraction.
[0072] More specifically, when extracting prior bounding boxes, the settings are mainly based on two aspects: scale (size) and aspect ratio. Because the gesture detection in this algorithm is a single-object detection and the target occupies approximately 70% of the area of the gesture image, the prior bounding boxes are actually generated on the feature map output at the last three scales, thereby speeding up the detection process. For example, the sk value for layer 09 is 0.43, with initial prior bounding boxes (83,83), (117,117), (117,59), and (59,117). The default number of prior bounding boxes is 4, and the total number of prior bounding boxes is 144. The sk value for layer 10 is 0.67, with initial prior bounding boxes (128,128), (181,181), (181,91), and (91,181). The default number of prior bounding boxes is 4, and the total number of prior bounding boxes is 36. The sk value for layer 11 is 0.9, with initial prior bounding boxes (173,173), (245,245), (181,122), and (128,122). The default number of prior bounding boxes is 4, and the total number of prior bounding boxes is 4. This adapts to the size changes of gestures at different angles and distances, ensuring the comprehensiveness of feature extraction.
[0073] In this embodiment, the algorithm loss function is defined as a weighted sum of the classification confidence loss and the location regression loss. Specifically, the classification confidence loss function is: ,in, This is used to indicate that the i-th prior box matches the j-th ground truth bounding box and is of category k. The value used to represent the confidence that the i-th prior box is predicted to be of category k is denoted by c, where c represents the i-th prior box. The location regression loss function is expressed as: ,in, , , , , , , , , , , , These are used to represent the offset of the predicted bounding box relative to the center coordinates (x, y) of the positive sample prior box and the width and height (w, h) of the logarithmic space, respectively. , , , These are used to represent the offset of the true bounding box relative to the center coordinates (x, y) of the positive sample prior box and the width and height (w, h) of the logarithmic space, respectively. , , y, w, and h represent the x-coordinates of the predicted bounding box, the positive sample prior box, and the ground truth bounding box, respectively.
[0074] Based on this, the algorithm loss function is expressed as: .
[0075] During training, data collection begins by identifying several candidates who place their hands inside an occlusion shell, making corresponding gestures and slowly rotating their palms. A camera inside the shell records one minute of video for each tagged gesture. The detected video frames are saved as PNG files, and the marked location information is saved as TXT files, named according to the tag type. The tags for gestures "0-9" are assigned ten-digit integers from 0 to 9; "10" is labeled with the integer 10 (representing the correct key on a keyboard); "20" with the integer 11 (representing the confirm key on a keyboard); and "80" with the integer 12 (representing the cancel key on a keyboard). During training, the ground truth labels are assigned to a specific output from a fixed set of detector outputs. Training also includes selecting a series of default bounding boxes and the feature map scale used for detection, as well as hard-negative-mining and data augmentation strategies. In terms of matching strategy, each real label box is initially matched with the default box that has the best Jacobian overlap value, and the default box is paired with all real label boxes as long as the Jacobian overlap value between the two is greater than the threshold of 0.5.
[0076] When determining the gesture recognition result based on this gesture recognition algorithm, for multiple frames of gesture images at each angle, basic features are first extracted through depthwise separable convolution, and then the feature representation is enhanced by an inverse residual structure (combined with Linear Bottlenecks to solve the information loss problem) to accurately capture subtle differences in gestures. Subsequently, gesture details of different size feature layers are extracted by a multi-scale detection layer, and the category confidence of each frame image is generated by combining the default box mechanism. After fusing the results of multiple frames, the sub-recognition result of that angle (including the probability of each preset gesture) is obtained. The sub-recognition result is the intermediate output of single-angle gesture recognition, which is used to indicate the probability of each preset gesture being recognized. Finally, all angle sub-recognition results are summarized, the average probability of each preset gesture at different angles is calculated, and the preset gesture with the highest average probability is determined as the final recognition result, thereby reducing the impact of single-angle recognition bias on the overall result.
[0077] Based on the aforementioned gesture recognition algorithm, depthwise separable convolution reduces computational cost by splitting standard convolution into depthwise convolution and pointwise convolution, preserving detailed gesture features (such as finger texture and joint shape). The inverse residual structure enhances the non-linear representation of features, more accurately distinguishing gestures with similar shapes. Multi-scale detection and default box mechanisms adapt to the shape changes of gestures at different angles and sizes, solving the problem of traditional algorithms being sensitive to viewing angles. The multi-angle fusion strategy further reduces recognition bias caused by limitations in shooting, occlusion, or lighting interference from a single viewpoint, making the final result more closely match the user's true input intent. This algorithm retains the efficiency of MobileNetV2 for mobile deployment, meeting real-time requirements, while also improving recognition accuracy and robustness through multi-dimensional optimization. It synergizes with the system's multi-angle acquisition scheme, effectively solving the problems of traditional algorithms being sensitive to viewing angles and having difficulty distinguishing similar gestures, providing reliable recognition assurance for password input.
[0078] In this embodiment, the information transmission device receives the gesture recognition result and inputs it into the service processing terminal. Specifically, the information transmission device can establish a data transmission link with the service processing terminal via a wired connection (such as a USB interface or serial port) or a wireless communication method (such as Bluetooth, Wi-Fi, or NFC). As a preferred example, after receiving the gesture recognition result output by the information processing device, the data is encrypted to prevent it from being stolen or tampered with during transmission. Then, the encrypted recognition result is sent to the service processing terminal according to a preset communication protocol. After receiving the data, the service processing terminal restores the information through a corresponding decryption mechanism, completes the verification of the user's input password and subsequent service processing, thereby achieving a secure closed loop for password input.
[0079] The password input system provided in this embodiment protects user privacy by setting up a shielding shell to form a closed or semi-closed space. It uses an information acquisition device to capture images of the user's hand gestures from n angles to comprehensively capture features. It processes the images using a gesture recognition algorithm that integrates depthwise separable convolution, inverted residual structure, multi-scale feature map detection, and default box mechanism. The recognition results are transmitted to the business processing terminal through an information transmission device, thus constructing a complete gesture password input scheme.
[0080] The system described in this embodiment effectively solves the problems in known technologies, such as low password recognition accuracy due to limited gesture acquisition angles, insufficient recognition algorithm for complex gestures, and environmental interference, as well as the security risks caused by the vulnerability of traditional input methods to eavesdropping. It improves the convenience of password input while taking into account both recognition accuracy and operational security.
[0081] As a further explanation of the password input system of this application, this application also provides an embodiment of the password input system, which details how the information processing device determines the gesture recognition result and how the information transmission device transmits the gesture recognition result.
[0082] In this embodiment, the information processing device is further configured to determine the final gesture recognition result based on the gesture recognition results corresponding to the gesture images of the previous m frames before the current time when the gesture recognition result indicates that the current gesture is a preset confirmation gesture, a preset cancellation gesture, or a preset correction gesture. Correspondingly, the information transmission device is specifically configured to: receive the final gesture recognition result and transmit the final gesture recognition result to the service processing terminal.
[0083] Specifically, when determining the final gesture recognition result based on the gesture recognition results corresponding to the gesture images of the previous m frames, the following situations apply:
[0084] ① If the gesture recognition result indicates that the current gesture is a preset confirmation gesture, then the gesture recognition result corresponding to the gesture images of the previous m frames before the current moment is taken as the final gesture recognition result. More specifically, m is 30-60 frames (corresponding to 0.5-1 seconds of video content). The system first sorts the recognized digital gestures in the previous m frames by timestamp, and removes momentary misidentified frames by checking the consistency of consecutive frames (if more than 3 consecutive frames are recognized consistently, they are considered valid); then, the valid frames are deduplicated (if the same digital gesture appears in no more than 5 consecutive frames, it is merged into 1 input), and finally, they are combined into an ordered digital sequence according to the time order. The overall confidence of the sequence is calculated synchronously (the average confidence of all valid frames must be ≥0.8). If the standard is met, it is used as the confirmed password input; otherwise, a re-entry prompt is triggered.
[0085] ② If the gesture recognition result indicates that the current gesture is a preset cancellation gesture, then the gesture recognition results corresponding to the previous m frames of gesture images are discarded, and the final gesture recognition result is set to null or a cancellation command is sent. More specifically, the system immediately clears the cached previous m frames of gesture feature data, recognition results, and timestamp records, and simultaneously sends an encrypted cancellation command (including the device's unique identifier and timestamp) to the business processing terminal through the information transmission device; after receiving the command, the terminal clears the input sequence on the display interface, pops up a "Input Cancelled" prompt box, and automatically closes after a 3-second countdown, during which new gesture input is prohibited to avoid accidental operation, and the input state is restored after the countdown ends.
[0086] ③ If the gesture recognition result indicates that the current gesture is a preset correction gesture, then discard the gesture recognition result corresponding to the gesture image of the previous frame or the previous k frames before the current time, and determine the final gesture recognition result based on the gesture recognition results of the remaining frames; where k is a positive integer less than m. More specifically, combine the gesture recognition results of the remaining frames in chronological order to form a new input sequence to be confirmed; if a preset confirmation gesture is recognized again within a preset time, then the new input sequence to be confirmed is used as the final gesture recognition result; if a preset cancellation gesture is recognized again within a preset time, then discard the new input sequence to be confirmed and set the final gesture recognition result to empty; if a new numeric gesture is recognized within a preset time, then add the new numeric gesture to the end of the new input sequence to be confirmed to update the input sequence to be confirmed.
[0087] In this embodiment, the information processing device realizes complete interactive control of the gesture input process by logically processing three control gestures: confirmation gesture is used to submit a valid input sequence, cancellation gesture is used to fully reset the input, and correction gesture is used to partially correct the input content. The three operations cooperate with each other to form a closed-loop input error correction mechanism.
[0088] The method in this embodiment retains the convenience of contactless gesture input while improving input accuracy through multi-frame verification, abnormal frame removal, and hierarchical control logic. It solves the problems of easy accidental touches and difficulty in correction in traditional gesture input. At the same time, the directional transmission of information through the information transmission device ensures the security of password information, thereby improving the overall efficiency and reliability of password input.
[0089] This application also provides a password input method. Figure 4 This is a flowchart illustrating a password input method provided in an embodiment of this application. Figure 4 As shown, the password input method of this application includes:
[0090] S401 captures multiple frames of hand gesture images from n angles during the user's password input process when the user's hand is inserted into the cover.
[0091] The shielding shell is used to create a closed or semi-closed space to shield the user's hands.
[0092] Specifically, in this embodiment, n takes the value of 3-5, and each angle is evenly distributed along the circumference of the inner wall of the shielding shell (e.g., 0°, 90°, 180°, 270°). Each angle is equipped with a high-definition camera (resolution not less than 1080P) with a shooting frame rate of 30-60fps. The inner wall of the shielding shell is made of diffuse reflective material, and with the embedded infrared fill light (wavelength 850nm), it can clearly capture gesture details (such as finger bending angle and joint contour) without external light interference. The camera lens is facing the central area inside the shell, covering the main range of the user's hand activities.
[0093] Furthermore, for an explanation of the occlusion shell and multi-frame gesture images, please refer to the aforementioned system embodiments, which will not be repeated here.
[0094] S402 uses a preset gesture recognition algorithm to process multiple frames of gesture images from each angle to obtain gesture recognition results.
[0095] The gesture recognition algorithm employs depthwise separable convolution, inverted residual structure, multi-scale feature map detection, and a default bounding box mechanism, with the default bounding box scale increasing linearly as the feature map size decreases.
[0096] In this embodiment, for multiple frames of images from each angle, basic features are first extracted through depthwise separable convolution, and then the feature representation is enhanced through an inverse residual structure (adapting to subtle differences in gestures through "dimensional expansion-dimensional compression"). Subsequently, features of different sizes are extracted using a multi-scale detection layer, and the category confidence of each frame is generated by combining the default bounding box mechanism. After fusing the results of multiple frames from the same angle, sub-recognition results are obtained. Finally, all sub-recognition results from all angles are summarized, and the average probability is calculated to determine the final gesture recognition result, effectively reducing the recognition bias of a single angle.
[0097] S403, input the gesture recognition result into the business processing terminal.
[0098] It should be understood that the descriptions of each feature can be found in the aforementioned system embodiments, and will not be repeated here.
[0099] In this embodiment, physical privacy is achieved by covering the outer shell. Combined with multi-angle shooting and optimized gesture recognition algorithm, the privacy of the password input process is ensured while improving the accuracy and robustness of gesture recognition. This solves the problems of easy peeping, unhygienic contact input, and the influence of viewing angle on gesture recognition in traditional password input methods, providing users with a safe, convenient, and contactless password input experience.
[0100] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0101] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0102] The above embodiments describe a password input method from the perspective of process flow. The following embodiments describe a password input device from the perspective of virtual module or virtual unit. For details, please refer to the following embodiments.
[0103] This application provides a password input device. Figure 5 This is a schematic diagram of the structure of a password input device provided in an embodiment of this application, as shown below. Figure 5 As shown, the device includes:
[0104] The acquisition module 51 is used to capture multiple frames of gesture images of the user's hand from n angles during the password input process when the user's hand is inserted into the cover shell; wherein, the cover shell is used to form a closed or semi-closed space to cover the user's hand.
[0105] The processing module 52 is used to process multi-frame gesture images from each angle using a preset gesture recognition algorithm to obtain gesture recognition results. The gesture recognition algorithm uses depthwise separable convolution, inverted residual structure, multi-scale feature map detection and default box mechanism, and the default box scale increases linearly as the feature map size decreases.
[0106] Input module 53 is used to input the gesture recognition result into the business processing terminal.
[0107] The password input device provided in this application is applicable to the above-described password input method embodiments, and will not be described again here.
[0108] This application provides an electronic device. Figure 6 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 6 As shown, Figure 6 The illustrated electronic device includes a processor 61 and a memory 62. The processor 61 and the memory 62 are connected, for example, via a bus 63. Optionally, the electronic device may also include a transceiver 64. It should be noted that in practical applications, the transceiver 64 is not limited to one type, and the structure of this electronic device does not constitute a limitation on the embodiments of this application.
[0109] Processor 61 may be a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 61 may also be a combination that implements computational functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, etc.
[0110] Bus 63 may include a pathway for transmitting information between the aforementioned components. Bus 63 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Bus 63 may be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus 63 is represented by only one thick line, but this does not mean that there is only one bus 63 or one type of bus 63.
[0111] The memory 62 may be a read-only memory (ROM) or other type of static storage device capable of storing static information and instructions, random access memory (RAM) or other type of dynamic storage device capable of storing information and instructions, or electrically erasable programmable read-only memory (EEPROM), compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0112] The memory 62 stores computer execution instructions for implementing the scheme of this application, and the processor 61 controls the execution. The processor 61 executes the computer execution instructions stored in the memory 62 to implement the content shown in the foregoing method embodiments.
[0113] This application also provides a computer-readable storage medium, which may include various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. Specifically, the computer-readable storage medium stores computer-executable instructions, which are used to implement the methods in the above embodiments.
[0114] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the technical solution of the above method embodiments. Its implementation principle and technical effects are similar, and will not be repeated here.
[0115] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0116] It should be understood that this application is not limited to the precise structure described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this application is limited only by the appended claims.
Claims
1. A password entry system, characterized by The system comprises: An information acquisition device configured to capture multiple gesture pictures of a user's hand during a password input process from n angles when the user's hand is inserted into a shielding shell; wherein the shielding shell is configured to form a closed or semi-closed space to shield the user's hand; An information processing device configured to process the multiple gesture pictures of each angle by using a preset gesture recognition algorithm to obtain a gesture recognition result; wherein the gesture recognition algorithm uses a depth separable convolution, an inverted residual structure, a multi-scale feature map detection, and a default box mechanism, and the scale of the default box increases linearly with the decrease of the feature map size; An information transmission device configured to receive the gesture recognition result and input the gesture recognition result into a service processing terminal.
2. The system of claim 1, wherein, The information processing device is specifically configured to: process the multiple gesture pictures of each angle by using the gesture recognition algorithm to obtain a corresponding sub-recognition result; wherein the sub-recognition result is used to indicate the probability of each preset gesture being recognized; average the sub-recognition results corresponding to each angle to obtain an average probability of each preset gesture being recognized, and obtain the gesture recognition result based on the average probability.
3. The system of claim 1 or 2, wherein, The gesture recognition algorithm is a gesture recognition algorithm based on MobileNetV2 and SSD target detection network.
4. The system of claim 3, wherein, The information processing device is specifically configured to: for the multiple gesture pictures of each angle, apply a single convolution filter to each input channel of each gesture picture to perform lightweight filtering by using the depth separable convolution, and apply a 1x1 point convolution to linearly combine the features of each input channel after depth convolution processing to construct new fusion features; process the fusion features by using the inverted residual structure, and extract gesture features of different size feature layers by using the multi-scale feature map detection; generate gesture bounding boxes and class confidence by using the default box mechanism to obtain the sub-recognition result corresponding to the current angle.
5. The system of claim 4, wherein, When the inverted residual structure processes the fusion features, the first layer has no short connection branch, the second layer has a step distance of 1 and an input feature matrix channel number equal to an output feature matrix channel number of the previous layer, and the output feature and the input feature are added through a short connection.
6. The system of claim 4, wherein, The multi-scale feature map detection is realized by adding convolution feature layers with gradually decreasing sizes at the tail end of the truncated base network, and the default box mechanism is that a fixed group of default boxes is associated with each feature unit of multiple feature maps at the top of the network, the offset relative to the default box and the class confidence score of each position are predicted, and the scale of the default box increases linearly with the decrease of the feature map size.
7. The system of claim 3, wherein, The information processing device is further configured to: when the gesture recognition result indicates that the current gesture is a preset confirmation gesture, a preset cancellation gesture, or a preset correction gesture, determine a final gesture recognition result according to gesture recognition results corresponding to the previous m gesture pictures before the current time; correspondingly, the information transmission device is specifically configured to: receive the final gesture recognition result and input the final gesture recognition result into the service processing terminal.
8. The system of claim 7, wherein, The information processing device is specifically configured to: If the gesture recognition result indicates that the current gesture is the preset confirmation gesture, gesture recognition results corresponding to the previous m frames of gesture pictures before the current time are taken as the final gesture recognition result; If the gesture recognition result indicates that the current gesture is the preset cancel gesture, gesture recognition results corresponding to the previous m frames of gesture pictures before the current time are discarded, and the final gesture recognition result is emptied or a cancel instruction is sent; If the gesture recognition result indicates that the current gesture is the preset correction gesture, gesture recognition results corresponding to the previous one frame or the previous k frames of gesture pictures before the current time are discarded, and the final gesture recognition result is determined according to gesture recognition results of the remaining frames; k is a positive integer less than m.
9. The system of claim 8, wherein, The information processing device is specifically used for: If the gesture recognition result indicates that the current gesture is the preset correction gesture, gesture recognition results of the remaining frames are combined in time sequence to form a new to-be-confirmed input sequence; If the preset confirmation gesture is recognized again within a preset time, the new to-be-confirmed input sequence is taken as the final gesture recognition result; If the preset cancel gesture is recognized again within a preset time, the new to-be-confirmed input sequence is discarded and the final gesture recognition result is emptied; If a new digital gesture is recognized within a preset time, the new digital gesture is added to the end of the new to-be-confirmed input sequence to update the to-be-confirmed input sequence.
10. A password input method characterized by comprising: The method comprises: When a user's hand is inserted into a shielding shell, multiple frames of gesture pictures of the user's hand during inputting a password are captured from n angles; the shielding shell is used to form a closed or semi-closed space to shield the user's hand; A preset gesture recognition algorithm is used to process the multiple frames of gesture pictures of each angle to obtain gesture recognition results; the gesture recognition algorithm uses deep separable convolution, reverse residual structure, multi-scale feature map detection, and default frame mechanism, and the scale of the default frame increases linearly with the decrease of the feature map size; The gesture recognition results are input into a business processing terminal.
11. A password input device, characterized by The device comprises: A collection module is configured to capture multiple frames of gesture pictures of a user's hand during inputting a password from n angles when the user's hand is inserted into a shielding shell; the shielding shell is used to form a closed or semi-closed space to shield the user's hand; A processing module is configured to use a preset gesture recognition algorithm to process the multiple frames of gesture pictures of each angle to obtain gesture recognition results; the gesture recognition algorithm uses deep separable convolution, reverse residual structure, multi-scale feature map detection, and default frame mechanism, and the scale of the default frame increases linearly with the decrease of the feature map size; An input module is configured to input the gesture recognition results into a business processing terminal.
12. An electronic device, comprising: Comprise: At least one processor, and a memory connected in communication with the processor; The memory stores computer execution instructions; The at least one processor executes the computer execution instructions stored in the memory to realize the method of claim 10.
13. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer-executable instructions, which, when executed by a processor, implement the method of claim 10.
14. A computer program product, characterised in that, A computer program is included, which, when executed by a processor, implements the method of claim 10.