Payment input method and device based on intelligent glasses, intelligent glasses and medium

By capturing and parsing images of the payment page, the complexity and compatibility issues of smart glasses payment operations have been resolved, enabling convenient payment without the need for external devices.

CN121810285APending Publication Date: 2026-04-07CHINA UNIONPAY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-11
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing smart glasses require external terminal devices or bound devices to operate during the payment process, making it impossible to complete the payment when the user's hands are occupied, and they cannot adapt to diverse front-end payment pages, resulting in complicated and inconvenient payment operations.

Method used

Smart glasses capture images of the payment page, analyze the images and interact with the page to establish a mapping relationship between the image position and the page position, receive voice commands to determine the keyboard key positions, generate simulated click event commands to complete the payment information input, and realize payment operations without external devices.

Benefits of technology

It lowers the requirements for payment operations on smart glasses, making it convenient to pay even when the user's hands are occupied. It is compatible with a variety of payment pages, solves the technical problems of user payments, and improves payment accuracy and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121810285A_ABST
    Figure CN121810285A_ABST
Patent Text Reader

Abstract

The invention discloses a payment input method and device based on intelligent glasses, the intelligent glasses and a medium, and belongs to the field of intelligent terminals. The method comprises the following steps: acquiring and analyzing a payment page image, and interacting with a payment page to obtain image position information of keyboard keys in the payment page image; according to the mapping relation between the image position in the payment page image and the page position in the payment page, obtaining the page position information of the keyboard key corresponding to the image position information of the keyboard key; receiving a payment voice instruction of a user, and determining page position information of keyboard keys corresponding to payment information obtained by converting the payment voice instruction; and generating a simulated click event instruction and sending the simulated click event instruction to the payment page to perform simulated click on the keyboard keys in the payment page to complete input of the payment information, the simulated click event instruction including page position information of the keyboard keys corresponding to the payment information. According to the embodiment of the invention, the payment operation based on the intelligent glasses is simpler and more convenient.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of smart terminals, and in particular relates to a payment input method, device, smart glasses and medium based on smart glasses. Background Technology

[0002] With the continuous development of smart terminal technology, the application areas of smart wearable devices such as smart glasses are becoming increasingly widespread. For example, smart glasses can be used for offline payments. Smart glasses can scan a merchant's QR code and transmit the information obtained from the scanned QR code to a terminal device or cloud backend bound to the smart glasses. The terminal device or cloud backend then calls the user's payment application and connects with the merchant's payment system through a backend interface to complete the amount confirmation and deduction. However, the operation of entering the amount during the payment process relies on the terminal device bound to or connected to the smart glasses. If the user's hands are occupied and it is inconvenient to operate, payment cannot be made, making the user's requirements for payment operations based on smart glasses relatively high and complex. Summary of the Invention

[0003] This application provides a payment input method, device, smart glasses, and medium based on smart glasses, which makes payment operations based on smart glasses more convenient.

[0004] In a first aspect, embodiments of this application provide a payment input method based on smart glasses, applied to smart glasses. The method includes: acquiring a payment page image, parsing the payment page image, and interacting with the payment page to obtain image position information of keyboard keys in the payment page image; obtaining page position information of keyboard keys corresponding to the image position information of keyboard keys according to the mapping relationship between the image position in the payment page image and the page position in the payment page; receiving a user's payment voice command, determining the page position information of keyboard keys corresponding to the payment information converted from the payment voice command; generating a simulated click event command and sending it to the payment page to simulate clicking the keyboard keys in the payment page to complete the input of payment information. The simulated click event command includes the page position information of the keyboard keys corresponding to the payment information.

[0005] Secondly, embodiments of this application provide a payment input device based on smart glasses, applied to smart glasses. The device includes: an image processing module, used to acquire a payment page image, parse the payment page image, and interact with the payment page to obtain image position information of keyboard keys in the payment page image; and used to obtain page position information of the keyboard keys corresponding to the image position information of the keyboard keys according to the mapping relationship between the image position in the payment page image and the page position in the payment page; a voice processing module, used to receive the user's payment voice command and determine the page position information of the keyboard keys corresponding to the payment information converted from the payment voice command; and a simulated click module, used to generate a simulated click event command and send it to the payment page to simulate clicking the keyboard keys in the payment page to complete the input of payment information. The simulated click event command includes the page position information of the keyboard keys corresponding to the payment information.

[0006] Thirdly, embodiments of this application provide a smart glasses, including: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the payment input method based on smart glasses according to the first aspect.

[0007] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the payment input method based on smart glasses as described in the first aspect.

[0008] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the payment input method based on smart glasses as described in the first aspect.

[0009] This application provides a payment input method, device, smart glasses, and medium based on smart glasses. The smart glasses can capture and parse payment page images, and interact with the payment page to obtain the image position information of keyboard keys in the payment page image. Based on the mapping relationship between the image positions in the payment page image and the page positions in the payment page, the page position information of the keyboard keys in the payment page is obtained. The user's payment voice command can be converted into payment information, and the page position information of the corresponding keyboard key can be determined. A simulated click event command is generated based on the page position information of the keyboard key corresponding to the payment information and sent to the payment page. The payment page responds to the simulated click event command and can perform simulated clicks on the corresponding keyboard keys in the payment page, thereby inputting payment information on the payment page. This payment input process does not require the user to operate any external or bound terminal devices to the smart glasses. Even if the user's hands are occupied, as long as the user gives a payment voice command, the simulated clicks of the keyboard keys on the payment page can be triggered to complete the payment process, thereby reducing the requirements for payment operations based on smart glasses and making payment operations based on smart glasses more convenient. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of this application, the accompanying drawings used in the embodiments of this application will be briefly introduced below. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 A flowchart illustrating a payment input method based on smart glasses provided in an embodiment of this application; Figure 2 A schematic diagram illustrating an example of a payment page provided in an embodiment of this application; Figure 3 A schematic diagram illustrating an example of a payment input process based on smart glasses provided in an embodiment of this application; Figure 4 A schematic diagram of the structure of a payment input device based on smart glasses provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of smart glasses provided in an embodiment of this application. Detailed Implementation

[0012] The features and exemplary embodiments of various aspects of this application will be described in detail below. To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain this application and not to limit it. For those skilled in the art, this application can be implemented without some of these specific details. The following description of the embodiments is merely to provide a better understanding of this application by illustrating examples. It should be noted that the acquisition, storage, use, and processing of information and data in the embodiments of this application are all authorized by users or relevant organizations and comply with the relevant provisions of national laws and regulations.

[0013] With the continuous development of smart terminal technology, the application fields of smart wearable devices such as smart glasses are becoming increasingly widespread. For example, smart glasses can be used for offline payments. Smart glasses can scan a merchant's QR code, transmitting the information obtained from the scanned QR code to a terminal device or cloud backend bound to the smart glasses. The terminal device or cloud backend then calls the user's payment application and connects with the merchant's payment system through a backend interface to complete the amount confirmation and deduction. The payment page displayed on the front end is often developed independently by the payment page developer. The distribution of view elements in payment pages developed by different payment page developers may differ. The backend interface can only process preset, structured options, resulting in insufficient adaptability of the payment architecture based on smart glasses. It relies entirely on a fixed backend interface to achieve a closed transaction loop and cannot be compatible with payment pages developed by different payment page developers. The operation of entering payment information during the payment process requires a terminal device bound to or connected to the smart glasses. If the user's hands are occupied and it is inconvenient to operate, payment cannot be completed, making the user's requirements for payment operations based on smart glasses relatively high and complex. Furthermore, the smart glasses themselves lack the ability to interact with the payment page and cannot trigger the click operation of the keyboard keys on the payment page. Without the assistance of an external terminal device, the payment process cannot be completed.

[0014] This application provides a payment input method, device, smart glasses, medium, and program product based on smart glasses. It breaks through the limitations of traditional smart glasses, which rely on fixed backend interfaces, cannot adapt to diverse front-end payment pages, and only support system-level virtual keyboards for voice input. It can adaptively recognize images on payment pages and non-intrusively simulate key clicks on payment page keyboards. It enables payment input based on smart glasses without the need for external terminal devices. When the user's hands are occupied and inconvenient to operate, payment can be completed through smart glasses, reducing the requirements for payment operations based on smart glasses, making payment operations based on smart glasses more convenient, and adapting to diverse front-end payment pages, compatible with multiple payment pages.

[0015] The following describes the payment input method, device, smart glasses, medium, and program products based on smart glasses provided in this application.

[0016] This application provides a payment input method based on smart glasses, applicable to scenarios where users wear and use smart glasses for payments. This payment input method based on smart glasses can be executed by the smart glasses themselves. Figure 1 A flowchart of a payment input method based on smart glasses provided in an embodiment of this application is shown below. Figure 1 As shown, the payment input method based on smart glasses may include steps S101 to S104.

[0017] In step S101, the payment page image is acquired, parsed, and interacted with to obtain the image position information of the keyboard keys in the payment page image.

[0018] Smart glasses can capture images of the payment page displayed on the glasses themselves, as well as images of payment pages displayed on other devices. The payment page image is an image of the payment page itself, which may include view elements such as payment information input boxes and a keyboard. Payment information input boxes may include, but are not limited to, input boxes for payment amounts. The keyboard may include multiple keyboard keys; for example, it may include, but is not limited to, number keys, symbol keys, confirmation keys, and delete keys. Figure 2 A schematic diagram illustrating an example of a payment page provided in an embodiment of this application, as shown below. Figure 2 As shown, the payment page 20 includes a payment amount input box 21 and a keyboard 22. The keyboard 22 may include number keys "0" to "9", a symbol key "." (i.e., decimal point), a confirmation key "Confirm Payment", and a delete key "←". The payment page can be triggered by scanning a payment code with smart glasses, but is not limited to this.

[0019] The smart glasses can interact with the payment page to obtain its information. This interaction may include interaction between the smart glasses and the front-end payment page, and / or, interaction between the smart glasses and the back-end of the payment page, but is not limited thereto. In some examples, the smart glasses can obtain the payment page's information by sending a request to the payment page; this information may include the positional information of view elements within the payment page. The page information of the payment page may include the Document Object Model (DOM). For example, the DOM information may include the page information of the payment information input box "class="amount-input", style="top:100px;left:50px;width:200px;height:40px"" and the page information of the number keys "data-type="digit-key", left:115px;top:175px, width:40px;height:50px, data-value="2"". Here, class="amount-input" and data-type="digit-key" are the type information of the visual elements, style="top:100px;left:50px;width:200px;height:40px", left:115px;top:175px, width:40px;height:50px" are the position information of the visual elements, data-value="2" is the character information of the visual elements, top and left can represent coordinates, width and height can represent size, and data-value represents the number on the number keys.

[0020] Smart glasses can analyze payment page images, identify the keyboard and its keys, and combine this with page information to determine the image position information of the keyboard keys within the payment page image. This image position information includes the position of the keyboard keys within the payment page image, representing their location. For example, if the payment page looks like... Figure 2 As shown, the obtained keyboard key image position information can include the position information of each of the number keys "0" to "9" in the payment page image, the position information of the "." key, the position information of the "Confirm Payment" key, and the position information of the "←" delete key. In some examples, the payment page image can be converted to grayscale to reduce the amount of image data to be processed and improve processing efficiency.

[0021] In step S102, the page position information of the keyboard keys corresponding to the image position information of the keyboard keys is obtained according to the mapping relationship between the image position in the payment page image and the page position in the payment page.

[0022] The image position in the payment page image refers to the location of a specific object within the payment page image. The page position within the payment page refers to the location of a specific object within the payment page. Because smart glasses are worn by the user, the payment page image captured by the smart glasses is affected by factors such as lighting and user actions, which may cause the image position of the same view element in the payment page image to differ from its actual page position on the payment page. A mapping relationship between the image position in the payment page image and the actual page position on the payment page can be established when the smart glasses capture the payment page image. In some examples, when obtaining the payment page image, the image position of a reference view element can be identified from the payment page image, and the page position of the reference view element can be obtained from the payment page. Based on the image position and page position of the reference view element, a mapping relationship between the image position in the payment page image and the actual page position on the payment page can be established. The reference view element is the view element used as the mapping reference. Any fixed view element from the payment page and the payment page image can be selected as the reference view element. For example, the payment page... Figure 2 As shown, the "Confirm Payment" button can be selected as the reference view element. A mapping relationship is established between the image position in the payment page image and the page position in the payment page image by the deviation between the image position of the reference view element in the payment page image and its page position in the payment page. After obtaining this mapping relationship, the page position information of any pixel in the payment page image can be obtained based on its image position information. Correspondingly, after obtaining the image position information of the keyboard keys, the page position information of the keyboard keys can be obtained using the mapping relationship. Even if user actions cause changes in the image of the payment page in the smart glasses during subsequent use, the change parameters between the payment page image and the initially acquired payment page image can be adjusted to ensure that the image position information of the keyboard keys can be accurately mapped to the page position information of the keyboard keys, avoiding page position misalignment that could lead to failed simulated clicks in subsequent processes. The change parameters can characterize changes in the payment page image; for example, transformation parameters may include, but are not limited to, page image scaling factor and page image rotation angle. The change parameters can be represented in matrix form, but are not limited to this. The image position information and page position information can be implemented as coordinates, size, or other information that can characterize position, and are not limited here. Step S102 can obtain the page position information corresponding to each keyboard key in the payment page image.

[0023] In step S103, the user's payment voice command is received, and the page position information of the keyboard keys corresponding to the payment information converted from the payment voice command is determined.

[0024] Users can issue commands to the smart glasses via voice. The smart glasses receive the user's payment voice command, which is a voice instruction from the user to make a payment. The smart glasses can convert the payment voice command into payment information, which can be in text format, and is not limited thereto. The payment information is the information that the user wants to input into the payment information input box on the payment page. In some examples, the payment information includes the payment amount, and the characters corresponding to the payment amount converted from the payment voice command can be numbers. The payment information converted from the payment voice command includes characters. From the page position information of each keyboard key obtained in step S102, the page position information of the keyboard key corresponding to the character in the payment information is obtained as the page position information of the keyboard key corresponding to the payment information. For example, if the user inputs the payment voice command "pay 25 yuan", the payment information converted from the payment voice command includes the payment amount "25", "25" includes two characters, the two characters are the numbers "2" and "5", and the page position information of the keyboard key corresponding to the character information includes the page position information of the number key "2" and the page position information of the number key "5".

[0025] In this embodiment of the application, the image location information, page location information, etc. that need to be stored can be stored using an 8-bit unsigned integer, i.e., uint8 data type. Compared with the storage of 32-bit floating-point numbers, i.e., float32 data type, the memory usage can be reduced by 75%.

[0026] In step S104, a simulated click event instruction is generated and sent to the payment page to simulate clicking the keyboard keys on the payment page and complete the input of payment information.

[0027] Simulated click event commands can be generated based on the page position information of the keyboard keys corresponding to the payment information. These commands include the page position information of the keyboard keys corresponding to the payment information and can be used to simulate clicks on these keys on the payment page. The payment page responds to the simulated click event commands by simulating a click at the position indicated by the page position information in the command, thereby allowing payment information to be entered into the payment information input box. In some examples, smart glasses can establish a temporary connection with the payment page via WebSocket, sending simulated click event commands to simulate the user's manual click interaction logic. The smart glasses interact with the payment page via WebSocket without modifying the original code of the payment page; simulated clicks between the smart glasses and the payment page can be achieved through the native Application Programming Interface (API) of the browser or other applications within the smart glasses.

[0028] In some examples, if the payment information includes multiple characters, the input of the payment information can be achieved through a simulated click event instruction. This simulated click event instruction can include the page position information of the keyboard keys corresponding to all characters in the payment message, and the page position information is arranged according to the order of the characters in the payment message, so that the payment page can respond to the simulated click event instruction, simulate clicking the keyboard keys corresponding to the characters in the order of the characters, and input the correct payment information on the payment page.

[0029] In other examples, if the payment information includes multiple characters, the input of the payment information can be achieved through multiple simulated click event instructions. Each simulated click event instruction can include the page position information of the keyboard key corresponding to one character in the payment message. Multiple simulated click event instructions are sent to the payment page in the order of the characters in the payment message, so that the payment page can respond to the simulated click event instructions, simulate clicking the keyboard key corresponding to the character in the order of the characters, and input the correct payment information on the payment page.

[0030] In this embodiment, the smart glasses can capture and parse the payment page image, and interact with the payment page to obtain the image position information of the keyboard keys in the payment page image. Based on the mapping relationship between the image position in the payment page image and the page position on the payment page, the page position information of the keyboard keys on the payment page is obtained. The user's payment voice command can be converted into payment information, and the page position information of the corresponding keyboard key can be determined. A simulated click event command is generated based on the page position information of the keyboard key corresponding to the payment information and sent to the payment page. The payment page responds to the simulated click event command by performing a simulated click on the corresponding keyboard key on the payment page, thereby inputting payment information on the payment page. This payment input process does not require the user to operate any external or bound terminal device on the smart glasses. Even if the user's hands are occupied, simply giving a payment voice command can trigger the simulated click of the keyboard keys on the payment page to complete the payment process, thereby reducing the requirements for payment operations based on smart glasses and making payment operations based on smart glasses more convenient. The adaptive recognition of the payment page image also makes the payment input method based on smart glasses in this embodiment adaptable to diverse front-end payment pages, exhibiting stronger compatibility.

[0031] Through testing, the payment input method based on smart glasses provided in this application embodiment can control the total latency of voice input payment information to within 500 milliseconds, approaching that of manual input payment. Furthermore, the error rate of input payment information has been reduced from over 15% to below 3%, avoiding repeated corrections and significantly improving payment efficiency and accuracy. This application embodiment can overcome the limitations of smart glasses adapting to front-end payment pages such as JS (i.e., JavaScript) payment pages, realizing a closed-loop process of scanning, voice input of amount, and simulated click confirmation. This fills the gap in flexible offline payment for smart glasses, enabling voice payment for smart glasses to cover, but not limited to, payment pages developed using Vue, React, and JS languages, such as over 95% of mainstream offline payment pages.

[0032] In some embodiments, the image position information of the keyboard keys in the payment page image can be obtained by comprehensively considering three aspects: the structural features of the keyboard keys, the visual features of the keyboard keys, and the page element features of the keyboard keys, thereby improving the accuracy of keyboard key recognition. The smart glasses can identify the keyboard area through a grid structure based on the payment page image and obtain the first position information of the keyboard keys; based on the image in the keyboard area, distinguish the keyboard keys from the keyboard background to obtain the second position information of the keyboard keys; obtain the third position information of the keyboard keys in the document object model of the payment page; and obtain the image position information of the keyboard keys in the payment page image based on the first, second, and third position information.

[0033] In some embodiments, to reduce the computational load and resource consumption of smart glasses, the image containing the predicted keyboard area can be cropped from the payment page image. Further image processing is then performed on the cropped image area to obtain the image position information of the keyboard keys. The smart glasses can acquire the position information of the payment information input box in the payment page image; based on the positional relationship between the payment information input box and the keyboard, a region to be processed is cropped from the payment page image; for the region to be processed, the keyboard area is identified through a grid structure, and the first position information of the keyboard keys is obtained. In the payment page, the payment information input box and the keyboard generally have a relatively fixed positional relationship. After obtaining the position information of the payment information input box, the region where the keyboard is located can be predicted based on the positional relationship between the payment information input box and the keyboard. The predicted keyboard region is then cropped from the payment page image, and subsequent image processing is performed on the cropped image region. In the above embodiments, the image region to be processed is the predicted keyboard region cropped from the payment page image.

[0034] In some examples, smart glasses can interact with the payment page, requesting the page location information of the payment information input box. The smart glasses can also locate the payment information input box in the payment page image using visual edge detection technology, obtaining its image location information. The position information of the payment information input box can be obtained based on both its page location information and its image location information. The smart glasses can request the page location information of the payment information input box in the DOM from the payment page. The payment information input box is usually located at the top or center of the payment page, making it easily identifiable using visual edge detection. Visual edge detection can utilize the Canny operator algorithm to identify the rectangular outline of the payment information input box. The edge gradient threshold in the Canny operator algorithm can be set from 50 to 150, but is not limited to this. The system can obtain the deviation between the page position information and the image position information of the payment information input box. If the deviation is within the standard threshold range, either the page position information or the image position information can be determined as the position information of the payment information input box; alternatively, the average of the page position information and the image position information can be used as the position information. If the deviation exceeds the standard threshold range, either the page position information or the image position information can be determined as the position information, depending on the actual situation. This dual confirmation of page position information and image position information improves the accuracy of the payment information input box's positioning.

[0035] Payment information input boxes are typically located at the top or center of the payment page, with the keyboard usually positioned below them. The area to be processed can be cropped below the payment information input box according to a preset keyboard size, assuming that the image area to be processed includes the keyboard area. For example, the area to be processed can be defined as the region extending downwards from the bottom center of the payment information input box in the payment page image, and then extending to the left and right by a second distance. The first distance can be the maximum height covering the keyboard, and twice the second distance can be the maximum width covering the keyboard. Recognizing the keyboard area and determining the first position information of the keyboard keys within the image area to be processed can significantly reduce the computational resources consumed in these steps. For example, if the payment page image is 1280×720 pixels (approximately 921,600 pixels), and the cropped image area is 325×300 pixels (approximately 97,500 pixels), the size of the image area to be processed is 10.6% of the payment page image size, significantly reducing the amount of data required for subsequent image processing. In some examples, the image region to be processed can also be grayscaled and Gaussian blurred. The convolution kernel for Gaussian blurring can be 3×3, but is not limited to this. Grayscaled and Gaussian blurred processing can further reduce the amount of data that needs to be processed and reduce noise in the image.

[0036] The keyboard on the payment page is generally a numeric keypad, which typically follows a 3-column x 4-row grid layout. Special numeric keypads may have one more row, one more column, one less row, or one less column than the standard 3-column x 4-row grid layout. The keyboard area and the keys within that area can be determined by the characteristics of the keyboard's grid structure, thus obtaining the first position information. The keyboard area refers to the region where the keyboard is located. The first position information includes the position information of the keyboard keys identified through the grid structure. This first position information may include coordinates, size, and other information that characterizes position, but is not limited here.

[0037] In some examples, the keyboard area and keyboard keys can be determined by the complete contour detection method. This can find the locations in the payment page image where the pixel values ​​change drastically, thereby obtaining the entire outer border of the keyboard and the dividing lines of the keyboard keys. Based on the entire outer border of the keyboard and the dividing lines of the keyboard keys, the keyboard area and keyboard keys are determined, and thus the first position information of the keyboard keys is obtained.

[0038] In other examples, the dividing lines between keyboard keys can be detected using Hough transform based on the payment page image; the keyboard area and the first position information of the keyboard keys can then be obtained from these dividing lines. Hough transform detection can be used to detect lines, but the outer border of the keyboard area is not detected; instead, the dividing lines between the keyboard keys within the keyboard are detected. The keyboard area and the keyboard keys are determined by these dividing lines. For example, for a 3-column × 4-row keyboard, this example only needs to detect the 3 dividing lines separating the 4 rows of keys and the 2 dividing lines separating the 3 columns, without needing to detect the 4 outer border lines. This reduces the number of lines detected by 50%, significantly reducing the resource consumption for keyboard key recognition. The position of the keyboard keys can be determined based on the dividing lines between them, thus obtaining the first position information. For example, the center coordinates of the squares formed by the dividing lines (i.e., the keyboard keys) can be used as the first position information of the keyboard keys. In Hough transform detection, a length threshold is needed to detect lines, filtering out excessively short, discontinuous, or indistinct lines. In this example, a more lenient length threshold can be set, which can be lower than the standard length threshold in Hough transform detection. A lenient length threshold reduces the filtering of weak lines by Hough transform detection, thereby reducing the number of computation iterations. For example, the commonly used standard length threshold for Hough transform detection is 50px, so in this example, the length threshold for Hough transform detection is set to 40px, meaning that Hough transform detection will obtain lines with a length greater than or equal to 40px.

[0039] In some embodiments, there is a significant visual difference between the keyboard keys and the keyboard background. For example, the keyboard keys are light-colored and the keyboard background is dark-colored, or vice versa. The keyboard keys can be located through visual contrast. The smart glasses can perform Otsu adaptive threshold segmentation on the image in the keyboard area to obtain the keyboard keys and the keyboard background; use morphological opening operations to remove noise from the edges of the keyboard keys to obtain their outlines; and determine the second position information of the keyboard keys based on their outlines. The second position information includes the position information of the keyboard keys identified through the principle of visual difference. The second position information may include coordinates, size, and other information that characterizes position, but is not limited thereto.

[0040] Otsu adaptive thresholding segmentation calculates a threshold that distinguishes keyboard keys from the keyboard background by maximizing the inter-class variance between the foreground and background classes in the image. In this embodiment, the foreground is the keyboard keys, and the background is the keyboard background. This threshold is determined based on the grayscale distribution of the payment page image itself, and it can automatically distinguish between keyboard keys and the keyboard background. After distinguishing the keyboard keys, morphological opening operations can be used to remove the edges of the keyboard keys, thereby preserving the main outline of the keyboard keys. The convolution kernel for the morphological opening operation can be 3×3, but is not limited to this. For each keyboard key, the position information of a point on the keyboard key can be used as the second position information of the keyboard key. For example, the position information of the center of the keyboard key can be used as the second position information of the keyboard key.

[0041] In some examples, excessively strong or weak ambient light may cause Otsu adaptive thresholding segmentation to fail. For instance, if the grayscale difference between the keyboard keys and the keyboard background is less than 30, Otsu adaptive thresholding segmentation is highly likely to fail. In this case, color features can be used to assist in distinguishing between the keyboard keys and the keyboard background. When Otsu adaptive thresholding segmentation fails, the smart glasses can distinguish between the keyboard keys and the keyboard background in the keyboard area based on the dominant color values ​​of the keyboard keys, obtaining the second position information of the keyboard keys. The RGB dominant color values ​​of the keyboard keys can be extracted, and the similarity between the color values ​​of pixels in the keyboard area and the RGB dominant color values ​​of the keyboard keys can be calculated to distinguish between the keyboard keys and the keyboard background.

[0042] The third location information includes the keyboard key position information obtained from the DOM of the payment page. This third location information may include coordinates, size, and other information that characterizes location, and is not limited here. The smart glasses can send a page information retrieval request to the payment page to obtain the keyboard key position information in the DOM, i.e., the third location information. This is particularly useful in situations where the keyboard may dynamically load but not be displayed on the payment page, or where the payment page is partially obscured; the third location information can supplement the visual recognition limitations of the second location information. If the second location information of a certain keyboard key is not obtained, it can be inferred from the third location information of its neighboring keys and the positional relationship between the keyboard key and its neighbors.

[0043] In some embodiments, for a keyboard key, the average value of the first position information, the second position information, and the third position information can be determined as the image position information of the keyboard key. Alternatively, weights can be pre-set for the first position information, the second position information, and the third position information, and the image position information of the keyboard key can be calculated using a weighted algorithm based on the first position information, the second position information, and the third position information.

[0044] In other embodiments, to improve the obtained image position information of the keyboard keys, different calculation methods can be used to obtain the image position information of the keyboard keys, depending on whether the first, second, and third position information are valid or invalid. Invalid information is missing position information among the first, second, and third position information, or position information whose absolute value of the deviation from two position information other than itself is greater than an error standard threshold. Conversely, if a position information is not invalid, then that position information is valid. For example, if the value of the second position information is missing among the first, second, and third position information, then the first and third position information are valid information, and the second position information is invalid information. The absolute value of the deviation between any two position information is the absolute value of the difference between the two position information. The error standard threshold can be set according to the scenario, requirements, experience, etc., and is not limited here. For example, if the error standard threshold is 5px, and the absolute value of the deviation between the first position information and the second position information is greater than 5px, the absolute value of the deviation between the second position information and the third position information is greater than 5px, and the absolute value of the deviation between the first position information and the third position information is less than or equal to 5px, then the second position information is invalid information, and the first position information and the third position information are valid information.

[0045] If the first, second, and third position information are all valid, the image position information of the keyboard keys is calculated based on the first, second, and third position information, their corresponding weights, and the image position information of the keyboard keys. That is, a weighted average can be calculated using a weighted algorithm based on the first, second, and third position information and their corresponding weights, and this weighted average is determined as the image position information of the keyboard keys. The sum of the weights corresponding to the first, second, and third position information is 1. The specific weights corresponding to the first, second, and third position information can be set according to the scenario, requirements, and experience, and are not limited here. For example, the weight corresponding to the first position information could be 40%, the weight corresponding to the second position information could be 35%, and the weight corresponding to the third position information could be 25%.

[0046] If invalid information exists in the first, second, and third position information, the image position information of the keyboard keys is calculated based on the other position information (excluding invalid information) and their corresponding weights. That is, a weighted average algorithm can be used to calculate the weighted average of the valid information in the first, second, and third position information and its weights. This weighted average is then used as the image position information of the keyboard keys, where the sum of the weights of the valid information is 1. For example, if the third position information is missing, the image position information of the keyboard keys can be calculated using a weighted algorithm based on the first and second position information and their respective weights. The weight of the first position information is 40%, and the weight of the second position information is 60%.

[0047] Each keyboard key corresponds to a character on it, and the image position information of the keyboard keys can also be understood as the image position information corresponding to the characters on the keyboard keys. The obtained image position information of the keyboard keys can be implemented as a correspondence between the characters on the keyboard keys and the image position information. For example, the obtained image position information of the keyboard keys may include: {"2":{"x":118,"y":178},"5":{"x":158,"y":178},…}; where "2":{"x":118,"y":178} represents the image position coordinates of the number key "2" as (118,178), and "5":{"x":158,"y":178} represents the image position coordinates of the number key "5" as (158,178).

[0048] In some embodiments, payment information includes characters, and keyboard keys also contain characters. After processing the payment page image to obtain the page position information of all keyboard keys in the payment page image, a target character matching the payment information can be found among the characters corresponding to the keyboard keys. The page position information of the keyboard key corresponding to the target character is then determined as the page position information of the keyboard key corresponding to the payment information. For example, the payment page may look like this: Figure 2As shown, the user's voice payment command is "Pay 36 yuan". The payment information includes the characters "3" and "6". Correspondingly, the numbers "3" and "6" can be selected from the numbers "0" to "9" on the numeric keys as target characters. The page position information of the keyboard keys corresponding to the number "3" and "6" is determined as the page position information of the keyboard keys corresponding to the payment information. The characters corresponding to the keyboard keys include numbers, which can be identified by the number of closed contours and the stroke direction. The number of closed contours refers to the number of closed contours; for example, the number "8" contains two closed contours, the number "9" contains one closed contour, and the number "7" contains zero closed contours. The stroke direction represents the direction of a stroke in the number; for example, the straight line in the number "7" is a right-slanted line. The number of closed contours and the stroke direction are the core features of numbers. Identifying numbers through the number of closed contours and the stroke direction simplifies number feature extraction, ignoring information such as font style, shadows, and animation effects, thus reducing the computational complexity of number recognition. Moreover, by recognizing numbers through the number of closed contours and the direction of strokes, the character matching method without load has a single number feature extraction time of less than or equal to 3 milliseconds. Compared with the traditional character template matching method (i.e., comparing with 10 number templates to determine the number, with each number template comparison taking about 10 milliseconds), the number recognition in this embodiment can be accelerated by 70%, thereby improving the efficiency of payment input based on smart glasses.

[0049] By cropping the payment page image, using Hough transform to detect and identify keyboard keys, setting a relaxed threshold in Hough transform detection, identifying numbers by the number of closed contours and stroke direction, using a lightweight interface to call algorithms for image processing, and using 8-bit unsigned integers to store position information, the occupancy rate of the central rationalizer used for image recognition in smart glasses can be reduced to 15% to 20%, effectively reducing lag issues. The power consumption of a single payment can be reduced to 1 / 4 of the traditional solution, and the impact of 10 payments per day on the battery life of smart glasses is shortened to within 30 minutes, which is compatible with the hardware conditions of smart glasses.

[0050] In some embodiments, during the simulated click process on the payment page, attention can be paid to timing verification of the simulated clicks to improve accuracy. Specifically, if the characters corresponding to the recognized keyboard keys do not contain any characters from the payment information, the smart glasses perform a waiting action until the characters corresponding to the recognized keyboard keys contain every character from the payment information, then generate a simulated click event command and send it to the payment page. Due to page loading delays, the positions of the keyboard keys in the payment page image may be temporarily unclear, potentially resulting in situations where the characters corresponding to the recognized keyboard keys do not contain any characters from the payment information. In such cases, a waiting action can be performed, i.e., pausing the simulated click, until the recognition stabilizes and the keyboard keys corresponding to the characters in the payment information are recognized, before generating a simulated click event command and sending it to the payment page. This avoids out-of-order input and improves the accuracy of the simulated clicks.

[0051] In some embodiments, semantic verification of the simulated clicks on the payment page can be performed to improve the accuracy of the simulated clicks. Specifically, the smart glasses can acquire characters from the payment information input box in the payment page image; if the characters in the payment information input box are inconsistent with the characters in the payment information corresponding to the simulated click, a simulated click of the correct keyboard key is triggered to correct the characters in the payment information input box. After each simulated click, the characters in the payment information input box in the payment page image can be acquired as feedback information and compared with the characters of the simulated click in the payment information converted from the payment voice command; if they match, it means that the correct keyboard key was simulated; if they do not match, a simulated click of the delete key can be automatically triggered, and the keyboard key corresponding to the character of the simulated click in the payment information can be simulated again. That is, a simulated click event command containing the page position information of the character of the simulated click in the payment information is generated again and sent to the payment page to complete the correction of the simulated keyboard key click, thereby updating the incorrect characters in the payment information input box. For example, if the payment information is "23", and the first simulated click returns "2" to the payment information input box, then the character in the input box matches the character in the payment information corresponding to the simulated click. A second simulated click is then performed. If the second simulated click returns "4", then the character in the input box does not match the character in the payment information corresponding to the simulated click, and the delete button is automatically triggered, deleting the "4" entered in the second simulated click. The simulation is then repeated. If the second simulated click returns "3", then the payment input ends. By triggering corrections to simulated keyboard clicks, the accuracy of simulated clicks can be further improved.

[0052] By performing dual verification of the timing and semantics of simulated clicks, the final error rate of payment information input can be controlled to below 1%, which is far lower than the 20% error rate of traditional solutions. This effectively reduces the misoperation rate of payment input based on smart glasses and ensures the accuracy of payments.

[0053] In the above embodiments, the operations related to payment voice command parsing, keyboard key recognition, timing verification of simulated clicks, and semantic verification can be implemented by calling the algorithm model through lightweight interfaces such as the OpenCV interface, avoiding the need to call deep learning models. Compared with the traditional solution that requires loading a deep learning model of more than 50MB when calling the YOLO interface, the inference time of the embodiments of this application can be significantly shortened.

[0054] To facilitate understanding, the following will be combined with... Figure 3 To illustrate the logic of the payment input method based on smart glasses in the embodiments of this application, Figure 3 A schematic diagram illustrating an example of a payment input process based on smart glasses provided in this application embodiment, as shown below. Figure 3 As shown, smart glasses 301 scan the payment code 302 to obtain a payment page image 303. The user can issue a payment voice command 304, which is then processed by the voice recognition function 305 to obtain text 306. The voice recognition function 305 can then automatically recognize speech. The text 306 is processed by semantic analysis function 307 to obtain payment information 308. Keyboard parsing 309 is performed on the payment information 308 to determine the keyboard key corresponding to the payment information 308. The payment page image 303 is preprocessed by image preprocessing 310 to obtain a more easily processed payment page image 303. Then, through the image recognition big model 311, keyboard key positioning 312 and character recognition 313 corresponding to the keyboard key can be obtained. The image recognition big model 311 can realize the recognition of the keyboard area and keyboard key in the above embodiment, which will not be elaborated here. The image recognition big model 311 can be called through a lightweight interface such as the OpenCV interface to obtain the keyboard key positioning 312 and the character 313 of the keyboard key. The input of payment information 308 on the payment page is realized by combining the mapping relationship between image position and page position 314 and the simulated click event instruction 315.

[0055] In this application embodiment, the parsing of payment voice commands, the dynamic layout recognition of keyboard keys, the timing verification of simulated clicks, and the semantic verification of simulated clicks can be considered as Artificial Intelligence (AI) foundational capabilities; the cross-end communication between smart glasses and the payment page (i.e., smart glasses connecting to the payment page via WebSocket), cropping the image region to be processed, the mapping between image position information and page position information, the execution of simulated clicks, the timing control and interactive feedback during the payment input process, etc., belong to the edge-side foundational capabilities of smart glasses.

[0056] This application also provides a payment input device based on smart glasses, which is applied to smart glasses and corresponds to the payment input method based on smart glasses in the above embodiments. Figure 4 This is a schematic diagram of the structure of a payment input device based on smart glasses provided in an embodiment of this application, as shown below. Figure 4 As shown, the payment input device 400 based on smart glasses may include an image processing module 401, a voice processing module 402, and a simulated click module 403.

[0057] The image processing module 401 can be used to acquire payment page images, parse payment page images, and interact with the payment page to obtain image position information of keyboard keys in the payment page images; and, based on the mapping relationship between image positions in the payment page images and page positions in the payment page, to obtain page position information of the keyboard keys corresponding to the image position information of the keyboard keys.

[0058] The voice processing module 402 can be used to receive the user's payment voice command and determine the page position information of the keyboard key corresponding to the payment information converted from the payment voice command.

[0059] The simulated click module 403 can be used to generate simulated click event instructions and send them to the payment page to simulate clicking the keyboard keys on the payment page to complete the input of payment information. The simulated click event instructions include the page position information of the keyboard keys corresponding to the payment information.

[0060] In some embodiments, the image processing module 401 may be used to: identify a keyboard area through a grid structure based on a payment page image, and obtain first position information of the keyboard keys; distinguish between keyboard keys and keyboard background in the keyboard area based on the image in the keyboard area, and obtain second position information of the keyboard keys; obtain third position information of the keyboard keys in the document object model of the payment page from the payment page; and obtain image position information of the keyboard keys in the payment page image based on the first position information, the second position information, and the third position information.

[0061] In some examples, the image processing module 401 can be used to: obtain the position information of the payment information input box in the payment page image; crop the image area to be processed in the payment page image according to the position relationship between the payment information input box and the keyboard; and identify the keyboard area through a grid structure for the image area to be processed, and obtain the first position information of the keyboard keys.

[0062] In some examples, the image processing module 401 can be used to: detect the dividing lines of keyboard keys using Hough transform based on the payment page image; and obtain the keyboard area and the first position information of the keyboard keys based on the dividing lines of the keyboard keys.

[0063] In some examples, the image processing module 401 can be used to: perform Otsu adaptive thresholding on the image in the keyboard area to obtain keyboard keys and keyboard background; remove noise from the edges of the keyboard keys using morphological opening operations to obtain the outline of the keyboard keys; and determine the second position information of the keyboard keys based on the outline of the keyboard keys.

[0064] In some examples, the image processing module 401 can be used to: in the event that Otsu adaptive thresholding fails, distinguish between keyboard keys and keyboard background in the keyboard area based on the primary color value of the keyboard keys, and obtain the second position information of the keyboard keys.

[0065] In some examples, the image processing module 401 can be used to: if the first position information, the second position information, and the third position information are all valid information, then calculate the image position information of the keyboard keys based on the first position information, the second position information, the third position information, the weight corresponding to the first position information, the weight corresponding to the second position information, and the weight corresponding to the third position information; if there is invalid information among the first position information, the second position information, and the third position information, then calculate the image position information of the keyboard keys based on the other position information among the first position information, the second position information, and the third position information excluding the invalid information, and the weight corresponding to the other information; wherein, the invalid information is the missing position information among the first position information, the second position information, and the third position information, or the position information among the first position information, the second position information, and the third position information whose absolute value of the deviation from the two position information other than itself is greater than the error standard threshold.

[0066] In some embodiments, the image processing module 401 can also be used to: when obtaining the payment page image, identify and determine the image position of the reference view element from the payment page image, and obtain the page position of the reference view element from the payment page; and establish a mapping relationship between the image position in the payment page image and the page position in the payment page based on the image position of the reference view element and the page position of the reference view element.

[0067] In some embodiments, the voice processing module 402 can be used to: search for a target character that matches the payment information among the characters corresponding to the keyboard keys; determine the page position information of the keyboard key corresponding to the target character as the page position information of the keyboard key corresponding to the payment information; the characters corresponding to the keyboard keys include numbers, which are identified by the number of closed contours and the direction of the strokes.

[0068] In some embodiments, the simulated click module 403 can be used to: if the characters corresponding to the identified keyboard keys do not contain any characters from the payment information, perform a waiting action until the characters corresponding to the identified keyboard keys contain each character from the payment information, generate a simulated click event instruction, and send it to the payment page.

[0069] In some embodiments, the simulated click module 403 can also be used to: obtain characters in the payment information input box in the payment page image; and, if the characters in the payment information input box are inconsistent with the characters in the payment information corresponding to the simulated click, trigger a simulated click of the correction keyboard keys to correct the characters in the payment information input box.

[0070] It should be noted that the payment input device 400 based on smart glasses is a device corresponding to the payment input method based on smart glasses described above. All implementation methods in the above method embodiments are applicable to the embodiments of this device and can achieve the same technical effect.

[0071] This application also provides a smart glasses. Figure 5 This is a schematic diagram of the structure of smart glasses provided in an embodiment of this application, as shown below. Figure 5 As shown, the smart glasses 500 includes a memory 501, a processor 502, and a computer program stored in the memory 501 and capable of running on the processor 502.

[0072] In some examples, the processor 502 described above may include a central processing unit (CPU), or an application-specific integrated circuit (ASIC), or one or more integrated circuits that may be configured to implement the embodiments of this application.

[0073] Memory 501 may include read-only memory (ROM), random access memory (RAM), disk storage media device, optical storage media device, flash memory device, electrical, optical, or other physical / tangible memory storage device. Therefore, typically, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the payment input method based on smart glasses according to embodiments of this application.

[0074] The processor 502 runs a computer program corresponding to the executable program code by reading the executable program code stored in the memory 501, so as to implement the payment input method based on smart glasses in the above embodiment.

[0075] In some examples, the smart glasses 500 may also include a communication interface 503 and a bus 504. For example, Figure 5 As shown, the memory 501, processor 502, and communication interface 503 are connected through bus 504 and complete communication with each other.

[0076] The communication interface 503 is mainly used to enable communication between various modules, devices, units, and / or equipment in the embodiments of this application. Input devices and / or output devices can also be connected through the communication interface 503.

[0077] Bus 504 includes hardware, software, or both, that couples the components of smart glasses 500 together. For example, and not limitingly, bus 504 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a Hyper Transport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-E) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local Bus (VLB) bus, or other suitable buses, or a combination of two or more of these. Where appropriate, bus 504 may include one or more buses. Although specific buses are described and illustrated in the embodiments of this application, this application considers any suitable bus or interconnection.

[0078] This application also provides a computer-readable storage medium storing computer program instructions. When these computer program instructions are executed by a processor, they can implement the payment input method based on smart glasses in the above embodiments and achieve the same technical effect. To avoid repetition, further details are omitted here. The aforementioned computer-readable storage medium may include non-transitory computer-readable storage media, such as read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks, etc., and is not limited thereto.

[0079] This application also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the payment input method based on smart glasses in the above embodiments and can achieve the same technical effect. To avoid repetition, it will not be described again here.

[0080] It should be clarified that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. For the device embodiments, smart glasses embodiments, computer-readable storage medium embodiments, and computer program product embodiments, the relevant parts can be referred to the description section of the method embodiments. This application is not limited to the specific steps and structures described above and shown in the figures. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application. Furthermore, for the sake of brevity, detailed descriptions of known methods and techniques are omitted here.

[0081] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0082] Those skilled in the art will understand that the above embodiments are exemplary and not restrictive. Different technical features appearing in different embodiments can be combined to achieve beneficial effects. Based on a study of the drawings, specification, and claims, those skilled in the art should be able to understand and implement other variations of the disclosed embodiments. In the claims, the term "comprising" does not exclude other means or steps; the quantifier "a" does not exclude a plurality; the terms "first" and "second" are used to identify names and not to indicate any particular order. No reference numerals in the claims should be construed as limiting the scope of protection. The functionality of multiple parts appearing in the claims can be implemented by a single hardware or software module. The appearance of certain technical features in different dependent claims does not mean that these technical features cannot be combined to achieve beneficial effects.

Claims

1. A payment input method based on smart glasses, characterized in that, The method, applied to the smart glasses, includes: The payment page image is captured, parsed, and interacted with to obtain the image position information of the keyboard keys in the payment page image; Based on the mapping relationship between the image position in the payment page image and the page position in the payment page, the page position information of the keyboard keys corresponding to the image position information of the keyboard keys is obtained; Receive the user's payment voice command and determine the page position information of the keyboard keys corresponding to the payment information converted from the payment voice command; A simulated click event instruction is generated and sent to the payment page to simulate clicking on keyboard keys on the payment page to complete the input of payment information. The simulated click event instruction includes the page position information of the keyboard keys corresponding to the payment information.

2. The method according to claim 1, characterized in that, The process of parsing the payment page image and interacting with the payment page to obtain the image position information of the keyboard keys in the payment page image includes: Based on the payment page image, the keyboard area is identified through a grid structure, and the first position information of the keyboard keys is obtained. Based on the image in the keyboard area, keyboard keys and keyboard background are distinguished in the keyboard area to obtain the second position information of the keyboard keys; Obtain the third position information of the keyboard keys in the document object model of the payment page; Based on the first location information, the second location information, and the third location information, the image location information of the keyboard keys in the payment page image is obtained.

3. The method according to claim 2, characterized in that, The step of identifying the keyboard area through a grid structure based on the payment page image and obtaining the first position information of the keyboard keys includes: Obtain the position information of the payment information input box in the payment page image; Based on the positional relationship between the payment information input box and the keyboard, the image region to be processed is cropped from the payment page image. For the image region to be processed, the keyboard region is identified through a grid structure, and the first position information of the keyboard keys is obtained.

4. The method according to claim 2, characterized in that, The step of identifying the keyboard area through a grid structure based on the payment page image and obtaining the first position information of the keyboard keys includes: Based on the payment page image, the separator lines of the keyboard keys are detected using Hough transform; The first position information of the keyboard area and the keyboard keys is obtained based on the dividing lines of the keyboard keys.

5. The method according to claim 2, characterized in that, The step of distinguishing keyboard keys and keyboard background based on the image in the keyboard area to obtain second position information of the keyboard keys includes: Otsu adaptive thresholding is performed on the image in the keyboard area to obtain the keyboard keys and keyboard background. Morphological opening operations are used to remove noise from the edges of keyboard keys, thus obtaining the outline of the keyboard keys. The second position information of the keyboard keys is determined based on the outline of the keyboard keys.

6. The method according to claim 5, characterized in that, Also includes: In the event that Otsu adaptive threshold segmentation fails, the keyboard keys and the keyboard background are distinguished in the keyboard area based on the primary color value of the keyboard keys, and the second position information of the keyboard keys is obtained.

7. The method according to claim 2, characterized in that, The step of obtaining the image location information of the keyboard keys in the payment page image based on the first location information, the second location information, and the third location information includes: If the first location information, the second location information, and the third location information are all valid information, then the image location information of the keyboard keys is calculated based on the first location information, the second location information, the third location information, the weight corresponding to the first location information, the weight corresponding to the second location information, and the weight corresponding to the third location information. If there is invalid information in the first location information, the second location information, and the third location information, then the image location information of the keyboard keys is calculated based on the other location information in the first location information, the second location information, and the third location information excluding invalid information, and the weights corresponding to the other information. Invalid information includes missing location information from the first location information, the second location information, and the third location information, or location information from the first location information, the second location information, and the third location information whose absolute deviation from the other two location information (excluding itself) is greater than the error standard threshold.

8. The method according to claim 1, characterized in that, Also includes: When the payment page image is obtained, the image position of the reference view element is identified from the payment page image, and the page position of the reference view element is obtained from the payment page; Based on the image position and page position of the reference view element, a mapping relationship is established between the image position in the payment page image and the page position in the payment page.

9. The method according to claim 1, characterized in that, The step of determining the page position information of the keyboard keys corresponding to the payment information obtained from the payment voice conversion includes: Search for the target character that matches the payment information among the characters corresponding to the keyboard keys; The page position information of the keyboard key corresponding to the target character is determined as the page position information of the keyboard key corresponding to the payment information; The characters corresponding to the keyboard keys include numbers, which are identified by the number of closed contours and the direction of the strokes.

10. The method according to claim 1, characterized in that, The step of generating a simulated click event instruction and sending it to the payment page includes: If the characters corresponding to the identified keyboard keys do not contain any of the characters in the payment information, a waiting action is performed until the characters corresponding to the identified keyboard keys contain each of the characters in the payment information, then the simulated click event instruction is generated and sent to the payment page.

11. The method according to claim 1, characterized in that, Also includes: Obtain the characters in the payment information input box in the payment page image; If the characters in the payment information input box are inconsistent with the characters in the payment information corresponding to the simulated click, a simulated click of the correction keyboard key is triggered to correct the characters in the payment information input box.

12. A payment input device based on smart glasses, characterized in that, The device, applied to the smart glasses, includes: The image processing module is used to acquire payment page images, parse the payment page images, and interact with the payment page to obtain image position information of keyboard keys in the payment page images; and to obtain page position information of keyboard keys corresponding to the image position information of keyboard keys based on the mapping relationship between the image positions in the payment page images and the page positions in the payment page. The voice processing module is used to receive the user's payment voice command and determine the page position information of the keyboard key corresponding to the payment information converted from the payment voice command; The simulated click module is used to generate simulated click event instructions and send them to the payment page to simulate clicking on the keyboard keys on the payment page to complete the input of the payment information. The simulated click event instructions include the page position information of the keyboard keys corresponding to the payment information.

13. A type of smart glasses, characterized in that, include: Processor and memory storing computer program instructions; When the processor executes the computer program instructions, it implements the payment input method based on smart glasses as described in any one of claims 1 to 11.

14. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the payment input method based on smart glasses as described in any one of claims 1 to 11.

15. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the payment input method based on smart glasses as described in any one of claims 1 to 11.