Gesture recognition method and system based on AR glasses interaction

By identifying static and dynamic gestures in AR glasses, using three-dimensional coordinates and similarity judgments to generate gesture operation renderings, the limitations of gesture recognition in existing AR glasses interactions are solved, and the user interaction experience and system stability are improved.

CN120233885APending Publication Date: 2025-07-01GUANGZHOU GUDONG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510392418.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Among the existing AR glasses interaction methods, gesture recognition technology is mainly limited to simple gesture movements, which is difficult to cope with complex and changeable interaction needs, resulting in poor user interaction experience.

Method used

By obtaining the images taken by AR glasses, identifying preset gestures and obtaining continuous images within the preset duration, using three-dimensional coordinates and similarity to determine whether the gesture is static or dynamic, inputting the static and dynamic gesture databases respectively for matching, generating gesture operation renderings, supporting the recognition of static and dynamic gestures.

Benefits of technology

It improves the accuracy of gesture recognition and user interaction experience, can recognize complex and changeable gesture actions, meet diverse interaction needs, reduce the rate of misidentification and improve system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120233885A_ABST
    Figure CN120233885A_ABST
Patent Text Reader

Abstract

The invention discloses a gesture recognition method and system based on AR glasses interaction, and relates to the technical field of augmented reality glasses. Acquiring a first image; if the preset gesture exists in the first image, obtaining a second image corresponding to a preset duration according to the gesture interaction operation; processing the second image to obtain a third image and a fourth image; processing the third image and the fourth image to obtain a target distance and a first similarity; when the target distance is less than or equal to a preset distance and the first similarity is greater than or equal to a preset first similarity, summarizing the second image into a static image set; acquiring a preset first gesture database according to the static image set, and inputting the second image into the preset first gesture database for matching to obtain first identification information; and generating a first gesture operation effect picture according to the first identification information, and sending the first gesture operation effect picture to the target AR glasses. By implementing the technical scheme provided by the invention, the interaction experience between the user and the AR glasses is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of augmented reality glasses, and particularly to a gesture recognition method and system based on AR glasses interaction. Background Art

[0002] With the continuous progress of technology and the increasing market demand, the market potential of AR glasses is gradually emerging and accelerating. AR glasses, relying on intelligent high-definition technology, have successfully superimposed virtual information onto the real world, greatly enriching people's visual experience and cognitive scope.

[0003] In current AR glasses applications, users do not need to rely on handheld devices. They can easily interact with virtual information only through natural body movements such as eye rotation and head movement. In addition, AR glasses also widely support various interaction modes such as voice commands, gesture recognition, and touchpads, providing users with more rich and flexible interaction options. However, the existing AR glasses interaction methods still have certain limitations. Specifically, gesture recognition technology is currently mainly limited to the recognition of simple gesture actions and is difficult to meet complex and changeable interaction requirements, resulting in a poor interaction experience between users and AR glasses.

[0004] Therefore, there is an urgent need for a gesture recognition method and system based on AR glasses interaction that can solve the above technical problems. Summary of the Invention

[0005] This application provides a gesture recognition method and system based on AR glasses interaction. This method can solve the limitations currently faced by gesture recognition technology and improve the interaction experience between users and AR glasses.

[0006] First aspect, the present application provides a gesture recognition method based on AR glasses interaction, which is applied to a server. The method includes: obtaining a first image, where the first image is an image captured by a target AR glasses of a target hand of a target user; if a preset gesture exists in the first image, determining that the target user triggers a gesture interaction operation, and obtaining a second image corresponding to a preset duration according to the gesture interaction operation; processing the second image to obtain a third image and a fourth image, where the third image is an image corresponding to a first time point in the second image, the fourth image is an image corresponding to a second time point in the second image, the first time point is the start time point corresponding to the preset duration, the second time point is the end time point corresponding to the preset duration, and the first time point is earlier than the second time point; processing the third image and the fourth image to obtain a target distance and a first similarity, where the target distance is the distance between a first coordinate and a second coordinate, the first coordinate is the three-dimensional coordinate corresponding to a first gesture in the third image, the second coordinate is the three-dimensional coordinate corresponding to a second gesture in the fourth image, and the first similarity is the similarity corresponding to the first gesture and the second gesture; determining whether the target distance is less than or equal to a preset distance and whether the first similarity is greater than or equal to a preset first similarity; when the target distance is less than or equal to the preset distance and the first similarity is greater than or equal to the preset first similarity, incorporating the second image into a static image set; obtaining a preset first gesture database according to the static image set, inputting the second image into the preset first gesture database for matching to obtain first recognition information; generating a first gesture operation effect diagram according to the first recognition information, and sending the first gesture operation effect diagram to the target AR glasses so as to display the first gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

[0007] By adopting the above technical solution, gesture recognition is performed on the first image to determine the triggering of the gesture interaction operation, and continuous images within a preset duration, that is, the second image, are obtained. The third image and the fourth image are obtained from the second image, and the third image and the fourth image are processed respectively to obtain the target distance and the first similarity. The target distance uses three-dimensional coordinates to accurately locate the position of the gesture in space, which helps to accurately identify the position change of the gesture. The first similarity calculates the similarity of the gesture at different time points (the third image and the fourth image), and combines the distance change of the three-dimensional coordinates to determine whether the gesture in the second image is static or dynamic. When it is determined that the gesture in the second image is static, the preset first gesture database is retrieved, and the second image is input into the preset first gesture database for matching to obtain the first recognition information. Once the gesture is recognized, a corresponding gesture operation effect diagram is immediately generated and displayed to the target user through the AR glasses, enabling the user to intuitively see the result of the gesture operation, improving the accuracy of gesture recognition while further enhancing the interaction experience between the user and the AR glasses.

[0008] Optionally, after determining whether the target distance is less than or equal to a preset distance and whether the first similarity is greater than or equal to a preset first similarity, the method further includes: when the target distance is greater than the preset distance and the first similarity is less than the preset first similarity, incorporating the second image into the dynamic image set; obtaining a preset second gesture database according to the dynamic image set, inputting the second image into the preset second gesture database for matching to obtain second recognition information; generating a second gesture operation effect diagram according to the second recognition information, and sending the second gesture operation effect image to the target AR glasses, so as to display the second gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

[0009] By adopting the above technical solution, when the target distance is greater than the preset distance and the first similarity is less than the preset first similarity, it is determined that the second image is a dynamic image set, and different recognition strategies are adopted for the gesture actions of the static image set and the dynamic image set. This classification helps to improve the accuracy of gesture recognition. Determining a preset second gesture database according to the dynamic image set, inputting the second image into the preset second gesture database for matching to obtain second recognition information, can recognize both static and dynamic types of gesture actions, and flexibly adjust according to the actual needs of the user, thereby solving the main limitation in gesture recognition technology, which is limited to the recognition of simple gesture actions, enabling the user to interact with the AR glasses through more complex and diverse gesture actions, and thus meeting more diverse interaction needs.

[0010] Optionally, before inputting the second image into the preset second gesture database for matching to obtain second recognition information, the method further includes: performing frame-by-frame analysis on the second image to obtain a plurality of sub-images; obtaining a first occupancy ratio corresponding to the first sub-image, where the first sub-image is any one of the plurality of sub-images, and the first occupancy ratio is the proportion of the target hand in the first sub-image; determining whether the first occupancy ratio is greater than or equal to a preset occupancy ratio; when the first occupancy ratio is greater than or equal to the preset occupancy ratio, incorporating the first sub-image into the first image set and outputting the first image set as the second image.

[0011] By adopting the above technical solution, performing frame-by-frame analysis on the second image can capture the subtle changes in the hand movement, thereby more accurately recognizing the gesture. Then, obtaining the first sub-image from the plurality of sub-images and screening according to the occupancy ratio of the hand in the first sub-image can exclude the image frames with a small hand occupancy ratio that are difficult to recognize due to poor shooting angles, further improving the accuracy of hand recognition; and through screening and classification, the number of image frames to be processed can be reduced, the running time of the algorithm can be reduced, and the consumption of computing resources can be reduced.

[0012] Optionally, after determining whether the first occupancy ratio is greater than or equal to a preset occupancy ratio, the method further includes: when the first occupancy ratio is less than the preset occupancy ratio, classifying the first sub-image into the second image set, and obtaining a second sub-image from the multiple sub-images, where the second sub-image is any one of the multiple sub-images other than the first sub-image; obtaining a second occupancy ratio corresponding to the second sub-image; determining whether the second occupancy ratio is greater than or equal to the preset occupancy ratio; when the second occupancy ratio is greater than or equal to the preset occupancy ratio, classifying the second sub-image into the first image set.

[0013] By adopting the above technical solution, when the hand occupancy ratio of the first sub-image does not meet the preset condition, the first sub-image is classified into the second image set, and then the remaining sub-images are further analyzed. This way improves the utilization rate of image data, filters out images whose part occupancy ratio meets the preset condition from multiple sub-images, and can ensure that the image data used for gesture recognition has high quality and clarity, which helps to reduce the misrecognition rate and improve the accuracy and reliability of gesture recognition.

[0014] Optionally, after classifying the first sub-image into the first image set when the first occupancy ratio is greater than or equal to the preset occupancy ratio, the method further includes: obtaining a third sub-image and a fourth sub-image from the first image set; calculating a similarity between the third sub-image and the fourth sub-image to obtain a second similarity; determining whether the second similarity is greater than or equal to a preset second similarity; when the second similarity is greater than or equal to the preset second similarity, determining that the third sub-image and the fourth sub-image are images of the same gesture, and deleting the fourth sub-image from the first image set.

[0015] By adopting the above technical solution, calculating the similarity between the third sub-image and the fourth sub-image and deleting duplicate images with high similarity can significantly reduce the redundant data in the first image set, which not only helps to save storage space but also improves the efficiency of subsequent gesture recognition or processing. After removing the duplicate images, the remaining images in the first image set are more representative, can more accurately reflect the user's hand movements, and also help to reduce misrecognition caused by duplicate images and improve the accuracy of gesture recognition.

[0016] Optionally, inputting the second image into a preset second gesture database for matching to obtain second recognition information, specifically including: sorting each first sub-image in the first image set according to the chronological order to obtain a target sorting result; obtaining a fifth sub-image and a sixth sub-image, where the fifth sub-image is the first sub-image corresponding to the top position in the target sorting result, and the sixth sub-image is the first sub-image corresponding to the bottom position in the target sorting result; inputting both the fifth sub-image and the sixth sub-image into the preset second gesture database for matching to obtain second recognition information.

[0017] By adopting the above technical solution, sorting the sub-images in the first image set in terms of time can clearly capture the start and end states of the gesture action. Selecting the fifth sub-image ranked first and the sixth sub-image ranked last for matching can represent the start and end postures of the gesture. Analyzing the start and end states of the gesture can more comprehensively identify and understand the user's gesture intention. This helps reduce misrecognition caused by incomplete or ambiguous gesture actions and improves the accuracy and integrity of gesture recognition.

[0018] Optionally, after presenting the first gesture operation effect diagram on the display field of view of the target AR glasses to the target user, the method further includes: responding to a click operation of the target user on the first gesture operation effect diagram; determining the recognition information corresponding to the click operation, where the recognition information includes recognition success information and recognition failure information; when the recognition information is recognition success information, executing the corresponding target operation instruction according to the first gesture recognition information; when the recognition information is recognition failure information, sending a prompt message to the target user to prompt the target user to perform the gesture recognition operation again.

[0019] By adopting the above technical solution, obtaining the recognition result based on the first click gesture operation effect diagram and executing the target operation instruction or receiving the prompt for re-recognition accordingly, this interaction method is intuitive and easy to understand. It enables the user to clearly know whether the gesture is correctly recognized and take corresponding actions according to the recognition result, thereby enhancing the user's interaction experience. Through the clear recognition success or failure information, the user can clearly understand whether the gesture is correctly recognized. This helps reduce misoperations caused by misrecognition and improves the stability and reliability of the system.

[0020] In the second aspect of the present application, a gesture recognition system based on AR glasses interaction is provided. The system is a server, and the server includes an acquisition unit, a processing unit, and a sending unit; the acquisition unit acquires a first image, which is an image of the target hand of the target user captured by the target AR glasses; the processing unit, if a preset gesture exists in the first image, determines that the target user triggers a gesture interaction operation, and acquires a second image corresponding to a preset duration according to the gesture interaction operation; processes the second image to obtain a third image and a fourth image. The third image is the image corresponding to the first time point in the second image, and the fourth image is the image corresponding to the second time point in the second image. The first time point is the start time point corresponding to the preset duration, and the second time point is the end time point corresponding to the preset duration. The first time point is earlier than the second time point; processes the third image and the fourth image to obtain a target distance and a first similarity. The target distance is the distance between the first coordinate and the second coordinate. The first coordinate is the three-dimensional coordinate corresponding to the first gesture in the third image, and the second coordinate is the three-dimensional coordinate corresponding to the second gesture in the fourth image. The first similarity is the similarity corresponding to the first gesture and the second gesture; determines whether the target distance is less than or equal to a preset distance, and whether the first similarity is greater than or equal to a preset first similarity; when the target distance is less than or equal to the preset distance, and the first similarity is greater than or equal to the preset first similarity, incorporates the second image into the static image set; Obtains a preset first gesture database according to the static image set, inputs the second image into the preset first gesture database for matching to obtain first recognition information; the sending unit generates a first gesture operation effect diagram according to the first recognition information, and sends the first gesture operation effect diagram to the target AR glasses, so as to display the first gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

[0021] Optionally, the processing unit is used to incorporate the second image into the dynamic image set when the target distance is greater than the preset distance and the first similarity is less than the preset first similarity; obtains a preset second gesture database according to the dynamic image set, inputs the second image into the preset second gesture database for matching to obtain second recognition information; the sending unit is used to generate a second gesture operation effect diagram according to the second recognition information, and send the second gesture operation effect image to the target AR glasses, so as to display the second gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

[0022] Optionally, the processing unit is configured to perform frame-by-frame analysis on the second image to obtain a plurality of sub-images; the acquisition unit is configured to acquire a first occupancy ratio corresponding to a first sub-image, where the first sub-image is any one of the plurality of sub-images, and the first occupancy ratio is the occupancy ratio of the target hand in the first sub-image; the processing unit is configured to determine whether the first occupancy ratio is greater than or equal to a preset occupancy ratio; when the first occupancy ratio is greater than or equal to the preset occupancy ratio, the first sub-image is classified into the first image set, and the first image set is output as the second image.

[0023] Optionally, when the first occupancy ratio is less than the preset occupancy ratio, the processing unit is configured to classify the first sub-image into the second image set, and acquire a second sub-image from the plurality of sub-images, where the second sub-image is any one of the plurality of sub-images other than the first sub-image; the acquisition unit is configured to acquire a second occupancy ratio corresponding to the second sub-image; the processing unit is configured to determine whether the second occupancy ratio is greater than or equal to the preset occupancy ratio; when the second occupancy ratio is greater than or equal to the preset occupancy ratio, the second sub-image is classified into the first image set.

[0024] Optionally, the acquisition unit is configured to acquire a third sub-image and a fourth sub-image from the first image set; The processing unit is configured to calculate the similarity between the third sub-image and the fourth sub-image to obtain a second similarity; determine whether the second similarity is greater than or equal to a preset second similarity; when the second similarity is greater than or equal to the preset second similarity, determine that the third sub-image and the fourth sub-image are the same gesture image, and delete the fourth sub-image from the first image set.

[0025] Optionally, the processing unit is configured to sort the first sub-images in the first image set according to the chronological order to obtain a target sorting result; the acquisition unit is configured to acquire a fifth sub-image and a sixth sub-image, where the fifth sub-image is the first sub-image corresponding to the first position in the target sorting result, and the sixth sub-image is the first sub-image corresponding to the last position in the target sorting result; the processing unit is configured to input both the fifth sub-image and the sixth sub-image into a preset second gesture database for matching to obtain second recognition information.

[0026] Optionally, the processing unit is configured to respond to a click operation of the target user on the first gesture operation effect diagram; determine the recognition information corresponding to the click operation, where the recognition information includes a recognition success message and a recognition failure message; when the recognition information is a recognition success message, execute a corresponding target operation instruction according to the first gesture recognition information; the sending unit is configured to send a prompt message to the target user when the recognition information is a recognition failure message, so as to prompt the target user to perform a gesture recognition operation again.

[0027] In a third aspect of the present application, an electronic device is provided. The electronic device includes a processor, a memory, a user interface, and a network interface. The memory is used to store instructions. The user interface and the network interface are used to communicate with other devices. The processor is used to execute the instructions stored in the memory, so that an electronic device executes the method of any one of the above in the present application.

[0028] In a fourth aspect of the present application, a computer-readable storage medium is provided. The computer-readable storage medium stores instructions, and when the instructions are executed, the method of any one of the above in the present application is executed.

[0029] In summary, one or more technical solutions provided in the embodiments of the present application have at least the following technical effects or advantages: 1. Perform gesture recognition on the first image to determine the triggering of a gesture interaction operation, obtain consecutive images within a preset duration, that is, the second image, obtain the third image and the fourth image from the second image, and process the third image and the fourth image respectively to obtain the target distance and the first similarity. The target distance uses three-dimensional coordinates to accurately locate the position of the gesture in space, which helps to accurately identify the position change of the gesture. The first similarity calculates the similarity of the gesture at different time points (the third image and the fourth image), and combines the distance change of the three-dimensional coordinates to determine whether the gesture in the second image is static or dynamic. When it is determined that the gesture in the second image is static, retrieve the preset first gesture database, input the second image into the preset first gesture database for matching, and obtain the first recognition information. Once the gesture is recognized, immediately generate the corresponding gesture operation effect diagram and display it to the target user through the AR glasses, enabling the user to intuitively see the result of the gesture operation, while improving the accuracy of gesture recognition and further enhancing the interaction experience between the user and the AR glasses.

[0030] 2. Obtain the recognition result according to the first click gesture operation effect diagram, and accordingly execute the target operation instruction or receive a prompt for re-recognition. This interaction method is intuitive and easy to understand. It enables the user to clearly know whether the gesture is correctly recognized and take corresponding actions according to the recognition result, thereby enhancing the user's interaction experience. Through clear recognition success or failure information, the user can clearly understand whether the gesture is correctly recognized. This helps to reduce misoperations caused by misrecognition and improve the stability and reliability of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a schematic flowchart of a gesture recognition method based on AR glasses interaction provided by an embodiment of the present application; Figure 2 is a schematic structural diagram of a gesture recognition system based on AR glasses interaction provided by an embodiment of the present application; Figure 3It is a schematic structural diagram of an electronic device disclosed in an embodiment of the present application.

[0032] Explanation of reference numerals: 201, acquisition unit; 202, processing unit; 203, sending unit; 300, electronic device; 301, processor; 302, memory; 303, user interface; 304, network interface; 305, communication bus. Detailed implementation manners

[0033] In order to enable those skilled in the art to better understand the technical solutions in this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this specification. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments.

[0034] In the description of the embodiments of the present application, words such as "for example" or "for instance" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "for example" or "for instance" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "for example" or "for instance" is intended to present relevant concepts in a specific manner.

[0035] In the description of the embodiments of the present application, the meaning of the term "a plurality of" refers to two or more. For example, a plurality of systems refers to two or more systems, and a plurality of screen terminals refers to two or more screen terminals. In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the technical features indicated. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. The terms "include", "comprise", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0036] With the continuous progress of technology and the continuous increase in market demand, the market potential of AR glasses is gradually emerging and accelerating its release. AR glasses, relying on intelligent high-definition technology, have successfully superimposed virtual information onto the real world, greatly enriching people's visual experience and cognitive scope.

[0037] In current AR glasses applications, users do not need to rely on handheld devices. They can easily interact with virtual information through natural body movements such as eye rotation and head movement. In addition, AR glasses widely support various interaction modes such as voice commands, gesture recognition, and touchpads, providing users with richer and more flexible interaction options. However, the existing AR glasses interaction methods still have certain limitations. Specifically, gesture recognition technology is currently mainly limited to the recognition of simple gesture actions and is difficult to handle complex and changing interaction requirements, resulting in a poor interaction experience between users and AR glasses.

[0038] Therefore, how to solve the limitations currently faced by gesture recognition technology is an urgent problem to be solved. A gesture recognition method based on AR glasses interaction provided by an embodiment of this application is applied to a server. The server of this application can be a platform that provides interaction services for AR glasses. Figure 1 It is a schematic flowchart of a gesture recognition method based on AR glasses interaction provided by an embodiment of this application. Refer to Figure 1 and this method includes the following steps S101 - step S108.

[0039] S101: Obtain a first image, where the first image is an image captured by a target AR glasses of a target user's target hand.

[0040] In the above S101, the target user refers to the person wearing the target AR glasses. The target AR glasses are equipped with a camera for capturing the hand image of the target user. The target hand refers to the two hands of the target user. When the user wears AR glasses and extends their hands, the camera of the glasses will automatically or according to the user's instruction start capturing the hand image, and this image is called the first image. The process of image capture may involve image processing technologies such as automatic exposure and white balance adjustment to ensure image quality.

[0041] S102: If there is a preset gesture in the first image, determine that the target user triggers a gesture interaction operation, and obtain a second image corresponding to a preset duration according to the gesture interaction operation.

[0042] In the above S102, after obtaining the first image, an image recognition algorithm (such as a deep learning model) will be used to detect whether there is a preset gesture in the image. The preset gesture is the starting gesture for triggering the gesture interaction operation, such as making a fist or waving. When it is detected that the gesture of the target user in the first image is a preset gesture, it is defaulted that the user triggers the gesture interaction operation. At this time, the gesture interaction operation refers to recognizing the user's gesture action and then triggering the operation corresponding to the gesture action. Once the gesture interaction operation is triggered, the server will start recording an image sequence of a preset duration (such as 1 second to 30 seconds), and this image sequence is called the second image. During this period, the user's hand movements will be continuously captured to form an image sequence.

[0043] S103: Process the second image to obtain a third image and a fourth image. The third image is the image corresponding to the first time point in the second image, and the fourth image is the image corresponding to the second time point in the second image.

[0044] In the above S103, the first time point is the start time point corresponding to a preset duration, and the second time point is the end time point corresponding to the preset duration. The first time point is earlier than the second time point. From the second image, select the time points of two key points, the start time point (the first time point) and the end time point (the second time point). Then extract the images corresponding to these two time points and use them as the third image and the fourth image respectively. For example, if the preset duration is 30 seconds, the start time point is the 1st second, that is, intercept the image corresponding to the 1st second from the second image as the third image for output. The end time point is the 30th second, that is, intercept the image corresponding to the 30th second from the second image as the fourth image for output. When extracting the third image and the fourth image from the second image, image preprocessing such as denoising and enhancing contrast will be involved.

[0045] S104: Process the third image and the fourth image to obtain a target distance and a first similarity.

[0046] In the above S104, after obtaining the third image and the fourth image from the second image, identify a first gesture in the third image. At this time, the first gesture refers to the hand movement of the target user in the third image, and then identify a second gesture in the fourth image. The second gesture refers to the hand movement of the target user in the fourth image. An infrared depth sensor can be used to obtain the first coordinate of the first gesture in the three-dimensional space, and then use the infrared depth sensor to obtain the second coordinate of the second gesture in the three-dimensional space, and calculate the distance between the first coordinate and the second coordinate as the target distance. The target distance is the distance between the first coordinate and the second coordinate. The first coordinate is the three-dimensional coordinate corresponding to the first gesture in the third image, and the second coordinate is the three-dimensional coordinate corresponding to the second gesture in the fourth image. Since the first gesture and the second gesture are in the same three-dimensional space, the Euclidean distance between two points in the space can be used to calculate the straight-line distance between the first coordinate and the second coordinate, or other distance calculation formulas can be selected. However, the specific choice of which calculation formula can be based on the actual situation, and no more limitations are made here. At the same time, use an image similarity algorithm (such as feature matching, structural similarity index, etc.) to calculate the similarity between the first gesture in the third image and the second gesture in the fourth image as the first similarity.

[0047] S105: Determine whether the target distance is less than or equal to a preset distance, and whether the first similarity is greater than or equal to a preset first similarity.

[0048] In the above S105, after obtaining the target distance and the first similarity corresponding to the third image and the fourth image, the target distance is compared with a preset distance, and the first similarity is further compared with a preset first similarity. The preset distance and the preset first similarity are set for static or dynamic classification of the gestures of the target user. According to different comparison results, it is further determined whether the gesture of the target user in the second image is static or dynamic. Static means that the gesture does not change within a preset duration, and dynamic means that the gesture changes significantly within a preset duration.

[0049] S106: When the target distance is less than or equal to the preset distance and the first similarity is greater than or equal to the preset first similarity, the second image is classified into the static image set.

[0050] In the above S106, when the target distance is less than or equal to the preset distance and the first similarity is greater than or equal to the preset first similarity, it is considered that the first gesture and the second gesture are close enough in space and form to meet the conditions of a static gesture. The second image (or more specifically, the gesture sequence represented by the third image and the fourth image) that meets the conditions of the static gesture is classified into a static image set.

[0051] S107: Obtain a preset first gesture database according to the static image set, input the second image into the preset first gesture database for matching, and obtain the first recognition information.

[0052] In the above S107, a preset first gesture database is further obtained according to the static image set. The preset first gesture database stores images of various static gestures and corresponding operation instructions. The images in the static image set are input into the preset first gesture database for matching to find the most similar gesture and its corresponding operation instructions as the first recognition information. Since the static image set means that the hand movements of the target user do not change within a preset duration, a complete hand movement can be intercepted from the static image set and then input into the preset first gesture database for matching, and then the most similar gesture and the corresponding operation instructions can be matched. The operation instructions include turn on, turn off, pause, and start search, etc., and then the operation instructions are output as the first recognition information.

[0053] S108: Generate a first gesture operation effect diagram according to the first recognition information, and send the first gesture operation effect diagram to the target AR glasses so as to display the first gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

[0054] In the above S108, after the gesture action is recognized, the AR glasses can provide feedback to the user by displaying an operation effect diagram. Through the feedback, it is convenient for the user to determine whether the gesture action recognition is correct, and then perform subsequent interaction operations according to the user's feedback operation. According to the first recognition information, a gesture operation effect diagram is generated, which can be a text prompt, an animation demonstration, or other forms of visual feedback. Finally, this effect diagram is sent to the target user through the display system of the AR glasses for display in the user's display field of view.

[0055] In addition, after the target user views the first gesture operation effect diagram, it is necessary to monitor the first gesture operation effect diagram and process the user's click operation in a timely manner. Then, corresponding operations are performed according to the recognition information corresponding to the click operation, so as to realize the response and processing of the user's gesture operation, and the user can also know whether the operation corresponding to the gesture action is being executed currently. Specifically, it includes: responding to the click operation of the target user on the first gesture operation effect diagram; determining the recognition information corresponding to the click operation, where the recognition information includes recognition success information and recognition failure information; when the recognition information is recognition success information, execute the corresponding target operation instruction according to the first gesture recognition information; when the recognition information is recognition failure information, send a prompt message to the target user to prompt the target user to perform the gesture recognition operation again. Specifically, when displaying the first gesture operation effect diagram to the target user, options for recognition success and recognition failure will also be displayed in the display field of view of the AR glasses, and then the click operations of the target user on the two options will be monitored in real time. After detecting the click operation, capture the specific position of the click, that is, whether it is the position corresponding to the recognition success option or the position corresponding to the recognition failure option. Then, determine the recognition information according to the specific position. At this time, the recognition information is that if the click operation successfully recognizes the option, recognition success information is generated; if the click operation fails to recognize the option, recognition failure information is generated. When the recognition information is recognition success information, execute the corresponding target operation instruction according to the first gesture recognition information. The target operation instruction may be to open a certain application program, execute a certain function, send a message, etc. Ensure that the target operation instruction matches the first gesture recognition information to achieve the function expected by the user. When the recognition information is recognition failure information, send a prompt message to the target user. The prompt message can be displayed through the user interface of the application program, such as popping up a dialog box, displaying an error message, etc. The prompt message should be clear and understandable, informing the user of the reason for the recognition failure (such as the gesture does not meet the requirements, etc.), and guiding the user to perform the gesture recognition operation again. At this time, the gesture recognition operation is that the user needs to perform the gesture interaction operation again.

[0056] Further, when the target distance is greater than the preset distance and the first similarity is less than the preset first similarity, the second image is incorporated into the dynamic image set; the preset second gesture database is obtained according to the dynamic image set, the second image is input into the preset second gesture database for matching to obtain the second recognition information; the second gesture operation effect diagram is generated according to the second recognition information, and the second gesture operation effect image is sent to the target AR glasses so as to display the second gesture operation effect diagram to the target user in the display field of view of the target AR glasses. Specifically, first obtain the target distance between the first gesture in the third image and the second gesture in the fourth image, which can be achieved through sensors (such as depth sensors, infrared sensors, etc.) built in the AR glasses. Compare the target distance with the preset distance. The preset distance is usually set according to the design and usage scenario of the AR glasses to ensure the accuracy and effectiveness of gesture recognition. If the target distance is greater than the preset distance, it is defaulted that the positions of the first gesture in the third image and the second gesture in the fourth image change in the three-dimensional space. Then calculate the first similarity between the first gesture in the third image and the second gesture in the fourth image. The preset first similarity is used to determine whether the gestures in the third image and the fourth image are similar enough. If the first similarity is less than the preset first similarity, it is considered that the first gesture in the third image and the second gesture in the fourth image have changed, that is, the gesture in the second image has changed, so the second image is incorporated into the dynamic image set. The dynamic image set is used to store continuous gesture images that need to be recognized frame by frame. According to the image features in the dynamic image set, the preset second gesture database is retrieved. The preset second gesture database is a database containing various gesture features and data, which is used to match the images in the dynamic image set. In other words, the dynamic image set refers to the recognition of continuous hand movements, and the static image set refers to the recognition of fixed hand movements. The second image is input into the preset second gesture database for feature comparison and matching. The second recognition information, that is, the recognized gesture type or action, is generated according to the matching result. The corresponding second gesture operation effect diagram is generated according to the second recognition information. The generated second gesture operation effect diagram is sent to the target AR glasses. The target AR glasses will receive the effect diagram and display it to the target user in its display field of view. The user can see the recognized gesture operation effect diagram through the AR glasses and perform corresponding operations or give feedback as needed.

[0057] Furthermore, after incorporating the second image into the dynamic image set, the second image needs to be analyzed to obtain a first image set. Then, the images in the first image set are sequentially input into a preset second gesture database for matching, in order to better identify the information corresponding to the hand movement. Specifically, it includes: analyzing the second image frame by frame to obtain multiple sub-images; obtaining a first occupancy ratio corresponding to a first sub-image, where the first sub-image is any one of the multiple sub-images, and the first occupancy ratio is the proportion of the target hand in the first sub-image; determining whether the first occupancy ratio is greater than or equal to a preset occupancy ratio; when the first occupancy ratio is greater than or equal to the preset occupancy ratio, incorporating the first sub-image into the first image set and outputting the first image set as the second image. Specifically, use a video processing or image processing library (such as OpenCV, PIL, etc.) to analyze the second image frame by frame. If the second image is a video stream, extract each frame image sequentially; if the second image is a sequence of static images, process each image in order. After each frame or each image is extracted, it is used as a sub-image for subsequent processing. Obtain the multiple extracted sub-images, then obtain the first sub-image from the multiple sub-images, and then use an image recognition algorithm (such as a deep learning model, edge detection, color recognition, etc.) to identify the target hand in the first sub-image. This may involve preprocessing the sub-image, such as grayscale conversion, binarization, denoising, etc., to improve the accuracy of recognition. Once the target hand is identified, calculate its area or the number of pixels in the first sub-image. At the same time, calculate the total area or the total number of pixels of the sub-image. Divide the area or the number of pixels of the target hand by the total area or the total number of pixels of the sub-image to obtain the first occupancy ratio. According to the application scenario and requirements, preset an occupancy ratio threshold (i.e., the preset occupancy ratio). This threshold is used to determine whether the proportion of the target hand in the first sub-image is large enough for subsequent processing. Compare the calculated first occupancy ratio with the preset occupancy ratio. If the first occupancy ratio is greater than or equal to the preset occupancy ratio, it is considered that the target hand occupies a sufficient proportion in the first sub-image and meets the requirements for subsequent processing. When the first occupancy ratio meets the condition (i.e., is greater than or equal to the preset occupancy ratio), incorporate the corresponding first sub-image into a new image set, which is called the first image set. The first image set is used to store all sub-images that meet the conditions. After processing all sub-images, output the first image set as the new "second image". Here, the "second image" is actually an image set, rather than a single image.

[0058] For example, the first occupancy ratio of the target hand in the first sub-image is 0.8, and the preset occupancy ratio is set to 0.7. At this time, since the first occupancy ratio is greater than the preset occupancy ratio, the first sub-image is defaultly classified into the first image set. Then, according to the above process, continue to calculate and compare the occupancy ratios of other sub-images, and then obtain the first image set. Subsequently, the first image set can be used as the input of the preset second gesture database for output. That is, when inputting the second image into the preset second gesture database for matching, it is necessary to analyze the second image to ensure that the images input into the preset second gesture database are all complete and non-repetitive hand movements.

[0059] To ensure that each sub-image in the first image set represents a complete and non-redundant hand gesture, it is necessary to calculate the similarity of each sub-image in the first image set to identify and remove identical gesture images, which helps reduce the complexity of subsequent image matching. Specifically, it includes: obtaining a third sub-image and a fourth sub-image from the first image set; calculating the similarity between the third sub-image and the fourth sub-image to obtain a second similarity; determining whether the second similarity is greater than or equal to a preset second similarity; when the second similarity is greater than or equal to the preset second similarity, determining that the third sub-image and the fourth sub-image are identical gesture images and deleting the fourth sub-image from the first image set. Specifically, two sub-images are randomly or sequentially selected from the first image set and used as the third sub-image and the fourth sub-image respectively. The third sub-image and the fourth sub-image are loaded for subsequent similarity calculation. If the sub-images are of different sizes or formats, preprocessing such as resizing and grayscale conversion may be required to ensure the accuracy of similarity calculation. According to the application scenario and requirements, a suitable similarity calculation method is selected. Commonly used methods include histogram comparison, cosine similarity, hashing algorithms, mean squared error (MSE), structural similarity (SSIM), and feature matching. Here, since we are dealing with gesture images, we may be more concerned about the structure and features of the images, so SSIM or feature matching may be a suitable choice. The selected similarity calculation method is used to calculate the third sub-image and the fourth sub-image. This may involve converting the images into feature vectors, calculating the distance or similarity between features, and mapping the results to similarity scores. The calculated similarity score is used as the second similarity. According to the application scenario and requirements, a preset second similarity threshold is set. This threshold is used to determine whether two sub-images are similar enough to be considered identical gesture images. The calculated second similarity is compared with the preset second similarity. If the second similarity is greater than or equal to the preset second similarity, it is considered that the third sub-image and the fourth sub-image are similar enough and may be identical gesture images. When the second similarity meets the condition (i.e., is greater than or equal to the preset second similarity), it is determined that the third sub-image and the fourth sub-image are identical gesture images. This means that they may be images of the same gesture taken at the same time point. The fourth sub-image is deleted from the first image set to avoid duplicate or redundant data in subsequent processing. This can be achieved by removing the corresponding element from the image set list or marking it as deleted.

[0060] In addition, when the first occupancy ratio is less than the preset occupancy ratio, the first sub-image is classified into the second image set, and a second sub-image is obtained from multiple sub-images. The second sub-image is any one of the multiple sub-images other than the first sub-image; the second occupancy ratio corresponding to the second sub-image is obtained; it is determined whether the second occupancy ratio is greater than or equal to the preset occupancy ratio; when the second occupancy ratio is greater than or equal to the preset occupancy ratio, the second sub-image is classified into the first image set. Specifically, in the previous step, the first occupancy ratio of the first sub-image has been calculated and compared with the preset occupancy ratio. If the first occupancy ratio is less than the preset occupancy ratio, it means that the proportion of the target hand in the first sub-image is insufficient and does not meet the requirements of subsequent processing. The first sub-image is classified into a new image set, which is called the second image set. The second image set is used to store all sub-images that do not meet the occupancy ratio requirements. After processing the first sub-image, one of the remaining sub-images is selected as the second sub-image. Ensure that the selected second sub-image is not the previously processed first sub-image. The target hand is recognized in the second sub-image, which is the same as the steps when processing the first sub-image. The same image recognition algorithm and preprocessing steps are used to recognize the target hand. Once the target hand is recognized, calculate its area or the number of pixels in the second sub-image. At the same time, calculate the total area or the total number of pixels of the second sub-image. Divide the area or the number of pixels of the target hand by the total area or the total number of pixels of the second sub-image to obtain the second occupancy ratio. Compare the calculated second occupancy ratio with the preset occupancy ratio. If the second occupancy ratio is greater than or equal to the preset occupancy ratio, it is considered that the target hand occupies a sufficient proportion in the second sub-image and meets the requirements of subsequent processing. When the second occupancy ratio meets the condition (i.e., is greater than or equal to the preset occupancy ratio), the corresponding second sub-image is classified into the previously mentioned first image set. The first image set is used to store all sub-images that meet the conditions, and these sub-images will be used for subsequent processing or analysis. By analyzing sub-images frame by frame, calculating occupancy ratios, judging conditions, and classifying sub-images into different image sets, etc., the precise control of the proportion of the target hand in the sub-image and the basis for subsequent processing are achieved.

[0061] Further, input the second image into a preset second gesture database for matching to obtain second recognition information, which specifically includes: sorting each first sub-image in the first image set according to the chronological order to obtain a target sorting result; acquiring a fifth sub-image and a sixth sub-image, where the fifth sub-image is the first sub-image corresponding to the first position in the target sorting result, and the sixth sub-image is the first sub-image corresponding to the last position in the target sorting result; inputting both the fifth sub-image and the sixth sub-image into the preset second gesture database for matching to obtain second recognition information. Specifically, when processing each first sub-image, first extract its corresponding timestamp. The timestamp can be metadata read from the image file or time information recorded during image acquisition. Ensure that each first sub-image has a unique and accurate timestamp for subsequent sorting. Use a sorting algorithm to sort each first sub-image in the first image set according to the timestamp. The sorting algorithm can be a simple bubble sort, selection sort, insertion sort, etc., or more efficient quicksort, merge sort, etc. After sorting, obtain a list of first sub-images arranged in chronological order, that is, the target sorting result. The first element in the target sorting result is the first sub-image with the earliest timestamp, and the last element is the first sub-image with the latest timestamp. Extract the first element from the target sorting result as the fifth sub-image, that is, the first sub-image ranked first. Extract the last element from the target sorting result as the sixth sub-image, that is, the first sub-image ranked last. Load the fifth sub-image and the sixth sub-image, and select a suitable matching algorithm according to the application scenario and requirements. Commonly used matching algorithms include template matching, feature point matching, deep learning model matching, etc. Since this application processes gesture images and matches them with gesture actions pre-stored in the preset second gesture database, algorithms such as template matching or feature point matching can be selected. Input both the fifth sub-image and the sixth sub-image into the preset second gesture database for matching. That is, find in the preset second gesture database which gesture template best matches both the fifth sub-image and the sixth sub-image, with the fifth sub-image in front and the sixth sub-image behind. After matching, obtain the gesture information that best matches both the fifth sub-image and the sixth sub-image from the preset second gesture database as the second recognition information, and the second recognition information refers to the operation instruction corresponding to the current hand movement. To improve the accuracy of dynamic image recognition, each sub-image can also be input into the preset second gesture database for matching according to the chronological order of the target sorting result, and then find the gesture information that best matches the consecutive sub-images in the target sorting result in the preset second gesture database, and then output the operation instruction corresponding to the gesture information. It can be understood that the preset first gesture database recognizes fixed hand movements and then outputs recognition information. The preset second gesture database recognizes consecutive hand movements and then outputs recognition information.

[0062] Using the above method, gesture recognition is performed on the first image to determine the triggering of a gesture interaction operation, and continuous images within a preset duration, i.e., the second image, are obtained. The third image and the fourth image are obtained from the second image, and the third image and the fourth image are processed respectively to obtain the target distance and the first similarity. Based on the target distance and the first similarity, it is determined whether the gesture in the second image is static or dynamic. Different recognition strategies are adopted for the gesture actions of the static image set and the dynamic image set. This classification helps to improve the accuracy of gesture recognition. If the second image is determined to be a static image set, a preset first gesture database is determined according to the static image set, and the second image is input into the preset first gesture database for matching to obtain the first recognition information; if the second image is determined to be a dynamic image set, a preset second gesture database is determined according to the dynamic image set, and the second image is input into the preset second gesture database for matching to obtain the second recognition information. It can recognize both static and dynamic types of gesture actions, and can be flexibly adjusted according to the actual needs of users, thereby solving the main limitation in gesture recognition technology, which is the recognition of simple gesture actions, enabling users to interact with the AR glasses through more complex and diverse gesture actions, and thus meeting more diverse interaction needs.

[0063] An embodiment of the present application further provides a gesture recognition system based on AR glasses interaction. Figure 2 It is a schematic structural diagram of a gesture recognition system based on AR glasses interaction provided by an embodiment of the present application. Refer to Figure 2 In the figure, the system is a server, and the server includes an acquisition unit 201, a processing unit 202, and a sending unit 203.

[0064] The acquisition unit 201 acquires a first image, where the first image is an image captured by the target AR glasses of the target hand of the target user.

[0065] The processing unit 202, if there is a preset gesture in the first image, determines that the target user triggers a gesture interaction operation, and obtains a second image corresponding to a preset duration according to the gesture interaction operation; processes the second image to obtain a third image and a fourth image, where the third image is the image corresponding to the first time point in the second image, and the fourth image is the image corresponding to the second time point in the second image. The first time point is the start time point corresponding to the preset duration, and the second time point is the end time point corresponding to the preset duration. The first time point is earlier than the second time point; processes the third image and the fourth image to obtain a target distance and a first similarity. The target distance is the distance between the first coordinate and the second coordinate. The first coordinate is the three-dimensional coordinate corresponding to the first gesture in the third image, and the second coordinate is the three-dimensional coordinate corresponding to the second gesture in the fourth image. The first similarity is the similarity corresponding to the first gesture and the second gesture; determines whether the target distance is less than or equal to a preset distance, and whether the first similarity is greater than or equal to a preset first similarity; when the target distance is less than or equal to the preset distance, and the first similarity is greater than or equal to the preset first similarity, incorporates the second image into the static image set; obtains a preset first gesture database according to the static image set, and inputs the second image into the preset first gesture database for matching to obtain first recognition information.

[0066] The sending unit 203 generates a first gesture operation effect diagram according to the first recognition information, and sends the first gesture operation effect diagram to the target AR glasses, so as to display the first gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

[0067] In a possible implementation manner, the processing unit 202 is configured to, when the target distance is greater than the preset distance and the first similarity is less than the preset first similarity, incorporate the second image into the dynamic image set; obtain a preset second gesture database according to the dynamic image set, and input the second image into the preset second gesture database for matching to obtain second recognition information; the sending unit 203 is configured to generate a second gesture operation effect diagram according to the second recognition information, and send the second gesture operation effect diagram to the target AR glasses, so as to display the second gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

[0068] In a possible implementation manner, the processing unit 202 is configured to perform frame-by-frame analysis on the second image to obtain a plurality of sub-images; the obtaining unit 201 is configured to obtain a first occupancy ratio corresponding to a first sub-image, where the first sub-image is any one of the plurality of sub-images, and the first occupancy ratio is the occupancy ratio of the target hand in the first sub-image; the processing unit 202 is configured to determine whether the first occupancy ratio is greater than or equal to a preset occupancy ratio; when the first occupancy ratio is greater than or equal to the preset occupancy ratio, incorporate the first sub-image into the first image set and output the first image set as the second image.

[0069] In a possible implementation, the processing unit 202 is configured to, when the first occupancy ratio is less than a preset occupancy ratio, classify the first sub-image into the second image set, and obtain a second sub-image from multiple sub-images, where the second sub-image is any one of the multiple sub-images other than the first sub-image; the obtaining unit 201 is configured to obtain a second occupancy ratio corresponding to the second sub-image; the processing unit 202 is configured to determine whether the second occupancy ratio is greater than or equal to the preset occupancy ratio; when the second occupancy ratio is greater than or equal to the preset occupancy ratio, classify the second sub-image into the first image set.

[0070] In a possible implementation, the obtaining unit 201 is configured to obtain a third sub-image and a fourth sub-image from the first image set; the processing unit 202 is configured to calculate a similarity between the third sub-image and the fourth sub-image to obtain a second similarity; determine whether the second similarity is greater than or equal to a preset second similarity; when the second similarity is greater than or equal to the preset second similarity, determine that the third sub-image and the fourth sub-image are the same gesture images, and delete the fourth sub-image from the first image set.

[0071] In a possible implementation, the processing unit 202 is configured to sort each first sub-image in the first image set according to the chronological order to obtain a target sorting result; the obtaining unit 201 is configured to obtain a fifth sub-image and a sixth sub-image, where the fifth sub-image is the first sub-image corresponding to the first position in the target sorting result, and the sixth sub-image is the first sub-image corresponding to the last position in the target sorting result; the processing unit 202 is configured to input both the fifth sub-image and the sixth sub-image into a preset second gesture database for matching to obtain second recognition information.

[0072] In a possible implementation, the processing unit 202 is configured to respond to a click operation of a target user on the first gesture operation effect diagram; determine recognition information corresponding to the click operation, where the recognition information includes recognition success information and recognition failure information; when the recognition information is recognition success information, execute a corresponding target operation instruction according to the first gesture recognition information; the sending unit 203 is configured to, when the recognition information is recognition failure information, send a prompt message to the target user to prompt the target user to perform a gesture recognition operation again.

[0073] It should be noted that: when the system provided in the above embodiments implements its functions, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the system and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0074] This application also discloses an electronic device. Refer to Figure 3 , Figure 3 which is a schematic structural diagram of an electronic device provided by an embodiment of this application. The electronic device 300 may include: at least one processor 301, at least one network interface 304, a user interface 303, a memory 302, and at least one communication bus 305.

[0075] Among them, the communication bus 305 is used to realize the connection and communication between these components.

[0076] Among them, the user interface 303 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 303 may further include a standard wired interface and a wireless interface.

[0077] Among them, the network interface 304 may optionally include a standard wired interface and a wireless interface (such as a WI-FI interface).

[0078] Among them, the processor 301 may include one or more processing cores. The processor 301 connects various parts within the entire server through various interfaces and lines, and by running or executing instructions, programs, code sets, or instruction sets stored in the memory 302, as well as calling data stored in the memory 302, it executes various functions of the server and processes data. Optionally, the processor 301 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 301 may integrate one or several combinations of a central processing unit (CPU), a graphics processing unit (GPU), and a modem, etc. Among them, the CPU mainly processes operating systems, user interfaces, and application requests, etc.; the GPU is responsible for the rendering and drawing of the content to be displayed on the display screen; the modem is used to process wireless communications. It can be understood that the above modem may also not be integrated into the processor 301 and may be implemented separately by a single chip.

[0079] Among them, the memory 302 may include a Random Access Memory (RAM), or may include a Read-Only Memory. Optionally, the memory 302 includes a non-transitory computer-readable storage medium. The memory 302 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 302 may include a program storage area and a data storage area. Among them, the program storage area can store instructions for implementing the operating system, instructions for at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the above-mentioned method embodiments, etc.; the data storage area can store the data involved in the above-mentioned method embodiments. Optionally, the memory 302 may also be at least one storage device located far from the aforementioned processor 301.

[0080] As Figure 3 shown, the memory 302, as a computer storage medium, may include an operating system, a network communication module, a user interface module, and an application program for gesture recognition based on AR glasses interaction.

[0081] In Figure 3 the electronic device 300 shown, the user interface 303 is mainly used to provide an input interface for the user to obtain the data input by the user; while the processor 301 can be used to call the application program for gesture recognition based on AR glasses interaction stored in the memory 302. When executed by one or more processors, the electronic device is enabled to execute one or more of the methods as described in the above embodiments.

[0082] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that this application is not limited by the described action sequence, because according to this application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by this application.

[0083] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0084] In several embodiments provided by the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some service interfaces. The indirect coupling or communication connection of the device or unit can be in an electrical or other form.

[0085] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0086] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0087] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. And the aforementioned memory includes: various media such as USB flash drives, mobile hard disks, magnetic disks or optical discs that can store program codes.

[0088] The above are only exemplary embodiments of the present disclosure, and the scope of the present disclosure cannot be limited thereby. That is, all equivalent changes and modifications made in accordance with the teachings of the present disclosure still fall within the scope covered by the present disclosure. Those skilled in the art will readily think of other implementation schemes of the present disclosure after considering the specification and practice of the present disclosure. The present application aims to cover any variations, uses or adaptive changes of the present disclosure, and these variations, uses or adaptive changes follow the general principles of the present disclosure and include the common general knowledge or conventional technical means in the technical field not recorded in the present disclosure.

Claims

1. A gesture recognition method based on AR glasses interaction, characterized in that: Applied in a server, the method comprises: Acquire a first image, where the first image is an image of a target hand of a target user captured by a target AR glasses; If there is a preset gesture in the first image, determining that the target user triggers a gesture interaction operation, and acquiring a second image corresponding to a preset duration according to the gesture interaction operation; Processing the second image to obtain a third image and a fourth image, the third image being an image corresponding to a first time point in the second image, the fourth image being an image corresponding to a second time point in the second image, the first time point being a corresponding start time point in the preset duration, the second time point being a corresponding end time point in the preset duration, and the first time point being earlier than the second time point; Processing the third image and the fourth image to obtain a target distance and a first similarity, wherein the target distance is a distance between a first coordinate and a second coordinate, the first coordinate is a three-dimensional coordinate corresponding to a first gesture in the third image, the second coordinate is a three-dimensional coordinate corresponding to a second gesture in the fourth image, and the first similarity is a similarity between the first gesture and the second gesture; Determine whether the target distance is less than or equal to a preset distance, and whether the first similarity is greater than or equal to a preset first similarity; When the target distance is less than or equal to the preset distance, and the first similarity is greater than or equal to the preset first similarity, classifying the second image into a static image set; Acquire a preset first gesture database according to the static image set, input the second image into the preset first gesture database for matching, and obtain first recognition information; A first gesture operation effect diagram is generated according to the first recognition information, and the first gesture operation effect diagram is sent to the target AR glasses so as to display the first gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

2. The method according to claim 1, characterized in that: After determining whether the target distance is less than or equal to a preset distance and whether the first similarity is greater than or equal to a preset first similarity, the method further includes: When the target distance is greater than the preset distance and the first similarity is less than the preset first similarity, the second image is included in the dynamic image set; Acquire a preset second gesture database according to the dynamic image set, input the second image into the preset second gesture database for matching, and obtain second recognition information; A second gesture operation effect diagram is generated according to the second recognition information, and the second gesture operation effect diagram is sent to the target AR glasses so as to display the second gesture operation effect diagram to the target user in the display field of view of the target AR glasses.

3. The method according to claim 2, characterized in that Before inputting the second image into the preset second gesture database for matching to obtain second recognition information, the method further includes: Analyzing the second image frame by frame to obtain a plurality of sub-images; Acquire a first proportion corresponding to a first sub-image, where the first sub-image is any one of the plurality of sub-images, and the first proportion is a proportion of the target hand in the first sub-image; Determine whether the first proportion is greater than or equal to a preset proportion; When the first proportion is greater than or equal to the preset proportion, the first sub-image is summarized into a first image set, and the first image set is output as the second image.

4. The method according to claim 3, characterized in that After determining whether the first proportion is greater than or equal to a preset proportion, the method further includes: When the first proportion is less than the preset proportion, the first sub-image is included in a second image set, and a second sub-image is obtained from the plurality of sub-images, where the second sub-image is any sub-image among the plurality of sub-images except the first sub-image; Obtaining a second proportion corresponding to the second sub-image; Determine whether the second proportion is greater than or equal to the preset proportion; When the second proportion is greater than or equal to the preset proportion, the second sub-image is included in the first image set.

5. The method according to claim 4, characterized in that After incorporating the first sub-image into a first image set when the first proportion is greater than or equal to the preset proportion, the method further includes: Acquire a third sub-image and a fourth sub-image from the first image set; Calculating similarity between the third sub-image and the fourth sub-image to obtain a second similarity; Determining whether the second similarity is greater than or equal to a preset second similarity; When the second similarity is greater than or equal to the preset second similarity, it is determined that the third sub-image and the fourth sub-image are the same gesture image, and the fourth sub-image is deleted from the first image set.

6. The method according to claim 3, characterized in that The step of inputting the second image into the preset second gesture database for matching to obtain second recognition information specifically includes: Sorting each of the first sub-images in the first image set according to a chronological order to obtain a target sorting result; Acquire a fifth sub-image and a sixth sub-image, wherein the fifth sub-image is the first sub-image corresponding to the first position in the target ranking result, and the sixth sub-image is the first sub-image corresponding to the last position in the target ranking result; The fifth sub-image and the sixth sub-image are both input into the preset second gesture database for matching to obtain the second recognition information.

7. The method according to claim 1, characterized in that After displaying the first gesture operation effect diagram to the target user on the display field of view of the target AR glasses, the method further includes: In response to a click operation of the target user on the first gesture operation effect diagram; Determine identification information corresponding to the click operation, wherein the identification information includes identification success information and identification failure information; When the recognition information is the recognition success information, executing a corresponding target operation instruction according to the first gesture recognition information; When the recognition information is the recognition failure information, a prompt message is sent to the target user to prompt the target user to perform the gesture recognition operation again.

8. A gesture recognition system based on AR glasses interaction, characterized in that: The system is a server, and the server comprises an acquisition unit (201), a processing unit (202) and a sending unit (203); The acquisition unit (201) acquires a first image, where the first image is an image of a target hand of a target user photographed by a target AR glasses; The processing unit (202) determines that the target user triggers a gesture interaction operation if a preset gesture exists in the first image, and acquires a second image corresponding to a preset duration according to the gesture interaction operation; Processing the second image to obtain a third image and a fourth image, the third image being an image corresponding to a first time point in the second image, the fourth image being an image corresponding to a second time point in the second image, the first time point being a corresponding start time point in the preset duration, the second time point being a corresponding end time point in the preset duration, and the first time point being earlier than the second time point; Processing the third image and the fourth image to obtain a target distance and a first similarity, wherein the target distance is a distance between a first coordinate and a second coordinate, the first coordinate is a three-dimensional coordinate corresponding to a first gesture in the third image, the second coordinate is a three-dimensional coordinate corresponding to a second gesture in the fourth image, and the first similarity is a similarity between the first gesture and the second gesture; determining whether the target distance is less than or equal to a preset distance, and whether the first similarity is greater than or equal to a preset first similarity; When the target distance is less than or equal to the preset distance, and the first similarity is greater than or equal to the preset first similarity, classifying the second image into a static image set; Acquire a preset first gesture database according to the static image set, input the second image into the preset first gesture database for matching, and obtain first recognition information; The sending unit (203) generates a first gesture operation effect diagram according to the first recognition information, and sends the first gesture operation effect diagram to the target AR glasses, so as to display the first gesture operation effect diagram to the target user in a display field of view of the target AR glasses.

9. An electronic device, characterized in that: The electronic device (300) comprises a processor (301), a memory (302), a user interface (303) and a network interface (304), wherein the memory (302) is used to store instructions, the user interface (303) and the network interface (304) are used to communicate with other devices, and the processor (301) is used to execute the instructions stored in the memory (302) so that the electronic device (300) executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores instructions, and when the instructions are executed, the method according to any one of claims 1 to 7 is executed.

Citation Information

Cited By

  • Intelligent glasses gesture interaction method, device and equipment

    CN120743119A