Method for sign language recognition using AR glasses, AR glasses and storage medium

By using the YOLOv8 model in AR glasses to detect lip words and combined with gesture statements, the problem of low sign language recognition accuracy in the prior art is solved, and more efficient and accurate sign language recognition and communication is achieved.

CN118015696BActive Publication Date: 2025-05-13GUANGZHOU GUDONG INTELLIGENT TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410044354.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-05-13
Estimated Expiration
2044-01-11

AI Technical Summary

Technical Problem

In the prior art, the text translated through sign language is relatively stiff, and due to the rapid transformation of recognition angles or gestures, it is difficult to correctly recognize the intermediate gestures, resulting in a large difference between the concatenated sentences and the meanings they want to express, and slow gestures are required to correctly recognize them.

Method used

AR glasses are used to obtain communication videos, and the first YOLOv8 model is used to detect whether there is a lip word. If it exists, gesture statements and lip word statements are generated and combined into recognition results. If it does not exist, gesture statements are generated as recognition results, and displayed on AR glasses or played in voice form. Improve the accuracy of sign language recognition by combining lip language.

Benefits of technology

It improves the accuracy of sign language recognition, reduces recognition errors, enhances the efficiency and accuracy of communication with sign language people, and provides better privacy protection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118015696B_ABST
    Figure CN118015696B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence, and specifically to a method for sign language recognition using AR glasses, AR glasses and storage media, aiming to improve the accuracy of sign language recognition. The sign language recognition method of the present invention comprises: obtaining a communication video through AR glasses; inputting the communication video into a first YOLOv8 model to detect whether there is lip reading; if there is lip reading, generating gesture sentences and lip reading sentences according to the communication video respectively and combining them into a recognition result; if there is no lip reading, generating gesture sentences according to the communication video as a recognition result; and displaying the recognition result on AR glasses in the form of text or playing it in the form of voice. The present invention uses a method for dual recognition of lip reading and sign language to improve the accuracy of sign language recognition, and using AR glasses for recognition increases the real-time nature of sign language recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular to a method for performing sign language recognition using AR glasses, AR glasses, and a storage medium. Background Art

[0002] Sign language is the use of hand gestures to measure movements, imitating images or syllables based on changes in gestures to form certain meanings or words. It is a language of the hands for people with hearing impairments or who cannot speak to communicate and exchange ideas with each other. It is "an important auxiliary tool for spoken language", and for people with hearing impairments, it is the main tool for communication.

[0003] For users who have no sign language knowledge, when communicating with sign language speakers, they can first use a camera to collect video data of sign language speakers communicating in sign language, then extract image frames including gesture actions from the video data, and then use a gesture recognition model to recognize the gesture actions in the key image frames to obtain semantic information corresponding to the gesture actions. Finally, AR (Augmented Reality) technology can be used to display the semantic information to users communicating with sign language speakers in three-dimensional space.

[0004] However, in the prior art, the text translated through sign language is rather stiff, and due to the rapid change of recognition angle or gesture, sometimes the middle gesture cannot be correctly recognized, and even the concatenated sentences are very different from the intended meaning. It is necessary to express the gestures at a slower speed in order to more accurately recognize all the gestures. Summary of the invention

[0005] In order to solve the above problems in the prior art, the present invention proposes a method for sign language recognition using AR glasses, AR glasses and a storage medium, which improves the accuracy of sign language recognition.

[0006] A first aspect of the present invention provides a method for sign language recognition using AR glasses, the method comprising:

[0007] Get the communication video through AR glasses;

[0008] Input the communication video into a first YOLOv8 model to detect whether there is lip movement;

[0009] If lip reading exists, respectively generating gesture sentences and lip reading sentences according to the communication video and combining them into a recognition result;

[0010] If there is no lip reading, generating a gesture sentence as a recognition result according to the communication video;

[0011] The recognition result is displayed on the AR glasses in the form of text or played in the form of voice.

[0012] Preferably, the step of "generating gesture sentences and lip language sentences respectively according to the communication video and combining them into recognition results" includes:

[0013] Input the communication video into a second YOLOv8 model to obtain an atomic action sequence and a corresponding first discrete sentence and a gesture confidence mean;

[0014] Inputting the atomic action sequence and the first discrete sentence into a two-gate LSTM model to generate a second discrete sentence corresponding to the gesture;

[0015] Input the communication video into the first YOLOv8 model to obtain a third discrete sentence corresponding to the lip reading and a mean lip reading confidence value;

[0016] Inputting the second discrete sentence into the first Transformer model, so that the first Transformer model performs semantic understanding on the second discrete sentence in combination with the context, and then supplements or adjusts the missing or incoherent parts to obtain the gesture sentence;

[0017] Inputting the third discrete sentence into the second Transformer model, so that the second Transformer model performs semantic understanding on the third discrete sentence in combination with the context, and then supplements or adjusts the missing or incoherent parts to obtain the lip reading sentence;

[0018] Combining the gesture sentence and the lip reading sentence according to the gesture confidence mean, the lip reading confidence mean and a preset weight to obtain a recognition result;

[0019] The recognition results are fed back to the first Transformer model and the second Transformer model respectively, and integrated into the context to participate in subsequent semantic understanding.

[0020] Preferably, the preset weight includes a gesture weight and a lip reading weight, and the sum of the gesture weight and the lip reading weight is 1;

[0021] The step of “combining the gesture sentence and the lip reading sentence according to the gesture confidence mean, the lip reading confidence mean and a preset weight to obtain a recognition result” includes:

[0022] If the product of the gesture confidence mean and the gesture weight is greater than or equal to the product of the lip reading confidence mean and the lip reading weight, the gesture sentence is taken as the recognition result; otherwise, the lip reading sentence is taken as the recognition result.

[0023] Preferably, the step of "generating a gesture sentence as a recognition result according to the communication video" includes:

[0024] Input the communication video into a second YOLOv8 model to obtain an atomic action sequence and a corresponding first discrete sentence;

[0025] Inputting the atomic action sequence and the first discrete sentence into a two-gate LSTM model to generate a second discrete sentence corresponding to the gesture;

[0026] Inputting the second discrete sentence into the first Transformer model, so that the first Transformer model performs semantic understanding on the second discrete sentence in combination with the context, and then supplements or adjusts the missing or incoherent parts to obtain a recognition result;

[0027] The recognition result is fed back to the first Transformer model and integrated into the context to participate in subsequent semantic understanding.

[0028] Preferably, the step of “obtaining communication video through AR glasses” includes:

[0029] If blur, ghosting or residual image exists in the current image, the current image is discarded and a new frame is acquired;

[0030] If the brightness of the current image is lower than a preset brightness range, turning on the fill light of the AR glasses;

[0031] If the brightness of the current image is higher than the preset brightness range, adjusting the brightness of the current image to the preset brightness range;

[0032] If the resolution of the current image is higher than the preset resolution range, the current image is compressed to the preset resolution range.

[0033] Preferably, the first YOLOv8 model is pre-trained using a first training set containing lip reading videos; and the second YOLOv8 model is pre-trained using a second training set containing gesture videos.

[0034] Preferably, the two-gate LSTM model includes: 2 input gates, 2 forget gates and 1 output gate;

[0035] The training method of the first Transformer model and the second Transformer model includes:

[0036] Constructing a third training set, wherein the third training set includes: discrete sentences generated by gestures and discrete sentences generated by lip reading;

[0037] Selecting discrete sentences from the third training set and inputting them into the Transformer model to be trained;

[0038] Compare the output result of the Transformer model to be trained with the true label and calculate the loss function;

[0039] Using a back propagation algorithm to update the parameters of the Transformer model to be trained;

[0040] Repeat the process until the loss function is minimized.

[0041] The trained Transformer model is copied as the first Transformer model and the second Transformer model.

[0042] Preferably, the method for determining the gesture weight and the lip reading weight includes:

[0043] Constructing a data set, wherein the data set includes a video clip and a corresponding correct sentence; the video clip also includes gestures and lip readings;

[0044] Selecting a video clip from the data set each time;

[0045] Input the selected video clip into the second YOLOv8 model to obtain a sign language action sequence and a corresponding first discrete sentence;

[0046] Inputting the sign language action sequence and the first discrete sentence into a two-gate LSTM model to generate a second discrete sentence corresponding to the gesture;

[0047] Input the selected video clip into the first YOLOv8 model to obtain a third discrete sentence corresponding to the lip reading;

[0048] Inputting the second discrete sentence and the third discrete sentence into the first Transformer model and the second Transformer model respectively to obtain a gesture sentence and a lip reading sentence respectively;

[0049] Calculate the similarity between the gesture sentence and the lip reading sentence and the correct sentence respectively by using the TF-IDF algorithm to obtain the sign language similarity and the lip reading similarity;

[0050] Normalizing the sign language similarity and the lip reading similarity respectively;

[0051] If the sign language similarity is greater than or equal to the lip reading similarity, the value of the first weight is increased by a preset step size and the value of the second weight is decreased by the preset step size; otherwise, the value of the first weight is decreased by the preset step size and the value of the second weight is increased by the preset step size;

[0052] Repeating the steps of selecting video clips, calculating similarities, and adjusting weights until all video clips in the data set are used up;

[0053] Setting the values ​​of the first weight and the second weight to the values ​​of the gesture weight and the lip reading weight respectively;

[0054] The initial values ​​of the first weight and the second weight are both 0.5.

[0055] According to a second aspect of the present invention, a pair of AR glasses is provided, wherein the AR glasses perform sign language recognition according to the method described above.

[0056] According to a third aspect of the present invention, a computer-readable storage medium is provided, storing a computer program that can be loaded by a processor and execute the method described above.

[0057] The present invention has the following beneficial effects:

[0058] The method of using AR glasses for sign language recognition proposed in the present invention is mainly aimed at users who want to communicate with sign language professionals but lack sign language skills. After wearing AR glasses, first obtain the communication video through the camera on the glasses, and then use the first YOLOv8 model to detect whether there is lip reading in the video. If there is lip reading, gesture sentences and lip reading sentences are generated according to the communication video and combined into recognition results, and displayed on AR glasses or played in voice form. The accuracy of sign language recognition is effectively improved by combining lip reading. In addition, the use of AR glasses for shooting, recognition and display not only avoids the transmission and storage of personal data on the network, provides better privacy protection, but also increases the real-time nature of sign language recognition and improves the efficiency of communication.

[0059] When the gesture sentences and lip reading sentences are generated according to the communication video and combined into the recognition result, the present invention inputs the atomic action sequence output by the second YOLOv8 model and the corresponding first discrete sentence into the dual-gate LSTM model, accurately interprets the emotional factors of the sign language, captures more details to generate the second discrete sentence corresponding to the gesture; and inputs the second discrete sentence corresponding to the gesture and the third discrete sentence corresponding to the lip reading into the first Tranformer model and the first Tranformer model respectively, makes full use of the context to understand, and supplements or adjusts the missing or incoherent parts; finally, the gesture sentences and lip reading sentences are combined according to the gesture confidence mean, the lip reading confidence mean and the preset weight to obtain the recognition result. Because the gesture confidence mean and the lip reading confidence mean are calculated by the second YOLOv8 model and the first YOLOv8 model respectively, these two values ​​will change with the change of the input video, thereby realizing the dynamic combination of the recognition results according to the sign language sentence and the lip reading sentence, so that the sign language and the lip reading complement each other, and further improve the recognition effect. Through the above means, the sentence output at the end is smoother and more accurate.

[0060] When determining the preset weights (gesture weights and lip reading weights), the present invention uses the TF-IDF algorithm to calculate the similarities between the gesture sentence and the lip reading sentence and the correct sentence, and adjusts and finally determines the values ​​of the two weights accordingly. This is more accurate than the conventional method of setting by experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] Figure 1 1 is a schematic diagram of the main steps of an embodiment of a method for sign language recognition using AR glasses in the present invention;

[0062] Figure 2 It is a schematic diagram of the main steps of joint training of a dual-gate LSTM model and a Transformer model in an embodiment of the present invention;

[0063] Figure 3 It is a schematic diagram of the main steps of predetermining the gesture weight and the lip reading weight in an embodiment of the present invention. DETAILED DESCRIPTION

[0064] The preferred embodiments of the present invention are described below with reference to the accompanying drawings. It should be understood by those skilled in the art that these embodiments are only used to explain the technical principles of the present invention and are not intended to limit the protection scope of the present invention.

[0065] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0066] It should be noted that, in the description of the present invention, the terms "first" and "second" are only for the convenience of description, and do not indicate or imply the relative importance of the devices, elements or parameters, and therefore cannot be understood as limiting the present invention. In addition, the term "and / or" in the present invention is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article, unless otherwise specified, generally indicates that the associated objects before and after are in an "or" relationship.

[0067] Figure 1 Schematic diagram of the main steps of the method for sign language recognition using AR glasses in the present invention. Figure 1 As shown, the method of this embodiment includes steps A10-A50:

[0068] Step A10: Acquire the communication video through the AR glasses.

[0069] Specifically, step A10 may further include steps A11-A14:

[0070] Step A11: If blur, ghosting or afterimage exists in the current image, the current image is discarded and a new frame is acquired.

[0071] Step A12: If the brightness of the current image is lower than the preset brightness range, turn on the fill light of the AR glasses.

[0072] Step A13: if the brightness of the current image is higher than the preset brightness range, adjust the brightness of the current image to the preset brightness range.

[0073] Step A14: if the resolution of the current image is higher than the preset resolution range, compress the current image to the preset resolution range.

[0074] Step A20, input the communication video into the first YOLOv8 model to detect whether there is lip reading. If there is lip reading, go to step A30; otherwise, go to step A40.

[0075] Step A30, respectively generate gesture sentences and lip reading sentences according to the communication video and combine them into recognition results, and then go to step A50.

[0076] Specifically, step A30 may further include steps A31-A37:

[0077] Step A31: Input the communication video into the second YOLOv8 model to obtain the atomic action sequence and the corresponding first discrete sentence and the gesture confidence mean. The gesture confidence mean represents the recognition accuracy mean of the current gesture sequence.

[0078] Step A32, inputting the atomic action sequence and the first discrete sentence into the dual-gate LSTM model to generate a second discrete sentence corresponding to the gesture.

[0079] Step A33: Input the communication video into the first YOLOv8 model to obtain the third discrete sentence corresponding to the lip reading and the average lip reading confidence. The average lip reading confidence represents the average recognition accuracy of the current lip reading sequence.

[0080] Step A34, input the second discrete sentence into the first Transformer model, so that the first Transformer model can semantically understand the second discrete sentence in combination with the context, and then supplement or adjust the missing or incoherent parts to obtain the gesture sentence.

[0081] When encountering missing words, the Transformer model uses a self-attention mechanism to capture the relationship between different positions in the sequence. At the missing position, the model can focus on words at other positions to predict the missing words, and the self-attention mechanism allows the Transformer model to focus on different positions in the sequence when processing the input sequence, thereby capturing long-distance dependencies in the language. This enables the model to find key information in the sentence and correctly establish semantic connections through the self-attention mechanism when processing sentences with incorrect word order. By adding position encoding to the input sequence, the Transformer model can learn semantic information at different positions. Position encoding can provide the model with information about the relative position of words or tokens in the input, and typos are also fine-tuned in the same way.

[0082] Step A35, input the third discrete sentence into the second Transformer model, so that the second Transformer model can perform semantic understanding on the third discrete sentence in combination with the context, and then supplement or adjust the missing or incoherent parts to obtain the lip reading sentence.

[0083] Step A36, combining the gesture sentence and the lip reading sentence according to the gesture confidence mean, the lip reading confidence mean and the preset weights to obtain a recognition result.

[0084] In this embodiment, the preset weights include: a gesture weight and a lip reading weight, and the sum of the gesture weight and the lip reading weight is 1.

[0085] Specifically, step A36 may be: if the product of the gesture confidence mean and the gesture weight is greater than or equal to the product of the lip reading confidence mean and the lip reading weight, the gesture sentence is taken as the recognition result; otherwise, the lip reading sentence is taken as the recognition result.

[0086] In step A37, the recognition results are fed back to the first Transformer model and the second Transformer model respectively, and integrated into the context to participate in subsequent semantic understanding.

[0087] The core component of Transformer is the self-attention mechanism, which can capture the relationship between elements in the input sequence, including syntactic structure, semantic information, and contextual information. The self-attention mechanism can discover the interdependence between different elements, thereby helping the model to better understand the structure and semantic information of the text. Multi-head attention maps the input data to different subspaces for self-attention calculations and concatenates the results. This allows the model to learn different types of interdependencies in the input sequence in parallel. Specifically, the input sequence is linearly mapped to different subspaces, self-attention operations are performed separately in each subspace, and then the outputs of each subspace are concatenated into the final output.

[0088] Step A40, generating a gesture sentence as a recognition result according to the communication video.

[0089] Specifically, step A40 may further include steps A41-A44:

[0090] Step A41: input the communication video into the second YOLOv8 model to obtain an atomic action sequence and a corresponding first discrete sentence.

[0091] Step A42, inputting the atomic action sequence and the first discrete sentence into the dual-gate LSTM model to generate a second discrete sentence corresponding to the gesture.

[0092] Step A43, input the second discrete sentence into the first Transformer model, so that the first Transformer model performs semantic understanding of the second discrete sentence in combination with the context, and then supplements or adjusts the missing or incoherent parts to obtain a recognition result.

[0093] Step A44, feeding back the recognition result to the first Transformer model, integrating it into the context to participate in subsequent semantic understanding.

[0094] Step A50, display the recognition result in text form on the AR glasses or play it in voice form.

[0095] In this embodiment, the first YOLOv8 model is pre-trained using a first training set containing lip reading videos; the second YOLOv8 model is pre-trained using a second training set containing gesture videos.

[0096] In this embodiment, the two-gate LSTM model includes: 2 input gates, 2 forget gates and 1 output gate.

[0097] Figure 2 Schematic diagram of the main steps of joint training of the dual-gate LSTM model and the Transformer model in the embodiment of the present invention. Figure 2 As shown, the joint training method of this embodiment includes steps B10-B60:

[0098] Step B10, constructing a third training set, wherein the third training set includes: discrete sentences generated by gestures and discrete sentences generated by lip reading.

[0099] Step B20, selecting discrete sentences from the third training set and inputting them into the Transformer model to be trained.

[0100] Step B30, comparing the output result of the Transformer model to be trained with the true label and calculating the loss function.

[0101] Step B40, using the back propagation algorithm to update the parameters of the Transformer model to be trained.

[0102] Step B50, go to step B10, until the loss function no longer decreases, and obtain the trained Transformer model.

[0103] Step B60, copying the trained Transformer model as the first Transformer model and the second Transformer model.

[0104] In this embodiment, a Transformer model is trained, and two Transformer models are arranged in the recognition application, namely a first Transformer model and a second Transformer model, which are used to generate gesture sentences and lip reading sentences respectively.

[0105] Figure 3 FIG. 1 is a schematic diagram of the main steps of predetermining the weights of gestures and lip readings in an embodiment of the present invention. Figure 3 As shown, the determination method of this embodiment includes steps C10-C110:

[0106] Step C10, constructing a data set, wherein the data set includes video clips and corresponding correct sentences; the video clips also include gestures and lip readings.

[0107] In step C20, a video clip is selected from the data set each time.

[0108] Step C30: input the selected video clip into the second YOLOv8 model to obtain the sign language action sequence and the corresponding first discrete sentence.

[0109] Step C40, inputting the sign language action sequence and the first discrete sentence into the dual-gate LSTM model to generate a second discrete sentence corresponding to the gesture.

[0110] Step C50: input the selected video clip into the first YOLOv8 model to obtain a third discrete sentence corresponding to the lip reading.

[0111] Step C60: input the second discrete sentence and the third discrete sentence into the first Transformer model and the second Transformer model respectively to obtain a gesture sentence and a lip reading sentence respectively.

[0112] Step C70, using the TF-IDF algorithm to calculate the similarities between the gesture sentence and the lip reading sentence and the correct sentence, to obtain the sign language similarity and the lip reading similarity.

[0113] Step C80: normalize the sign language similarity and the lip reading similarity respectively.

[0114] Step C90, if the sign language similarity is greater than or equal to the lip reading similarity, the value of the first weight is increased by a preset step size and the value of the second weight is decreased by a preset step size; otherwise, the value of the first weight is decreased by a preset step size and the value of the second weight is increased by a preset step size.

[0115] Step C100, go to step C20, until all the video clips in the data set are used up. In this embodiment, the data set constructed in step C10 includes dozens of video clips, and when all the video clips are used up, the two weight values ​​will basically tend to be stable.

[0116] Step C110, setting the values ​​of the first weight and the second weight to the values ​​of the gesture weight and the lip reading weight respectively.

[0117] In this embodiment, the preset step size is 0.001; the initial values ​​of the first weight and the second weight are both 0.5; the first YOLOv8 model, the second YOLOv8 model, the dual-gate LSTM model, the first Transformer model and the second Transformer model used here are all trained models.

[0118] Although the various steps in the above embodiment are described in the above-mentioned order, those skilled in the art can understand that in order to achieve the effect of this embodiment, different steps do not have to be executed in such an order. They can be executed simultaneously (in parallel) or in a reverse order. These simple changes are within the scope of protection of the present invention.

[0119] Based on the above method embodiment, the present invention further provides an embodiment of AR glasses. The AR glasses of this embodiment perform sign language recognition according to the above method.

[0120] The present invention also provides an embodiment of a computer-readable storage medium. The storage medium of this embodiment stores a computer program that can be loaded by a processor and execute the method described above.

[0121] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.

[0122] Those skilled in the art should be able to appreciate that the method steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of electronic hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in electronic hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.

[0123] So far, the technical solutions of the present invention have been described in conjunction with the preferred embodiments shown in the accompanying drawings. However, it is easy for those skilled in the art to understand that the protection scope of the present invention is obviously not limited to these specific embodiments. Without departing from the principle of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions after these changes or substitutions will fall within the protection scope of the present invention.

Claims

1. A method for sign language recognition using AR glasses, characterized in that: The method comprises: Get the communication video through AR glasses; Input the communication video into a first YOLOv8 model to detect whether there is lip movement; If lip reading exists, respectively generating gesture sentences and lip reading sentences according to the communication video and combining them into a recognition result; If there is no lip reading, generating a gesture sentence as a recognition result according to the communication video; Displaying the recognition result in the form of text on the AR glasses or playing it in the form of voice; in, The step of "generating gesture sentences and lip language sentences respectively according to the communication video and combining them into recognition results" includes: Input the communication video into a second YOLOv8 model to obtain an atomic action sequence and a corresponding first discrete sentence and a gesture confidence mean; Inputting the atomic action sequence and the first discrete sentence into a two-gate LSTM model to generate a second discrete sentence corresponding to the gesture; Input the communication video into the first YOLOv8 model to obtain a third discrete sentence corresponding to the lip reading and a mean lip reading confidence value; Inputting the second discrete sentence into the first Transformer model, so that the first Transformer model performs semantic understanding on the second discrete sentence in combination with the context, and then supplements or adjusts the missing or incoherent parts to obtain the gesture sentence; Inputting the third discrete sentence into the second Transformer model, so that the second Transformer model performs semantic understanding on the third discrete sentence in combination with the context, and then supplements or adjusts the missing or incoherent parts to obtain the lip reading sentence; Combining the gesture sentence and the lip reading sentence according to the gesture confidence mean, the lip reading confidence mean and a preset weight to obtain a recognition result; The recognition results are fed back to the first Transformer model and the second Transformer model respectively, and integrated into the context to participate in subsequent semantic understanding.

2. The method for sign language recognition using AR glasses according to claim 1, characterized in that: The preset weight includes a gesture weight and a lip reading weight, and the sum of the gesture weight and the lip reading weight is 1; The step of "combining the gesture sentence and the lip reading sentence according to the gesture confidence mean, the lip reading confidence mean and a preset weight to obtain a recognition result" includes: If the product of the gesture confidence mean and the gesture weight is greater than or equal to the product of the lip reading confidence mean and the lip reading weight, the gesture sentence is taken as the recognition result; otherwise, the lip reading sentence is taken as the recognition result.

3. The method for sign language recognition using AR glasses according to claim 1, characterized in that: The step of "generating a gesture sentence as a recognition result according to the communication video" includes: Input the communication video into a second YOLOv8 model to obtain an atomic action sequence and a corresponding first discrete sentence; Inputting the atomic action sequence and the first discrete sentence into a two-gate LSTM model to generate a second discrete sentence corresponding to the gesture; Inputting the second discrete sentence into the first Transformer model, so that the first Transformer model performs semantic understanding on the second discrete sentence in combination with the context, and then supplements or adjusts the missing or incoherent parts to obtain a recognition result; The recognition result is fed back to the first Transformer model and integrated into the context to participate in subsequent semantic understanding.

4. The method for sign language recognition using AR glasses according to any one of claims 1 to 3, characterized in that: The steps of "obtaining communication video through AR glasses" include: If blur, ghosting or residual image exists in the current image, the current image is discarded and a new frame is acquired; If the brightness of the current image is lower than a preset brightness range, turning on the fill light of the AR glasses; If the brightness of the current image is higher than the preset brightness range, adjusting the brightness of the current image to the preset brightness range; If the resolution of the current image is higher than the preset resolution range, the current image is compressed to the preset resolution range.

5. The method for sign language recognition using AR glasses according to claim 1, characterized in that: The first YOLOv8 model is pre-trained using a first training set containing lip reading videos; The second YOLOv8 model is pre-trained using a second training set containing gesture videos.

6. The method for sign language recognition using AR glasses according to claim 1, characterized in that: The dual-gate LSTM model includes: 2 input gates, 2 forget gates and 1 output gate; The training method of the first Transformer model and the second Transformer model includes: Constructing a third training set, wherein the third training set includes: discrete sentences generated by gestures and discrete sentences generated by lip reading; Selecting discrete sentences from the third training set and inputting them into the Transformer model to be trained; Compare the output result of the Transformer model to be trained with the true label and calculate the loss function; Using a back propagation algorithm to update the parameters of the Transformer model to be trained; Repeat the process until the loss function is minimized, and a trained Transformer model is obtained. The trained Transformer model is copied as the first Transformer model and the second Transformer model.

7. The method for sign language recognition using AR glasses according to claim 2, characterized in that: The method for determining the gesture weight and the lip reading weight includes: Constructing a data set, wherein the data set includes a video clip and a corresponding correct sentence; the video clip also includes gestures and lip readings; Selecting a video clip from the data set each time; Input the selected video clip into the second YOLOv8 model to obtain a sign language action sequence and a corresponding first discrete sentence; Inputting the sign language action sequence and the first discrete sentence into a two-gate LSTM model to generate a second discrete sentence corresponding to the gesture; Input the selected video clip into the first YOLOv8 model to obtain a third discrete sentence corresponding to the lip reading; Inputting the second discrete sentence and the third discrete sentence into the first Transformer model and the second Transformer model respectively to obtain a gesture sentence and a lip reading sentence respectively; Calculate the similarity between the gesture sentence and the lip reading sentence and the correct sentence respectively by using the TF-IDF algorithm to obtain the sign language similarity and the lip reading similarity; Normalizing the sign language similarity and the lip reading similarity respectively; If the sign language similarity is greater than or equal to the lip reading similarity, the value of the first weight is increased by a preset step size and the value of the second weight is decreased by the preset step size; otherwise, the value of the first weight is decreased by the preset step size and the value of the second weight is increased by the preset step size; Repeating the steps of selecting video clips, calculating similarities, and adjusting weights until all video clips in the data set are used up; Setting the values ​​of the first weight and the second weight to the values ​​of the gesture weight and the lip reading weight respectively; The initial values ​​of the first weight and the second weight are both 0.

5.

8. An AR glasses, characterized in that: The AR glasses perform sign language recognition according to the method as described in any one of claims 1-7.

9. A computer-readable storage medium, characterized in that: A computer program is stored which can be loaded by a processor and execute the method according to any one of claims 1 to 7.