Lightweight BiLSTM-based sign language recognition system and application

By adopting lightweight BiLSTM network, knowledge distillation and model pruning technologies in the sign language recognition system, the problem that existing sign language recognition technology cannot adapt to the performance requirements of embedded devices is solved, efficient and accurate sign language recognition is achieved, and the deployment flexibility of the equipment is improved.

CN120220236APending Publication Date: 2025-06-27NANJING UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510342147.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Due to the high computational complexity, existing sign language recognition technology cannot adapt to the performance requirements of embedded devices, which limits its application scope.

Method used

The sign language recognition system based on lightweight BiLSTM is adopted to reduce the computational complexity and storage requirements through knowledge distillation and model pruning, so that the model can run efficiently on embedded devices.

Benefits of technology

It realizes the performance requirements of embedded hardware while maintaining high recognition accuracy, reducing computing pressure and storage requirements, and improving inference speed and device deployment flexibility.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220236A_ABST
    Figure CN120220236A_ABST
Patent Text Reader

Abstract

The invention discloses a sign language recognition system based on lightweight BiLSTM and application thereof, and the system comprises a skeleton point coordinate extraction module 100 which is used for extracting upper limb skeleton point coordinate data in a sign language video image; the recognition module 200 is used for acquiring a sign language recognition result according to the coordinates of the upper limb skeleton points of the continuous image frames; the recognition module inputs a sequence (p1,..., pK-1, pK) formed by upper limb skeleton point coordinates in continuous K frames of video images based on a lightweight BiLSTM network, and outputs vocabularies corresponding to input sign language videos in a vocabulary list; pk is an upper limb skeleton point coordinate in the kth frame of video image, k = 1, 2,..., K; the value of K is the number of key frames of a single sign language vocabulary; and the large language model 300 is used for generating reasonable statements according to the plurality of recognized vocabularies. The system can meet the performance requirements of embedded hardware while keeping high recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sign language recognition system based on lightweight BiLSTM and the application of this system, belonging to the technical field of continuous sign language recognition. Background Art

[0002] With the progress of society and the development of technology, sign language, as the main means of communication for the deaf, has received increasing attention. At present, most sign language recognition technologies rely on video image processing and use computer vision algorithms to recognize gestures or hand movements. The Chinese patent application with the application number 2023117478912 discloses a dynamic gesture recognition method based on hand key points and a double-layer bidirectional LSTM network. This recognition technology requires building a relatively large model, has high requirements for computing resources, and is not suitable for deployment on embedded devices. Existing sign language recognition technologies usually adopt complex image processing algorithms, with a large amount of calculation, requiring high-performance hardware, resulting in high energy consumption and being unable to be applied to portable devices, thus limiting the scope of use. Summary of the Invention

[0003] Object of the Invention: The technical problem to be solved by the present invention is to provide a sign language recognition system based on lightweight BiLSTM in view of the deficiencies of the prior art. While maintaining a high recognition accuracy, this system can adapt to the performance requirements of embedded hardware.

[0004] To solve the above technical problem, the present invention discloses a sign language recognition system based on lightweight BiLSTM, including:

[0005] A skeletal point coordinate extraction module 100, which is used to extract the coordinate data of the upper limb skeletal points in the sign language video image;

[0006] A recognition module 200, which is used to obtain the sign language recognition result according to the coordinate of the upper limb skeletal points of consecutive image frames;

[0007] The recognition module is based on a lightweight BiLSTM network. The input is a sequence (p1, …, p K-1 , p K ) composed of the coordinates of the upper limb skeletal points in K consecutive video images, and the output is the vocabulary in the vocabulary table corresponding to the input sign language video; p1 is the coordinate of the upper limb skeletal points in the first video image, and p k is the coordinate of the upper limb skeletal points in the kth video image, where k = 1, 2, …, K; the value of K is the number of key frames of a single sign language vocabulary;

[0008] A large language model 300, which is used to generate reasonable sentences according to multiple recognized vocabularies.

[0009] Furthermore, the construction and training of the recognition module include the steps:

[0010] S1. Establish a student model and a teacher model. The recognition module is built based on the student model;

[0011] The teacher model includes an input layer, two hidden layers, and an output layer. The input layer is used to receive the coordinate sequence of upper limb bone points in continuous K-frame video images and pass it to the hidden layer of the teacher model. The hidden layer includes two LSTM layers. Each LSTM layer includes a forward LSTM layer and a backward LSTM layer. Both the forward LSTM layer and the backward LSTM layer include 1024 neurons. The output layer adopts a fully connected layer to output the logits obtained by the teacher model according to the input coordinate sequence of upper limb bone points. The student model has the same overall structure as the student model, including an input layer, two hidden layers, and an output layer. The input layer is used to receive the coordinate sequence of upper limb bone points in continuous K-frame video images and pass it to the hidden layer. The hidden layer includes two LSTM layers. Each LSTM layer includes a forward LSTM layer and a backward LSTM layer. Both the forward LSTM layer and the backward LSTM layer include 256 neurons. The output layer adopts a fully connected layer to output the logits obtained by the student model according to the input coordinate sequence of upper limb bone points.

[0012] S2. Use the sign language video sample set to train the teacher model, and adopt the cross-entropy loss function (Cross-Entropy Loss) as the optimization objective to improve the recognition ability of the teacher model for sign language sequences. After training, freeze all the parameters of the teacher model to ensure its stable knowledge representation during the subsequent training process of the student model.

[0013] S3. Input the samples in the sign language video sample set into the trained teacher model, and extract the logits generated by the teacher model for each sample as soft labels (Soft Targets) to guide the knowledge distillation training of the student model.

[0014] S4. Under the guidance of the soft labels provided by the teacher model, train the student model. During the training process, the goal of the student model is not only to optimize its own classification performance through the cross-entropy loss function, but also to minimize the gap between its output distribution and the output of the teacher model through the distillation loss, so as to effectively inherit the knowledge of the teacher model.

[0015] Train the student model using the soft labels provided by the teacher model. The loss function L of the training sum is the weighted sum of the cross-entropy loss L KD and the distillation loss L CE :

[0016] L sum = αL CE +(1 - α)LKD

[0017] α is a weight parameter, where 0 < α < 1;

[0018] Among them, the cross-entropy loss L KD is:

[0019]

[0020] N is the total number of samples in the sign language video sample set, y i is the true label of the i-th sample, is the predicted probability of the student model for the i-th sample, Z s,i is the logits obtained by the student model for the i-th sample;

[0021] The distillation loss L CE is:

[0022]

[0023] Among them, S teacher,i and S student,i are the temperature softmax values of the teacher model and the student model for the i-th sample respectively; S teacher,i = softmax(Z t,i / T), Z t,i is the logits obtained by the teacher model for the i-th sample, and T is the preset distillation temperature; S student,i = softmax(Z s,i / T);

[0024] Further, after the training of the student model is completed, it also includes pruning the student model, and using the pruned student model as the recognition module. The model pruning includes: pruning 20% of the weights with the smallest absolute values in the student model.

[0025] Further, after the pruning of the student model is completed, the trained student model is quantized; the quantized student model, the softmax function, and the vocabulary determination module form the recognition module; the softmax function converts the output of the student model into a predicted probability, and the vocabulary determination module is used to select the vocabulary with the largest predicted probability as the recognition result of the recognition module.

[0026] Further, the skeleton point coordinate extraction module 100 uses the Mediapipe architecture to perform human skeleton point recognition and extract the upper limb skeleton point coordinates in the video image.

[0027] Further, the upper limb skeleton points include: left shoulder coordinates, right shoulder coordinates, left elbow coordinates, right elbow coordinates, 21 left hand key points, and 21 right hand key points.

[0028] Further, the skeletal point coordinate extraction module 100 first preprocesses the sign language video image and then extracts the upper limb skeletal point coordinates in the image; the preprocessing includes: denoising and image segmentation; and normalizes the extracted upper limb skeletal point coordinates.

[0029] Further, the skeletal point coordinate extraction module 100 and the recognition module 200 are implemented based on the embedded system platform RDKX5.

[0030] On the other hand, the present invention also discloses a sign language dialogue system, including:

[0031] A video acquisition module for acquiring sign language videos;

[0032] The above-mentioned sign language recognition system based on lightweight BiLSTM is used to recognize the vocabulary in the acquired sign language video and generate sentences.

[0033] Further, the above-mentioned sign language dialogue system further includes a sound output device for converting the sentences generated by the sign language recognition system into speech and playing them.

[0034] Further, the above-mentioned sign language dialogue system further includes a display device for displaying the sentences generated by the sign language recognition system.

[0035] Beneficial effects: Compared with the prior art, the sign language recognition system based on lightweight BiLSTM disclosed in the present invention reduces the computational complexity and storage requirements through knowledge distillation and model pruning, enabling the model to operate efficiently on embedded devices or edge computing platforms. At the same time, a BiLSTM neural network of an appropriate scale is used to reduce redundant calculations while ensuring accuracy and improve the inference speed. In addition, the sign language recognition system constructed by combining the lightweight BiLSTM recognition model with a large prediction model processes the recognition of isolated sign language vocabulary and the subsequent task of connecting words into sentences separately. The edge device can focus on the sign language recognition part, while the sentence generation depends on an efficient language model in the cloud or locally, reducing the computational pressure of local deployment and improving the overall efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0037] Figure 1 It is a schematic diagram of the composition of the sign language recognition system based on lightweight BiLSTM disclosed in Embodiment 1;

[0038] Figure 2 It is a structural block diagram of the recognition module in Embodiment 1;

[0039] Figure 3 It is the structural block diagram of the student model in Embodiment 1;

[0040] Figure 4 It is the composition schematic diagram of the sign language dialogue system disclosed in Embodiment 2;

[0041] Figure 5 It is the application scenario schematic diagram of the sign language dialogue system disclosed in Embodiment 2. Specific implementation manners

[0042] Embodiment 1

[0043] This embodiment discloses a sign language recognition system based on lightweight BiLSTM, as Figure 1 shown, including:

[0044] A skeletal point coordinate extraction module 100, which is used to extract the upper limb skeletal point coordinate data in the sign language video image; in this embodiment, the upper limb skeletal points include: left shoulder coordinates, right shoulder coordinates, left elbow coordinates, right elbow coordinates, 21 left hand key points, and 21 right hand key points; a total of 46 points of coordinates. In order to remove the influence of image size, clarity, etc. on the recognition result, the sign language video image is preprocessed first, and then the upper limb skeletal point coordinates in the image are extracted; the preprocessing includes: denoising and image segmentation, and in this embodiment, OpenCV is used to preprocess the image; in the preprocessed image, the Mediapipe architecture is used to recognize human skeletal points, extract the upper limb skeletal point coordinates in the video image, and standardize the extracted upper limb skeletal point coordinates to the [0,1] interval to overcome the influence of different sizes.

[0045] A recognition module 200, which is used to obtain the sign language recognition result according to the upper limb skeletal point coordinates of consecutive image frames;

[0046] The recognition module is based on a lightweight BiLSTM network, and the input is a sequence (p1,…,p K-1 ,p K ) composed of the upper limb skeletal point coordinates in K consecutive video images, and the output is the vocabulary corresponding to the input sign language video in the vocabulary; p1 is the upper limb skeletal point coordinates in the first frame of the video image, p kLet \(x_k\) be the coordinates of the upper limb bone points in the \(k\)-th frame of the video image, where \(k = 1, 2, \ldots, K\); the value of \(K\) is the number of key frames of a single sign language word. The value of \(K\) is related to the video frame rate. \(K\) consecutive frames in the video stream can completely represent a single sign language word. For a video stream with a relatively high frame rate, it is not necessary to use all the image frames of a sign language word for calculation. To further save computing resources, the following steps are adopted in this embodiment to extract key frames: For the video stream of a sign language word, traverse each frame of the image in turn, calculate the absolute pixel difference between each frame and the previous frame, and select the \(K\) frames with the most obvious changes as key frames (i.e., sort in descending order according to the diff value), and take out the \(K\) image frames with the largest changes as key frames. In this embodiment, \(K = 36\), and 36 is determined by experiments and can be adjusted according to different software and hardware systems. Extracting key frames is only to reduce the amount of calculation, and a reasonable number of key frames will not affect the final recognition accuracy.

[0047] The large language model 300 is used to generate reasonable sentences according to multiple recognized words.

[0048] The recognition module 200 adopts a lightweight BiLSTM network, and its structure is as Figure 2 shown, and its construction and training include the steps:

[0049] S1. Establish a teacher model and a student model, and the recognition module is built based on the student model;

[0050] The teacher model includes an input layer, two hidden layers, and an output layer. The input layer is used to receive the sequence of coordinates of the upper limb bone points in \(K\) consecutive frames of the video image and pass it to the hidden layer of the teacher model; the hidden layer includes two LSTM layers, and each LSTM layer includes a forward LSTM layer and a backward LSTM layer. Both the forward LSTM layer and the backward LSTM layer include 1024 neurons; the output layer adopts a fully connected layer to output the logits obtained by the teacher model according to the input sequence of upper limb bone point coordinates (logits is the last layer of the neural network, usually the raw output of the fully connected layer, and the unnormalized score);

[0051] The student model is consistent with the overall structure of the student model, including an input layer, two hidden layers, and an output layer. The input layer is used to receive the sequence of coordinates of the upper limb bone points in \(K\) consecutive frames of the video image and pass it to the hidden layer; the hidden layer includes two LSTM layers, and each LSTM layer includes a forward LSTM layer and a backward LSTM layer. Both the forward LSTM layer and the backward LSTM layer include 256 neurons; the output layer adopts a fully connected layer to output the logits obtained by the student model according to the input sequence of upper limb bone point coordinates.

[0052] S2. Use the sign language video sample set to train the teacher model, and adopt the cross-entropy loss function (CrossEntropy Loss) as the optimization objective to improve the recognition ability of the teacher model for sign language sequences; after the training is completed, freeze all the parameters of the teacher model to ensure its stable knowledge representation during the subsequent training process of the student model;

[0053] In this embodiment, CSL-500 provided by the University of Science and Technology of China is used as the sign language video sample set, which includes 500 common Chinese sign language words, covering most sign language expressions in daily communication, that is, the number of recognition categories is 500. Before training, the labels need to be converted into one-hot encoding format to adapt to the classification task.

[0054] In this embodiment, the training loss function L of the teacher model teacher Adopt cross-entropy loss (Cross-EntropyLoss):

[0055]

[0056] where N is the total number of samples in the sign language video sample set, y i is the true label of the i-th sample, is the predicted probability of the teacher model for the i-th sample, Z t,i is the logits obtained by the teacher model for the i-th sample.

[0057] In the training, the optimizer adopts Adam, the initial value of the learning rate is 0.01, and cosine annealing decay is adopted. The training batch size Batch Size value is set to 256; the number of training epochs Epochs value is 100; the training sampling early stopping strategy (Early Stopping) is adopted, that is, if the validation set accuracy does not improve within 10 epochs, the training will be stopped in advance.

[0058] The training process is to read the preprocessed skeleton point data; send it into the BiLSTM model for forward propagation to calculate the predicted probability; calculate the cross-entropy loss and update the weights using the Adam optimizer; after the training is completed, freeze the parameters of the teacher model for guiding the training of the student model.

[0059] S3. Input the samples in the sign language video sample set into the trained teacher model, and obtain the logits produced by the teacher model for each sample as soft labels (Soft Targets) for guiding the knowledge distillation training of the student model;

[0060] S4. Under the guidance of the soft labels provided by the teacher model, train the student model; during the training process, the goal of the student model is not only to optimize its own classification performance through the cross-entropy loss function, but also to minimize the gap between its output distribution and the output of the teacher model through the distillation loss, so as to effectively inherit the knowledge of the teacher model.

[0061] Train the student model using the soft labels provided by the teacher model, and the loss function L of the training sum is the cross-entropy loss L KD and the distillation loss L CE weighted sum of:

[0062] L sum = αL CE +(1 - α)L KD

[0063] α is the weight parameter, 0 < α < 1, and is set to 0.5 in this embodiment; where the cross-entropy loss L KD is:

[0064]

[0065] N is the total number of samples in the sign language video sample set, y i is the true label of the i-th sample, is the predicted probability of the student model for the i-th sample, Z s,i is the logits obtained by the student model for the i-th sample;

[0066] The distillation loss L CE is:

[0067]

[0068] where S teacher,i and S student,i are the temperature softmax values of the teacher model and the student model for the i-th sample respectively; S teacher,i = softmax(Z t,i / T), Z t,i is the logits obtained by the teacher model for the i-th sample, T is the preset distillation temperature, and the value of T is set to 3 in this embodiment; S student,i = softmax(Z s,i / T);

[0069] The structures of the teacher model and the student model are roughly the same. The difference is that the number of neurons in each LSTM layer of the teacher model is 1024, while the number of neurons in each LSTM layer of the student model is 256. For neural networks with roughly the same structure, having more neurons means better feature extraction effect and better recognition results, but it also requires more parameters to be trained, and running this model has higher requirements for hardware. Each neural network unit has 4 groups of weights, which are used for the input gate (inputgate, i), the forget gate (forget gate, f), the cell state update (cell gate, g), and the output gate (output gate, o). Therefore, for the teacher model, among the 1024 neurons in each LSTM layer, there are 4096 groups of weights, while the student model has 1024 groups of weights in each LSTM layer, which is more suitable for running on embedded devices. The present invention adopts the knowledge distillation technology, and uses the output result of the teacher model with better recognition effect as the soft label to guide the training of the student model, so that the recognition effect of the smaller-scale student model is close to that of the larger-scale teacher model.

[0070] In the present invention, in order to further optimize the inference speed of the recognition model on the embedded device, model pruning is performed on the trained student model, and 20% of the weights with the smallest absolute values in the student model are pruned to further reduce the model size.

[0071] In the present invention, the pruned student model is quantized; the quantized student model, the softmax function, and the vocabulary determination module constitute the recognition module; the softmax function converts the output of the student model into prediction probabilities, and the vocabulary determination module is used to select the vocabulary with the largest prediction probability as the recognition result of the recognition module.

[0072] In this embodiment, the skeleton point coordinate extraction module 100 and the recognition module 200 are implemented based on the embedded system platform RDK X5. Horizon RDK X5 is equipped with an efficient BPU (Neural Network Processing Unit), which provides dedicated hardware support for AI inference tasks, can significantly accelerate the processing speed of the Chinese sign language recognition task, and reduce the power consumption of the system.

[0073] The pruned student model is quantized with 8 bits. First, the model trained under the Pytorch framework is converted to the ONNX framework, and then the Horizon toolchain hb_mapper is used for mapping and transplantation to run on the BPU.

[0074] In this embodiment, the teacher model and the student model are used as recognition models for comparison. The size of the teacher model is 300MB, and the pruned student model is only 70MB, with a significantly reduced scale and a 60% reduction in storage requirements, making it suitable for resource-constrained embedded systems. Compared with the teacher model, the pruned student model has significantly fewer parameters, lower computational complexity, and a reduced inference time of approximately 30%, significantly improving the inference speed.

[0075] Sign language recognition is usually a multi-class classification task. In this embodiment, Recall and F1-score are compared. The recall rate and F1-score of the teacher model are 89.8% and 89.4% respectively, and the test recall rate and F1-score of the pruned student model drop to 84.3% and 84.2%. However, it should be noted that the subsequent large language model has a certain error correction ability, which performs semantic reasoning and automatic correction on individual misrecognized sign language words, thereby improving the accuracy and readability of the final output. In the actual context, the proportion of the final long sentence output that is consistent with the user's actual sign language input is measured to be 95.4%. It can be found that the large language model can effectively repair the model inference errors in sign language recognition and has stronger error correction ability in long sentences and specific contexts.

[0076] The pruned and optimized student model can run smoothly on an embedded platform of 10Tops RDK X5 without the need for a high-performance GPU, enhancing the deployment flexibility of the device.

[0077] Embodiment 2

[0078] This embodiment discloses a sign language dialogue system, as Figure 4 shown, including:

[0079] A video acquisition module for acquiring sign language videos; in this embodiment, an RGB camera (USB-CAM or MIPI-CAM) is used to acquire sign language image data or skeleton point data. The video stream is input into the skeleton point coordinate extraction module through a USB 3.0 interface, with a frame rate of 30FPS and an image resolution of 1280×720 (16:9, SAR 1:1).

[0080] The sign language recognition system based on lightweight BiLSTM disclosed in Embodiment 1 is used to recognize the words in the acquired sign language videos and generate reasonable sentences.

[0081] The sentences generated by the sign language recognition system can be output through a sound output device or a display device. The sound output device converts the sentences generated by the large language model into speech and plays them, and the display device is used to display the sentences generated by the large language model, so that the conversation between deaf-mute people and non-deaf-mute people can be realized without relying on typing, reducing the time delay and risk of misunderstanding during communication. The two parties in the conversation can be in the same space, that is, on-site instant interaction, or in different spaces, that is, remote interaction. This conversation system can be applied to scenarios such as public services, special education, medical services, and smart homes, enhancing the sense of participation and social convenience of deaf-mute people.

[0082] In addition, the sign language conversation system also includes a voice input device. The voice input device converts the voice of non-deaf-mute people into text and displays it on the display screen to show the deaf-mute people the conversation sentences of the other party. The display screen is connected to the RDK X5 development board through the HDMI interface. As Figure 5 shown, in a certain application scenario, the sign language conversation system disclosed in this embodiment converts sign language into text, constructs the text into sentences in the large language model of the cloud server, and displays them to able-bodied people through the interactive display screen. The conversation of able-bodied people is converted into text through the voice-to-text embedded system and displayed to deaf-mute people on the interactive display screen, thus realizing the conversation between the two parties.

[0083] The large language model uses deepseek-r1:7b-Ollama. When using it, the corresponding prompt is provided, and the example is as follows: You are an expert in Chinese sign language translation, and you are good at translating sign language into accurate and fluent Chinese sentences. Please generate natural and grammatically correct complete Chinese sentences according to the provided Chinese sign language vocabulary. You need to pay attention to the logical coherence and grammatical correctness of the sentences, and at the same time maintain the accuracy and conciseness of the translation, so as to convert isolated sign language vocabulary into coherent Chinese sentences.

[0084] The present invention deploys a lightweight sign language recognition system and a large language model DeepSeek locally for the recognition of isolated sign language vocabulary and the processing of subsequent tasks of connecting words into sentences. The task of isolated vocabulary recognition is relatively simple and does not require complex context reasoning. Therefore, a lighter model can be adopted to reduce the difficulty of embedded deployment, enabling the system to operate efficiently on low-resource devices. The sentence construction is handed over to off-the-shelf large language models, which have been trained on a large scale and can efficiently and accurately complete Chinese grammar processing and sentence generation tasks, transferring the computational burden from the edge device to the cloud or large servers, reducing the local deployment computational pressure. Especially in an embedded system, the local recognition part uses a lightweight model, which can operate efficiently on devices with limited resources, reduce power consumption, and improve the system response speed. This solution can not only reduce the complexity of the sign language recognition project but also significantly reduce the computational overhead of the model and the deployment difficulty of the embedded system. The large language model requires the service provider to provide a server that communicates with the embedded device; however, the performance requirements for this server are extremely low, and an ordinary personal computer can be used instead; even a lower parameter version of 1.5b can also complete the task.

[0085] The present invention provides an idea and method for sign language recognition based on lightweight BiLSTM. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred implementation manner of the present invention. It should be noted that for those of ordinary skill in the art in this technical field, without departing from the principle of the present invention, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented using existing technologies.

Claims

1. A sign language recognition system based on lightweight BiLSTM, characterized by: include: A skeleton point coordinate extraction module (100) is used to extract upper limb skeleton point coordinate data in a sign language video image; A recognition module (200) is used to obtain a sign language recognition result according to the coordinates of the upper limb bone points in the continuous image frames; The recognition module is based on a lightweight BiLSTM network, and its input is a sequence of upper limb bone point coordinates in K consecutive video frames (p1,…,p K-1 ,p K ), the output is the vocabulary corresponding to the input sign language video; p1 is the coordinates of the upper limb bone point in the first frame of the video image, p k is the coordinate of the upper limb bone point in the k-th frame of the video image, k = 1, 2, ..., K; the value of K is the number of key frames of a single sign language vocabulary; A large language model (300) is used to generate reasonable sentences based on multiple recognized vocabularies.

2. The sign language recognition system based on lightweight BiLSTM according to claim 1, characterized in that: The construction and training of the recognition module includes the steps of: S1. Establish teacher model and student model, and the recognition module is built based on the student model; The teacher model includes an input layer, two hidden layers and an output layer, wherein the input layer is used to receive a sequence of upper limb bone point coordinates in a continuous K-frame video image and pass it to the teacher model hidden layer; the hidden layer includes two LSTM layers, each LSTM layer includes a forward LSTM layer and a reverse LSTM layer, and the forward LSTM layer and the reverse LSTM layer each include 1024 neurons; the output layer adopts a fully connected layer, and outputs the logits obtained by the teacher model according to the input sequence of upper limb bone point coordinates; The student model is consistent with the overall structure of the student model, including an input layer, two hidden layers and an output layer, wherein the input layer is used to receive a sequence of upper limb bone point coordinates in a continuous K-frame video image and pass it to the hidden layer; the hidden layer includes two LSTM layers, each LSTM layer includes a forward LSTM layer and a reverse LSTM layer, and the forward LSTM layer and the reverse LSTM layer each include 256 neurons; the output layer adopts a fully connected layer, and outputs the logits obtained by the student model according to the input sequence of upper limb bone point coordinates; S2. Use the sign language video sample set to train the teacher model, and use the cross entropy loss function as the optimization target to improve the teacher model's ability to recognize sign language sequences; after the training is completed, freeze all parameters of the teacher model; S3, input the samples in the sign language video sample set into the trained teacher model, extract the logits generated by the teacher model for each sample as soft labels, and use them to guide the knowledge distillation training of the student model; S4. Use the soft labels provided by the teacher model to train the student model. The training loss function L sum is the cross entropy loss L KD and distillation loss L CE The weighted sum of: L sum =αL CE +(1-α)L KD α is the weight parameter, 0<α<1; The cross entropy loss L KD for: N is the total number of samples in the sign language video sample set, y i is the true label of the i-th sample, is the predicted probability of the student model for the i-th sample, Z s,i The logits obtained by the student model for the i-th sample; Distillation loss L CE for: Among them, S teacher,i and S student,i are the temperature softmax values ​​of the teacher model and the student model for the i-th sample; S teacher,i =softmax(Z t,i / T), Z t,i is the logits obtained by the teacher model for the i-th sample, T is the preset distillation temperature; S student,i =softmax(z s,i / T); The trained student model is quantized; the quantized student model, the softmax function and the vocabulary determination module constitute a recognition module; the softmax function converts the output of the student model into a prediction probability, and the vocabulary determination module is used to select the vocabulary with the largest prediction probability as the recognition result of the recognition module.

3. The sign language recognition system based on lightweight BiLSTM according to claim 2, characterized in that: It also includes model pruning of the student model, which prunes the 20% of the student models with the smallest absolute weight values.

4. The sign language recognition system based on lightweight BiLSTM according to claim 1, characterized in that: The skeleton point coordinate extraction module (100) adopts the Mediapipe architecture to perform human skeleton point recognition and extract the coordinates of the upper limb skeleton points in the video image.

5. The sign language recognition system based on lightweight BiLSTM according to claim 1, characterized in that: The upper limb skeleton points include: left shoulder coordinates, right shoulder coordinates, left elbow coordinates, right elbow coordinates, 21 left hand key points, and 21 right hand key points.

6. The sign language recognition system based on lightweight BiLSTM according to claim 1, characterized in that: The skeleton point coordinate extraction module (100) first pre-processes the sign language video image, and then extracts the upper limb skeleton point coordinates in the image; the pre-processing includes: denoising, image segmentation; and standardizing the extracted upper limb skeleton point coordinates.

7. The sign language recognition system based on lightweight BiLSTM according to claim 1, characterized in that: The skeleton point coordinate extraction module (100) and the recognition module (200) are implemented based on the embedded system platform RDK X5.

8. A sign language dialogue system, characterized in that: include: Video acquisition module, used to acquire sign language videos; The sign language recognition system based on lightweight BiLSTM as described in claims 1-7 is used to recognize words in the collected sign language video and generate sentences.

9. The sign language dialogue system according to claim 8, further comprising a sound output device for converting the sentences generated by the sign language recognition system into voice and playing them.

10. The sign language dialogue system according to claim 8, further comprising a display device for displaying the sentences generated by the sign language recognition system.