Sign language translation methods, devices, computer equipment and storage media
By combining hand region detection, sign language recognition, and semantic understanding with adaptive recognition and grammar correction technologies in the text generation module, the problems of accuracy and naturalness in sign language translation are solved, achieving high-quality multilingual translation and personalized sign language communication.
Patent Information
- Application Number
- CN202511384637.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-26
- Publication Date
- 2026-01-06
- Estimated Expiration
- 2045-09-26
AI Technical Summary
Existing sign language translation technologies suffer from low accuracy and unnatural expression, mainly because they simplify sign language into a one-way sequence conversion task from "video to text," failing to effectively capture the visual-spatial-temporal linguistic features of sign language, resulting in semantic comprehension bias and word order disorder.
This approach combines a hand region detection network, a sign language recognition module, and a semantic understanding and text generation module. Through adaptive recognition and grammar correction techniques, including natural language understanding, grammar correction, multimodal sentiment analysis, and adaptive learning, it processes sign language videos to generate accurate and natural target language text.
It improves the accuracy and naturalness of sign language translation, supports bidirectional translation between multiple sign languages and spoken languages, enhances international communication capabilities across languages and hearing-impaired groups, adapts to different sign language styles, and provides personalized output.
Smart Images

Figure CN120877390B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of sign language translation technology, and in particular to a sign language translation method, apparatus, computer device and storage medium. Background Technology
[0002] As the primary natural means of communication for the hearing impaired, sign language's automatic recognition technology has significant application value in fields such as human-computer interaction and barrier-free communication.
[0003] Currently, most sign language translation technologies follow the model of speech recognition, simplifying sign language recognition into a one-way sequence conversion task from video to text. These methods typically treat sign language videos as a sequence of consecutive frames of images, directly mapping them to the corresponding text output through a deep learning model.
[0004] However, sign language is an independent visual-spatial-temporal language. Its expression relies on the coordination of multiple modalities such as gestures, facial expressions, body postures, and spatial trajectories. This makes there are profound linguistic differences between sign language and spoken language. Therefore, simply viewing sign language as a "video to text" sequence conversion task can easily lead to problems such as semantic comprehension bias, word order disorder, and missing references, which in turn affect the accuracy and naturalness of the translation. Summary of the Invention
[0005] This application provides a sign language translation method, apparatus, computer device, and storage medium, aiming to solve the problems of low accuracy and unnatural expression in traditional sign language translation methods.
[0006] In a first aspect, embodiments of this application provide a sign language translation method, the sign language translation method comprising:
[0007] Obtain the sign language video to be processed;
[0008] The sign language video to be processed is input into a preset hand region detection network to obtain the video stream output by the hand region detection network;
[0009] The video stream is input into a preset sign language recognition module for adaptive recognition processing to obtain a preliminary text representation output by the sign language recognition module;
[0010] The preliminary text representation is input into a preset semantic understanding and text generation module, and the semantic understanding and text generation module is used to perform grammatical correction on the preliminary text representation to obtain the target language text output by the semantic understanding and text generation module.
[0011] A further technical solution is that the semantic understanding and text generation module includes a natural language understanding submodule, a grammar correction submodule, and a text generation submodule. The step of using the semantic understanding and text generation module to perform grammar correction on the preliminary text representation to obtain the target language text output by the semantic understanding and text generation module includes:
[0012] The natural language understanding submodule is used to perform semantic analysis on the preliminary text representation to obtain a phrase representation.
[0013] The phrase representation is input into the grammar correction submodule for grammar correction processing to obtain the corrected text;
[0014] The corrected text is input into the text generation submodule for processing to obtain the target language text output by the text generation submodule.
[0015] A further technical solution is that the sign language recognition module is equipped with a trained deep neural network model. The step of inputting the video stream into the sign language recognition module for adaptive recognition processing to obtain the preliminary text representation output by the sign language recognition module includes:
[0016] The video stream is input into the trained deep neural network model;
[0017] The trained deep neural network model is used to preprocess the video stream to obtain a structured video stream;
[0018] The structured video stream is subjected to feature extraction to obtain multiple hand features;
[0019] By fusing multiple hand features, a preliminary text representation output by the trained deep neural network model is obtained.
[0020] A further technical solution is that the preprocessing of the video stream using the trained deep neural network model to obtain a structured video stream includes:
[0021] The video stream is processed into video frames to obtain a sequence of hand image frames;
[0022] The pose of the hand image frame sequence is estimated to obtain a key point coordinate sequence, wherein the key point coordinate sequence includes hand key points, facial key points and body key points;
[0023] The key point coordinate sequence is normalized to obtain a structured video stream.
[0024] A further technical solution is that the multiple hand features include spatial features and spatiotemporal features, and the structured video stream is subjected to feature extraction to obtain multiple hand features, including:
[0025] The spatial features are obtained by independently extracting static hand features from each frame of the structured video stream.
[0026] Hand dynamic features are extracted from multiple consecutive frames in the structured video stream to obtain spatiotemporal features.
[0027] The further technical solution is that the grammar correction process includes non-manual feature processing, spatial grammar processing, word order processing, morphological processing, and conciseness processing.
[0028] A further technical solution is that the sign language recognition module adopts a semi-supervised learning mechanism of pseudo-labels and / or knowledge distillation.
[0029] Secondly, embodiments of this application also provide a sign language translation device, which includes a unit for performing the above-described method.
[0030] Thirdly, embodiments of this application also provide a computer device, which includes a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the above-described method.
[0031] Fourthly, embodiments of this application also provide a computer-readable storage medium storing a computer program that, when executed by a processor, can implement the above-described method.
[0032] This application provides a sign language translation method, apparatus, computer device, and storage medium. The method includes: acquiring a sign language video to be processed; inputting the sign language video to be processed into a preset hand region detection network to obtain a video stream output by the hand region detection network; inputting the video stream into a preset sign language recognition module for adaptive recognition processing to obtain a preliminary text representation output by the sign language recognition module; inputting the preliminary text representation into a preset semantic understanding and text generation module, and using the semantic understanding and text generation module to perform grammatical correction on the preliminary text representation to obtain target language text output by the semantic understanding and text generation module.
[0033] This application establishes a semantic understanding and text generation module capable of grammatical correction. After inputting the sign language video to be processed into a hand region detection network to obtain the video stream output by the hand region detection network, and then inputting the video stream into a sign language recognition module for adaptive recognition processing to obtain a preliminary text representation output by the sign language recognition module, the preliminary text representation is input into the semantic understanding and text generation module. The semantic understanding and text generation module then performs grammatical correction on the preliminary text representation to eliminate structural differences between sign language and spoken language. This results in higher accuracy and more natural expression of the target language text output by the semantic understanding and text generation module. Attached Figure Description
[0034] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0035] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0036] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0037] Figure 1 A flowchart illustrating a first embodiment of a sign language translation method provided in this application;
[0038] Figure 2 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0039] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0040] The following disclosure provides numerous different embodiments or examples for implementing various structures of this application. To simplify the disclosure, specific examples of components and arrangements are described below. These are merely examples and are not intended to limit the scope of this application. Furthermore, reference numerals and / or letters may be repeated in different examples. Such repetition is for simplification and clarity and does not in itself indicate a relationship between the various embodiments and / or arrangements discussed.
[0041] It should be understood that, when used in this specification and the appended claims, the terms "comprising" and "including" indicate the presence of the described features, integrals, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0042] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of the application. As used in this specification and the appended claims, the singular forms “a,” “an,” and “the” are intended to include the plural forms unless the context clearly indicates otherwise.
[0043] It should also be further understood that the term “and / or” as used in this application specification and the appended claims means any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0044] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0045] To address the aforementioned issues, this application provides a sign language translation method that utilizes the semantic understanding and text generation module to perform grammatical correction on the preliminary text representation, thereby eliminating structural differences between sign language and spoken language. This results in higher accuracy and more natural expression of the target language text output by the semantic understanding and text generation module.
[0046] See Figure 1 , Figure 1 This application provides a flowchart illustrating a first embodiment of a sign language translation method, which includes the following steps:
[0047] Step 110: Obtain the sign language video to be processed.
[0048] Step 120: Input the sign language video to be processed into a preset hand region detection network to obtain the video stream output by the hand region detection network.
[0049] Step 130: Input the video stream into a preset sign language recognition module for adaptive recognition processing to obtain the preliminary text representation output by the sign language recognition module.
[0050] Step 140: Input the preliminary text representation into a preset semantic understanding and text generation module, and use the semantic understanding and text generation module to perform grammatical correction on the preliminary text representation to obtain the target language text output by the semantic understanding and text generation module.
[0051] This embodiment sets up a semantic understanding and text generation module capable of grammatical correction. The sign language video to be processed is input into a preset hand region detection network to obtain the video stream output by the hand region detection network. This video stream is then input into a sign language recognition module for adaptive recognition processing, resulting in a preliminary text representation output by the sign language recognition module. This preliminary text representation is then input into the preset semantic understanding and text generation module, which performs grammatical correction on the preliminary text representation to eliminate structural differences between sign language and spoken language. This results in higher accuracy and more natural expression of the target language text output by the semantic understanding and text generation module.
[0052] The sign language translation system described in this application supports mutual translation between multiple mainstream sign languages, such as Chinese Sign Language, American Sign Language, and French Sign Language, and enables bidirectional translation with multiple spoken languages (such as Chinese and English) (e.g., Chinese to English and English to Chinese). This design aims to meet the international communication needs across languages and hearing-impaired groups, and enhance the system's global applicability and inclusivity.
[0053] In some possible implementations, the semantic understanding and text generation module includes a natural language understanding submodule, a grammar correction submodule, and a text generation submodule. Specifically, refer to the second embodiment of a sign language translation method provided in this application, which includes the following steps:
[0054] Step 210: Obtain the sign language video to be processed.
[0055] Step 220: Input the sign language video to be processed into the hand region detection network to obtain the video stream output by the hand region detection network.
[0056] Step 230: Input the video stream into the sign language recognition module for adaptive recognition processing to obtain the preliminary text representation output by the sign language recognition module.
[0057] Step 240: Input the preliminary text representation into the semantic understanding and text generation module.
[0058] Step 250: Perform semantic analysis on the preliminary text representation using the natural language understanding submodule to obtain phrase representations.
[0059] The Natural Language Understanding submodule is primarily responsible for interpreting and identifying the semantic intent and context of sign language sequences, including but not limited to disambiguation and inferring implicit meanings.
[0060] In some possible implementations, this application can integrate multimodal sentiment analysis technology for sentiment recognition and contextual understanding. For example, by recognizing the sign language user's facial expressions, body posture, and tone of voice (if any), its emotional state can be accurately captured, and the corresponding emotional coloring can be preserved in the translation results (e.g., translating a gesture expressing joy as "Great!" instead of the neutral statement "I'm so happy"). Simultaneously, a dialogue history modeling mechanism is introduced to enhance the understanding of long-distance contexts, ensuring the accurate referencing of pronouns and time words (such as "tomorrow") in different contexts.
[0061] Step 260: Input the phrase representation into the grammar correction submodule for grammar correction processing to obtain the corrected text.
[0062] The grammar correction submodule is a key submodule, mainly responsible for converting the input text that reflects sign language grammar into grammatically correct and natural-sounding spoken / written language.
[0063] Step 270: Input the corrected text into the text generation submodule for processing to obtain the target language text output by the text generation submodule.
[0064] The text generation submodule is responsible for polishing the corrected text to generate target language text (such as Chinese text) that is semantically accurate, grammatically correct, and naturally fluent.
[0065] In some possible implementations, the following key grammatical differences exist between sign language and spoken language:
[0066] 1) Non-manual features in sign language (such as raising eyebrows, frowning, head tilting, body posture, and lip movements) are crucial components of sign language, conveying grammatical information, discourse function, and semantic nuances, which are often expressed through intonation or particles in spoken language. For example, in American Sign Language (ASL), raising eyebrows indicates a yes / no question, while frowning indicates specific questions such as "who," "what," or "where." The meaning of a single sign language word can change due to these non-manual signals. For instance, when the sign "high" is accompanied by the lip movement "cha," the meaning immediately jumps from "high" to "huge." Therefore, non-manual features in sign language are not merely secondary expressions, but independent linguistic dimensions that determine grammatical structure and lexical meaning.
[0067] 2) Sign language extensively utilizes the three-dimensional sign language space for grammatical expression, including establishing reference, confirming verb agreement, demonstrating spatial relationships, and conveying temporal sequences. For example, directional verbs can clearly define the roles of different participants in a spatial context.
[0068] 3) Word order: While many spoken languages (such as English and Chinese) typically follow a subject-verb-object (SVO) structure, sign language often uses a topic-comment structure. For example, in British Sign Language (BSL), "I eat the apple" is a common structure. American Sign Language also accepts an object-subject-verb (OSV) word order.
[0069] Time markers usually appear at the beginning of a sentence (e.g., "See you tomorrow" rather than "I will see you tomorrow").
[0070] 4) Morphology: Sign language has rich morphological features, and its word formation is significantly different from the linear sequence structure of spoken language. In sign language, the combination of affixes is usually simultaneous (such as the simultaneous presentation of non-manual features such as gestures, facial expressions, and body postures), rather than the sequential arrangement in spoken language.
[0071] Common morphological phenomena include:
[0072] Derivative forms: Word class conversion is achieved through adjustments to gestures or actions. For example, the noun "chair" is derived from the action pattern of the verb "sit";
[0073] Inflectional morphology: expressing grammatical meaning through changes in non-manual features (such as facial expressions, mouth shapes, head movements), such as the aspect of verbs (perfective, progressive) or mood;
[0074] Repetition in word formation: The repetition of gestures is widely used to express semantic intensity, verb aspect (such as continuous or habitual actions), or to achieve morphological derivation from verbs to nouns.
[0075] 5) Conciseness: Sign language is usually more concise than spoken language, often omitting auxiliary verbs commonly found in spoken grammar (e.g., "is", "will").
[0076] Therefore, if these profound grammatical differences are not taken into account, and simple spoken grammar rules are directly applied to translate sign language into spoken language literally, the output translation will have problems such as incorrect grammar, semantic distortion, and unnaturalness, making it difficult for native speakers of the target language to understand.
[0077] This application employs a grammar correction submodule to perform non-manual feature processing, spatial grammar processing, word order processing, morphological processing, and conciseness processing, ensuring that the output translation retains the original meaning while conforming to the grammatical rules and natural fluency of the target language, thereby improving the quality and accuracy of the translation results.
[0078] In some possible implementations, this application can construct a model decision visualization module, such as using attention heatmaps and grammar correction path annotations to intuitively display key judgment criteria during the translation process (e.g., "eyebrow raising action detected, judged as interrogative tone"). This function not only enhances system transparency but can also serve as a teaching tool to help sign language learners understand the conversion logic and grammatical rules between sign language and spoken language.
[0079] In some possible implementations, the sign language recognition module is equipped with a trained deep neural network model. Step 230, which involves inputting the video stream into the sign language recognition module for adaptive recognition processing to obtain the preliminary text representation output by the sign language recognition module, may include the following steps:
[0080] Step 231: Input the video stream into the trained deep neural network model.
[0081] Step 232: Preprocess the video stream using the trained deep neural network model to obtain a structured video stream.
[0082] In some possible implementations, the preprocessing of the video stream using the trained deep neural network model to obtain a structured video stream includes:
[0083] The video stream is processed into video frames to obtain a sequence of hand image frames.
[0084] Specifically, a continuous video stream can be divided into segments (e.g., 15-30 frames per second) according to a time window.
[0085] The pose of the hand image frame sequence is estimated to obtain a key point coordinate sequence, wherein the key point coordinate sequence includes hand key points (such as fingertips and joint positions), facial key points (such as mouth and eyebrows), and body key points (such as shoulder and elbow positions to help determine orientation).
[0086] Among them, a pre-trained human pose estimation model can be used to detect key points and estimate pose, thereby obtaining a sequence of key point coordinates.
[0087] 3) Normalize the coordinates of the key points to obtain a structured video stream.
[0088] Specifically, by converting the coordinates of key points into positions relative to reference points such as the torso, the impact of differences in user position and body size can be effectively reduced.
[0089] In addition, since many sign language gestures are essentially specific motion trajectories (such as waving, drawing circles, pushing and pulling), and optical flow can directly capture these motion patterns, based on this, in some possible implementations, optical flow calculations can be further performed on the structured video stream to capture motion information, thereby helping to more accurately recognize dynamic gestures.
[0090] Step 233: Extract features from the structured video stream to obtain multiple hand features.
[0091] Step 234: Perform feature fusion on multiple hand features to obtain a preliminary text representation output by the trained deep neural network model.
[0092] For step 233, in some possible implementations, the multiple hand features include spatial features and spatiotemporal features. The feature extraction of the structured video stream yields multiple hand features, including:
[0093] The static hand features are extracted independently from each frame of the structured video stream to obtain spatial features.
[0094] Specifically, convolutional neural networks (CNNs) can be used to process visual features of hand shape, orientation, and static gestures. For example, the target region (such as the hand region) can be cropped from the preprocessed keypoint heatmap / coordinate vector / original image.
[0095] In addition, graph convolutional networks (GCNs) can be used to process the topological relationships of key points, thereby obtaining the spatial dependencies between hand joints / body parts (such as the relative positions of fingers).
[0096] 2) Extract hand dynamic features from multiple consecutive frames in the structured video stream to obtain spatiotemporal features.
[0097] Among these methods, recurrent neural networks (such as RNN, LSTM, or GRU) can be used to capture the dynamic changes of gestures in the time dimension (such as movement trajectory and speed).
[0098] A three-dimensional convolutional neural network (3D-CNN) is used to learn spatial and short-term motion features (such as the direction of hand movement) simultaneously.
[0099] We utilize a self-attention mechanism to efficiently capture the dependencies between any two frames in a long sequence.
[0100] In some embodiments, spatial features and spatiotemporal features can be fused (e.g., splicing, attention weighting).
[0101] In some possible implementations, the sign language recognition module employs an adaptive learning mechanism based on an efficient teacher-student framework that has achieved significant success in improving natural language understanding (NLU) tasks using massive amounts of unlabeled data. Its core idea is to guide the learning of the "student" model through a "teacher" model, thereby fully utilizing unlabeled data. For details, please refer to the following:
[0102] 1) Pseudo-label generation and iterative optimization
[0103] Pseudo-label generation: First, an initial "teacher" model trained on limited labeled sign language data is used to predict the sign language content of a large number of unlabeled videos. Predictions with high confidence are then used as "pseudo-labels" and assigned to the corresponding unlabeled data.
[0104] Expanded Training and Model Iteration: Subsequently, the "student" model is trained on an expanded dataset that combines the original labeled data with unlabeled data containing pseudo-labels. As the student model's performance improves, it will take over as the new "teacher" model in subsequent iterations, generating more accurate pseudo-labels for the unlabeled data. This self-training cycle enables continuous optimization of model performance without requiring additional manual annotation.
[0105] Sample selection strategy: To ensure the quality of pseudo-labels and learning efficiency, the algorithm employs an intelligent sample selection strategy to filter out the most valuable samples for self-supervised learning (SSL) from massive amounts of unlabeled data. For example:
[0106] The committee's choice: Use multiple models to predict the same data, and only generate pseudo-labels when the models reach a consensus.
[0107] Sub-model optimization selection: Prioritize the selection of diverse samples that are rich in information and representative to improve the effectiveness of training.
[0108] Thus, this adaptive learning mechanism achieves continuous improvement and performance enhancement of the model without additional labeling costs through iterative pseudo-label learning driven by the teacher-student framework and combined with intelligent sample selection.
[0109] Knowledge distillation and model refinement
[0110] Knowledge distillation (KD) is a model compression technique designed to transfer the knowledge learned by a complex, high-performance "teacher" model (or model ensemble) to a smaller, lighter "student" model. Unlike using traditional hard labels (i.e., deterministic class labels), knowledge distillation utilizes the soft probabilities or logits output by the teacher model as the training targets for the student model, thereby conveying richer information about inter-class relationships.
[0111] This technology is significant for deploying sign language recognition algorithms to edge devices such as "sign language bridges." These devices are typically based on embedded systems with limited computing power and storage resources. Through knowledge distillation, compact models that are small in size, fast inference speed, and suitable for offline operation can be built, while preserving the high accuracy performance of the teacher model to the greatest extent possible.
[0112] Because sign language involves not only fine hand movements, facial expressions, and body postures, but also rich temporal dynamics and non-manual grammatical markers, accurate annotation requires a large amount of manpower and time from professionals. This makes the collection and annotation of sign language data extremely complex and time-consuming, leading to the problem of data sparsity.
[0113] The adaptive learning framework in the sign language recognition module of this application adopts a semi-supervised learning method of pseudo-labels and knowledge distillation, which enables the algorithm to effectively utilize a large amount of unlabeled sign language data. This expanded data contact helps to train a more robust, more generalizable and more accurate sign language recognition model, thereby reducing overfitting to limited labeled datasets and improving the problem of scarce labeled sign language datasets.
[0114] In addition, this application can be based on advanced paradigms of meta-learning and prompt learning, enabling the model to quickly adapt to new vocabulary, local sign language variants or emerging expressions (such as gestures of internet slang) with only a very small number of samples or even no labeled samples, significantly improving the system's responsiveness to dynamic language environments.
[0115] In some possible implementations, other semi-supervised learning methods for evaluation, such as Virtual Adversarial Training (VAT) and Cross-View Training (CVT), can also be employed. These methods each possess unique regularization properties and mechanisms to enhance model robustness. Based on this, these techniques can be selectively integrated into the overall framework according to the specific characteristics of the sign language data and the needs of the actual application scenario, in order to further improve the model's generalization ability and performance.
[0116] In addition, to improve robustness and generalization ability in diverse real-world scenarios, the sign language recognition module of this application also introduces a domain adaptation (DA) mechanism and personalized learning to ensure that the model can maintain high-precision recognition performance when facing different application environments and individual user differences.
[0117] Specifically, in actual deployments, the visual environment, lighting conditions, background complexity, or sign language expression habits of the target application scenario (such as hospitals, schools, and government service windows) may differ significantly from the source domain data used for model training, leading to a decline in recognition performance, i.e., the "domain shift" problem. To address this, this application employs domain adaptation technology, which, by jointly utilizing labeled source domain data and unlabeled target domain data, achieves feature space alignment and distribution transfer, effectively reducing inter-domain differences and improving the model's generalization ability in new environments.
[0118] In addition, the system can implicitly collect users' sign language video data (such as gesture amplitude, movement speed, expression clarity, etc.) during device operation, and use these small number of user-specific samples to perform lightweight fine-tuning on the basic recognition model, so as to achieve continuous learning and adaptation, gradually master the speed, movement amplitude, clarity and unique style characteristics of individual users in sign language expression, thereby significantly improving the recognition accuracy and user experience of the same user in subsequent use.
[0119] In some possible implementations, this application can upgrade the virtual sign language synthesis system to support user-defined sign language expression styles, including gesture amplitude, sign language speed, facial expressions (such as smiling, focused, serious) and body movements, to achieve personalized output that is "one face for one person", enhance the emotional warmth and affinity of the interaction, and improve the user experience.
[0120] It also allows users to customize translation rules for specific gestures or phrases, such as technical terms, family expressions, or expressions specific to social circles. The system can quickly update its personalized vocabulary database through user feedback mechanisms, enabling dynamic learning and localization adaptation to meet customized communication needs in special scenarios such as education, healthcare, and family settings (e.g., exclusive gesture instructions used by deaf teachers in the classroom).
[0121] In this way, users can use it immediately without changing their own expression habits, and the system can continuously optimize itself during use, achieving an intelligent experience that becomes more accurate the more it is used.
[0122] In summary, this application combines adaptive learning (to enable individual sign language users to adapt) and grammar correction (to generate natural and fluent target language), enabling the system to not only understand diverse sign language styles, but also to communicate with deaf and hearing people in a clear, standard, and easy-to-understand manner, thereby improving the user experience.
[0123] Furthermore, by leveraging the adaptive mechanism within the sign language recognition module and the synergy between the grammar correction module, these two modules can influence each other. For example, improvements in one sign language recognition module can directly enhance the performance of the other, thus forming a self-reinforcing and positive feedback loop learning ecosystem. The improvements in the translation system's output and the results of grammar correction can, in turn, be used to generate higher-quality pseudo-labels for the adaptive learning components, or even as a form of weak supervision to further refine the model. This self-reinforcing ecosystem means that the algorithm can continuously learn and improve its performance, becoming more accurate and natural, as usage and data exposure increase.
[0124] Furthermore, the core design philosophy of the "Sign Language Bridge" product is to support offline operation, aiming to fundamentally protect user privacy and system stability. To adapt to the embedded system used by "Sign Language Bridge" (which typically has limited computing resources), this application employs several lightweight and efficient inference techniques in its algorithm.
[0125] For example, knowledge distillation (KD) enables model compression, transferring knowledge from large models to smaller, more efficient, lightweight models; simultaneously, techniques such as non-autoregressive (NAR) GANs are explored to support parallel inference and significantly improve processing speed. These optimizations ensure that complex AI models can run efficiently locally on devices without relying on cloud computing.
[0126] Furthermore, by achieving complete offline functionality and employing a "privacy design," sensitive user communication content, especially video and interaction data related to personal sign language expressions, is retained entirely on the local device and is never uploaded to any remote server. This fundamentally eliminates the risk of data leakage and is suitable for scenarios with extremely high privacy requirements, such as hospitals, government agencies, educational institutions, or homes.
[0127] Furthermore, by employing non-autoregressive GAN parallel generation technology, the end-to-end translation latency can be controlled within 0.5 seconds, meeting the low-latency operation requirements for real-time communication.
[0128] Furthermore, because it does not rely on a network connection, the system can operate stably in environments with no or weak network coverage, avoiding service failures caused by network latency, interruptions, or server malfunctions. This offline robustness ensures continuous service availability, significantly improving the reliability and user experience of "Sign Language Bridge" in various practical application environments.
[0129] In conjunction with the above embodiments, this application provides a comprehensive, end-to-end two-way sign language translation system. The system architecture design closely matches the actual application needs of intelligent barrier-free interactive products such as "sign language bridges," fully embodying the core concepts of real-time performance, interactivity, and intelligence.
[0130] The two-way sign language translation system supports two core communication processes, enabling seamless two-way communication between deaf and hearing people:
[0131] 1. The process of translating sign language into spoken language:
[0132] 1) The camera collects real-time video of sign language from deaf and mute users.
[0133] 2) Input the sign language video to be processed into the hand region detection network to obtain the video stream output by the hand region detection network;
[0134] 3) The sign language recognition module performs adaptive recognition processing on the video stream to obtain the preliminary text representation output by the sign language recognition module;
[0135] 3) The initial text representation input semantic understanding and text generation module, through the natural language understanding submodule, grammar correction submodule and text generation submodule, converts the original output that conforms to sign language grammar into grammatically correct, semantically complete and naturally expressed spoken text, i.e. target language text.
[0136] 4) The final target language text can be read aloud via text-to-speech (TTS) or displayed on the screen in real time for hearing people to read.
[0137] 2. The process of translating spoken language into sign language:
[0138] 1) The microphone collects the voice input of the hearing user;
[0139] 2) The speech recognition module (ASR) converts speech into text;
[0140] 3) The text-to-sign language translation module maps spoken text into a sequence of action parameters that conform to the grammatical rules of sign language, including gestures, directions, non-manual features (such as facial expressions and lip movements), and spatial trajectories;
[0141] 4) The virtual sign language generation module drives the AI virtual human to synthesize realistic, standard, and imitable sign language videos, which are presented to deaf and mute users in real time.
[0142] Corresponding to the above sign language translation method, this application also provides a sign language translation device. The sign language translation device includes a unit for performing the above sign language translation method, and the device can be configured in a desktop computer, tablet computer, laptop computer, or other terminal.
[0143] like Figure 2 As shown in the figure, this application provides a computer device including a processor 111, a communication interface 112, a memory 113, and a communication bus 114, wherein the processor 111, the communication interface 112, and the memory 113 communicate with each other through the communication bus 114.
[0144] Memory 113 is used to store computer programs;
[0145] In one embodiment of this application, when the processor 111 executes the program stored in the memory 113, it implements the sign language translation method provided in any of the foregoing method embodiments, including:
[0146] Obtain the sign language video to be processed;
[0147] The sign language video to be processed is input into the hand region detection network to obtain the video stream output by the hand region detection network;
[0148] The video stream is input into the sign language recognition module for adaptive recognition processing to obtain the preliminary text representation output by the sign language recognition module;
[0149] The preliminary text representation is input into the semantic understanding and text generation module, and the semantic understanding and text generation module is used to perform grammatical correction on the preliminary text representation to obtain the target language text output by the semantic understanding and text generation module.
[0150] It will be understood by those skilled in the art that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program may be stored in a storage medium, which is a computer-readable storage medium. The computer program is executed by at least one processor in the computer system to implement the process steps of the embodiments of the above methods.
[0151] Therefore, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, wherein the computer program, when executed by a processor, implements the steps of the sign language translation method provided in any of the foregoing method embodiments, including:
[0152] Obtain the sign language video to be processed;
[0153] The sign language video to be processed is input into the hand region detection network to obtain the video stream output by the hand region detection network;
[0154] The video stream is input into the sign language recognition module for adaptive recognition processing to obtain the preliminary text representation output by the sign language recognition module;
[0155] The preliminary text representation is input into the semantic understanding and text generation module, and the semantic understanding and text generation module is used to perform grammatical correction on the preliminary text representation to obtain the target language text output by the semantic understanding and text generation module.
[0156] The storage medium is a physical, non-transient storage medium, such as a USB flash drive, external hard drive, read-only memory (ROM), magnetic disk, or optical disk, or any other physical storage medium capable of storing program code. The computer-readable storage medium can be non-volatile or volatile.
[0157] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0158] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative. For example, the division of each unit is merely a logical functional division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed.
[0159] The steps in the methods of this application embodiment can be adjusted, merged, or deleted according to actual needs. The units in the apparatus of this application embodiment can be merged, divided, or deleted according to actual needs. Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0160] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, a terminal, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application.
[0161] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0162] Obviously, those skilled in the art can make various modifications and variations to this application without departing from the spirit and scope of this application. Since these modifications and variations fall within the scope of the claims and their equivalents, this application also intends to include these modifications and variations.
[0163] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A sign language translation method, characterized by, The sign language translation method comprises: acquiring a sign language video to be processed; inputting the sign language video to be processed into a preset hand region detection network to obtain a video stream output by the hand region detection network; inputting the video stream into a preset sign language recognition module for adaptive recognition processing to obtain a preliminary text representation output by the sign language recognition module; inputting the preliminary text representation into a preset semantic understanding and text generation module, and performing grammar correction on the preliminary text representation by using the semantic understanding and text generation module to obtain a target language text output by the semantic understanding and text generation module; wherein the semantic understanding and text generation module comprises a natural language understanding submodule, a grammar correction submodule and a text generation submodule, and the grammar correction on the preliminary text representation by using the semantic understanding and text generation module to obtain the target language text output by the semantic understanding and text generation module comprises: performing semantic analysis on the preliminary text representation by using the natural language understanding submodule to obtain a phrase representation; inputting the phrase representation into the grammar correction submodule for grammar correction processing to obtain a corrected text; inputting the corrected text into the text generation submodule for processing to obtain the target language text output by the text generation submodule; wherein the grammar correction processing comprises non-manual feature processing, spatial grammar processing, order processing, morphological processing and conciseness processing; wherein the sign language recognition module is provided with a trained deep neural network model, and the adaptive recognition processing on the video stream by inputting the video stream into the sign language recognition module to obtain the preliminary text representation output by the sign language recognition module comprises: inputting the video stream into the trained deep neural network model; preprocessing the video stream by using the trained deep neural network model to obtain a structured video stream; extracting features from the structured video stream to obtain a plurality of hand features; fusing the plurality of hand features to obtain the preliminary text representation output by the trained deep neural network model.
2. The method of claim 1, wherein, The preprocessing of the video stream by using the trained deep neural network model to obtain the structured video stream comprises: performing video frame processing on the video stream to obtain a sequence of hand image frames; performing pose estimation on the sequence of hand image frames to obtain a sequence of key point coordinates, wherein the sequence of key point coordinates comprises hand key points, face key points and body key points; performing coordinate normalization processing on the sequence of key point coordinates to obtain the structured video stream.
3. The method of claim 1, wherein, The plurality of hand features comprises spatial features and spatio-temporal features, and the feature extraction from the structured video stream to obtain the plurality of hand features comprises: independently performing hand static feature extraction on each frame in the structured video stream to obtain spatial features; performing hand dynamic feature extraction on a plurality of consecutive frames in the structured video stream to obtain spatio-temporal features.
4. The method of claim 1, wherein, The sign language recognition module adopts a semi-supervised learning mechanism of pseudo-labeling and / or knowledge distillation.
5. A sign language translation device, characterized by, The method comprises a unit for performing the method according to any one of claims 1-4.
6. A computer device, comprising: The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1-4 when executing the computer program.
7. A computer readable storage medium characterized by The storage medium stores a computer program, and the computer program can implement the method according to any one of claims 1-4 when being executed by a processor.
Citation Information
Patent Citations
Robot sign language communication method, related device and storage medium
CN120578294A