Systems, methods, computer accessible medium, and devices for electronically converting between text / spoken language and sign language
The OneWorldAI architecture integrates sign language translation and production using advanced tokenization and cross-attention mechanisms, addressing the inefficiencies of separate systems by ensuring semantic and temporal alignment, thus enhancing communication between the Deaf and hearing communities.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- NEW YORK UNIV IN ABU DHABI CORP
- Filing Date
- 2026-01-16
- Publication Date
- 2026-07-23
AI Technical Summary
Existing sign language translation and production systems treat translation and production as distinct tasks, leading to increased computational demands, resource inefficiencies, and a disconnect between the two, failing to leverage shared semantic and temporal dependencies, thus hindering seamless bidirectional communication.
A unified multimodal generative framework, such as the OneWorldAI architecture, integrates sign language translation and production using advanced tokenization and a SignXFormer model with cross-attention mechanisms to align sign and natural language tokens within a shared representation space, capturing temporal and semantic coherence.
This approach reduces computational redundancy and enhances seamless bidirectional communication by preserving semantic and temporal alignment, enabling accurate and fluent sign-to-text translation and text-to-sign generation, bridging the gap between the Deaf and hearing communities.
Smart Images

Figure IB2026050419_23072026_PF_FP_ABST
Abstract
Description
PATENT APPLICATION SYSTEMS, METHODS, COMPUTER ACCESSIBLE MEDIUM, AND DEVICES FOR ELECTRONICALLY CONVERTING BETWEEN TEXT / SPOKEN LANGUAGE AND SIGN LANGUAGECROSS-REFERENCE TO RELATED APPLICATION(S)
[0001] This application relates to and claims priority from U. S. Patent Application No.63 / 746,709, filed on January 17, 2025, the entire disclosure of which is incorporated herein by reference.FIELD OF THE DISCLOSURE
[0002] The present disclosure relates generally to communications using sign language, and more specifically to systems, methods and devices that can convert between sign language and text / spoken language, e.g., sign language translation and production.BACKGROUND INFORMATION
[0003] Globally, hearing loss and deafness impact millions of individuals, with over 430 million people, including 34 million children, requiring rehabilitation for disabling hearing loss. (See, e.g., Kushalnagar, 2019). By 2050, this number is expected to exceed 700 million, representing one in every ten people, alongside an estimated 2.5 billion experiencing varying degrees of hearing loss (Bauman & Murray, 2009). Communication barriers significantly affect education, employment, and social inclusion for the Deaf community and those with hearing loss, particularly in low- and middle-income countries, where 80% of cases are concentrated. (See, e.g., Luey et al., 1995). To address these challenges, the development of sign language production and translation technologies is critical. These techniques facilitate more accessible communication, enhance societal integration, and improve the quality of life for individuals with hearing impairments, bridging the gap between the Deaf community and the hearing world.
[0004] Sign language production and translation can be pivotal areas in bridging the communication gap between the Deaf community and non-sign language users. Sign language production involves the creation of synthetic sign language representations, typically using avatars or video-based models, to convey information in a visual format that adheres to the linguistic rules of sign languages. (See, e.g., Saunders et al., 2021, 2020). This process requires precise modeling of gestures, facial expressions, and body movements to ensure linguisticPATENT APPLICATION accuracy and naturalness. On the other hand, sign language translation focuses on converting sign language, whether in video or sensor-based inputs, into written or spoken languages. (See, e.g., Chen et al,, 2022).
[0005] Previous work in sign language production primarily relies on autoregressive models, such as the progressive transformer introduced by Saunders et al., which utilized encoder-decoder architectures to map gloss sequences to sign poses in a sequential manner (Saunders et al., 2020). To mitigate the issues like an exposure bias during inference (Schmidt et al., 2019), non-autoregressive methods were later developed, including Huang et al.’s G2P model with monotonic alignment search, which enabled parallel prediction of alignment lengths for gloss sequences. (See, e.g., Huang et al., 2021). More recent advancements leverage discrete diffusion models and vector quantization techniques, as exemplified by the G2P-DDM framework, which combines Pose-VQVAE and CodeUnet architectures to utilize spatial and temporal information, enhancing the generation of sign pose sequences. (See, e.g., Xie et al., 2024).
[0006] Recent work in sign language translation (SLT) can be categorized into gloss-based and gloss-free approaches. Gloss-based methods utilize intermediate representations called glosses, which are sequential annotations of sign language videos. These methods often employ encoderdecoder frameworks, such as SLRT (see, e.g., Camgoz et al., 2020), which uses Connectionist Temporal Classification (CTC) loss for soft alignment between glosses and natural language text. Similarly, STMC-T (see, e.g., Zhou et al., 2021b) incorporates multi-cue learning to enhance the translation process, and SignBack (see, e.g., Zhou et al., 2021) introduces back-translation to improve the model’s performance. However, gloss-based approaches face challenges due to the labor-intensive nature of gloss annotation and the bottleneck effect of glosses limiting the richness of translation. In contrast, gloss-free SLT methods aim to bypass gloss annotations entirely. For example, NSLT (see, e.g., Camgoz et al., 2018) employs CNNs for visual feature extraction and RNNs with attention mechanisms for text modeling, while GFSLT-VLP (see, e.g., Zhou et al., 2023) utilizes visual-language pretraining to align visual and textual features in a joint semantic space, achieving competitive results on datasets like PHOENIX14T.
[0007] Previous works, as exemplified above, treated sign language translation (e.g., sign-to-text) and production (e.g., text-to-sign) as distinct tasks, often requiring separate models and training processes. This separation not only increases computational and resource demands butPATENT APPLICATION also creates a disconnect between the two tasks, which inherently share semantic and temporal dependencies. Without a unified framework, these models are unable to leverage shared knowledge between translation and production, such as temporal alignment and semantic relationships, which can enhance both tasks simultaneously. Additionally, existing approaches often struggle to capture the complex and multimodal nature of sign language and natural language, including the temporal dependencies of signing gestures and the syntactic nuances of text. These limitations hinder the creation of a seamless and cohesive system that can handle bidirectional communication, reducing the usability of such models in real-world scenarios
[0008] Accordingly, there is a need to address and / or improve at least the above-described deficiencies which exist in the previous systems, methods and devices by providing systems, methods and devices by providing, e.g., a unified multimodal generative framework designed to integrate sign language translation and production into a cohesive model.SUMMARY OF EXEMPLARY EMBODIMENTS
[0009] The following is intended to be a brief summary of the exemplary embodiments of the present disclosure and is not intended to limit the scope of the exemplary embodiments.
[0010] In some exemplary embodiments of the present disclosure, the exemplary systems, methods, computer accessible medium, and devices can be provided for generating bi-directional real-time speech and sign language conversion. The exemplary systems, methods, computer accessible medium, and devices can receive speech in real-time, convert the speech into a plurality of speech tokens, identify a relevance between the speech tokens and a plurality of sign tokens, and generate, in real-time and based on the identified relevance, an animated avatar performing the generated sign language corresponding to the received speech. A learning model configured to convert the speech into the plurality of speech tokens and identifying the relevance can be trained to align semantically equivalent concepts within a shared unified representation space. Further, the exemplary trained learning model can, e.g., capture (i) static semantics of the speech tokens and sign tokens, and (ii) temporal dependencies of the speech tokens and sign tokens. The conversion of speech into a plurality of speech tokens can involve progressively refining token predictions to enhance representation learning. Given its global context information in speech, this process remains robust even in the presence of missing or corrupted speech tokens.PATENT APPLICATION
[0011] The animated avatar performing the generated sign language corresponding to the received speech can be displayed on a device proximal to a person creating the received speech. For example, the generated sign language can be displayed on a podium, from which a speaker is creating the speech. Further, the animated avatar performing the generated sign language corresponding to the received speech is displayed on a pair of smart glasses.
[0012] In some exemplary embodiments of the present disclosure, the exemplary systems, methods, computer accessible medium, and devices can be provided for generating bi-directional real-time speech and sign language conversion. The exemplary systems, methods, computer accessible medium, and devices can observe a plurality of sign language signs in real-time, convert the observed signs into a plurality of sign tokens, identify a relevance between the sign tokens and a plurality of speech tokens, and generate, in real-time and based on the identified relevance, one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs. An exemplary learning model configured to convert the observed signs into the plurality of sign tokens and identifying the relevance can be trained to align semantically equivalent concepts within a shared unified representation space. Further, the trained learning model can: (i) progressively learn speech and sign tokens, (ii) capture the hierarchical motion structure of sign tokens by incorporating motion frequencies from low to high to enhance representation, and (iii) seamlessly account for the unique temporal and gestural characteristics of sign language signs to maintain consistent temporal alignment with speech tokens.Additionally, the conversion of speech into sign tokens involves refining token predictions by capturing both high- and low-frequency components of the representation. The model is capable of generating 3D sign motion gestures in a 3D pose skeleton in real time and rendering realistic human sign gestures in real time.
[0013] Additionally, such one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs can be, e.g., (i) displayed on a device proximal to a person creating the received speech for written text, or (ii) played through a speaker proximal to a person creating observed sign language signs. Further, such one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs can be, e.g., (i) displayed on a pair of smart glasses for written text, or (ii) played through a speaker on the smart glasses for the spoken text.
[0014] These and other objects, features and advantages of the exemplary embodiments of thePATENT APPLICATION present disclosure will become apparent upon reading the following detailed description of the exemplary embodiments of the present disclosure, when taken in conjunction with the appended claims.BRIEF DESCRIPTION OF THE DRAWINGS
[0013] Further objects, features and advantages of the present disclosure will become apparent from the following detailed description taken in conjunction with the accompanying Figures showing illustrative embodiments of the present disclosure, in which:
[0014] Figure 1 is an exemplary diagram of a OneWorldAI model for sign language translation and production according to an exemplary embodiment of the present disclosure;
[0015] Figure 2 is an exemplary system diagram for an exemplary Hierarchical Sign Language Tokenizer according to an exemplary embodiment of the present disclosure;
[0016] Figure 3 is an exemplary diagram of an exemplary SignXFormer according to an exemplary embodiment of the present disclosure;
[0017] Figure 4 is an exemplary diagram of an exemplary HierSignFormer according to an exemplary embodiment of the present disclosure; and
[0018] Figure 5 is an exemplary diagram illustrating an exemplary inference time data flow of work according to an exemplary embodiment of the present disclosure.
[0019] Throughout the drawings, the same reference numerals and characters, unless otherwise stated, are used to denote like features, elements, components or portions of the illustrated embodiments. Moreover, while the present disclosure will now be described in detail with reference to the figures, it is done so in connection with the illustrative embodiments and is not limited by the particular embodiments illustrated in the figures and the appended claims.DETAILED DESCRIPTION OF EXEMPLARY EMBODIMENTS
[0020] The following description of exemplary embodiments provides non-limiting representative examples referencing numerals to particularly describe features and teachings of different exemplary aspects and exemplary embodiments of the present disclosure. The exemplary embodiments described should be recognized as capable of implementation separately, or in combination, with other exemplary embodiments from the description of the exemplary embodiments. A person of ordinary skill in the art reviewing the description of the exemplary embodiments should be able to learn and understand the different described aspects ofPATENT APPLICATION the present disclosure. The description of the exemplary embodiments should facilitate understanding of the exemplary embodiments of the present disclosure to such an extent that other implementations, not specifically covered but within the knowledge of a person of skill in the art having read the description of embodiments, would be understood to be consistent with an application of the exemplary embodiments of the present disclosure.
[0021] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide a unified multimodal generative framework 100 configured to integrate sign language translation and production into a single cohesive model (e.g., the OneWorldAI architecture) as shown in Figure 1. This exemplary framework 100 can utilize advanced tokenization techniques through tokenizers 130 and 140, as well as aligned token space 150, to transform complex input modalities into compact, linguistically meaningful tokens while preserving critical features like temporal alignment, semantic coherence, and gestural nuances. For sign language, the exemplary model according to the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can process video inputs into discrete tokens using the Sign Language Tokenization Module 130, ensuring that gestures, hand shapes, and facial expressions are accurately captured.
[0022] On the natural language side, off-the-shelf tokenization 140 can be used to preserve syntactic and semantic structures. These tokens can be aligned within a shared representation space 150 using the novel SignXFormer model, which leverages a cross-attention mechanism (see, e.g., A. Vaswani et al., 2017 and E. Pinyoanuntapong et al., 2024) to align multimodal representations, including latent embedding of speech, text, and video. It effectively incorporates temporal information to maintain coherence in sequential data. Upon aligning the representation space, SignXFormer 160 and / or 170, a bidirectional transformer, can further bridge the two modalities, enabling accurate and fluent sign-to-text translation and text-to-sign production. The exemplary systems, methods, computer-accessible medium, and devices described in this disclosure leverage this approach to enhance multimodal integration.
[0023] As shown in Figure 1, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide additional functional components, such as SignXFormer for processing noisy inputsand HierSignFormer for refining token predictions (together 160 and / or 170), further enhance thePATENT APPLICATION model’s robustness and contextual sensitivity. By unifying translation and production tasks, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can significantly reduce computational redundancy and ensures seamless bidirectional communication, making it a powerful tool for bridging the gap between the Deaf and hearing communities.
[0024] For example, Figure 1 illustrates that given a sign avatar input 115 or spoken language input 125, the Tokenizer module 130 and 140 respectively can first process the input into an aligned token space 150, preserving semantic and temporal coherence. Using these aligned tokens, the Bidirectional SignXFormer 160 and 170 respectively can generate the corresponding modality. The system 100 can output a visual sign 165 from a spoken language 120 and / or a textual output 175 from a sign motion 110.Exemplary Related WorkExemplary Sign Language Production
[0025] Sign language production (SLP) has garnered significant attention, with various approaches aiming to improve the naturalness and accuracy of generated sign sequences. Early works, such as those by Stoll et al., introduced neural machine translation methods combined with generative adversarial networks (GANs) to produce realistic sign pose sequences (see, e.g., Stoll et al., 2018). Saunders et al. further advanced the field by proposing a progressive transformer based model for Grapheme-to-Phoneme Conversion (G2P) tasks, which used an encoder-decoder framework to map gloss sequences to sign poses (see, e.g., Saunders et al., 2020). Additionally, their later work incorporated a mixture of motion primitives (MoMP) to enhance the fluency of continuous sign sequences (see, e.g., Saunders et al., 2021). To improve the realism of sign videos, Saunders et al. also introduced SIGNGAN, a pose-conditioned human synthesis model, capable of generating photo-realistic sign videos directly from skeleton poses (see, e.g., Saunders et al., 2022).
[0026] Non-autoregressive methods have also emerged, exemplified by Huang et al.’s G2P model, which employed a monotonic alignment search for one-shot decoding, reducing inference latency and error accumulation. (See, e.g., Huang et al., 2021). Recent advancements include discrete diffusion models, such as G2P-DDM, which leverage Pose VQVAE for latent space tokenization and CodeUnet for spatial-temporal modeling, achieving promising results onPATENT APPLICATION datasets like RWTH-PHOENIX-WEATHER-2014T. (See, e.g., Xie et al., 2024). Despite this progress, challenges persist in creating universally adaptable models that account for linguistic, cultural, and signer variability, as well as the need for larger annotated datasets to support robust training.Exemplary Sign Language Translation
[0027] Sign Language Translation (SLT) research has advanced through gloss-based and gloss-free approaches, each addressing unique challenges. Gloss-based methods leverage intermediate gloss annotations to facilitate translation, with frameworks like SLRT (see, e.g., Camgoz et al., 2020) employing Transformer- based encoder-decoder models and Connectionist Temporal Classification (CTC) loss for aligning sign representations with gloss sequences.STMC-T (see, e.g., Zhou et al., 2021b) further enhances temporal modeling through multi cue learning, while SignBack (see, e.g., Zhou et al., 2021) applies back-translation techniques to augment data and improve accuracy. Transfer learning has also been introduced in SLT, with Chen et al. (see, e.g., Chen et al., 2022) utilizing large language models to boost performance. However, gloss annotations are resource-intensive and create bottlenecks in capturing rich contextual information. To overcome these limitations, gloss-free methods bypass intermediate representations, as seen in NSLT (see, e.g., Camgoz et al., 2018), which uses CNNs and RNNs for direct translation, and GFSLT VLP (see, e.g., Zhou et al., 2023), which employs Visual-Language Pretraining (VLP) to align visual and textual representations in a shared semantic space. This gloss-free approach achieves competitive results while addressing scalability and annotation challenges, marking a significant step forward in SLT development.Exemplary MethodsExemplary Overall OneWorldAI Model Architecture
[0028] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide architecture that is a unified, multimodal generative framework with a specific initial focus on enabling seamless translation and generation between sign language and natural language. By leveraging the complementary strengths of advanced tokenization, temporal alignment, and a transformer-based architecture, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can bridge the communication gap between thePATENT APPLICATION deaf and hearing communities. The exemplary model provided by the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can be designed to handle multi-model generative tasks — sign-to-text / speech and text / speech-to-sign — with high accuracy, fluency, and contextual sensitivity.
[0029] One aspect of the exemplary architecture according to exemplary embodiments is its ability to transform complex and informative input modalities into a compact yet representative representation while maintaining semantic coherence and temporal correspondence. For sign language, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide a model that can take video input, which can be preprocessed into discrete, linguistically meaningful tokens using a sign Language tokenization module. This tokenization performed by the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can respect the unique temporal and gestural nature of sign language, ensuring that critical features such as hand shapes, movements, and facial expressions are accurately captured. These tokens can produce robust representations tailored for downstream tasks like alignment and generation. On the natural language side, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide architecture capable of employing off-the-shelf tokenization mechanisms for text, leveraging domain-specific vocabularies and positional encodings to preserve semantic nuances and syntactic structures, which can allow for alignment with the sign language tokens.
[0030] The exemplary tokenization process of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure described above is not just a preprocessing step but can be a fundamental enabler for effective alignment between sign language and natural language. By converting complex input modalities into discrete, representative tokens, the architecture of exemplary embodiments can establish a unified foundation that can bridge the differences in how these modalities encode meaning. This tokenization according to the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure, ensures that sign language’s temporal and gestural nuances, as well as the semantic and syntactic structures of text, can be preserved in a way that allows for meaningful cross-modal interactions.PATENT APPLICATION With this foundation in place, the next important step for the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure is to achieve seamless alignment of these tokens within a shared representation space, ensuring that equivalent concepts in sign and text are closely mapped while respecting their unique temporal dependencies.
[0031] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide anovel SignXFormer model is provided, which leverages a cross-attention mechanism (see, e.g., A. Vaswani et al., 2017; and E. Pinyoanuntapong et al., 2024) to facilitate the alignment of multimodal representations, including latent embeddings of speech, text, and video.The SignXFormer model is further configured to incorporate temporal information, thereby ensuring coherence and consistency in sequential data processing. Temporal coherence is a modality-invariant attribute, making it essential for bridging the gap between sign language and text representations. This mechanism ensures that temporal dependencies and semantic relationships are preserved, laying the foundation for effective cross-modal alignment.
[0032] Upon alignment of the representation space, the SignXFormer model, functionized as a bidirectional transformer (see, e.g., H. Bao et al., 2021), is configured to bridge multiple modalities, thereby enabling accurate and fluent sign-to-text translation for Sign Language Translation (SLT) and text-to-sign generation for Sign Language Production (SLP). The exemplary systems, methods, computer-accessible media, and devices described herein leverage this approach to enhance multimodal integration and improve overall system performance. The SignXFormer can employ a cross-attention mechanism to compute relevance between the two modalities, aggregating information from one modality based on the context provided by the other. The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can train the model to learn to align semantically equivalent concepts within a shared unified representation space, ensuring that both the static semantics of tokens and their temporal dependencies are effectively captured.
[0033] The architecture according to the exemplary systems, methods, computer-accessible medium, and devices described in the exemplary embodiments of the present disclosure incorporates two novel modules, SignXFormer and HierSignFormer, which function cooperatively to enhance multimodal language processing, including but not limited to speech,PATENT APPLICATION text, and video. SignXFormer is implemented utilizing a cross-attention mechanism and a masked bidirectional transformer architecture, enabling the prediction of masked, missing, or corrupted components within a given modality, such as natural language tokens and sign tokens, by leveraging a maximum-likelihood guidance framework. This architecture facilitates multiway transition pathways, allowing for flexible and contextually adaptive transformations across different modalities. Meanwhile, HierSignFormer is structured to progressively reconstruct high-frequency components of sign language signals through a hierarchical, incremental refinement process. Utilizing signal decomposition techniques (see, e.g., N. Zeghidour et al. 2021; and C. Guo et al., 2024), HierSignFormer iteratively refines sign token predictions to maintain contextual consistency and fluency in both sign-to-text translation (SLT) and text-to-sign generation (SLP). By integrating SignXFormer and HierSignFormer, the disclosed system enhances robustness, accuracy, and multimodal alignment, thereby facilitating seamless and coherent sign language translation across multiple communication modalities.
[0034] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide a design that emphasizes modularity and extensibility, allowing for future expansion to include speech and other modalities. However, the exemplary model’s initial focus on sign to-text and text-to-sign translation ensures that it can serve as a powerful tool for improving accessibility and communication for the deaf community. By integrating advanced tokenization, alignment mechanisms, and conditional generation, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure provides a state-of-the-art solution for cross-modal translation and communication.Exemplary Multimodal Tokenization
[0035] Sign language and natural language representations are inherently temporal and continuous in nature. These data modalities can often exhibit high dimensionality, such as the intricate temporal and gestural dynamics of sign language or the symbolic structure and syntactic dependencies of natural language. Directly learning patterns from such data is challenging, as the continuous and high-dimensional nature of these inputs can result in redundant information and inefficiencies in the learning process.
[0036] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure’s use of tokenization serves as a crucialPATENT APPLICATION step in addressing these challenges by transforming raw input modalities into discrete, structured representations. By discretizing the continuous latent space embeddings into discrete tokens, tokenization can reduce the complexity of raw data while retaining its essential semantic and structural features. This lower-dimensional, discrete representation can not only simplify the learning process but can also facilitate early multimodal alignment by operating on compact tokens rather than directly aligning high-dimensional continuous representations.EXEMPLARY SIGN LANGUAGE TOKENIZATION
[0037] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide a Hierarchical Sign Language Tokenizer 200, which leverages signal frequency decomposition and a residual implicit hierarchical latent representation, drawing upon theoretical foundations similar to those described in prior works (see, e.g., N. Zeghidour et al., 2021, and C. Guo et al., 2024). This tokenizer 200 is configured to transform continuous and temporal sign sequences 210 into discrete, compact, and semantically meaningful representations. By capturing multiple levels of detail from low to high frequency, it effectively preserves both global structural patterns and fine-grained motion details, thereby enhancing the expressiveness, robustness, and accuracy of sign language processing. The architecture is illustrated in Figure 2.
[0038] Latent Embedding Representation: Let the input sign sequence (210) m =[m1, m2,..., mT] ∈ ℝT x Dconsist of T frames, where each frame mt∈ ℝDis a D-dimensional sign feature vector derived from preprocessed skeletal data or visual sign features. An encoder (220) ℰ(·) transforms this sequence into a latent representation z = [z1, z2,..., zT] ∈ ℝT x d(230) where d is the dimensionality of the latent space:z=
[0039] Exemplary Hierarchical Vector Quantization: The representation z can be encoded iteratively into L discrete levels (e.g., 230, 232... 234). At each level ℓ, a code vector cℓcan be selected from a pre-defined dictionary= {dℓk}풟ℓk=1(e.g., 240, 250... 260) which contains all possible code vectors for that level. Formally, the encoding process at level ℓ is expressed as:cℓt= argmin||rℓt— dℓk||22, t = 1, . . .PATENT APPLICATION where: rℓtrepresents the Cross-to-fine feature vector for frame t at level ℓ, dℓkis the k-th vector in the code dictionary for level ℓ, T is the total number of frames in the input sequence.The hierarchical embedding for different level is obtained iteratively as:rℓ+1t= rℓt− cℓt, with r0t= zt.
[0040] Exemplary Sign Language Generation from Tokens: After processing through all L levels, the final discrete representation ẑ (270) can be reconstructed by aggregating the codes across all levels. The decoder 풟(·) (280) is responsible for transforming ẑ back into the original sign sequence m̂. Formally, the tokenized feature embedding for the / -th frame is expressed as:ẑt= ∑where cℓtrepresents the quantized code vector from level ℓ for frame t. The reconstructed sign sequence can then be obtained as:where 풟(·) denotes the decoder 280 that takes the aggregated latent representation ẑ and generates the reconstructed sequence m̂ = [m̂1, m̂2,..., m̂T. The decoder 풟(·) (280) is parameterized as a sequence modeling network, which can be implemented using a stack of feedforward layers or recurrent modules such as transformers. Formally:t = 1,..., T,where g( ; ΘD) is the decoding function with parameters ΘD.
[0041] Exemplary Sign Motion Dictionary Via Sign Temporal and Contrastive Learning:The process of understanding and generating sign language involves complex temporal and spatial coordination that captures the interplay of hand gestures, facial expressions, and body movements. These sign language are inherently high-dimensional and continuous, posing significant challenges in representation and generalization. Directly learning sign language from such raw data often leads to inefficiencies, as the resulting representations are entangled withPATENT APPLICATION irrelevant variations and lack interpretability.
[0042] To address these limitations, the exemplary embodiments of the present disclosure can provide the Sign Motion Dictionary, a structured latent representation framework designed to enforce a one-to-one temporal mapping between latent encodings and their corresponding motions. The Sign Motion Dictionary can ensure that each latent representation at a given time step decodes exclusively to its corresponding motion, preventing interference from non-temporally aligned motions. To further maintain temporal consistency, the exemplary framework employs contrastive learning, ensuring that corresponding latent-motion pairs remain closely aligned, while non-corresponding pairs are explicitly separated, thereby improving motion clarity and representation precision. By structuring motion representations in a temporally ordered latent space, the Sign Motion Dictionary enhances motion consistency, recognition accuracy, and synthesis fidelity, significantly improving sign-to-text (SLT) and text-to-sign (SLP) translation. This novel approach advances sign language processing, human-computer interaction, and embodied AI applications, ensuring a more structured, interpretable, and temporally coherent representation of sign language motions..
[0043] The exemplary systems, methods, computer-accessible medium, and devices described in the present disclosure can incorporate a contrastive loss formulation 290 to develop a novel loss function for Sign Motion Dictionary discovery. In particular, a Contrastive Loss function can serve as an effective optimization objective to promote the discovery of diverse and informative motion representations in sign language processing. By leveraging mutual information as a guiding principle, the disclosed system can facilitate the learning of structured and reusable sign motion patterns while ensuring predictability and consistency in motion representation. This approach is particularly advantageous for sign language skill abstraction, as it is essential that the discovered motion patterns are not only diverse and representative but also maintain semantic coherence and temporal alignment with the observed sign sequences.
[0044] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can consider the information-theoretic paradigm of mutual information maximization framework to enforce structured motion representation learning in the Sign Motion Dictionary. One of the objectives can be to ensure that each latent encoding uniquely corresponds to its temporally aligned motion, while distinguishing non-corresponding motions through contrastive separation. This formulation enhances thePATENT APPLICATION predictability and consistency of motion representations by preserving semantic and temporal structure within the latent space. To achieve this, the contrastive loss is designed to (i) maximize mutual information between reconstructed motions and their corresponding ground-truth annotations (positive pairs), and (ii) minimize mutual information between reconstructed motions and non-corresponding ground-truth annotations (negative pairs), thereby enforcing contrastive separation. Given a latent embedding Z = {z1, z2,..., zn} and a decoder function D: Z → M that reconstructs motion representations M = {m1, m2,..., mn}, the contrastive loss can be defined as follows:where:- D(zi) is the reconstructed motion representation from the latent embedding Zi,- mi is the corresponding ground-truth motion annotation,- I(D(zi); mj) represents the mutual information between the reconstructed motion D(zi) and the ground-truth motion annotation mj,- T is a temperature scaling parameter that controls the contrastive separation,- The numerator maximizes mutual information for positive pairs (D(zi), mi), ensuring strong alignment,- The denominator minimizes mutual information for negative pairs (D(zi), mj) where j ≠ i, ensuring that non-matching motions remain distinct.
[0045] This mutual information-based contrastive loss ensures that each reconstructed motion remains closely aligned with its intended ground-truth annotation, enforcing a one-to-one temporal mapping, while preventing interference from non-corresponding motions. By structuring motion representations in a temporally ordered latent space, the Sign Motion Dictionary enforces temporal coherence, motion consistency, and structured sign language representation, significantly improving both sign-to-text (SLT) and text-to-sign (SLP) translation accuracy.
[0046] Exemplary Training Objective: The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the presentPATENT APPLICATION disclosure can train the proposed Sign Language Tokenizer along with the encoder ℰ(·) and decoder 풟(·) to minimize a loss function that maximize reconstruction accuracy, embedding consistency, commitment to the codebook and skill abstraction accuracy. The total loss Lvqfor each quantization layer is defined as:&vq ™ £rec 4“ ^emb 4“ ^com + ^contrastive-Lrecis the reconstruction loss that ensures the reconstructed sign motion sequence closely matches the original input:= ||m̂ − m||where m̂ = 풟(ẑ) is the reconstructed sign sequence obtained from the decoder.
[0047] In addition to the reconstruction loss, there is Lembwhich is embedding loss (see, e.g., P. Esser, et al., 2021). Lembpenalizes the discrepancy between the latent embedding and the quantized representation b for each layer ℓ.= stop_grad(bℓ).where the stop grad is the stop gradient flag.
[0048] The third term of Lcomis the commitment loss (see, e.g., P. Esser, et al., 2021) which encourages the encoder to commit to a code in the codebook and reduces codebook utilization inefficiency:= ||stop_grad(rℓt) − bℓ||and the last term is the skill abstraction loss from above.EXEMPLARY NATURAL LANGUAGE TOKENIZATION
[0049] Natural language tokenization can be a fundamental step in transforming textual input into discrete representations suitable for downstream processing. This exemplary process converts sentences, denoted as w = [w1, w2, ..., wL], where wirepresents the i-th word in aPATENT APPLICATION sentence of length L, into a sequence of subword tokens. Each word wiis decomposed into smaller units derived from a predefined vocabulary V, resulting in a tokenized sequence t = {t1, t2,..., tN}, where ti∈ V and N > L. This subword tokenization ensures the ability to handle rare or out- of- vocabulary words while preserving semantic and syntactic structures.
[0050] To retain the positional relationships within the input sequence, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can add a positional encoding to each token embedding (see, e.g., T. Pires, et al., 2019):ht= f(ti) + ptwhere f(ti) represents the embedding of the token ti, and pt denotes its positional encoding. The resulting embeddings, hi, capture both the semantic content of the tokens and their sequential order, making them suitable for integration with cross-modal tasks.
[0051] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can utilize the pretrained tokenizer from multilingual BERT (mBERT) (see, e.g., T. Pires, et al., 2019) to perform natural language tokenization. Specifically, the input sentence w can be mapped into subword tokens using mBERT's pretrained vocabulary VmBERTas follows:This tokenizer has been optimized on large multilingual corpora, ensuring robust handling of linguistic diversity and compatibility with pretrained models.
[0052] The decision to use mBERT's tokenizer is motivated by its compatibility with large language models trained on this tokenization scheme. By aligning the natural language tokenization of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure with mBERT's vocabulary, exemplary embodiments can fully leverage the semantic and syntactic representations encoded during pretraining, ensuring that the resulting embeddings are rich and well-suited for downstream tasks. Additionally, mBERT's sub word-based approach ensures effective handling of multilingual inputs, accommodating the morphological and structural variations across languages. This design choice allows the exemplary systems, methods, computer accessiblePATENT APPLICATION medium, and devices according to the exemplary embodiments of the present disclosure to focus on the cross-modal alignment of natural language tokens with sign language representations without the need to retrain or customize a tokenizer, thereby preserving computational efficiency and model compatibility.
[0053] By leveraging mBERT's tokenizer, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can ensure that the natural language tokens in the framework are not only semantically meaningful but also aligned with the tokenization standards of state-of-the-art pretrained language models.
[0054] This alignment can be critical for achieving robust and effective integration of sign language and natural language modalities in the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure.Exemplary Bidirectional Transformer for Multimodal Alignment
[0055] The quantization-based hierarchical design of the exemplary systems, methods, computer-accessible medium, and devices described in the present disclosure is facilitated by two specialized transformer-based modules: SignXFormer and HierSignFormer. These components are specifically designed to enhance multimodal language processing, including but not limited to speech, text, and video, ensuring a structured and efficient representation of sign language signals. The SignXFormer module is implemented utilizing a cross-attention mechanism and a masked bidirectional transformer architecture, enabling the prediction of masked, missing, or corrupted components within a given modality, such as natural language tokens and sign tokens, by leveraging a maximum-likelihood guidance framework. This module is further configured to facilitate multi-way transition pathways, thereby enabling flexible and contextually adaptive transformations across different modalities. In particular, SignXFormer generates sign tokens corresponding to the base vector quantization (VQ) layer, establishing a foundational representation that preserves the integrity of sign language semantics. Meanwhile, the HierSignFormer module is structured to progressively reconstruct high-frequency components of sign language signals through a hierarchical and incremental refinement process. By leveraging signal decomposition techniques (see, e.g., N. Zeghidour et al., 2021, and H. Bao et al., 2024), HierSignFormer iteratively refines sign token predictions, thereby maintainingPATENT APPLICATION contextual consistency and fluency across sign-to-text translation (SLT) and text-to-sign generation (SLP) tasks. This module is particularly configured to handle subsequent hierarchical layers, ensuring that higher-order linguistic structures and motion dynamics are systematically captured and preserved in sign language representation. Accordingly, the integration of SignXFormer and HierSignFormer within this dual-transformer architecture facilitates a comprehensive and seamless fusion of multimodal information, thereby enhancing the robustness, accuracy, and hierarchical alignment of sign token representations to support improved sign language translation, recognition, and synthesis.
[0056] The exemplary SignXFormer of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can focus on generating sign tokens for the base layer of the VQ quantizer. In this component, a portion of the sign token sequence from the level 0 VQ quantizer can be randomly masked out during training. The transformer can be trained to predict all the masked sign tokens concurrently, conditioned on both the unmasked tokens and the input text. By attending to sign language and text tokens in all directions, the SignXFormer of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can explicitly capture the inherent correlations within the sign language tokens as well as the semantic mappings between sign language and text. This mechanism can enable text-driven parallel decoding during inference. Specifically, in each iteration, the model can concurrently predict multiple high-quality sign language tokens that are consistent with both the textual description and the sign language dynamics. This bidirectional attention can ensure that the generated sign language tokens are semantically aligned with the input text while preserving temporal coherence. The ability of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure to predict multiple tokens in parallel not only enhances the generation speed but also ensures high fidelity, making the SignXFormer an effective component for producing realistic and semantically accurate sign language sequences.
[0057] Once the coarse tokens are generated, the exemplary HierSignFormer of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can take over to predict the hierarchical tokens for the subsequent VQ layers. The exemplary HierSignFormer can operate progressively, using thePATENT APPLICATION token sequence at the current layer as input to predict the tokens at the next layer. By leveraging the hierarchical quantization structure, this transformer can refine the sign language representation at each layer, progressively adding finer details to the sign language sequence. This layered refinement can ensure that the final sign language sequence captures both high-level semantic alignment and low-level dynamic precision. The exemplary HierSignFormer of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can achieve this by iteratively enriching the sign language representation with information preserved in the deep layers of the VQ quantizer. This progressive refinement mechanism can complement the SignXFormer, enabling the bidirectional transformer system to generate high-quality sign language sequences that are temporally coherent and semantically aligned with the textual input.
[0058] Together, the SignXFormer and HierSignFormer form a cohesive framework for aligning multimodal inputs. This bidirectional transformer design provided by the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure, bridges the semantic gap between text and sign language while ensuring high-fidelity and efficient sign language generation across all quantization layers.EXEMPLARY SignXFormer
[0059] The SignXFormer 310 of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can take as input, a combination of sign language tokens 320 and text embeddings 330, as shown in Figure 3. During training, sign language tokens 320 can be derived from the encoder's output via a vector quantizer, resulting in a sequence Y = {yt}Lt=1where L is the sequence length and etrepresents quantized embeddings. Sentence embeddings (340) Wsentencecan be extracted from mBERT (see, e.g., T. Pires, et al., 2019) to capture global textual context and can be prepended to the sign language tokens 320. Additionally, word embeddings Wword= [wf]Ff=1, also obtained from mBERT, can be used to establish local relationships between individual words and sign language tokens via a cross-attention mechanism 350. The complete input sequence is formulated as:X = [Wsentence, s1, s2, ..., sL, [END]]PATENT APPLICATION where [END] is a special token indicating the sign language sequence's endpoint.
[0060] Exemplary Cross-Attention Mechanism: A cornerstone of the SignXFormer module 310 in the exemplary systems, methods, computer-accessible medium, and devices described in the present disclosure is its novel cross-attention mechanism 350, which is the first to be utilized for cross-attending across multiple modalities in sign language processing, including sign language tokens, natural language representations, and spoken language signals. This innovative cross-attention mechanism 350 enables a unified representation by dynamically integrating information across these modalities, thereby addressing the inherent challenges in multimodal sign language understanding and generation. This unique novelty of this cross-attention mechanism 350 lies in its ability to go beyond traditional approaches that treat text as a monolithic input. Unlike conventional models, the exemplary systems, methods, computer-accessible medium, and devices described in the present disclosure leverage a dual-level text representation, comprising both sentence-level and word-level embeddings, thereby enabling a precise alignment of textual semantics with sign language dynamics. This dual-level encoding allows for more granular and context-aware multimodal integration, ensuring that linguistic structures are accurately mapped to the corresponding sign language expressions in both sign-to-text translation (SLT) and text-to-sign generation (SLP).
[0061] The motivation for this exemplary dual-level attention design stems from the distinct roles that words and sentences play in conveying semantic information. Words correspond to specific actions or sub-sequences of sign language, providing fine-grained instructions for localized sign language dynamics. In contrast, sentences offer overarching semantic and temporal constraints, ensuring coherence and logical flow across the entire sign language sequence. By decoupling these two aspects, the attention mechanism of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can capture both local specificity and global continuity, which are critical for generating high-fidelity sign language aligned with text.
[0062] To align word embeddings 360 with sign language tokens 320, the framework of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can employ a cross-attention mechanism 350. Word embeddings 360 derived from mBERT (see, e.g., T. Pires, et al., 2019), denoted as W = [wi, W2,..., Wm], can serve as keys (K) (362) and values (V) (364) in the attention mechanism 350,PATENT APPLICATION while the sign language token embeddings Y^ = [e1(, e2,..., eL] where masked tokens are represented by [MASK], act as the queries (Q). The cross-attention output is computed as:fOKT\ CrossAttention(Q, K, V ) = sottoiax — | V.\ / where dk is the dimensionality of the keys. This operation enables each sign language token to selectively attend to the most relevant words, providing a fine-grained semantic grounding for the sign language generation process.
[0063] For the exemplary sentence-level context, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can integrate a global representation by prepending the sentence embedding s (also derived from mBERT) to the sign language token sequence. This updated sequence, expressed as Y = [s, e1(e2,..., eL], facilitates the sentence embedding to act as a global context vector. Through self-attention 370 within the transformer layers, each sign language token 320, including the masked ones, can interact with both the sentence embedding 340 and other sign language tokens. The self-attention mechanism 370 of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can ensure global coherence and temporal alignment, and can be computed as, e.g.,: / QKT\K, V ) = ■ softmax | j V.where Q = K = V = ¥,
[0064] This combined mechanism can establish a unified flow of information. The crossattention mechanism 350 of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can capture localized dependencies between text and sign language by aligning individual words with corresponding sign language tokens. Meanwhile, the sentence embedding 340 can propagate global context throughout the sign language sequence, ensuring that the predictions are semantically meaningful and temporally consistent. The exemplary systems, methods, computer accessible medium, andPATENT APPLICATION devices according to the exemplary embodiments of the present disclosure’s hierarchical alignment of sign language and text tokens underpins the effectiveness of the proposed transformer and represents a significant advancement in bridging the gap between natural language and sign language synthesis.
[0065] Exemplary Dynamic Masking and Corruption: To train the model effectively, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can apply a dynamic masking strategy to corrupt sign language sequences. The exemplary masking ratio r is determined using a cosine schedule:T:(T) = cos ( ••••"••• ) * r ~ ZO, 1),where T is a random variable sampled from a uniform distribution. The number of masked tokens can be computed as:with L being the sequence length. Masked tokens can be replaced based on the following probabilities: e.g., 80% are replaced with [MASK], 10% with random tokens, and 10% remain unchanged. This corruption process can generate a masked sequence Y^, from the original sequence T
[0066] Exemplary Training Objective: The SignXFormer of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can learn to reconstruct the original sign language sequence K from the corrupted sequence Y^ conditioned on text embeddings W. The reconstruction probability for each token is defined as:Lpa: v^.w) = (J pie,: ya,w),where Y^ represents the unmasked sequence. The exemplary model can be optimized to minimize the negative log-likelihood of the predicted masked tokens:PATENT APPLICATION
[0067] During inference, the exemplary process can begin with a fully masked sequence Y = [ [MASK] ]=1. The exemplary model according to the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can predict tokens iteratively, progressively refining the sequence. In each iteration t, tokens with the lowest prediction confidence can be re-masked and re-predicted. The masking ratio can decrease with each iteration, following a decaying schedule:where ' / 'max is the maximum number of iterations. This iterative decoding strategy can ensure that the generated sign language sequence is semantically consistent with the text input while maintaining temporal coherence.
[0068] The Exemplary SignXFormer of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can serve as an essential component for text-conditioned sign language generation. By combining dynamic masking, semantic cross-attention, and iterative decoding, the model can achieve high-fidelity sign language reconstruction that aligns closely with textual descriptions and preserves sign language dynamics.Exemplary HierSignF ormer
[0069] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can provide an exemplary HierSignFormer 400 that can be designed to model and predict tokens for the Hierarchical Vector Quantization layers 410 beyond the base layer in a hierarchical manner, as shown in Figure 4. This exemplary architecture can build upon the exemplary SignXFormer, incorporating distinct mechanisms to account for hierarchical dependencies among Hierarchical Vector Quantization layers 410.PATENT APPLICATION
[0070] For a sign language sequence, let the tokens for the exemplary Hierarchical Vector Quantization layer be denoted as bf.-r= [bf, b|,..., b ] where T represents the sequence length. The embeddings from all preceding quantization layers up to layer — 1 are aggregated as:providing cumulative context from all prior layers. Additionally, the text embedding w, derived from mBERT, can encode the semantic context of the associated natural language input. A learnable layer indicator embedding r, corresponding to the one-hot encoded index of the Hierarchical Vector Quantization layer ■£, can also be included to specify the layer being modeled.
[0071] The aggregated input token embedding for the - -th layer is represented as:X, 'p — 7“ F 7“ W.This embedding integrates hierarchical, semantic, and positional context for the current layer.
[0072] The HierSignFormer 400 predicts the tokens for the current layer by minimizing the negative log-likelihood of the predicted tokens. Formally, the exemplary training objective is given by:T^Hier = -JEbe?^ 10gp(bf | xf.r, W)t=lwhere bf is the true token at timestep t, and D denotes the dataset.
[0073] To enhance contextual reasoning, a masking strategy is employed during training. A subset of tokens b^ in the current hierarchical layer i is masked, following a cosine schedule:PATENT APPLICATION where y(r) determines the masking ratio. The corrupted token sequence b^ unmasked tokens bgj, and aggregated input x.rare then fed into the HierSignFormer 400 for reconstruction. The predicted sequence b.rapproximates the true tokens as:I X5; T, W).Exemplary Conditional GenerationEXEMPLARY TEXT-TO-SIGN GENERATION
[0074] The exemplary text-to-sign generation process 500 of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can involve a hierarchical pipeline that combines the SignXF ormer 310 and HierSignFormer 400 to progressively generate sign language tokens aligned with textual input, as shown in Figure 5. This methodology ensures a seamless transformation of linguistic descriptions into high-fidelity sign language sequences while maintaining semantic and temporal consistency.
[0075] The exemplary process 500 shown in Figure 5 can begin with a textual input sequence 510, which can be tokenized 520 using the pretrained tokenizer from mBERT. This tokenization 520 can yield both sentence-level embeddings 525 that capture global semantic context and word-level embeddings 530 that can provide localized, temporally relevant information. These embeddings can be essential for aligning textual semantics with sign language dynamics.Simultaneously, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can prepare an initial sequence of masking tokens 535, representing an empty sign language canvas to be filled during the generation process.
[0076] In the first exemplary stage, the exemplary SignXFormer 310 of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can predict the base-level sign language tokens 540 corresponding to the fundamental features of the sign language. By attending to the sentence embeddings 525, the model can ensure global semantic consistency, while the cross-attention mechanism with word embeddings 530 can capture the localized relationships betweenPATENT APPLICATION individual words and segments of the sign language. This stage enables the model according to the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure, to establish the foundation of the sign language sequence, leveraging the interplay between masked tokens, unmasked sign language features, and text.
[0077] Once the base sign language tokens 540 are generated by the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure, the HierSignFormer 400 of exemplary embodiments can be employed to predict tokens 545 for the hierarchical quantization layers. For each layer, the tokens from all preceding layers can be embedded and aggregated to form the input representation. This aggregated input, along with the text embeddings and a layer-specific indicator, can guide the HierSignFormer 400 of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure in predicting tokens 545 for the current Hierarchical layer. This iterative process can continue until tokens for all layers have been generated, ensuring that finer sign language details are progressively incorporated.
[0078] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can then map tokens across all layers to their corresponding entries in the quantization codebooks learned during training. Each predicted token can index a specific vector in the codebook, and the vectors from all layers can be aggregated to reconstruct the sign language feature sequence 550. This aggregation ensures that the hierarchical structure of the quantization process is preserved, resulting in a comprehensive representation of the sign language.
[0079] Further, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can pass the reconstructed sign language features through a decoder 560, which can transform them into the final sign language sequence 570. This end-to-end process can integrate the textual semantics with sign language dynamics, enabling the generation of sign language that are both semantically faithful and temporally coherent.
[0080] By leveraging this structured pipeline, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the presentPATENT APPLICATION disclosure can achieve efficient and accurate text-to-sign generation, highlighting the interplay between hierarchical quantization, advanced transformer architectures, and learned sign language representations. This approach ensures a robust mapping from linguistic input to sign language output, facilitating effective communication across modalities.EXEMPLARY SIGN-TO-TEXT GENERATION
[0081] The process of Sign-to-Text generation for the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can leverage the reverse dynamics of the text-to-sign framework, adapting the pipeline to map sign language tokens into corresponding linguistic representations. By switching the roles of text and sign language tokens, the generation methodology of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can align sign language inputs with text outputs, ensuring semantic fidelity and fluency in the resulting language.
[0082] The exemplary process can begin by encoding the input sign language sequence into discrete sign language tokens using the learned vector quantization codebooks from the training phase. The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can map each segment of the sign language sequence to a corresponding token at each quantization layer. The aggregated tokens from all quantization layers can form the initial representation of the sign language, encapsulating both coarse and fine-grained details.
[0083] To initiate the text generation process, the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can input the sequence of sign language tokens into the SignXFormer according to exemplary embodiments, that can now serve as a decoder to map sign language semantics into text representations. The sign language tokens can act as the contextual basis for the transformer, while the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can iteratively predict masked text tokens. The self-attention mechanism can ensure global coherence across the predicted text tokens, while the cross-attention mechanism can dynamically attend to sign language tokens at both global (sentence embedding) and local (word embedding) levels. This can allow the model to accurately capture the temporal and semantic relationships between the sign language and correspondingPATENT APPLICATION linguistic expressions.
[0084] For enhanced granularity, the exemplary HierSignFormer of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can refine the token predictions by processing the quantization layers of the sign language representation. At each layer, the transformer can integrate the encoded information from preceding sign language layers, progressively aligning it with the evolving textual context. This iterative refinement can ensure that complex and nuanced sign language details are faithfully translated into text, preserving the richness of the original input.
[0085] The exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can subsequently map the predicted text tokens to their corresponding indices in the linguistic vocabulary. Each token index can represent a specific word or symbol, which can be sequentially decoded to form the final textual output. The decoder can generate the complete text representation, ensuring both grammatical correctness and semantic alignment with the input sign language.
[0086] By leveraging the hierarchical quantization structure and the exemplary bidirectional transformer architecture, the Sign-to-Text generation process of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure can effectively bridge the gap between sign language and natural language. This approach can enable the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure to handle the inherent complexity of sign language, producing textual outputs that are coherent, meaningful, and linguistically accurate. The reverse generation pipeline not only underscores the flexibility of the exemplary systems, methods, computer accessible medium, and devices according to the exemplary embodiments of the present disclosure, but also demonstrates its efficacy in facilitating multimodal translation.
[0087] According to the exemplary embodiments of the present disclosure, numerous specific details have been set forth. It is to be understood, however, that implementations of the disclosed technology can be practiced without these specific details. In other instances, well-known methods, structures and techniques have not been shown in detail in order not to obscure an understanding of this description. References to “some examples,” “other examples,” “one example,” “an example,” “various examples,” “one embodiment,” “an embodiment,” “somePATENT APPLICATION embodiments,” “example embodiment,” “various embodiments,” “one implementation,” “an implementation,” “example implementation,” “various implementations,” “some implementations,” etc., indicate that the implementation(s) of the disclosed technology so described may include a particular feature, structure, or characteristic, but not every implementation necessarily includes the particular feature, structure, or characteristic. Further, repeated use of the phrases “in one example,” “in one exemplary embodiment,” or “in one implementation” does not necessarily refer to the same example, exemplary embodiment, or implementation, although it may.
[0088] As used herein, unless otherwise specified the use of the ordinal adjectives “first,” “second,” “third,” etc., to describe a common object, merely indicate that different instances of like objects are being referred to, and are not intended to imply that the objects so described must be in a given sequence, either temporally, spatially, in ranking, or in any other manner.
[0089] While certain implementations of the disclosed technology have been described in connection with what is presently considered to be the most practical and various implementations, it is to be understood that the disclosed technology is not to be limited to the disclosed implementations, but on the contrary, is intended to cover various modifications and equivalent arrangements included within the scope of the appended claims. Although specific terms are employed herein, they are used in a generic and descriptive sense only and not for purposes of limitation.
[0090] The foregoing merely illustrates the principles of the disclosure. Various modifications and alterations to the described embodiments will be apparent to those skilled in the art in view of the teachings herein. It will thus be appreciated that those skilled in the art will be able to devise numerous systems, arrangements, and procedures which, although not explicitly shown or described herein, embody the principles of the disclosure and can be thus within the spirit and scope of the disclosure. Various different exemplary embodiments can be used together with one another, as well as interchangeably therewith, as should be understood by those having ordinary skill in the art. In addition, certain terms used in the present disclosure, including the specification and drawings, can be used synonymously in certain instances, including, but not limited to, for example, data and information. It should be understood that, while these words, and / or other words that can be synonymous to one another, can be used synonymously herein, that there can be instances when such words can be intended to not be used synonymously.PATENT APPLICATION Further, to the extent that the prior art knowledge has not been explicitly incorporated by reference herein above, it is explicitly incorporated herein in its entirety. All publications referenced are incorporated herein by reference in their entireties.
[0091] Throughout the disclosure, the following terms take at least the meanings explicitly associated herein, unless the context clearly dictates otherwise. The term “or” is intended to mean an inclusive “or.” Further, the terms “a,” “an,” and “the” are intended to mean one or more unless specified otherwise or clear from the context to be directed to a singular form.
[0092] This written description uses examples to disclose certain implementations of the disclosed technology, including the best mode, and also to enable any person skilled in the art to practice certain implementations of the disclosed technology, including making and using any devices or systems and performing any incorporated methods. The patentable scope of certain implementations of the disclosed technology is defined in the appended claims, and may include other examples that occur to those skilled in the art. Such other examples are intended to be within the scope of the appended claims if they have structural elements that do not differ from the literal language of the appended claims, or if they include equivalent structural elements with insubstantial differences from the literal language of the appended claims.PATENT APPLICATION EXEMPLARY REFERENCES
[0093] The following references are hereby incorporated by reference, in their entireties:Bauman, H. D. and Murray, J. Reframing: From hearing loss to deaf gain. Deaf studies digital journal, 1(1):1-10, 2009.Camgoz, N. C., Hadfield, S., Koller, 0., Ney, H., and Bowden, R. Neural sign language translation.In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 7784- 7793, 2018.Camgoz, N. C., Koller, O., Hadfield, S., and Bowden, R. Sign language transformers: Joint end- to-end sign language recognition and translation. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp. 10023-10033, 2020.Chen, Y., Wei, F., Sun, X., Wu, Z., and Lin, S. A simple multi-modality transfer learning baseline for sign language translation. In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition, pp. 5120--5130, 2022a.Chen, Y., Zuo, R., Wei, F., Wu, Y., Liu, S., and Mak, B. Two-stream network for sign language recognition and translation. Advances in Neural Information Processing Systems, 35:17043- 17056, 2022b.Huang, W., Pan, W., Zhao, Z., and Tian, Q. Towards fast and high-quality sign language production. In Proceedings of the 29th ACM International Conference on Multimedia pp.3172-3181,2021.Kushalnagar, R. Deafness and hearing loss. Web accessibility: A foundation for research, pp. 35- 47, 2019.Luey, H. S., Glass, L., and Elliott, H. Hard-of-hearing or deaf: Issues of ears, language, culture, and identity. Social Work, 40(2): 177-182, 1995.PATENT APPLICATION Saunders, B., Camgoz, N. C., and Bowden, R. Progressive transformers for end-to-end sign language production. In Computer Vision-ECCV 2020: 16th European Conference, Glasgow, UK, August 23-28, 2020, Proceedings, Part XI 16, pp. 687-705. Springer, 2020.Saunders, B., Camgoz, N. C., and Bowden, R. Mixed signals: Sign language production via a mixture of motion primitives. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 1919-1929, 2021.Saunders, B., Camgoz, N. C., and Bowden, R. Signing at scale: Learning to co-articulate signs for large-scale photo-realistic sign language production. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 5141-5151, 2022.Stoll, S., Camgoz, N. C., Hadfield, S., and Bowden, R. Sign language production using neural machine translation and generative adversarial networks. In BMVC, volume 2019, pp. 1-12, 2018.Xie, P., Zhang, Q., Taiying, P., Tang, H., Du, Y., and Li, Z. G2p-ddm: Generating sign pose sequence from gloss sequence with discrete diffusion model. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 38, pp. 6234-6242, 2024.Zhou, B., Chen, Z., Gapes, A., Wan, J., Liang, Y., Escalera, S., Lei, Z., and Zhang, D. Gloss-free sign language translation: Improving from visual-language pretraining. In Proceedings of the IEEE / CVF International Conference on Computer Vision, pp. 20871-20881, 2023.Zhou, H., Zhou, W., Qi, W., Pu, J., and Li, H. Improving sign language translation with monolingual data by sign back-translation. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition, pp. 1316-1325, 2021a.Zhou, H., Zhou, W., Zhou, Y., and Li, H. Spatial-temporal multi-cue network for sign language recognition and translation. IEEE Transactions on Multimedia, 24:768- 779, 2021b.PATENT APPLICATION Schmidt, F., Generalization in generation: A closer look at exposure bias. arXiv preprint arXiv: 1910.00292, 2019.Vaswani, A., N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. Gomez, L. Kaiser, and I.Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017Pinyoanuntapong, E., Wang, P., Lee, M., & Chen, C. Mmm: Generative masked motion model.In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 1546-1555), 2024.Bao, H., Dong, L., Piao, S., & Wei, F. BEiT: BERT Pre-Training of Image Transformers.In International Conference on Learning Representations, 2022Esser, P., Rombach, R., & Ommer, B.. Taming transformers for high-resolution image synthesis.In Proceedings of the IEEE / CVF conference on computer vision and pattern recognition (pp. 12873-12883), 2021.Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., & Tagliasacchi, M.. Soundstream: An end- to-end neural audio codec. IEEE / ACM Transactions on Audio, Speech, and Language Processing, 30, 495-507, 2021.Pires, T.. How multilingual is multilingual BERT. arXiv preprint arXiv: 1906.01502, 2019Guo, C., Mu, Y., Javed, M. G, Wang, S., & Cheng, L.. Momask: Generative masked modeling of 3d human motions. In Proceedings of the IEEE / CVF Conference on Computer Vision and Pattern Recognition (pp. 1900-1910), 2024
Claims
PATENT APPLICATION WHAT IS CLAIMED IS:
1. A method for generating bi-directional real-time speech and sign language conversion, comprising:receiving speech data in real-time;converting the speech data into a plurality of speech tokens;identifying a relevance between the speech tokens and a plurality of sign tokens; and generating, in real-time and based on the identified relevance, a visual avatar performing the generated sign language corresponding to the received speech data.
2. The method of claim 1, wherein the converting of the speech data into the plurality of speech tokens and the identifying of the relevance is performed by a learning model which is trained to align semantically equivalent concepts within a shared unified representation space.
3. The method of claim 2, wherein the trained learning model is configured to capture (i) static semantics of the speech tokens and sign tokens, and (ii) temporal dependencies of the speech tokens and sign tokens.
4. The method of claim 1, further comprising predicting one or more missing or corrupted speech tokens which are associated with the plurality of speech tokens.
5. The method of claim 1, wherein the converting of the speech into the plurality of speech tokens comprises progressively refining a token prediction by capturing a high-frequency and a low-frequency component of a representation.
6. The method of claim 1, wherein the visual avatar performing the generated sign language corresponding to the received speech is displayed on a device proximal to a user generating the received speech data.PATENT APPLICATION 7. The method of claim 1, wherein the visual avatar performing the generated sign language corresponding to the received speech is integrated within a podium from which a user generating the received speech data is located.
8. The method of claim 1, wherein the visual avatar performing the generated sign language corresponding to the received speech data is displayed on one or more smart glasses.
9. A method for generating bi-directional real-time speech and sign language conversion, comprising:observing a plurality of sign language signs in real-time;converting the observed signs into a plurality of sign tokens;identifying a relevance between the sign tokens and a plurality of speech tokens; and generating, in real-time and based on the identified relevance, one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs.
10. The method of claim 9, wherein the converting of the speech data into the plurality of speech tokens and the identifying of the relevance is performed by a learning model which is trained to align semantically equivalent concepts within a shared unified representation space.
11. The method of claim 10, wherein the trained learning model is configured to capture (i) static semantics of the speech tokens and sign tokens, and (ii) temporal dependencies of the speech tokens and sign tokens.
12. The method of claim 9, wherein the observing the plurality of sign language signs in realtime includes a learning model capable of detecting 3 -dimensional skeletal pose data.
13. The method of claim 9, wherein converting the observed signs into a plurality of sign tokens distinguishes and considers a unique temporal and gestural nature of the plurality of sign language signs.PATENT APPLICATION 14. The method of claim 9, further comprising predicting one or more missing or corrupted sign tokens which are associated with the plurality of sign tokens.
15. The method of claim 9, wherein the converting of the observed signs into the plurality of sign tokens comprises progressively refining a token prediction by capturing a high-frequency and a low-frequency component of a representation.
16. The method of claim 9, wherein the one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs is (i) displayed on a device proximal to a user generating the plurality of sign language signs, or (ii) played through a speaker proximal to the user generating the plurality of sign language signs.
17. The method of claim 9, wherein the one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs is (i) displayed on one or more smart glasses for written text, or (ii) played through a speaker on the one or more smart glasses for spoken text.
18. A system for generating bi-directional real-time speech and sign language conversion, comprising:at least one processor configured to:receive speech data in real-time;convert the speech data into a plurality of speech tokens;identify a relevance between the speech tokens and a plurality of sign tokens; and generate, in real-time and based on the identified relevance, a visual avatar performing the generated sign language corresponding to the received speech data.
19. The system of claim 18, wherein the converting of the speech data into the plurality of speech tokens and the identifying of the relevance is performed by a learning model which is trained to align semantically equivalent concepts within a shared unified representation space.PATENT APPLICATION 20. The system of claim 19, wherein the trained learning model is configured to capture (i) static semantics of the speech tokens and sign tokens, and (ii) temporal dependencies of the speech tokens and sign tokens.
21. The system of claim 18, further comprising predicting one or more missing or corrupted speech tokens which are associated with the plurality of speech tokens.
22. The system of claim 18, wherein the converting of the speech into the plurality of speech tokens comprises progressively refining a token prediction by capturing a high-frequency and a low-frequency component of a representation.
23. The system of claim 18, wherein the visual avatar performing the generated sign language corresponding to the received speech is displayed on a device proximal to a user generating the received speech data.
24. The system of claim 18, wherein the visual avatar performing the generated sign language corresponding to the received speech is integrated within a podium from which a user generating the received speech data is located.
25. The system of claim 18, wherein the visual avatar performing the generated sign language corresponding to the received speech data is displayed on one or more smart glasses.
26. A system for generating bi-directional real-time speech and sign language conversion, comprising:at least one processor configured to:observe a plurality of sign language signs in real-time;convert the observed signs into a plurality of sign tokens;identify a relevance between the sign tokens and a plurality of speech tokens; and generate, in real-time and based on the identified relevance, one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs.PATENT APPLICATION27. The system of claim 26, wherein the converting of the speech data into the plurality of speech tokens and the identifying of the relevance is performed by a learning model which is trained to align semantically equivalent concepts within a shared unified representation space.
28. The system of claim 27, wherein the trained learning model is configured to capture (i) static semantics of the speech tokens and sign tokens, and (ii) temporal dependencies of the speech tokens and sign tokens.
29. The system of claim 26, wherein the observing the plurality of sign language signs in realtime includes a learning model capable of detecting 3 -dimensional skeletal pose data.
30. The system of claim 26, wherein converting the observed signs into a plurality of sign tokens distinguishes and considers a unique temporal and gestural nature of the plurality of sign language signs.
31. The system of claim 26, further comprising predicting one or more missing or corrupted sign tokens which are associated with the plurality of sign tokens.
32. The method of claim 26, wherein the converting of the observed signs into the plurality of sign tokens comprises progressively refining a token prediction by capturing a high-frequency and a low-frequency component of a representation.
33. The method of claim 26, wherein the one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs is (i) displayed on a device proximal to a user generating the plurality of sign language signs, or (ii) played through a speaker proximal to the user generating the plurality of sign language signs.
34. The method of claim 26, wherein the one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs is (i) displayed on one or morePATENT APPLICATION smart glasses for written text, or (ii) played through a speaker on the one or more smart glasses for spoken text.
35. A non-transitory computer-accessible medium having stored thereon computer-executable instructions for generating bi-directional real-time speech and sign language conversion, which when executed by a computer arrangement, configure the computer arrangement to perform procedures comprising:receiving speech data in real-time;converting the speech data into a plurality of speech tokens;identifying a relevance between the speech tokens and a plurality of sign tokens; and generating, in real-time and based on the identified relevance, a visual avatar performing the generated sign language corresponding to the received speech data.
36. The non-transitory computer-accessible medium of claim 35, wherein the converting of the speech data into the plurality of speech tokens and the identifying of the relevance is performed by a learning model which is trained to align semantically equivalent concepts within a shared unified representation space.
37. The non-transitory computer-accessible medium of claim 36, wherein the trained learning model is configured to capture (i) static semantics of the speech tokens and sign tokens, and (ii) temporal dependencies of the speech tokens and sign tokens.
38. The non-transitory computer-accessible medium of claim 35, further comprising predicting one or more missing or corrupted speech tokens which are associated with the plurality of speech tokens.
39. The non-transitory computer-accessible medium of claim 35, wherein the converting of the speech into the plurality of speech tokens comprises progressively refining a token prediction by capturing a high-frequency and a low-frequency component of a representation.PATENT APPLICATION 40. The non-transitory computer-accessible medium of claim 35, wherein the visual avatar performing the generated sign language corresponding to the received speech is displayed on a device proximal to a user generating the received speech data.
41. The non-transitory computer-accessible medium of claim 35, wherein the visual avatar performing the generated sign language corresponding to the received speech is integrated within a podium from which a user generating the received speech data is located.
42. The non-transitory computer-accessible medium of claim 35, wherein the visual avatar performing the generated sign language corresponding to the received speech data is displayed on one or more smart glasses.
43. A non-transitory computer-accessible medium having stored thereon computer-executable instructions for generating bi-directional real-time speech and sign language conversion, which when executed by a computer arrangement, configure the computer arrangement to perform procedures comprising:observing a plurality of sign language signs in real-time;converting the observed signs into a plurality of sign tokens;identifying a relevance between the sign tokens and a plurality of speech tokens; and generating, in real-time and based on the identified relevance, one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs.
44. The non-transitory computer-accessible medium of claim 43, wherein the converting of the speech data into the plurality of speech tokens and the identifying of the relevance is performed by a learning model which is trained to align semantically equivalent concepts within a shared unified representation space.
45. The non-transitory computer-accessible medium of claim 44, wherein the trained learning model is configured to capture (i) static semantics of the speech tokens and sign tokens, and (ii) temporal dependencies of the speech tokens and sign tokens.PATENT APPLICATION 46. The non-transitory computer-accessible medium of claim 43, wherein the observing the plurality of sign language signs in real-time includes a learning model capable of detecting 3-dimensional skeletal pose data.
47. The non-transitory computer-accessible medium of claim 43, wherein converting the observed signs into a plurality of sign tokens distinguishes and considers a unique temporal and gestural nature of the plurality of sign language signs.
48. The non-transitory computer-accessible medium of claim 43, further comprising predicting one or more missing or corrupted sign tokens which are associated with the plurality of sign tokens.
49. The non-transitory computer-accessible medium of claim 43, wherein the converting of the observed signs into the plurality of sign tokens comprises progressively refining a token prediction by capturing a high-frequency and a low-frequency component of a representation.
50. The non-transitory computer-accessible medium of claim 43, wherein the one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs is (i) displayed on a device proximal to a user generating the plurality of sign language signs, or (ii) played through a speaker proximal to the user generating the plurality of sign language signs.
51. The non-transitory computer-accessible medium of claim 43, wherein the one or more of a spoken text or a written text corresponding to the observed plurality of sign language signs is (i) displayed on one or more smart glasses for written text, or (ii) played through a speaker on the one or more smart glasses for spoken text.