Method and apparatus for generating regression-based gloss-free sign language pose using discrete representation

The method addresses the need for initial pose information in autoregressive models by converting spoken sentences into discretized tokens for sign language generation, achieving accurate and natural sign language translation using a Transformer-based encoder-decoder architecture.

WO2026089111A1PCT designated stage Publication Date: 2026-04-30KOREA ADVANCED INST OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
KOREA ADVANCED INST OF SCI & TECH
Filing Date
2024-11-22
Publication Date
2026-04-30

AI Technical Summary

Technical Problem

Existing autoregressive sign language generation models require initial sign language pose and timing information during inference, complicating true autoregression on continuous datasets without auxiliary information.

Method used

A regression-based method using a Transformer-based encoder-decoder architecture that converts spoken sentences into discretized sign language pose tokens through a self-supervised learning model, enabling direct translation into sign language without commentary, utilizing discrete representations and Beam Search for optimal pose generation.

Benefits of technology

Generates accurate and natural sign language poses by aligning linguistic features with spatial and temporal characteristics, facilitating smoother communication between hearing and deaf individuals.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2024018644_30042026_PF_FP_ABST
    Figure KR2024018644_30042026_PF_FP_ABST
Patent Text Reader

Abstract

The present invention relates to a method and apparatus for generating a regression-based gloss-free sign language pose using discrete representation that generates a sign language pose on the basis of (gloss-free) spoken sentence data, and the method comprises the steps of: receiving spoken sentence data as an input and transforming the spoken sentence data into a discretized sign language pose token according to spatial and temporal characteristics by using a self-supervised learning model; and automatically recursively generating a sign language pose through a transformer-based encoder-decoder architecture by using the discretized sign language pose token as an input.
Need to check novelty before this filing date? Find Prior Art

Description

Regression-based Gloss-Free Sign Pose Generation Method and Device Using Discrete Representation

[0001] The present invention relates to a regression-based gloss-free sign language pose generation method and apparatus utilizing discrete representation, and more specifically, to a technology for generating sign language poses based on spoken sentence data without commentary (gloss-free).

[0002] Sign language is an essential language for communication among the Deaf. It possesses a unique grammatical system and, unlike spoken language, incorporates spatial and temporal aspects. Sign language interpreters play a crucial role in facilitating communication between the Deaf and hearing populations, yet their numbers are significantly insufficient compared to the deaf population. To address this issue, much research related to sign language has investigated methods to generate sign language poses from spoken language.

[0003] Sign language glosses are character-based representations that correspond one-to-one with actual sign language movements, serving as specific annotations for real sign language sequences. However, obtaining these annotations requires significant effort, time, and expertise in sign language. Because this demands substantial resources, there is growing interest in research methods that do not use sign language glosses. This is referred to as gloss-free sign language research.

[0004] In the field of gloss-free sign language generation, two approaches are used: search models and generative models. Search models extract relevant samples from a dataset based on given spoken language, whereas generative models can generate new sign language sequences by utilizing patterns acquired during the training process.

[0005] Existing autoregression-based sign language generation models require initial sign language pose and timing information during the inference process. This dependency complicates performing true autoregression generation on continuous datasets, such as sign language poses (keypoints), without auxiliary information.

[0006] Accordingly, the present invention proposes introducing a quantization method that converts continuous sign language poses into a discrete form.

[0007] The objective of the present invention is to generate sign language poses using an autoregressive method through a Transformer-based encoder-decoder architecture, after converting spoken sentences into discretized sign language pose tokens based on spatial and temporal characteristics using a self-supervised learning model during the process of directly translating speech sentences into sign language without commentary (gloss-free).

[0008] The objective of the present invention is to resolve the problem of discrepancy between spoken sentences and sign language while maintaining the natural flow of sign language poses through automatic regressive sign language pose generation by utilizing discretized sign language pose tokens.

[0009] However, the technical problems that the present invention aims to solve are not limited to the above problems, and can be expanded in various ways without departing from the technical concept and scope of the present invention.

[0010] A regression-based gloss-free sign language pose generation method utilizing discrete representation comprising at least one processor according to an embodiment of the present invention comprises the steps of: receiving spoken sentence data as input and converting the spoken sentence data into discrete sign language pose tokens according to spatial and temporal characteristics using a self-supervised learning model; and automatically regressively generating sign language poses through a transformer-based encoder-decoder architecture using the discrete sign language pose tokens as input.

[0011] A regression-based gloss-free sign language pose generation device utilizing discrete representation according to an embodiment of the present invention comprises: a token conversion unit that receives spoken sentence data as input and converts the spoken sentence data into discrete sign language pose tokens according to spatial and temporal characteristics using a self-supervised learning model; and a sign language pose generation unit that automatically and regressively generates sign language poses through a transformer-based encoder-decoder architecture using the discrete sign language pose tokens as input.

[0012] According to an embodiment of the present invention, in the process of directly translating spoken sentences into sign language without commentary (gloss-free), a self-supervised learning model is used to convert them into discretized sign language pose tokens based on spatial and temporal characteristics, and then sign language poses are generated in an autoregressive manner through a transformer-based encoder-decoder architecture, thereby enabling the generation of accurate sign language translation and natural sign language poses.

[0013] According to an embodiment of the present invention, linguistic characteristics and the characteristics of quantized sign language poses are directly matched to enable more accurate and natural sign language generation.

[0014] According to an embodiment of the present invention, unlike existing sign language pose generation models, the introduction of a knowledge distillation technique directly intervenes knowledge within the sign language pose distribution in the mapping process to generate realistic and accurate sign language poses, thereby enabling hearing people without knowledge of sign language to communicate freely with deaf people, as well as facilitating smoother and real-time communication with deaf people.

[0015] However, the effects of the present invention are not limited to the above effects and can be extended in various ways without departing from the technical concept and scope of the present invention.

[0016] Figure 1 illustrates a flowchart of the operation of a regression-based gloss-free sign pose generation method utilizing discrete representation according to an embodiment of the present invention.

[0017] Figure 2 illustrates a schematic diagram of a sign language vector quantization network (SignVQNet) according to an embodiment of the present invention.

[0018] FIG. 3 illustrates a schematic diagram of an STGP block according to an embodiment of the present invention.

[0019] Figure 4 is a table showing the statistics of a sign language data set according to an embodiment of the present invention.

[0020] Figure 5 is a table showing the results of comparing the existing method and the method of the present invention (SignVQNet).

[0021] Figure 6 illustrates a graph illustrating the performance discrepancy between SLP indicators according to the change in beam size of beam search according to an embodiment of the present invention.

[0022] Figures 7a and 7b illustrate a visual comparison between the baselines of the conventional method and the method of the present invention (SignVQNet).

[0023] Figures 8a to 8d show the results of the ablation experiment for PHOENIX14T in a table.

[0024] FIG. 9 is a block diagram illustrating the detailed configuration of a regression-based gloss-free sign pose generation device utilizing discrete representation according to an embodiment of the present invention.

[0025] The advantages and features of the present invention and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the accompanying drawings. However, the present invention is not limited to the embodiments disclosed below but may be implemented in various different forms. These embodiments are provided merely to ensure that the disclosure of the present invention is complete and to fully inform those skilled in the art of the scope of the invention, and the present invention is defined only by the scope of the claims.

[0026] The terms used herein are for describing the embodiments and are not intended to limit the invention. In this specification, the singular form includes the plural form unless specifically stated otherwise in the text. As used herein, "comprises" and / or "comprising" do not exclude the presence or addition of one or more other components, steps, actions, and / or elements to the mentioned components, steps, actions, and / or elements.

[0027] Unless otherwise defined, all terms used in this specification (including technical and scientific terms) may be used in a meaning commonly understood by those skilled in the art to which the present invention pertains. Additionally, terms defined in commonly used dictionaries are not to be interpreted ideally or excessively unless explicitly and specifically defined otherwise.

[0028] Hereinafter, preferred embodiments of the present invention will be described in more detail with reference to the attached drawings. Identical components in the drawings are denoted by the same reference numerals, and redundant descriptions of identical components are omitted.

[0029]

[0030] The gist of the present invention is that it enables natural sign language generation without a gloss (an intermediary for commentary or interpretation) by receiving spoken sentence data as input, converting the spoken sentence into discretized sign language pose tokens using a self-supervised learning autoencoder model, and automatically and regressively generating sign language poses through a transformer-based encoder-decoder architecture.

[0031] Existing autoregressive Sign Language Production (SLP) methods have not fully achieved true autoregression because they often rely on ground truth data during inference. To fill this gap, the present invention introduces the Sign Language Vector Quantization Network (SignVQNet) to utilize discrete spatiotemporal representations of sign language poses. Through these discrete representations, the present invention integrates Beam Search, a decoding strategy widely used in natural language processing. Furthermore, it aligns the discrete representations with the linguistic features of pre-trained language models such as BERT.

[0032] Accordingly, the present invention is characterized by generating accurate sign language translation and natural poses, and effectively reflecting the spatial and temporal characteristics of sign language data.

[0033] The present invention will be described in detail below with reference to FIGS. 1 to 9.

[0034]

[0035] Figure 1 illustrates a flowchart of the operation of a regression-based gloss-free sign pose generation method utilizing discrete representation according to an embodiment of the present invention.

[0036] The method of FIG. 1 is performed by a regression-based gloss-free sign pose generation device (900) utilizing a discrete representation according to an embodiment of the present invention illustrated in FIG. 9.

[0037] Referring to Fig. 1, in step S110, spoken sentence data is received.

[0038] Spoken sentence data refers to sentence data of spoken language, which is the actual language used by people in everyday conversation and may include sentences employed to convey specific meanings. Examples include conversations used in daily life, speeches, presentations, or lectures delivered in public, questions and answers in interviews, and the transmission of information through broadcasting or radio.

[0039] In step S120, a self-supervised learning model is used to convert spoken sentence data into discretized sign language pose tokens based on spatial and temporal characteristics.

[0040] In step S120, the present invention can convert spoken sentence data into discretized sign language pose tokens by using an autoencoder of a self-supervised learning model based on a Spatio-Temporal Graph Pyramid (STGP) to convert sign language poses into discrete tokens by processing the spatial and temporal characteristics (spatio-temporal) of the sign language poses.

[0041] Here, the autoencoder of a self-supervised learning model based on a Spatio-Temporal Graph Pyramid (STGP) is characterized by enabling automatic regression generation by utilizing vector quantization to convert sign language poses into discretized sign language pose tokens of individual tokens.

[0042] The spatiotemporal graph pyramid can effectively discretize sign language poses during the process of converting spoken sentence data into sign language. More specifically, the spatiotemporal graph pyramid is used to convert sign language pose sequences into discretized sign language pose tokens, thereby effectively encapsulating the spatial and temporal characteristics of sign language and enabling the Sign Language Vector Quantization Network (SignVQNet) to receive clearer input. In this regard, the present invention may utilize STGP blocks necessary to process both spatial and temporal aspects and preserve important information. The STGP blocks process sign language pose sequences spatially and temporally to capture complex movements of sign language, and can pursue more stable token representations by using Gumbel-Softmax during training.

[0043] In step S130, a sign language pose is automatically and regressively generated through a transformer-based encoder-decoder architecture using a discretized sign language pose token as input.

[0044] In step S130, the present invention can generate sign language poses using a latent alignment loss that performs alignment between the linguistic features of a pre-trained language model and discretized sign language pose tokens. This process includes encoding sign language poses into tokens using a previously trained autoencoder, and using these tokens, sign language poses can be generated automatically and regressively.

[0045] In this case, the present invention can select the optimal sign language pose from among various sign language pose sequences by applying Beam Search in a Transformer-based encoder-decoder architecture. Here, Beam Search is a search algorithm for generating the optimal sign language pose sequence; it helps find the optimal output while maintaining the most probable candidates and can assist in effective conversion between natural language and sign language by using discrete representations. More specifically, Beam Search can explore multiple possible sign language pose tokens from the current state and calculate a probability for each token. In this case, the probability may be a value predicted by the Transformer-based encoder-decoder architecture. Accordingly, the optimal sign language pose can be selected and output by selecting candidates with the highest probability from the Beam Search sign language pose sequences.

[0046] Accordingly, the present invention can generate a more sophisticated and detailed high-quality sign language pose sequence by performing beam search using discretized sign language pose tokens. The discretized sign language pose tokens are converted via STGP, and beam search can be performed based thereon.

[0047]

[0048] FIG. 2 illustrates a schematic diagram of a sign vector quantization network (SignVQNet) according to an embodiment of the present invention. FIG. 3 also illustrates a schematic diagram of an STGP block according to an embodiment of the present invention.

[0049] The present invention, using the SignVQNet vector quantization network illustrated in Fig. 2, converts sign language pose sequences into individual tokens to enable automatic regression generation. This approach supports beam search, which is commonly used in Natural Language Processing (NLP) tasks, and introduces latent-level alignment to directly associate linguistic features with sign language pose features.

[0050] Referring to FIG. 2, the present invention relates to spoken sentence data composed of U words. Considers the SLPsms sign language pose sequence of the present invention. We intend to generate a , where V represents the number of vertices and C represents the feature dimension of the skeletal pose data. Instead of direct modeling, the present invention uses an intermediary representation z composed of individual tokens. These tokens can encapsulate both the spatial and temporal attributes of the sign language. Then, the generation process is defined by the co-distribution p(y, z|x) = q(y|z, x)p(z|x). Here, p(z|x) represents the probability of generating the discrete representation z from the input x, and q(y|z, x) represents the probability of generating the continuous sign language pose sequence y based on z and x. Furthermore, the first term q(y|z, x) is processed by a vector quantization model, and the second term p(z|x) is modeled in an autoregressive manner.

[0051]

[0052] In addition, the present invention proposes a dVAE based on the Spatio-Temporal Graph Pyramid (STGP) illustrated in FIG. 2 to convert a sign language pose sequence into individual tokens. Referring to the graph pyramid, the present invention proposes an STGP block, which is the basic building block of an encoder (210) and a decoder (220), to resolve the complexity of downsampling and upsampling within a skeletal graph characterized by a non-uniform grid. A detailed overview of the STGP block is illustrated in FIG. 3. The STGP block processes the input sequence sequentially to handle both spatial and temporal aspects. Each processing ends with a residual connection to preserve downsampled spatiotemporal information.

[0053] The model of the present invention It is designed to process fixed-length sign language segments denoted by , where L represents the window size. A sign language segment refers to a small part of an entire sign language pose sequence. The center of STGP-dVAE is a codebook composed of latent variable categories, It is denoted as. Here, K represents the number of latent variable categories, and D C represents the embedding size. Then, the output of the STGP encoder is discretized using Gumbel-Softmax relaxation. This process is the encoder e i Latent variables can be sampled from the output as follows.

[0054] [Mathematical Formula 1]

[0055]

[0056] Here, g i represents the independent i-th sample of the Gumbel distribution. Parameter adjusts the approximation for the categorical distribution, and w i represents the weights for the codebook vectors. These discretized representations are used for “sign language tokens (or discretized sign language pose tokens)” in subsequent training. Then, the sampled latent vectors are It is given by.

[0057] The model of the present invention is optimized by minimizing a combined loss function composed of reconstruction and various losses. The L2 loss is used as the reconstruction loss and is defined as follows.

[0058] [Mathematical Formula 2]

[0059]

[0060] Here, and and represent the actual value (ground truth) and the reconstructed sign language segment, respectively. Various losses support the model in effectively using the codebook and can be expressed as follows.

[0061] [Mathematical Formula 3]

[0062]

[0063] Accordingly, the final loss can be defined as follows.

[0064] [Mathematical Formula 4]

[0065]

[0066] Here, α represents a hyperparameter that determines the magnitude of diversity loss.

[0067]

[0068] Additionally, the present invention uses a transformer-based encoder-decoder architecture (230) as illustrated in FIG. 2 to generate sign language tokens (or discretized sign language pose tokens) from given spoken sentence data and to generate autoregressive sign language poses using them. The first step is to convert a sign language pose sequence into sign language tokens. This process involves dividing the input sign language pose sequence y into several sign language segments, each of which has a length L. Thus, the number of sign language tokens is This can be, and as a result It can be. For computational efficiency, the remaining sign language poses are simply removed. Afterwards, each segment is encoded by a pre-trained STGP encoder (210) through argmax operations z i = argmax(h i Generates sign language tokens indicated by ), where h i represents the i-th hidden representation of the STGP encoder (210). To indicate the start and end of the sine, z is <bos>and <eox>It is padded.

[0069] This model is optimized by minimizing a combined loss function composed of cross-entropy (CE) and latent alignment loss. The CE loss can be expressed as follows.

[0070] [Mathematical Formula 5]

[0071]

[0072] Here, is the previous token z <i와 입력 x가 주어졌을 때 i번째 토큰 zi를 생성할 확률을 나타낸다.

[0073] The potential loss using L2 loss aligns the output of the transformer decoder with the output of the pre-trained STGP encoder (210) to provide a supplementary potential level signal. This can be defined as follows.

[0074] [Mathematical Formula 6]

[0075]

[0076] Here, represents the i-th hidden representation of the transformer decoder. The total loss is the sum of the two previously mentioned losses and is expressed as follows.

[0077] [Mathematical Formula 7]

[0078]

[0079] Here, β represents a hyperparameter that extends potential loss.

[0080]

[0081] FIG. 4 is a table illustrating the statistics of a sign language data set according to an embodiment of the present invention. In FIG. 4, NoF represents the number of frames.

[0082] To evaluate the regression-based gloss-free sign language pose generation technology utilizing discrete representation proposed in this invention, two sign language datasets, RWTH-PHOENIX-WEATHER-2014T (PHOENIX14T) and How2Sign, were used. Details of each dataset are shown in Figure 4. PHOENIX14T is a weather forecast dataset in German Sign Language (DGS). This dataset contains 8,257 pairs of German words and word-level annotations of the corresponding DGS videos. How2Sign is a large-scale American Sign Language (ASL) dataset containing 2,500 educational videos. Since keypoints are not provided for PHOENIX14T, OpenPose and a skeletal correction model were used.

[0083] The present invention preprocessed the aforementioned dataset. During preprocessing, joint coordinates were centered and normalized with respect to the shoulder joint to ensure consistency across all poses. This step ensures that the length from one shoulder to the other is consistently scaled to 1. Additionally, to further refine the data, noise frames were removed, and the remaining joint coordinates were normalized within the range [-1, 1] to maintain consistent scaling and positioning. The text was converted to lowercase and then tokenized using Byte-Pair Encoding (BPE). The vocabulary size of this encoding was set to 3,000 for the PHOENIX14T dataset and 10,000 for the How2Sign dataset.

[0084] In the present invention, the Gumbel-Softmax relaxation technique was used during the experiment, and the temperature (τ) was gradually reduced from 0.9 to 0.1. The parameters α and β were set to 0.1 and 0.001, respectively. In addition, the present invention used 4 STGP blocks, and when constructing the Transformer model, the hidden size was set to 768, with 4 layers and 8 attention heads, a dropout rate of 0.1, and an intermediate size of 1,024. Furthermore, the AdamW optimizer was used, and the learning rate was set to 0.0001. To encode spoken sentences, the present invention used pre-trained BERT1 (bert-base-cased and bert-base-german-cased) and fine-tuned during training. The present invention selected checkpoints that minimized the score for the FGD metric, the entire training process was conducted for about 24 hours on a Tesla A100 GPU, and the batch size was 64.

[0085] The present invention evaluated the method of the present invention using various evaluation metrics.

[0086] Sign language pose sequences generated using Back-Translation (BT) were translated back into spoken language, and BLEU-4 was calculated by comparing them with the original text. Joint-SLT was trained on PHOENIX14T and How2Sign as the Back-Translation model. Additionally, the present invention measured the difference between the predicted sign language pose sequences and the actual sign language pose sequences using DTW-MJE, which combines Dynamic Time Warping (DTW) and Mean Joint Error (MJE). Furthermore, the visual fidelity of the generated sign language pose sequences was evaluated by comparing the distribution of the actual sequences and the generated sequences using Frechet Gesture Distance (FGD).

[0087]

[0088] FIG. 5 is a table showing the results of comparing the conventional method and the method of the present invention (SignVQNet), FIG. 6 is a graph showing the performance discrepancy between SLP indicators according to the change in beam size of beam search according to an embodiment of the present invention, FIG. 7a and FIG. 7b are visual comparisons between the baselines of the conventional method and the method of the present invention (SignVQNet), and FIG. 8a to FIG. 8d are tables showing the results of ablation experiments on PHOENIX14T.

[0089] The present invention compared the method of the present invention with existing methods (PT and NSLP-G), which are SLP methods for generating gloss-free sign language. PT used default settings and GN&FP (Gaussian and Future Prediction) settings and was modified to exclude the use of additional real data for a fair comparison. NSLP-G used frozen and fine-tuning options during training.

[0090] As shown in Fig. 5, it can be seen that the present invention (SignVQNet) demonstrates improved performance compared to existing methods in terms of FGD and BLEU-4 in two datasets (in Fig. 5, the best results are indicated in bold, and the next best results are indicated underlined). In How2Sign (see Fig. 4), which features longer frame sequences, it can be seen that the present invention (SignVQNet) demonstrates strength in generating extended pose sequences. However, while the present invention (SignVQNet) showed superior performance in FGD and BLEU-4, it was found to be inferior to NSLP-G in DTW-MJE, which is predicted to be mainly due to the unique evaluation method of DTW-MJE.

[0091] In the field of SLP (Sign Language Generation), selecting a reliable metric is crucial for comprehensively evaluating the performance of generative models. Currently, metrics such as FGD, DTW-MJE, and BT are used, but finding the optimal metric for comprehensive evaluation remains an unresolved challenge. This issue is evident in the contrasting results between the present invention (SignVQNet) and NSLP-G. These differences stem primarily from the differences in the loss functions used by the two models. For example, the present invention (SignVQNet) uses CE loss, focusing on sequential prediction accuracy and sign language structure preservation; this emphasis on sequential structure can affect model performance in the DTW-MJE metric, which primarily evaluates the spatial accuracy of joint positions. On the other hand, NSLP-G tends to achieve better scores in DTW-MJE by using MSE loss, which focuses on spatial accuracy on a frame-by-frame basis.

[0092] To investigate these differences more deeply, the present invention further analyzed the impact on model performance by varying the beam size in the PHOENIX14T. As shown in Fig. 6, FGD and BLEU-4 generally improve as the beam size increases, whereas DTW-MJE tends to decrease. Additionally, DTW was included in the analysis for a more comprehensive evaluation. Since DTW-MJE operates by forcibly aligning the generated sign language pose sequence with the actual sign language pose sequence, it may not always provide an accurate comparison. Specifically, DTW attempts to minimize distance by aligning the sequences, but this may not always reflect actual temporal alignment. When combined with MJE, this can lead to inconsistent error measurements. This demonstrates the need for careful consideration when developing new metrics tailored to specific problems in SLP evaluation.

[0093] The present invention visually compared the sign language pose sequences generated by the present invention (SignVQNet) with existing methods in PHOENIX14T and How2Sign. As highlighted in the dotted boxes in Figures 7a and 7b, it can be seen that the method of the present invention (Ours) generates sign language pose sequences such as "Morning," "Sunday," "Me," and "Work" that are more realistic and accurate than existing methods. Figure 7a illustrates a visual comparison between the baselines of PHOENIX14T and the present invention (Ours), and Figure 7b illustrates a visual comparison between the baselines of How2Sign and the present invention (Ours).

[0094] This invention investigated the effects of various components of the method and design choices on PHOENIX14T, the most widely used dataset in sign language research. Figure 8a shows the window size, Figure 8b shows the loss type, Figure 8c shows the architecture type, and Figure 8d shows the codebook size. According to Figure 8c, it can be seen that the Transformer encoder-decoder model demonstrates superior performance compared to GRU-based networks commonly used for human motion tasks, and that performance is significantly improved when using a pre-trained BERT. Regarding the loss function, and The combination showed the best performance, which can be seen in Fig. 8b. Optimal performance was achieved when the window size L was set to 32 (see Fig. 8a), and the codebook size of 1,024 showed the best performance (see Fig. 8d).

[0095]

[0096] As described above, the present invention proposes SignVQNet, a gloss-free sign language generation (SLP) model that derives discrete tokens from sign language pose sequences using vector quantization. This enables true autoregression without essential data during inference and resolves the shortcomings of previous autoregressive SLP models. In experiments, the present invention demonstrated state-of-the-art performance by achieving superior performance compared to existing methods such as PHOENIX14T and How2Sign. Furthermore, the reliability of BT and FGD as evaluation metrics was emphasized, and discrepancies in the DTW-MJE metric were also mentioned.

[0097]

[0098] FIG. 9 is a block diagram illustrating the detailed configuration of a regression-based gloss-free sign pose generation device utilizing discrete representation according to an embodiment of the present invention.

[0099] Referring to FIG. 9, a regression-based gloss-free sign language pose generation device utilizing discrete representation according to an embodiment of the present invention generates sign language poses based on spoken sentence data without gloss (commentary).

[0100] To this end, a regression-based gloss-free sign language pose generation device (900) utilizing a discrete representation according to an embodiment of the present invention includes a token conversion unit (910) and a sign language pose generation unit (920).

[0101] The token conversion unit (910) receives spoken sentence data.

[0102] Spoken sentence data refers to sentence data of spoken language, which is the actual language used by people in everyday conversation and may include sentences employed to convey specific meanings. Examples include conversations used in daily life, speeches, presentations, or lectures delivered in public, questions and answers in interviews, and the transmission of information through broadcasting or radio.

[0103] Subsequently, the token conversion unit (910) uses a self-supervised learning model to convert spoken sentence data into discretized sign language pose tokens according to spatial and temporal characteristics.

[0104] The token conversion unit (910) can convert spoken sentence data into discrete sign language pose tokens by using an autoencoder of a self-supervised learning model based on a spatio-temporal graph pyramid (STGP) to convert sign language poses into discrete sign language pose tokens by processing the spatial and temporal characteristics (spatio-temporal) of the sign language poses.

[0105] Here, the autoencoder of a self-supervised learning model based on a Spatio-Temporal Graph Pyramid (STGP) is characterized by enabling automatic regression generation by utilizing vector quantization to convert sign language poses into discretized sign language pose tokens of individual tokens.

[0106] The spatiotemporal graph pyramid can effectively discretize sign language poses during the process of converting spoken sentence data into sign language. More specifically, the spatiotemporal graph pyramid is used to convert a sign language pose sequence into discretized sign language pose tokens, thereby effectively encapsulating the spatial and temporal characteristics of sign language, allowing the Sign Language Vector Quantization Network (SignVQNet) to receive clearer input. At this time, the token conversion unit (910) may use STGP blocks necessary to process both spatial and temporal aspects and preserve important information. The STGP blocks process the sign language pose sequence spatially and temporally to capture complex movements of sign language, and can pursue a more stable token representation by using Gumbel-Softmax during training.

[0107] The sign language pose generation unit (920) takes a discretized sign language pose token as input and automatically and regressively generates a sign language pose through a transformer-based encoder-decoder architecture.

[0108] The sign language pose generation unit (920) can generate sign language poses using a latent alignment loss that performs alignment between the linguistic features of a pre-trained language model and discretized sign language pose tokens. This process includes encoding sign language poses into tokens using a previously trained autoencoder, and using these tokens, can generate sign language poses automatically and regressively.

[0109] At this time, the sign language pose generation unit (920) can select the optimal sign language pose among various sign language pose sequences by applying beam search in a transformer-based encoder-decoder architecture. Here, beam search is a search algorithm for generating the optimal sign language pose sequence, which helps find the optimal output while maintaining the most likely candidates, and can assist in effective conversion between natural language and sign language using discrete representations.

[0110] More specifically, Beam Search can explore multiple possible sign language pose tokens from the current state and calculate the probability for each token. In this case, the probability may be a value predicted by a Transformer-based encoder-decoder architecture. Accordingly, the optimal sign language pose can be selected and output by choosing candidates with the highest probabilities from the Beam Search sign language pose sequences.

[0111] Accordingly, the regression-based gloss-free sign language pose generation device (900) utilizing discrete representation according to an embodiment of the present invention can generate a more sophisticated and detailed high-quality sign language pose sequence by performing beam search using discrete sign language pose tokens. The discrete sign language pose tokens are converted through STGP, and beam search can be performed based thereon.

[0112]

[0113] Although the description of the device in FIG. 9 has been omitted, each constituent means constituting FIG. 9 may include all the contents described in FIG. 1 to FIG. 8d, which is obvious to those skilled in the art.

[0114] The system or device described above may be implemented as a hardware component, a software component, and / or a combination of a hardware component and a software component. For example, the device and component described in the embodiments may be implemented using one or more general-purpose or special-purpose computers, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a Field Programmable Gate Array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing and responding to instructions. The processing unit may execute an operating system (OS) and one or more software applications executed on said operating system. Additionally, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit may be described as being used as a single unit, but those skilled in the art will understand that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, the processing unit may include multiple processors or one processor and one controller. Additionally, other processing configurations, such as parallel processors, are also possible.

[0115]

[0116] Software may include computer programs, code, instructions, or a combination of one or more of these, and may configure a processing unit to operate as desired or command the processing unit independently or collectively. Software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave so as to be interpreted by the processing unit or to provide instructions or data to the processing unit. Software may be distributed over networked computer systems and may be stored or executed in a distributed manner. Software and data may be stored on one or more computer-readable recording media.

[0117]

[0118] The method according to the embodiment may be implemented in the form of program instructions that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program instructions, data files, data structures, etc., either individually or in combination. The program instructions recorded on the medium may be those specifically designed and configured for the embodiment, or they may be those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include hard disks, magneto-optical media, solid-state drives (SSDs), and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. The hardware devices described above may be configured to operate as one or more software modules to perform the operation of the embodiment, and vice versa.

[0119] Although the embodiments have been described above with reference to limited examples and drawings, those skilled in the art can make various modifications and variations from the description above. For example, suitable results can be achieved even if the described techniques are performed in a different order than described, and / or the components of the described system, structure, device, circuit, etc. are combined or assembled in a form different from described, or replaced or substituted by other components or equivalents.

[0120]

[0121] Therefore, other implementations, other embodiments, and equivalents to the claims also fall within the scope of the claims set forth below.< / eox> < / bos>

Claims

1. A regression-based gloss-free sign pose generation method utilizing a discrete representation including at least one processor, A step of receiving spoken sentence data as input and converting the spoken sentence data into discretized sign language pose tokens according to spatial and temporal characteristics using a self-supervised learning model; and A step of automatically and regressively generating sign language poses through a Transformer-based encoder-decoder architecture using the above discretized sign language pose tokens as input. A regression-based gloss-free sign pose generation method utilizing a discrete representation including 2. In Paragraph 1, The step of converting into the above-mentioned discretized sign language pose tokens A regression-based gloss-free sign language pose generation method utilizing discrete representations, wherein the spoken sentence data is processed into the spatio-temporal characteristics of sign language poses and converted into discrete sign language pose tokens using an autoencoder of the self-supervised learning model based on the Spatio-Temporal Graph Pyramid (STGP).

3. In Paragraph 2, The autoencoder of the above self-supervised learning model is Enabling automatic regression generation by utilizing vector quantization to convert sign language poses into the aforementioned discretized sign language pose tokens of individual tokens. A regression-based gloss-free sign pose generation method utilizing discrete representations, characterized by 4. In Paragraph 1, The step of generating the above sign language pose is A regression-based gloss-free sign language pose generation method utilizing discrete representations, which generates the sign language pose using a latent alignment loss that performs alignment between the linguistic features of a pre-trained language model and the discrete sign language pose token.

5. In Paragraph 1, The step of generating the above sign language pose is Selecting the optimal sign language pose from among various sign language pose sequences by applying beam search in the above-mentioned transformer-based encoder-decoder architecture. A regression-based gloss-free sign pose generation method utilizing discrete representations, characterized by 6. In a regression-based gloss-free sign pose generation device utilizing discrete representation, A token conversion unit that receives spoken sentence data as input and converts the spoken sentence data into discretized sign language pose tokens according to spatial and temporal characteristics using a self-supervised learning model; and A sign language pose generation unit that automatically and regressively generates sign language poses through a Transformer-based encoder-decoder architecture using the above-mentioned discretized sign language pose token as input. A regression-based gloss-free sign pose generation device utilizing a discrete representation including 7. In Paragraph 6, The above token conversion unit A regression-based gloss-free sign language pose generation device utilizing discrete representations, which processes spoken sentence data into the spatio-temporal characteristics of sign language poses and converts them into discrete sign language pose tokens using an autoencoder of the self-supervised learning model based on the Spatio-Temporal Graph Pyramid (STGP).

8. In Paragraph 7, The autoencoder of the above self-supervised learning model is Enabling automatic regression generation by utilizing vector quantization to convert sign language poses into the aforementioned discretized sign language pose tokens of individual tokens. A regression-based gloss-free sign pose generation device utilizing discrete representation, characterized by 9. In Paragraph 6, The above sign language pose generating unit A regression-based gloss-free sign language pose generation device utilizing discrete representations, which generates the sign language pose using a latent alignment loss that performs alignment between the linguistic features of a pre-trained language model and the discrete sign language pose token.

10. In Paragraph 6, The above sign language pose generating unit Selecting the optimal sign language pose from among various sign language pose sequences by applying beam search in the above-mentioned transformer-based encoder-decoder architecture. A regression-based gloss-free sign pose generation device utilizing discrete representation, characterized by