Robust sign language method for common visual impairment in combination with information bottleneck and disturbance self-alignment
Through the method of self-alignment of information bottlenecks and perturbations, the robustness problem of sign language translation model facing multiple visual damage in the real world is solved, and good robustness is maintained without predicting the damaged information, which is suitable for real-life scenarios.
Patent Information
- Application Number
- CN202510581040.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-07
- Publication Date
- 2025-08-26
AI Technical Summary
Existing sign language translation models lack robustness in the face of common visual damage in the real world and cannot adapt to a variety of complex scenarios of visual damage.
Using the method of self-aligning information bottleneck and perturbation, a robust sign language translation model is generated through sign language corpus modeling, video enhancement, visual feature extraction, filtering, encoding and decoding, perturbation calculation and alignment loss calculation.
Without the need to obtain visual damage information in advance, the robustness of the sign language translation model to various common visual damage in the real world is improved, and is suitable for real-life scenarios.
Smart Images

Figure CN120544263A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of sign language translation, and is directed to research on robustness against visual impairments, particularly a robust sign language method for common visual impairments that combines information bottlenecks with perturbation self-alignment. Background Art
[0002] Sign language is a visual language used by the hearing-impaired in daily community communication. Sign language translation (SLT) plays a vital role in bridging the communication gap between the hearing-impaired and hearing-disabled. Consequently, SLT has garnered increasing research attention. In the field of artificial intelligence (AI), SLT models typically take a sign language video as input and generate a natural-sounding spoken sentence. A practical SLT model should be robust to common visual impairments in the real world, such as rain, snow, lighting, noise, blur, and camera rotation.
[0003] However, existing robustness methods under common visual impairments are not applicable to sign language translation. This is because: these methods only focus on robustness to common visual impairments at the single image level, while the SLT model takes video as input; they can only make the SLT model robust to specific visual impairments, but not to any unknown visual impairments. However, in reality, the SLT model may need to process videos that cover one or more visual impairments. Summary of the Invention
[0004] The main purpose of the present invention is to provide a robust sign language method for common visual impairments that combines information bottlenecks and perturbation self-alignment, thereby solving the problems existing in the prior art and maintaining robustness in the face of various common visual impairments in the real world.
[0005] In order to achieve the above object, the solution of the present invention is: A sign language method that is robust to common visual impairments by combining information bottleneck and perturbation self-alignment, including: (1) Sign language corpus selection and modeling In the video input module, sign language corpus is selected and modeled, and the sign language videos in the sign language corpus are input into the model in the form of video frames; (2) Sign language video enhancement In the video enhancement module, the spatial and temporal dimensions of the sign language video input to the model are enhanced; (3) Sign language visual feature extraction In the visual feature extraction module, a convolutional neural network is used to extract features from each frame of the sign language video to obtain sequence features containing visual information of the sign language; (4) Sign language visual feature filtering Based on the information bottleneck theory, the non-robust information of sign language in the sequence features is filtered out in the information filtering module to obtain the hidden layer features. ; (5) End-to-end sign language video conversion The hidden layer features The data is sent to the sign language encoding module to obtain encoding features, and the sign language decoding module generates the translation text based on the encoding features; (6) Calculation of sign language translation loss The sign language translation loss calculation module calculates the sign language translation loss ;in, Representative model based on sign language video Predict corresponding translation text The probability of Represents the parameters of the model; (7) Calculation of sign language visual feature perturbation The perturbation calculation module generates adversarial samples based on the gradient, and adds the sign language translation gradient to the hidden layer features. ; (8) Frame-level alignment loss calculation The perturbed visual features are aligned with their corresponding clean visual features by a frame-level alignment loss module; (9) Video-level alignment loss calculation The video-level alignment module aligns the clean video with the visually corrupted video by calculating the Kullback-Leibler divergence between the output distribution of the reduced clean visual features and the output distribution of the perturbed visual features; (10) Output of sign language translation results The final translated text is output by the output module.
[0006] Preferably, the sign language corpus Includes a collection of sign language videos , translation text collection , recorded as ;in, The representative frame number is Sign language video, to Represents the sign language video To Frame picture; The representative length is The translated text, to Represents the translated text To words.
[0007] Preferably, the enhancement of the spatial dimension refers to performing data enhancement on each frame of the sign language video to diversify the sign language video; the enhancement of the temporal dimension refers to randomly sampling all frames of the sign language video, and some of the frames obtained by random sampling are used for training.
[0008] Preferably, the information filtering module reduces the filtering loss To increase , and increase by reducing the loss of sign language translation ; Filtration loss The calculation formula is ;in, Represents Gaussian distribution The standard deviation of Represents Gaussian distribution The mean of Representative distribution variance; Representative distribution The mean of the distribution Upsampling to obtain hidden layer features .
[0009] Preferably, the sign language corpus Sign language videos in After modeling, a transformer-based encoder-decoder structure is used for video-to-text conversion; a sign language encoding module using a self-attention mechanism is used to learn meaningful spatiotemporal representations and sign language representations.
[0010] Preferably, the coding features output by the sign language encoding module carry a begin identifier. When the begin identifier is recognized, the sign language decoding module generates a sign language translation based on the coding features output by the sign language encoding module. The sign language decoding module outputs a word from the output sequence at each step in the decoding phase. The output of each step is input to the bottom decoder in the next time step, so that its decoding result is output upward to a higher layer. Position codes are embedded and added to these decoder inputs to indicate the position of each word. This process is repeated until the end identifier appears, indicating that the sign language decoding module has completed the output.
[0011] Preferably, the perturbation calculation module only adds perturbations to the hidden layer features 15% of visual features , the generated perturbed visual features are ;in, , , represents a configurable hyperparameter to indicate the size of the perturbation, Represented by visual features Calculate sign language translation loss for variables gradient; Represents element-wise multiplication of matrices; Represents the mask matrix, the elements in the matrix to When the value of is 1, it means that the frame corresponding to its subscript is selected to add disturbance, otherwise the value is 0.
[0012] Preferably, the set of disturbed visual features is set to , the size of the set is , then there exists Perturbed visual features of the frame and its corresponding clean visual features , ; After the information filtering module and the sign language encoding module, the clean visual features and disturbed visual features The corresponding coding features are 、 , expressed as: ; ; in, represents the encoder; Thus we get the Perturbed coding features of frames and its corresponding clean encoding features ; The frame-level alignment loss module reduces the size of the corresponding encoded features. Distance alignment is achieved, but there is a loss in frame-level alignment .
[0013] Preferably, the video level alignment loss calculated by the video level alignment module is for: ; in, stands for calculating the Kullback-Leibler divergence; Represents the translated text words; Represents the front of the translated text words, not including indivual.
[0014] By adopting the above technical solution, the present invention can maintain good robustness to any common visual impairment in the real world without obtaining relevant information about visual impairment in advance. This allows the SLT model to be used in the real world to cope with various situations where visual impairment may occur in the real world, specifically: ① The present invention filters out the non-robust parts of sign language in sign language video features through an information filtering module, obtaining information that is truly beneficial to sign language translation, thereby improving the robustness of the sign language translation model to common visual impairments. Compared with other robustness methods for common visual impairments, this is a self-supervised training method that does not require obtaining impairment information in advance and is applicable to any common visual impairment.
[0015] ② The present invention simulates the scene of visual damage in sign language video by adding model-generated gradient perturbations to video features, and performs two levels of alignment on such disturbed visual features and corresponding clean visual features, thereby improving the robustness of the sign language translation model to common visual damage.
[0016] ③The present invention fills the gap in the robustness of common visual impairments in SLT tasks, improves the robustness of the SLT model to common visual impairments in the real world, and enables the SLT model to be used in the real world.
[0017] ④ This invention fills the gap in the existing robustness to common visual impairments in video tasks, and provides a method for improving the robustness to common visual impairments for tasks with video as input, providing a reference for other tasks with video as input (such as video retrieval, video-level action recognition, etc.). BRIEF DESCRIPTION OF THE DRAWINGS
[0018] Figure 1 Schematic diagram of the model architecture of a specific embodiment of the present invention.
[0019] Figure 2 Schematic diagram of the system framework of a specific embodiment of the present invention.
[0020] Figure 3 Schematic diagram of the working process of a specific embodiment of the present invention. DETAILED DESCRIPTION
[0021] In order to further explain the technical solution of the present invention, the present invention is described in detail below through specific embodiments.
[0022] refer to Figure 1-3 As shown, the present invention discloses a robust sign language method for common visual impairments that combines information bottleneck and perturbation self-alignment, including: (1) Sign language corpus selection and modeling In the video input module, sign language corpus is selected and modeled, and the sign language videos in the sign language corpus are input into the model in the form of video frames.
[0023] Furthermore, sign language corpus Includes a collection of sign language videos , translated text collection (naturalized spoken sentences) , recorded as: ; in, The representative frame number is Sign language video, to Represents the sign language video To Frame picture; The representative length is The translated text, to Represents the translated text To words.
[0024] (2) Sign language video enhancement The video enhancement module enhances the spatial and temporal dimensions of the sign language video input to the model. Spatial enhancement involves data augmentation of each frame to diversify the video. Temporal enhancement involves random sampling of all frames, with some of the sampled frames used for training. Sign language video enhancement improves the robustness of the present invention.
[0025] Furthermore, the above data enhancement operation includes modifying the color, brightness and other attributes of each frame.
[0026] (3) Sign language visual feature extraction In the visual feature extraction module, a convolutional neural network is used to extract features from each frame of the sign language video to obtain sequence features containing sign language visual information.
[0027] (4) Sign language visual feature filtering Generally speaking, the information contained in the video features of sign language is redundant and needs to be filtered to remove the non-robust information of sign language and retain the information that is truly beneficial to sign language translation. Therefore, based on the information bottleneck method proposed by Naftali Tishby et al. (this is a public prior art), the information filtering module (IB layer) filters out the non-robust information of sign language in the sequence features and obtains the hidden layer features. , hidden layer features Information that is truly beneficial to sign language interpreters is retained. Figure 2 The information filtering module is placed after the visual feature extraction module and before the sign language encoding module. The Transformer-based sign language encoding module can then transform the hidden features into Translate into corresponding translation text .
[0028] Furthermore, inspired by the information bottleneck theory, in step 4, by reducing the hidden layer features Video with sign language Mutual information , increase the hidden layer features With translated text Mutual information Train the information filtering module so that the hidden features obtained by filtering As much as possible, only features that are beneficial to SLT are included, that is, the information filtering module reduces the filtering loss To increase , and increase by reducing the sign language translation loss (see step 6 below) , where the filtration loss The calculation formula is: ; in, Represents Gaussian distribution The standard deviation of Represents Gaussian distribution The mean of Representative distribution variance; Representative distribution The mean of the distribution Upsampling to obtain hidden layer features .
[0029] (5) End-to-end sign language video conversion The hidden layer features The sign language encoding module is used to obtain encoding features, and the sign language decoding module generates translation text based on the encoding features.
[0030] Furthermore, the above sign language corpus Sign language videos in After modeling, a transformer-based encoder-decoder architecture is used for video-to-text conversion. The sign language encoding module, which employs a self-attention mechanism (a technique that enables the model to focus on and fully absorb important information), overcomes the limitations of input sequence length and is used to learn meaningful spatiotemporal and sign language representations. The encoded features output by the sign language encoding module are accompanied by a begin identifier. Upon recognizing the begin identifier, the sign language decoding module generates a sign language translation based on the encoded features output by the sign language encoding module. The sign language decoding module outputs a word from the output sequence at each decoding step. The output of each step is fed into the bottom decoder at the next time step, and the decoded result is output to the next higher layer. Positional encodings are embedded and added to these decoder inputs to indicate the position of each word. This process is repeated until the end identifier is reached, indicating that the sign language decoding module has completed its output.
[0031] (6) Calculation of sign language translation loss The sign language translation loss calculation module calculates the sign language translation loss , expressed as: ; in, Representative model based on sign language video Predict corresponding translation text The probability of Represents the parameters of the model.
[0032] (7) Calculation of sign language visual feature perturbation The perturbation calculation module generates adversarial samples based on the gradient, and adds the sign language translation gradient to the hidden layer features. This is to simulate the situation where sign language videos are subject to common visual perturbations in real life.
[0033] Furthermore, the above perturbation calculation module only adds perturbations to the hidden layer features Visual features The part (accounting for 15%), the generated disturbed visual features are ;in, , , represents a configurable hyperparameter to indicate the size of the perturbation, Represented by visual features Calculate sign language translation loss for variables gradient; Represents element-wise multiplication of matrices; Represents the mask matrix, the elements in the matrix to When the value of is 1, it means that the frame corresponding to its subscript is selected to add disturbance, otherwise the value is 0.
[0034] (8) Frame-level alignment loss calculation A frame-level alignment loss module aligns the perturbed visual features with their corresponding clean visual features, thereby improving the robustness of the SLT model to general visual corruptions.
[0035] Furthermore, the set of disturbed visual features is set to , the size of the set is , then there exists Perturbed visual features of the frame and its corresponding clean visual features , ; After the information filtering module and the sign language encoding module, the clean visual features and disturbed visual features The corresponding coding features are 、 , expressed as: ; ; in, represents the encoder; Thus we get the Perturbed coding features of frames and its corresponding clean encoding features ; The frame-level alignment loss module reduces the size of the corresponding encoded features. Distance alignment is achieved, but there is a loss in frame-level alignment .
[0036] The sign language encoding module uses a self-attention mechanism to model the contextual relationships between frames in a sign language video. The output encoding features include the video frame itself as well as the contextual information between the frames. When a video frame is perturbed, the encoding features obtained after sign language encoding should remain robust to the unperturbed contextual information. With this in mind, the frame-level alignment loss module calculates the perturbed encoding features and their corresponding clean encoding features in the aligned encoding features to improve the robustness of the SLT model to common visual impairments.
[0037] (9) Video-level alignment loss calculation The clean video and the visually corrupted video are aligned by the video-level alignment module by computing the Kullback-Leibler divergence between the output distribution of the reduced clean visual features and the output distribution of the perturbed visual features.
[0038] Furthermore, the video level alignment loss calculated by the video level alignment module is for: ; in, stands for calculating the Kullback-Leibler divergence; Represents the translated text words; Represents the front of the translated text words, not including indivual.
[0039] (10) Output of sign language translation results The final translated text is output by the output module.
[0040] Through the above scheme, the present invention can maintain good robustness to any common visual impairment in the real world without obtaining relevant information about visual impairment in advance. This allows the SLT model to be used in the real world to cope with various situations where visual impairment may occur in the real world, specifically: ① The present invention filters out the non-robust parts of sign language in sign language video features through an information filtering module, obtaining information that is truly beneficial to sign language translation, thereby improving the robustness of the sign language translation model to common visual impairments. Compared with other robustness methods for common visual impairments, this is a self-supervised training method that does not require obtaining impairment information in advance and is applicable to any common visual impairment.
[0041] ② The present invention simulates the scene of visual damage in sign language video by adding model-generated gradient perturbations to video features, and performs two levels of alignment on such disturbed visual features and corresponding clean visual features, thereby improving the robustness of the sign language translation model to common visual damage.
[0042] ③The present invention fills the gap in the robustness of common visual impairments in SLT tasks, improves the robustness of the SLT model to common visual impairments in the real world, and enables the SLT model to be used in the real world.
[0043] ④ This invention fills the gap in the existing robustness to common visual impairments in video tasks, and provides a method for improving the robustness to common visual impairments for tasks with video as input, providing a reference for other tasks with video as input (such as video retrieval, video-level action recognition, etc.).
[0044] The above embodiments and drawings do not limit the product form and style of the present invention. Any appropriate changes or modifications made by ordinary technicians in the relevant technical field should be deemed to be within the patent scope of the present invention.
Claims
1. A robust sign language method for common visual impairments that combines information bottleneck and perturbation self-alignment, characterized by include: (1) Sign language corpus selection and modeling In the video input module, sign language corpus is selected and modeled, and the sign language videos in the sign language corpus are input into the model in the form of video frames; (2) Sign language video enhancement In the video enhancement module, the spatial and temporal dimensions of the sign language video input to the model are enhanced; (3) Sign language visual feature extraction In the visual feature extraction module, a convolutional neural network is used to extract features from each frame of the sign language video to obtain sequence features containing visual information of the sign language; (4) Sign language visual feature filtering Based on the information bottleneck theory, the non-robust information of sign language in the sequence features is filtered out in the information filtering module to obtain the hidden layer features. ; (5) End-to-end sign language video conversion The hidden layer features The data is sent to the sign language encoding module to obtain encoding features, and the sign language decoding module generates the translation text based on the encoding features; (6) Calculation of sign language translation loss The sign language translation loss calculation module calculates the sign language translation loss ;in, Representative model based on sign language video Predict corresponding translation text The probability of Represents the parameters of the model; (7) Calculation of sign language visual feature perturbation The perturbation calculation module generates adversarial samples based on the gradient, and adds the sign language translation gradient to the hidden layer features. ; (8) Frame-level alignment loss calculation The perturbed visual features are aligned with their corresponding clean visual features by a frame-level alignment loss module; (9) Video-level alignment loss calculation The video-level alignment module aligns the clean video with the visually corrupted video by calculating the Kullback-Leibler divergence between the output distribution of the reduced clean visual features and the output distribution of the perturbed visual features; (10) Output of sign language translation results The final translated text is output by the output module.
2. The robust sign language method for common visual impairments combining information bottleneck and perturbation self-alignment as claimed in claim 1, characterized in that: The sign language corpus Includes a collection of sign language videos , translation text collection , recorded as ;in, The representative frame number is Sign language video, to Represents the sign language video To Frame picture; The representative length is The translated text, to Represents the translated text To words.
3. The robust sign language method for common visual impairments combining information bottleneck and perturbation self-alignment as claimed in claim 2, characterized in that: The enhancement of the spatial dimension refers to performing data enhancement on each frame of the sign language video to diversify the sign language video; the enhancement of the temporal dimension refers to randomly sampling all frames of the sign language video, and some of the randomly sampled frames are used for training.
4. The robust sign language method for common visual impairments combining information bottleneck and perturbation self-alignment according to claim 2, characterized in that: The information filtering module reduces the filtering loss by To increase , and increase by reducing the loss of sign language translation ; Filtration loss The calculation formula is ;in, Represents Gaussian distribution The standard deviation of Represents Gaussian distribution The mean of Representative distribution variance; Representative distribution The mean of the distribution Upsampling to obtain hidden layer features .
5. The robust sign language method for common visual impairments combining information bottleneck and perturbation self-alignment according to claim 4, characterized in that: The sign language corpus Sign language videos in After modeling, a transformer-based encoder-decoder structure is used for video-to-text conversion; a sign language encoding module using a self-attention mechanism is used to learn meaningful spatiotemporal representations and sign language representations.
6. The robust sign language method for common visual impairments combining information bottleneck and perturbation self-alignment according to claim 5, characterized in that: The encoding features output by the sign language encoding module carry a begin identifier. When the begin identifier is recognized, the sign language decoding module generates a sign language translation based on the encoding features output by the sign language encoding module. The sign language decoding module outputs a word from the output sequence at each step in the decoding phase. The output of each step is input to the bottom decoder in the next time step, so that its decoding result is output upward to a higher layer. Position codes are embedded and added to these decoder inputs to indicate the position of each word. This process is repeated until the end identifier appears, indicating that the sign language decoding module has completed the output.
7. The robust sign language method for common visual impairments combining information bottleneck and perturbation self-alignment according to claim 5, characterized in that: The perturbation calculation module only adds perturbations to the hidden layer features. 15% of visual features , the generated perturbed visual features are ;in, , , represents a configurable hyperparameter to indicate the size of the perturbation, Represented by visual features Calculate sign language translation loss for variables gradient; Represents element-wise multiplication of matrices; Represents the mask matrix, the elements in the matrix to When the value of is 1, it means that the frame corresponding to its subscript is selected to add disturbance, otherwise the value is 0.
8. The robust sign language method for common visual impairments combining information bottleneck and perturbation self-alignment according to claim 7, characterized in that: Set the set of disturbed visual features to , the size of the set is , then there exists Perturbed visual features of the frame and its corresponding clean visual features , ; After the information filtering module and the sign language encoding module, the clean visual features and disturbed visual features The corresponding coding features are 、 , expressed as: ; ; in, represents the encoder; Thus we get the Perturbed coding features of frames and its corresponding clean encoding features ; The frame-level alignment loss module reduces the size of the corresponding encoded features. Distance alignment is achieved, but there is a loss in frame-level alignment .
9. The robust sign language method for common visual impairments combining information bottleneck and perturbation self-alignment according to claim 8, characterized in that: The video level alignment loss calculated by the video level alignment module for: ; in, stands for calculating the Kullback-Leibler divergence; Represents the translated text words; Represents the front of the translated text words, not including indivual.