Communication method and system for hearing-impaired people based on user instruction emphasis

By using a visual model based on the Transformer architecture and a temporal and spatial emphasis method, combined with deblurring and context distillation techniques, the problem of poor discernibility of sign language movements in sign language videos was solved, and the accuracy of sign language recognition was improved.

CN120412103BActive Publication Date: 2025-10-21COMMUNICATION UNIVERSITY OF CHINA
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510847138.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-21
Estimated Expiration
2045-06-24

AI Technical Summary

Technical Problem

In the existing technology, the sign language movements in sign language video images are poorly distinguishable and the accuracy of sign language movement feature extraction is insufficient, resulting in low sign language recognition accuracy. This is especially evident in real environments where sign language movements are too fast, the shooting is not professional enough, and there are interference from environmental factors.

Method used

A visual model based on the Transformer architecture is used for feature extraction, combined with the temporal difference operator and spatial emphasis method, and the sign language action feature matrix is ​​emphasized through instructions. The deblurred video is processed using a deblurred model, and the context distillation model and speech synthesis technology are combined to achieve accurate extraction and recognition of sign language actions.

Benefits of technology

It significantly improves the accuracy of sign language recognition, solves the problems of poor quality of sign language video images and insufficient mining of key area features, and realizes the precise extraction and recognition of sign language movement features.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412103B_ABST
    Figure CN120412103B_ABST
Patent Text Reader

Abstract

The application provides a deaf-mute communication method and system based on user instruction emphasis, which comprises the following steps: obtaining a sign language video to be processed and user instruction information; using a visual model based on a Transformer architecture to perform feature extraction on the sign language video to be processed to obtain a sign language action feature matrix; based on the sign language action feature matrix, obtaining a sign language action feature vector emphasized by the instruction and a sign language action feature matrix emphasized in space and time respectively; performing feature fusion on the sign language action feature vector emphasized by the instruction and the sign language action feature matrix emphasized in space and time to obtain a fused sign language action feature matrix; and based on the fused sign language action feature matrix, using a preset speech synthesis model to obtain speech information corresponding to the sign language video to be processed. The application achieves the technical effect of significantly improving the accuracy of sign language recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of sign language video technology, and more specifically, to a communication method and system for hearing-impaired people based on user instruction emphasis. Background Art

[0002] In a scenario where the hearing-impaired communicate with ordinary people using sign language, on the one hand, the sign language of the hearing-impaired needs to be recognized and synthesized into speech before being conveyed to ordinary people; on the other hand, the ordinary people's voice messages also need to be recognized and synthesized into sign language animation before being conveyed to the hearing-impaired.

[0003] Existing technology typically trains sign language recognition models based on deep neural networks using sign language training sets from different regions and countries to achieve sign language recognition. However, the following drawbacks remain: 1) In real-world sign language data collection scenarios, the discernibility of sign language movements in sign language video images is poor due to factors such as rapid sign language movements, inadequate filming, and interference from environmental factors, leading to a need for improved accuracy in sign language movement feature extraction. 2) Sign language videos contain features from multiple dimensions, and during the sign language movement feature extraction process, insufficient attention is paid to feature mining in key, information-rich areas such as the hands and face, resulting in a need for improved sign language movement recognition accuracy.

[0004] Therefore, there is an urgent need for a highly accurate communication method for the hearing-impaired. Summary of the Invention

[0005] In view of the above problems, an object of the present invention is to provide a communication method and system for hearing-impaired persons based on user instruction emphasis, so as to solve at least one problem existing in the prior art.

[0006] According to one aspect of the present invention, a method for communication with hearing-impaired persons based on user instruction emphasis is provided, which is applied to electronic devices, including: obtaining a sign language video to be processed and user instruction information; obtaining a sign language video to be processed and user instruction information; using a visual model based on a Transformer architecture to perform feature extraction on the sign language video to be processed to obtain a sign language action feature matrix; obtaining an instruction-emphasized sign language action feature vector and a spatiotemporal-emphasized sign language action feature matrix based on the sign language action feature matrix; wherein the instruction-emphasized sign language action feature vector is obtained by processing the user instruction information to obtain a feature extraction instruction, and the sign language action feature matrix is ​​processed based on the feature extraction instruction. The matrix is ​​obtained after instruction emphasis is performed; the method for obtaining the time-space emphasized sign language action feature matrix includes: using a time difference operator to perform action frame emphasis and still frame suppression on the sign language action feature matrix to obtain the time-emphasized sign language action feature matrix; performing spatial emphasis on hand movements and facial expressions on the time-emphasized sign language action feature matrix to obtain the time-space emphasized sign language action feature matrix; performing feature fusion on the instruction emphasized sign language action feature vector and the time-space emphasized sign language action feature matrix to obtain a fused sign language action feature matrix; based on the fused sign language action feature matrix, a preset speech synthesis model is used to obtain the speech information corresponding to the sign language video to be processed.

[0007] In addition, an optional technical solution is that the method for obtaining the voice information corresponding to the sign language video to be processed based on the fused sign language motion feature matrix through a preset speech synthesis model includes inputting the fused sign language motion feature matrix into a preset context distillation model; using the context distillation model to perform context information mining on the sign language motion feature matrix to obtain text data corresponding to the sign language motion feature matrix; wherein, the context distillation model is obtained after optimization through a cross-temporal knowledge distillation loss function and a cross-modal knowledge distillation loss function; and inputting the text data into a preset speech synthesis model to obtain the voice information corresponding to the sign language video to be processed.

[0008] In addition, an optional technical solution is to emphasize the sign language action feature matrix based on the feature extraction instruction, and the method for obtaining the instruction-emphasized sign language action feature vector includes using Q-Former to generate a corresponding interactive instruction vector according to the feature extraction instruction; interacting the feature extraction instruction with the interactive instruction vector in Q-Former; using the key information of the feature extraction instruction captured by the interactive instruction vector as the interaction result; and screening the interaction result and the sign language action feature matrix using a cross-attention mechanism to obtain the instruction-emphasized sign language action feature vector.

[0009] In addition, an optional technical solution is that before the step of extracting features from the sign language video to be processed using a visual model based on the Transformer architecture, the step further includes deblurring the sign language video to be processed using a trained deblurring model; including,

[0010] A blur kernel is trained using a preset blur and clear dataset; wherein the preset blur and clear dataset includes image pairs consisting of a clear image and a blurry image; a sign language video to be processed is divided into two sets: an unknown blur frame set and a clear frame set according to a function; and the blur kernel is implanted into the clear frame set to obtain a known blur frame set;

[0011] Constructing a discriminator network for distinguishing between real blurred images in a known blurred frame set and blurred images generated by a fuzzy transformation network; optimizing the fuzzy transformation network through adversarial loss and reconstruction loss to obtain a trained fuzzy transformation network; wherein the fuzzy transformation network is used to allow the unknown blurred frame set to learn the fuzzy features of the known blurred frame set;

[0012] By performing supervised learning on the known blurred frame set, a deblurring model capable of learning the mapping relationship from blurred to clear images is established;

[0013] The deblurring model is used to deblur an image of a new fuzzy set generated by a trained fuzzy conversion network to obtain a deblurred sign language video; wherein the new fuzzy set includes image features of the unknown fuzzy frame set and fuzzy features of the known fuzzy frame set.

[0014] In addition, an optional technical solution is to use a time difference operator to emphasize action frames and suppress still frames on the sign language action feature matrix to obtain a time-emphasized sign language action feature matrix, including using a time difference operator to perform differential calculation on the sign language action feature matrix in the time dimension to determine the time difference value of each frame of the sign language action feature matrix; dividing each frame of the sign language action feature matrix into still frames or action frames according to the time difference value; increasing the weight of the eigenvalue of the action frame and reducing the weight of the eigenvalue of the still frame; and obtaining a time-emphasized sign language action feature matrix.

[0015] In addition, an optional technical solution is to perform spatial emphasis on hand movements and facial expressions on the time-emphasized sign language movement feature matrix. The method for obtaining the time-space emphasized sign language movement feature matrix includes: performing projection dimensionality reduction processing on the time-emphasized sign language movement feature matrix, and then performing multi-branch feature extraction to obtain features extracted by each branch; wherein the multi-branch features include hand movement branches and facial expression branches with different spatial expansion rates; assigning corresponding preset weights to the features extracted from each branch and summing them to obtain a fused feature matrix and perform dimensionality increase; using the sigmoid function to map the dimensionality increased feature matrix and generate a spatial attention map; multiplying the value range of the spatial attention map by the feature matrix element by element to obtain a time-space emphasized sign language movement feature matrix.

[0016] In addition, an optional technical solution is to further include performing voice recognition on the voice information to be processed to obtain text content corresponding to the voice information; performing multimodal emotion recognition on the text content and the non-text content in the voice information to be processed to obtain the emotion category of the voice information to be processed and the probability distribution of the emotion category; determining expression parameters, lip shape parameters and body movement parameters based on the pre-acquired sign language document, the emotion category and the probability distribution of the emotion category; wherein the sign language document is obtained by performing semantic summary extraction on the text content; based on the expression parameters, the lip shape parameters and the body movement parameters, the expression, lip shape and body movement of the virtual image are respectively driven to obtain the sign language communication video corresponding to the voice information to be processed.

[0017] On the other hand, the present invention also provides a hearing-impaired communication system based on user instruction emphasis, which utilizes the above-mentioned hearing-impaired communication method based on user instruction emphasis for communication; the system comprises:

[0018] An acquisition unit, configured to acquire the sign language video and user instruction information to be processed;

[0019] An emphasis unit is used to extract features from a sign language video to be processed using a visual model based on a Transformer architecture to obtain a sign language action feature matrix; based on the sign language action feature matrix, an instruction-emphasized sign language action feature vector and a spatiotemporal-emphasized sign language action feature matrix are respectively obtained; wherein the instruction-emphasized sign language action feature vector is obtained by processing the user instruction information to obtain a feature extraction instruction, and the sign language action feature matrix is ​​instruction-emphasized based on the feature extraction instruction; a method for obtaining the spatiotemporal-emphasized sign language action feature matrix comprises: using a time difference operator to perform action frame emphasis and still frame suppression on the sign language action feature matrix to obtain a time-emphasized sign language action feature matrix; and performing spatial emphasis on hand movements and facial expressions on the time-emphasized sign language action feature matrix to obtain the spatiotemporal-emphasized sign language action feature matrix;

[0020] a fusion unit, configured to fuse the feature vector of the sign language action emphasized by the instruction and the feature matrix of the sign language action emphasized in time and space to obtain a fused feature matrix of the sign language action;

[0021] A generation unit is used to input the fused sign language motion feature matrix into a preset context distillation model; use the context distillation model to perform context information mining on the sign language motion feature matrix to obtain text data corresponding to the sign language motion feature matrix; input the text data into a preset speech synthesis model to obtain speech information corresponding to the sign language video to be processed.

[0022] The communication method and system for hearing-impaired people based on user command emphasis of the present invention addresses the problems of poor discernibility of sign language movements and low accuracy of movement feature extraction in sign language video images in real environments. By adopting feature extraction methods with time emphasis and space emphasis, the movement trajectory of sign language can be effectively captured. In addition, by adopting a feature extraction method with command emphasis, the present invention enables the model to focus on specific features according to the commands, so as to purposefully improve the relevance of features and the targetedness of recognition. Ultimately, it not only solves the problems of poor quality of sign language video images and insufficient mining of key area features, but also realizes the precise extraction and recognition of sign language movement features, thereby achieving the technical effect of significantly improving the accuracy of sign language recognition.

[0023] In order to achieve the above and related purposes, one or more aspects of the present invention include the features that will be described in detail later. The following description and the accompanying drawings describe some exemplary aspects of the present invention in detail. However, these aspects indicate only some of the various ways in which the principles of the present invention can be used. In addition, the present invention is intended to include all of these aspects and their equivalents. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] By referring to the following description in conjunction with the accompanying drawings, and with a more complete understanding of the present invention, other objects and results of the present invention will become more clear and easy to understand. In the accompanying drawings:

[0025] Figure 1 2. A flowchart of a method for communication with hearing-impaired persons based on user instruction emphasis according to an embodiment of the present invention;

[0026] Figure 2 2. A schematic diagram showing the principle of sign language recognition for a communication method for hearing-impaired persons based on user instruction emphasis according to an embodiment of the present invention;

[0027] Figure 3 2. It is a flowchart of feature matrix emphasis processing of sign language recognition according to a method for communication for hearing-impaired persons based on user instruction emphasis according to an embodiment of the present invention;

[0028] Figure 4 A schematic diagram of the principle of a method for communication for hearing-impaired persons based on user instruction emphasis according to an embodiment of the present invention;

[0029] Figure 5 A schematic diagram of modules of a communication system for hearing-impaired persons based on user instruction emphasis provided according to an embodiment of the present invention;

[0030] Figure 6 A schematic diagram of the internal structure of an electronic device for implementing a communication method for hearing-impaired persons based on user instruction emphasis according to an embodiment of the present invention.

[0031] The same reference numerals throughout the drawings indicate similar or corresponding features or functions. DETAILED DESCRIPTION

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts shall fall within the scope of protection of the present invention.

[0033] The following is a clear and detailed description of the technical solutions in the embodiments of the present application, with reference to the accompanying drawings. In the description of the embodiments of the present application, unless otherwise specified, the term "and / or" is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, "A and / or B" can represent the following three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0034] In the following, the terms "first" and "second" are used for descriptive purposes only and should not be construed as implying or suggesting relative importance or implicitly indicating the number of the technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of such features. Furthermore, in the description of the embodiments of this application, "plurality" means two or more than two.

[0035] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0036] The Vision Transformer (ViT) is a Transformer-based visual model. It utilizes a self-attention mechanism to capture global relationships between different image regions, overcoming the limitations of traditional CNNs. ViT is broadly divided into three parts: image segmentation, positional encoding, and Transformer processing. It is frequently used in scenarios such as image classification, object detection, image segmentation, and video processing. It can simultaneously consider both spatial and temporal information in videos, demonstrating promising results for video classification, action recognition, and video generation. ViT divides an image into a series of fixed-size "patches," with each patch treated as an independent unit, similar to a word or subword in natural language processing. These patches are linearly mapped to vectors, and spatial information is infused through positional encoding to form a serialized representation that serves as the Transformer input. ViT offers the advantages of global context modeling, parallel computing, and flexibility.

[0037] To describe in detail the hearing-impaired communication method and system based on user command emphasis of the present invention, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Example

[0038] Figure 1 A flow chart of a method for communication for hearing-impaired persons based on user instruction emphasis according to an embodiment of the present invention is shown.

[0039] like Figure 1As shown, the method for communication for the hearing-impaired based on user command emphasis in an embodiment of the present invention includes steps S110-S140. Steps S110-S140 describe the communication implementation process of recognizing the hearing-impaired's sign language, synthesizing it into speech, and then conveying it to the ordinary person in a scenario where the hearing-impaired person communicates with the ordinary person using sign language.

[0040] Specifically, S110: Obtain the sign language video and user instruction information to be processed.

[0041] In specific implementations, user command information is used to guide the model to focus on specific features, such as facial expression details and background information of hearing-impaired individuals. User command information can be voice commands picked up by a recording device, input text commands, or voice commands converted to text by a speech recognition system. This is determined by the specific application scenario and is not specifically limited here.

[0042] In addition, the sign language video to be processed can be a sign language video shot by professional video acquisition equipment in a controlled environment, a sign language video stored from a storage device, or a sign language video of communicating with the hearing-impaired that is shot in real time through the cameras of smart wearable devices, smart phones, tablet computers and other devices. In addition, the video content of the sign language video to be processed mainly includes images of the hearing-impaired person expressing through sign language, including sign language movements, facial expressions of the hearing-impaired person, and background information. Sign language movements can include gestures, hand movements, changes in arm position, etc. The facial expressions of the hearing-impaired person can include smiling, frowning, wide eyes, etc. It should be noted that the video may also contain background information, such as indoor room layout, outdoor street scenes, background music, etc. This background information can help provide context, but sometimes it may also interfere with sign language recognition.

[0043] In specific implementation scenarios, the sign language video being processed may be blurry due to differences in video capture equipment, varying sign language expression speeds, and environmental factors. Therefore, before extracting features from the sign language video using a Transformer-based visual model, the method also includes deblurring the sign language video using a trained deblurring model, including steps S111-S114.

[0044] S111. Train the blur kernel using a preset blur and clear dataset; wherein the preset blur and clear dataset includes image pairs consisting of a clear image and a blurred image; divide the sign language video to be processed into two sets, an unknown blur frame set B and a clear frame set S, according to a function; and implant the blur kernel into the clear frame set S to obtain a known blur frame set K. It should be noted that, in the specific implementation process, the RB2V dataset can be used to extract the blur kernel. Blur in sign language videos is mainly motion blur and defocus blur. Motion blur is also caused by relative motion during shooting; defocus blur is caused by inaccurate camera focus or intentionally setting a shallow depth of field, which makes some parts of the image clear while other parts are blurred. The RB2V dataset is established to address these two types of blur. The sign language video that is subsequently filmed is used to distinguish between the unknown blur frame set B and the clear frame set S using features such as image gradient and variance. Implant the aforementioned blur kernel into the clear frame set S to obtain a known blur frame set K. S112. Construct a discriminator network for distinguishing between real blurred images in a known blurred frame set and blurred images generated by a fuzzy conversion network; optimize the fuzzy conversion network through adversarial loss and reconstruction loss to obtain a trained fuzzy conversion network; the fuzzy conversion network is used to allow the unknown fuzzy frame set B to learn the fuzzy features of the known fuzzy frame set, while retaining the image features of the unknown fuzzy frame set B. S113. By performing supervised learning on the known fuzzy frame set K, establish a deblurring model that can learn the image mapping relationship from blur to clear. S114. Use the deblurring model to deblur the image of the new fuzzy set N generated by the trained fuzzy conversion network G to obtain a deblurred sign language video; wherein, the new fuzzy set includes the image features of the unknown fuzzy frame set and the fuzzy features of the known fuzzy frame set.

[0045] In general, a fuzzy transformation network G (i.e., a generator) is trained using an unknown fuzzy frame set B and a known fuzzy frame set K. Generator G generates a new fuzzy set N that is primarily based on the various image features of the unknown fuzzy frame set B and also includes the various fuzzy features of the known fuzzy frame set K. This allows the images in the unknown fuzzy frame set B to learn the fuzzy features of the known fuzzy frame set K while maintaining visual consistency with the images in the unknown fuzzy frame set B. A discriminator network D is established to distinguish between real images from the known fuzzy frame set K and images from the new fuzzy set N generated by the fuzzy transformation network G. The fuzzy transformation network G is continuously optimized using an adversarial loss and a reconstruction loss. The images from the new fuzzy set N generated by the trained fuzzy transformation network G are then passed through the deblurring model to produce a deblurred sign language video. In summary, the present invention converts unknown fuzziness into known fuzziness, making subsequent deblurring operations easier. Because mature deblurring techniques and pre-trained models already exist in the known fuzzy domain, these existing techniques and models can be leveraged to deblur images, improving deblurring accuracy and effectiveness.

[0046] Figure 2 FIG. 1 is a schematic diagram showing the principle of sign language recognition for a hearing-impaired communication method based on user instruction emphasis according to an embodiment of the present invention. Figure 2 As shown, the first step is video acquisition and preprocessing. Specifically, native sign language video is captured through a camera and processed by the blur reduction module to produce a deblurred video, providing a clear image for subsequent feature extraction. The second step is feature extraction. Specifically, the deblurred video is input into the Vision Transformer to extract a feature matrix. The third step is simultaneous user command emphasis and temporal and spatial emphasis. Specifically, user commands are input into the command emphasis module, which generates a command feature vector. This performs preliminary screening and emphasis on sign language features to better meet user needs. The feature matrix is ​​then input into the temporal and spatial emphasis modules. Temporal emphasis highlights key action frames and extracts temporal features of sign language movements. Spatial emphasis focuses on areas such as the hands and face, strengthening spatial features and improving feature expressiveness. The fourth step is the fusion of the command-emphasized feature vector and the temporally and spatially emphasized feature matrix. Specifically, the command feature vector and the emphasized feature matrix are combined in the feature fusion module to generate a processed feature matrix. This fusion of multi-dimensional feature information forms a more complete and expressive sign language feature representation. The final step involves sign language recognition and speech synthesis. Specifically, a contextual distillation model is used to convert feature representations into textual expressions and emotional features. The textual expressions are used to generate corresponding text data, while the emotional features convey the signer's emotions. The speech synthesis module combines the textual expressions and emotional features with the timbre characteristics of the reference speech to synthesize natural and expressive speech output, achieving effective sign language-to-speech conversion.

[0047] like Figure 2 As shown, after the video capture and pre-processing steps, the feature extraction step S120 is continued. The processing is divided into two branches: one branch is to enhance the partial video frame and the position of the face and hands, and the other branch is to highlight special features according to the input user instructions.

[0048] Figure 3 1 is a flow chart of feature matrix emphasis processing in sign language recognition of a hearing-impaired communication method based on user command emphasis according to an embodiment of the present invention.

[0049] S120: Using a Transformer-based visual model to extract features from the sign language video being processed, a sign language motion feature matrix is ​​obtained. Based on the sign language motion feature matrix, a command-emphasized sign language motion feature vector and a spatiotemporal-emphasized sign language motion feature matrix are obtained. The command-emphasized sign language motion feature vector is obtained by processing the user command information to obtain a feature extraction instruction, and the sign language motion feature matrix is ​​then command-emphasized based on the feature extraction instruction. The spatiotemporal-emphasized sign language motion feature matrix is ​​obtained by using a temporal difference operator to emphasize motion frames and suppress still frames on the sign language motion feature matrix to obtain a temporal-emphasized sign language motion feature matrix; and spatially emphasizing hand movements and facial expressions on the temporal-emphasized sign language motion feature matrix to obtain a spatiotemporal-emphasized sign language motion feature matrix. Specifically, the user command information can be in the form of voice or text. If it is a voice command, it must be converted into text by a speech recognition system (such as a large language model) for subsequent processing. If it is a text command, it can proceed directly to the next step.

[0050] The method for performing instruction emphasis on the sign language action feature matrix based on the feature extraction instruction to obtain the instruction emphasized sign language action feature vector includes steps S1201 to S1203.

[0051] S1201. Use Q-Former to generate a corresponding instruction vector after interaction based on the feature extraction instruction. Specifically, the instruction in the instruction emphasis is an instruction expressed in natural language that clearly tells the model that a task needs to be completed, which is equivalent to user instruction information. It has clear semantics and intentions, directly guiding the model to mine information from the image, and conforms to the definition and characteristics of the instruction, such as "Please pay attention to the emotional information expressed by the characters in the video." The instruction vector after interaction is a learnable vector in Q-Former. It is not in the form of natural language itself and does not have the ability to directly issue task instructions. It mainly guides the model to extract information related to the instruction from the image through interaction with instructions and image features. The converted text instructions or the original text instructions are encoded and mapped to a numerical space that the model can recognize and process to form an instruction vector, i.e., a feature extraction instruction. For example, text information can be converted into a vector representation of a fixed dimension through technologies such as word embedding.

[0052] S1202. The feature extraction instruction and the instruction vector after interaction are interacted in the Q-Former; the key information of the feature extraction instruction captured by the instruction vector after interaction is used as the interaction result. It should be noted that the fusion calculation is performed through the self-attention mechanism. The instruction vector after interaction is the key vector used to guide feature extraction in the model. Through interaction with the instruction vector, the model can understand the feature extraction intention contained in the user instruction.

[0053] S1203: The interaction result and the sign language action feature matrix are filtered using a cross-attention mechanism to obtain a feature vector of the sign language action emphasized by the instruction. The interactive instruction vector is then cross-attended with the feature matrix of the sign language video. During this process, the interactive instruction vector guides the model to filter out instruction-related features from the feature matrix of the sign language video, such as facial expression details of the hearing-impaired person and background information, while ignoring other irrelevant information. The cross-attention mechanism calculates attention weights based on the interactive instruction vector and each element in the feature matrix. This attention weight is used to filter and emphasize the feature matrix, highlighting the parts closely related to the user's instruction. Finally, through forward propagation, a feature vector is generated that reflects the user's instruction. This process, rather than indiscriminately processing all visual information, guides the model to focus on features such as facial expression details and character movements, while ignoring other irrelevant information in the image, thereby more accurately extracting the features of interest. This step achieves targeted screening and refinement of the feature matrix, ensuring that the final output feature vector better reflects the features of interest to the user's instruction. The interaction between the instruction and query embeddings through the Q-Former enables the model to focus on image features closely related to the task.

[0054] like Figure 3As shown, the feature matrix F[T, C, H, W] is serialized to obtain the serialized feature [T X ,H X, W, C]; where T represents the time dimension, C represents the number of channels, H represents the height, and W represents the width. After the user command and the interactive command vector are input, they are processed by the cross-attention mechanism and enter the forward propagation phase. After passing through the fully connected layer and tensor operations, the feature vector of the sign language action emphasized by the command is obtained [T, N, C, H, W].

[0055] A method for obtaining a time-emphasized sign language motion feature matrix by using a time difference operator to emphasize action frames and suppress still frames, includes: S1211: using a time difference operator to perform difference calculation on the sign language motion feature matrix in the time dimension to determine the time difference value of each frame of the sign language motion feature matrix; S1212: dividing each frame of the sign language motion feature matrix into still frames or action frames according to the time difference value; S1213: increasing the weight of the eigenvalue of the action frame and reducing the weight of the eigenvalue of the still frame; obtaining the time-emphasized sign language motion feature matrix.

[0056] In other words, temporal emphasis is used to appropriately enhance action frames and suppress static frames. Not all frames in a sign language video contribute equally to recognition; some are more discriminative than others. Temporal emphasis can enhance action frames to better identify the body trajectory of sign language. In the temporal emphasis pipeline architecture, the sign language feature matrix extracted by the Vision Transformer in step 2 is taken as input. The feature matrix first passes through a global average pooling layer to eliminate spatial dimensions. After global average pooling, the feature matrix passes through a convolutional layer with a kernel size of 1, reducing the number of channels to 1 by a reduction factor r. To better utilize sign language motion to identify action frames, a temporal difference operator is used to calculate the differences between adjacent frames as approximate motion information. This is then concatenated with the feature matrix to combine appearance features with motion information, enabling the model to better understand sign language motion. The concatenated feature matrix is ​​then processed using a convolutional layer with a kernel size of 3×1 for further feature extraction and integration. The number of channels is then restored through a convolutional layer with a kernel size of 1. Finally, a sigmoid activation function introduces nonlinearity to enhance the model's expressiveness and generate a temporal attention map. The resulting attention map is element-wise multiplied with the input feature matrix in a residual manner. Residual connections better preserve the information in the original features, resulting in a feature matrix that emphasizes action frames and suppresses still frames.

[0057] like Figure 3As shown, the feature matrix F[T, C, H, W] is pooled to obtain [T, C], which is then convolved to obtain [T, C / r], where r is the compression ratio of the convolution kernel. Based on the feature sequence, a temporal difference operator is used to obtain the feature sequences F(t) and F(t+1) in the time dimension. The feature sequences F(t) and F(t+1) are concatenated to obtain [T, 2C / r], which is then convolved to obtain [T, C / r]. Convolution is then performed again to obtain [T, C]. After processing with the activation function, the resulting attention map is element-wise multiplied with the input feature matrix in a residual manner to obtain the time-emphasized sign language action feature matrix G[T, C, H, W].

[0058] As an improvement to this embodiment, a method for performing spatial emphasis on hand movements and facial expressions on the time-emphasized sign language action feature matrix to obtain the time-space-emphasized sign language action feature matrix includes the following steps.

[0059] S1221: After performing projection dimensionality reduction processing on the time-emphasized sign language action feature matrix, multi-branch feature extraction is performed to obtain features extracted by each branch; wherein, the multi-branch features include hand action branches and facial expression branches with different spatial expansion rates; S1222: Assign corresponding preset weights to the features extracted by each branch and sum them up to obtain a fused feature matrix and perform dimensionality upgrade; S1223: Use the sigmoid function to map the dimensionality-enhanced feature matrix and generate a spatial attention map; S1224: Multiply the value range of the spatial attention map by the feature matrix element by element to obtain a time-space emphasized sign language action feature matrix.

[0060] A spatial emphasis method is used to highlight information-rich hand and facial regions. Specifically, the spatial emphasis process first reduces the dimensionality of the weighted sign language feature matrix obtained from temporal emphasis through convolution projection with a kernel size of 1. This reduces subsequent computational complexity and linearly combines features from different channels to enhance feature representation. The spatial emphasis method learns the importance of each spatial location relative to other locations without directly operating on the temporal dimension. This is because it implicitly utilizes temporal information while learning spatial features. Information such as hand movements and facial expressions in sign language expressions resides in spatial regions of varying scales, making it difficult to fully capture this information with a single convolutional kernel size and dilation rate. Multiple parallel branches, each with a different spatial dilation rate, perceive features from different receptive fields. Branches with smaller dilation rates can focus on local details, while branches with larger dilation rates can capture broader contextual information. This multi-branch approach allows the input feature matrix to be analyzed from multiple receptive fields, enriching the feature representation. Different weights are assigned to the importance of features from different branches, and then summed to form a fused new feature matrix. This fused feature matrix integrates the diverse spatial information extracted by each branch. Next, the fused feature matrix is ​​fed into a 1×1×1 convolutional layer to restore the feature matrix dimension. The sigmoid function is then used to map the feature matrix values ​​to the [0, 1] interval to generate a spatial attention map. To avoid damaging the original representation and reducing accuracy, the input features are emphasized using a residual approach. The resulting attention map is then resized to [-0.5, 0.5] and element-wise multiplied with the feature matrix to emphasize key information areas and suppress unnecessary areas, achieving greater focus on facial and hand features. Finally, the feature matrix is ​​output after both temporal and spatial emphasis. Figure 3 As shown in the figure, the time-emphasized sign language action feature matrix G[T, C, H, W] is convolved to obtain [T, C / r, H, W], and then a convolution operation of multi-branch feature extraction with convolution dilation rate 1, dilation rate 2... dilation rate N is performed respectively; the convolved features are weighted summed to obtain the weighted summed features; the weighted summed features are convolved to obtain [T, C, H, W]; finally, after activation function processing, the space-emphasized sign language action feature matrix R[T, C, H, W] is obtained.

[0061] S130: Fusing the feature vector of the sign language action emphasized by the instruction and the feature matrix of the sign language action emphasized in time and space to obtain a fused feature matrix of the sign language action. Figure 3As shown, the spatially emphasized sign language motion feature matrix R[T, C, H, W] and the command-emphasized sign language motion feature vector [T, N, C, H, W] are fused, ultimately outputting the fused sign language motion feature matrix S[T, C, H, W]. Specifically, because the spatiotemporally emphasized sign language motion feature matrix and the command-emphasized sign language motion feature vector in step S120 have different dimensions, dimension matching is required. The feature vector is first processed through a fully connected layer to match the number of channels with the feature matrix. Then, tensor operations are used to project the user-command-aware feature vector onto the same feature dimension as the feature matrix after the temporal and spatial emphasis processing. An attention mechanism is used to dynamically fuse the feature vector and feature matrix. The attention mechanism adaptively assigns weights based on feature relevance. After fusion, the fused feature matrix is ​​further processed through convolutional layers, fully connected layers, and nonlinear activation function layers to enhance the model's expressive power. Ultimately, a newly fused feature matrix is ​​obtained.

[0062] S140: Based on the fused sign language action feature matrix, a preset speech synthesis model is used to obtain speech information corresponding to the sign language video to be processed.

[0063] The method for obtaining the speech information corresponding to the sign language video to be processed based on the fused sign language motion feature matrix through a preset speech synthesis model includes the following steps: S141, inputting the fused sign language motion feature matrix into a preset context distillation model, using the context distillation model to perform context information mining on the sign language motion feature matrix, and obtaining text data corresponding to the sign language motion feature matrix; wherein the context distillation model is obtained after optimization through a cross-temporal knowledge distillation loss function and a cross-modal knowledge distillation loss function; S142, inputting the text data into a preset speech synthesis model to obtain the speech information corresponding to the sign language video to be processed. Specifically, a context knowledge distillation method is used to train a prediction model based on the above-mentioned new fused feature matrix, and finally the trained model is used to analyze the sign language, identify its corresponding vocabulary, sentences or meaning, and convert it into a textual expression, i.e., text data.

[0064] In addition, it should be noted that the generation of voice information corresponding to the sign language video to be processed from the above-mentioned text expression needs to be implemented through a speech synthesis module. Specifically, the speech synthesis module can convert text into natural speech output; in the specific implementation process, its working process is reference timbre feature extraction, feature embedding, model training, and speech output based on text. In the specific implementation process, the features that need to be embedded in this module include the reference timbre features selected by the user and the emotional feature matrix obtained by the sign language recognition module. The trained model then performs speech synthesis based on the text expression output by the sign language recognition module. Exemplarily, the speech synthesis module can be implemented through an existing TTS model.

[0065] Figure 4 Schematic diagram of the principle of a communication method for hearing-impaired persons based on user command emphasis according to an embodiment of the present invention.

[0066] Specifically, in scenarios where a hearing-impaired person communicates with a non-hearing person using sign language, in addition to recognizing the hearing-impaired person's sign language and synthesizing it into speech for transmission to the non-hearing person, the communication process also includes recognizing the non-hearing person's voice message and synthesizing it into sign language animation for transmission to the hearing-impaired person. It should be noted that the steps of recognizing the non-hearing person's voice message and synthesizing it into sign language animation do not need to be performed after steps S110-S140; they can also be performed before steps S110-S140 or in parallel with steps S110-S140. In other words, the processing needs to be tailored to the actual scenario in which the hearing-impaired person communicates with the non-hearing person using sign language, and specific limitations are not provided herein. If the non-hearing person's voice message is initiated in the communication scenario, then the non-hearing person's voice message is first recognized and synthesized into sign language animation. After the hearing-impaired person obtains the sign language communication video corresponding to the non-hearing person's voice message, when the sign language communication occurs, steps S110-S140 are then performed to recognize the hearing-impaired person's sign language and synthesize it into speech for transmission to the non-hearing person. That is to say, the present invention can realize real-time communication between the hearing-impaired and ordinary people.

[0067] like Figure 4 As shown, during the reception process, the hearing-impaired person's sign language movements are captured by the sign language recognition module. This module extracts key features from the sign language video using methods including temporal emphasis, spatial emphasis, and command emphasis. A contextual distillation model identifies the meaning of the sign language based on these feature matrices, converts it into textual representations, and passes it to the speech synthesis module. The speech synthesis module converts the textual representations into speech output, which is broadcast to the staff (ordinary people) via headphones. The textual representations can also be displayed on the screen, allowing the staff to understand and respond to the hearing-impaired person's expressions. During the output process, the staff's speech input is received by the speech recognition module. The speech recognition module converts the speech information into textual representations and extracts emotional features to ensure that the converted sign language accurately conveys tone and emotion. These textual representations and emotional features are then passed to the sign language output module, which converts the text into sign language animations and displays them as sign language movements on the screen, allowing the hearing-impaired person to intuitively understand and receive the information. Through the collaborative work of the reception and output processes, this invention enables efficient two-way communication between the hearing-impaired person and staff, improving the accuracy and convenience of communication.

[0068] In a specific implementation process, steps S101-S104 implement a communication process in which ordinary people's voice messages are voice-recognized and synthesized into sign language animations, and then conveyed to hearing-impaired people. S101: Perform voice recognition on the voice information to be processed to obtain the text content corresponding to the voice information; S102: Perform multimodal emotion recognition on the text content and the non-text content in the voice information to be processed to obtain the emotion category of the voice information to be processed and the probability distribution of the emotion category; S103: Determine expression parameters, lip shape parameters, and body movement parameters based on the pre-acquired sign language document, the emotion category, and the probability distribution of the emotion category; wherein the sign language document is obtained by performing semantic summary extraction on the text content; S104: Based on the expression parameters, the lip shape parameters, and the body movement parameters, respectively drive the expression, lip shape, and body movement of the virtual image to obtain a sign language communication video corresponding to the voice information to be processed.

[0069] The task of the speech recognition module is to convert the speech signal input by the staff into text, which generally includes the steps of speech signal acquisition, feature extraction, model construction and decoding output text. Exemplarily, the speech recognition module can be implemented using the whisper algorithm. The sign language output module requires the system to be able to understand the text content and simulate sign language expression through AI digital human (virtual image) animation so that the hearing-impaired can understand the content through vision. The implementation steps are NLP semantic analysis, gesture sequence generation, action expression generation, body posture generation, animation timing optimization and output 3D animation. Exemplarily, the sign language output module can be implemented using a combination of MediaPipe + Blender.

[0070] The communication method for hearing-impaired people based on user command emphasis of the present invention addresses the problems of poor discernibility of sign language movements and low accuracy of movement feature extraction in sign language video images in real environments. The present invention adopts feature extraction methods with time emphasis and space emphasis to effectively capture the movement trajectory of sign language. In addition, the present invention adopts a feature extraction method with command emphasis, so that the model can focus on specific features according to the commands, so as to purposefully improve the relevance of features and the targetedness of recognition. Ultimately, it not only solves the problems of poor quality of sign language video images and insufficient mining of key area features, but also realizes the precise extraction and recognition of sign language movement features, thereby achieving the technical effect of significantly improving the accuracy of sign language recognition.

[0071] like Figure 5As shown, the present invention provides a communication system for hearing-impaired persons based on user instruction emphasis, which uses the communication method for hearing-impaired persons based on user instruction emphasis as described above for communication. Depending on the functions implemented, the communication system 500 for hearing-impaired persons based on user instruction emphasis may include an acquisition unit 510, an emphasis unit 520, a fusion unit 530, and a generation unit 540. The unit of the present invention may also be referred to as a module, which refers to a series of computer program segments that can be executed by an electronic device processor and can perform fixed functions, and is stored in the memory of the electronic device.

[0072] In this embodiment, the functions of each module / unit are as follows:

[0073] An acquisition unit 510 is configured to acquire the sign language video to be processed and user instruction information;

[0074] The emphasis unit 520 is configured to perform feature extraction on the sign language video to be processed using a visual model based on the Transformer architecture to obtain a sign language action feature matrix; based on the sign language action feature matrix, respectively obtain an instruction-emphasized sign language action feature vector and a spatiotemporal-emphasized sign language action feature matrix; wherein the instruction-emphasized sign language action feature vector is obtained by processing the user instruction information to obtain a feature extraction instruction, and the sign language action feature matrix is ​​instruction-emphasized based on the feature extraction instruction; a method for obtaining the spatiotemporal-emphasized sign language action feature matrix includes performing action frame emphasis and still frame suppression on the sign language action feature matrix using a time difference operator to obtain a time-emphasized sign language action feature matrix; and performing spatial emphasis on hand movements and facial expressions on the time-emphasized sign language action feature matrix to obtain the spatiotemporal-emphasized sign language action feature matrix;

[0075] A fusion unit 530 is configured to fuse the instruction-emphasized sign language action feature vector and the spatiotemporal-emphasized sign language action feature matrix to obtain a fused sign language action feature matrix;

[0076] The generation unit 540 is used to input the fused sign language motion feature matrix into a preset context distillation model; use the context distillation model to perform context information mining on the sign language motion feature matrix to obtain text data corresponding to the sign language motion feature matrix; input the text data into a preset speech synthesis model to obtain speech information corresponding to the sign language video to be processed.

[0077] The communication system for hearing-impaired people based on user command emphasis of the present invention addresses the problems of poor discernibility of sign language movements and low accuracy of movement feature extraction in sign language video images in real environments. The present invention adopts feature extraction methods with time emphasis and space emphasis to effectively capture the movement trajectory of sign language. In addition, the present invention adopts a feature extraction method with command emphasis, so that the model can focus on specific features according to the commands, so as to purposefully improve the relevance of features and the targetedness of recognition. Ultimately, it not only solves the problems of poor quality of sign language video images and insufficient mining of key area features, but also realizes the precise extraction and recognition of sign language movement features, thereby achieving the technical effect of significantly improving the accuracy of sign language recognition.

[0078] More specific implementation methods of the above-mentioned hearing-impaired communication system based on user command emphasis can refer to the above-mentioned embodiment description of the hearing-impaired communication method based on user command emphasis, and will not be described in detail here.

[0079] like Figure 6 As shown, the present invention also provides an electronic device 1 for a communication method for hearing-impaired persons based on user instruction emphasis.

[0080] The electronic device 1 may include a processor 10, a memory 11, and a bus. It may also include a computer program stored in the memory 11 and executable on the processor 10, such as a user-command-based communication program 12 for the hearing-impaired. The memory 11 may include both an internal storage unit of the user-command-based communication system for the hearing-impaired and an external storage device. The memory 11 may be used not only to store installed application software and various data, such as the code for the user-command-based communication program, but also to temporarily store data that has been output or is about to be output.

[0081] The memory 11 includes at least one type of readable storage medium, including flash memory, a removable hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 11 may be an internal storage unit of the electronic device 1, such as a removable hard disk of the electronic device 1. In other embodiments, the memory 11 may also be an external storage device of the electronic device 1, such as a plug-in removable hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a flash memory card, etc. Furthermore, the memory 11 may include both an internal storage unit of the electronic device 1 and an external storage device. The memory 11 can be used not only to store application software installed in the electronic device 1 and various types of data, such as the code of a communication program for the hearing-impaired emphasized based on user instructions, but also to temporarily store data that has been output or is about to be output.

[0082] In some embodiments, the processor 10 may be comprised of an integrated circuit, such as a single packaged integrated circuit or a combination of multiple packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips. The processor 10 is the control core (control unit) of the electronic device, connecting the various components of the electronic device using various interfaces and circuits. It executes programs or modules stored in the memory 11 (e.g., a communication program for the hearing-impaired based on user instructions) and accesses data stored in the memory 11 to perform various functions and process data.

[0083] The bus may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. The bus is configured to enable connection and communication between the memory 11 and at least one processor 10, etc.

[0084] Figure 6 Only the electronic device with components is shown, and it can be understood by those skilled in the art that Figure 6The structure shown does not constitute a limitation on the electronic device 1 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.

[0085] For example, although not shown, the electronic device 1 may further include a power source (e.g., a battery) to power various components. Preferably, the power source may be logically connected to the at least one processor 10 via a power management system, thereby enabling functions such as charge management, discharge management, and power consumption management through the power management system. The power source may further include any components such as one or more DC or AC power sources, a recharging system, a power failure detection circuit, a power converter or inverter, and a power status indicator. The electronic device 1 may further include various sensors, Bluetooth modules, Wi-Fi modules, etc., which are not further described here.

[0086] Furthermore, the electronic device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is generally used to establish a communication connection between the electronic device 1 and other electronic devices.

[0087] Optionally, the electronic device 1 may further include a user interface, which may be a display or an input unit (such as a keyboard). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a display screen or a display unit, and is used to display information processed by the electronic device 1 and to display a visual user interface.

[0088] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.

[0089] The communication program 12 for hearing-impaired persons based on user instruction emphasis stored in the memory 11 of the electronic device 1 is a combination of multiple instructions. When running in the processor 10, it can achieve the following: obtaining the sign language video to be processed and the user instruction information; using the visual model based on the Transformer architecture to perform feature extraction on the sign language video to be processed to obtain a sign language action feature matrix; obtaining the instruction-emphasized sign language action feature vector and the spatiotemporal-emphasized sign language action feature matrix based on the sign language action feature matrix; wherein the instruction-emphasized sign language action feature vector is obtained by processing the user instruction information to obtain a feature extraction instruction, and the sign language action feature matrix is ​​extracted based on the feature extraction instruction. The method for obtaining the spatiotemporal emphasized sign language action feature matrix comprises the following steps: using a time difference operator to perform action frame emphasis and still frame suppression on the sign language action feature matrix to obtain a time-emphasized sign language action feature matrix; performing spatial emphasis on hand movements and facial expressions on the time-emphasized sign language action feature matrix to obtain a spatiotemporal emphasized sign language action feature matrix; performing feature fusion on the instruction-emphasized sign language action feature vector and the spatiotemporal emphasized sign language action feature matrix to obtain a fused sign language action feature matrix; and obtaining speech information corresponding to the sign language video to be processed through a preset speech synthesis model based on the fused sign language action feature matrix.

[0090] Specifically, the specific implementation method of the processor 10 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments is not repeated here. Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium may include: any entity or system capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM).

[0091] The embodiment of the present invention also provides a computer-readable storage medium, which can be non-volatile or volatile, and stores a computer program. When the computer program is executed by a processor, it realizes the following: obtaining a sign language video to be processed and user instruction information; using a visual model based on a Transformer architecture to perform feature extraction on the sign language video to be processed to obtain a sign language action feature matrix; based on the sign language action feature matrix, respectively obtaining a sign language action feature vector emphasized by the instruction and a sign language action feature matrix emphasized in time and space; wherein the sign language action feature vector emphasized by the instruction is obtained by processing the user instruction information to obtain a feature extraction instruction, and based on the feature extraction instruction, the feature extraction instruction is used to extract the sign language action feature matrix. The sign language action feature matrix is ​​obtained after instruction emphasis; the method for obtaining the spatiotemporal emphasized sign language action feature matrix includes: using a time difference operator to emphasize action frames and suppress still frames on the sign language action feature matrix to obtain a time-emphasized sign language action feature matrix; performing spatial emphasis on hand movements and facial expressions on the time-emphasized sign language action feature matrix to obtain the spatiotemporal emphasized sign language action feature matrix; performing feature fusion on the instruction-emphasized sign language action feature vector and the spatiotemporal emphasized sign language action feature matrix to obtain a fused sign language action feature matrix; based on the fused sign language action feature matrix, obtaining voice information corresponding to the sign language video to be processed through a preset voice synthesis model.

[0092] Specifically, the specific implementation method when the computer program is executed by the processor can refer to the description of the relevant steps in the communication method for the hearing-impaired person emphasized based on user instructions in the embodiment, and will not be repeated here.

[0093] In the several embodiments provided herein, it should be understood that the disclosed devices, systems, and methods may be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the module division is merely a logical functional division, and actual implementation may employ other division methods.

[0094] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.

[0095] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.

[0096] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.

[0097] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.

[0098] In addition, it is obvious that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or systems stated in the system claim can also be implemented by one unit or system through software or hardware.

[0099] However, those skilled in the art will appreciate that various improvements may be made to the above-mentioned method for communication with the hearing-impaired based on user command emphasis and the system for communication with the hearing-impaired based on user command emphasis without departing from the scope of the present invention. Therefore, the scope of protection of the present invention shall be determined by the contents of the appended claims.

Claims

1. A method for communication with hearing-impaired persons based on user instruction emphasis, applied to electronic devices, characterized in that: include: Obtain the sign language video and user command information to be processed; The user instruction is an instruction expressed in natural language that guides the model to focus on specific features; A visual model based on a Transformer architecture is used to extract features from a sign language video to be processed to obtain a sign language action feature matrix; based on the sign language action feature matrix, an instruction-emphasized sign language action feature vector and a spatiotemporal-emphasized sign language action feature matrix are respectively obtained; wherein the instruction-emphasized sign language action feature vector is obtained by processing the user instruction information to obtain a feature extraction instruction, and the sign language action feature matrix is ​​instruction-emphasized based on the feature extraction instruction; a method for obtaining the spatiotemporal-emphasized sign language action feature matrix comprises: using a time difference operator to perform action frame emphasis and still frame suppression on the sign language action feature matrix to obtain a time-emphasized sign language action feature matrix; and performing spatial emphasis on hand movements and facial expressions on the time-emphasized sign language action feature matrix to obtain the spatiotemporal-emphasized sign language action feature matrix; Performing feature fusion on the sign language action feature vector emphasized by the instruction and the sign language action feature matrix emphasized in time and space to obtain a fused sign language action feature matrix; The fused sign language action feature matrix is ​​converted into text expression and emotional features through a context distillation model, and the voice information corresponding to the sign language video to be processed is obtained through a preset speech synthesis model; the sign language action feature matrix is ​​emphasized based on the feature extraction instruction, and the method for obtaining the instruction-emphasized sign language action feature vector includes: using Q-Former to generate a corresponding interactive instruction vector according to the feature extraction instruction; interacting the feature extraction instruction with the interactive instruction vector in Q-Former; using the key information of the feature extraction instruction captured by the interactive instruction vector as the interaction result; and screening the interaction result and the sign language action feature matrix using a cross-attention mechanism to obtain the instruction-emphasized sign language action feature vector.

2. The method for communication with the hearing-impaired based on user instruction emphasis according to claim 1, characterized in that: The method of converting the fused sign language action feature matrix into text expression and emotional features through a context distillation model and obtaining speech information corresponding to the sign language video to be processed through a preset speech synthesis model includes: Inputting the fused sign language action feature matrix into a preset context distillation model; using the context distillation model to perform context information mining on the sign language action feature matrix to obtain text data corresponding to the sign language action feature matrix; wherein the context distillation model is obtained by optimizing the cross-temporal knowledge distillation loss function and the cross-modal knowledge distillation loss function; The text data is input into a preset speech synthesis model to obtain speech information corresponding to the sign language video to be processed.

3. The method for communication with the hearing-impaired based on user instruction emphasis according to claim 1, characterized in that: Before the step of extracting features from the sign language video to be processed using a visual model based on a Transformer architecture, the step also includes deblurring the sign language video to be processed using a trained deblurring model; include, A blur kernel is trained using a preset blur and clear dataset; wherein the preset blur and clear dataset includes image pairs consisting of a clear image and a blurry image; a sign language video to be processed is divided into two sets: an unknown blur frame set and a clear frame set according to a function; and the blur kernel is implanted into the clear frame set to obtain a known blur frame set; Constructing a discriminator network for distinguishing between real blurred images in a known blurred frame set and blurred images generated by a fuzzy transformation network; optimizing the fuzzy transformation network through adversarial loss and reconstruction loss to obtain a trained fuzzy transformation network; wherein the fuzzy transformation network is used to allow the unknown blurred frame set to learn the fuzzy features of the known blurred frame set; By performing supervised learning on the known blurred frame set, a deblurring model capable of learning the mapping relationship from blurred to clear images is established; The deblurring model is used to deblur an image of a new fuzzy set generated by a trained fuzzy conversion network to obtain a deblurred sign language video; wherein the new fuzzy set includes image features of the unknown fuzzy frame set and fuzzy features of the known fuzzy frame set.

4. The method for communication with the hearing-impaired based on user instruction emphasis according to claim 1, characterized in that: The method of using a time difference operator to perform action frame emphasis and still frame suppression on the sign language action feature matrix to obtain a time-emphasized sign language action feature matrix includes: Performing differential calculation on the sign language motion feature matrix in a time dimension using a time differential operator to determine a time differential value of each frame of the sign language motion feature matrix; Dividing each frame of the sign language action feature matrix into a still frame or an action frame according to the time difference value; The weight of the feature value of the action frame is increased, and the weight of the feature value of the static frame is reduced; and a time-emphasized sign language action feature matrix is ​​obtained.

5. The method for communication with the hearing-impaired based on user instruction emphasis according to claim 1, characterized in that: The method of performing spatial emphasis on hand movements and facial expressions on the time-emphasized sign language action feature matrix to obtain the time-space-emphasized sign language action feature matrix includes: After performing projection dimensionality reduction processing on the time-emphasized sign language action feature matrix, multi-branch feature extraction is performed to obtain features extracted from each branch; wherein the multi-branch feature includes a hand action branch and a facial expression branch with different spatial expansion rates; Assign corresponding preset weights to the features extracted by each branch and sum them up to obtain the fused feature matrix and perform dimensionality increase; Use the sigmoid function to map the dimension-enhanced feature matrix and generate a spatial attention map; The value range of the spatial attention map is multiplied element-by-element by the feature matrix to obtain a spatiotemporal emphasized sign language action feature matrix.

6. The method for communication with the hearing-impaired based on user instruction emphasis according to claim 1, characterized in that: Also includes, Performing speech recognition on the speech information to be processed to obtain text content corresponding to the speech information; Performing multimodal emotion recognition on the text content and the non-text content in the voice information to be processed to obtain the emotion category of the voice information to be processed and the probability distribution of the emotion category; Determining expression parameters, lip shape parameters, and body movement parameters based on a pre-acquired sign language document, the emotion category, and a probability distribution of the emotion category; wherein the sign language document is obtained by performing semantic summary extraction on the text content; The expression, lip shape and body movement of the virtual image are driven respectively based on the expression parameters, the lip shape parameters and the body movement parameters to obtain a sign language communication video corresponding to the voice information to be processed.

7. A hearing-impaired communication system based on user command emphasis, utilizing the hearing-impaired communication method based on user command emphasis as claimed in any one of claims 1 to 6 for communication; the system comprising: An acquisition unit, configured to acquire the sign language video and user instruction information to be processed; The user instruction is an instruction expressed in natural language that guides the model to focus on specific features; The emphasis unit is used to extract features from the sign language video to be processed using a visual model based on the Transformer architecture to obtain a sign language action feature matrix; based on the sign language action feature matrix, an instruction-emphasized sign language action feature vector and a spatiotemporal-emphasized sign language action feature matrix are respectively obtained; wherein, the instruction-emphasized sign language action feature vector is obtained by processing the user instruction information to obtain a feature extraction instruction, and the sign language action feature matrix is ​​obtained after instruction emphasis is performed on the feature extraction instruction; the method for obtaining the spatiotemporal-emphasized sign language action feature matrix includes using a time difference operator to perform action frame emphasis and still frame suppression on the sign language action feature matrix to obtain a time-emphasized sign language action feature matrix; A time-emphasized sign language action feature matrix performs spatial emphasis on hand actions and facial expressions to obtain the time-space-emphasized sign language action feature matrix; a method for performing instruction emphasis on the sign language action feature matrix based on the feature extraction instruction to obtain an instruction-emphasized sign language action feature vector includes: using a Q-Former to generate a corresponding interactive instruction vector according to the feature extraction instruction; interacting the feature extraction instruction with the interactive instruction vector in a Q-Former; using key information of the feature extraction instruction captured by the interactive instruction vector as an interaction result; and screening the interaction result and the sign language action feature matrix using a cross-attention mechanism to obtain an instruction-emphasized sign language action feature vector. a fusion unit, configured to fuse the feature vector of the sign language action emphasized by the instruction and the feature matrix of the sign language action emphasized in time and space to obtain a fused feature matrix of the sign language action; A generation unit is used to input the fused sign language motion feature matrix into a preset context distillation model; use the context distillation model to perform context information mining on the sign language motion feature matrix to obtain text expressions and emotional features corresponding to the sign language motion feature matrix; input the text expressions and emotional features into a preset speech synthesis model to obtain speech information corresponding to the sign language video to be processed.

8. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a communication program for the hearing-impaired based on user instruction emphasis stored in the memory and runnable on the processor. When the communication program for the hearing-impaired based on user instruction emphasis is executed by the processor, the communication method for the hearing-impaired based on user instruction emphasis as described in any one of claims 1 to 6 is implemented.

9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method for communication for the hearing-impaired based on user instruction emphasis as claimed in any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Sign language interaction method and device and computer medium

    CN110598576A

  • Continuous sign language recognition method based on multi-clue mutual distillation and self-distillation

    CN114821802A

  • Space-time Transform action recognition method for sign language recognition

    CN115205966A

  • Display device, control method of display device and sign language interaction method

    CN120164466A

  • System

    JP2025055245A