Display-Based Communication Systems

A display-based communication aid with a sign language recognition module enhances sign language communication accuracy and convenience, addressing barriers faced by hearing-impaired individuals, particularly in non-face-to-face environments.

JP2025525409AActive Publication Date: 2025-08-05BATONERS INC
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2024576377
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-09-04
Filing Date
2023-09-04
Publication Date
2025-08-05
Estimated Expiration
2043-09-04

AI Technical Summary

Technical Problem

Hearing-impaired individuals face challenges in communicating effectively due to barriers that obstruct visual cues, such as transparent barriers installed during the COVID-19 pandemic, and the lack of sign language accessibility in services, necessitating improved sign language communication systems.

Method used

A display-based communication aid that includes a sign language recognition module to extract sign language sentences from user movements, a display to show the extracted sentences, and additional modules for voice-to-text conversion, word card selection, and text input to enhance communication accuracy and convenience.

Benefits of technology

The system improves the accuracy and convenience of sign language communication, enabling hearing-impaired individuals to interact without professional interpreters and facilitating easy access to services.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025525409000001_ABST
    Figure 2025525409000001_ABST
Patent Text Reader

Abstract

The present disclosure provides a communication aid for communication using sign language, the communication aid including a sign language recognition module that extracts sign language sentences from user movements analyzed from video data, and a transparent display that displays the extracted sign language sentences.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to communication systems, and more particularly to display-based communication systems. [Background technology]

[0002] The term "hearing impaired" is a general term for people who have reduced hearing or lost hearing function. The hearing impaired can communicate in three main ways depending on the degree of hearing impairment. First, if the degree of hearing impairment is low, they can communicate orally with non-disabled people by using assistive hearing devices to augment their hearing. Second, they can communicate with non-disabled people using speech reading, which involves inferring what the other person is saying by looking at their mouths. Finally, they can communicate with non-disabled people using sign language.

[0003] The hearing impaired have always had difficulty communicating, and the COVID-19 pandemic that began in 2020 exacerbated these difficulties. For example, as COVID-19 spread, transparent barriers were installed at reception and consultation desks to block droplets, making it difficult for the hearing impaired to hear what the other person was saying. Furthermore, if the barriers become contaminated, hearing impaired people who use speechreading or sign language are unable to properly see the mouth or sign language movements, making communication even more difficult.

[0004] In addition, since the majority of services are not based on sign language, various technological developments are required to increase the accessibility of services to sign language users. Summary of the Invention [Problem to be solved by the invention]

[0005] The present disclosure aims to provide a display-based sign language communication system for increasing the accuracy and convenience of sign language communication. [Means for solving the problem]

[0006] The present disclosure provides a communication aid for communication using sign language, the communication aid including a sign language recognition module that extracts sign language sentences from analyzed user movements of video data, and a display that displays the extracted sign language sentences.

[0007] According to one embodiment, the communication assistance device may further include an STT module that converts voice data into text data, and a sign language generation module that converts voice data into sign language data.

[0008] According to one embodiment, the communication assistance device may further include a word card selection module that provides word cards selectable by the user on the display, and the sign language recognition module may extract the sign language sentence based on the selected word card.

[0009] According to one embodiment, the communication assistance device further includes a text input module that provides a user interface on the display for a user to input text, and the text input module can be activated when the sign language recognition module fails to extract the sign language sentence.

[0010] According to one embodiment, the communication assistance device may further include a communication module that controls the communication assistance device to be communicatively connected to an external device, and when the sign language recognition module fails to extract the sign language sentence, the communication module controls the communication assistance device to be connected to the external device.

[0011] According to one embodiment, the sign language recognition module divides the video data into a plurality of segments, determines a recognition accuracy for each of the glosses of the plurality of segments, and extracts sign language sentences based on glosses of the plurality of segments whose recognition accuracy is greater than a predetermined value.

[0012] According to one embodiment, the recognition accuracy is determined based on the similarity between the gloss of the segment and a similar gloss, and the similar gloss can be characterized as being the gloss that is most similar to the gloss of the segment.

[0013] According to one embodiment, the sign language recognition module may be characterized by detecting the user's joints from the video data, extracting skeletal information for tracking the user's movements, and comparing the user's gloss based on the skeletal information with the similar gloss.

[0014] According to one embodiment, the display may be characterized by displaying a message requesting retransmission of the sign language sentence if the recognition accuracy of the glosses of the plurality of segments is all less than a predetermined value.

[0015] According to one embodiment, the sign language recognition module may extract sign language sentences based on glosses and previous dialogue content whose recognition accuracy is greater than a predetermined value.

[0016] According to one embodiment, when the glosses of the plurality of segments include a first gloss whose recognition accuracy is greater than a predetermined value and a second gloss whose recognition accuracy is less than the predetermined value, the sign language recognition module determines a plurality of candidate glosses to replace the second gloss based on the first gloss, and extracts a sign language sentence based on the gloss candidate selected from the plurality of candidate glosses and the first gloss.

[0017] According to one embodiment, when the glosses of the plurality of segments include a first gloss whose recognition accuracy is greater than a predetermined value and a second gloss whose recognition accuracy is less than the predetermined value, the sign language recognition module determines a plurality of candidate glosses to replace the second gloss based on the first gloss and the content of the previous dialogue, and extracts a sign language sentence based on the gloss candidate selected from the plurality of candidate glosses and the first gloss.

[0018] According to one embodiment, the sign language recognition module may prioritize the plurality of gloss candidates based on their similarity to the second gloss, and the display may display the plurality of gloss candidates according to the priority.

[0019] According to one embodiment, the display may be characterized as a transparent display.

[0020] According to one embodiment, the communication assistance device can constitute a communication system together with an input device that receives the user's voice or video.

[0021] The present disclosure provides a program for realizing various functions and commands of the communication assistance device, and a recording medium on which the program is stored. [Effects of the Invention]

[0022] The communication assisting device of the present disclosure can improve the accuracy and convenience of sign language-based communication. In particular, by improving the accuracy and convenience of communication for sign language users, sign language users can easily receive services provided to non-disabled people without the assistance of a professional sign language interpreter.

[0023] Furthermore, the device for assisting sign language recognition disclosed herein allows the user to easily control the start and end of sign language input, thereby increasing the convenience of sign language communication by allowing the user to input sign language images into the communication assistance device at a time of their choice. [Brief explanation of the drawings]

[0024] [Figure 1] 1 is a diagram illustrating a video display device for communication using sign language and a system including the same; [Figure 2] The manner in which the communication assistance device is used will now be described. [Figure 3] The manner in which the communication assistance device is used will now be described. [Figure 4] 1 illustrates one embodiment of an image input to a display. [Figure 5] An example of a method for inferring sign language sentences based on recognition accuracy will be described. [Figure 6] 1 is an example of skeleton information extracted from sign language video. DETAILED DESCRIPTION OF THE INVENTION

[0025] The present disclosure provides a communication aid for communication using sign language, the communication aid including a sign language recognition module that extracts sign language sentences from analyzed user movements of video data, and a display that displays the extracted sign language sentences.

[0026] Because the present disclosure can be modified in various ways and can have various embodiments, specific embodiments will be illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the disclosure to the specific embodiments, but rather to encompass all modifications, equivalents, or alternatives falling within the spirit and scope of the present disclosure. In the drawings, like reference numerals denote the same or similar functions throughout the various aspects. The shape and size of elements in the drawings may be exaggerated for clarity. The detailed description of exemplary embodiments below refers to the accompanying drawings, which show specific embodiments by way of example. These embodiments are described in sufficient detail to enable those skilled in the art to practice the embodiments. It should be understood that various embodiments, although different from one another, are not necessarily mutually exclusive. For example, specific shapes, structures, and characteristics described herein in connection with one embodiment can be implemented in other embodiments without departing from the spirit and scope of the present disclosure. It should also be understood that the location or arrangement of individual components within each disclosed embodiment can be changed without departing from the spirit and scope of the embodiment. Accordingly, the following detailed description is not intended to be taken in a limiting sense, and the scope of the exemplary embodiments is limited only by the appended claims, along with the full scope of equivalents to which such claims are entitled, if properly recited.

[0027] In this disclosure, terms such as "first," "second," etc. may be used to describe various components, but these components should not be limited by these terms. These terms are used only to distinguish one component from another. For example, a first component can be named a second component, and similarly, a second component can be named a first component, without departing from the scope of this disclosure. The term "and / or" includes a combination of multiple associated listed items or any of multiple associated listed items.

[0028] When a component of the present disclosure is referred to as being "coupled" or "connected" to another component, it should be understood that the component may be directly coupled or connected to the other component, or that there may be other components intervening between them. In contrast, when a component is referred to as being "directly coupled" or "directly connected" to another component, it should be understood that there are no other components intervening between them.

[0029] The components shown in the embodiments of the present disclosure are illustrated independently to show different characteristic functions, and do not mean that each component is composed of separate hardware or a single software unit. That is, each component is included side by side in each component for convenience of explanation, and at least two of the components may be combined into a single component, or one component may be divided into multiple components to perform a function. Such integrated and separated embodiments of each component are also within the scope of the present disclosure as long as they do not deviate from the essence of the present disclosure.

[0030] The terms used in this disclosure are merely used to describe particular embodiments and are not intended to limit the disclosure. A singular expression includes a plural expression unless the context clearly dictates otherwise. In this disclosure, terms such as "comprise" or "have" are intended to specify the presence of features, numbers, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to preclude the presence or possibility of addition of one or more other features, numbers, steps, operations, components, parts, or combinations thereof. In other words, in this disclosure, a description of a specific configuration as "comprising" does not exclude configurations other than the specified configuration, but means that additional configurations may be included within the scope of the implementation of the disclosure or the technical idea of the disclosure.

[0031] Some components of the present disclosure may not be essential components that perform essential functions in the present disclosure, but may be optional components simply for improving performance. The present disclosure can be realized by including only components essential for realizing the essence of the present disclosure, excluding components used simply for improving performance, and a structure including only essential components, excluding optional components used simply for improving performance, is also included in the scope of the present disclosure.

[0032] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings. When describing embodiments of the present specification, if it is determined that a detailed description of related known configurations or functions may obscure the gist of the specification, the detailed description will be omitted, and the same reference numerals will be used for the same components in the drawings, and duplicate descriptions of the same components will be omitted.

[0033] The present disclosure provides a display-based method and system for communicating in sign language and AAC. By providing various embodiments of the display-based method for communicating in sign language and AAC, the communication ability of sign language and AAC users can be enhanced.

[0034] Here, "sign language" refers to a language expressed using hands. And "AAC" stands for Augmentative and Alternative Communication. Specifically, "AAC" aims to improve the expressive ability of people who lack language skills by using images to represent sentences or words. Sign language and AAC are communication methods used by people who have difficulty communicating using speech. Sign language sentences can be divided into glosses, which are headwords in sign language. A gloss is the smallest unit of sign language, i.e., a semantic element.

[0035] Hereinafter, the video display device according to the present disclosure is used to assist communication between two or more people. In this disclosure, for the sake of convenience, a speaker who expresses his / her intention is described as a "user" and a listener to whom the user's intention is conveyed is described as a "conversation partner." Therefore, the positions of the "user" and the "conversation partner" can be interchanged in a conversation.

[0036] FIG. 1 shows an assistive device 110 for communication using sign language and a system 100 including the same.

[0037] The communication system 100 may include a communication aid 110 , an audio input unit 130 , a video input unit 140 , and a control device 150 .

[0038] The communication aid 110 may include a display 112. The display 112 may be a transparent display. Thus, two or more users may be located on opposite sides of the display 112 of the communication aid 110 and communicate with each other using sign language and AAC. Alternatively, if the display 112 is a general display, two or more users may be communicatively connected to each other, allowing two or more users located far apart to communicate with each other using sign language and AAC. Thus, a person who finds voice communication inconvenient in a non-face-to-face environment may convey text to the other party using sign language and / or AAC via the communication aid 110.

[0039] The display 112 can display different screen UI / UX (User Interface / User eXperience) depending on the user's characteristics (communication method). The display 112 also continues to display a certain portion of the existing dialogue content on the display 112, allowing the user to easily view the existing dialogue content at any time.

[0040] The communication aid 110 may additionally include certain modules to facilitate communication using sign language. Specifically, the communication aid 110 may include some of a Speech to Text (STT) module 114, a sign language generation module 116, a sign language recognition module 118, a flashcard selection module 120, a text input module 122, and a communication module 124.

[0041] The STT module 114 can convert voice data into text data. Specifically, the STT module 114 can convert voice data input via the voice input unit 130 into text data and transmit the text data to the display 112. The display 112 can then display the text data.

[0042] The sign language generation module 116 can convert audio data into sign language data. Specifically, the sign language generation module 116 can convert audio data input from the audio input unit 130 into sign language data and transmit the sign language data to the display 112. Then, the display 112 can display a sign language image based on the sign language data.

[0043] The sign language recognition module 118 analyzes the user's movements in the video data and extracts a sign language sentence of the user's intention from the user's movements. Furthermore, the sign language recognition module 118 can convert the extracted sign language sentence into text data and transmit the text data to the display 112. Then, the display 112 can display the text data.

[0044] The flashcard selection module 120 can provide flashcards on the display 112 so that the user can express simple meanings using AAC. Therefore, a user can communicate using voice, text, and sign language, and can more accurately convey his or her intentions to the other party by selecting a flashcard provided by the flashcard selection module 120. Here, the flashcard includes an image in which a word is visualized, so the other party can understand the user's intentions by looking at the image on the flashcard. For example, the flashcard selection module 120 can provide flashcards indicating the user's moods, such as happiness, boredom, sadness, annoyance, and anger, and can display the image included on the flashcard on the display 112 according to the user's selection. Therefore, the other party can easily understand the user's intentions by referring to the image included on the flashcard along with the text or sign language image entered by the user.

[0045] The text input module 122 can provide a text input user interface (UI) for the user to directly input text on the display 112. If the user finds it difficult to express his or her exact intentions in sign language or AAC, the text input UI of the text input module 122 can be used to directly communicate text to the other party.

[0046] The communication module 124 may allow the communication aid 110 to be communicatively connected to an external device. If mutual communication is difficult, the communication module 124 may be used to allow a third party, such as a sign language interpreter, to participate in the dialogue.

[0047] The audio input unit 130 can be realized by a device that receives input of audio information, such as a microphone. The video input unit 140 can be realized by a device that receives input of video information, such as a camera. The control device 150 can be used to control the start and end of sign language video recording.

[0048] FIG. 2 illustrates an embodiment of a usage mode of the communication aid device 110. As shown in FIG.

[0049] 2, a communication assistant device 110 implemented as a transparent display is positioned in the center of a desk, and two users 200 and 202 are positioned on opposite sides of the desk. Thus, users 200 and 202 can communicate using communication assistant device 110 without immediately meeting face-to-face. This not only prevents infection between users 200 and 202, but also allows people who have difficulty communicating to easily communicate their intentions to others by using the data input, conversion, and display functions of communication assistant device 110.

[0050] FIG. 3 illustrates an embodiment of a usage mode of the communication aid device 110. As shown in FIG.

[0051] Referring to FIG. 3, a first communication aid 300 and a second communication aid 310, each implemented with a non-transparent general display, are positioned in the center of each desk. User 200 uses the first communication aid 300, and user 202 uses the second communication aid 310. The first communication aid 300 and the second communication aid 310 are communicatively connected to each other, allowing users 200 and 202 to communicate with each other. Therefore, similar to the embodiment of FIG. 2, not only does this prevent infection between users 200 and 202, but the data input, conversion, and display functions of the communication aids 300 and 310 allow people who have communication difficulties to easily communicate their intentions to others. The first communication aid 300 and the second communication aid 310 of FIG. 3 may have a configuration similar to that of the communication aid 110 of FIG. 1.

[0052] FIG. 4 shows one embodiment of the video input to the display 112.

[0053] In FIG. 4, the display 112 can be realized as a transparent display. In this case, the other party can see the user's state as it is. In addition, the display 112 can display various functions provided by the communication assistance device 110 on the left side. The functions of the communication assistance device 110 are displayed as icons, and the user can activate the function corresponding to the icon by pressing an icon. In addition, the display 112 can display existing dialogue content on the right side. Therefore, the existing dialogue content continues to be displayed on the display 112, allowing the user to easily view the existing dialogue content at any time. The positions of the function icons and dialogue content on the display 112 can be determined differently depending on the embodiment.

[0054] Hereinafter, an embodiment of a method for converting sign language sentences into appropriate text when users communicate using sign language is provided.

[0055] When an actual sign language user inputs a sign language sentence into the communication assistant device 110, the analysis result is derived by analyzing words on a gloss basis. At this time, if the user's sign language movements are inaccurate or the sign language movement image is distorted due to the surrounding environment, the gloss may be recognized with a different meaning. Therefore, the sign language sentence may be interpreted differently from the user's intention.

[0056] To solve this problem, the present disclosure provides a method for calculating the recognition accuracy of each gloss, and if the recognition accuracy of a specific gloss is determined to be below a predetermined value, ignoring the result value for the specific gloss and inferring the meaning of the entire sign language sentence based on other sign language glosses with higher recognition accuracy. This method for inferring the meaning of the sign language sentence based on the recognition accuracy can be applied to the sign language recognition module 118.

[0057] In this disclosure, recognition accuracy refers to the degree of similarity between the current gloss and the most similar gloss that has already been learned. That is, if the current gloss closely matches a particular most similar gloss, the recognition accuracy can be determined to be close to 100%. Conversely, if the current gloss does not clearly correspond to any gloss, the recognition accuracy can be determined to be low.

[0058] The predetermined value is any value between 10% and 90%. As the predetermined value is lower, the sign language recognition module 118 may generate a sign language sentence using glosses with low recognition accuracy, which may increase the error rate. Conversely, as the predetermined value is higher, the sign language recognition module 118 may generate a sign language sentence using only glosses with high recognition accuracy, which may reduce the error rate. However, filtering out too many glosses may make it difficult to infer and complete the entire sign language sentence. Therefore, in order to reduce errors in sign language sentence interpretation while increasing convenience, it is necessary to determine the predetermined value within an appropriate range.

[0059] FIG. 5 illustrates an example of a method for inferring sign language sentences based on recognition accuracy.

[0060] In step 510, a sign language sentence meaning "Where is the toilet?" is input. The sign language sentence is composed of a gloss meaning "toilet" and a gloss meaning "where." However, if either of the sign language expressions "toilet" or "where" is incorrectly recognized, the sign language sentence may be translated into a completely different meaning. FIG. 5 illustrates a method for inferring a sign language sentence based on recognition accuracy, assuming that the sign language action corresponding to "where" is inaccurate.

[0061] In steps 520 and 530, an existing sign language sentence construction method that is not based on recognition accuracy is described. As described above, in step 520, if the sign language action corresponding to "where" is inaccurate, the sign language action may be erroneously recognized as "eat." Therefore, in step 530, the sign language sentence may be translated as "Do you want to eat the toilet?"

[0062] To solve this problem, in steps 540 and 550, the meaning of the sign language sentence can be inferred using only sign language glosses with a recognition accuracy of 50% or higher.

[0063] In step 540, the recognition accuracy for the two glosses can be calculated. At this time, the most similar word for the two glosses is determined. For example, the most similar word for the gloss corresponding to "toilet" may be correctly recognized as "toilet," while the most similar word for the gloss corresponding to "where" may be incorrectly recognized as "eat." Then, the recognition accuracy between the most similar words for each gloss is calculated. For example, the recognition accuracy for the gloss corresponding to "toilet" may be calculated to be 80%, and the recognition accuracy for "eat" may be calculated to be 35%.

[0064] In step 550, the entire sign language sentence is inferred based on glosses with a recognition accuracy of more than 50%. Therefore, "eat" (eat) with a recognition accuracy of less than 50% is ignored in the sign language sentence inference process. That is, the meaning of the sign language sentence can be inferred based on the sign for "toilet" with a recognition accuracy of more than 50%. For example, sentence candidates such as "Where is the toilet?" and "Please show me the toilet" can be suggested. Then, the sign language sentence can be translated into text according to the user's selection.

[0065] According to one embodiment, if the recognition accuracy of all glosses of a sign language sentence is below a predetermined value, the meaning of the sign language sentence cannot be inferred. The communication assistant device 110 can then request a retransmission of the sign language sentence. For example, the display 112 can display a message such as, "The sign language could not be properly recognized. Please sign again."

[0066] According to one embodiment, when glosses of a plurality of segments include a first gloss whose recognition accuracy is greater than a predetermined value and a second gloss whose recognition accuracy is less than a predetermined value, the sign language recognition module 118 can determine multiple gloss candidates to replace the second gloss based on the first gloss. Alternatively, in the above case, the sign language recognition module 118 can determine multiple gloss candidates to replace the second gloss based on the first gloss and the content of the previous dialogue. Then, the sign language recognition module 118 can extract a sign language sentence based on a gloss candidate selected from the multiple gloss candidates and the first gloss. In this case, the sign language recognition module 118 can prioritize the multiple gloss candidates based on their similarity to the second gloss, and the display 112 can display the multiple gloss candidates in order of priority.

[0067] According to one embodiment, the sign language recognition module 118 can infer the meaning of a sign language sentence by taking into account the existing dialogue content. For example, if the only gloss in a sign language sentence with a recognition accuracy of 50% or higher is "toilet," the sign language recognition module 118 can complete a sign language sentence containing "toilet" by taking into account the existing dialogue content.

[0068] According to one embodiment, if a user asks the question "Where is the restroom?" in sign language, the sign language recognition module 118 can recognize the user's gender through video and provide guidance on the location of the men's restroom if the user is male, or the location of the women's restroom if the user is female.

[0069] A method for deriving the gross recognition accuracy will be described below.

[0070] First, a sign language video of a user is input via the video input unit 140. Then, the sign language recognition module 118 uses artificial intelligence technology to recognize the user in the sign language video and extracts skeletal information for tracking the user's movements by detecting the user's joints. FIG. 6 of the present application shows an example of skeletal information extracted from the sign language video. The sign language recognition module 118 can also compare the user's movements based on the skeletal information with previously stored gross movements having specific meanings. The degree of similarity between the two is then determined as the recognition accuracy of the current gross.

[0071] The sign language recognition module 118 may include an AI learning model for inferring glosses from glosses and an AI learning model for inferring natural language sentences from glosses. The AI learning model may be configured with a convolutional neural network (CNN) and a transformer model. The AI learning model may be trained using training data composed of sign language actions and glosses, and training data composed of glosses and natural language sentences.

[0072] The amount of training data can be increased up to 100 times or more using proprietary data augmentation techniques (shift, resize, frame manipulation, etc.). In addition, to prevent overfitting in each sign language translation step, motion data that is not the target of translation and the results of a general natural language model can be used to train the AI learning model.

[0073] The AI learning model is trained based on continuous glosses. Specifically, continuous glosses can be divided into multiple segments. Then, in the training step, the label probability for each segment is calculated. An unknown label is assigned to unlearned actions.

[0074] The sign language recognition module 118 can infer the glossary meaning of the video using the trained AI learning model. In this case, the sign language recognition module 118 can divide the input sign language video into a plurality of segments. The sign language recognition module 118 can then determine the expression with the highest ranking among the sign language expression probabilities of each segment. After grasping all the sign language expressions for each action, the sign language recognition module 118 can translate the entire sign language expressions into a general natural language sentence. The inference result of the sign language recognition module 118 can be output as two items: an array of sign language expressions and a general natural language sentence string.

[0075] The control device 150 will now be described.

[0076] In order to recognize sign language, it is important to recognize when the user starts and stops signing. Sign language video recording can be started by inputting a signal to the start button of the control device 150. Sign language video recording can be automatically ended one second after both hands are out of the camera's view. Once recording is finished, inference can be made about the sign language expression based on the recorded sign language video.

[0077] The control device 150 may be a personal smartphone. In this case, the smartphone can be used as a remote controller. Alternatively, the control device 150 may be a dedicated device including a shooting or recording button. The control device 150 can be used to control the start and end of sign language image recognition. To improve the user experience, a remote control webpage can be developed for the user's smartphone. Then, the webpage can be easily accessed by using the control device 150 to take a photo of a webpage in a tablet or PC environment or a QR marker installed in the real space.

[0078] If user authorization is performed by scanning a QR marker, automatic login can be performed using the ID of the attached PC or tablet. If not, each device can log in with its own ID / PW and connect uniquely. In addition, to prevent duplicate connections, simultaneous connections can be limited so that if a device has already connected, other devices cannot connect.

[0079] When the capture button on the control device 150 is pressed, the same process as pressing the capture button on a tablet or PC can be performed. When the record button on the control device 150 is pressed, the same process as pressing the record button on a tablet or PC can be performed. However, the subject of video capture is a PC or tablet installed in front of the user, and the subject of audio recording is the microphone of the smartphone held by the user. The subjects of video capture and audio recording can be changed.

[0080] Alternatively, the control device 150 can be implemented as a foot button, which can be used to determine the start and end points of sign language recognition.

[0081] According to one embodiment, the start and end points of sign language recognition can be determined from the recognition of a particular hand shape, independent of the control device 150. For example, sign language recognition can begin when a hand suddenly rises from outside the lower screen and enters the screen, and can end when a hand drops down from within the screen and exits the screen.

[0082] The various embodiments of the present disclosure are not intended to enumerate all possible combinations, but rather to describe representative aspects of the present disclosure, and the matters described in the various embodiments may be applied independently or in combination of two or more.

[0083] Furthermore, various embodiments of the present disclosure may be realized by hardware, firmware, software, or a combination thereof. When realized by hardware, the hardware may be realized by one or more Application Specific Integrated Circuits (ASICs), Digital Signal Processors (DSPs), Digital Signal Processing Devices (DSPDs), Programmable Logic Devices (PLDs), Field Programmable Gate Arrays (FPGAs), general processors, controllers, microcontrollers, microprocessors, etc. It is clear that the hardware may be realized in the form of a program stored in a non-transitory computer-readable medium that can be used at a terminal or edge, or in the form of a program stored in a non-transitory computer-readable medium that can be used at an edge or in the cloud.

[0084] For example, the information display method according to one embodiment of the present disclosure can be realized in the form of a program stored on a non-transitory computer-readable medium, and the method for performing phase spreading in a directionally based block unit described above can also be realized in the form of a computer program.

[0085] The scope of the present disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that cause operations according to the methods of the various embodiments to be performed on a device or computer, as well as non-transitory computer-readable media on which such software or commands can be stored and executed on a device or computer.

[0086] The present invention described above is susceptible to various substitutions, modifications, and alterations by those skilled in the art without departing from the technical spirit of the present invention, and therefore the scope of the present invention is not limited to the above-described embodiments and accompanying drawings.

Claims

1. A communication aid for communication using sign language, a sign language recognition module that extracts sign language sentences from the analyzed user movements in the video data; a display that displays the extracted sign language sentence.

2. The communication aid device is an STT module that converts voice data into text data; The communication aid according to claim 1 , further comprising: a sign language generation module that converts voice data into sign language data.

3. The communication aid device is a word card selection module that provides word cards selectable by the user on the display; The communication assistance device according to claim 1 , wherein the sign language recognition module extracts the sign language sentence based on a selected word card.

4. The communication aid device is a text input module for providing a user interface on the display for a user to input text; 2. The communication aid according to claim 1, wherein the text input module is activated when the sign language recognition module fails to extract the sign language sentence.

5. The communication aid device is a communication module that controls the communication aid device to be communicatively connected to an external device; The communication assistance device according to claim 1 , wherein when the sign language recognition module fails to extract the sign language sentence, the communication module controls the communication assistance device to be connected to the external device.

6. The sign language recognition module Dividing the video data into a plurality of segments; determining a recognition accuracy for each gloss of the plurality of segments; 2. The communication assistance device according to claim 1, wherein the sign language sentence is extracted based on a gloss of the plurality of segments whose recognition accuracy is greater than a predetermined value.

7. The recognition accuracy is determined based on a similarity between the gloss of the segment and a similar gloss; 7. The communication aid according to claim 6, wherein the similar gloss is a gloss that is most similar to the gloss of the segment.

8. The sign language recognition module The communication assistance device of claim 7, characterized in that it extracts skeleton information for tracking the user's movements by detecting the user's joints from the video data, and compares the user's gross based on the skeleton information with the similar gross.

9. The display includes:

7. The communication assistance device according to claim 6, wherein if the recognition accuracy of the glosses of the plurality of segments is lower than a predetermined value, a message requesting retransmission of the sign language sentence is displayed.

10. The sign language recognition module 7. The communication assistance device according to claim 6, wherein the sign language sentence is extracted based on the gloss having a recognition accuracy greater than a predetermined value and the content of a previous dialogue.

11. The sign language recognition module If the glosses of the plurality of segments include a first gloss having a recognition accuracy greater than a predetermined value and a second gloss having a recognition accuracy less than the predetermined value, determining a plurality of gloss candidates to replace the second gloss based on the first gloss; 7. The communication assistance device according to claim 6, wherein a sign language sentence is extracted based on a gloss candidate selected from the plurality of gloss candidates and a first gloss.

12. The sign language recognition module If the glosses of the plurality of segments include a first gloss having a recognition accuracy greater than a predetermined value and a second gloss having a recognition accuracy less than a predetermined value, determining a plurality of candidate glosses to replace the second gloss based on the first gloss and the content of the previous dialogue; 7. The communication assistance device according to claim 6, wherein a sign language sentence is extracted based on a gloss candidate selected from the plurality of gloss candidates and a first gloss.

13. The sign language recognition module determining a priority order for the plurality of gloss candidates based on a similarity to the second gloss; The display includes:

12. The communication assistance device according to claim 11, wherein the plurality of gloss candidates are displayed in accordance with the priority order.

14. 2. The communication aid according to claim 1, wherein the display is a transparent display.

Citation Information

Patent Citations

  • Translation method, mobile terminal and computer readable storage medium

    CN109960813A

  • Sign language translation device

    KR101915088B1

  • Sign language translator, system and method

    KR1020160109708A

  • Method of Casting using Molten Metal Pouring and Additive

    KR1020230055064A

  • Sign language translation apparatus using gloss and translation model learning apparatus

    KR102115551B1