Display-based communication systems
The display-based communication assistance device addresses communication barriers for hearing-impaired individuals by enhancing sign language recognition and conversion, ensuring accurate and convenient interactions.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- BATONERS INC
- Filing Date
- 2023-09-04
- Publication Date
- 2026-07-24
AI Technical Summary
Hearing-impaired individuals face challenges in communication due to transparent shielding films obstructing lip-reading and sign language recognition, especially during the COVID-19 pandemic, necessitating improved accessibility for sign language users.
A display-based communication assistance device with modules for sign language recognition, generation, and conversion, including a sign language recognition module that extracts text from user movements, a display for output, and additional modules for voice-to-text and text input, ensuring accurate and convenient communication.
Enhances the accuracy and convenience of sign language communication, allowing hearing-impaired individuals to interact effectively without professional interpreters, even in non-face-to-face settings, by improving recognition and conversion processes.
Smart Images

Figure 0007894958000001 
Figure 0007894958000002 
Figure 0007894958000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a communication system, and more specifically, to a display-based communication system.
Background Art
[0002] Hearing-impaired persons are a general term for people with reduced hearing or lost hearing function. Hearing-impaired persons can communicate in three main ways according to the degree of hearing impairment. First, when the degree of hearing impairment is low, it is possible to communicate verbally with non-disabled persons by enhancing hearing using a hearing aid or the like. Second, it is possible to communicate with non-disabled persons by using a lip-reading method of inferring what is being said by looking at the mouth shape of the other person. Finally, it is possible to communicate with non-disabled persons by using sign language.
[0003] For hearing-impaired persons, communication has already been difficult, and due to COVID-19 that began to spread in 2020, the difficulty of communication has further deepened. For example, due to the spread of COVID-19, transparent shielding films for droplet prevention are installed at reception desks and counseling counters, etc., making it difficult for hearing-impaired persons to hear the voices of others. Also, especially when the transparent shielding film is contaminated, hearing-impaired persons who use the lip-reading method or sign language cannot clearly see the mouth shape or sign language movements,加重 the difficulty of communication.
[0004] Also, since most services are not based on sign language, various technological developments are required to improve the service accessibility for sign language users.
Summary of the Invention
Problems to be Solved by the Invention
[0005] An object of the present disclosure is to provide a display-based sign language communication system for increasing the accuracy and convenience of sign language communication. <000002३>
Means for Solving the Problems
[0006] This disclosure provides a communication assistance device for communication using sign language, which includes a sign language recognition module that extracts sign language text from user movements analyzed from video data, and a display that displays the extracted sign language text.
[0007] According to one embodiment, the communication assistance device may further include an STT module that converts voice data into text data, and a sign language generation module that converts voice data into sign language data.
[0008] According to one embodiment, the communication assistance device further includes a word card selection module that provides the user with selectable word cards on the display, and the sign language recognition module extracts the sign language text based on the selected word card.
[0009] According to one embodiment, the communication assistance device further includes a text input module that provides a user interface for the user to input text to the display, and is characterized in that the text input module is activated when the sign language recognition module fails to extract the sign language text.
[0010] According to one embodiment, the communication assistance device further includes a communication module that controls the communication assistance device to be communicatively connected to an external device, and is characterized in that when the sign language recognition module fails to extract the sign language text, the communication module controls the communication assistance device to be connected to the external device.
[0011] According to one embodiment, the sign language recognition module divides the video data into a plurality of segments, determines the recognition accuracy of each of the plurality of segments, and extracts sign language text from the plurality of segments based on the segment whose recognition accuracy is greater than a predetermined value.
[0012] According to one embodiment, the recognition accuracy is determined based on the similarity between the segment's gross and a similar gross, wherein the similar gross is the gross that is most similar to the segment's gross.
[0013] According to one embodiment, the sign language recognition module may be characterized by detecting the user's joint locations from the video data, extracting skeleton information for tracking the user's movements, and comparing the user's gross based on the skeleton information with similar gross.
[0014] According to one embodiment, the display may display a message requesting retransmission of the sign language text if the recognition accuracy of the gross of the plurality of segments is lower than a predetermined value.
[0015] According to one embodiment, the sign language recognition module may be characterized by extracting sign language text based on gross and previous dialogue content where the recognition accuracy is greater than a predetermined value.
[0016] According to one embodiment, the sign language recognition module may, when the gross of the plurality of segments includes a first gross whose recognition accuracy is greater than a predetermined value and a second gross whose recognition accuracy is less than a predetermined value, determine a plurality of gross candidates to replace the second gross based on the first gross, and extract a sign language sentence based on the gross candidate selected from the plurality of gross candidates and the first gross.
[0017] According to one embodiment, the sign language recognition module, when the gross of the plurality of segments includes a first gross whose recognition accuracy is greater than a predetermined value and a second gross whose recognition accuracy is less than a predetermined value, determines a plurality of gross candidates to replace the second gross based on the first gross and the previous dialogue content, and extracts a sign language text based on the gross candidate selected from the plurality of gross candidates and the first gross.
[0018] According to one embodiment, the sign language recognition module determines the priority order of the plurality of gross candidates based on their similarity to the second gross, and the display displays the plurality of gross candidates according to the priority order.
[0019] According to one embodiment, the display can be characterized as a transparent display.
[0020] According to one embodiment, the communication assistance device can constitute a communication system together with an input device that receives the user's voice or video.
[0021] This disclosure provides a program that implements various functions and commands of the communication assistance device, and a recording medium on which the program is stored. [Effects of the Invention]
[0022] The communication assistance device disclosed herein can improve the accuracy and convenience of sign language-based communication. In particular, by improving the accuracy and convenience of communication for sign language users, sign language users will be able to receive services provided to non-disabled persons without inconvenience, even without the assistance of a professional sign language interpreter.
[0023] In addition, a device for sign language recognition assistance according to the present disclosure enables a user to easily control the start and end of sign language input. Thus, by allowing the user to input sign language video into the communication assistance device at their desired time, the convenience of sign language communication can be increased.
Brief Description of the Drawings
[0024] [Figure 1] It is a diagram illustrating a video display device for communication using sign language and a system including the same. [Figure 2] The usage mode of the communication assistance device is described. [Figure 3] The usage mode of the communication assistance device is described. [Figure 4] An embodiment of the video input to the display is shown. [Figure 5] An example of a method for analogizing sign language sentences based on recognition accuracy is described. [Figure 6] It is an example of skeleton information extracted from sign language video.
Modes for Carrying Out the Invention
[0025] In the present disclosure, there is provided a communication assistance device for communication using sign language, including a sign language recognition module that extracts sign language sentences from analyzed user movements of video data, and a display that displays the extracted sign language sentences.
[0026] This disclosure is subject to various modifications and can have a variety of embodiments; therefore, specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this should not be understood as limiting the disclosure to specific embodiments, but rather as including all modifications, equivalents, or substitutions that fall within the spirit and technical scope of this disclosure. Similar reference numerals in the drawings refer to the same or similar functions across various aspects. The shape and size of elements in the drawings may be exaggerated for clarity. Detailed descriptions of exemplary embodiments described later refer to accompanying drawings illustrating specific embodiments. These embodiments are described in sufficient detail so that those skilled in the art can carry them out. It should be understood that the various embodiments are different from one another but do not necessarily have to be mutually exclusive. For example, certain shapes, structures, and characteristics described herein can be realized in other embodiments in relation to one embodiment without departing from the spirit and scope of this disclosure. It should also be understood that the position or arrangement of individual components within each disclosed embodiment can be changed without departing from the spirit and scope of the embodiment. Therefore, the detailed descriptions set forth below are not intended to be restrictive, and the scope of the exemplary embodiments, along with all equivalents to those claimed, are limited only by the appended claims, if appropriately described.
[0027] In this disclosure, terms such as “First,” “Second,” etc., may be used to describe various components, but these components should not be limited by these terms. These terms are used solely for the purpose of distinguishing one component from another. For example, without exceeding the scope of the rights of this disclosure, the first component may be named the second component, and similarly, the second component may be named the first component. The term “and / or” includes a combination of multiple related descriptions or any of multiple related descriptions.
[0028] When a component of this disclosure is described as being "linked" or "connected" to another component, it should be understood that it may be directly linked or connected to the other component, or that another component may be interposed between them. Conversely, when a component is described as being "directly linked" or "directly connected" to another component, it should be understood that there is no other component interposed between them.
[0029] The components shown in the embodiments of this disclosure are illustrated independently to illustrate distinct and characteristic functions, and do not imply that each component consists of a separate hardware or software unit. That is, each component is included alongside the others for convenience of explanation, and at least two of the components may be combined to form a single component, or a single component may be divided into multiple components to perform functions. Such integrated and separated embodiments of each component are included within the scope of this disclosure, as long as they do not deviate from the essence of this disclosure.
[0030] The terms used in this disclosure are used solely to describe specific embodiments and are not intended to limit the disclosure. Singular expressions include plural expressions unless the context clearly indicates otherwise. In this disclosure, terms such as “includes” or “has” are intended to specify the presence of features, figures, steps, operations, components, parts, or combinations thereof described in the specification, and should be understood not to pre-exist the presence or possibility of adding one or more other features, figures, steps, operations, components, parts, or combinations thereof. In other words, when a particular configuration is described as “includes” in this disclosure, it does not exclude other configurations, but rather means that additional configurations may be included within the scope of the implementation of this disclosure or the technical idea of this disclosure.
[0031] Some components of this disclosure may not be essential components that perform an essential function in this disclosure, but may be optional components that merely improve performance. This disclosure may be implemented by including only the components that are essential to realizing the essence of this disclosure, excluding components used merely for performance improvement, and a structure including only the essential components, excluding optional components used merely for performance improvement, is also included within the scope of rights of this disclosure.
[0032] Embodiments of this disclosure will be described below with reference to the drawings. In describing embodiments of this specification, if it is determined that a specific description of a related known configuration or function would obscure the gist of this specification, such detailed description will be omitted, the same reference numerals will be used for the same components in the drawings, and redundant descriptions of the same components will be omitted.
[0033] This disclosure provides a method and system for communicating in sign language and AAC based on a display. By providing various embodiments of the display-based sign language and AAC communication method in this disclosure, the communication capabilities of sign language and AAC users can be enhanced.
[0034] Here, "sign language" refers to a language expressed using hands. "AAC" stands for Augmentative and Alternative Communication. Specifically, AAC aims to improve the ability of people with insufficient language skills to express themselves using images that represent sentences or words. Sign language and AAC are communication methods used by people who have difficulty communicating using speech. Sign language sentences can be divided into gross units, which are the headwords of sign language. A gross unit refers to the smallest unit of sign language, i.e., a word, or semantic element.
[0035] Hereinafter, the video display device described herein is used to assist communication between two or more people. For the sake of clarity, in this disclosure, the speaker who expresses their intentions is referred to as the "user," and the listener who receives the user's intentions is referred to as the "other party." Therefore, the positions of the "user" and the "other party" can be reversed in a dialogue.
[0036] Figure 1 shows an auxiliary device 110 for communication using sign language, and a system 100 including it.
[0037] The communication system 100 may include a communication assistance device 110, an audio input unit 130, a video input unit 140, and a control device unit 150.
[0038] The communication assistance device 110 may include a display 112. The display 112 can be a transparent display. Therefore, two or more users can communicate with each other using sign language and AAC while being located on opposite sides of the display 112 of the communication assistance device 110. Alternatively, if the display 112 is a standard display, multiple communication assistance devices 110 can be communicatively connected, allowing two or more users located at a distance from each other to communicate using sign language and AAC. Thus, in a non-face-to-face environment, a person who has difficulty with voice communication can transmit text to another person using sign language and / or AAC via the communication assistance device 110.
[0039] Display 112 can display other screen UI / UX (User Interface / User Experience) depending on the user's characteristics (communication method). Furthermore, by continuously displaying a certain portion of the existing conversation content on Display 112, the user can easily view the existing conversation content at any time.
[0040] The communication assistance device 110 may include additional modules to facilitate communication using sign language. Specifically, the communication assistance device 110 may include some of the following: an STT (Speech to Text) module 114, a sign language generation module 116, a sign language recognition module 118, a word card selection module 120, a text input module 122, and a communication module 124.
[0041] The STT module 114 can convert audio data into text data. Specifically, the STT module 114 can convert audio data input by the audio input unit 130 into text data and transmit that text data to the display 112. The display 112 can then display the text data.
[0042] The sign language generation module 116 can convert audio data into sign language data. Specifically, the sign language generation module 116 can convert audio data input from the audio input unit 130 into sign language data and transmit that sign language data to the display 112. The display 112 can then display sign language video using the sign language data.
[0043] The sign language recognition module 118 analyzes the user's movements in the video data and extracts sign language sentences that represent the user's intentions from those movements. Furthermore, the sign language recognition module 118 can convert the extracted sign language sentences into text data and transmit that text data to the display 112. The display 112 can then display the text data.
[0044] The word card selection module 120 can provide word cards on the display 112 so that the user can express simple semantic expressions in AAC. Therefore, the user can communicate their intentions more accurately to the other party by selecting word cards provided by the word card selection module 120, in addition to communicating by voice, text, and sign language. Here, since the word cards include images in which the words are visualized, the other party can understand the user's intentions by looking at the images on the word cards. For example, the word card selection module 120 can provide word cards that represent the user's moods, such as happiness, boredom, sadness, disgust, and anger, and can display the images contained in the word cards on the display 112 according to the user's selection. Therefore, the other party can easily understand the user's intentions by referring to the images contained in the word cards along with the text or sign language images entered by the user.
[0045] The text input module 122 can provide a text input UI (User Interface) on the display 112 for the user to directly input text. If a user finds it difficult to express their precise intentions using sign language or AAC, they can use the text input UI of the text input module 122 to directly communicate text to the other party.
[0046] The communication module 124 allows the communication assistance device 110 to communicate with an external device. If there are difficulties in mutual communication, a third party, such as a sign language interpreter, can participate in the conversation using the communication module 124.
[0047] The audio input unit 130 can be implemented using a device that receives audio information, such as a microphone. The video input unit 140 can be implemented using a device that receives video information, such as a camera. The control device 150 can be used to control the start and end of sign language video recording.
[0048] Figure 2 illustrates one embodiment of how the communication assistance device 110 is used.
[0049] As shown in Figure 2, a communication assistance device 110, implemented with a transparent display, is positioned in the center of the desk, with two users 200 and 202 positioned on the opposite side of the desk. Therefore, users 200 and 202 can communicate using the communication assistance device 110 without immediately facing each other. Thus, not only is infection prevented between users 200 and 202, but the data input, conversion, and display functions of the communication assistance device 110 allow people who have difficulty communicating to easily convey their intentions to the other party.
[0050] Figure 3 illustrates one embodiment of how the communication assistance device 110 is used.
[0051] As shown in Figure 3, a first communication aid device 300 and a second communication aid device 310, both implemented using non-translucent general displays, are positioned in the center of each desk. User 200 can use the first communication aid device 300, and user 202 can use the second communication aid device 310. The first communication aid device 300 and the second communication aid device 310 are communicatively connected to each other, allowing users 200 and 202 to communicate with one another. Therefore, similar to the embodiment in Figure 2, not only is infection prevented between users 200 and 202, but the data input, conversion, and display functions of the communication aid devices 300 and 310 allow people who have difficulty communicating to easily convey their intentions to others. The first communication aid device 300 and the second communication aid device 310 in Figure 3 can have the same configuration as the communication aid device 110 in Figure 1.
[0052] Figure 4 shows one embodiment of the video input to the display 112.
[0053] In Figure 4, the display 112 can be implemented as a transparent display. In this case, the other party can see the user's actions directly. The display 112 can also display various functions provided by the communication assistance device 110 on the left side. The functions of the communication assistance device 110 are displayed as icons, and the user can activate the function corresponding to the icon by pressing it. The display 112 can also display the existing conversation content on the right side. Therefore, by continuously displaying the existing conversation content on the display 112, the user can easily view the existing conversation content at any time. The positions of the function icons and conversation content on the display 112 can be determined differently depending on the embodiment.
[0054] The following provides an embodiment of a method for converting sign language text into appropriate text when a user communicates using sign language.
[0055] When an actual sign language user inputs sign language text into the communication assistance device 110, the words are analyzed in gross units, and the analysis results are derived. At this time, if the user's sign language movements are inaccurate, or if the sign language movement video is distorted by the surrounding environment, the gross units may be perceived with a different meaning. Therefore, the sign language text may be interpreted in a way that differs from the user's intention.
[0056] To address this, this disclosure provides a method for calculating the recognition accuracy of each gross sign language sign and, if the recognition accuracy of a particular gross sign language sign is determined to be below a predetermined value, ignoring that result for that particular gross sign language sign and inferring the meaning of the entire sign language text based on other sign language gross sign language signs with higher recognition accuracy. Such a method for inferring the meaning of a sign language text based on recognition accuracy can be applied to the sign language recognition module 118.
[0057] In this disclosure, recognition accuracy refers to the similarity between the current gross and the closest similar gross that has already been learned. That is, if the current gross closely matches a particular closest similar gross, the recognition accuracy can be determined to be close to 100%. Conversely, if the current gross does not clearly correspond to any gross, the recognition accuracy can be determined to be low.
[0058] The predetermined value is any value between 10% and 90%. The lower the predetermined value, the more the sign language recognition module 118 can generate sign language sentences using gross signs with low recognition accuracy, which can increase the error rate. Conversely, the higher the predetermined value, the more the sign language recognition module 118 will generate sign language sentences using only gross signs with high recognition accuracy, which can decrease the error rate. However, filtering out too many gross signs can make it difficult to infer and complete the entire sign language sentence. Therefore, in order to reduce errors in interpreting sign language sentences while increasing convenience, it is necessary to determine the predetermined value within an appropriate range.
[0059] Figure 5 illustrates an example of a method for inferring sign language text based on recognition accuracy.
[0060] In step 510, a sign language sentence meaning "Where is the toilet?" is entered. This sign language sentence consists of a gross sign meaning "toilet" and a gross sign meaning "where". However, if either of the sign language expressions for "toilet" or "where" is misrecognized, the sign language sentence can be translated with a completely different meaning. Figure 5 explains a method for inferring the sign language sentence based on recognition accuracy, assuming that the sign language action corresponding to "where" is inaccurate.
[0061] Steps 520 and 530 describe existing sign language sentence construction methods that are not based on recognition accuracy. In step 520, as mentioned above, if the sign language action corresponding to "where" is inaccurate, the sign language action may be misrecognized as "eat." Therefore, in step 530, the sign language sentence may be translated as "Do you eat the toilet?"
[0062] To address this, in steps 540 and 550, the meaning of the sign language text can be inferred using only sign language gross with a recognition accuracy of 50% or higher.
[0063] In step 540, the recognition accuracy for the two gross words can be calculated. At this time, the word most similar to the two gross words is determined. For example, the word most similar to the gross word corresponding to "toilet" may be accurately recognized as "toilet," while the word most similar to the gross word corresponding to "where" may be incorrectly recognized as "eat." Then, the recognition accuracy between the words most similar to each gross word is calculated. For example, the recognition accuracy for the gross word corresponding to "toilet" can be calculated as 80%, and the recognition accuracy for "eat" can be calculated as 35%.
[0064] In step 550, the entire sign language sentence is inferred based on gross signs with a recognition accuracy of 50% or higher. Therefore, signs like "eat" with a recognition accuracy of 50% or lower are ignored in the sign language sentence inference process. In other words, the meaning of the sign language sentence can be inferred based on the sign for "toilet" with a recognition accuracy of 50% or higher. For example, sentence candidates such as "Where is the toilet?" and "Please show me to the toilet." can be suggested. Then, depending on the user's selection, the sign language sentence can be translated into text.
[0065] According to one embodiment, if the recognition accuracy of all gross signs in a sign language text is below a predetermined value, the meaning of the sign language text cannot be inferred. The communication assistance device 110 can then request retransmission of the sign language text. For example, the display 112 may display a message such as, "We were unable to properly recognize your sign language. Please sign again."
[0066] According to one embodiment, if the gross of a plurality of segments includes a first gross whose recognition accuracy is greater than a predetermined value and a second gross whose recognition accuracy is less than a predetermined value, the sign language recognition module 118 can determine a plurality of gross candidates to replace the second gross based on the first gross. Alternatively, in the above case, the sign language recognition module 118 can determine a plurality of gross candidates to replace the second gross based on the first gross and the content of the previous dialogue. The sign language recognition module 118 can then extract a sign language text based on the selected gross candidate and the first gross. At this time, the sign language recognition module 118 can determine the priority order of the plurality of gross candidates based on their similarity to the second gross, and the display 112 can display the plurality of gross candidates according to the priority order.
[0067] According to one embodiment, the sign language recognition module 118 can infer the meaning of a sign language sentence by considering the existing dialogue content. For example, if the only sign language sentence with a recognition accuracy of 50% or more is "toilet," the module can complete a sign language sentence including "toilet" by considering the existing dialogue content.
[0068] According to one embodiment, if a user asks the question "Where is the toilet?" in sign language, the sign language recognition module 118 can recognize the user's gender via video and guide them to the location of the men's toilet if the user is male, or to the location of the women's toilet if the user is female.
[0069] The method for deriving Gross's recognition accuracy is described below.
[0070] First, a video of the user's sign language is input via the video input unit 140. The sign language recognition module 118 then uses artificial intelligence technology to recognize the user in the sign language video and detect the user's joints, thereby extracting skeleton information for tracking the user's movements. Figure 6 of this application shows an example of skeleton information extracted from a sign language video. The sign language recognition module 118 can also compare the user's movements based on the skeleton information with existing stored gross movements that have a specific meaning. The degree of similarity between the two is then determined as the recognition accuracy of the gross.
[0071] The sign language recognition module 118 may include an AI learning model for inferring sign language from sign language, and an AI learning model for inferring natural language sentences from sign language. The AI learning models can consist of a CNN (Convolutional Neural Network) and a Transformer model. The AI learning models can be trained using training data consisting of sign language actions and sign language, and training data consisting of sign language and natural language sentences.
[0072] The aforementioned training data can be augmented by more than 100 times using proprietary data augmentation techniques (such as shift, resize, and frame manipulation). Furthermore, to prevent overfitting in each sign language translation step, non-translatable motion data and results from general natural language models can be used to train the AI learning model.
[0073] AI learning models are trained based on a continuous gross. Specifically, a continuous gross can be divided into multiple segments. Then, in the learning step, the probability of a label for each segment is calculated. In addition, an "Unknown Label" is assigned to actions that have not been learned.
[0074] The sign language recognition module 118 can then infer the overall meaning of the video using the trained AI learning model. At this time, the sign language recognition module 118 can divide the input sign language video into multiple segments. The sign language recognition module 118 can then determine the highest-ranking sign language expression among the probability of each segment. After grasping all the sign language expressions for each action, the sign language recognition module 118 can translate the overall sign language expression into a natural language sentence. The inference result of the sign language recognition module 118 can output two things: an array of sign language expressions and a string of natural language sentences.
[0075] The control device 150 will be described below.
[0076] For sign language to be recognized, it is important that the user recognizes when they begin and end their sign language. Sign language video recording can be started by inputting a signal to the start button on the control device 150. Sign language video recording can then be automatically stopped one second after both hands have completely disappeared from the camera's view. Once recording is complete, inferences about the sign language expression can be made based on the recorded sign language video.
[0077] The control device 150 may be a personal smartphone. In this case, the smartphone can be used as a remote controller. Alternatively, the control device 150 may be a dedicated device including a shooting or recording button. The start and end of sign language video recognition can be controlled using the control device 150. To improve the user experience, a remote control webpage for the user's smartphone can be developed. Then, by using the control device 150 to photograph a webpage in a tablet or PC environment, or a QR marker installed in the real world, it is possible to easily connect to the said webpage.
[0078] If user authorization is performed by scanning a QR code, the system can automatically log in using the ID on the provided PC or tablet. Otherwise, users can log in with their own unique ID / password and connect independently. Furthermore, to prevent duplicate connections, simultaneous connections can be restricted so that if a device is already connected, other devices cannot connect.
[0079] When the capture button on the control device 150 is pressed, the same process as pressing the capture button on a tablet or PC can be performed. Similarly, when the recording button on the control device 150 is pressed, the same process as pressing the recording button on a tablet or PC can be performed. However, the primary device for capturing images is the PC or tablet placed in front of the user, and the primary device for recording audio is the microphone of the user's smartphone. The primary devices for video capture and audio recording can be changed.
[0080] Alternatively, the control device 150 can be implemented using foot buttons. In this case, the start and end points of sign language recognition can be determined using the foot buttons.
[0081] According to one embodiment, the start and end times of sign language recognition can be determined from the recognition of a specific hand shape, independently of the control device 150. For example, sign language recognition can start when a hand is suddenly raised from outside the screen and enters the screen. Sign language recognition can end when a hand is lowered from inside the screen and exits the screen.
[0082] The various embodiments of this disclosure are intended to illustrate representative aspects of the disclosure, rather than listing all possible combinations. The matters described in the various embodiments may be applied independently or in combination of two or more.
[0083] Furthermore, various embodiments of this disclosure can be implemented by hardware, firmware, software, or a combination thereof. In the case of hardware implementation, it can be implemented by one or more ASICs (Application Specific Integrated Circuits), DSPs (Digital Signal Processors), DSPDs (Digital Signal Processing Devices), PLDs (Programmable Logic Devices), FPGAs (Field Programmable Gate Arrays), general processors, controllers, microcontrollers, microprocessors, etc. For example, it is obvious that it can also be implemented in the form of a program stored on a non-temporary computer-readable medium that can be used at the terminal or edge, or in the form of a program stored on a non-temporary computer-readable medium that can be used at the edge or in the cloud.
[0084] For example, the information display method according to one embodiment of this disclosure can be implemented as a program stored in a non-temporary computer-readable medium, and the method of performing phase expansion in directional-based block units as described above can also be implemented as a computer program.
[0085] The scope of this disclosure includes software or machine-executable commands (e.g., operating systems, applications, firmware, programs, etc.) that enable operation by various embodiments to be performed on a device or computer, and non-transitory computer-readable medium on which such software or commands etc. are stored and can be executed on a device or computer.
[0086] As described above, the present invention can be substituted, modified, and altered in various ways without departing from the technical spirit of the invention, and therefore the scope of the present invention is not limited by the embodiments and accompanying drawings described above.
Claims
1. A communication aid device for communication using sign language, A sign language recognition module that extracts sign language text from user movements analyzed from video data, The system includes a display that shows the extracted sign language text, The aforementioned sign language recognition module is The aforementioned video data is divided into multiple segments, The recognition accuracy of each of the aforementioned segments' gross is determined based on the similarity between the segment's gross and a similar gross, wherein the similar gross is the gross that is most similar to the segment's gross. The system is configured to determine the sign language text based on the gross of the multiple segments whose recognition accuracy is greater than a predetermined value. The aforementioned sign language recognition module further, If the gross of the plurality of segments includes a first gross whose recognition accuracy is greater than a predetermined value and a second gross whose recognition accuracy is less than a predetermined value, then a plurality of candidate gross to replace the second gross is determined based on the first gross and the previous dialogue content. A communication assistance device configured to extract sign language text based on a gross candidate selected from the aforementioned plurality of gross candidates and the first gross.
2. The aforementioned communication assistance device is An STT module that converts audio data into text data, The communication assistance device according to claim 1, further comprising a sign language generation module that converts voice data into sign language data.
3. The aforementioned communication assistance device is The system further includes a word card selection module that provides the user-selectable word cards to the display, The communication assistance device according to claim 1, characterized in that the sign language recognition module extracts the sign language sentence based on the selected word card.
4. The aforementioned communication assistance device is The system further includes a text input module that provides a user interface for the user to enter text into the display, The communication assistance device according to claim 1, characterized in that the text input module is activated when the sign language recognition module fails to extract the sign language text.
5. The aforementioned communication assistance device is The communication assistance device further includes a communication module that controls the device to communicate with an external device, The communication assistance device according to claim 1, characterized in that when the sign language recognition module fails to extract the sign language text, the communication module controls the communication assistance device to connect to the external device.
6. The aforementioned sign language recognition module is The communication assistance device according to claim 1, characterized in that it detects the user's joint locations from the video data, extracts skeleton information for tracking the user's movements, and compares the user's gross based on the skeleton information with similar gross.
7. The aforementioned display is The communication assistance device according to claim 1, characterized in that if the recognition accuracy of the gross of the plurality of segments is all less than a predetermined value, a message requesting retransmission of the sign language text is displayed.
8. The aforementioned sign language recognition module is The priority order of the multiple gross candidates is determined based on their similarity to the second gross. The aforementioned display is The communication assistance device according to claim 1, characterized in that it displays the plurality of gross candidates in accordance with the aforementioned priority order.
9. The communication assistance device according to claim 1, characterized in that the display is a transparent display.
Citation Information
Patent Citations
Translation method, mobile terminal and computer readable storage medium
CN109960813A
Sign language translation device
KR101915088B1
Sign language translator, system and method
KR1020160109708A
Apparatus for generating parallel corpus data between text language and sign language and method therefor
KR1020210138311A
Sign language translation apparatus using gloss and translation model learning apparatus
KR102115551B1