Sign language translation device, sign language translation method and program
The sign language translation device addresses grammar differences by converting sign language movements into boilerplate text, enabling seamless communication between deaf and hearing individuals.
Patent Information
- Application Number
- JP2024024584
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-21
- Publication Date
- 2025-09-02
AI Technical Summary
Existing sign language translation devices fail to accurately convey meaning due to differences in grammar between Japanese Sign Language and spoken languages, leading to communication barriers for hearing individuals.
A sign language translation device that includes a video data acquisition unit, image recognition unit, sign language translation unit, and text conversion unit, which converts sign language movements into predetermined boilerplate text for better understanding by non-sign language users.
Facilitates effective communication between deaf and hearing individuals by translating sign language into understandable text.
Smart Images

Figure 2025127714000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sign language translation device, a sign language translation method, and a program. [Background technology]
[0002] Conventionally, sign language has been used for communication among deaf people. However, many hearing people cannot understand sign language, so sign language translation is required for communication between deaf and hearing people. Patent Document 1 discloses a sign language translation device that determines a fixed phrase to be used based on the similarity between a sentence to be translated and the fixed phrase, and synthesizes a sign language CG animation based on motion data associated with the fixed phrase to be used. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-156082 Summary of the Invention [Problem to be solved by the invention]
[0004] As mentioned above, Patent Document 1 discloses a sign language translation device that translates text into sign language, but does not disclose a device that translates sign language into text, etc. Japanese Sign Language has a different grammar from languages such as Japanese, and there is a problem in that the meaning may not be conveyed simply by listing the meanings of the sign language words.
[0005] The present invention has been made to solve the above-mentioned problems, and has as its object to provide a sign language translation device that enables appropriate communication even with people who do not understand sign language. [Means for solving the problem]
[0006] In order to solve the above problem, the present invention is characterized by comprising a video data acquisition means for acquiring video data, an image recognition means for recognizing the hand movements of a sign language user from the video data acquired by the video data acquisition means, a sign language translation means for performing a sign language translation on the recognition result by the image recognition means, and a text conversion means for converting the translation result of the sign language translation means into a predetermined boilerplate text. [Effects of the Invention]
[0007] According to the present invention, it is possible to provide a sign language translation device that allows appropriate communication even with people who do not understand sign language. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing a functional configuration of a sign language translation device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the hardware configuration of the sign language translation device. [Figure 3] 10 is a flowchart showing the processing of the sign language translation device. [Figure 4] 10 is a diagram showing an example of a sign language translation process in a sign language translation unit 12. FIG. [Figure 5] FIG. 10 is a diagram showing an example of a word group table showing the correspondence between word groups and scenario numbers. [Figure 6] FIG. 10 is a diagram showing an example of a fixed phrase text table showing the correspondence between scenario numbers and fixed phrases. DETAILED DESCRIPTION OF THE INVENTION
[0009] A sign language translation device according to the present invention will be described in detail below with reference to the drawings. The embodiments described below are preferred examples of the system according to the present invention, and may include various limitations based on typical hardware and software configurations. However, the technical scope of the present invention is not limited to these aspects unless otherwise specified. Furthermore, the components in the embodiments described below can be appropriately replaced with existing components, and various variations, including combinations with other existing components, are possible. Therefore, the description of the embodiments described below does not limit the content of the invention described in the claims.
[0010] 1 is a block diagram showing the functional configuration of a sign language translation device according to one embodiment of the present invention. Sign language translation device 1 of this embodiment includes a video data acquisition unit 10 that acquires video data captured by a video capture device 2, an image recognition unit 11 that performs image recognition processing on the video data acquired by video data acquisition unit 10, a sign language translation unit 12 that translates the recognition result by image recognition unit 11 into sign language, a text conversion unit 13 that converts the translation result of sign language translation unit 12 into text, and a text output unit 14 that outputs the text converted by text conversion unit 13 to an output device 3.
[0011] The sign language translation device 1 of this embodiment is applicable, for example, to counter operations in stores, government offices, etc., where a store clerk or employee at the counter assists a sign language user who comes to the counter. The imaging device 2 is a camera capable of capturing video, and captures the sign language of the sign language user who comes to the counter. The sign language translation device 1 acquires the video data captured by the imaging device 2, converts it into text in a language that the counter clerk can understand, and outputs it to a text output unit 14, which is, for example, a display or a printing device.
[0012] The image recognition unit 11 performs image recognition processing on the video data acquired by the video data acquisition unit 10 to recognize the sign language user and the movements of the sign language user's hands, etc. This image recognition processing may use any known method. The sign language translation unit 12 performs sign language translation by inferring the words meant by the sign language from the recognition results by the image recognition unit 11, for example, using machine learning. The sign language translation unit 12 outputs a word group, which is a collection of words meant by the sign language, as the translation result.
[0013] However, the output of the sign language translation unit 12, which simply translates sign language, is in sign language grammar and may be difficult for the counter staff to understand. Therefore, in this embodiment, the text conversion unit 13 converts the group of words output by the sign language translation unit 12 into fixed phrase text, which is a fixed phrase used in counter work.
[0014] 2 is a block diagram showing the hardware configuration of the sign language translation device 1. The sign language translation device 1 is an information processing device (computer) operated by a user, such as a PC (personal computer), a PDA (personal digital assistant), or a smartphone.
[0015] The sign language translation device 1 includes a control means 20 that controls the entire sign language translation device 1, an input / output means 22 that receives input from a user who operates the sign language translation device 1 and outputs to the user, a storage means 21 that stores programs executed by the control means 20 and various data, and a communication means 23 that communicates with the image capture device 2 and the output device 3. The communication means 23 may be one that performs wired communication or one that performs wireless communication. The image capture device 2 and the output device 3 may also be included in the sign language translation device 1. Each component shown in FIG. 1 is realized by the control means 20 executing a program stored in the storage means 21. At least a part of the components shown in FIG. 1 may be provided in a housing separate from the sign language translation device 1, such as a server, and data may be exchanged by communication between the sign language translation device 1 and the server.
[0016] 3 is a flowchart showing the processing of the sign language translation device 1. The user turns on the power to the sign language translation device 1, the photographing device 2, and the output device 3 to start the processing. When the photographing device 2 is turned on, it starts photographing the set range. The video data photographed by the photographing device 2 is transmitted to the sign language translation device 1. In step S31, the video data acquisition unit 10 acquires the video data from the photographing device 2.
[0017] In step S32, the image recognition unit 11 performs image recognition processing to recognize the sign language user and the movements of the sign language user's hands, etc., on the video data acquired by the video data acquisition unit 10. In step S33, the sign language translation unit 12 performs sign language translation by inferring the words that the sign language means from the recognition result by the image recognition unit 11.
[0018] 4 is a diagram showing an example of sign language translation processing in the sign language translation unit 12. The sign language translation unit 12 has a trained model 40 that has been trained in advance using image recognition results for sign language and words that the sign language means as training data. The sign language translation unit 12 inputs input data 41, which is the recognition result by the image recognition unit 11, to the trained model 40, and outputs a group of words, which is the inference result of the trained model 40, as an inference result 42.
[0019] 3, in step S34, the text conversion unit 13 converts the output of the sign language translation unit 12, i.e., the word group that is the sign language translation result, into a fixed phrase text. In this embodiment, the text conversion unit 13 converts the word group into a fixed phrase text using the word group table shown in FIG. 5 and the fixed phrase text table shown in FIG. 6.
[0020] Fig. 5 is a diagram showing an example of a word group table showing the correspondence between word groups and scenario numbers. Fig. 6 is a diagram showing an example of a fixed phrase text table showing the correspondence between scenario numbers and fixed phrase texts. The word group table is a table that stores in advance the correspondence between word groups and scenario numbers. The fixed phrase text table is a table that stores in advance the correspondence between scenario numbers and fixed phrase texts.
[0021] The text conversion unit 13 first refers to the word group table and finds the scenario number corresponding to the word group output by the sign language translation unit 12. Next, the text conversion unit 13 refers to the fixed phrase text table and finds the fixed phrase text corresponding to the scenario number found in the word group table. For example, if the word group output by the sign language translation unit 12 is "body," "disability," "person," "certificate," "lost," and "want," the text conversion unit 13 looks up the word group table and finds "C-1302" as the corresponding scenario number. Next, since the scenario number is "C-1302," the text conversion unit 13 looks up the fixed phrase text table and finds "I lost my disability certificate" as the corresponding fixed phrase text.
[0022] In this embodiment, a word group table and a fixed phrase text table are used, but the present invention is not limited to this, and the word group table and the fixed phrase text table may be combined into one table.
[0023] Furthermore, in this embodiment, the relationship between the word groups and the fixed phrase texts in the word group table and the fixed phrase text table is one to one, but the present invention is not limited to this, and may be many to one.
[0024] Furthermore, in this embodiment, a scenario number that includes all of the word group output by the sign language translation unit 12 is selected using the word group table, but the present invention is not limited to this, and a scenario number that includes some of the word group output by the sign language translation unit 12 may be selected. For example, a key word from the word group may be determined in advance, and a corresponding scenario number may be determined when the sign language includes the key word.
[0025] Returning to the explanation of FIG. 3, in step S35, the text output unit 14 outputs the fixed phrase text converted by the text conversion unit 13 to the output device 3, and the processing of the sign language translation apparatus 1 ends.
[0026] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments. The object of the present invention can also be achieved by providing a storage medium storing program code (computer program) that realizes the functions of the above-described embodiments to a system or device, and having a computer in the system or device read and execute the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the above-described embodiments, and the storage medium storing the program code constitutes the present invention. Furthermore, in the above-described embodiments, the functions of each component are realized by a computer executing the program, but some or all of the processing may be implemented using dedicated electronic circuits (hardware). The present invention is not limited to the specific examples described, and various modifications and variations are possible within the spirit and scope of the present invention as defined by the claims. [Explanation of symbols]
[0027] 1. Sign language translation device 2. Imaging equipment 3 Output Devices 10 Video data acquisition unit 11 Image Recognition Unit 12 Sign Language Translation Department 13 Text conversion section 14 Text output section
Claims
1. video data acquisition means for acquiring video data; an image recognition means for recognizing hand movements of a sign language user from the video data acquired by the video data acquisition means; a sign language translation means for executing a sign language translation on the recognition result by the image recognition means; a text conversion means for converting a translation result of the sign language translation means into a predetermined fixed phrase text; A sign language translation device comprising:
2. the sign language translation means takes a group of words that are a set of words that are sign language included in the video data as a translation result, The text conversion means refers to a table that stores in advance the correspondence between the word group and the fixed phrase, and converts the translation result of the sign language translation means into the fixed phrase text.
2. The sign language translation device according to claim 1.
3. a video data acquisition step of acquiring video data; an image recognition step of recognizing hand movements of a sign language user from the video data acquired in the video data acquisition step; a sign language translation step of executing a sign language translation on the recognition result from the image recognition step; a text conversion step of converting the translation result of the sign language translation step into a predetermined fixed phrase text; A sign language translation method comprising:
4. Computer, video data acquisition means for acquiring video data; an image recognition means for recognizing hand movements of a sign language user from the video data acquired by the video data acquisition means; a sign language translation means for executing a sign language translation on the recognition result by the image recognition means; and a text conversion means for converting the translation result of the sign language translation means into a predetermined fixed phrase text; A program characterized by functioning as
Citation Information
Patent Citations
Sign language translation device and program
JP2023156082A