Sign language translation device, sign language translation method, and program
The sign language translation device optimizes processing by initiating image recognition only when sign language actions are identified, reducing unnecessary load and improving efficiency.
Patent Information
- Application Number
- JP2024032833
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-03-05
- Publication Date
- 2025-09-18
AI Technical Summary
Existing sign language translation devices impose unnecessary processing loads by continuously performing image recognition for sign language movements, even when no actions are being made, which is undesirable.
A sign language translation device that includes a video data acquisition means, a start timing determination means to identify when a sign language action starts, an image recognition means to recognize hand movements only when necessary, and a sign language translation means to perform translation based on the recognition results.
Reduces processing load by performing image recognition only when sign language actions are detected, thereby optimizing resource utilization.
Smart Images

Figure 2025135166000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sign language translation device, a sign language translation method, and a program. [Background technology]
[0002] Conventionally, sign language has been used for communication among deaf people. However, many hearing people cannot understand sign language, so sign language translation is required for communication between deaf and hearing people. Patent Document 1 discloses a sign language translation device that determines a fixed phrase to be used based on the similarity between a sentence to be translated and the fixed phrase, and synthesizes a sign language CG animation based on motion data associated with the fixed phrase to be used. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Publication No. 2023-156082 Summary of the Invention [Problem to be solved by the invention]
[0004] As mentioned above, Patent Document 1 discloses a sign language translation device that translates text into sign language, but does not disclose a device that translates sign language into text, etc. To translate sign language, it is necessary to perform image recognition processing to recognize sign language actions on video data of a sign language user.
[0005] In the past, it was unclear when a sign language user would start signing, which meant that image recognition processing to recognize sign language movements had to be performed at all times. However, if image recognition processing to recognize sign language movements were performed even when no sign language movements were being made, this would impose an unnecessary processing load and is not desirable.
[0006] The present invention has been made to solve the above-mentioned problems, and has an object to provide a sign language translation device that can reduce the processing load. [Means for solving the problem]
[0007] In order to solve the above problem, the present invention is characterized by comprising a video data acquisition means for acquiring video data, a start timing determination means for determining whether the video data acquired by the video data acquisition means includes a signal for a sign language user to start a sign language action, an image recognition means for recognizing the hand movements of the sign language user from the video data acquired by the video data acquisition means when the start timing determination means determines that the video data acquired by the video data acquisition means includes the signal, and a sign language translation means for performing a sign language translation on the recognition result by the image recognition means. [Effects of the Invention]
[0008] According to the present invention, it is possible to provide a sign language translation device that can reduce the processing load. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a block diagram showing a functional configuration of a sign language translation device according to an embodiment of the present invention. [Figure 2] FIG. 2 is a block diagram showing the hardware configuration of the sign language translation device. [Figure 3] 10 is a flowchart showing the processing of the sign language translation device. [Figure 4] 10 is a diagram showing an example of a sign language translation process in a sign language translation unit 12. FIG. [Figure 5] FIG. 10 is a diagram showing an example of a word group table showing the correspondence between word groups and scenario numbers. [Figure 6] FIG. 10 is a diagram showing an example of a fixed phrase text table showing the correspondence between scenario numbers and fixed phrases. DETAILED DESCRIPTION OF THE INVENTION
[0010] A sign language translation device according to the present invention will be described in detail below with reference to the drawings. The embodiments described below are preferred examples of the system according to the present invention, and may include various limitations based on typical hardware and software configurations. However, the technical scope of the present invention is not limited to these aspects unless otherwise specified. Furthermore, the components in the embodiments described below can be appropriately replaced with existing components, and various variations, including combinations with other existing components, are possible. Therefore, the description of the embodiments described below does not limit the content of the invention described in the claims.
[0011] 1 is a block diagram showing the functional configuration of a sign language translation device according to one embodiment of the present invention. The sign language translation device 1 of this embodiment includes a video data acquisition unit 10 that acquires video data captured by a camera 2, an image recognition unit 11 that performs image recognition processing on the video data acquired by the video data acquisition unit 10, a sign language translation unit 12 that performs sign language translation on the recognition result by the image recognition unit 11, a text conversion unit 13 that converts the translation result of the sign language translation unit 12 into text, a text output unit 14 that outputs the text converted by the text conversion unit 13 to an output device 3, and a start timing determination unit 15 that superimposes a marker on the video data acquired by the video data acquisition unit 10 and outputs the video data to an output device 4, and determines the timing to start image recognition processing that recognizes sign language actions based on the video data acquired by the video data acquisition unit 10.
[0012] The sign language translation device 1 of this embodiment is applicable, for example, to counter operations in stores, government offices, etc., where a store clerk or employee at the counter assists a sign language user who comes to the counter. The imaging device 2 is a camera capable of capturing video, and captures the sign language user who comes to the counter. The sign language translation device 1 acquires the video data captured by the imaging device 2, converts it into text in a language understandable to the counter clerk, and outputs it to an output device 3, such as a display or a printer. The sign language translation device 1 also superimposes a marker at a predetermined position on the video data captured by the imaging device 2, and outputs it to an output device 4, such as a display.
[0013] The image recognition unit 11 performs image recognition processing on the video data acquired by the video data acquisition unit 10 to recognize the sign language user and the movements of the sign language user's hands, etc. This image recognition processing may use any known method. The sign language translation unit 12 performs sign language translation by inferring the words meant by the sign language from the recognition results by the image recognition unit 11, for example, using machine learning. The sign language translation unit 12 outputs a word group, which is a collection of words meant by the sign language, as the translation result.
[0014] The start timing determination unit 15 superimposes a marker at a predetermined position on the video data captured by the imaging device 2 and outputs the superimposed video to the output device 4. As a result, the output device 4 displays the video with the marker superimposed on the video data captured by the imaging device 2. The sign language translation device 1 according to this embodiment can know that the sign language user will start a sign language action when a signal is given by the sign language user. The sign language user looks at the display on the output device 4 and moves their hand so that it overlaps with the marker displayed on the output device 4, which is a signal to start a sign language action.
[0015] The marker is, for example, a red circle displayed next to the face of the sign language user on the display of the output device 4. It is desirable to display the marker in a position that does not overlap with the sign language user's hand unless the sign language user consciously moves their hand.
[0016] In this embodiment, the cue for a sign language user to start a sign language action is when the marker overlaps with the hand, but the present invention is not limited to this. For example, the cue for a sign language user to start a sign language action may be when one hand of the sign language user overlaps with the marker, or when both hands of the sign language user overlap with the marker. The shape and size of the marker may be any as long as the cue for starting a sign language action can be identified.
[0017] Furthermore, the display of a marker is not essential, and any cue to start a sign language action can be used as long as it is possible to identify the cue to start a sign language action by the action of a sign language user. For example, the cue to start a sign language action may be a specific hand shape made by a sign language user with one or both hands. The specific hand shape may be a hand sign such as a peace sign or an OK sign, or may be a specific sign language, such as a sign for a greeting.
[0018] However, the output of the sign language translation unit 12, which simply translates sign language, is in sign language grammar and may be difficult for the counter staff to understand. Therefore, in this embodiment, the text conversion unit 13 converts the group of words output by the sign language translation unit 12 into fixed phrase text, which is a fixed phrase used in counter work.
[0019] 2 is a block diagram showing the hardware configuration of the sign language translation device 1. The sign language translation device 1 is an information processing device (computer) operated by a user, such as a PC (personal computer), a PDA (personal digital assistant), or a smartphone.
[0020] The sign language translation device 1 includes control means 20 that controls the entire sign language translation device 1, input / output means 22 that accepts input from a user who operates the sign language translation device 1 and outputs to the user, storage means 21 that stores programs executed by the control means 20 and various data, and communication means 23 that communicates with the image capture device 2, output device 3, and output device 4. The communication means 23 may be one that performs wired communication or one that performs wireless communication. Furthermore, the image capture device 2, output device 3, and output device 4 may be included in the sign language translation device 1. Each component shown in FIG. 1 is realized by the control means 20 executing a program stored in the storage means 21. At least a part of the components shown in FIG. 1 may be provided in a housing separate from the sign language translation device 1, such as a server, and data may be exchanged between the sign language translation device 1 and the server through communication.
[0021] 3 is a flowchart showing the processing of the sign language translation device 1. The user turns on the power to the sign language translation device 1, the photographing device 2, the output device 3, and the output device 4 to start the processing. When the photographing device 2 is turned on, it starts photographing the set range. The video data photographed by the photographing device 2 is transmitted to the sign language translation device 1. In step S31, the video data acquisition unit 10 acquires the video data from the photographing device 2.
[0022] In step S32, the start timing determination unit 15 superimposes a marker at a predetermined position on the video data acquired by the video data acquisition unit 10 and outputs the superimposed marker to the output device 4. The sign language user looks at the display on the output device 4 and, when ready to start sign language, gives a signal to start sign language action. In this embodiment, when ready to start sign language, the sign language user moves their hand so that it overlaps with the marker displayed on the output device 4.
[0023] In step S33, the start timing determination unit 15 performs image recognition processing on the video data acquired by the video data acquisition unit 10 to recognize the hand of the sign language user, and determines whether the recognized hand overlaps with the marker superimposed on the video data in step S32. If the start timing determination unit 15 determines that the hand overlaps the marker, i.e., that a signal to start a sign language action has been given, the process proceeds to step S34. If the start timing determination unit 15 determines that the hand does not overlap the marker, i.e., that a signal to start a sign language action has not been given, the process proceeds to step S32. In other words, if the start timing determination unit 15 determines that the video data acquired by the video data acquisition unit 10 includes a signal by the sign language user to start a sign language action, the process proceeds to step S32.
[0024] In step S34, the image recognition unit 11 performs image recognition processing to recognize the sign language user and the movements of the sign language user's hands, etc., on the video data acquired by the video data acquisition unit 10. In step S35, the sign language translation unit 12 performs sign language translation by inferring the words that the sign language means from the recognition result by the image recognition unit 11.
[0025] 4 is a diagram showing an example of sign language translation processing in the sign language translation unit 12. The sign language translation unit 12 has a trained model 40 that has been trained in advance using image recognition results for sign language and words that the sign language means as training data. The sign language translation unit 12 inputs input data 41, which is the recognition result by the image recognition unit 11, to the trained model 40, and outputs a group of words, which is the inference result of the trained model 40, as an inference result 42.
[0026] 3, in step S36, the text conversion unit 13 converts the output of the sign language translation unit 12, i.e., the word group that is the sign language translation result, into a fixed phrase text. In this embodiment, the text conversion unit 13 converts the word group into a fixed phrase text using the word group table shown in FIG. 5 and the fixed phrase text table shown in FIG. 6.
[0027] Fig. 5 is a diagram showing an example of a word group table showing the correspondence between word groups and scenario numbers. Fig. 6 is a diagram showing an example of a fixed phrase text table showing the correspondence between scenario numbers and fixed phrase texts. The word group table is a table that stores in advance the correspondence between word groups and scenario numbers. The fixed phrase text table is a table that stores in advance the correspondence between scenario numbers and fixed phrase texts.
[0028] The text conversion unit 13 first refers to the word group table and finds the scenario number corresponding to the word group output by the sign language translation unit 12. Next, the text conversion unit 13 refers to the fixed phrase text table and finds the fixed phrase text corresponding to the scenario number found in the word group table. For example, if the word group output by the sign language translation unit 12 is "body," "disability," "person," "certificate," "lost," and "want," the text conversion unit 13 looks up the word group table and finds "C-1302" as the corresponding scenario number. Next, since the scenario number is "C-1302," the text conversion unit 13 looks up the fixed phrase text table and finds "I lost my disability certificate" as the corresponding fixed phrase text.
[0029] In this embodiment, a word group table and a fixed phrase text table are used, but the present invention is not limited to this, and the word group table and the fixed phrase text table may be combined into one table.
[0030] Furthermore, in this embodiment, the relationship between the word groups and the fixed phrase texts in the word group table and the fixed phrase text table is one to one, but the present invention is not limited to this, and may be many to one.
[0031] Furthermore, in this embodiment, a scenario number that includes all of the word group output by the sign language translation unit 12 is selected using the word group table, but the present invention is not limited to this, and a scenario number that includes some of the word group output by the sign language translation unit 12 may be selected. For example, a key word from the word group may be determined in advance, and a corresponding scenario number may be determined when the sign language includes the key word.
[0032] Returning to the explanation of FIG. 3, in step S37, the text output unit 14 outputs the fixed phrase text converted by the text conversion unit 13 to the output device 3, and the processing of the sign language translation apparatus 1 ends.
[0033] Although preferred embodiments of the present invention have been described above, the present invention is not limited to these embodiments. The object of the present invention can also be achieved by providing a storage medium storing program code (computer program) that realizes the functions of the above-described embodiments to a system or device, and having a computer in the system or device read and execute the program code stored in the storage medium. In this case, the program code read from the storage medium itself realizes the functions of the above-described embodiments, and the storage medium storing the program code constitutes the present invention. Furthermore, in the above-described embodiments, the functions of each component are realized by a computer executing the program, but some or all of the processing may be implemented using dedicated electronic circuits (hardware). The present invention is not limited to the specific examples described, and various modifications and variations are possible within the spirit and scope of the present invention as defined by the claims. [Explanation of symbols]
[0034] 1. Sign language translation device 2. Imaging equipment 3 Output Devices 10 Video data acquisition unit 11 Image Recognition Unit 12 Sign Language Translation Department 13 Text conversion section 14 Text output section 15 Start timing determination section
Claims
1. video data acquisition means for acquiring video data; start timing determination means for determining whether the video data acquired by the video data acquisition means includes a cue from a sign language user to start a sign language action; an image recognition means for recognizing hand movements of a sign language user from the video data acquired by the video data acquisition means when the start timing determination means determines that the video data acquired by the video data acquisition means includes the signal; and a sign language translation means for executing a sign language translation on the recognition result by the image recognition means; A sign language translation device comprising:
2. The signal is given by the sign language user placing his / her hand on a marker superimposed on the video data acquired by the video data acquisition means.
2. The sign language translation device according to claim 1.
3. a video data acquisition step of acquiring video data; a start timing determination step of determining whether the video data acquired in the video data acquisition step includes a cue from a sign language user to start a sign language action; an image recognition step of recognizing hand movements of a sign language user from the video data acquired in the video data acquisition step when it is determined in the start timing determination step that the video data acquired in the video data acquisition step includes the signal; a sign language translation step of executing a sign language translation on the recognition result in the image recognition step; A sign language translation method comprising:
4. Computer, video data acquisition means for acquiring video data; start timing determination means for determining whether the video data acquired by the video data acquisition means includes a cue from a sign language user to start a sign language action; an image recognition means for recognizing hand movements of a sign language user from the video data acquired by the video data acquisition means when the start timing determination means determines that the video data acquired by the video data acquisition means includes the signal; and A sign language translation means for executing a sign language translation on the recognition result by said image recognition means. A program characterized by functioning as
Citation Information
Patent Citations
Sign language translation device and program
JP2023156082A