Sign language translation device and program

The sign language translation device addresses inaccuracies and unnaturalness in conventional systems by using fixed phrases and motion data for phrases or sentences, ensuring coherent and meaningful sign language animation.

JP7807169B2Active Publication Date: 2026-01-27NIPPON HOSO KYOKAI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2022065731
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-04-12
Publication Date
2026-01-27
Estimated Expiration
2042-04-12

AI Technical Summary

Technical Problem

Conventional real-time sign language translation technology generates inaccurate and unnatural CG animation due to the lack of motion data for certain words, resulting in interrupted translations and loss of meaning, with insufficient consideration for facial expressions and mouth shapes.

Method used

A sign language translation device that stores fixed phrases and motion data, determines the most similar fixed phrase for a given sentence, and synthesizes CG animation using motion data corresponding to phrases or sentences, incorporating facial expressions and mouth shapes.

Benefits of technology

Generates natural sign language CG animation that preserves the meaning of the original spoken language by reducing interruptions and improving the coherence of hand and facial movements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007807169000001
    Figure 0007807169000001
  • Figure 0007807169000002
    Figure 0007807169000002
  • Figure 0007807169000003
    Figure 0007807169000003
Patent Text Reader

Abstract

To allow a natural sign language CG animation to be generated while keeping minimal meaning expressed in an original spoken language.SOLUTION: A sign language translation device is provided, comprising a storage unit configured to store fixed phrases in association with motion data, a fixed phrase determination unit configured to determine a used fixed phrase on the basis of similarity between a translation target phrase and fixed phrases, and a video synthesis unit configured to synthesize a sign language CG animation according to motion data associated with the used fixed phrase.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a sign language translation device and a program. [Background technology]

[0002] Real-time sign language translation technology that uses computer graphics (CG) animation is currently in use. Conventional real-time sign language translation technology first translates spoken language converted into text using speech recognition or other methods into a string of sign language labels for generating sign language CG animation. Next, motion data corresponding to each word in the sign language label string is read and interpolated to synthesize motion data for each sentence. Finally, the motion data for each sentence is played back on a sign language CG model to generate sign language CG animation.

[0003] For example, Patent Document 1 discloses an invention for real-time translation from a spoken language such as Japanese to a sign language. The invention described in Patent Document 1 generates sign language CG animation using machine translation that uses a corpus of bilingual data between Japanese and sign language. [Prior art documents] [Patent documents]

[0004] [Patent Document 1] Japanese Patent Application Laid-Open No. 2014-21180 Summary of the Invention [Problem to be solved by the invention]

[0005] However, conventional real-time sign language translation technology has the problem of generating inaccurate and unnatural sign language CG animation. For example, if a sign language label string translated from a spoken language contains a word for which no motion data exists, the sign language CG animation will be interrupted. Also, for example, an unnatural feeling may appear where motion data has been interpolated. Furthermore, the translation may output sign language labels that express content different from the meaning of the original spoken language, resulting in a significant change in the meaning of the sentence.

[0006] In view of the above technical problems, one aspect of the present invention aims to generate natural sign language CG animation while preserving the meaning of the original spoken language to a minimum extent. [Means for solving the problem]

[0007] In order to solve the above problem, a sign language translation device of one embodiment of the present invention includes a memory unit that stores fixed phrases and motion data in association with each other, a fixed phrase determination unit that determines the fixed phrase to be used based on the similarity between the sentence to be translated and the fixed phrase, and an image synthesis unit that synthesizes sign language CG animation based on the motion data associated with the fixed phrase to be used. [Effects of the Invention]

[0008] According to one aspect of the present invention, natural sign language CG animation can be generated while preserving the meaning of the original spoken language to a minimum extent. [Brief explanation of the drawings]

[0009] [Figure 1] 1 is a block diagram showing an example of the overall configuration of a sign language translation system. [Figure 2] FIG. 2 is a block diagram illustrating an example of a hardware configuration of a computer. [Figure 3] FIG. 1 is a diagram illustrating an example of a functional configuration of a sign language translation system. [Figure 4] 10 is a flowchart illustrating an example of a processing procedure of a sign language translation method. [Figure 5] FIG. 1 is a diagram for explaining a specific example of a conventional technique. [Figure 6] FIG. 1 is a diagram for explaining a specific example of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0010] Hereinafter, embodiments of the present invention will be described with reference to the accompanying drawings. In this specification and drawings, components having substantially the same functional configuration are designated by the same reference numerals, and redundant description will be omitted.

[0011] [overview] Real-time sign language translation technology using CG animation has the potential to be widely used in a variety of fields. Hereafter, CG animation generated by real-time sign language translation technology will be referred to as "sign language CG animation."

[0012] Conventional real-time sign language translation technology typically uses a translation system that translates spoken language into a sequence of sign language labels, which are used to generate sign language CG animations.

[0013] In the real-time sign language translation technology described above, the speech signal to be translated is first converted into text in a spoken language such as Japanese using speech recognition or other methods. Next, the converted text is input into a pre-trained translation system to generate a sign language label string. After that, motion data corresponding to each word in the sign language label string is read, and the motion data synthesized into sentence units by interpolating between each word is played on a sign language CG model that represents the sign language speaker.

[0014] However, there are problems with improving the accuracy of translation from any spoken language sentence to a sequence of sign language labels. Even if accurate translation from spoken language to sign language is possible, there is a problem that the sign language CG animation will be interrupted if there is no motion data corresponding to the translation result.

[0015] For example, in conventional technology, there is no large-scale sign language corpus that can be used as training data for translation systems. As a result, the accuracy of the translation system does not improve, and translation results that differ in meaning from the original spoken text are often output. Even if the translation is accurate, there may be unnaturalness in the parts where word motion data is interpolated and connected. Furthermore, if there is no motion data corresponding to each sign language label in the translation result, the sign language CG animation will be interrupted in that part.

[0016] Furthermore, conventional sign language labels often do not include information that is important in sign language, such as facial expressions and mouth shapes. This is because, when building a corpus, only hand and finger movements below the neck are transcribed as sign language labels for sign language videos. As a result, facial information is missing from the sign language label strings resulting from translation, and facial expressions and mouth shapes that should change depending on the context cannot be reproduced in sign language CG animations.

[0017] In order to solve the above technical problems, one aspect of the present invention prepares a database of multiple patterns of fixed phrases, determines the similarity between an input arbitrary sentence in spoken language and each fixed phrase, and generates a sign language CG animation using the fixed phrase with the highest similarity.

[0018] One aspect of the present invention is not a literal translation, so the resulting sign language may contain less information than the original spoken language. However, one aspect of the present invention is that motion data corresponds to standard phrases or sentences, so the sign language CG animation is continuous. Also, fewer connections between words result in smoother movements.

[0019] Furthermore, in one aspect of the present invention, since motion data corresponds to phrases or sentences, facial expressions and mouth shapes according to the context can be included in the motion data, making it possible to generate sign language CG animation in which facial expressions, mouth shapes, and hand movements match.

[0020] Therefore, according to one aspect of the present invention, it is possible to provide a user with information that preserves the minimum content of the original text using sign language that is more natural than conventional sign language.

[0021] [Embodiment] <Overall configuration of the sign language translation system> First, the overall configuration of the sign language translation system in this embodiment will be described with reference to Fig. 1. Fig. 1 is a block diagram showing an example of the overall configuration of the sign language translation system in this embodiment.

[0022] 1, a sign language translation system 1 in this embodiment includes a sign language translation device 10 and a user terminal 30. The sign language translation device 10 and the user terminal 30 are connected to each other so as to be able to communicate data with each other via a communication network N1 such as a LAN (Local Area Network) or the Internet.

[0023] The sign language translation device 10 is an information processing device such as a PC (Personal Computer), workstation, or server that generates sign language CG animation in response to a request from a user terminal 30. The sign language translation device 10 receives text representing the content of the spoken language to be translated (hereinafter also referred to as "sentence to be translated") from the user terminal 30. The sign language translation device 10 also transmits the sign language CG animation generated based on the sentence to be translated to the user terminal 30.

[0024] The user terminal 30 is an information processing terminal such as a PC, tablet terminal, or smartphone operated by a user. In response to a user operation, the user terminal 30 transmits a sentence to be translated to the sign language translation device 10. The user terminal 30 also outputs the sign language CG animation received from the sign language translation device 10 to a display device or the like.

[0025] 1 is just one example, and various system configurations are possible depending on the application and purpose. For example, the sign language translation device 10 may be realized by multiple computers, or may be realized as a cloud computing service. Furthermore, for example, the sign language translation system 1 may be realized by a standalone information processing device that combines the functions that the sign language translation device 10 and the user terminal 30 should each have.

[0026] An example of a sign language translation device 10 realized by a stand-alone information processing device will be described below.

[0027] <Hardware configuration of sign language translation system> Next, the hardware configuration of the sign language translation system 1 in this embodiment will be described with reference to FIG.

[0028] <Computer hardware configuration> The sign language translation device 10 and the user terminal 30 in this embodiment are realized by, for example, a computer. Fig. 2 is a block diagram showing an example of the hardware configuration of a computer in this embodiment.

[0029] 2, a computer 500 in this embodiment includes a CPU (Central Processing Unit) 501, a ROM (Read Only Memory) 502, a RAM (Random Access Memory) 503, an HDD (Hard Disk Drive) 504, an input device 505, a display device 506, a communication I / F (Interface) 507, and an external I / F 508. The CPU 501, the ROM 502, and the RAM 503 form a so-called computer. The hardware components of the computer 500 are connected to each other via a bus line 509. The input device 505 and the display device 506 may be connected to the external I / F 508 for use.

[0030] The CPU 501 is a computing device that controls the entire computer 500 and realizes its functions by reading programs and data from a storage device such as the ROM 502 or HDD 504 onto the RAM 503 and executing the processes.

[0031] The ROM 502 is an example of a non-volatile semiconductor memory (storage device) that can retain programs and data even when the power is turned off. The ROM 502 functions as a main storage device that stores various programs, data, etc. required for the CPU 501 to execute various programs installed in the HDD 504. Specifically, the ROM 502 stores boot programs such as a Basic Input / Output System (BIOS) and an Extensible Firmware Interface (EFI) that are executed when the computer 500 starts up, as well as data such as OS (Operating System) settings and network settings.

[0032] The RAM 503 is an example of a volatile semiconductor memory (storage device) in which programs and data are erased when the power is turned off. The RAM 503 is, for example, a dynamic random access memory (DRAM) or a static random access memory (SRAM). The RAM 503 provides a working area in which various programs installed in the HDD 504 are expanded when executed by the CPU 501.

[0033] The HDD 504 is an example of a non-volatile storage device that stores programs and data. The programs and data stored in the HDD 504 include an OS, which is basic software that controls the entire computer 500, and applications that provide various functions on the OS. Note that the computer 500 may use a storage device that uses flash memory as a storage medium (e.g., an SSD (Solid State Drive)) instead of the HDD 504.

[0034] The input device 505 includes a touch panel, operation keys and buttons, a keyboard and mouse, a microphone for inputting sound data such as voice, and the like, which are used by the user to input various signals.

[0035] The display device 506 is composed of a display such as a liquid crystal display or organic EL (Electro-Luminescence) display for displaying a screen, a speaker for outputting sound data such as voice, and the like.

[0036] The communication I / F 507 is an interface that connects to a communication network and enables the computer 500 to perform data communication.

[0037] The external I / F 508 is an interface with external devices, such as a drive device 510.

[0038] The drive device 510 is a device for loading a recording medium 511. The recording medium 511 here includes media that record information optically, electrically, or magnetically, such as a CD-ROM, a flexible disk, or a magneto-optical disk. The recording medium 511 may also include semiconductor memories that record information electrically, such as ROMs and flash memories. This allows the computer 500 to read from and / or write to the recording medium 511 via the external I / F 508.

[0039] The various programs to be installed in the HDD 504 are installed, for example, by setting the distributed recording medium 511 in a drive device 510 connected to the external I / F 508 and reading out the various programs recorded on the recording medium 511 by the drive device 510. Alternatively, the various programs to be installed in the HDD 504 may be installed by being downloaded via the communication I / F 507 from a network different from the communication network.

[0040] <Functional configuration of the sign language translation system> Next, the functional configuration of the sign language translation system in this embodiment will be described with reference to Fig. 3. Fig. 3 is a block diagram showing an example of the functional configuration of the sign language translation system 1 in this embodiment.

[0041] <Functional configuration of the sign language translation device> As shown in FIG. 3, the sign language translation device 10 in this embodiment receives as input an audio signal that records the speech of a sentence to be translated, and outputs a sign language CG animation that reproduces the content of the sentence to be translated in sign language.

[0042] The sign language translation device 10 in this embodiment includes a voice input unit 11, a voice recognition unit 12, a sign language translation unit 13, a motion synthesis unit 14, a video synthesis unit 15, and an animation output unit 16. The sign language translation unit 13 includes a fixed phrase storage unit 130, a morphological analysis unit 131, a proper noun extraction unit 132, a fixed phrase determination unit 133, and a proper noun insertion unit 134. The motion synthesis unit 14 includes a motion data storage unit 140, a motion data acquisition unit 141, and a sentence motion synthesis unit 142. The video synthesis unit 15 includes a sign language model storage unit 150 and an animation generation unit 151.

[0043] The voice input unit 11 is realized by the processing that the CPU 501 and the input device 505 execute in response to a program loaded onto the RAM 503 from the HDD 504 shown in FIG.

[0044] The speech recognition unit 12, the sign language translation unit 13, the motion synthesis unit 14, and the video synthesis unit 15 are realized by the processing that the CPU 501 executes by the program loaded from the HDD 504 onto the RAM 503 shown in FIG.

[0045] The animation output unit 16 is realized by the processing that the CPU 501 and the display device 506 execute in response to a program loaded onto the RAM 503 from the HDD 504 shown in FIG.

[0046] The fixed phrase storage unit 130, the motion data storage unit 140, and the sign language model storage unit 150 are realized using, for example, the HDD 504 shown in FIG.

[0047] (Audio input section) The voice input unit 11 receives an input of a voice signal that is a recording of a human speech. The voice input unit 11 sends the received voice signal to the voice recognition unit 12.

[0048] (Voice recognition section) The speech recognition unit 12 performs speech recognition on the speech signal received from the speech input unit 11. The speech recognition unit 12 sends the speech recognition result to the sign language translation unit 13 as a sentence to be translated.

[0049] (Sign Language Translation Department) The sign language translation unit 13 translates the translation target sentence received from the speech recognition unit 12 into a sign language sentence. The sign language sentence is an arrangement in which information representing motion data is arranged in chronological order. The sign language translation unit 13 sends the sign language sentence to the motion synthesis unit 14.

[0050] A plurality of fixed phrases and sign language fixed phrases are stored in advance in the fixed phrase storage unit 130 in association with each other. Fixed phrases are sentences written in natural language such as Japanese. Variables for inserting proper nouns may be embedded in the fixed phrases.

[0051] A fixed sign language phrase is information in which the file names of motion data are arranged in chronological order to represent the content of the phrase. If a fixed phrase includes variables, the fixed sign language phrase also includes the position information and attribute information of the variables.

[0052] The motion data is obtained by capturing actual sign language movements, including hand and facial expressions, and saving them in a format such as BVH (Biovision Hierarchy). The motion data is recorded in units of words, phrases, or sentences.

[0053] In a sign language fixed phrase, one piece of motion data is defined for each phrase when the fixed phrase is divided at the position of a variable. For example, if the fixed phrase does not contain a variable (or if the fixed phrase contains one variable at the beginning or end), one piece of motion data is defined for the sign language fixed phrase. Also, for example, if the fixed phrase contains one variable other than at the beginning or end, two pieces of motion data are defined for the sign language fixed phrase: one piece of motion data corresponding to the phrase before the variable, and one piece of motion data corresponding to the phrase after the variable.

[0054] The morphological analysis unit 131 performs morphological analysis on the sentence to be translated received from the speech recognition unit 12. The morphological analysis unit 131 sends the result of the morphological analysis to the proper noun extraction unit 132 and the fixed phrase determination unit 133.

[0055] The proper noun extraction unit 132 extracts proper nouns from the morphological analysis result received from the morphological analysis unit 131. The proper noun extraction unit 132 sends the extracted proper nouns to the proper noun insertion unit .

[0056] The fixed phrase determination unit 133 determines a fixed phrase to be used for synthesizing sign language CG animation (hereinafter also referred to as a "fixed phrase to be used") from the fixed phrases stored in the fixed phrase storage unit 130, based on the morphological analysis result received from the morphological analysis unit 131. The fixed phrase determination unit 133 sends the fixed phrase to be used and a sign language fixed phrase corresponding to the fixed phrase to be used to the proper noun insertion unit 134.

[0057] The proper noun insertion unit 134 inserts a sign language label corresponding to the proper noun received from the proper noun extraction unit 132 into the sign language fixed phrase received from the fixed phrase determination unit 133. The proper noun insertion unit 134 sends the sign language fixed phrase into which the proper noun has been inserted to the motion synthesis unit 14 as a sign language sentence.

[0058] (Motion synthesis section) The motion synthesis unit 14 synthesizes motion data representing a sentence (hereinafter also referred to as "sentence motion data") based on the sign language sentence received from the sign language translation unit 13. The motion synthesis unit 14 sends the sentence motion data to the video synthesis unit 15.

[0059] A plurality of motion data are pre-stored in the motion data storage unit 140. The motion data includes motion data corresponding to words representing proper nouns, motion data corresponding to phrases of fixed phrases in which variables are embedded, and motion data corresponding to sentences of fixed phrases in which variables are not embedded.

[0060] The motion data corresponding to the fixed phrase or sentence can be motion-captured data of a set of facial expressions and mouth shapes expressed simultaneously with hand and finger movements.

[0061] In conventional technology, sign language CG animation is synthesized by connecting motion data corresponding to words. As a result, the motion data is given a neutral facial expression. By using motion data corresponding to phrases or sentences, it is possible to reproduce sign language expressions that are more faithful to the whole-body movements of actual sign language speakers.

[0062] Furthermore, if motion data corresponding to phrases or sentences is used, since the majority of the sentence is made up of phrases, there are fewer places where the motion data needs to be interpolated and connected compared to connecting multiple words arranged in units smaller than a phrase, making it possible to synthesize more natural sign language CG animation.

[0063] The motion data acquisition unit 141 acquires motion data from the motion data storage unit 140 based on the sign language sentence received from the sign language translation unit 13. The motion data acquisition unit 141 sends the motion data to the sentence motion synthesis unit 142.

[0064] The sentence motion synthesis unit 142 generates sentence motion data by chronologically connecting the motion data received from the motion data acquisition unit 141. The sentence motion synthesis unit 142 sends the sentence motion data to the video synthesis unit 15.

[0065] (Video synthesis section) The video synthesis unit 15 generates sign language CG animation based on the sentence motion data received from the motion synthesis unit 14. The video synthesis unit 15 sends the sign language CG animation to the animation output unit 16.

[0066] The sign language model storage unit 150 stores in advance a sign language CG model representing a sign language speaker. The sign language CG model is human data modeled using CG, and has a skeletal structure corresponding to the joint information defined in the motion data. There are no restrictions on the gender, age, or race of the sign language CG model. There are also no restrictions on the method of expression, such as two-dimensional, three-dimensional, or drawing style.

[0067] The animation generation unit 151 combines the sentence motion data received from the motion synthesis unit 14 with the sign language CG model stored in the sign language model storage unit 150 to render sign language CG animation. The animation generation unit 151 sends the sign language CG animation to the animation output unit 16.

[0068] (Animation output section) The animation output unit 16 outputs the sign language CG animation received from the video synthesis unit 15 to the sign language translation device 10.

[0069] <Functional configuration of user terminal 30> When the sign language translation system 1 has a client-server configuration including the sign language translation device 10 and the user terminal 30, the user terminal 30 only needs to include the voice input unit 11, the voice recognition unit 12, and the animation output unit 16.

[0070] <Processing procedure for sign language translation system> Next, the processing steps of the sign language translation method executed by the sign language translation system 1 in this embodiment will be described with reference to Fig. 4. Fig. 4 is a flowchart showing an example of the processing steps of the sign language translation method in this embodiment.

[0071] In step S1, the speech input unit 11 included in the sign language translation device 10 accepts input of a speech signal. The speech signal is an acoustic signal recorded by a microphone or the like when a person speaks a sentence to be translated. Next, the speech input unit 11 sends the accepted speech signal to the speech recognition unit 12.

[0072] In step S2, the speech recognition unit 12 included in the sign language translation device 10 receives a speech signal from the speech input unit 11. Next, the speech recognition unit 12 recognizes the received speech signal using a predetermined speech recognition method. Any speech recognition technology can be used as the speech recognition method. Next, the speech recognition unit 12 sends the speech recognition result to the sign language translation unit 13 as a sentence to be translated.

[0073] In step S3, the sign language translation unit 13 included in the sign language translation device 10 receives the sentence to be translated from the speech recognition unit 12. Next, the sign language translation unit 13 sends the received sentence to be translated to the morphological analysis unit 131.

[0074] The morphological analysis unit 131 receives a sentence to be translated from the sign language translation unit 13. Next, the morphological analysis unit 131 performs morphological analysis on the received sentence to be translated using a predetermined morphological analysis method. Any morphological analysis technology can be used as the morphological analysis method. The morphological analysis result assigns a part of speech to each morpheme (word) included in the sentence to be translated. If the morpheme is a proper noun, attribute information of the proper noun is assigned. The attribute information of the proper noun is, for example, a person's name, a country name, etc. Next, the morphological analysis unit 131 sends the result of the morphological analysis to the proper noun extraction unit 132 and the fixed phrase determination unit 133.

[0075] In step S4, the proper noun extraction unit 132 receives the morphological analysis result from the morphological analysis unit 131. Next, the proper noun extraction unit 132 extracts proper nouns from the received morphological analysis result. Subsequently, the proper noun extraction unit 132 pairs the extracted proper nouns with attribute information of the proper nouns and sends them to the proper noun insertion unit 134.

[0076] If the morphological analysis results contain multiple proper nouns with the same attribute information, they can be numbered in order to distinguish them. For example, if the morphological analysis results contain two proper nouns that are personal names, they can be labeled as "Personal Name 1," "Personal Name 2," etc.

[0077] In step S5, the fixed phrase determination unit 133 receives the morphological analysis result from the morphological analysis unit 131. Next, the fixed phrase determination unit 133 acquires a fixed phrase from the fixed phrase storage unit .

[0078] Next, if the morphological analysis result includes a proper noun, the fixed phrase determination unit 133 replaces the proper noun with a variable. Next, the fixed phrase determination unit 133 converts the morphological analysis result in which the proper noun has been replaced and each fixed phrase into a vector. For example, TF-IDF (Term Frequency - Inverse Document Frequency) can be used as a method for converting into a vector. Next, the fixed phrase determination unit 133 calculates the similarity between the vector of the morphological analysis result and each of the fixed phrase vectors. The calculated similarity is, for example, cosine similarity.

[0079] Next, the fixed phrase determination unit 133 determines the fixed phrase with the largest cosine similarity (in other words, the fixed phrase most similar to the morphological analysis result) as the fixed phrase to be used. Subsequently, the fixed phrase determination unit 133 acquires a sign language fixed phrase associated with the fixed phrase to be used from the fixed phrase storage unit 130. Then, the fixed phrase determination unit 133 sends the fixed phrase to be used and the sign language fixed phrase corresponding to the fixed phrase to be used to the proper noun insertion unit 134.

[0080] In step S6, the proper noun insertion unit 134 receives the sign language fixed phrase from the fixed phrase determination unit 133. The proper noun insertion unit 134 also receives the proper nouns and attribute information from the proper noun extraction unit 132.

[0081] Next, if a variable is embedded in the sign language fixed phrase received from the fixed phrase determination unit 133, the proper noun insertion unit 134 inserts the proper noun received from the proper noun extraction unit 132 into each variable of the sign language fixed phrase. At this time, the proper noun insertion unit 134 inserts a proper noun with matching attribute information into each variable. The proper noun insertion unit 134 sends the sign language fixed phrase into which the proper noun has been inserted to the motion synthesis unit 14 as a sign language sentence.

[0082] If a sign language template contains multiple variables with the same attribute information, the proper nouns are inserted in the order numbered by the proper noun extraction unit 132. For example, if two variables for personal names are embedded in a sign language template, a proper noun labeled "Personal Name 1" is inserted into the first variable, and a proper noun labeled "Personal Name 2" is inserted into the second variable.

[0083] In step S7, the motion synthesis unit 14 included in the sign language translation device 10 receives the sign language sentence from the sign language translation unit 13. Next, the motion synthesis unit 14 sends the received sign language sentence to the motion data acquisition unit 141.

[0084] The motion data acquisition unit 141 receives the sign language sentence from the motion synthesis unit 14. Next, the motion data acquisition unit 141 acquires motion data from the motion data storage unit 140 based on the file name included in the received sign language sentence. Then, the motion data acquisition unit 141 sends the motion data to the sentence motion synthesis unit 142.

[0085] In step S8, the sentence motion synthesis unit 142 receives motion data from the motion data acquisition unit 141. Next, the sentence motion synthesis unit 142 aligns the received motion data according to the time series expressed in the sign language sentence. Next, the sentence motion synthesis unit 142 interpolates and connects the motion data together. This generates sentence motion data. Then, the sentence motion synthesis unit 142 sends the sentence motion data to the video synthesis unit 15.

[0086] In step S9, the video synthesis unit 15 receives the sentence motion data from the motion synthesis unit 14. Next, the video synthesis unit 15 sends the received sentence motion data to the animation generation unit 151.

[0087] The animation generation unit 151 receives the sentence motion data from the video composition unit 15. Next, the animation generation unit 151 combines the received sentence motion data with the sign language CG model stored in the sign language model storage unit 150 to render a sign language CG animation. Next, the animation generation unit 151 sends the sign language CG animation to the animation output unit 16.

[0088] The animation generation unit 151 may previously read a sign language CG model from the sign language model storage unit 150 and keep it in a standby state. By configuring in this way, the animation generation unit 151 can reduce delays in generating sign language CG animation.

[0089] In step S10, the animation output unit 16 included in the sign language translation device 10 receives the sign language CG animation from the video synthesis unit 15. Next, the animation output unit 16 outputs the received sign language CG animation to the user.

[0090] For example, animation output unit 16 displays the sign language CG animation on display device 506 or the like. Also, for example, animation output unit 16 may use communication I / F 507 to transmit the sign language CG animation to an external device.

[0091] <Example> The specific processing contents of the sign language translation system in this embodiment will be described while comparing it with the conventional technology. Fig. 5 is a conceptual diagram showing the processing contents in the conventional sign language translation system. Fig. 6 is a conceptual diagram showing the processing contents in the sign language translation system of this embodiment.

[0092] 5 and 6 show an example in which live commentary from a sports broadcast is used as input and sign language CG animation is synthesized in real time from the speech recognition results.

[0093] As shown in Figure 5, suppose that during a live basketball broadcast, a commentator says, "Japan's Yamada gets the rebound himself and scores!" In this case, the speech recognition result is "Japan's Yamada gets the rebound himself and scores!"

[0094] In the conventional technology, a sign language label string is generated by converting each word into a sign language label after morphological analysis of the speech recognition results. In this example, the result is assumed to be "Japan, Yamada, rebound, after, myself, hold, carry, decide."

[0095] Then, motion data representing each word is obtained based on the sign language label string of the translation result.The motion data are then interpolated and connected to generate a sign language CG animation as shown in Figure 5.

[0096] As shown in Figure 6, the live commentary and the speech recognition results are the same as those of the prior art. In this embodiment, the phrase most similar to the speech recognition results is used from among the prepared phrases. In this example, it is assumed that the phrase "Shoot, goal from player [name] from [country name]!" is selected. Note that words enclosed in [ ] represent variables. Words enclosed in [ ] represent attribute information of the variables.

[0097] Then, motion data representing each word and phrase is obtained based on the sign language template corresponding to the template used. Next, each motion data is interpolated and connected to generate a sign language CG animation as shown in Figure 6.

[0098] As mentioned above, compared to translation between spoken languages ​​(e.g., Japanese and English), translation from spoken language to sign language, which is a visual language, and generation of sign language CG animation based on the translation results can raise issues that do not arise when translating between spoken languages.

[0099] First, there is the issue of difficulty in improving translation accuracy when translating any spoken language sentence into a sequence of sign language labels. In other words, there can be significant differences in meaning between the input spoken language text and the translated sign language label sequence. One reason for this is that sign language is a minority language, and there are no large-scale language corpora that can serve as sufficient training data.

[0100] In the examples shown in Figures 5 and 6, the commentary corresponds to a spoken language meaning such as "The player dribbled the ball, shot, and scored." In conventional technology, the sign language label string of the translation result simply replaces words with sign language expressions. As a result, the translation result loses the nuances of "ball," "dribble," and "score" that were semantically contained in the original text. As a result, the meaning of the sign language expressions in the final generated sign language CG animation also changes significantly.

[0101] Another issue with conventional technology is that the motion data corresponding to sign language labels such as "Japan" or "decide" only represents hand movements, and does not include information such as facial expressions or mouth shape. In sign language, facial information that co-occurs with hand movements is also an important element that carries grammatical meaning. With conventional technology, facial information is not added to the translation results, and the facial movements during motion capture are used as is.

[0102] Even if the translation result of a sign language label string for hand movements can be output with perfect accuracy, if the facial movements in the motion data do not match the facial expressions or mouth shapes in the context of the input spoken text, the meaning of the generated sentence will be different.Furthermore, even if the translation result of a sign language label string can be output with perfect accuracy, if there is insufficient motion data corresponding to each sign language label that makes up the sign language label string, the motion will be missing where the corresponding sign language is expressed, as in "rebound" in Figure 5, and the generated sentence will be cut off midway.

[0103] In addition, when connecting multiple motion data loaded according to sign language labels, interpolation is performed based on the previous and following motion data, making it difficult to completely eliminate unnaturalness. With conventional technology, the number of connection points increases, making the unnaturalness worse.

[0104] As shown in the sign language label string in Figure 5, conventional technology requires interpolation between all of the words "Japan," "Yamada," "rebound," "after," "myself," "have," etc., in order to connect multiple sign language words, which are the smallest units. As a result, the natural rhythm and pauses that ensure the continuity of a single, coherent sign language sentence are lost.

[0105] On the other hand, the sign language translation system of this embodiment can solve the above-mentioned problem by performing the following operations. First, the speech recognition unit 12 outputs a speech recognition result such as "Yamada from Japan, take it from the rebound yourself and score!" Next, the morphological analysis unit 131 generates a morphological analysis result such as "Japan," "of," "Yamada," "rebound," etc., along with part-of-speech information. Furthermore, the proper noun extraction unit 132 extracts "Japan" and "Yamada," which have been determined to be proper nouns by the morphological analysis, and outputs information that a proper noun has been included to the fixed phrase determination unit 133 together with the morphological analysis result.

[0106] The fixed phrase determination unit 133 determines the similarity with all the fixed phrases stored in the fixed phrase storage unit 130 based on the information obtained by vectorizing the speech recognition result on a sentence-by-sentence basis. As a result, the fixed phrase with the highest similarity to the input speech recognition result, "Shoot, goal by player [name] from [country name]!", is output. As a preprocessing step for determining the similarity, the target fixed phrases may be filtered based on the number of proper nouns included in the morphological analysis result and the number of variable parts in the fixed phrase.

[0107] The proper noun insertion unit 134 replaces the [country name] and [person's name] parts of the fixed phrase output by the fixed phrase determination unit 133 with "Japan" and "Yamada" extracted as proper nouns. As a result, a sign language sentence "Japan, of, Yamada, player's shot, goal!" corresponding to the sign language word label string is output. If no proper nouns are included at the time of morphological analysis, the proper noun extraction unit 132 and the proper noun insertion unit 134 are not executed.

[0108] The motion data acquisition unit 141 reads each piece of motion data that makes up the input sign language sentence "Japan, from, Yamada, player shoots, goal!" from the motion data storage unit 140. The sentence motion synthesis unit 142 interpolates between the motion data to generate sentence motion data "Japan's Yamada player shoots, goal!" By using a fixed phrase, the phrase portion excluding proper nouns such as "player shoots, goal!" can be reproduced as a single piece of motion data. Therefore, the number of places requiring interpolation can be minimized.

[0109] Furthermore, when creating motion data using motion capture, by recording the entire phrase sentence by sentence, it is possible to transcribe motion data that takes into account the meaning of the sentence, including natural facial expressions and mouth shapes. Therefore, for complete phrases such as "The player shoots, goal!", natural facial movements can be reproduced simply by using the motion data as is.

[0110] Finally, the animation generation unit 151 combines the generated sentence motion data with the sign language CG model to render a sign language CG animation.

[0111] As described above, according to this embodiment, by selecting the fixed phrase with the highest similarity on a sentence-by-sentence basis for an input arbitrary sentence in a spoken language, it is possible to generate a sign language CG animation that reliably conveys the minimum meaning of the original sentence without interruptions. Furthermore, according to this embodiment, it is possible to generate a more natural sign language CG animation with fewer word connections and in which facial expressions, mouth shapes, and hand movements match.

[0112] <Effects of the embodiment> The sign language translation system of this embodiment synthesizes sign language CG animation based on a fixed phrase similar to the sentence to be translated. The fixed phrase is associated with motion data corresponding to a phrase or sentence via the sign language fixed phrase. By preparing motion data for each phrase or sentence, the number of connection points for the motion data can be reduced. Furthermore, it is possible to reproduce sign language expressions that are faithful to the whole-body movements of an actual sign language speaker, including facial expressions, mouth shape, etc. Therefore, the sign language translation system of this embodiment can generate natural sign language CG animation.

[0113] [supplement] Each function of the above-described embodiments can be realized by one or more processing circuits. Here, the term "processing circuit" in this specification includes a processor programmed to perform each function by software, such as a processor implemented by an electronic circuit, as well as devices such as an ASIC (Application Specific Integrated Circuit), a DSP (Digital Signal Processor), an FPGA (Field Programmable Gate Array), and conventional circuit modules designed to perform each of the above-described functions.

[0114] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to these embodiments, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims. [Explanation of symbols]

[0115] 1. Sign language translation system 10 Sign language translation device 11 Audio input section 12 Voice recognition unit 13 Sign Language Translation Department 130 Fixed phrase memory section 131 Morphological analysis section 132 Proper noun extraction unit 133 Template Determination Department 134 Proper Noun Insertion 14 Motion synthesis section 140 Motion data storage unit 141 Motion data acquisition unit 142 Text Motion Synthesis Unit 15 Video synthesis section 150 Sign Language Model Memory 151 Animation Generation Unit 16 Animation Output Section 30 User terminals

Claims

1. a storage unit that stores template text and motion data in association with each other; a proper noun extraction unit that extracts proper nouns from a sentence to be translated; a template determination unit that determines a template to be used based on the similarity between the translation target sentence and the template; a proper noun insertion unit that inserts the proper noun into the fixed phrase; an image synthesis unit that synthesizes sign language CG animation based on the motion data associated with the template; Equipped with the template includes a variable for inserting the proper noun; the fixed phrase determination unit calculates the similarity by replacing the proper nouns included in the translation target sentence with the variables; Sign language translation device.

2. 2. The sign language translation device according to claim 1, the motion data corresponds to a phrase obtained by dividing the template at the position of the variable; Sign language translation device.

3. 2. The sign language translation device according to claim 1, The motion data represents at least one of a hand movement, a facial expression, and a mouth shape. Sign language translation device.

4. A program for causing a computer to function as the sign language translation device according to any one of claims 1 to 3.

Citation Information

Patent Citations

  • Device and method for presenting information, and device and method for creating information to be presented

    JP2008216397A

  • Sign language translation device and sign language translation program

    JP2014021180A

  • Motion video generation device and motion video generation program

    JP2014109988A

  • Translation device, translation model learning device, method, and program

    JP2015121992A

  • Voice interpretation device, voice interpretation method, and voice interpretation program

    JP2015125499A