Sign language operation processing device and program
The sign language motion processing device addresses high generation costs by selecting and synthesizing sign language components using pre-partitioned motion data, enhancing naturalness and reducing costs in sign language CG animation.
Patent Information
- Application Number
- JP2023216859
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-22
- Publication Date
- 2025-07-03
AI Technical Summary
Existing methods for generating sign language CG animation, such as motion capture and kineme-based approaches, result in high generation costs due to the need for extensive capture systems, post-processing, and continuous capture of new expressions, leading to inefficiencies and reduced naturalness and smoothness.
A sign language motion processing device that selects sign language components and uses pre-partitioned or existing motion data to generate new motion data for unknown words, reducing the need for continuous motion capture.
This approach significantly reduces the generation cost of sign language CG animation while maintaining naturalness and smoothness by utilizing pre-partitioned or existing motion data to synthesize new sign language motions.
Smart Images

Figure 2025099881000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a sign language operation processing device and a program.
Background Art
[0002] Real-time sign language translation technology using computer graphics (CG) animation is being used. In conventional real-time sign language translation technology, first, voice language texturized by voice recognition or the like is translated into a sign language word sequence for generating sign language CG animation. Next, motion data corresponding to each word in the sign language word sequence is read, and these are interpolated to synthesize (generate) motion data in units of sentences (also referred to as sentence motion data). Then, sign language CG animation is generated by reproducing the motion data in units of sentences with a sign language CG model. Note that motion data means data representing motion (movement), and represents rotation data or rotation information of a plurality of points (for example, joints) for each video frame. Motion data may be simply read as motion as appropriate.
[0003] When generating sign language CG animation, a method has been proposed in which data of motion elements (kinemes) is extracted based on parameters necessary for the input sign presentation, and using this data, the fingers of a character are moved to synthesize sign language animation (see, for example, Patent Document 1). In this method, the posture of the hand determined by the shape of the hand, the position of the hand, and the direction of the hand is defined as a kineme, and each of these is combined to generate a sign language operation. According to this method, the degree of freedom of the sign language operations that can be generated is high, but since animation is generated by specifying kinemes for each operation point and interpolating between those operation points, the generation result is like keyframe animation without manual intervention. For example, when generating a trajectory of an arc motion, it is necessary to set three or more states, and since the movement between those states is generated by interpolation processing with each state as an operation point, the naturalness of the operation is reduced in exchange for the degree of freedom of operation generation.
[0004] In addition, a method of using motion capture is known for generating sign language CG animations. When using motion capture, interpolation processing between frames according to the recorded frame rate occurs when reproducing the movement with an avatar. However, since motion capture data is basically recorded at a high frame rate such as 60 to 120 fps (frames per second), the influence of interpolation is not so great, and the motion points of all frames become the actual movements of humans. Therefore, when comparing the motion generation results using the above-described conventional method and the motion generation results using motion capture data, the conventional method is inferior in terms of naturalness and smoothness.
[0005] As a method of using motion capture, a method has been proposed in which, for the CG pattern of sign language words from the wrist obtained by a glove-type sensor interface, the joint angles of the elbow and shoulder are calculated and presented at a specified position (see, for example, Patent Document 2). However, also in this method, since all the movements of each joint up to the wrist in the animation are calculated and generated, although the reproduction accuracy of only the hand shape is high, the naturalness and smoothness are inferior compared to the method of using motion capture data of the whole body joints including the movements of the arm and elbow.
Prior Art Documents
Patent Documents
[0006]
Patent Document 1
Patent Document 2
Summary of the Invention
Problems to be Solved by the Invention
[0007] As described above, by using the motion data for sign language obtained through motion capture, it becomes possible to reproduce natural sign language motions similar to human motions. On the other hand, in motion capture, in post-processing such as equipment construction, noise removal of recorded data, and interpolation of missing parts, the cost of generating sign language CG animation (the generation cost of sign language CG animation) becomes high.
[0008] In view of the above, an aspect of the present invention aims to provide a technology capable of reducing the generation cost of sign language CG animation.
Means for Solving the Problem
[0009] A sign language motion processing apparatus according to an aspect of the present invention includes a selection unit that selects sign language components constituting a sign language motion, and based on the selected sign language components, uses pre-partitioned sign language component part motion data or existing motion data corresponding to known sign language words to generate new motion data corresponding to an unknown sign language word, and a generation unit.
Effect of the Invention
[0010] According to an aspect of the present invention, the generation cost of sign language CG animation can be reduced.
Brief Description of the Drawings
[0011]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Mode for Carrying Out the Invention
[0012] (Background of the Invention) When generating sign language CG animation using motion capture data, it is necessary to construct a large-scale capture system for recording fine finger movements for sign language and perform manual post-processing by an operator related to the recorded data. Therefore, in addition to the implementation cost, it takes time from the execution of the motion capture operation until the motion capture data can be used, resulting in high generation costs for sign language CG animation.
[0013] Furthermore, not only is it necessary to digitize the operations of new sign language expressions, but even when only changing the handshape, hand orientation, or hand position of existing sign language words, it is necessary to entirely re-perform motion capture as new. Therefore, in methods that use motion data, there is also the problem that the capture work must be continuously performed semi-permanently in order to increase the vocabulary. In this regard, FIG. 1 is a diagram showing an example of motion data in which the handshape and / or hand orientation has been changed. For Japanese sign language "iku" (go), as shown in the part surrounded by the circle in FIG. 1, depending on the subject and means, there are (a) [go alone], (b) [go with two people], (c) [go by car], (d) [go by plane], etc. Thus, even when the movement of the arm is the same but the handshape and / or hand orientation is different, it is necessary to individually capture them as different motion data.
[0014] Furthermore, from the perspective of the understandability of the sign language expressions of the sign language CG animations to be generated, the above-mentioned problems will be supplemented and explained below.
[0015] For example, consider generating a sign language CG animation from a Japanese text such as "Tomorrow, I will go to Tokyo by car." When translating this Japanese text into sign language to generate a sign language CG animation, as the sign language word sequence of the translation result (a sequence in which Japanese headings of sign language words are arranged in time series), "[Tomorrow][Drive][Tokyo][Go]" is output. However, in order to reproduce natural sign language, it is necessary to concisely express the parts expressed by two independent sign languages, "[Drive]" and "[Go]", with a single sign language, "[Go by car]", where one hand is in the shape of a car and the "[Go]" sign language motion is performed. This expression is linguistically called a CL (Classifier) handshape in sign language, and is a unique expression in sign language that represents the object of indication, its shape, size, operation method, etc. in the shape of the hand. Conventionally, when reproducing a CL handshape in a sign language CG animation, it is necessary to newly motion-capture a new sign language expression such as "[Go by car]". By registering the motion capture data motion-captured in this way in a sign language word motion database, it becomes possible to use it for the first time to correct the translation result into a more appropriate expression. Since various cases occur depending on the context and vocabulary combination even for just one CL handshape, it is not realistic to motion-capture the expressions that are lacking each time according to the translation result just because the motion reproduction quality of the motion capture data is high.
[0016] In consideration of the above points, the inventor has arrived at the present invention for reducing the generation cost of sign language CG animations.
[0017] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. However, the scope of the invention is not limited to the illustrated examples. In the following description, those having the same functions and configurations may be denoted by the same reference numerals, and the description thereof may be omitted.
[0018] (Embodiment) <Overview> First, referring to FIG. 2, the outline of the processing of the sign language motion synthesis device according to an embodiment of the present invention will be described. Note that the sign language motion synthesis device may also be referred to as a sign language motion processing device or the like.
[0019] As shown in FIG. 2, a sign language motion synthesis device 10 according to an embodiment of the present disclosure acquires a Japanese text as an input. The sign language motion synthesis device 10 generates sign language sentence motion data corresponding to the acquired Japanese text by performing conversion of the acquired Japanese text into a sign language word sequence, reading motion data corresponding to each word in the sign language word sequence, generating new motion data, modifying the read motion data and the generated new motion data, and the like. Then, the sign language motion synthesis device 10 generates a sign language CG animation based on the generated sentence motion data, and outputs the generated sign language CG animation as shown in FIG. 2.
[0020] More specifically, the sign language motion synthesis device 10 selects sign language components that constitute a sign language motion, and uses sign language component part motion data that has been previously partitioned to generate motion data corresponding to the sign language motion based on the selected sign language components, or existing motion data corresponding to known sign language words, to generate new motion data corresponding to unknown sign language words.
[0021] Here, the sign language components may include a "handshape" sign language component, a "motion" sign language component, a "hand orientation" sign language component, and a "position" sign language component. "Handshape" represents the shape of the hand, "motion" represents the motion from the shoulder to the fingers, "hand orientation" represents the orientation of the hand (e.g., the palm of the hand), and "position" represents the presentation position of the motion data. The sign language motion synthesis device 10 can generate new motion data by synthesizing (combining) sign language component part motion data that has been previously partitioned and / or existing motion data corresponding to the "handshape" sign language component and / or the "motion" sign language component, and rotating or moving the motion data according to the "hand orientation" sign language component and / or the "position" sign language component.
[0022] According to the above, the sign language motion synthesis device 10 generates new motion data using sign language component part motion data or existing motion data that has been pre-partitioned without performing motion capture, so that the generation cost of sign language CG animation can be reduced.
[0023] <Configuration of Sign Language Motion Synthesis Device> FIG. 3 is a block diagram showing a hardware configuration example of the sign language motion synthesis device according to the present embodiment. The sign language motion synthesis device 10 may be realized by any computer (computer) or information processing device such as a server, a PC, a tablet, or a smartphone.
[0024] For example, as shown in FIG. 3, the sign language motion synthesis device 10 may include a storage device 11, a processing device 12, a user interface (UI) device 13, a communication device 14, and a bus 15. The storage device 11, the processing device 12, the UI device 13, and the communication device 14 may be interconnected via the bus 15.
[0025] The program or instruction for realizing the functions and processes described above and below in the sign language motion synthesis device 10 may be downloaded from some external device (for example, a server) via a network or the like. Further, such a program or instruction may be provided from a removable storage medium such as a CD-ROM (Compact Disc Read Only Memory) or a flash memory.
[0026] The storage device 11 is realized by a RAM (Random Access Memory), a flash memory, a hard disk drive, etc., and stores files, data, etc. used for the execution of the installed program or instruction together with the installed program or instruction. The storage device 11 may include a non-transitory storage medium.
[0027] The processing device 12 may be realized by, for example, a general-purpose processor or controller (circuit), or may be realized by a dedicated processor or controller (circuit). When realized by a general-purpose processor, the processing device 12 may be realized by one or more CPUs (Central Processing Units), GPUs (Graphics Processing Units), processing circuitry, etc., which may be composed of one or more processor cores. In this case, the processing device 12 executes the functions and processes of the sign language motion synthesis device 10 described above and below according to programs or instructions stored in the storage device 11, data such as parameters used to execute the programs or instructions, etc. Such a program may be a program that causes the computer or the processing device 12 to function as the sign language motion synthesis device 10 according to the embodiments of the present invention when executed by the computer or the processing device 12. Further, such a program may be a program that causes the computer or the processing device 12 to execute the sign language motion synthesis method according to the embodiments of the present invention when executed by the computer or the processing device 12.
[0028] The UI device 13 may be composed of input devices such as a keyboard, mouse, camera, microphone, output devices such as a display, speaker, headset, printer, input / output devices such as a haptic device, touch panel, etc., and realizes an interface between the user of the sign language motion synthesis device 10 and the sign language motion synthesis device 10. For example, the user of the sign language motion synthesis device 10 operates the sign language motion synthesis device 10 by operating a graphical user interface (GUI) displayed on the display or touch panel using a keyboard, mouse, etc.
[0029] The communication device 14 may be realized by various communication circuits that execute communication processing with external devices, communication networks such as the Internet, intranet, LAN, VPN, etc.
[0030] The hardware configuration of the above sign language motion synthesis device 10 is only an example, and the sign language motion synthesis device 10 according to the embodiment of the present invention may be realized by other appropriate hardware configurations.
[0031] FIG. 4 is a block diagram showing a functional configuration example of the sign language motion synthesis device according to the present embodiment.
[0032] For example, as shown in FIG. 4, the sign language motion synthesis device 10 includes a sign language translation unit 101, a motion synthesis unit 102, an animation generation unit 103, a sign language word motion data storage unit 131, a sign language component part motion data storage unit 132, and an avatar data storage unit 133. The sign language translation unit 101, the motion synthesis unit 102, and the animation generation unit 103 may be realized as hardware by the processing device 12, may be realized as software, or may be realized as a combination of hardware and software. The sign language word motion data storage unit 131, the sign language component part motion data storage unit 132, and the avatar data storage unit 133 may be realized by the storage device 11.
[0033] A Japanese text is input to the sign language translation unit 101. For example, the input Japanese text may be a voice language texturized by voice recognition or the like. The sign language translation unit 101 converts (translates) the input Japanese text into a sign language word sequence in which sign language words written in Japanese are arranged in time series. The sign language translation unit 101 outputs the converted sign language word sequence to the motion synthesis unit 102 (specifically, the motion data reading unit 111 included in the motion synthesis unit 102).
[0034] The motion synthesis unit 102 reads sign language word motion data (also referred to as motion data, sign language, or sign language expression) corresponding to each word in the converted sign language word sequence, generates new motion data as necessary, modifies these motion data, interpolates between the modified motion data, and synthesizes the motion data to generate sentence motion data corresponding to the input Japanese sentence text. The motion synthesis unit 102 includes a motion data reading unit 111, a motion modification unit 112, and a motion interpolation unit 113.
[0035] The motion data reading unit 111 receives, as input, the sign language word sequence output by the sign language translation unit 101. The motion data reading unit 111 uses the headings of the sign language words arranged in time series in the sign language word sequence as indexes to read the corresponding sign language word motion data from the sign language word motion data storage unit 131 (which may be referred to as a sign language word motion data database (DB)). The motion data reading unit 111 outputs the read sign language word motion data to the motion modification unit 112 (specifically, the motion adjustment unit 121 included in the motion modification unit 112). The sign language word motion data stored in the sign language word motion data storage unit 131 corresponds to the existing motion data corresponding to known sign language words.
[0036] The motion modification unit 112 modifies the read sign language word motion data in units of sign language word motion data and, as necessary, generates new motion data in parallel. The generated new motion data corresponds to unknown sign language words. The motion modification unit 112 includes a motion adjustment unit 121, a sign language component selection unit 122 for generating new motion data based on sign language components, a sign notation interpretation unit 123, and a new motion generation unit 124. The sign language components may include "hand shape", "motion", "hand orientation", and "position".
[0037] The motion adjustment unit 121 receives, as input, the sign language word motion data output by the motion data reading unit 111. The motion adjustment unit 121 corrects the received sign language word motion data by adjusting the received sign language word motion data in units of sign language word motion data, including reordering the received sign language word motion data, deleting the received sign language word motion data, and changing the playback speed of the received sign language word motion data, the connection interval between motions, etc.
[0038] The motion adjustment unit 121 receives, as input, the new motion data output by the new motion generation unit 124. When the motion adjustment unit 121 receives the new motion data, it performs the above-described processing including the received new motion data.
[0039] The motion adjustment unit 121 outputs the corrected motion data (which may include newly generated new motion data) to the motion interpolation unit 113.
[0040] For example, the motion adjustment unit 121 may replace two or more pieces of motion data in a plurality of pieces of motion data respectively corresponding to a plurality of words in a sign language word sequence converted from the input Japanese sentence text with one piece of new motion data (generated by the new motion generation unit 124 described later).
[0041] Also, for example, when one piece of new motion data is generated, the motion adjustment unit 121 may use the one piece of new motion data as existing motion data.
[0042] The sign language component selection unit 122 selects sign language components such as "handshape", "motion", "hand orientation", "position", etc., input (instructed) by the user of the sign language motion synthesis device 10 during the generation of new motion data. The sign language component selection unit 122 outputs information (e.g., parameters, indexes, identification information, etc.) indicating the selected sign language components to the new motion generation unit 124.
[0043] The sign language notation interpretation unit 123 receives a symbol string according to the sign language notation. Examples of sign language notations include HamNoSys (The Hamburg Sign Language Notation System; Hanke, 2004, see Fig. 6). When generating new motion data, the sign language notation interpretation unit 123 (automatically) selects the sign language components associated with each symbol in the input symbol string according to the sign language notation. That is, the sign language notation interpretation unit 123 selects sign language components based on the symbols according to the input sign language notation. The sign language notation interpretation unit 123 outputs information (such as an index, identification information, etc.) indicating the selected sign language components to the new motion generation unit 124. In this way, by using the sign language notation, it is possible to automate the generation of new motions.
[0044] The new motion generation unit 124 receives, as input, information indicating the sign components output by the sign component selection unit 122 or the sign notation interpretation unit 123. The new motion generation unit 124 reads the corresponding "hand shape" part motion data and / or "motion" part motion data from the sign component part motion data storage unit 132 (which may be referred to as the sign component part motion DB) according to the information indicating the "hand shape" and / or "motion" of the input sign components. The new motion generation unit 124 generates new motion data by synthesizing the selected "hand shape" and "motion", for example, by replacing the information of the joints from the wrist forward of the read "hand shape" part motion data with the information of the corresponding joints of the "motion" part motion data. Also, the new motion generation unit 124 rotates the hand of the generated new motion data by changing the rotation information of the wrist joint according to the information indicating the "orientation of the hand" of the input sign components, for example. Then, the new motion generation unit 124 changes the presentation position of the new motion data by moving the joint positions such as the wrist of the generated and rotated new motion data to the specified position of the absolute coordinates in the three-dimensional space according to the information indicating the "position" of the input sign components, for example. Here, regarding the method of changing the presentation position of the sign, for a single-handed sign, a method of moving based on the joint coordinates of the wrist, or for a two-handed sign, a method of using the center of gravity of the hand position (described in Japanese Patent Application Laid-Open No. 2015-203860) or a method of taking the intermediate value of the trajectory of the sign motion section can be applied, but it is not limited to this.
[0045] When the synthesis, rotation, and position specification processing of the part motion data are completed, the final new motion data is generated. The new motion generation unit 124 outputs the generated final new motion data to the motion adjustment unit 121 and stores it in the sign word motion data storage unit 131. As a result, hereafter, the motion adjustment unit 121 can operate the new motion data in the same form as the existing sign word motion data.
[0046] In the above description, an example of generating new motion data by applying the "handshape", "motion", "hand orientation", and "position", which are sign language components, has been described. However, depending on the sign language word motion data to be generated, it may not be necessary to apply all of these sign language components. That is, it is also possible to apply only the selected sign language components in combination to generate (final) new motion data. As will be described below, the source motion data is not limited to the "handshape" part motion data and the "motion" part motion data, and may be the sign language word motion data stored in the sign language word motion data storage unit 131.
[0047] The new motion generation unit 124 may generate new motion data corresponding to an unknown sign language word by using the sign language component part motion data that has been pre-partitioned or the existing motion data corresponding to a known sign language word based on the sign language components selected by the sign language component selection unit 122 or the sign language notation interpretation unit 123.
[0048] Further, the new motion generation unit 124 may generate new motion data by using, as the sign language component part motion data that has been pre-partitioned, the sign language component part motion data corresponding to the hand shape ("handshape") and / or the motion from the shoulder to the fingers ("motion") selected as the sign language components.
[0049] Further, the new motion generation unit 124 may generate new motion data according to the hand orientation ("hand orientation") and / or the hand position ("position") selected as the sign language components.
[0050] Further, when the new motion generation unit 124 generates new motion data, it may use the new motion data as the existing motion data.
[0051] For example, the new motion generation unit 124 may generate one new motion data for replacing two or more pieces of motion data out of a plurality of pieces of motion data respectively corresponding to a plurality of words in a sign language word sequence converted from a Japanese sentence text, using sign language component part motion data that has been pre-partitioned. Alternatively, the new motion generation unit 124 may generate one new motion data for replacing two or more pieces of motion data out of a plurality of pieces of motion data respectively corresponding to a plurality of words in a sign language word sequence converted from a Japanese sentence text, using existing motion data.
[0052] Also, for example, the new motion generation unit 124 may generate one new motion data using, as sign language component part motion data that has been pre-partitioned, sign language component part motion data corresponding to the shape of the hand ("handshape") and / or sign language component part motion data corresponding to the movement from the shoulder to the fingers ("movement").
[0053] Also, for example, the new motion generation unit 124 may generate one new motion data according to the direction of the hand ("hand direction") and / or position ("position") selected by the sign language component selection unit 122 or the sign language notation interpretation unit 123.
[0054] Also, for example, when one new motion data is generated, the new motion generation unit 124 may use the one new motion data as existing motion data.
[0055] The motion interpolation unit 113 receives, as input, the corrected motion data (which may include new motion data) output by the motion adjustment unit 121. The motion interpolation unit 113 generates sentence motion data corresponding to the input Japanese sentence text by interpolating and connecting between the corrected motion data. The motion interpolation unit 113 outputs the generated sentence motion data to the animation generation unit 103.
[0056] The animation generation unit 103 receives, as input, the sentence motion data output by the motion interpolation unit 113. The animation generation unit 103 generates (renders) a CG animation by combining the sentence motion data with the sign language avatar data that has been pre-loaded from the avatar data storage unit 133 (which may be referred to as an avatar DB) and is in a standby state. The animation generation unit 103 outputs the generated CG animation to a display device (not shown). The display device as the output destination may be included in the sign language motion synthesis device 10 as the UI device 13 of the sign language motion synthesis device 10, or may be a device connected to the sign language motion synthesis device 10 via a network.
[0057] The sign language word motion data storage unit 131 stores sign language word motion data corresponding to sign language words in association with the headings of the sign language words. Further, the sign language word motion data storage unit 131 stores the new motion data generated by the new motion generation unit 124 as sign language word motion data.
[0058] The sign language component part motion data storage unit 132 stores sign language component part motion data in association with information indicating sign language components.
[0059] The avatar data storage unit 133 stores avatar data imitating a person for sign language.
[0060] Two or more of the sign language word motion data storage unit 131, the sign language component part motion data storage unit 132, and the avatar data storage unit 133 may be integrated. For example, the sign language word motion data storage unit 131 and the sign language component part motion data storage unit 132 may be integrated as a motion data storage unit.
[0061] As described above, it becomes possible to generate in real time as necessary the sign language word motion data that is not stored in the sign language word motion data storage unit 131 without performing motion capture.
[0062] In the present application, "A outputs B to C" may mean not only "A directly outputs B to C" but also the following. · "A transmits B to C". · "A causes B to be stored in a storage area accessible by C".
[0063] Next, with reference to FIG. 5, a specific example of sign language motion synthesis when a Japanese sentence text is input to the sign language motion synthesizer 10 will be described. In this example, as shown in FIG. 5(a), the case where the above-described Japanese sentence text "Tomorrow, I will go to Tokyo by car." is input to the sign language motion synthesizer 10 is considered.
[0064] When the Japanese sentence text is input to the sign language motion synthesizer 10, as shown in FIG. 5(b), the sign language translation unit 101 outputs "[tomorrow][drive][Tokyo][go]" as a sign language word sequence to the motion data reading unit 111.
[0065] The motion data reading unit 111 reads the sign language word motion data corresponding to [tomorrow], [drive], [Tokyo], and [go] respectively from the sign language word motion data storage unit 131, and outputs the read sign language word motion data to the motion adjustment unit 121.
[0066] In order to reproduce natural sign language, the motion adjustment unit 121 deletes the sign language word motion data corresponding to [drive] and [go] among the plurality of sign language word motion data corresponding to [tomorrow], [drive], [Tokyo], and [go] respectively, as shown in FIG. 5(c), in order to replace the sign language word motion data corresponding to [drive] and [go] with one sign language word motion data corresponding to "[go by car]".
[0067] Here, if the sign language word motion data corresponding to "[go by car]" has already been motion captured and is stored in the sign language word motion data storage unit 131, the motion adjustment unit 121 deletes the sign language word motion data corresponding to "[drive]" and "[go]", and adds the sign language word motion data corresponding to "[go by car]" for replacing them, so that the target sign language sentence motion data can be generated. On the other hand, if the sign language word motion data corresponding to "[go by car]" is not stored in the sign language word motion data storage unit 131, for example, the following processing is executed.
[0068] The new motion generation unit 124 reads the sign language component part motion data corresponding to the sign language components selected by the sign language component selection unit 122 or the sign language notation interpretation unit 123 from the sign language component part motion data storage unit 132 in order to generate one sign language word motion data corresponding to "[go by car]". The processing by the sign language component selection unit 122 and the sign language notation interpretation unit 123 will be described later with reference to FIG. 6. The new motion generation unit 124 synthesizes the read sign language component part motion data to generate one new motion data corresponding to "[go by car]" as shown in FIG. 5(d). Then, the new motion generation unit 124 outputs the generated new motion data to the motion adjustment unit 121.
[0069] As shown in FIG. 5(e), the motion adjustment unit 121 replaces the sign language word motion data corresponding to "[drive]" and "[go]" in the plurality of sign language word motion data corresponding to "[tomorrow]", "[drive]", "[Tokyo]" and "[go]" respectively with one new motion data corresponding to "[go by car]". Then, the motion adjustment unit 121 outputs the motion data corresponding to "[tomorrow]", "[Tokyo]" and "[go by car]" to the motion interpolation unit 113.
[0070] The motion interpolation unit 113 generates sign language sentence motion data corresponding to the input Japanese sentence text based on the motion data corresponding to [[tomorrow]], [[Tokyo]], and [[going by car]].
[0071] Next, with reference to FIG. 6, the processing of the sign language component selection unit 122 and the sign language notation interpretation unit 123 will be described.
[0072] When a sign language component is selected by the sign language component selection unit 122, a screen for the user of the sign language motion synthesizer 10 to select the "hand shape" as shown in FIG. 6(a), the "motion" as shown in FIG. 6(b), the "hand orientation" as shown in FIG. 6(c), and the "position" as shown in FIG. 6(d) is displayed to the user. In addition, the screen may include information for allowing the user of the sign language motion synthesizer 10 to select a sign language corresponding to the sign language word motion data stored in the sign language word motion data storage unit 131. Then, the sign language component selection unit 122 receives information indicating the sign language components ("hand shape", "motion", "hand orientation", "position", etc.) selected (input) by the user of the sign language motion synthesizer 10, and outputs the received information to the new motion generation unit 124.
[0073] On the other hand, when a sign language component is selected by the sign language notation interpretation unit 123, for example, a symbol string according to the sign language notation "HamNoSys" shown in FIG. 6(e) is input to the sign language notation interpretation unit 123. The sign language notation interpretation unit 123 selects the sign language components ("hand shape", "motion", "hand orientation", "position", etc.) associated with each symbol in the input symbol string. Then, the sign language notation interpretation unit 123 outputs information indicating the selected sign language components to the new motion generation unit 124.
[0074] Next, with reference to FIGS. 7 to 10, specific examples of the new motion data generated by the new motion generation unit 124 will be described.
[0075] The new motion data [Sorry 2] (International Sign Language) is generated by reading and synthesizing the "handshape" part motion data and the "motion" part motion data corresponding to the "handshape: both hands with 5 fingers extended and aligned" and "motion: the motion of rotating the arm clockwise" shown in Figure 7, and finally generated by rotating and moving the generated new motion data according to the "hand orientation: the palm facing down" and "position: centered left and right relative to the body and at chest height" shown in Figure 7. More specifically, first, according to the "handshape" and "motion" shown in Figure 7, new motion data of "the motion of rotating the arm in front of the body in the state of the handshape with 5 fingers extended and aligned" can be generated. Next, according to the "hand orientation" shown in Figure 7, by rotating the wrist joint of the generated new motion data to make the palm face down, it can be deformed into motion data of "the motion of rotating the arm clockwise in the state of the palm facing down with 5 fingers extended and aligned". Finally, according to the "position" shown in Figure 7, by moving the presentation position of the deformed motion data to "the center of the body, height = chest", the final new motion data corresponding to the target [Sorry 2] (International Sign Language) of "the motion of rotating the arm clockwise at the center left and right of the body and at chest height in the state of the palm facing down with 5 fingers extended and aligned" can be generated.
[0076] Also, the new motion data [IMPORTANT 1] (German Sign Language) is generated by reading and synthesizing the sign language word motion data corresponding to [Study] (Japanese Sign Language) shown in Figure 8 and the "handshape" part motion corresponding to the "handshape" shown in Figure 8.
[0077] Also, the new motion data [SAY 1] (German Sign Language) is generated by rotating the sign language word motion data corresponding to [Say] (Japanese Sign Language) shown in Figure 9 according to the "hand orientation" shown in Figure 9.
[0078] Also, the new motion data [I see] (Japanese Sign Language) is generated by moving the sign language word motion data corresponding to the [yellow] (Japanese Sign Language) shown in FIG. 10 according to the [position] shown in FIG. 10.
[0079] In this way, based on the sign language components, new motion data corresponding to unknown sign language words can be generated by using the pre-partitioned sign language component part motion data or the existing motion data corresponding to known sign language words.
[0080] <Operation of Sign Language Motion Synthesis Device> Next, with reference to FIG. 11, an operation example of the sign language motion synthesis device 10 according to the present embodiment will be described.
[0081] In step S11, the sign language motion synthesis device 10 receives a Japanese text as an input. Step S11 is executed by the sign language translation unit 101.
[0082] In step S12, the sign language motion synthesis device 10 performs translation by converting the received Japanese text into a sign language word sequence. Step S12 is executed by the sign language translation unit 101.
[0083] In step S13, the sign language motion synthesis device 10 reads the sign language word motion data corresponding to each word in the converted sign language word sequence from the sign language word motion data storage unit 131. Step S13 is executed by the motion data reading unit 111.
[0084] In step S14, the sign language motion synthesis device 10 corrects the read sign language word motion data (and the new motion data generated in step S15) in units of motion data, including reordering the sign language word motion data (and the new motion data generated in step S15), and adjusting the deletion / playback speed of the sign language word motion data and the connection interval between motions. Step S14 is executed by the motion correction unit 112 (motion adjustment unit 121).
[0085] In step S15, which is a sub-step of step S14, the sign language motion synthesizing apparatus 10 generates new motion data. Step S15 is executed by the motion correction unit 112 (sign language component selection unit 122, sign language notation interpretation unit 123, new motion generation unit 124).
[0086] In step S16, the sign language motion synthesizing apparatus 10 generates sentence motion data corresponding to the input Japanese sentence text by interpolating and connecting between the corrected motion data (including the newly generated new motion data). Step S16 is executed by the motion interpolation unit 113.
[0087] In step S17, the sign language motion synthesizing apparatus 10 generates a sign language CG animation by combining the sentence motion data and the avatar data stored in the avatar data storage unit 133. Step S16 is executed by the animation generation unit 103.
[0088] <Modification Example> In the above embodiment, the description is based on the premise of generating new motion data in real time, but it is not limited thereto. For example, the sign language component selection unit 122, the sign language notation interpretation unit 123, and the new motion generation unit 124 may be configured to generate new motions offline, independent of the other functional units shown in FIG. 4.
[0089] The functional units described and illustrated in the above embodiment may be divided into a plurality of functional units, or two or more functional units may be integrated into one functional unit. Also, the names of the functional units are merely examples and may be changed as appropriate. For example, the sign language component selection unit 122 and / or the sign language notation interpretation unit 123 may be referred to as a selection unit or the like, and the new motion generation unit 124 may be referred to as a generation unit or the like.
[0090] In the above-described embodiment, the steps in the flowchart described and illustrated may be executed with the order appropriately switched or may be executed in parallel.
[0091] <Effects of the Embodiment> As described above, the sign language motion synthesis device 10 according to the present embodiment selects sign language components that make up a sign language motion by the sign language component selection unit 122 or the sign language notation interpretation unit 123. Further, the sign language motion synthesis device 10 according to the present embodiment, based on the sign language components selected by the sign language component selection unit 122 or the sign language notation interpretation unit 123, uses pre-partitioned sign language component part motion data or existing motion data corresponding to known sign language words to generate new motion data corresponding to unknown sign language words. Thereby, the sign language motion synthesis device 10 generates new motion data using pre-partitioned sign language component part motion data or existing motion data without performing motion capture, so that the generation cost of sign language CG animation can be reduced, the motion data can be expanded, and it becomes possible to reproduce natural sign language expressions in various variations.
[0092] <Summary of the Embodiment> A sign language motion processing device according to an aspect of the present invention includes a selection unit that selects sign language components that make up a sign language motion, and a generation unit that generates new motion data corresponding to an unknown sign language word using pre-partitioned sign language component part motion data or existing motion data corresponding to known sign language words based on the selected sign language components.
[0093] In one example, the generation unit generates the new motion data using, as the pre-partitioned sign language component part motion data, sign language component part motion data corresponding to the hand shape and / or the motion from the shoulder to the finger selected as the sign language components, based on the hand shape and / or the motion from the shoulder to the finger.
[0094] In one example, the generation unit generates the new motion data according to the orientation and / or position of the hand selected as the sign language component.
[0095] In one example, the selection unit selects the sign language component based on a symbol according to a sign language notation method.
[0096] In one example, when the new motion data is generated, the generation unit uses the new motion data as existing motion data.
[0097] A program according to one aspect of the present invention causes a computer to execute a process including: selecting a sign language component that constitutes a sign language motion; and generating new motion data corresponding to an unknown sign language word using pre-partitioned sign language component part motion data or existing motion data corresponding to a known sign language word based on the selected sign language component.
[0098] As described above, the embodiments of the present invention have been described with reference to the drawings, but the present invention is not limited to such examples. It is obvious that those skilled in the art can conceive of various modification examples or correction examples within the scope described in the claims. Such modification examples or correction examples are also understood to belong to the technical scope of the present invention. Also, within the scope not departing from the gist of the present invention, the components and matters in the embodiments (including modification examples) may be arbitrarily combined.
[0099] The present invention can be realized by software, hardware, or software in cooperation with hardware.
Explanation of Reference Numerals
[0100] 10 Sign language motion synthesis device 11 Storage device 12 Processing device 13 User interface (UI) device 14 Communication device 101 Sign language translation unit 102 Motion synthesis unit 103 Animation generation unit 111 Motion data reading unit 112 Motion correction unit 113 Motion interpolation unit 121 Motion adjustment unit 122 Sign language component selection unit 123 Sign language notation interpretation unit 124 New motion generation unit 131 Sign language word motion data storage unit 132 Sign language component part motion data storage unit 133 Avatar data storage unit
Claims
1. A selection unit that selects sign language components constituting a sign language motion, A generation unit that generates new motion data corresponding to an unknown sign language word using pre-partitioned sign language component part motion data or existing motion data corresponding to known sign language words based on the selected sign language components, A sign language motion processing device comprising the above.
2. Based on the hand shape and / or the motion from the shoulder to the fingers selected as the sign language components, the generation unit uses, as the pre-partitioned sign language component part motion data, sign language component part motion data corresponding to the hand shape and / or the motion from the shoulder to the fingers to generate the new motion data. The sign language motion processing device according to Claim 1. The sign language motion processing device according to Claim 1.
3. The generation unit generates the new motion data according to the orientation and / or position of the hand selected as the sign language component. The sign language motion processing device according to Claim 1. The sign language motion processing device according to Claim 1.
4. The selection unit selects the sign language components based on symbols according to a sign language notation. The sign language motion processing device according to Claim 1. The sign language motion processing device according to Claim 1.
5. When the new motion data is generated, the generation unit uses the new motion data as existing motion data. The sign language motion processing device according to Claim 1. The sign language motion processing device according to Claim 1.
6. Selecting sign language components constituting a sign language motion, Generating new motion data corresponding to an unknown sign language word using pre-partitioned sign language component part motion data or existing motion data corresponding to known sign language words based on the selected sign language components, A program for causing a computer to execute a process including the above.
Citation Information
Patent Citations
Device for generating talking with hands
JP1994251123A
Finger language information presenting device
JP1999184370A