Sign language information transmission device, sign language information output device, sign language information transmission system, and program

The sign language information transmission system addresses the challenge of conveying sign language meaning by synchronizing visual, auditory, and tactile stimuli with sign language videos, enhancing communication clarity and understanding.

JP7696250B2Active Publication Date: 2025-06-20NIPPON HOSO KYOKAI
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
JP2021128187
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-08-04
Publication Date
2025-06-20
Estimated Expiration
2041-08-04

AI Technical Summary

Technical Problem

Existing sign language video technologies struggle to convey the original meaning of sign language due to the lack of synchronization with auditory and tactile stimuli, leading to decreased understanding for the listener.

Method used

A sign language information transmission system that includes a detection unit to identify the timing of auditory or tactile stimuli in sign language videos and a transmission unit to synchronize these stimuli with the video data, allowing for synchronized output of visual, auditory, and tactile stimuli.

Benefits of technology

The system effectively enhances the understanding of sign language by reproducing stimuli such as plosive sounds and vibrations in synchronization with the sign language video, thereby improving communication clarity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007696250000001
    Figure 0007696250000001
  • Figure 0007696250000002
    Figure 0007696250000002
  • Figure 0007696250000003
    Figure 0007696250000003
Patent Text Reader

Abstract

To output stimuli generated accompanying a sign language in synchronization with a picture of the sign language.SOLUTION: A sign language information transmission system has a sign language information transmission device and a sign language information output device. The sign language information transmission device includes a transmission unit for transmitting sign language video data and timing data showing a timing at which auditory or tactile stimuli of a prescribed kind is generated in the sign language indicated by the sign language video data. The sign language information output device includes: a first receiving unit for receiving video data; a second receiving unit for receiving the timing data; a reproduction control unit for reproducing the video data; and an output control unit for controlling a device so as to output stimuli of vibration, light, or sound at a timing indicated by the timing data.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sign language information transmission device, a sign language information output device, a sign language information transmission system, and a program.

Background Art

[0002] There is a demand for a system that translates from Japanese to Japanese sign language and presents the translation result in the form of sign language. For this realization, research has been conducted on techniques for synthesizing sign language CG (computer graphics) animations (see, for example, Patent Document 1). However, through the evaluation experiments of this technology, it has become clear that it is very difficult to convey the original meaning only with videos expressing sign language. Therefore, improvement of technology for making sign language easier to convey is required.

[0003] Looking at the scenes where sign language is actually used, the speaker and the listener are often in close proximity. If the distance during conversation is short, it is conceivable that not only visual images but also the sounds and vibrations emitted by the speaker are transmitted. On the other hand, what the sign language video reproduces is only the stimulus to the visual sense of the viewer. Along with the difference in the distance between the speaker and the listener, some important information is missing only with the sign language video, and it is considered that this affects the decrease in the listener's understanding.

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] Observing where two deaf people are conversing in sign language, it can be considered that the information transmitted as visual information is only a part of the whole, and communication is achieved using various means other than visual information. For example, in sign language spoken by a hard-of-hearing person or a deaf person, sound is not always completely absent. Sound may be generated when the hand touches the opposite hand or a body part, or sound may be generated from the mouth. From the flow of sign language conversation, it is considered that by generating a loud sound to emphasize in a particularly important context, the other party is being asked to understand.

[0006] Also, even if not recognized as hearing, the movement of the speaker at a very close distance can be transmitted as air vibrations and can be imagined to be felt by the listener as an "air sensation". In particular, people who are congenitally deaf have extremely sensitive senses throughout their bodies compared to hearing people, and there is a high possibility that communication is being carried out smoothly.

[0007] From the above, in order to reproduce easily transmissible sign language, it is considered desirable to reproduce not only the video for vision but also the stimuli generated by sign language utterances such as hearing and touch. Also, reproducing these stimuli separately not only makes no sense but may rather hinder understanding. Therefore, reproduction at the timing synchronized with sign language utterance is required.

[0008] The present invention has been made in consideration of such circumstances, and an object thereof is to provide a sign language information transmission device, a sign language information output device, a sign language information transmission system, and a program that can output stimuli generated along with sign language in synchronization with a sign language video.

Means for Solving the Problem

[0009] [1]One aspect of the present invention is a sign language information transmission device including a detection unit that detects a timing at which a predetermined type of auditory or tactile stimulus has occurred in the sign language shown in the video data based on one or both of the feature amounts obtained from the video data of the sign language and the feature amounts obtained from the audio data of the sign language, and a transmission unit that transmits the video data and timing data indicating the timing detected by the detection unit.

[0010] [2]One aspect of the present invention is the above-described sign language information transmission device, wherein the predetermined type of stimulus is the utterance of a plosive sound, the detection unit detects the timing at which the plosive sound is uttered and the type of the plosive sound, and the transmission unit transmits the timing data indicating the timing and the type of the plosive sound detected by the detection unit.

[0011] [3]One aspect of the present invention is a sign language information output device including a first reception unit that receives video data of sign language, a second reception unit that receives timing data indicating a timing at which a predetermined type of auditory or tactile stimulus has occurred in the sign language shown in the video data, a reproduction control unit that reproduces the video data, and an output control unit that controls a device to output a stimulus of vibration, light, or sound at the timing indicated by the timing data.

[0012] [4]One aspect of the present invention is the above-described sign language information output device, wherein the stimulus is the utterance of a plosive sound, the timing data includes the timing at which the plosive sound is uttered and information on the type of the plosive sound, and the output control unit controls the device to output a stimulus of vibration, light, or sound corresponding to the type of the plosive sound at the timing.

[0013] [5]One aspect of the present invention is a sign language information transmission system having a sign language information transmission device and a sign language information output device. The sign language information transmission device includes a transmission unit that transmits video data of sign language and timing data indicating a timing at which a predetermined type of auditory or tactile stimulus occurs in the sign language indicated by the video data. The sign language information output device includes a first reception unit that receives the video data, a second reception unit that receives the timing data, a reproduction control unit that reproduces the video data, and an output control unit that controls a device to output a vibration, light, or sound stimulus at the timing indicated by the timing data. It is a sign language information transmission system.

[0014] [6]One aspect of the present invention is a program for causing a computer to function as any of the above-described sign language information transmission devices.

[0015] [7]One aspect of the present invention is a program for causing a computer to function as any of the above-described sign language information output devices.

Effects of the Invention

[0016] According to the present invention, it becomes possible to output a stimulus that occurs along with sign language in synchronization with the video of sign language.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Mode for Carrying Out the Invention

[0018] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings. In this embodiment, in order to supplement the visual stimulus by the sign language video, stimuli by hearing and touch are reproduced in synchronization with the sign language video. Examples of the stimuli of sign language other than the visual stimulus by the sign language video include stimuli of hearing and touch such as sounds generated from the mouth and sounds generated when the hand touches other parts of the body. Hereinafter, stimuli of hearing and touch generated in sign language other than the visual stimulus of the sign language video are simply referred to as stimuli. In this embodiment, it can be used for all stimuli of hearing and touch in sign language. Hereinafter, for the sake of clarity, the case where the stimulus is the utterance of the plosive sounds "pa·pi·pu·pe·po" accompanying the mouth shapes of Japanese sign language will be mainly used as an example for explanation. The utterance of plosive sounds is the most prominent example of sign language information that cannot be conveyed in video (see, for example, Reference 1).

[0019] (Reference 1) Edited by the NPO Bilingual Bicultural Deaf Education Center, "Understanding the Structure of Japanese Sign Language from the Basics of Grammar", Daisukan Publishing Co., Ltd., 2011, pp. 89-92

[0020] [First Embodiment] FIG. 1 is a diagram showing the configuration of a sign language information transmission system 1 according to a first embodiment of the present invention. The sign language information transmission system 1 includes a sign language information transmission device 100 and a sign language information output device 300. The sign language information transmission device 100 and the sign language information output device 300 are connected via a transmission network 500. The transmission network 500 may be a broadcast, a communication network such as the Internet, or both a broadcast and a communication network. In FIG. 1, only one sign language information output device 300 is shown, but the number of sign language information output devices 300 is arbitrary. In the present embodiment, a case where a sign language video distributor who distributes sign language videos or the like has the sign language information transmission device 100 and a viewer of the sign language videos has the sign language information output device 300 will be described as an example.

[0021] In order to perform reproduction closer to actual sign language, the sign language information transmission device 100 transmits not only the video of sign language recognized visually but also the timing of generation of stimuli such as hearing and touch in sign language to the sign language information output device 300. The timing of generation of stimuli in sign language may be manually input by the sign language video distributor to the sign language information transmission device 100, but the sign language information transmission device 100 can assist the sign language video distributor by detecting it using video recognition technology or voice recognition technology. The sign language information transmission device 100 transmits sign language video data, which is video data of sign language, and timing data in which the timing of generation of stimuli in the sign language indicated by the sign language video data is set, to the sign language information output device 300. In this way, the timing of generation of stimuli is transmitted in a form that can be electronically handled.

[0022] The sign language information output device 300 plays the sign language video based on the sign language video data, and outputs stimuli other than the sign language video in synchronization with the sign language video at the occurrence timing described in the timing data. The device that outputs the stimuli is referred to as a stimulus output device. The stimulus output device is, for example, an auditory device, a tactile device, a light-emitting device, etc. The auditory device is a device that gives a sound stimulus to the sense of hearing, such as a speaker. The tactile device is a device that gives a vibration stimulus to the sense of touch, such as a vibrator. The light-emitting device is a device that gives a light stimulus to the sense of vision, such as a lamp.

[0023] In the first embodiment, the sign language video data is actual video data of a signer being photographed. During operation, the sign language information transmission device 100 distributes the sign language video data to the sign language information output device 300. As preparation before the operation, detection conditions for detecting the occurrence of a stimulus are set in the sign language information transmission device 100. To determine the detection conditions, the sign language video distributor prepares video data of the mouth of the signer being photographed while the signer is signing, and the voice data of that sign language. The sign language video distributor manually adds information on the timing when a stimulus such as the utterance of a popping sound occurs in the video data and the voice data, and generates timed sign language video data and timed sign language voice data. The sign language information transmission device 100 detects a feature amount representing the occurrence of a stimulus, a change in the feature amount, etc. based on the feature amounts obtained from the timed sign language video data and the timed sign language voice data. The sign language information transmission device 100 stores the detection conditions of the stimulus based on the detected feature amounts and the change in the feature amounts.

[0024] During operation, a sign language user signs, and sign language video data obtained by photographing the sign language or sign language voice data obtained by recording the voice of the sign language is input to the sign language information transmission device 100. When the feature amount of the mouth video obtained from the sign language video data or the feature amount obtained from the sign language voice data satisfies the detection condition of the stimulus, the sign language information transmission device 100 detects the occurrence of the stimulus. The sign language information transmission device 100 generates timing data describing the occurrence timing of the detected stimulus. The sign language video distributor may modify the timing data as necessary. The sign language information transmission device 100 transmits the sign language video data, the sign language voice data, and the timing data to the sign language information output device 300 of each viewer. Note that the transmission of the sign language voice data may be optional. The sign language information output device 300 displays a sign language video based on the sign language video data and outputs a sign language voice based on the sign language voice data. Further, the sign language information output device 300 outputs a stimulus other than the sign language video in synchronization with the sign language video by the stimulus output device at the occurrence timing indicated by the timing data.

[0025] As described above, the sign language information transmission system 1 can not only return a voice signal generated by sign language, such as represented by papipupepo, to a voice signal of the same format at the sign language information output device 300 on the playback side, but also convert it into other forms of stimuli output by a light emitting device, a vibration device, etc. and output it. In this way, the sign language information output device 300 outputs a voice signal specialized for speaking sign language as a stimulus in synchronization with the video of the sign language, thereby improving the understanding of the person viewing the sign language video reproduced by the sign language information output device 300.

[0026] FIG. 2 is a functional block diagram showing the configuration of the sign language information transmission device 100. In FIG. 2, only the functional blocks related to the present embodiment are extracted and shown. The sign language information transmission device 100 is realized by, for example, one or more computer devices. The sign language information transmission device 100 includes a data input unit 101, a storage unit 102, an input unit 103, a display unit 104, an audio output unit 105, a playback control unit 111, an analysis unit 112, a detection condition setting unit 113, a detection unit 114, a distribution data generation unit 115, a correction unit 116, and a transmission unit 117.

[0027] The data input unit 101 inputs sign language video data from the camera 201 and sign language audio data from the microphone 202. The sign language video data is video data captured by the camera 201 of the sign language of the signer. The video data includes time information and a video frame at that time. The time information may be, for example, time using UTC (Coordinated Universal Time), or relative time with the time when a predetermined video frame such as the first frame is captured as the start time. Also, the frame number of the video frame may be used as the time information. The sign language audio data is audio data recorded by the microphone 202 when sign language is being performed by the signer. The sign language audio data includes time information and data representing the audio at that time. The time information of the sign language video data and the time information of the sign language audio data are synchronized.

[0028] The storage unit 102 stores various data. The storage unit 102 stores sign language video data, sign language audio data, time-stamped video data, time-stamped sign language audio data, and detection condition data. The time-stamped video data is sign language video data to which information on the occurrence timing of a stimulus is added. The time-stamped sign language audio data is sign language audio data to which information on the occurrence timing of a stimulus is added. The detection condition data indicates the conditions for detecting the occurrence of a stimulus.

[0029] The input unit 103 is configured using existing input devices such as a keyboard, a pointing device (mouse, tablet, etc.), buttons, and a touch panel. The input unit 103 is operated by the user when inputting various instructions to the sign language information transmission device 100. The input unit 103 may be an interface for connecting the input device to the sign language information transmission device 100. In this case, the input unit 103 inputs an input signal generated in response to the user's input in the input device to the sign language information transmission device 100.

[0030] The display unit 104 displays images. For example, the display unit 104 is an image display device such as a liquid crystal display, an organic EL (Electro Luminescence) display, or a CRT (Cathode Ray Tube) display. The display unit 104 may be an interface for connecting the image display device to the sign language information transmission device 100. In this case, the display unit 104 generates a video signal for displaying a video and outputs the video signal to the image display device connected to the sign language information transmission device 100.

[0031] The audio output unit 105 outputs audio. The audio output unit 105 is, for example, an audio output device such as a speaker. The audio output unit 105 may be an interface for connecting the audio output device to the sign language information transmission device 100. In this case, the audio output unit 105 generates an audio signal for outputting audio and outputs the audio signal to the audio output device connected to the sign language information transmission device 100.

[0032] The playback control unit 111 displays video data on the display unit 104 and outputs audio data from the audio output unit 105. The analysis unit 112 acquires the timing at which a stimulus occurs and one or more types of feature amounts before and after the timing at which the stimulus occurs from the timed video data and the timed sign language audio data. The analysis unit 112 detects a change in the feature amount at the timing at which the stimulus occurs or the feature amount in a time period including the timing at which the stimulus occurs. The detection condition setting unit 113 generates detection condition data indicating a feature amount or a change in the feature amount for determining that the occurrence of a stimulus has been detected based on the detection result by the analysis unit 112 and writes the detection condition data to the storage unit 102. The detection condition setting unit 113 may generate and correct the detection condition data based on the information input by the video distributor through the input unit 103.

[0033] The detection unit 114 acquires time-series feature quantities from the sign language video data and sign language voice data to be distributed, and determines whether or not the acquired feature quantities satisfy the detection conditions indicated by the detection condition data. When the detection unit 114 determines that the detection conditions are satisfied, it detects the occurrence of a stimulus. The detection unit 114 acquires, as the occurrence timing of the stimulus, the information on the time when a feature quantity satisfying the detection conditions is obtained in the sign language video data or sign language voice data to be distributed. The distribution data generation unit 115 generates timing data indicating the occurrence timing of the stimulus detected by the detection unit 114. The correction unit 116 corrects the timing data in accordance with the instruction of the sign language video distributor input by the input unit 103. The transmission unit 117 transmits the sign language video data, the sign language voice data, and the timing data to the sign language information output device 300. The transmission unit 117 may transmit the sign language video data, the sign language voice data, and the timing data by either broadcasting or communication. The sign language video data and the sign language voice data may be transmitted by broadcasting, and the timing data may be transmitted by communication. In the present embodiment, the timing data is set in the sign language information distribution data shown in FIG. 5 described later and transmitted. The sign language information distribution data indicates the correspondence between the sign language video data and the sign language voice data and the timing data.

[0034] FIG. 3 is a functional block diagram showing the configuration of the sign language information output device 300. In FIG. 3, only the functional blocks related to the present embodiment are extracted and shown. The sign language information output device 300 is, for example, a smartphone, a tablet terminal, a personal computer, a television receiver, or the like. The sign language information output device 300 includes a receiving unit 301, a storage unit 302, a display unit 303, an audio output unit 304, an input unit 305, a playback control unit 306, a delay addition unit 307, an output control unit 308, and an output device 309. The sign language information output device 300 may have a plurality of output devices 309 of different types or the same type.

[0035] The receiving unit 301 has the function of a first receiving unit that receives sign language video data and sign language audio data, and the function of a second receiving unit that receives timing data. The receiving unit 301 receives sign language video data, sign language audio data, and sign language information distribution data with timing data set from the sign language information transmission device 100, and writes them to the storage unit 302. The storage unit 302 stores various data. The storage unit 302 stores sign language video data, sign language audio data, sign language information distribution data, usage device information, and device-specific stimulus information. The usage device information indicates the types of one or more stimulus output devices used by the sign language information output device 300 for outputting stimuli. The types of stimulus output devices are, for example, auditory devices, tactile devices, light-emitting devices, etc. The device-specific stimulus information includes stimulus information for each stimulus output device. The stimulus information indicates the stimuli output from the stimulus output device to represent the occurrence of stimuli other than the sign language video in sign language. For example, the stimulus information indicates sounds output from the auditory device, vibration patterns of the tactile device, light emission patterns of the light-emitting device, etc.

[0036] The display unit 303 displays video. For example, the display unit 303 is an image display device such as a liquid crystal display, an organic EL display, or a CRT display. The display unit 303 may be an interface for connecting an image display device to the sign language information output device 300. In this case, the display unit 303 generates a video signal for displaying video and outputs the video signal to the image display device connected to the sign language information output device 300.

[0037] The audio output unit 304 outputs audio. The audio output unit 304 is an audio output device (audio output device) such as a speaker, for example. The audio output unit 304 may be an interface for connecting an audio output device to the sign language information output device 300. In this case, the audio output unit 304 generates an audio signal for outputting audio and outputs the audio signal to the audio output device connected to the sign language information output device 300. When outputting a stimulus by audio, the audio output unit 304 also serves as the output device 309.

[0038] The input unit 305 is configured using existing input devices such as a touch panel, buttons, a keyboard, and pointing devices (mouse, tablet, etc.). The input unit 305 is operated by the user when inputting various instructions to the sign language information output device 300. The input unit 305 may be an interface for connecting the input device to the sign language information output device 300. In this case, the input unit 305 inputs an input signal generated in response to the user's input in the input device to the sign language information output device 300. For example, when the sign language information output device 300 is a television receiver, the input unit 305 may receive the instructions input to the remote controller by infrared rays or receive the instructions input to the smartphone wirelessly.

[0039] The playback control unit 306 displays the sign language video data on the display unit 303 and outputs the sign language audio data from the audio output unit 304. The delay addition unit 307 adds a delay to the sign language video data displayed on the display unit 303 by the playback control unit 306 and the sign language audio data output from the audio output unit 304. The output control unit 308 controls the output device 309 of the type indicated by the used device information to output a stimulus at the output timing of the stimulus indicated by the timing data. The output device 309 is a stimulus output device. The output device 309 is, for example, a vibration device such as a vibrator or a light emitting device such as a lamp. When the output device 309 is a vibration device, the stimulus is output as vibration. When the output device 309 is a light emitting device, the stimulus is output as light. When the audio output unit 304 is used as the output device 309, the stimulus is output as sound.

[0040] FIG. 4 is a diagram showing an example of sign language information distribution data. A language for multimedia description such as SMIL (Synchronized Multimedia Integration Language) is used for the sign language information distribution data. The sign language information distribution data includes timing data described by XML (extensible markup language). The sign language transmission device 100 can transmit the sign language information distribution data described in a general-purpose language to the sign language information output device 300, and output a stimulus from the stimulus output device in synchronization with the sign language video data.

[0041] The sign language information distribution data includes the display position of the sign language video data, the sign language video data name, the sign language video playback timing, the sign language audio data name, the sign language audio playback timing, and the timing data. The sign language audio playback timing may be omitted if it is the same as the sign language video playback timing. The timing data includes the type of stimulus and the stimulus playback timing. In the present embodiment, the type of popping sound is set as the type of stimulus. Instead of the type of stimulus, a stimulus information name may be set. The stimulus information name is an example of information for specifying stimulus information. Further, information on the type of stimulus output device used for output of the stimulus may be set in association with the type of stimulus. Note that if the stimuli output from the stimulus output device are the same regardless of the type of stimulus, the setting of the type of stimulus and the stimulus information name can be omitted. The stimulus playback timing may be set based on the occurrence timing of the stimulus detected by the sign language transmission device 100, or may be manually set by the sign language video distributor.

[0042] FIG. 5 is a diagram showing an example of stimulus information for each device. The stimulus information for each device is information that associates the type of stimulus and the stimulus information for each type of stimulus output device included in the sign language information output device 300. FIG. 5 shows an example in which the sign language information output device 300 can use a speaker, a vibrator, and a lamp as the stimulus output device. In the present embodiment, the type of the crack sound is set as the type of the stimulus. When the stimulus information name is set in the timing data, the stimulus information name is set in the stimulus information for each device instead of the type of the stimulus. When the type of the stimulus output device is an auditory device such as a speaker, the stimulus information is crack sound voice data. The crack sound voice data is data of the voice of the crack sound or the sound representing the generation of the crack sound. When the stimulus output device is a tactile device such as a vibrator, the stimulus information is vibration pattern data. The vibration pattern data indicates either or a combination of the vibration pattern represented by the timing of the start and stop of the vibration and the frequency of the vibration. When the stimulus output device is a light emitting device such as a lamp, the stimulus information is light emission pattern data. The light emission pattern data indicates either or a combination of the lighting pattern represented by the timing of lighting and extinguishing and the lighting color. The same stimulus information may be associated with different types of stimuli. When the sign language information output device 300 has only one type of stimulus output device, the storage unit 312 may store the stimulus information for each type of stimulus instead of the stimulus information for each device.

[0043] FIG. 6 is a flowchart showing the detection condition data generation process by the sign language information transmission device 100. The sign language video distributor captures the sign language of the signer with the camera 201 and picks up the voice of the sign language with the microphone 202. The camera 201 may capture only the mouth of the signer when the signer is signing. The data input unit 101 associates the sign language video data input from the camera 201, the sign language voice data input from the microphone 202, and the identification information of the signer input by the input unit 103, and writes them to the storage unit 102 (step S105).

[0044] After a plurality of sign language video data and sign language voice data are stored in the storage unit 102, the sign language video distributor inputs, via the input unit 103, the selection of the sign language video data to be played. The playback control unit 111 reads out the selected sign language video data and the sign language voice data corresponding to the sign language video data from the storage unit 102. The playback control unit 111 displays the sign language video data on the display unit 104 and outputs the sign language voice data to the voice output unit 105 (step S110). At this time, the playback control unit 111 synchronizes the display of the video frame of the sign language video data and the output of the voice of the sign language voice data at the same time.

[0045] When the sign language video distributor detects the generation of a stimulus such as the utterance of a plosive sound based on the video of the sign language video data or the voice of the sign language voice data, the distributor inputs, via the input unit 103, the generation of the stimulus and the type of the generated stimulus (step S115). The input of the type of the stimulus may also serve as the input of the generation of the stimulus. In the present embodiment, as the types of the stimulus, types of plosive sounds such as "pa", "pi", "pu", "pe", and "po" are input. The playback control unit 111 acquires the time of the video frame displayed when the generation of the stimulus is input as the generation timing of the stimulus. Further, the playback control unit 111 continues to display the sign language video data and output the sign language voice data, and receives the input of the generation of the stimulus and the type of the stimulus.

[0046] When the playback control unit 111 finishes playing the sign language video data to the end or when the sign language information distributor inputs the end of playback via the input unit 103, the playback control unit 111 generates stimulus generation information associating the generation timing and the type of the stimulus. The playback control unit 111 adds the stimulus generation information to the sign language video data to generate sign language video data with timing, and adds the stimulus generation information to the sign language voice data to generate sign language voice data with timing (step S120). The playback control unit 111 writes the sign language video data with timing and the sign language voice data with timing in association with each other in the storage unit 102.

[0047] When the sign language video distributor further inputs the selection of the sign language video data to be played back through the input unit 103 (step S125: NO), the sign language information transmission device 100 repeats the process from step S110. Note that the sign language information transmission device 100 may repeat the process from step S105. Then, when the sign language video distributor inputs an analysis instruction through the input unit 103 (step S125: YES), the analysis unit 112 reads out the timed sign language video data and the timed sign language audio data from the storage unit 102. The analysis unit 112 acquires a plurality of types of time-series feature amounts from the data of the mouth video indicated by the read timed sign language video data and the timed sign language audio data. The analysis unit 112 analyzes the feature amounts obtained from the timed sign language video data and the timed sign language audio data to which the identification information of the same sign language user is assigned for each type of stimulus, and obtains the feature amounts representing the occurrence of the stimulus for each sign language user for each type of stimulus (step S130).

[0048] For the acquisition of the feature amount representing the generation of the plosive sound, for example, any existing technology can be used. For example, when combining and using the feature amount obtained from the timed sign language video data and the feature amount obtained from the timed sign language audio data, the technology of Reference 2 can be used. In Reference 2, as the feature amount, the coordinates of the time-series feature points around the mouth obtained from the timed sign language video data and the MFCC (Mel-Frequency Cepstrum Coefficient) obtained from the timed sign language audio data are used. Also, when using the feature amount obtained from the timed sign language audio data, the technology of Reference 3 can be used.

[0049] (Reference 2) Yasuo Ariki et al., "Speech Recognition Integrating Image and Voice Information", Information Processing, Vol52, No.1, 2011, p.87-94

[0050] (Reference 3) Guomin Wang et al., "Acoustic Characteristic Evaluation of Glottal Stop", Internet <https: / / www.jstage.jst.go.jp / article / cleftpalate1976 / 16 / 1 / 16_37 / _pdf>

[0051] Also, as the feature amounts obtained from the signed speech audio data with timing, the loudness of the sound represented in decibels, the change in the loudness of the sound, the loudness of the sound for each frequency, the change in the loudness of the sound for each frequency, etc. may be used. Note that the types of feature amounts are not limited to these. In the case of plosive sounds, since a sudden large sound is generated from a silent or small sound, it is conceivable that the change in the loudness of the sound is used for the detection of the plosive sound utterance. Therefore, the analysis unit 112 detects the lower limit value when a plosive sound occurs in the time-series feature amounts, using the loudness of the sound as a feature amount for each signer. Also, the analysis unit 112 detects the lower limit value of the feature amount for each type of plosive sound, using the loudness of the sound for each frequency as a feature amount.

[0052] The detection condition setting unit 113 generates detection condition data for each signer based on the detection result by the analysis unit 112. Specifically, the sign language video distributor inputs the value of the coefficient for each type of feature amount through the input unit 103. The detection condition setting unit 113 multiplies the lower limit value of each feature amount detected by the analysis unit 112 by the value of the coefficient input for the type of that feature amount to obtain a threshold value. The detection condition setting unit 113 associates the identification information of the signer, the type of plosive sound, and the detection condition indicating the threshold value of each feature amount calculated for that type of plosive sound, and generates detection condition data, which is written into the storage unit 102 (step S135). Note that the detection condition setting unit 113 may use the feature amounts detected by the analysis unit 112 as the detection condition data as they are. Also, the detection condition setting unit 113 may set the threshold value of the feature amount input by the sign language video distributor as the detection condition data.

[0053] After the sign language information distribution process shown in FIG. 7 described later is performed, the sign language video distributor may rewrite the detection condition data. For example, the sign language video distributor inputs, via the input unit 103, information specifying the detection condition data to be changed and the value of the coefficient after the change. The detection condition setting unit 113 rewrites the threshold value indicated by the detection condition data to be changed to a new threshold value calculated by multiplying the lower limit value of the feature amount detected by the analysis unit 112 by the value of the coefficient after the change. Alternatively, the sign language video distributor may input, via the input unit 103, information specifying the detection condition data to be changed and the threshold value of the feature amount after the change. The detection condition setting unit 113 rewrites the threshold value indicated by the detection condition data to be changed to the input threshold value.

[0054] FIG. 7 is a flowchart showing the sign language information distribution process by the sign language information transmission device 100. The sign language video distributor inputs, via the input unit 103, the identification information of the sign language user. Further, the sign language video distributor captures the sign language of the sign language user with the camera 201 and picks up the voice of the sign language with the microphone 202. The data input unit 101 inputs sign language video data from the camera 201 and sign language voice data from the microphone 202 (step S205).

[0055] The detection unit 114 acquires a feature amount from the sign language video data and the sign language voice data input in step S205. When the detection unit 114 determines that the acquired feature amount satisfies the detection condition described in any of the detection condition data specified by the input identification information of the sign language user, the detection unit 114 detects that a stimulus of the type associated with the detection condition data has occurred (step S210). For example, when the threshold value of the feature amount is described in the detection condition data, the detection unit 114 determines that the detection condition is satisfied when the acquired feature amount exceeds the threshold value. The detection unit 114 outputs to the distribution data generation unit 115 information associating the type of stimulus with the generation timing indicating the time when the feature amount in which the occurrence of the stimulus is detected in the sign language video data or the sign language voice data is obtained.

[0056] When the distribution data generation unit 115 receives information on the type and generation timing of the stimulus from the detection unit 114, it generates sign language information distribution data (step S215). First, the distribution data generation unit 115 sets the sign language video data display position, the data name of the distribution sign language video data generated from the sign language video data input in step S205, the sign language video playback timing, the data name of the distribution sign language audio data generated from the sign language audio data input in step S205, and the sign language audio playback timing in the sign language information distribution data. The sign language video data display position, the sign language video playback timing, and the sign language audio playback timing are set based on information previously input by the sign language video distributor using the input unit 103, for example. The sign language audio playback timing may be a relative time based on a predetermined video frame or a presentation time represented by UTC. Note that the sign language video playback timing and the sign language audio playback timing are set so that the video frames of the sign language video data and the audio of the sign language audio data are output in synchronization.

[0057] The distribution data generation unit 115 further sets timing data including information associating the type of stimulus received from the detection unit 114 with the stimulus reproduction timing indicating the generation timing received from the detection unit 114 in the sign language information distribution data. When the information on the time set in the sign language video data used by the detection unit 114 for stimulus detection is different from the information on the time set in the sign language video data for distribution, the distribution data generation unit 115 converts the generation timing detected by the detection unit 114 into the information on the time at the time of reproduction of the sign language video data for distribution and sets it as the stimulus reproduction timing. For example, assume that UTC is used for the time information of the sign language video data input in step S205, the first video frame is at time a, and the generation timing is at time b. When relative time with the first video frame set to 0 is used for the time information of the sign language video data for distribution, the distribution data generation unit 115 sets time (b - a) as the stimulus reproduction timing. Also, when UTC time c is set for the sign language video reproduction timing, the distribution data generation unit 115 may set time (c + b - a) as the stimulus reproduction timing. Further, the distribution data generation unit 115 may set the stimulus information name in the timing data instead of the information on the type of stimulus, and may further set the information on the type of stimulus output device corresponding to the type of stimulus in the timing data. In this case, the stimulus information name and the information on the type of stimulus output device corresponding to the type of stimulus are stored in the storage unit 102 in advance.

[0058] The playback control unit 111 displays on the display unit 104 the sign language video data name of the sign language video data described in the sign language information distribution data and the information on the time of the video frame being played back, and outputs the sign language audio data with the sign language audio data name described in the sign language information distribution data from the audio output unit 105. Further, the correction unit 116 displays the timing data included in the sign language information distribution data on the display unit 104 (step S220). The sign language video distributor inputs a correction instruction for the timing data as needed using the input unit 103. The correction unit 116 rewrites the timing data according to the input correction instruction (step S225). The transmission unit 117 transmits the sign language video data, the sign language audio data, and the sign language information distribution data to the sign language information output device 300 (step S230).

[0059] Note that the sign language video distributor may manually set all of the timing data. In that case, the sign language information transmission device 100 does not have to perform the processing of FIG. 6 and the processing of step S210 of FIG. 7. In step S215, the distribution data generation unit 115 generates sign language information distribution data excluding the timing data. In step S220, the sign language video distributor inputs the timing data using the input unit 103, and the correction unit 116 sets the input timing data in the sign language information distribution data.

[0060] FIG. 8 is a flowchart showing the sign language information output process by the sign language information output device 300. The receiving unit 301 of the sign language information output device 300 receives the sign language video data, the sign language audio data, and the sign language information distribution data transmitted from the sign language information transmission device 100 and writes them in the storage unit 302 (step S305). The playback control unit 306 starts outputting video frames so that the sign language video data specified by the sign language video data name is displayed at the video display position at the sign language video playback time based on the sign language information distribution data, and starts outputting the sign language audio data specified by the sign language audio data name at the sign language audio generation time. The delay addition unit 307 delays by a predetermined time and causes the video frames output by the playback control unit 306 to be displayed on the display unit 303 and the sign language audio data output by the playback control unit 306 to be output from the audio output unit 304 (step S310).

[0061] The output control unit 308 detects that the stimulation reproduction timing described in the timing data has been reached (step S315). The output control unit 308 reads out the type of stimulation or the stimulation information name corresponding to the detected stimulation reproduction timing from the timing data. The output control unit 308 reads out stimulation information corresponding to the type of the stimulation output device indicated by the used device information stored in the storage unit 302 and the type of the read stimulation or the stimulation information name from the stimulation information for each device. Note that when the type of the stimulation output device is set in the timing data in association with the stimulation reproduction timing, the output control unit 308 reads out the stimulation information when the type of the stimulation output device set in the timing data matches the type of the stimulation output device indicated by the used device information stored in the storage unit 302. The output control unit 308 outputs a stimulation based on the stimulation information read out for the type of the stimulation output device from the voice output unit 304, which is the stimulation output device of the type indicated by the used device information, or from the output device 309 (step S320).

[0062] The reproduction control unit 306 determines whether or not the reproduction of the sign language video data has ended (step S325). When the reproduction control unit 306 determines that the reproduction has not ended (step S325: NO), the sign language information output device 300 repeats the process from step S315. When the reproduction control unit 306 determines that the reproduction has ended (step S325: YES), the sign language information output device 300 ends the process of FIG. 8.

[0063] Note that the sign language information output device 300 may rewrite the used device information stored in the storage unit 302 based on the information input by the viewer through the input unit 305. Further, the sign language information output device 300 may rewrite the stimulation information based on the information input by the viewer through the input unit 305. Further, the sign language information output device 300 may replace the stimulation information set in the stimulation information for each device with other pre-prepared stimulation information based on the information input by the viewer through the input unit 305.

[0064] Note that the sign language information transmission device 100 may not have the detection unit 114 and the correction unit 116, and the sign language information output device 300 may have the detection unit 114. In this case, in advance, the transmission unit 117 of the sign language information transmission device 100 transmits the detection condition data for each type of stimulus corresponding to the signer to the sign language information output device 300, and the sign language information output device 300 stores the received detection condition data in the storage unit 302. The distribution data generation unit 115 distributes the sign language video data, the sign language audio data, and the sign language information distribution data excluding the timing data to the sign language information output device 300. The detection unit 114 of the sign language information output device 300 acquires a feature amount from the sign language video data and the sign language audio data received from the sign language information transmission device 100, and when it is determined that the acquired feature amount satisfies the detection condition described in any of the detection condition data, it detects that a stimulus of the type associated with the detection condition data has occurred. The detection unit 114 generates timing data associating the type of stimulus with the stimulus reproduction timing indicating the time when the feature amount from which the occurrence of the stimulus is detected in the sign language video data or the sign language audio data is obtained, and writes it to the storage unit 302. The output control unit 308 performs the same processing as above using the timing data generated by the detection unit 114. Further, the sign language information output device 300 may further include a detection condition setting unit 113. The detection condition setting unit 113 rewrites the detection condition indicated by the detection condition data stored in the storage unit 302 based on the coefficient or threshold value input by the viewer through the input unit 305.

[0065] As described above, the sign language information output device 300 synchronizes the output of the sign language video and audio with the output of the stimulus by adding a delay to the sign language video data and the sign language audio data and outputting them. This is because when comparing the video with a stimulus such as vibration, it is usually considered that the video is reproduced earlier. By adding a delay to the sign language video data and the sign language audio data and outputting them, the delay addition unit 307 can reproduce the output timing of the stimulus in the original sign language as faithfully as possible. Note that the sign language information output device 300 may receive control information such as the delay time and whether to operate the delay addition unit 307 from the sign language information transmission device 100, or may be input by the viewer through the input unit 305.

[0066] [Second Embodiment] In the first embodiment, real-shot video was used as the sign language video data. In the second embodiment, CG video is used as the sign language video data. The second embodiment will be described centering on the differences from the first embodiment.

[0067] The configuration of the sign language information transmission device of this embodiment is the same as that of the sign language information transmission device 100 of the first embodiment shown in FIG. 2. However, the data input unit 101 inputs, in step S105 of FIG. 6 and step S205 of FIG. 7, CG sign language video data and sign language voice data output in synchronization with the sign language video data. The input may be received from another device connected to the sign language information transmission device 100 or read from a recording medium. When the sign language video data is sign language CG, the sign language video data is described by motion data. The motion data is generally recorded in the BVH (Biovision Hierarchy) format represented by the rotation angles of the joints of the whole body. Therefore, as the feature amount obtained by the analysis unit 112 from the sign language video data with timing and the feature amount obtained by the detection unit 114 from the sign language video data, the rotation angles indicated by the motion data for generating the facial expression of the head are used. Except for these points, the sign language information transmission device 100 operates in the same manner as in the first embodiment. Also, the configuration and operation of the sign language information output device 300 are the same as in the first embodiment.

[0068] [Third Embodiment] In the third embodiment, the device that outputs the sign language video data and the sign language voice data and the device that outputs the stimulus are different devices.

[0069] FIG. 9 is a functional block diagram showing the configuration of the sign language information output device 300a. In FIG. 9, only the functional blocks related to the present embodiment are extracted and shown. The sign language information transmission system 1 may include the sign language information output device 300a shown in FIG. 9 instead of some or all of the sign language information output devices 300. The sign language information output device 300a includes a sign language video output device 310 and a stimulus output device 320. For example, the sign language video output device 310 is a television receiver or a personal computer, and the stimulus output device 320 is a smartphone, a tablet terminal, or a dedicated device for stimulus output. The sign language video output device 310 and the stimulus output device 320 are synchronized in time. The transmission unit 117 of the sign language information transmission device 100 transmits sign language video data, sign language audio data, and sign language information distribution data to the sign language video output device 310 by broadcasting or communication, and transmits timing data to the stimulus output device 320 by communication. The sign language information distribution data may not have timing data set therein. The transmission unit 117 of the sign language information transmission device 100 may transmit sign language information distribution data with timing data set thereto to the stimulus output device 320. For the sign language video playback timing set in the sign language information distribution data and the stimulus playback timing set in the timing data, for example, the UTC time is used.

[0070] The sign language video output device 310 includes a receiving unit 311, a storage unit 312, a display unit 313, an audio output unit 314, an input unit 315, a playback control unit 316, a delay addition unit 317, and a communication unit 318. The receiving unit 311, the display unit 313, the audio output unit 314, the input unit 315, and the delay addition unit 317 each have the same functions as the receiving unit 301, the display unit 303, the audio output unit 304, the input unit 305, and the delay addition unit 317 of the sign language information output device 300 of the first embodiment shown in FIG. 3. The storage unit 312 stores sign language video data, sign language audio data, and sign language information distribution data. In addition to the same function as the playback control unit 306 of the first embodiment shown in FIG. 3, the playback control unit 316 has a function of notifying the stimulus output device 320 of the start and end of the playback of the sign language video data. The communication unit 318 communicates with the stimulus output device 340 wirelessly or by wire.

[0071] The stimulation output device 320 includes a receiving unit 321, a storage unit 322, a communication unit 323, an output control unit 324, an output device 325, and an input unit 326. The output control unit 324 and the input unit 326 each have the same functions as the output control unit 308 and the input unit 305 of the sign language information output device 300 of the first embodiment shown in FIG. 3. The receiving unit 321 receives timing data from the sign language information transmission device 100 and writes it into the storage unit 322. The storage unit 322 stores the timing data, device-specific stimulation information, and used device information. When there is one used device, the storage unit 322 may store stimulation information associated with the type of stimulation or the name of the stimulation information instead of the device-specific stimulation information. The communication unit 323 communicates with the sign language video output device 310 wirelessly or by wire. The output device 325 is a voice output device, a vibration device, or a light emitting device. The light emitting device may include a plurality of output devices 325 of different types or the same type.

[0072] FIG. 10 is a flowchart showing the sign language information output process by the sign language information output device 300a. The receiving unit 311 of the sign language video output device 310 receives the sign language video data, sign language audio data, and sign language information distribution data transmitted from the sign language information transmission device 100 and writes them into the storage unit 312. The receiving unit 321 of the stimulation output device 320 writes the timing data transmitted from the sign language information transmission device 100 into the storage unit 322 (step S405). The playback control unit 316 of the sign language video output device 310 notifies the stimulation output device 320 to start playing the sign language video data at the sign language video playback time indicated by the sign language information distribution data (step S410). The playback control unit 316 starts outputting video frames based on the sign language information distribution data so that the sign language video data specified by the sign language video data name is displayed at the video display position at the sign language video playback time, and starts outputting the sign language audio data specified by the sign language audio data name at the sign language audio generation time. The delay addition unit 317 delays it by a predetermined time and causes the video frames output by the playback control unit 316 to be displayed on the display unit 313, and causes the sign language audio data output by the playback control unit 316 to be output from the audio output unit 314 (step S415).

[0073] After receiving the notification of the start of playback of the sign language video data, the output control unit 324 of the stimulation output device 320 detects that the stimulation playback timing described in the timing data has been reached (step S420). The output control unit 324 performs the same processing as the output control unit 308 in FIG. 8, and outputs a stimulation based on the stimulation information from the output device 325 (step S425).

[0074] The playback control unit 316 of the sign language video output device 310 determines whether or not the playback of the sign language video data has ended (step S430). If the playback control unit 316 determines that it has not ended (step S430: NO), the sign language video output device 310 and the stimulation output device 320 repeat the processing from step S420. When the playback control unit 316 determines that it has ended (step S430: YES), it notifies the stimulation output device 320 of the end of the playback of the sign language video data (step S435). The sign language video output device 310 and the stimulation output device 320 end the processing of FIG. 10.

[0075] Note that the stimulation output device 320 does not necessarily have to have the receiving unit 321. In this case, the sign language video output device 310 transmits the sign language information distribution data or the timing data received from the sign language information transmission device 100 from the communication unit 318 to the stimulation output device 320. The storage unit 322 of the stimulation output device 320 stores the sign language information distribution data or the timing data received by the communication unit 323. Also, a receiving device may be provided between the transmission network 500 and the sign language video output device 310 and the stimulation output device 320. The receiving device is set at, for example, the viewer's home. The receiving device receives the sign language video data, the sign language audio data, and the sign language information distribution data from the sign language information transmission device 100, transmits the sign language video data, the sign language audio data, and the sign language information distribution data to the sign language video output device 310, and transmits the sign language information distribution data or the timing data in the sign language information distribution data to the stimulation output device 320.

[0076] In the above-described sign language information output device 300a, the device including the stimulation output device controls the output of the stimulation, but the device that displays the sign language video data may control the output of the stimulation.

[0077] FIG. 11 is a block diagram showing the configuration of the sign language information output device 300b. The sign language information transmission system 1 may include the sign language information output device 300b shown in FIG. 11 instead of some or all of the sign language information output devices 300. In the same figure, the same parts as those of the sign language information output device 300a shown in FIG. 9 are denoted by the same reference numerals, and the description thereof is omitted. The sign language information output device 300b includes a sign language video output device 330 and a stimulus output device 340.

[0078] The difference between the sign language video output device 330 and the sign language video output device 310 shown in FIG. 9 is that the sign language video output device 330 includes a storage unit 331 instead of the storage unit 312 and further includes an output control unit 332. The storage unit 331 stores sign language video data, sign language voice data, sign language information distribution data, usage device information, and device-specific stimulus information transmitted from the sign language information transmission device 100, in the same manner as the storage unit 302 of the sign language information output device 300 shown in FIG. 3. The output control unit 332 performs the same processing as the output control unit 308 of the sign language information output device 300 in the first embodiment. However, the output control unit 332 transmits a stimulus output instruction for instructing the output of the stimulus indicated by the stimulus information from the communication unit 318 to the stimulus output device 340.

[0079] The stimulus output device 340 includes a communication unit 323, an output control unit 341, and an output device 325. The stimulus output device 340 may include a plurality of output devices 325 of different types or the same type. The output control unit 341 controls the output device 325 to output a stimulus according to the stimulus output instruction received from the sign language video output device 330. Since the output control unit 341 receives the stimulus output instruction and controls the output device 325 to output the stimulus information in real time, the times of the sign language video output device 330 and the stimulus output device 340 do not have to be synchronized.

[0080] The sign language information transmission device 100 operates in the same manner as in the first embodiment. The sign language information output device 300b performs the same processing as the sign language information output processing of the first embodiment shown in FIG. 8, except for the processing in step S320. That is, in steps S305 to S315, the receiving unit 311, the reproduction control unit 316, the delay addition unit 317, and the output control unit 332 of the sign language video output device 330 perform the same processing as the receiving unit 301, the reproduction control unit 306, the delay addition unit 307, and the output control unit 308 of the sign language information output device 300. In step S315, the output control unit 332 detects that the stimulation reproduction timing described in the timing data has been reached.

[0081] In step S320, the output control unit 332 performs the same processing as the output control unit 308 of the first embodiment, and reads out the stimulation information corresponding to the type of the stimulation output device indicated by the used device information and the type of the stimulation or the stimulation information name read from the timing data from the device-specific stimulation information. The output control unit 332 transmits a stimulation output instruction to the stimulation output device 340 to instruct the output device 325, which is a stimulation output device of the type indicated by the used device information, to output a stimulation based on the stimulation information read for the type of the stimulation output device. The output control unit 341 of the stimulation output device 340 outputs a stimulation from the output device 325 according to the instruction of the stimulation output instruction received from the sign language video output device 330.

[0082] In step S325, when the reproduction control unit 316 of the sign language video output device 330 determines that the reproduction of the sign language video data has not ended, the processing from step S315 is repeated, and when it determines that the reproduction has ended, the processing in FIG. 8 is ended.

[0083] Note that the output control unit 332 of the sign language video output device 330 may transmit a stimulation output instruction in which the type of the stimulation or the stimulation information name is set. In this case, the output control unit 332 of the stimulation output device 340 stores in advance the stimulation information corresponding to the type of the stimulation or the stimulation information name. The output control unit 332 outputs a stimulation from the output device 325 based on the stimulation information corresponding to the type of the stimulation or the stimulation information name set in the stimulation output instruction.

[0084] According to the above embodiment, for example, the stimulation output devices 320 and 340 can be dedicated devices for reproducing vibrations, and by changing the frequency of the vibrations, the differences in pa, pi, pu, pe, and po can be output. Therefore, it is possible to improve the viewers' understanding of sign language.

[0085] Further, as the sign language information output devices 300, the stimulation output devices 320 and 340, for example, generally used smartphones can be used, and stimulations can be output by the vibrators, speakers, and lamps of the smartphones. In this case, the same stimulation may be output regardless of the type of stimulation such as the type of bursting sound. The smartphone vibrates at the timing instructed by a dedicated application. When the sign language information output device 300 is a smartphone, the dedicated application realizes the functions of the reproduction control unit 306, the delay addition unit 307, and the output control unit 308. Further, when the stimulation output devices 320 and 340 are smartphones, synchronization may be performed with the sign language video output devices 310 and 330 by communication means such as wireless LAN or Bluetooth. Alternatively, when the sign language video output devices 310 and 330 are smartphones, synchronization may be performed with the stimulation output devices 320 and 340 by communication means such as wireless LAN or Bluetooth.

[0086] According to the present embodiment, the sign language information transmission system 1 can reproduce easily understandable sign language by combining the sign language video and the output of stimulation.

[0087] At least some of the functions of the sign language information transmission device 100, the sign language information output device 300, the sign language video output devices 310 and 330, and the stimulation output devices 320 and 340 in the above-described embodiments can be realized by a computer. In that case, a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Further, all or part of these functions may be realized using hardware such as an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). Here, the "computer system" is assumed to include hardware such as an OS and peripheral devices. Further, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, a magneto-optical disk, a ROM, or a CD-ROM, or a storage device such as a hard disk built into a computer system. Furthermore, the "computer-readable recording medium" also includes, like a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, a medium that dynamically holds a program for a short period of time, and a volatile memory inside a computer system that becomes a server or a client in that case and holds a program for a certain period of time. Also, the above program may be for realizing a part of the above-described functions, and may further be a program that can be realized in combination with a program already recorded in a computer system for the above-described functions.

[0088] Further, the sign language information transmission device 100 may be realized by a plurality of computer devices connected to a network. In this case, it can be arbitrary which of these plurality of computer devices realizes each functional unit of the sign language information transmission device 100. Also, the same functional unit of the sign language information transmission device 100 may be realized by a plurality of computer devices.

[0089] FIG. 12 is a diagram showing the hardware configuration of each device of the sign language information transmission device 100 described above. The sign language information transmission device 100 includes a processor 701, a storage unit 702, a communication interface 703, and a user interface 704. The processor 701 is a central processing unit that performs operations and controls. The processor 701 is, for example, a CPU (central processing unit). The processor 701 reads a program from the storage unit 702 and executes it. The storage unit 702 further has a work area and the like when the processor 701 executes various programs. The storage unit 102 is realized by the storage unit 702. The communication interface 703 is connected to be communicable with other devices. The data input unit 101 and the transmission unit 117 are realized by the communication interface 703. The user interface 704 is an input device such as a button, a keyboard, a pointing device, and a display device such as a lamp and a display. Also, an artificial operation is input through the user interface 704. The input unit 103, the display unit 104, and the voice output unit 105 are realized by the user interface 704.

[0090] All or part of the functions of the playback control unit 111, the analysis unit 112, the detection condition setting unit 113, the detection unit 114, the distribution data generation unit 115, the correction unit 116, and the transmission unit 117 are realized by the processor 701 reading a program from the storage unit 702 and executing it. Note that all or part of these functions may be realized using hardware such as an ASIC or a PLD.

[0091] The hardware configuration of the sign language information output device 300 is the same as that in FIG. 12. The receiving unit 301 is realized by the communication interface 703. The storage unit 302 is realized by the storage unit 702. The display unit 303, the audio output unit 304, the input unit 305, and the output device 309 are realized by the user interface 704. All or part of the functions of the playback control unit 306, the delay addition unit 307, and the output control unit 308 are realized by the processor 701 reading a program from the storage unit 702 and executing it. Note that all or part of these functions may be realized using hardware such as an ASIC, a PLD, or an FPGA.

[0092] The hardware configurations of the sign language video output devices 310 and 330 and the stimulus output devices 320 and 340 are also the same as that in FIG. 12.

[0093] According to the above-described embodiment, the sign language information transmission system includes a sign language information transmission device and a sign language information output device. The sign language information transmission device includes a transmission unit. The transmission unit transmits video data of sign language and timing data indicating the timing at which a predetermined type of auditory or tactile stimulus occurs in the sign language indicated by the video data. The sign language information transmission device may further include a detection unit. The detection unit detects the timing at which a predetermined type of auditory or tactile stimulus occurs in the sign language based on one or both of the feature amount obtained from the video data of the sign language and the feature amount obtained from the audio data of the sign language. The transmission unit transmits the video data and the timing data indicating the timing detected by the detection unit.

[0094] The sign language information output device includes a first receiving unit, a second receiving unit, a playback control unit, and an output control unit. The first receiving unit receives video data from the sign language information transmission device. The second receiving unit receives timing data from the sign language information transmission device. The playback control unit plays back the video data. The playback control unit and the output control unit control the device to output a vibration, light, or sound stimulus at the timing indicated by the timing data.

[0095] A stimulus of a specified type is, for example, the pronunciation of a plosive sound. The detection unit of the sign language information transmission device detects the timing when the plosive sound is pronounced and the type of the plosive sound. The transmission unit transmits timing data indicating the timing detected by the detection unit and the type of the plosive sound. The output control unit of the sign language information output device controls the device to output a stimulus of vibration, light, or sound corresponding to the type of the plosive sound at the timing indicated by the timing data.

[0096] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and designs and the like within the scope not departing from the gist of the present invention are also included.

Description of Reference Numerals

[0097] 1 Sign language information transmission system 100 Sign language information transmission device 101 Data input unit 102, 302, 312, 322, 331, 702 Storage unit 103, 305, 315, 326 Input unit 104, 303, 313 Display unit 105, 304, 314 Voice output unit 111, 306, 316 Reproduction control unit 112 Analysis unit 113 Detection condition setting unit 114 Detection unit 115 Distribution data generation unit 116 Correction unit 117 Transmission unit 201 Camera 202 Microphone 300, 300a, 300b Sign language information output device 301, 311, 321 Reception unit 307, 317 Delay addition unit 308, 324, 332, 341 Output control unit 309, 325 Output device 310, 330 Sign language video output device 318, 323 Communication unit 320, 340 Stimulus output device 500 Transmission Network 701 Processor 703 Communication Interface 704 User Interface

Claims

1. Based on the feature amount obtained from the sign language video data and the feature amount obtained from the sign language audio data, or based on the feature amount obtained from the audio data, a detection unit that detects the timing at which a predetermined type of auditory or tactile stimulus occurs in the sign language shown in the video data; A transmission unit that transmits the video data and timing data indicating the timing detected by the detection unit; comprising: The predetermined type of stimulus is the utterance of a plosive sound, The detection unit detects the timing at which the plosive sound is uttered and the type of the plosive sound, The transmission unit transmits the timing data indicating the timing and the type of the plosive sound detected by the detection unit, A sign language information transmission device.

2. A first reception unit that receives sign language video data; A second reception unit that receives timing data indicating the timing at which a predetermined type of auditory or tactile stimulus occurs in the sign language shown in the video data; A reproduction control unit that reproduces the video data; An output control unit that controls a device to output a stimulus of vibration, light, or sound at the timing indicated by the timing data; comprising: The stimulus is the utterance of a plosive sound, The timing data includes the timing at which the plosive sound is uttered and information on the type of the plosive sound, The output control unit controls the device to output a stimulus of vibration, light, or sound corresponding to the type of the plosive sound at the timing, A sign language information output device.

3. A sign language information transmission system having a sign language information transmission device and a sign language information output device, The sign language information transmission device is A transmitting unit that transmits sign language video data and timing data indicating the timing at which a predetermined type of auditory or tactile stimulus occurs in the sign language indicated by the video data. The sign language information output device A first receiving unit that receives the video data, A second receiving unit that receives the timing data, A playback control unit that plays back the video data, An output control unit that controls a device to output a vibration, light, or sound stimulus at the timing indicated by the timing data. The stimulus is the utterance of a popping sound. The timing data includes the timing at which the popping sound is uttered and information on the type of the popping sound. The output control unit controls the device to output a vibration, light, or sound stimulus corresponding to the type of the popping sound at the timing. A sign language information transmission system.

4. A program for causing a computer to function as the sign language information transmission device according to Claim 1.

5. A program for causing a computer to function as the sign language information output device according to Claim 2.

Citation Information

Patent Citations

  • Lip animation synthesizing method and device therefor

    JP1997265253A

  • Voice indication system and voice indication program

    JP2002244841A

  • Interactive system for hearing-impaired person

    JP2003296753A

  • Voice visualization method and storage medium storing the same

    JP2005209000A

  • Sign language CG interpretation and edit system, and program

    JP2019197084A