Subtitle Prompt Control Device and Program
The subtitle presentation control device addresses the synchronization challenge between sign language animations and subtitle texts by determining split points in subtitle texts and synchronizing their display with sign language video progress, ensuring accurate and synchronized content presentation.
Patent Information
- Application Number
- JP2021130847
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-10
- Publication Date
- 2025-06-16
- Estimated Expiration
- 2041-08-10
AI Technical Summary
Existing technologies face challenges in synchronizing the display of sign language animations and subtitle texts, leading to deviations in content due to the inability to accurately estimate the length of sign language videos from subtitle text lengths.
A subtitle presentation control device that supplies sign language label sequences and corresponding subtitle texts, determines split points within the subtitle text based on display area size, and generates motion data for sign language animations. The device synchronizes the display of subtitle texts with the progress of sign language videos by controlling the switching of subtitle text displays based on determined time positions within the motion data.
This solution enables appropriate synchronization of sign language videos and subtitle texts, preventing content deviations and ensuring that the content of the sign language video matches the subtitle text at all times.
Smart Images

Figure 0007692759000001 
Figure 0007692759000002 
Figure 0007692759000003
Abstract
Description
Technical Field
[0001] The present invention relates to a subtitle presentation control device and a program.
Background Art
[0002] Techniques for providing sign language information in animations generated using computer graphics (CG) have been studied and utilized. Providing information using sign language when presenting video content is mainly intended for people who cannot hear. On the other hand, there are various people who cannot hear, including those who are congenitally deaf and use sign language as their mother tongue, and those who previously communicated in spoken language and then lost their hearing. In order to ensure information for such diverse viewers, it is desirable to provide not only sign language independently but also a set of sign language and text subtitles.
[0003] In the generation of sign language CG animations according to the prior art, a method is used in which, in response to the input information, motion capture is performed in advance and sign language motion data stored in a database is combined to reproduce the sign language motion with a CG avatar. As input information for generating sign language animations, either arbitrary text or fixed-form text is used. Arbitrary text is, for example, any Japanese text. Fixed-form text is a fixed text for distributing information such as weather information, stock price information, and sports game information.
[0004] In order to generate a sign language animation corresponding to arbitrary text, a method of translating the arbitrary text into a sign language label sequence is used. Translation from arbitrary text to a sign language label sequence can be realized using existing techniques. Translation from arbitrary text to a sign language label sequence can be performed using a translation model (such as a neural network) learned by machine learning. Motion data corresponding to the sign language labels constituting the sign language label sequence can be stored in a database in advance.
[0005] In order to generate sign language animations corresponding to fixed phrases, a method is adopted in which motion data for sentence units corresponding to fixed phrases are prepared in advance, and the motion data for replaceable parts within the sentence (proper nouns, nouns representing the weather, numerical values, etc.) are replaced.
[0006] As a prior art, there exists a technology for simultaneously displaying sign language animations and subtitle texts. For example, in FIGS. 1 and paragraphs 0034 - 0035 of Patent Document 1, a composite display unit (14) is described which synthesizes the generated Japanese subtitles and the generated sign language animations and displays them on a display. Also, in FIG. 5 etc. of Patent Document 1, examples of screens displaying sign language animations with subtitles are described.
Prior Art Documents
Patent Documents
[0007]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0008] When displaying sign language animations and subtitle texts, the temporal synchronization between the two becomes an issue. In order to display sign language animations and subtitle texts within a screen of a predetermined size, the size of the area for subtitles within the screen is limited. Therefore, when the subtitle text is long, methods such as reducing the size of the displayed font or splitting the text in the middle and switching the display become necessary. In the method of reducing the font size, the problem occurs that the longer the subtitle text becomes, the more difficult it is to read, which is not practical. In the method of splitting the text and switching the display (existing technology), since the display is switched at a fixed time regardless of the length of the sign language, there is a problem that the content of the displayed sign language and the content of the subtitles cannot be synchronized.
[0009] That is, since the length of the sign language video changes depending on the input information, it was not possible to switch the subtitles at an appropriate timing without any sense of discomfort for any input information. Since natural language written as text (such as Japanese) and sign language are completely different languages, it is also impossible to estimate the length of the sign language video from the length of the subtitle text. As a result, a deviation in content occurs between the sign language video and the subtitle being displayed at that time. For example, inconveniences such as the sign language still conveying the content of the previous subtitle display even after the display of the subtitle text has switched, or the sign language video of the next content being displayed before the display of the subtitle text has switched, occur.
[0010] The present invention has been made based on the above recognition of the problem, and aims to provide a subtitle presentation control device and a program capable of appropriately synchronizing and displaying a sign language video and subtitle text. [Means for Solving the Problem]
[0011] [1] To solve the above problems, a subtitle presentation control device according to an aspect of the present invention includes a data supply unit that supplies a sign language label sequence that is a sequence of sign language labels representing sign language, and subtitle text corresponding to the sign language label sequence; a split point determination unit that obtains a word corresponding to a split point, which is a position of switching of display within the subtitle text, when the subtitle text is displayed in the subtitle text display area based on the subtitle text and the size of the subtitle text display area for displaying the subtitle text; a sign language motion generation unit that generates motion data corresponding to the sign language label sequence by connecting motions corresponding to each of the sign language labels based on the sign language label sequence, and determines a time position within the motion data of the motion of the sign language label corresponding to the word corresponding to the split point obtained by the split point determination unit; and a subtitle presentation unit that controls to switch the display of the subtitle text to the next display at a timing based on the time position determined by the sign language motion generation unit when the subtitle text is displayed in the subtitle text display area and a sign language video based on the motion data is displayed.
[0012] [2] Further, in one aspect of the present invention, in the above-described subtitle presentation control device, a sign language video generation unit that generates and reproduces and outputs the sign language video based on the motion data and notifies the subtitle presentation unit of the current reproduction position of the sign language video is further provided, and the subtitle presentation unit controls the switching of the display of the subtitle text based on the notified current reproduction position and the time position.
[0013] [3] Further, in one aspect of the present invention, in the above-described subtitle presentation control device, the subtitle presentation unit receives a notification of the timing of the start of reproduction of the sign language video and controls the switching of the display of the subtitle text based on the timing of the start of reproduction and the time position.
[0014] [4] Further, in one aspect of the present invention, in the above-described subtitle presentation control device, the data supply unit supplies the subtitle text acquired from the outside and the sign language label sequence obtained by machine-translating the subtitle text.
[0015] [5] Further, in one aspect of the present invention, in the above-described subtitle presentation control device, the data supply unit supplies the subtitle text obtained by inserting the embedded data received from the outside into the subtitle text template stored in advance, and the sign language label sequence obtained by inserting the embedded label corresponding to the embedded data into the sign language label sequence template stored in advance corresponding to the subtitle text template.
[0016] [6] Further, one aspect of the present invention is a data supply unit that supplies a sign language label sequence that is a sequence of sign language labels representing sign language, and caption text corresponding to the sign language label sequence, and based on the caption text and the size of the caption text display area for displaying the caption text, a division point determination unit that obtains a word corresponding to a division point that is a position of switching of display within the caption text when the caption text is displayed in the caption text display area, and by connecting motions corresponding to each of the sign language labels based on the sign language label sequence, generates motion data corresponding to the sign language label sequence, and a sign language motion generation unit that determines a time position in the motion data of the motion of the sign language label corresponding to the word corresponding to the division point obtained by the division point determination unit, and a caption presentation unit that controls to switch the display of the caption text to the next display at a timing based on the time position determined by the sign language motion generation unit when the caption text is displayed in the caption text display area and a sign language video based on the motion data is displayed, and is a program for causing a computer to function as a caption presentation control device.
Effects of the Invention
[0017] According to the present invention, it is possible to switch the caption text at an appropriate timing synchronized with the progress of the sign language video.
Brief Description of the Drawings
[0018]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Figure 12
Figure 13
Embodiments for Carrying Out the Invention
[0019] Next, an embodiment of the present invention will be described with reference to the drawings. The subtitle presentation control device 1 of this embodiment determines a point (a split point on the text) for switching the display of subtitle text based on the size of the area where the subtitle text is to be displayed. The subtitle presentation control device 1 determines the timing for switching the subtitle text based on the determined split point. That is, thereby, it is possible to synchronize the presentation of sign language CG (computer graphics) animation and the timing of switching the subtitle text. That is, in this embodiment, the subtitle presentation control device 1 detects the words at the split point on the text from the count of the number of characters in the subtitle text and the size of the display area of the subtitle text. Further, the subtitle presentation control device 1 obtains the position (which frame number of the video it corresponds to) where the motion of the sign language label corresponding to the word is presented. Thereby, the subtitle presentation control device 1 performs control to switch the display of the subtitle text at the obtained position (timing).
[0020] The above control of the switching timing can be automatically performed according to the input data. That is, in this embodiment, prior manual setting of the switching time and the like are not required.
[0021] In the conventional technology, without considering the timing of presenting (playing) the sign language animation, for example, the display of the subtitle text was switched at a fixed time. For this reason, in the conventional technology, there could be a situation where the content of the sign language being displayed at a certain point in time did not match the content of the subtitle text. In contrast, in this embodiment, such a mismatch can be prevented from occurring.
[0022] Note that in order to implement the subtitle presentation control device 1 of the present embodiment, some conventional technologies can be used. For example, as a translation process from arbitrary Japanese text to a sign language label sequence for generating sign language CG animation, machine translation (existing technology) using a corpus that is parallel data of Japanese sentences and sign language label sequences can be used. This machine translation is described in Japanese Patent Application Laid-Open No. 2013-186673 and Japanese Patent Application Laid-Open No. 2014-021180. In addition, technologies for simultaneously presenting sign language CG animation and Japanese subtitles by combining fixed-form data and templates are described in Japanese Patent Application Laid-Open No. 2016-038748 (utilization of weather information data), Japanese Patent Application Laid-Open No. 2018-124934 (utilization of weather information data), Japanese Patent Application Laid-Open No. 2018-182591 (utilization of sports competition data), etc.
[0023] FIG. 1 is a block diagram showing a schematic functional configuration of the subtitle presentation control device 1 according to the present embodiment. As shown in the figure, the subtitle presentation control device 1 includes an input data supply unit 21, a segmentation point determination unit 22, a sign language motion generation unit 23, a subtitle switching timing determination unit 24, a sign language animation generation unit 25, and a subtitle presentation unit 26. Each of these functional units can be realized by, for example, a computer and a program. In addition, each functional unit has a storage means as necessary. The storage means is, for example, a variable in a program or a memory allocated by the execution of the program. Further, if necessary, non-volatile storage means such as a magnetic hard disk device or a solid state drive (SSD) may be used. Further, at least some of the functions of each functional unit may be realized as a dedicated electronic circuit instead of a program.
[0024] Note that the subtitle presentation control device 1 of this embodiment outputs while synchronizing Japanese subtitle text and Japanese sign language video (animation by computer graphics (CG)). However, the subtitle text may be text (sentences) in a language other than Japanese (natural language). Also, the sign language video may be of a sign language other than Japanese sign language. Even when the language and the type of sign language are different, the functional configuration and processing procedure of the subtitle presentation control device 1 are the same as those described below.
[0025] The input data supply unit 21 supplies a sign language label sequence, which is a sequence of sign language labels representing sign language, and Japanese text corresponding to the sign language label sequence. The Japanese text is the text for use as subtitles. Note that the input data supply unit is also simply called the "data supply unit". The detailed functional configuration of the input data supply unit 21 will be described later with reference to another figure.
[0026] The input data supply unit 21 can supply Japanese text in arbitrary sentences and a sign language label sequence, or Japanese text in fixed-form sentences and a sign language label sequence. That is, the input data supply unit 21 can supply subtitle text (arbitrary sentences) acquired from the outside and a sign language label sequence (translation result corresponding to the arbitrary sentences) obtained by machine-translating the subtitle text. Also, the input data supply unit 21 can supply subtitle text (fixed-form sentences) obtained by inserting embedded data received from the outside into a pre-stored subtitle text template, and a sign language label sequence (sign language label sequence of fixed-form sentences) obtained by inserting an embedded label corresponding to the embedded data into a pre-stored sign language label sequence template corresponding to the subtitle text template.
[0027] The segmentation point determination unit 22 determines a word corresponding to the segmentation point within the subtitle text based on the Japanese text (subtitle text) supplied by the input data supply unit 21 and the size of the subtitle text display area for displaying the subtitle text. This segmentation point is the point for appropriately dividing and displaying a subtitle text that is too long when displaying the subtitle text in the subtitle text display area. That is, the segmentation point is the position of the display switch within the subtitle text. The reason for determining a word as the segmentation point is to specify the sign language motion corresponding to the word and its frame position (time position) in subsequent processing.
[0028] In addition to the above processing, the segmentation point determination unit 22 estimates the readings of the Chinese characters included in the subtitle text. The estimated readings of the Chinese characters can be used to add ruby (furigana) to the subtitle text.
[0029] The sign language motion generation unit 23 generates motion data corresponding to the sign language label sequence based on the sign language label sequence passed from the input data supply unit 21. Specifically, the sign language motion generation unit 23 generates motion data by connecting the data of the motion corresponding to each sign language label included in the sign language label sequence.
[0030] Also, the sign language motion generation unit 23 determines the time position (frame) within the motion data corresponding to a specific word in the subtitle text in response to an inquiry from the subtitle switching timing determination unit 24. The specific word is the word corresponding to the segmentation point determined by the segmentation point determination unit 22. Specifically, the sign language motion generation unit 23 specifies the sign language label corresponding to the specified word. Then, the sign language motion generation unit 23 specifies the time position (frame) within the above motion data of the motion corresponding to the sign language label. The sign language motion generation unit 23 specifies, for example, the time position (frame) of the last part of the motion. This time position determined by the sign language motion generation unit 23 corresponds to the segmentation point of the subtitle. That is, this time position is a timing suitable for switching the display of the subtitle.
[0031] The subtitle switching timing determination unit 24 receives the subtitle text with the information of the word at the splitting point from the input data supply unit 21. Then, the subtitle switching timing determination unit 24 makes an inquiry to the sign language motion generation unit 23 using the word at the splitting point. The subtitle switching timing determination unit 24 receives the information of the time position (frame) corresponding to the word at the splitting point from the sign language motion generation unit 23, and adds the information of the time position as the information of the subtitle switching timing to the subtitle data. The subtitle switching timing determination unit 24 passes the subtitle text with the information of the subtitle switching timing added to the subtitle presentation unit 26.
[0032] Note that the subtitle switching timing determination unit 24 may also receive the information of the reading of the Chinese characters included in the subtitle text from the input data supply unit 21 and pass it to the subtitle presentation unit 26.
[0033] The sign language animation generation unit 25 receives the motion data generated by the sign language motion generation unit 23, generates and plays the sign language video corresponding to the motion data. As the sign language video, the sign language animation generation unit 25 generates the video of the animation using CG. However, the generated sign language video does not necessarily have to be the video of the animation by CG. In any case, the sign language video generated and played by the sign language animation generation unit 25 is based on the motion data generated by the sign language motion generation unit 23. And the length and time position of the video corresponding to each sign language label included in the sign language video are determined by the motion data. Note that the sign language animation generation unit 25 is also called the "sign language video generation unit".
[0034] When playing the sign language video, the sign language animation generation unit 25 notifies the subtitle presentation unit 26 of the information of the current playback position (time position, frame) at any time. Thereby, the subtitle presentation unit 26 knows the position (time position, frame) of the currently played video.
[0035] That is, the sign language animation generation unit 25 generates a sign language video (sign language CG animation) based on the motion data, plays it back and outputs it, and notifies the caption presentation unit 26 of the current playback position of the sign language video. The process of generating the sign language video by the sign language animation generation unit 25 will be described in more detail later.
[0036] The caption presentation unit 26 displays Japanese text in the caption display area in accordance with the timing of playing back the sign language animation output by the sign language animation generation unit 25. When the Japanese text is too long to fit in the caption display area at once, the caption presentation unit 26 switches the Japanese text displayed in the caption display area as time passes. The caption presentation unit 26 uses the split point determined by the split point determination unit 22 as the point for switching caption displays. The caption presentation unit 26 receives information on the current playback position (frame position) of the video from the sign language animation generation unit 25. The caption presentation unit 26 receives information on the timing of caption switching from the sign language animation generation unit 25 together with the caption text data. For this reason, the caption presentation unit 26 controls to switch the caption text to the next text and display it when the playback position of the video reaches the timing of caption switching.
[0037] That is, the caption presentation unit 26 displays the caption text in the caption text display area and controls to switch the caption text display to the next display at the timing based on the time position determined by the sign language motion generation unit 23 when the sign language video based on the motion data is displayed. That is, specifically, the caption presentation unit 26 of the present embodiment controls the switching of the display of the caption text based on the current playback position notified from the sign language animation generation unit 25 and the time position determined by the sign language motion generation unit 23.
[0038] By the caption display unit 26 performing synchronization control of the display between such sign language CG animation and caption text, the content of the sign language being displayed at a certain point in time matches the content of the caption text without deviation. That is, together with the video of the animation generated by the sign language animation generation unit 25, the caption display unit 26 displays caption text based on the supplied Japanese text.
[0039] In addition, when the Japanese text to be displayed as a caption includes Chinese characters, the caption display unit 26 may display the reading of the Chinese characters estimated by the segmentation point determination unit 22 as furigana above the Chinese characters.
[0040] In the present embodiment, for synchronization of the display between the sign language video and the caption text, information for specifying a specific frame in the video (information for specifying the time position) is used. The information for specifying the frame is represented, for example, in the format of "hh:mm:ss.nnn". Here, "hh:mm:ss" corresponds to the hour, minute, and second of the relative time starting from the start time of the video. Also, "nnn" is the frame number (001, 002,...) within the second. Note that data in other expression formats may be used to specify the frame.
[0041] FIG. 2 is a block diagram showing the internal functional configuration of the input data supply unit 21. As shown in the figure, the input data supply unit 21 includes an arbitrary sentence acquisition unit 211, a distribution data acquisition unit 212, a fixed sentence processing unit 213, a Japanese text supply unit 214, a translation unit 215, and a sign language label sequence supply unit 216. The functions of each unit are as follows.
[0042] The arbitrary sentence acquisition unit 211 acquires the text of an arbitrary Japanese sentence from the outside. The arbitrary sentence acquisition unit 211, for example, acquires the text of an arbitrary sentence written in a storage means such as a magnetic hard disk, acquires the text of an arbitrary sentence sent via a communication line, or acquires the text of an arbitrary sentence manually input from a keyboard or a touch panel. Further, the arbitrary sentence acquisition unit 211 may acquire the text of an arbitrary sentence by performing speech recognition processing on the input speech. Further, the arbitrary sentence acquisition unit 211 may acquire the text of an arbitrary sentence by other means. The arbitrary sentence acquisition unit 211 passes the acquired text of the arbitrary sentence to the Japanese text supply unit 214. This arbitrary sentence is also passed to the translation process by the translation unit 215.
[0043] The distribution data acquisition unit 212 receives data distributed from an external server device or the like. The data acquired by the distribution data acquisition unit 212 is, for example, fixed-form data such as weather information, stock price information, or sports game information. The distribution data acquisition unit 212 passes the received fixed-form data to the fixed-form sentence processing unit 213.
[0044] The fixed-form sentence processing unit 213 generates and outputs a Japanese text and a sign language label sequence using the data received by the distribution data acquisition unit 212 and a template of a fixed-form sentence prepared in advance. That is, the data received by the distribution data acquisition unit 212 is fixed-form data, and the fixed-form sentence processing unit 213 inserts words into the template of the Japanese text or inserts sign language labels into the template of the sign language label sequence corresponding to the data. The fixed-form sentence processing unit 213 passes the Japanese text generated based on the template to the Japanese text supply unit 214. Further, the fixed-form sentence processing unit 213 passes the sign language label sequence generated based on the template to the sign language label sequence supply unit 216.
[0045] Note that as a technique itself for generating a fixed-form sentence by combining fixed-form distribution data and a template, an existing technique can be used. Such a technique is also described in the following documents. References (fixed-form sentence generation): Japanese Patent Application Laid-Open No. 2016-38748, Japanese Patent Application Laid-Open No. 2018-124934, Japanese Patent Application Laid-Open No. 2018-182591
[0046] The Japanese text supply unit 214 supplies the Japanese text to the segmentation point determination unit 22. The Japanese text in the case of arbitrary sentences is passed from the arbitrary sentence acquisition unit 211. The Japanese text in the case of fixed-form sentences is passed from the fixed-form sentence processing unit 213.
[0047] The translation unit 215 converts the Japanese text (sentence) passed from the Japanese text supply unit 214 into a sign language label sequence. The translation unit 215 performs translation processing based on an example corpus that is a set of pairs of the Japanese text on the input side and the sign language label sequence on the output side. The translation unit 215 performs translation processing using the statistical relationship between these input and output sides. This translation processing itself can be performed using existing technologies. The translation unit 215 is realized, for example, using a machine learning method which is an existing technology. That is, the translation unit 215 can be realized by previously training a neural network using pairs of Japanese text and sign language label sequences (translation sentence pairs) as training data. Methods of machine translation (including translation into sign language) are described, for example, in the following documents. References (machine translation): Japanese Patent Application Laid-Open No. 2013-186673, Japanese Patent Application Laid-Open No. 2014-21180
[0048] The sign language label sequence supply unit 216 supplies the sign language label sequence corresponding to the Japanese text output by the Japanese text supply unit 214 to the sign language motion generation unit 23. The sign language label sequence in the case of arbitrary sentences is output from the translation unit 215. The sign language label sequence in the case of fixed-form sentences is passed from the fixed-form sentence processing unit 213.
[0049] As described above, the Japanese text output by the Japanese text supply unit 214 and the sign language label sequence output by the sign language label sequence supply unit 216 correspond to each other whether it is a non-fixed sentence or a fixed sentence. In other words, these Japanese texts and sign language label sequences represent the same content. In this embodiment, the input data supply unit 21 is configured to perform processing corresponding to both non-fixed sentences and fixed sentences. However, the input data supply unit 21 may be configured to process only one of non-fixed sentences or fixed sentences.
[0050] Figure 3 is a block diagram showing the internal functional configuration of the segmentation point determination unit 22. As shown in the figure, the segmentation point determination unit 22 includes a morphological analysis unit 221, a segmentation point calculation unit 222, and a text information output unit 223. The functions of each unit are as follows.
[0051] The morphological analysis unit 221 performs morphological analysis processing on the Japanese text passed from the input data supply unit 21. As a result, the morphological analysis unit 221 outputs the Japanese text (sequence of morphemes) segmented into units of morphemes. In addition, the morphological analysis unit 221 adds part-of-speech information to each morpheme. Further, the morphological analysis unit 221 estimates the readings of Chinese characters included in the Japanese text and adds the information of the reading kana to the morphemes including Chinese characters.
[0052] The morphological analysis unit 221 passes the information of the subtitle text segmented into morpheme units and the part-of-speech information of each morpheme to the segmentation point calculation unit 222. In addition, the morphological analysis unit 221 passes the information of the readings of Chinese characters to the text information output unit 223.
[0053] The morphological analysis unit 221 can be realized by using existing text analysis technologies such as KyTea (Kyoto Text Analysis Toolkit). The estimation of the Japanese readings of Chinese characters included in the results of morphological analysis can also be realized by using analysis technologies such as KyTea. KyTea is described on a web page accessible at the following URL. Note that the morphological analysis unit 221 may be realized by using analysis technologies other than KyTea. URL: http: / / www.phontron.com / kytea / index-ja.html
[0054] Based on the morpheme sequence output by the morpheme analysis unit 221, the segmentation point calculation unit 222 obtains the switching points (segmentation points) when displaying the subtitle text. Specifically, the segmentation point calculation unit 222 obtains the range that can be displayed at one time in the subtitle text display area from the subtitle text (morpheme sequence), the size of the subtitle text display area, and the information (such as size) of the character font used for display. This range is represented by the start position and end position of the subtitle text (both are the positions of characters, information representing which character). The size of the font (especially the horizontal width) may be the same size for all characters (so-called fixed-width font), or different sizes for each character. That is, the segmentation point calculation unit 222 divides the passed subtitle text at zero or more segmentation points. The segmentation point calculation unit 222 can obtain the segmentation points by simple arithmetic operations based on the information of the size of the subtitle text display area and the size of the character font. Note that the segmentation point calculation unit 222 may use the end position of the previous morpheme as the segmentation point so that the text is not split in the middle of one morpheme based on the information of the result of the morpheme analysis.
[0055] Based on the obtained segmentation points, the segmentation point calculation unit 222 adds and outputs information identifying the last word (morpheme) of the subtitle text displayed at one time to the subtitle text. When the subtitle text displayed at one time ends in the middle of a morpheme, the previous word may be treated as the last word.
[0056] The text information output unit 223 outputs text information. That is, the text information output unit 223 passes the text information to the subtitle switching timing determination unit 24. The information output by the text information output unit 223 may include the information of the subtitle text, the information obtained by splitting the subtitle text into morphemes, the information of the part-of-speech of each morpheme, and the information of the reading of Chinese characters (in the case of morphemes including Chinese characters).
[0057] Note that the information on the readings of Chinese characters (the estimated results of the readings of Chinese characters) is used to add ruby characters (furigana) to the feeling in the subsequent processing.
[0058] FIG. 4 is a block diagram showing the internal functional configuration of the sign language motion generation unit 23. As shown in the figure, the sign language motion generation unit 23 includes a sign language motion data synthesis unit 231 and a sign language motion database 232. The functions of each unit are as follows.
[0059] The sign language motion data synthesis unit 231 synthesizes sign language motion data by synthesizing the motion data read from the sign language motion database 232 based on the sign language label sequence passed from the input data supply unit 21. The sign language label sequence received by the sign language motion data synthesis unit 231 from the input data supply unit 21 is, for example, a sequence of numbers (numerical values) corresponding to individual sign language labels. This number is unique information for each sign language word. The sign language motion data synthesis unit 231 reads the motion data from the sign language motion database 232 using each number in the sign language label sequence as a key, arranges them in time series, and connects between the sign language words to generate sign language motion data corresponding to the sign language sentence. In the part where the motions between words are connected, for example, the motions are connected by using linear interpolation of the motions or the like. The length of the connection between this motion data (span length, number of frames) may be determined in advance, for example.
[0060] The sign language motion data synthesis unit 231 also returns information on the frame position in response to a query from the subtitle switching timing determination unit 24. Specifically, when receiving a query from the subtitle switching timing determination unit 24, the sign language motion data synthesis unit 231 receives information on a word that corresponds to a split point for switching the display of subtitle text. The sign language motion data synthesis unit 231 identifies the sign language label corresponding to that word. Then, the sign language motion data synthesis unit 231 returns to the subtitle switching timing determination unit 24, which is the query source, information identifying the frame corresponding to the presentation timing of the motion corresponding to that sign language label (specifically, for example, the end time of that motion). When there are two or more split points in the subtitle text, the sign language motion data synthesis unit 231 receives a query from the subtitle switching timing determination unit 24 for each of those split points and returns information on each frame position.
[0061] The word that the sign language motion data synthesis unit 231 receives from the subtitle switching timing determination unit 24 is associated with a specific sign language label in the sign language label sequence. The sign language motion data synthesis unit 231 may hold data representing this correspondence relationship in the text to be processed, or may refer to a dictionary representing the correspondence relationship between Japanese words and sign language labels. The relationship between a word in the Japanese text, a sign language label, and the start position and end position of the motion corresponding to that sign language label is as will be described later (see FIG. 11).
[0062] The sign language motion database 232 stores motion data of sign language movements corresponding to each sign language label (number for each of the above sign language labels). That is, when the sign language motion data synthesis unit 231 makes a query to the sign language motion database 232 using a specific sign language label as a key, the sign language motion database 232 can return the motion data corresponding to that sign language label. The sign language motion database 232 stores, for example, motion data obtained by motion capturing the sign language movements performed by a person and recorded in a format such as BVH (Biovision Hierarchy). The motion data corresponding to each individual sign language label has motion information including finger movements, mouth shapes, facial expressions, etc., and includes the information necessary for generating CG animations. The sign language motion database 232 is realized, for example, using a relational database management system (RDBMS).
[0063] FIG. 5 is a block diagram showing the internal functional configuration of the subtitle switching timing determination unit 24. As shown in the figure, the subtitle switching timing determination unit 24 includes a frame position query unit 241 and a switching timing information adding unit 242. The functions of each unit are as follows.
[0064] Based on the information passed from the segmentation point determination unit 22, the frame position query unit 241 identifies the words corresponding to the segmentation points in the Japanese text. The frame position query unit 241 queries for information on the position (information specifying the frame) corresponding to the word in the sign language animation by passing the information on the word at the segmentation point to the sign language motion generation unit 23. Thereby, the frame position query unit 241 can identify the frame of the sign language video corresponding to the segmentation point. This frame position information is used as the information on the switching timing when the Japanese text is displayed as subtitles. When there are two or more segmentation points in the Japanese text, the frame position query unit 241 makes a query to the sign language motion generation unit 23 for each segmentation point and obtains the frame position information.
[0065] The switching timing information adding unit 242 adds the frame position information (switching timing information) obtained by the frame position inquiry unit 241 to the subtitle text data. When there are two or more split points, the switching timing information adding unit 242 adds switching timing information for each split point. The switching timing information adding unit 242 passes the subtitle text data with the switching timing information added to the subtitle presentation unit 26.
[0066] FIG. 6 is a block diagram showing the internal functional configuration of the sign language animation generation unit 25. As shown in the figure, the sign language animation generation unit 25 includes a character data storage unit 251 and a playback unit 252. The functions of each unit are as follows.
[0067] The character data storage unit 251 stores data of a CG avatar (character). The avatar data includes data related to the appearance of the character, such as the face and clothing of the character. The character data storage unit 251 may store data of a plurality of avatars.
[0068] The playback unit 252 reads the sentence-unit sign language motion data passed from the sign language motion generation unit 23, applies the motion to the CG avatar, and generates a sign language CG animation by rendering each frame. The playback unit 252 appropriately reads the data of an appropriate avatar from the character data storage unit 251 and uses it for generating the animation. The playback unit 252 outputs the generated sign language CG animation to the outside at a predetermined frame rate. Also, during the playback of the animation video, the playback unit 252 passes information (frame position information) for specifying the currently played frame to the subtitle presentation unit 26 at any time.
[0069] Note that the playback unit 252 may output the sign language CG animation to the outside while generating it in real time, or may generate the data of the sign language CG animation corresponding to the text in advance, store it once, and then appropriately read and play the data.
[0070] FIG. 7 is a schematic diagram showing an example of subtitle text (Japanese text) to be processed by the subtitle presentation control device 1. As shown in the figure, the subtitle text in this example is "These are the starting members of both teams. The British team. Jersey number 4, forward, Fallcad Naomi, class. Jersey number 6, guard, Holland Vicky, class 2.0. Jersey number 8, center, Stanford Noras. Jersey number 10, guard, Ames David, class 1.0. Jersey number 12, forward, Brye Sophie, class 3.0. The head coach is Benson Gordon." This example is text related to sports. When this subtitle text is displayed in a predetermined character font size, it does not fit within the subtitle display area. Therefore, the subtitle presentation control device 1 divides this subtitle text at the splitting point and switches the display in accordance with the playback progress of the sign language CG animation image.
[0071] FIG. 8 is a schematic diagram showing the result of performing a morphological analysis process on the Japanese text example shown in FIG. 7. In FIG. 8, the delimiter of the morpheme is indicated by the full-width slash " / ". As described above, the morphological analysis unit 221 performs a process for dividing the Japanese text into units of morphemes.
[0072] FIG. 9 is a schematic diagram showing the reading of a morpheme when the morpheme contains Chinese characters in the result of the morphological analysis process shown in FIG. 8. As shown in the figure, data of the reading "ryou" is assigned to the morpheme "both". Also, data of the reading "sebangou" is assigned to the morpheme "jersey number". As described above, the morphological analysis unit 221 estimates the reading of the Chinese characters.
[0073] FIG. 10 is a schematic diagram showing an example of displaying the Japanese text shown in FIG. 7 as subtitles in its display area. There is one split point in this Japanese text. The subtitle presentation control device 1 switches the subtitle display in accordance with the completion of the display of the sign language motion corresponding to the split point. Specifically, FIG. 10(A) shows the display of the subtitle text before the display switch. Further, FIG. 10(B) shows the display of the subtitle text after the display switch. In each of FIGS. 10(A) and (B), the symbol SR is the subtitle display area. Depending on the size of the subtitle display area and the size of the character font to be displayed, the display (A) of the subtitle text before the display switch is "They are the starting members of both teams. The British team. Back number 4, forward, Fallcad Naomi, class. Back number 6, guard, Holland Bicky, class 2.0. Back number 8, center, Stanford No" up to. After the switch, the display (B) is "Ras. Back number 10, guard, Ames David, class 1.0. Back number 12, forward, Brye Sophie, class 3.0. The head coach is Benson Gordon." As described above, the split point calculation unit 222 can calculate this split point from the subtitle text, the size of the subtitle display area SR, and the size of the character font. That is, in the illustrated example, it is calculated in advance that up to "No" of "Noras" is displayed before the display switch. That is, not all of the morpheme "Noras" is displayed before the display switch. For this reason, the split point calculation unit 222 determines the split point in word units (morpheme units) at the word "Stanford" immediately before it. That is, the frame position inquiry unit 241 inquires the sign language motion data synthesis unit 231 about the frame at the time when the sign language motion corresponding to the word "Stanford" ends. Then, the frame position inquiry unit 241 acquires information for specifying the frame at which the sign language motion of "Stanford" ends. This frame information is passed to the subtitle presentation unit 26. That is, the subtitle presentation unit 26 switches the display of the split subtitle text at the timing when the frame is reproduced. That is, the display switch from (A) to (B) in FIG. 10 is performed.
[0074] FIG. 11 is a schematic diagram showing the relationship between Japanese words (morphemes), sign language labels, data of the sign language motions, information specifying the start frame of the sign language motion, and information specifying the end frame of the sign language motion. As shown in the drawing, the Japanese words are associated with the sign language labels. The correspondence between the Japanese words and the sign language labels does not necessarily have to be one-to-one. Even when the correspondence is not one-to-one, for a Japanese word, there exists the sign language label that is the last in terms of time (the last in the order of the sign language label sequence) among the sign language labels corresponding to the word. Also, corresponding to the sign language label, the data of the sign language motion is determined. The sign language motion data in terms of sentence units is, as described above, the one obtained by connecting all the sign language motions corresponding to each label (including the "wataru" described above). That is, for the sign language label, the position of the end frame of the sign language motion can be specified. That is, for the Japanese word, it is possible to uniquely determine the position of the final frame of the sign language motion corresponding to the word. Note that in order for the sign language motion data synthesis unit 231 to respond with the information specifying the frame at the split point to the query from the frame position inquiry unit 241, the sign language motion data synthesis unit 231 may refer to the data in the form shown in FIG. 11, or may refer to equivalent information in other forms.
[0075] FIG. 12 is a flowchart showing the procedure of the operation by the subtitle presentation control device 1. Note that this operation procedure is a procedure by a method of obtaining a sign language label sequence by translating Japanese text, which is optional text. Hereinafter, the operation procedure will be described according to this flowchart.
[0076] First, in step S1, the morpheme analysis unit 221 of the split point determination unit 22 reads the data of the Japanese text passed from the Japanese text supply unit 214 of the input data supply unit 21.
[0077] Next, in step S2, the morphological analysis unit 221 of the segmentation point determination unit 22 performs morphological analysis processing on the read Japanese text. That is, the morphological analysis unit 221 divides the Japanese text into morphological units and assigns part-of-speech information to each morphological unit. At this time, the morphological analysis unit 221 estimates the reading of the morphological unit including Chinese characters and assigns Chinese character reading information.
[0078] Next, in step S3, the segmentation point calculation unit 222 of the segmentation point determination unit 22 calculates the segmentation point of the Japanese text based on the size of the subtitle display area and the size of the character font.
[0079] Next, in step S4, the segmentation point calculation unit 222 of the segmentation point determination unit 22 obtains the word at the segmentation point corresponding to the segmentation point obtained in step S3. The word at the segmentation point is the last word that does not exceed the segmentation point obtained in step S3.
[0080] Next, in step S5, the translation unit 215 of the input data supply unit 21 translates the Japanese text supplied by the Japanese text supply unit 214 and converts it into a sign language label sequence. The sign language label sequence supply unit 216 passes this sign language label sequence output from the translation unit 215 to the sign language motion generation unit 23.
[0081] Next, in step S6, the sign language motion data synthesis unit 231 of the sign language motion generation unit 23 synthesizes sign language motion data based on the sign language label sequence passed in step S5.
[0082] Next, in step S7, the frame position inquiry unit 241 of the subtitle switching timing determination unit 24 makes an inquiry to the sign language motion data synthesis unit 231 based on the word at the segmentation point obtained in step S4. The sign language motion data synthesis unit 231 identifies the frame at the end point of the sign language motion corresponding to the word at the segmentation point in the sign language motion data generated in step S6. The sign language motion data synthesis unit 231 passes the information for identifying the frame to the frame position inquiry unit 241.
[0083] Next, in step S8, the switching timing information adding unit 242 of the subtitle switching timing determination unit 24 adds the switching timing information to the subtitle text data. The switching timing information is the timing of playing the frame specified by the information passed from the sign language motion data synthesizing unit 231 in step S7.
[0084] Next, in step S9, the playback unit 252 of the sign language animation generation unit 25 generates and plays a sign language animation by CG based on the sign language motion data synthesized in step S6.
[0085] Next, in step S10, the subtitle presenting unit 26 generates and displays subtitles. The subtitle presenting unit 26 controls the switching of subtitle display in accordance with the playback timing of the sign language CG animation played by the playback unit 252. That is, the subtitle display is switched based on the switching timing information added in step S8.
[0086] Note that the operation procedure of the subtitle presentation control device does not necessarily have to be in the exact order shown in this flowchart. That is, the processing order may be changed as long as there is no logical contradiction. For example, the processing of step S5 may be performed after reading the Japanese text and before synthesizing the sign language motion data based on the sign language label sequence. Also, for example, the processing of step S6 may be performed after the sign language label sequence is generated and before generating the sign language animation based on the sign language motion data. Also, for example, the processing of step S9 may be performed after the sign language motion data is generated.
[0087] FIG. 13 is a block diagram showing an example of the internal configuration of the subtitle presentation control device 1. The subtitle presentation control device 1 can be realized using a computer. As shown in the figure, the computer includes a central processing unit 901, a RAM 902, an input / output port 903, input / output devices 904 and 905, etc., and a bus 906. The computer itself can be realized using existing technologies. The central processing unit 901 executes instructions included in the programs read from the RAM 902 and the like. The central processing unit 901 writes data to the RAM 902, reads data from the RAM 902, and performs arithmetic operations and logical operations according to each instruction. The RAM 902 stores data and programs. Each element included in the RAM 902 has an address and can be accessed using the address. Note that RAM is an abbreviation for "Random Access Memory". The input / output port 903 is a port for the central processing unit 901 to exchange data with external input / output devices and the like. The input / output devices 904 and 905 are input / output devices. The input / output devices 904 and 905 exchange data with the central processing unit 901 via the input / output port 903. The bus 906 is a common communication path used inside the computer. For example, the central processing unit 901 reads and writes data in the RAM 902 via the bus 906. Also, for example, the central processing unit 901 accesses the input / output port via the bus 906.
[0088] At least some of the functions of the subtitle display control device 1 can be realized by a computer and a program. In that case, a program for realizing this function may be recorded on a computer-readable recording medium, and the program recorded on this recording medium may be read into a computer system and executed to realize it. Here, the "computer system" is assumed to include hardware such as an OS and peripheral devices. Also, the "computer-readable recording medium" refers to a portable medium such as a flexible disk, magneto-optical disk, ROM, CD-ROM, DVD-ROM, USB memory, etc., and a storage device such as a hard disk built into a computer system. That is, the "computer-readable recording medium" may be a non-transitory computer-readable recording medium. Furthermore, the "computer-readable recording medium" also includes those that temporarily and dynamically hold a program, such as a communication line when transmitting a program via a network such as the Internet or a communication line such as a telephone line, and those that hold a program for a certain period of time, such as a volatile memory inside a computer system that serves as a server or client in that case. Also, the above program may be for realizing a part of the functions described above, and may further be capable of realizing the functions described above in combination with a program already recorded in the computer system.
[0089] As described above, the embodiments have been explained, but the present invention can also be implemented in the following modification examples.
[0090] In the above embodiment, the subtitle display unit 26 continuously received information on the playback position (frame position) of the sign language animation from the sign language animation generation unit 25. Modification Example 1: The subtitle display unit 26 may receive a notification of the timing of the start of playback of the sign language video, and control the switching of the display of the subtitle text based on the timing of the start of playback and the switching timing (time position) determined by the subtitle switching timing determination unit 24. Modification Example 2: The sign language animation generation unit 25 may acquire in advance information on the timing of subtitle switching from the sign language motion generation unit 23, and when that timing arrives (when a predetermined frame is played), give an instruction to the subtitle presentation unit 26 to switch the display of the subtitle. In this case, the subtitle presentation unit 26 controls to switch the display of the subtitle text to the next display based on the above instruction from the sign language animation generation unit 25.
[0091] As described above, the embodiments of the present invention have been described in detail with reference to the drawings. However, the specific configuration is not limited to this embodiment, and designs and the like within the scope not departing from the gist of the present invention are also included.
Industrial Applicability
[0092] The present invention can be used, for example, in the distribution (including broadcasting) of video content. However, the scope of use of the present invention is not limited to what is exemplified here.
Explanation of Signs
[0093] 1 Subtitle presentation control device 21 Input data supply unit (data supply unit) 22 Division point determination unit 23 Sign language motion generation unit 24 Subtitle switching timing determination unit 25 Sign language animation generation unit (sign language video generation unit) 26 Subtitle presentation unit 211 Arbitrary sentence acquisition unit 212 Distribution data acquisition unit 213 Fixed sentence processing unit 214 Japanese text supply unit 215 Translation unit 216 Sign language label sequence supply unit 221 Morphological analysis unit 222 Division point calculation unit 223 Text information output unit 231 Sign language motion data synthesis unit 232 Sign language motion database 241 Frame position inquiry unit 242 Switching timing information adding unit 251 Character data storage unit 252 Reproduction unit 901 Central processing unit 902 RAM 903 Input / output port 904, 905 Input / output devices 906 Bus SR Subtitle display area
Claims
1. A data supply unit that supplies a sign language label sequence, which is a sequence of sign language labels representing sign language, and caption text corresponding to the sign language label sequence; Based on the caption text and the size of the caption text display area for displaying the caption text, a split point determination unit that obtains a word corresponding to a split point, which is a position of switching the display within the caption text, when displaying the caption text in the caption text display area; By connecting motions corresponding to each of the sign language labels based on the sign language label sequence, motion data corresponding to the sign language label sequence is generated, and a sign language motion generation unit that determines a time position within the motion data of the motion of the sign language label corresponding to the word corresponding to the split point obtained by the split point determination unit; A caption presentation unit that controls to switch the display of the caption text to the next display at a timing based on the time position determined by the sign language motion generation unit when the caption text is displayed in the caption text display area and a sign language video based on the motion data is displayed; A caption presentation control device comprising the above.
2. A sign language video generation unit that generates and reproduces and outputs the sign language video based on the motion data and notifies the caption presentation unit of the current reproduction position of the sign language video; further comprising: The caption presentation unit controls the switching of the display of the caption text based on the notified current reproduction position and the time position. The caption presentation control device according to Claim 1.
3. The caption presentation unit receives a notification of the timing of the start of reproduction of the sign language video and controls the switching of the display of the caption text based on the timing of the start of reproduction and the time position. The caption presentation control device according to Claim 1.
4. The data supply unit supplies the subtitle text acquired from the outside and the sign language label sequence obtained by machine-translating the subtitle text. The subtitle display control device according to any one of claims 1 to 3.
5. The data supply unit supplies the subtitle text obtained by inserting the embedded data received from the outside into a pre-stored subtitle text template, and the sign language label sequence obtained by inserting the embedded label corresponding to the embedded data into a pre-stored sign language label sequence template corresponding to the subtitle text template. The subtitle display control device according to any one of claims 1 to 3.
6. A data supply unit that supplies a sign language label sequence, which is a sequence of sign language labels representing sign language, and subtitle text corresponding to the sign language label sequence. Based on the subtitle text and the size of the subtitle text display area for displaying the subtitle text, a split point determination unit that obtains a word corresponding to the split point, which is the position of switching the display within the subtitle text, when displaying the subtitle text in the subtitle text display area. By connecting motions corresponding to each of the sign language labels based on the sign language label sequence, motion data corresponding to the sign language label sequence is generated, and a sign language motion generation unit that determines the time position in the motion data of the motion of the sign language label corresponding to the word corresponding to the split point obtained by the split point determination unit. A subtitle presentation unit that controls to switch the display of the subtitle text to the next display at a timing based on the time position determined by the sign language motion generation unit when the subtitle text is displayed in the subtitle text display area and a sign language video based on the motion data is displayed. A program that causes a computer to function as a subtitle display control device including the above.
Citation Information
Patent Citations
Special mobile telephone for deaf-mute
CN101605158A
Receiver, broadcast transmission apparatus and auxiliary content server
JP2004235734A
Caption production system
JP2004336606A
Caption display control apparatus
JP2007316613A
Machine translation device and machine translation program
JP2013186673A