Information processing device, information processing method, computer program, and data structure
The information processing device addresses the limitations of existing speech assistance technologies by displaying a moving image with relative character positioning and speech features, enabling intuitive control of intonation, speed, and emotional expression for enhanced speech performance.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NAT UNIV CORP TOKAI NAT HIGHER EDUCATION & RES SYST
- Filing Date
- 2025-06-25
- Publication Date
- 2026-07-23
AI Technical Summary
Existing speech assistance technologies, such as visual prompts for memorizing lines, fail to adequately support the performance of manzai and other forms of speech that require control of intonation, speech rate, and emotional expression.
An information processing device that acquires speech reference data to display a moving image showing the relative positional relationship between characters, intonation, speech speed, and emotional expressions, allowing speakers to intuitively grasp and perform these aspects without memorization.
The device effectively supports speakers by enabling appropriate intonation, speech speed, timing, and emotional expression through a visually represented speech reference image, enhancing the performance of manzai and similar forms of speech.
Smart Images

Figure 0007894180000001_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed in this specification relates to information processing for assisting speech.
Background Art
[0002] Opportunities for speech exist innumerable, from daily conversations to presentations at international conferences. Conventionally, various devices for assisting speech have been utilized. For example, a prompt that assists smooth speech by a speaker by allowing the speaker to visually recognize an image showing a character string representing the speech content is widely used (see, for example, Patent Document 1).
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] As one aspect of speech, there is a performance of manzai. Manzai is an art form in which, for example, two performers engage in comical exchanges for the purpose of making the audience laugh. A common way to enjoy manzai is to watch a performance by professional manzai artists. In recent years, as a new way to enjoy manzai, there has been an increasing trend for ordinary people who are not professional manzai artists to perform manzai.
[0005] It is not easy for ordinary people to perform manzai well, for example, because it requires memorizing lines. It is conceivable to use a prompt to allow the performers to visually recognize an image showing the lines, thereby assisting the performers in vocalizing the lines. However, in order to perform manzai well, it is necessary not only to simply utter the lines but also to appropriately control the intonation and speech rate during speech. Therefore, even by using a conventional prompt, it is not possible to sufficiently assist the performance of manzai.
[0006] These challenges are not limited to speech during stand-up comedy performances, but are common challenges when supporting speech in general, including everyday conversation and presentations.
[0007] This specification discloses a technology capable of solving the above-mentioned problems. [Means for solving the problem]
[0008] The technologies disclosed herein can be implemented, for example, in the following forms:
[0009] (1) The information processing device disclosed herein comprises a data acquisition unit and an image display unit. The data acquisition unit acquires speech reference data. The speech reference data includes object data that identifies an object containing a plurality of characters representing the content of an utterance, and feature data that identifies speech features including at least one of the intonation and speech speed at the time of utterance. The image display unit causes a speech reference image, which is a moving image showing the object, to be displayed on a display device based on the speech reference data. The speech reference image is an image that represents the speech features based on the relative positional relationship between the characters.
[0010] According to this information processing device, in a speech reference image that shows an object containing multiple characters representing the content of an utterance, speech features, including at least one of the intonation and speech speed, are represented by the relative positional relationship between the multiple characters. Therefore, by referring to the speech reference image, the speaker can utter with appropriate intonation and / or speed without having to memorize the string of characters representing the content of the utterance. Thus, this information processing device can effectively support the speaker's utterance.
[0011] (2) In the above-mentioned information processing device, the speech features may include intonation during speech, and the speech reference image may be an image that represents intonation based on the positional relationship along a direction that intersects the direction indicating the order in which the plurality of characters are spoken. With this configuration, the speaker can intuitively grasp the appropriate intonation by referring to the speech reference image. Therefore, this information processing device can more effectively support speech by the speaker.
[0012] (3) In the above-mentioned information processing device, the speech feature may include speech speed, and the speech reference image may be an image that represents speech speed based on the positional relationship along the direction indicating the order in which the plurality of characters are spoken. With this configuration, the speaker can intuitively grasp the appropriate speech speed by referring to the speech reference image. Therefore, this information processing device can more effectively support the speaker's speech.
[0013] (4) In the above-described information processing device, the speech reference data may include timing data that specifies the execution timing of each object, and the speech reference image may be an image that shows the execution timing of each object. With this configuration, the speaker can execute the object at the appropriate timing by referring to the speech reference image. Therefore, this information processing device can more effectively support the speaker's speech.
[0014] (5) In the above-mentioned information processing device, the object may include a graphic that represents the emotional expression at the time of utterance. With this configuration, the speaker can make a speech with appropriate emotional expression by referring to the speech reference image. Therefore, this information processing device can effectively support the speaker's speech.
[0015] (6) In the above-mentioned information processing device, the object may include a graphic representing a body movement during speech. With this configuration, the speaker can perform speech with appropriate body movements by referring to the speech reference image. Therefore, this information processing device can effectively support speech by the speaker.
[0016] (7) In the above-mentioned information processing device, the image display unit may include a playback speed changing unit that changes the playback speed of the speech reference image. With this configuration, the speaker can speak while referring to the speech reference image played back at their preferred speed. Therefore, this information processing device can effectively support the speaker's speech.
[0017] (8) In the above-mentioned information processing device, the image display unit displays an edited image including the object on the display device based on the speech reference data, and the information processing device may further include an editing unit that edits the speech reference data in accordance with instructions from the user while the edited image is being displayed. With this configuration, the speech reference image can be edited into an image that represents more appropriate speech features, and the speech made by the speaker can be supported more effectively.
[0018] (9) In the above-mentioned information processing device, the instruction may include an instruction to change the position of the characters in the edited image. With this configuration, the editor can intuitively give editing instructions so that the utterance reference image becomes an image that more appropriately represents the utterance features.
[0019] The technologies disclosed herein can be implemented in various forms, for example, in the form of information processing apparatus and methods, speech support apparatus and methods, computer programs that implement these methods, data structures, data having data structures, and non-temporary recording media on which such computer programs and data are recorded. [Brief explanation of the drawing]
[0020] [Figure 1] Explanatory drawing showing an example of the speech reference image RI in this embodiment [Figure 2] Explanatory drawing showing an example of the speech reference image RI in this embodiment [Figure 3] Explanatory drawing showing an example of the speech reference data RD representing the speech reference image RI [Figure 4] Explanatory drawing showing an example of the edited image EI in this embodiment [Figure 5] Explanatory drawing showing the schematic configuration of the information processing apparatus 100 [Figure 6] Flowchart showing the speech reference image reproduction process [Figure 7] Flowchart showing the editing process [Figure 8] Explanatory drawing showing an example of the speech reference image RI in the modified example
Mode for Carrying Out the Invention
[0021] (Embodiment) (Overview of the speech reference image RI) FIG. 1 and FIG. 2 are explanatory drawings showing an example of the speech reference image RI in this embodiment. The speech reference image RI is a moving image for assisting the actual performance of a comic dialogue including the voice of a line by a performer by allowing a non-professional comedian, for example, an ordinary person, to visually recognize it when performing the comic dialogue. The speech reference image RI functions as a score for the performer. The speech reference image RI of this embodiment is referred to when two performers (so-called "straight man" and "funny man") perform a comic dialogue. FIGS. 1 and 2 illustrate one scene of the speech reference image RI at different timings.
[0022] The speech reference image RI includes a background image BI and a score image SI. The background image BI is an image that serves as the background of the speech reference image RI and is an image without movement. The background image BI includes a first lane L1 and a second lane L2, a playback-related button PB, and a speed adjustment button SB. Hereinafter, the first lane L1 and the second lane L2 are also collectively referred to as lanes L1, L2.
[0023] Each lane L1 and L2 extends linearly from the right edge to the left edge of the speech reference image RI. The second lane L2 is located below the first lane L1. In this embodiment, the first lane L1 is an area for displaying the content of speech by the straight man, and the second lane L2 is an area for displaying the content of speech by the funny man. A linear timing bar TB extending vertically is placed on each lane L1 and L2. The timing bar TB is positioned slightly to the left of the center of the speech reference image RI. The left-right position of the timing bar TB on the first lane L1 is the same as the left-right position of the timing bar TB on the second lane L2.
[0024] The playback-related buttons PB include a play button P1, a pause button P2, a fast-forward button P3, and a rewind button P4. By operating the playback-related buttons PB, the user can play, pause, fast-forward, and rewind the speech reference image RI.
[0025] The speed adjustment button SB includes a constant speed button S1, an acceleration button S2, and a deceleration button S3. By operating the speed adjustment button SB, the user can adjust the playback speed of the speech reference image RI to their preferred speed. The constant speed button S1 is a button that matches the playback speed of the speech reference image RI to the actual performance speed of a professional comedian. Each time the acceleration button S2 is pressed, the playback speed of the speech reference image RI increases by, for example, 5%. Each time the deceleration button S3 is pressed, the playback speed of the speech reference image RI decreases by, for example, 5%.
[0026] The musical score image SI is an image that shows the performance content (speech content, emotional expression, and physical movements) by each performer. The musical score image SI contains multiple letters CH. The multiple letters CH are images that show the speech content by each performer. The multiple letters CH are arranged from left to right in the order of utterance.
[0027] In this embodiment, the musical score image SI further includes an effect image EF and a motion panel image MP. Hereinafter, the components of the musical score image SI (character CH, effect image EF, motion panel image MP) will be collectively referred to as an object.
[0028] Effect images EF are partial images that represent the emotional expression of each performer through their design. Figures 1 and 2 show examples of effect images EF1 representing the emotion of "anger" and effect image EF2 representing the emotion of "joy". The emotion of "anger" can be used, for example, when a performer playing the straight man role makes a retort, and the emotion of "joy" can be used, for example, when a performer playing the funny man role makes a joke. Effect images EF corresponding to other emotions ("sadness" or "happiness") may also be used. Effect images EF may be displayed alone or overlaid on text CH or motion panel images MP, which will be described later.
[0029] Motion panel images MP are partial images that show the physical movements of each performer, indicated by their patterns. Figures 1 and 2 illustrate motion panel images MP1, which shows a straight man performer making a retort, and motion panel image MP2, which shows a funny man performer leaning back and laughing. Motion panel images MP corresponding to other physical movements may also be used. Motion panel images MP may be displayed alone or overlaid on text CH or effect images EF.
[0030] During playback of the speech reference image RI, each object constituting the score image SI (character CH, effect image EF, motion panel image MP) moves synchronously at a constant speed from right to left along lanes L1 and L2 in the background image BI. The object for the straight man character moves along the first lane L1, and the object for the funny man character moves along the second lane L2. In other words, the objects associated with each character are displayed side by side in the vertical direction. In each lane L1 and L2, the timing at which an object reaches the timing bar TB is the execution timing of that object. The execution timing of an object is the timing of speech for character CH, the timing of expressing emotion for effect image EF, and the timing of performing a physical action for motion panel image MP. Each character can execute an object at the appropriate timing by referring to the positional relationship between each object in their own lane L1 and L2 on the speech reference image RI and the timing bar TB. In other words, each character can deliver lines with the appropriate emotion or perform appropriate physical actions at the appropriate timing.
[0031] Furthermore, in the speech reference image RI, the timing of the dialogue between the two performers is determined by the difference between the timing at which each object for one performer reaches the timing bar TB and the timing at which each object for the other performer reaches the timing bar TB. In this way, the speech reference image RI is an image that shows the execution timing of each object by each performer in sync with each other.
[0032] Each object, after reaching the timing bar TB, moves past the timing bar TB towards the left edge. In this embodiment, the display of an object after passing the timing bar TB is different from the display of an object before passing the timing bar TB. For example, an object after passing the timing bar TB is displayed in gray.
[0033] In each lane L1 and L2 of the speech reference image RI, the spacing between multiple characters CH in the left-right direction (hereinafter referred to as "left-right spacing") is variable. For example, a narrow left-right spacing between multiple characters CH constituting a line of dialogue indicates that the line is spoken relatively quickly. Conversely, a wide left-right spacing between multiple characters CH constituting a line of dialogue indicates that the line is spoken relatively slowly. Thus, in the speech reference image RI of this embodiment, speech speed is expressed by the relative positional relationship of multiple characters CH in the left-right direction. Each performer can speak lines at an appropriate speed by referring to the relative positional relationship of multiple characters CH in the left-right direction in the speech reference image RI. The left-right direction is the direction that indicates the order in which the multiple characters CH are spoken. Speech speed is an example of a speech feature.
[0034] In each lane L1 and L2 of the speech reference image RI, the vertical position of each character CH (hereinafter referred to as "vertical position") is variable. In this embodiment, the vertical position of character CH is set to five levels (top, slightly up, center, slightly down, and bottom). The vertical position of character CH represents the intonation (pitching) of the performer's speech. For example, if the vertical position of character CH is top, slightly up, center, slightly down, and bottom, it represents high, slightly high, normal, slightly low, and low pitch, respectively. Thus, in the speech reference image RI of this embodiment, the intonation of speech is expressed by the relative vertical positional relationship between multiple character CHs. Each performer can refer to the relative vertical positional relationship of multiple character CHs in the speech reference image RI and deliver their lines with appropriate intonation. The vertical direction intersects with the direction (left-right direction) that indicates the order in which the multiple character CHs are spoken. Intonation is one example of speech characteristics.
[0035] In the speech reference image RI, the font of each character CH is variable. The font of each character CH represents the emotional expression at the time of speech. In this embodiment, eight types of fonts (yorokobi / kanashimi / odoroki / horror / neutral / tsukkomi / boke / emphasize) are used. For example, among joy, anger, sorrow, and pleasure, the emotions of "joy" and "pleasure" are represented by the font "yorokobi," the emotion of "anger" is represented by the font "tsukkomi," and the emotion of "sorrow" is represented by the font "kanashimi." The emotion of "anger" is used, for example, in the lines of a straight man's retort, while the emotions of "joy" and "pleasure" are used, for example, in the lines of a funny man's joke. Each performer can refer to the font of each character CH in the speech reference image RI and deliver their lines with the appropriate emotion. The emotional expression at the time of speech is an example of speech features.
[0036] In this way, each performer can perform a comedy routine with appropriate timing, speed, intonation, emotional expression, and physical movements by referring to the speech reference image RI, without having to memorize lines.
[0037] Figure 3 is an explanatory diagram showing an example of speech reference data RD representing a speech reference image RI. In this embodiment, the speech reference data RD includes CSV data that defines a musical score image SI. As shown in Figure 3, the CSV data included in the speech reference data RD consists of multiple rows. Each row of this data consists of five elements: "Target", "Content", "Timing", "Option", and "Font".
[0038] The element "Target" indicates the target performer and object. This element uses the codes "A" and "B" to indicate the performer's classification (straight man, funny man), and "(unsigned)", "E", and "M" to indicate the object's classification (text CH, effect image EF, motion panel image MP). For example, the value of the element "Target" for text CH targeting the funny man is "B", the value of the element "Target" for effect image EF targeting the straight man is "AE", and the value of the element "Target" for motion panel image MP targeting the funny man is "BM".
[0039] The element "Content" indicates the content of the object. For text elements CH, this element shows the text itself; for effect images EF, it indicates which of the four emotions (Joy, Angry, Sad, Happy) it represents; and for motion panel images MP, it shows the file name. The element "Content" is an example of object data.
[0040] The "Timing" element indicates the execution timing of an object. This element shows the elapsed time (in seconds) from the start time of playback of the speech reference image RI until the object reaches the position of the timing bar TB. The difference in the value of the "Timing" element between two objects determines the left-right spacing between those two objects in the speech reference image RI. The "Timing" element is an example of timing data that identifies the execution timing of each object, and also an example of feature data that identifies the speech rate.
[0041] The element "Option" indicates the vertical position of an object. This element uses the symbols "H", "h", "m", "l", and "L" to represent five vertical positions (top, slightly up, center, slightly down, bottom). The element "Option" is an example of feature data that identifies intonation during speech.
[0042] The element "Font" indicates the font of the object. In this embodiment, eight fonts are available as options, and the element "Font" indicates one of these fonts. In this embodiment, for objects other than the character CH, the element "Font" is set to the default value of "neutral". The element "Font" is an example of feature data that identifies emotional expression during speech.
[0043] For example, in the second row of the CSV data shown in Figure 3, the value of the element "Target" is "B", the value of the element "Content" is "っ", the value of the element "Timing" is "2.480633", the value of the element "Option" is "m", and the value of the element "Font" is "yorokobi". Therefore, this second row indicates that, as a musical score image SI included in the speech reference image RI, the character "っ" represented by the font "yorokobi" will be displayed at the vertical position = center position in the second lane L2 for the comedic role, at a timing such that it reaches the timing bar TB 2.480633 seconds after the start of playback. The comedic actor, referring to the speech reference image RI, is prompted to pronounce the character "っ" with a moderate intonation (pitching) and expressing joy at the above timing.
[0044] Furthermore, in the 21st row of the CSV data shown in Figure 3, the value of the element "Target" is "BM", the value of the element "Content" is "gray_kasu_tegaki5", the value of the element "Timing" is "2.387979", the value of the element "Option" is "l", and the value of the element "Font" is "neutral". Therefore, this 21st row indicates that, as a musical score image SI included in the speech reference image RI, the motion panel image MP (backward laughter) with the file name "gray_kasu_tegaki5" will be displayed at a position slightly below the vertical direction in the second lane L2 for the comedic character, at a timing such that it reaches the timing bar TB 2.387979 seconds after the start of playback. The comedic character performer, referring to the speech reference image RI, will be prompted to perform the backward laughter motion at the above timing.
[0045] Furthermore, in row 23 of the CSV data shown in Figure 3, the value of the element "Target" is "AE", the value of the element "Content" is "Angry", the value of the element "Timing" is "2.006745", the value of the element "Option" is "L", and the value of the element "Font" is "neutral". Therefore, row 23 indicates that, as a musical score image SI included in the speech reference image RI, the effect image EF representing the emotion of anger will be displayed at the bottom position in the first lane L1 for the straight man, at a timing such that it reaches the timing bar TB 2.006745 seconds after the start of playback. The performer of the straight man, referring to the speech reference image RI, will be prompted to express the emotion of anger at the above timing.
[0046] Thus, the speech reference data RD representing the speech reference image RI includes object data that identifies an object containing multiple characters CH representing the content of the utterance, and feature data that identifies speech features including at least one of the intonation and speech rate during the utterance. The speech reference data RD causes the speech reference image RI, which represents the speech features based on the relative positional relationships between the multiple characters CH, to be displayed on a display device.
[0047] The speech reference data RD with the above configuration is generated, for example, by the following method. The characters CH that make up each performer's lines are identified by extracting audio data and separating speakers from videos of professional comedians performing stand-up comedy. The intonation of the lines is obtained from the above audio data using Fourier transform, etc., and these frequencies are classified into the five levels of intonation (pitching) described above.
[0048] The effect images (EF) and fonts representing the performers' emotional expressions will be identified by performing emotion recognition based on the content of the dialogue, voice, and facial expressions in the aforementioned comedy performance video. The motion panel images (MP) representing the performers' physical movements will be identified by performing motion capture on the aforementioned comedy performance video. For the effect images (EF) and motion panel images (MP), the closest image will be assigned from a pre-prepared set of images. If no close image exists, a new one may be created.
[0049] The execution timing of each object will be determined according to the execution timing of each object in the above-mentioned comedy performance video. For physical movements, for example, the start timing of movements deemed important for the success of the routine will be determined.
[0050] Based on the identified lines (strings of dialogue), emotional expressions, and bodily movements, along with the timing of their execution, speech reference data (RD) representing them is created.
[0051] (Overview of edited image EI) Figure 4 is an explanatory diagram showing an example of an edited image EI in this embodiment. The edited image EI is an image used for editing the speech reference image RI. The edited image EI includes a background image BIe and a musical score image SI, similar to the speech reference image RI described above. The layout of the background image BIe of the edited image EI is substantially the same as the layout of the background image BI of the speech reference image RI. However, the background image BIe of the edited image EI includes several interfaces, which are described below, for editing the speech reference image RI. Note that the musical score image SI of the edited image EI is the same as the musical score image SI of the speech reference image RI, and is therefore omitted from the illustration in Figure 4.
[0052] The background image Bie of the edited image EI includes a dialogue input field E1. When the editor selects either lane L1 or L2 and enters one or more characters CH representing dialogue into the dialogue input field E1, the entered characters CH are displayed in the selected lane. The editor can change the intonation of the dialogue by changing the vertical position of each displayed character CH, for example by dragging with the mouse. The editor can also change the speaking speed and timing by changing the horizontal position of each displayed character CH. Note that changing the horizontal position of a character CH will consequently change the horizontal spacing between that character CH and other adjacent character CHs, and thus change the speaking speed of those character CHs. The editor can also set the font of the character CH, thereby setting the emotional expression of the dialogue. The background image Bie of the edited image EI also includes a furigana input field E2 for entering the phonetic readings of the dialogue.
[0053] The background image Bie of the editing image EI includes the part addition button E5. When the editor selects the part addition button E5, a library screen (not shown) opens displaying candidate objects (effect image EF and motion panel image MP). When the editor selects an object from the library screen and also selects a lane, the selected object is displayed in the selected lane. The editor can change the execution timing of the object by changing its left-right position.
[0054] The background image Bie of the edited image EI includes a right-scroll button E6 and a left-scroll button E7. When the editor selects the right-scroll button E6, each object (score image SI) in the edited image EI scrolls, displaying the image corresponding to a later scene in time. Conversely, when the editor selects the left-scroll button E7, each object in the edited image EI scrolls in the opposite direction, displaying the image corresponding to an earlier scene in time.
[0055] The background image Bie of the edited image EI includes a play button E3 and a pause button E4. When the editor selects the play button E3, the latest image at that time will be played, and when the pause button E4 is selected during playback, playback will be paused.
[0056] The background image BIe of the edited image EI includes the data output button E8. When the editor selects the data output button E8, speech reference data RD is generated, representing the edited speech reference image RI at that time. This enables the updating of the speech reference data RD.
[0057] (Configuration of the information processing device 100) Next, the configuration of the information processing device 100 for performing various processes related to the speech reference image RI will be described. Figure 5 is an explanatory diagram showing the schematic configuration of the information processing device 100. The information processing device 100 is composed of, for example, a computer (PC, server, smartphone, tablet terminal, etc.).
[0058] The information processing device 100 comprises a control unit 110, a storage unit 120, a display unit 130, an operation input unit 140, an interface unit 150, and an audio input / output unit 160. Each of these units is connected to the others via a bus 190 so as to be able to communicate with each other.
[0059] The display unit 130 of the information processing device 100 is configured, for example, as a liquid crystal display, and displays various images and information. The display unit 130 is an example of an output device and display device. The operation input unit 140 is configured, for example, as a keyboard, mouse, buttons, microphone, trackpad, etc., and accepts user operations and instructions. The display unit 130 may also function as the operation input unit 140 by being equipped with a touch panel. The interface unit 150 is configured, for example, as a LAN interface or USB interface, and communicates with other devices by wired or wireless connection. The audio input / output unit 160 is configured, for example, as a microphone or speaker, and inputs and outputs audio.
[0060] The storage unit 120 is composed of, for example, ROM, RAM, a hard disk drive (HDD), and stores various programs and data, and is used as a workspace and temporary storage area for data when executing various programs. For example, the storage unit 120 stores a speech support processing program CP, which is a computer program for executing various processes described later. The speech support processing program CP is provided, for example, stored on a computer-readable recording medium (not shown) such as a CD-ROM, DVD-ROM, or USB memory, or is provided in a state that can be obtained from an external device (a server on a network or other terminal device) via the interface unit 150, and is stored in the storage unit 120 in a state that can be operated on the information processing device 100. The storage unit 120 also stores speech reference data RD, which represents a speech reference image RI.
[0061] The control unit 110 is composed of, for example, a CPU, and controls the operation of the information processing device 100 by executing a computer program read from the storage unit 120. For example, the control unit 110 functions as a speech support processing unit 111 for executing various processes described later by reading and executing a speech support processing program CP from the storage unit 120. The speech support processing unit 111 includes a data acquisition unit 112, an image display unit 113 including a playback speed change unit 114, and an editing unit 115. The functions of each of these units will be explained in accordance with the descriptions of the various processes described later.
[0062] (Speech reference image playback processing) Next, the speech reference image playback process performed by the information processing device 100 of this embodiment will be described. Figure 6 is a flowchart of the speech reference image playback process. The speech reference image playback process is a process that displays (plays) a speech reference image RI, for example, to support a stand-up comedy performance by performers. The speech reference image playback process is started when a user operates the operation input unit 140 of the information processing device 100 and inputs a start command.
[0063] First, the data acquisition unit 112 (Figure 5) of the information processing device 100 acquires speech reference data RD (S110). The data acquisition unit 112 may acquire speech reference data RD from an external device via the interface unit 150, or it may acquire speech reference data RD by generating it itself. The acquired speech reference data RD is stored in the storage unit 120.
[0064] Next, the image display unit 113 (Figure 5) of the information processing device 100 reproduces and displays the speech reference image RI on the display unit 130 based on the speech reference data RD (S120). The performer performs the comedy routine by referring to the reproduced speech reference image RI. The audio during the comedy routine performance may be recorded via the audio input / output unit 160. The physical movements during the comedy routine performance may be recorded via a camera or the like (not shown).
[0065] The speech support processing unit 111 (Figure 5) of the information processing device 100 monitors for any user requests to change the playback speed during playback of the speech reference image RI (S130). These playback speed change requests are input, for example, via the speed adjustment button SB on the speech reference image RI. If a playback speed change request is received (S130: YES), the playback speed change unit 114 changes the playback speed of the speech reference image RI according to the request (S140).
[0066] Although not shown in the diagram, if a user input (pause, fast forward, rewind) is received, for example via the playback-related button PB, during playback of the speech reference image RI, the image display unit 113 will pause, fast forward, or rewind the speech reference image RI according to the input.
[0067] The speech support processing unit 111 (Figure 5) of the information processing device 100 monitors whether the speech reference image RI has been played back to the end (S150). If the speech reference image RI is played back to the end (S150: YES), the speech support processing unit 111 terminates the speech reference image playback process.
[0068] (Editing process) Next, the editing process performed by the information processing device 100 of this embodiment will be described. Figure 7 is a flowchart of the editing process. The editing process is the process of editing and updating the speech reference image RI. The editing process is started when the user operates the operation input unit 140 of the information processing device 100 and inputs a start command.
[0069] First, the image display unit 113 (Figure 5) of the information processing device 100 displays the edited image EI on the display unit 130 based on the speech reference data RD stored in the storage unit 120 (S210).
[0070] Next, the editing unit 115 of the information processing device 100 updates the edited image EI according to the editing instructions from the user (S220). For example, if the user instructs the placement of an object, the editing unit 115 updates the edited image EI so that the object is displayed at the instructed position in the edited image EI. Also, if the user instructs the position of an object to be changed, the editing unit 115 updates the edited image EI so that the position of the object in the edited image EI is changed.
[0071] The speech support processing unit 111 (Figure 5) of the information processing device 100 monitors for the presence or absence of a data output instruction from the user (S230). A data output instruction is input, for example, via the data output button E8 in the edited image EI. When a data output instruction is received (S230: YES), the editing unit 115 updates the speech reference data RD according to the editing results at that time (S240) and terminates the editing process.
[0072] (Effects of this embodiment) As described above, the information processing device 100 of this embodiment comprises a data acquisition unit 112 and an image display unit 113. The data acquisition unit 112 acquires speech reference data RD having object data that identifies an object containing a plurality of characters CH representing the content of an utterance, and feature data that identifies speech features including at least one of the intonation and speech speed at the time of utterance. Based on the speech reference data RD, the image display unit 113 displays a speech reference image RI, which is a moving image representing an object, on the display unit 130. The speech reference image RI is an image that represents the above-mentioned speech features based on the relative positional relationship between the plurality of characters CH.
[0073] Thus, in the information processing device 100 of this embodiment, in the speech reference image RI which shows an object containing multiple characters CH representing the content of the utterance, speech features including at least one of the intonation and speech speed during utterance are represented by the relative positional relationship between the multiple characters CH. Therefore, by referring to the speech reference image RI, the speaker can utter (speak the lines) with appropriate intonation and / or speed without having to memorize the string of characters (lines) representing the content of the utterance. Accordingly, the information processing device 100 of this embodiment can effectively support the speaker's utterance.
[0074] In this embodiment, the speech reference image RI is an image that represents the intonation of speech based on the positional relationship along a direction (e.g., vertical direction) that intersects with the direction (e.g., left-right direction) indicating the pronunciation order of multiple characters CH. Therefore, the speaker can intuitively grasp the appropriate intonation by referring to the speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support the speaker's speech.
[0075] In this embodiment, the speech reference image RI is an image that represents speech speed based on the positional relationship of multiple characters CH along a direction (e.g., left-right direction) indicating the order in which they are spoken. Therefore, the speaker can intuitively grasp the appropriate speech speed by referring to the speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support the speaker's speech.
[0076] In this embodiment, the speech reference data RD includes timing data that specifies the execution timing of each object, and the speech reference image RI is an image that shows the execution timing of each object. Therefore, the speaker can execute objects at the appropriate timing by referring to the speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support the speaker's utterances.
[0077] In this embodiment, the object includes a graphic representing the emotional expression at the time of utterance. Therefore, the speaker can make an utterance with appropriate emotional expression by referring to the utterance reference image RI. Accordingly, the information processing device 100 of this embodiment can effectively support the speaker's utterance.
[0078] In this embodiment, the object includes a graphic representing a body movement during speech. Therefore, the speaker can perform speech with appropriate body movements by referring to the speech reference image RI. Accordingly, the information processing device 100 of this embodiment can effectively support speech by the speaker.
[0079] In this embodiment, the image display unit 113 includes a playback speed changing unit 114 that changes the playback speed of the speech reference image RI. Therefore, the speaker can speak while referring to the speech reference image RI played back at their preferred speed. Accordingly, the information processing device 100 of this embodiment can effectively support the speaker's speech.
[0080] In this embodiment, the image display unit 113 displays an edited image EI, which includes an object, on the display unit 130 based on the speech reference data RD. The information processing device 100 further includes an editing unit 115 that edits the speech reference data RD according to instructions from the user while the edited image EI is being displayed. As a result, the speech reference image RI can be edited to represent more appropriate speech features, thereby more effectively supporting the speaker's utterance. In this embodiment, the user's instructions include instructions to change the position of the character CH in the edited image EI. As a result, the editor can intuitively give editing instructions so that the speech reference image RI becomes an image that represents more appropriate speech features.
[0081] The data structure of the speech reference data RD in this embodiment includes object data that identifies an object containing multiple characters CH, and feature data that identifies speech features including at least one of intonation and speech rate during speech. The data structure is such that a speech reference image RI representing speech features based on the relative positional relationship between the multiple characters CH is displayed on the display unit 130. According to the data structure of the speech reference data RD in this embodiment, a speech reference image RI that can effectively support speech by the speaker can be displayed.
[0082] The information processing device 100 of this embodiment includes a data acquisition unit 112 and an image display unit 113. The data acquisition unit 112 acquires speech reference data RD, which includes object data that identifies an object containing multiple characters CH representing the content of each of multiple speakers' speeches, and timing data that identifies the execution timing of each object by each speaker. Based on the speech reference data RD, the image display unit 113 displays a speech reference image RI, which is a moving image showing an object, on the display unit 130. The speech reference image RI is an image that shows the execution timing of each object by each speaker synchronized with each other.
[0083] Thus, in the information processing device 100 of this embodiment, in the speech reference image RI which shows an object containing multiple characters CH representing the speech content of each of the multiple speakers, the execution timing of each object by each speaker is shown to be synchronized with each other. Therefore, by referring to a single speech reference image RI, each speaker can speak (speak lines) at appropriate, synchronized timings without having to memorize the string of characters (lines) that represent the speech content. Accordingly, the information processing device 100 of this embodiment can effectively support speech by multiple speakers.
[0084] In this embodiment, the speech reference image RI is an image in which an object associated with each speaker moves at a constant speed in a predetermined direction (for example, left or right). Therefore, each speaker can easily grasp the execution timing of each other's objects by referring to a single speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support speech by multiple speakers.
[0085] In this embodiment, the speech reference image RI is an image in which objects associated with each speaker are arranged in a direction that intersects the direction of movement of the objects (e.g., left-right direction). Therefore, each speaker can more easily grasp the execution timing of each other's objects by referring to a single speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support speech by multiple speakers.
[0086] The data structure of the speech reference data RD in this embodiment comprises object data that identifies an object containing multiple characters CH representing the content of each utterance by multiple speakers, and timing data that identifies the execution timing of each object by each speaker. The data structure is used to display a speech reference image RI on the display unit 130, which shows the execution timing of each object by each speaker synchronized with each other. According to the data structure of the speech reference data RD in this embodiment, it is possible to display a speech reference image RI that can effectively support utterances by multiple speakers.
[0087] (modified version) The technologies disclosed herein are not limited to the embodiments described above and can be modified in various forms without departing from their essence, for example, the following modifications are possible.
[0088] The configuration of the information processing device 100 in the above embodiment is merely an example and can be modified in various ways. Similarly, the content of the various processes in the above embodiment is merely an example and can be modified in various ways. For example, in the above embodiment, the speech reference image RI and the edited image EI are displayed on the display unit 130 of the information processing device 100, but these images may also be displayed on an external display device.
[0089] In the above embodiment, the speaker's body movements are shown using a motion panel image MP in the speech reference image RI. However, as shown in the modified example in Figure 8, the speaker's body movements may also be shown using a body movement video MC. The body movement video MC is, for example, a video generated by motion capture from a performance by a professional comedian.
[0090] In the speech reference image RI of the above embodiment, some or all of the objects associated with multiple speakers may be displayed overlapping each other.
[0091] In the above embodiment, the speech reference image RI is an image showing the content of speech by two speakers, but the speech reference image RI may also be an image showing the content of speech by one speaker, or an image showing the content of speech by three or more speakers.
[0092] In the above embodiment, at least one of the functional units included in the control unit 110 of the information processing device 100 may be included in another device instead of the control unit 110 of the information processing device 100. The various processes in the above embodiment do not necessarily have to be performed by a single device, but may be performed by different devices. Each step of the various processes in the above embodiment does not necessarily have to be performed by a single device, but may be performed by different devices. In the above embodiment, some of the configurations implemented by hardware may be replaced with software, and conversely, some of the configurations implemented by software may be replaced with hardware.
[0093] In the embodiments described above, a stand-up comedy performance was used as an example of one form of speech, but the technology disclosed herein is similarly applicable to support speech in general, including everyday conversation and presentations at international conferences. For example, the technology disclosed herein can be applied to karaoke, allowing users to practice imitating the singing style (intonation, etc.) of artists by referring to speech reference images. Furthermore, the technology disclosed herein can be applied to dubbing movies, allowing users to refer to speech reference images to encourage emotional expression that more closely resembles the atmosphere of the original work, and to make the timing of each performer's lines clearer. [Explanation of symbols]
[0094] 100: Information processing unit 110: Control unit 111: Speech support processing unit 112: Data acquisition unit 113: Image display unit 114: Playback speed change unit 115: Editing unit 120: Memory unit 130: Display unit 140: Operation input unit 150: Interface unit 160: Audio input / output unit 190: Bus BI: Background image BIe: Background image CH: Text CP: Speech support processing program E1: Dialogue input field E2: Furigana input field E3: Play button E4: Pause button E5: Add part button E6: Screen right scroll button E7: Screen left scroll button E8: Data output button EF: Effect image EI: Editing image L1: First lane L2: Second lane MC: Body movement video MP: Motion panel image P1: Play button P2: Pause button P3: Fast forward button P4: Rewind button PB: Playback related buttons RD: Speech reference data RI: Speech reference image S1: Constant speed button S2: Acceleration button S3: Deceleration button SB: Speed adjustment button SI: Score image TB: Timing bar
Claims
1. An information processing device for supporting speech, A data acquisition unit acquires speech reference data having object data that identifies an object containing multiple characters representing the content of an utterance, and feature data that identifies speech features including at least one of the intonation and speech speed during the utterance. An image display unit that displays an utterance reference image, which is a moving image representing the object based on the utterance reference data, and which represents the utterance features based on the relative positional relationships between a plurality of characters arranged in the utterance reference image, on a display device; Equipped with, The utterance reference data includes timing data that specifies the execution timing of each object, The aforementioned utterance reference image is an image that shows the execution timing of each of the aforementioned objects, The object is an information processing device that includes at least one of a graphic representing an emotional expression during speech and a graphic representing a bodily movement during speech.
2. An information processing apparatus according to claim 1, The aforementioned speech features include intonation during speech, The utterance reference image is an image that represents intonation based on the positional relationship of the plurality of characters arranged in the utterance reference image, along a direction that intersects the direction indicating the order in which the characters are uttered.
3. An information processing apparatus according to claim 1 or claim 2, The aforementioned speech features include speech rate, The utterance reference image is an image that represents the speech rate based on the positional relationship of the plurality of characters arranged in the utterance reference image, along the direction indicating the order in which they are spoken.
4. An information processing apparatus according to claim 1 or claim 2, The image display unit includes a playback speed changing unit that changes the playback speed of the speech reference image, and is an information processing device.
5. An information processing apparatus according to claim 1 or claim 2, The image display unit displays the edited image including the object on the display device based on the speech reference data. The information processing device further comprises an editing unit that edits the speech reference data in accordance with instructions from the user while the edited image is being displayed.
6. An information processing device according to claim 5, The instruction includes an instruction to change the position of the characters in the edited image, which is an information processing device.
7. An information processing method for supporting speech, A step of obtaining speech reference data having object data that identifies an object containing multiple characters representing the content of an utterance, and feature data that identifies speech features including at least one of the intonation and speech rate at the time of utterance, A step of displaying on a display device an utterance reference image, which is a moving image representing the object based on the utterance reference data, and which represents the utterance features based on the relative positional relationships between a plurality of characters arranged in the utterance reference image; Equipped with, The utterance reference data includes timing data that specifies the execution timing of each object, The aforementioned utterance reference image is an image that shows the execution timing of each of the aforementioned objects, The object includes at least one of a graphic representing an emotional expression during speech and a graphic representing a bodily movement during speech, in this information processing method.
8. A computer program for assisting speech, On the computer, A process for obtaining speech reference data having object data that identifies an object containing multiple characters representing the content of an utterance, and feature data that identifies speech features including at least one of the intonation and speech rate at the time of utterance, A process to display on a display device an utterance reference image, which is a moving image representing the object, based on the utterance reference data, and which represents the utterance features based on the relative positional relationships between a plurality of characters arranged in the utterance reference image; Make it run, The utterance reference data includes timing data that specifies the execution timing of each object, The aforementioned utterance reference image is an image that shows the execution timing of each of the aforementioned objects, The object is a computer program that includes at least one of a graphic representing an emotional expression during speech and a graphic representing a bodily movement during speech.
9. A data structure for speech reference data used in a computer comprising a display unit, a control unit, and a storage unit, wherein the speech reference data is data representing a speech reference image, which is a moving image that shows an object containing multiple characters representing the content of the speech, Object data that identifies the aforementioned object, Feature data that identifies speech features including at least one of intonation and speech rate, and is used in a process in which the control unit causes the display unit to display the speech reference image representing the speech features based on the relative positional relationships between a plurality of characters arranged in the speech reference image, Includes timing data that specifies the execution timing of each of the aforementioned objects, The aforementioned utterance reference image is an image that shows the execution timing of each of the aforementioned objects, The object is a data structure that includes at least one of a graphic representing an emotional expression during speech and a graphic representing a bodily movement during speech.