Information processing device, information processing method, computer program, and data structure
Patent Information
- Application Number
- JP2025106907
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2045-06-25
Smart Images

Figure 0007923584000001_ABST
Abstract
Description
Technical Field
[0001] The technology disclosed in the present specification relates to information processing for supporting utterance.
Background Art
[0002] There are countless opportunities to give utterances, ranging from daily conversations to presentations at international conferences. Conventionally, various devices for supporting utterance have been used. For example, prompters that support smooth utterance by a speaker by allowing the speaker to visually recognize an image showing a character string representing the content of utterance are widely used (see, for example, Patent Document 1).
Prior Art Literature
Patent Literature
[0003]
Patent Document 1
Summary of Invention
Problem to be Solved by the Invention
[0004] As one mode of utterance, there is manzai performance. Manzai is a form of entertainment in which, for example, two performers perform comical interactions for the purpose of making the audience laugh. The common way to enjoy manzai is to appreciate manzai performed by professional manzai comedians. In recent years, as a new way to enjoy manzai, ordinary people who are not professional manzai comedians have started performing manzai.
[0005] It is not easy for ordinary people to perform manzai well, for example, because it requires memorization of lines. It is also conceivable to support the utterance of lines by the performers by using a prompter to allow each performer to visually recognize an image showing each performer's lines. However, in order to perform manzai well, the timing of interaction between the two performers needs to be appropriate. Therefore, even if a conventional prompter is used, it cannot sufficiently support manzai performance.
[0006] These challenges are not limited to speech during stand-up comedy performances, but are common when supporting speech by multiple speakers in general, including everyday conversations.
[0007] This specification discloses a technology capable of solving the above-mentioned problems. [Means for solving the problem]
[0008] The technologies disclosed herein can be implemented, for example, in the following forms:
[0009] (1) The information processing device disclosed herein comprises a data acquisition unit and an image display unit. The data acquisition unit acquires speech reference data. The speech reference data includes object data that identifies an object containing a plurality of characters representing the content of each speaker's speech, and timing data that identifies the timing of each of the objects executed by each of the speakers. The image display unit causes a display device to display a speech reference image, which is a moving image showing the object, based on the speech reference data. The speech reference image is an image that shows the timing of each of the objects executed by each of the speakers in sync with each other.
[0010] According to this information processing device, in a speech reference image that shows an object containing multiple characters representing the content of each utterance by multiple speakers, the execution timing of each object by each speaker is shown in synchronous order. Therefore, each speaker can speak at the appropriate synchronous timing without having to memorize the string of characters representing the content of their utterance, simply by referring to a single speech reference image. Thus, this information processing device can effectively support speech by multiple speakers.
[0011] (2) In the above-described information processing device, the speech reference image may be an image in which the object associated with each speaker moves in a predetermined direction at a constant speed. With this configuration, each speaker can easily grasp the execution timing of each other's objects by referring to one speech reference image. Therefore, this information processing device can more effectively support speech by multiple speakers.
[0012] (3) In the above-described information processing device, the speech reference image may be an image in which the objects associated with each speaker are arranged in a direction that intersects the predetermined direction. With this configuration, each speaker can more easily grasp the execution timing of each other's objects by referring to one speech reference image. Therefore, this information processing device can more effectively support speech by multiple speakers.
[0013] (4) In the above-mentioned information processing device, the object may include a graphic that represents the emotional expression at the time of utterance. With this configuration, each speaker can make an utterance with an appropriate emotional expression by referring to the utterance reference image. Therefore, this information processing device can effectively support utterances by multiple speakers.
[0014] (5) In the above-mentioned information processing device, the object may include a graphic representing a body movement during speech. With this configuration, each speaker can perform speech with appropriate body movements by referring to the speech reference image. Therefore, this information processing device can effectively support speech by multiple speakers.
[0015] (6) In the above-described information processing device, the image display unit may include a playback speed changing unit that changes the playback speed of the speech reference image. With this configuration, each speaker can speak while referring to the speech reference image played back at their preferred speed. Therefore, this information processing device can effectively support speech by multiple speakers.
[0016] (7) In the above-mentioned information processing device, the speech reference data may have feature data that identifies speech features including at least one of intonation and speech speed, and the speech reference image may be an image that represents the speech features based on the relative positional relationship between the characters. With this configuration, each speaker can speak with appropriate intonation and / or speed by referring to the speech reference image. Therefore, this information processing device can effectively support speech by multiple speakers.
[0017] (8) In the above-described information processing device, the speech features may include intonation during speech, and the speech reference image may be an image that represents intonation based on the positional relationship along a direction that intersects the direction indicating the order in which the plurality of characters are spoken. With this configuration, each speaker can intuitively grasp the appropriate intonation by referring to the speech reference image. Therefore, this information processing device can more effectively support speech by multiple speakers.
[0018] (9) In the above-described information processing device, the speech feature may include speech speed, and the speech reference image may be an image that represents speech speed based on the positional relationship along the direction indicating the order in which the plurality of characters are spoken. With this configuration, each speaker can intuitively grasp an appropriate speech speed by referring to the speech reference image. Therefore, this information processing device can more effectively support speech by multiple speakers.
[0019] (10) In the above-mentioned information processing device, the image display unit displays an edited image including the object on the display device based on the speech reference data, and the information processing device may further include an editing unit that edits the speech reference data in accordance with instructions from the user while the edited image is being displayed. With this configuration, the speech reference image can be edited into an image that represents more appropriate speech features, and speech by multiple speakers can be supported more effectively.
[0020] The technology disclosed in the present specification can be implemented in various forms, for example, it can be implemented in the form of an information processing apparatus and method, a speech support apparatus and method, a computer program that implements these methods, a data structure, data having a data structure, a non-transitory recording medium that records the computer program or the data, and the like. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] [Figure 1] Explanatory diagram showing an example of the speech reference image RI in the present embodiment [Figure 2] Explanatory diagram showing an example of the speech reference image RI in the present embodiment [Figure 3] Explanatory diagram showing an example of speech reference data RD representing the speech reference image RI [Figure 4] Explanatory diagram showing an example of the edited image EI in the present embodiment [Figure 5] Explanatory diagram showing the schematic configuration of the information processing apparatus 100 [Figure 6] Flowchart showing speech reference image reproduction processing [Figure 7] Flowchart showing editing processing [Figure 8] Explanatory diagram showing an example of the speech reference image RI in a modified example MODES FOR CARRYING OUT THE INVENTION
[0022] Embodiment Overview of Speech Reference Image RI FIG. 1 and FIG. 2 are explanatory diagrams showing an example of the speech reference image RI in the present embodiment. The speech reference image RI is a moving image for supporting the performance of manzai including utterance of lines by a performer, by being visually recognized by the performer when, for example, an ordinary person who is not a professional manzai comedian performs manzai. The speech reference image RI functions as a musical score for the performer. The speech reference image RI of the present embodiment is referred to when two performers (so-called "boke" and "tsukkomi") perform manzai. FIG. 1 and FIG. 2 exemplify one scene of the speech reference image RI at mutually different timings.
[0023] The speech reference image RI includes the background image BI and the musical score image SI. The background image BI is the background image for the speech reference image RI and is a static image. The background image BI includes the first lane L1 and the second lane L2, the playback-related button PB, and the speed adjustment button SB. Hereafter, the first lane L1 and the second lane L2 will be collectively referred to as lanes L1 and L2.
[0024] Each lane L1 and L2 extends linearly from the right edge to the left edge of the speech reference image RI. The second lane L2 is located below the first lane L1. In this embodiment, the first lane L1 is an area for displaying the content of speech by the straight man, and the second lane L2 is an area for displaying the content of speech by the funny man. A linear timing bar TB extending vertically is placed on each lane L1 and L2. The timing bar TB is positioned slightly to the left of the center of the speech reference image RI. The left-right position of the timing bar TB on the first lane L1 is the same as the left-right position of the timing bar TB on the second lane L2.
[0025] The playback-related buttons PB include a play button P1, a pause button P2, a fast-forward button P3, and a rewind button P4. By operating the playback-related buttons PB, the user can play, pause, fast-forward, and rewind the speech reference image RI.
[0026] The speed adjustment button SB includes a constant speed button S1, an acceleration button S2, and a deceleration button S3. By operating the speed adjustment button SB, the user can adjust the playback speed of the speech reference image RI to their preferred speed. The constant speed button S1 is a button that matches the playback speed of the speech reference image RI to the actual performance speed of a professional comedian. Each time the acceleration button S2 is pressed, the playback speed of the speech reference image RI increases by, for example, 5%. Each time the deceleration button S3 is pressed, the playback speed of the speech reference image RI decreases by, for example, 5%.
[0027] The musical score image SI is an image that shows the performance content (speech content, emotional expression, and physical movements) by each performer. The musical score image SI contains multiple letters CH. The multiple letters CH are images that show the speech content by each performer. The multiple letters CH are arranged from left to right in the order of utterance.
[0028] In this embodiment, the score image SI further includes an effect image EF and a motion panel image MP. Hereinafter, the components of the score image SI (character CH, effect image EF, motion panel image MP) will be collectively referred to as an object.
[0029] Effect images EF are partial images that represent the emotional expression of each performer through their design. Figures 1 and 2 show examples of effect images EF1 representing the emotion of "anger" and effect image EF2 representing the emotion of "joy". The emotion of "anger" can be used, for example, when a performer playing the straight man role makes a retort, and the emotion of "joy" can be used, for example, when a performer playing the funny man role makes a joke. Effect images EF corresponding to other emotions ("sadness" or "happiness") may also be used. Effect images EF may be displayed alone or overlaid on text CH or motion panel images MP, which will be described later.
[0030] Motion panel images MP are partial images that show the physical movements of each performer, indicated by their patterns. Figures 1 and 2 illustrate motion panel images MP1, which shows a straight man performer making a retort, and motion panel image MP2, which shows a funny man performer leaning back and laughing. Motion panel images MP corresponding to other physical movements may also be used. Motion panel images MP may be displayed alone or overlaid on text CH or effect images EF.
[0031] During playback of the speech reference image RI, each object constituting the score image SI (character CH, effect image EF, motion panel image MP) moves synchronously at a constant speed from right to left along lanes L1 and L2 in the background image BI. The object for the straight man character moves along the first lane L1, and the object for the funny man character moves along the second lane L2. In other words, the objects associated with each character are displayed side by side in the vertical direction. In each lane L1 and L2, the timing at which an object reaches the timing bar TB is the execution timing of that object. The execution timing of an object is the timing of speech for character CH, the timing of expressing emotion for effect image EF, and the timing of performing a physical action for motion panel image MP. Each character can execute an object at the appropriate timing by referring to the positional relationship between each object in their own lane L1 and L2 on the speech reference image RI and the timing bar TB. In other words, each character can deliver lines with the appropriate emotion or perform appropriate physical actions at the appropriate timing.
[0032] Furthermore, in the speech reference image RI, the timing of the dialogue between the two performers is determined by the difference between the timing at which each object for one performer reaches the timing bar TB and the timing at which each object for the other performer reaches the timing bar TB. In this way, the speech reference image RI is an image that shows the execution timing of each object by each performer in sync with each other.
[0033] Each object, after reaching the timing bar TB, moves past the timing bar TB towards the left edge. In this embodiment, the display of an object after passing the timing bar TB is different from the display of an object before passing the timing bar TB. For example, an object after passing the timing bar TB is displayed in gray.
[0034] In each lane L1 and L2 of the speech reference image RI, the spacing between multiple characters CH in the left-right direction (hereinafter referred to as "left-right spacing") is variable. For example, a narrow left-right spacing between multiple characters CH constituting a line of dialogue indicates that the line is spoken relatively quickly. Conversely, a wide left-right spacing between multiple characters CH constituting a line of dialogue indicates that the line is spoken relatively slowly. Thus, in the speech reference image RI of this embodiment, speech speed is expressed by the relative positional relationship of multiple characters CH in the left-right direction. Each performer can speak lines at an appropriate speed by referring to the relative positional relationship of multiple characters CH in the left-right direction in the speech reference image RI. The left-right direction is the direction that indicates the order in which the multiple characters CH are spoken. Speech speed is an example of a speech feature.
[0035] In each lane L1 and L2 of the speech reference image RI, the vertical position of each character CH (hereinafter referred to as "vertical position") is variable. In this embodiment, the vertical position of character CH is set to five levels (top, slightly up, center, slightly down, and bottom). The vertical position of character CH represents the intonation (pitching) of the performer's speech. For example, if the vertical position of character CH is top, slightly up, center, slightly down, and bottom, it represents high, slightly high, normal, slightly low, and low pitch, respectively. Thus, in the speech reference image RI of this embodiment, the intonation of speech is expressed by the relative vertical positional relationship between multiple character CHs. Each performer can refer to the relative vertical positional relationship of multiple character CHs in the speech reference image RI and deliver their lines with appropriate intonation. The vertical direction intersects with the direction (left-right direction) that indicates the order in which the multiple character CHs are spoken. Intonation is one example of speech characteristics.
[0036] In the speech reference image RI, the font of each character CH is variable. The font of each character CH represents the emotional expression at the time of speech. In this embodiment, eight types of fonts (yorokobi / kanashimi / odoroki / horror / neutral / tsukkomi / boke / emphasize) are used. For example, among joy, anger, sorrow, and pleasure, the emotions of "joy" and "pleasure" are represented by the font "yorokobi," the emotion of "anger" is represented by the font "tsukkomi," and the emotion of "sorrow" is represented by the font "kanashimi." The emotion of "anger" is used, for example, in the lines of a straight man's retort, while the emotions of "joy" and "pleasure" are used, for example, in the lines of a funny man's joke. Each performer can refer to the font of each character CH in the speech reference image RI and deliver their lines with the appropriate emotion. The emotional expression at the time of speech is an example of speech features.
[0037] In this way, each performer can perform a comedy routine with appropriate timing, speed, intonation, emotional expression, and physical movements by referring to the speech reference image RI, without having to memorize lines.
[0038] Figure 3 is an explanatory diagram showing an example of speech reference data RD representing a speech reference image RI. In this embodiment, the speech reference data RD includes CSV data that defines a musical score image SI. As shown in Figure 3, the CSV data included in the speech reference data RD consists of multiple rows. Each row of this data consists of five elements: "Target", "Content", "Timing", "Option", and "Font".
[0039] The element "Target" indicates the target performer and object. This element uses the codes "A" and "B" to indicate the performer's classification (straight man, funny man), and "(unsigned)", "E", and "M" to indicate the object's classification (text CH, effect image EF, motion panel image MP). For example, the value of the element "Target" for text CH targeting the funny man is "B", the value of the element "Target" for effect image EF targeting the straight man is "AE", and the value of the element "Target" for motion panel image MP targeting the funny man is "BM".
[0040] The element "Content" indicates the content of the object. For text elements CH, this element shows the text itself; for effect images EF, it indicates which of the four emotions (Joy, Angry, Sad, Happy) it represents; and for motion panel images MP, it shows the file name. The element "Content" is an example of object data.
[0041] The "Timing" element indicates the execution timing of an object. This element shows the elapsed time (in seconds) from the start time of playback of the speech reference image RI until the object reaches the position of the timing bar TB. The difference in the value of the "Timing" element between two objects determines the left-right spacing between those two objects in the speech reference image RI. The "Timing" element is an example of timing data that identifies the execution timing of each object, and also an example of feature data that identifies the speech rate.
[0042] The element "Option" indicates the vertical position of an object. This element uses the symbols "H", "h", "m", "l", and "L" to represent five vertical positions (top, slightly up, center, slightly down, bottom). The element "Option" is an example of feature data that identifies intonation during speech.
[0043] The element "Font" indicates the font of the object. In this embodiment, eight fonts are available as options, and the element "Font" indicates one of these fonts. In this embodiment, for objects other than the character CH, the element "Font" is set to the default value of "neutral". The element "Font" is an example of feature data that identifies emotional expression during speech.
[0044] For example, in the second row of the CSV data shown in Figure 3, the value of the element "Target" is "B", the value of the element "Content" is "っ", the value of the element "Timing" is "2.480633", the value of the element "Option" is "m", and the value of the element "Font" is "yorokobi". Therefore, this second row indicates that, as a musical score image SI included in the speech reference image RI, the character "っ" represented by the font "yorokobi" will be displayed at the vertical position = center position in the second lane L2 for the comedic role, at a timing such that it reaches the timing bar TB 2.480633 seconds after the start of playback. The comedic actor, referring to the speech reference image RI, is prompted to pronounce the character "っ" with a moderate intonation (pitching) and expressing joy at the above timing.
[0045] Furthermore, in the 21st row of the CSV data shown in Figure 3, the value of the element "Target" is "BM", the value of the element "Content" is "gray_kasu_tegaki5", the value of the element "Timing" is "2.387979", the value of the element "Option" is "l", and the value of the element "Font" is "neutral". Therefore, this 21st row indicates that, as a musical score image SI included in the speech reference image RI, the motion panel image MP (backward laughter) with the file name "gray_kasu_tegaki5" will be displayed at a position slightly below the vertical direction in the second lane L2 for the comedic character, at a timing such that it reaches the timing bar TB 2.387979 seconds after the start of playback. The comedic character performer, referring to the speech reference image RI, will be prompted to perform the backward laughter motion at the above timing.
[0046] Furthermore, in row 23 of the CSV data shown in Figure 3, the value of the element "Target" is "AE", the value of the element "Content" is "Angry", the value of the element "Timing" is "2.006745", the value of the element "Option" is "L", and the value of the element "Font" is "neutral". Therefore, row 23 indicates that, as a musical score image SI included in the speech reference image RI, the effect image EF representing the emotion of anger will be displayed at the bottom position in the first lane L1 for the straight man, at a timing such that it reaches the timing bar TB 2.006745 seconds after the start of playback. The performer of the straight man, referring to the speech reference image RI, will be prompted to express the emotion of anger at the above timing.
[0047] Thus, the speech reference data RD representing the speech reference image RI includes object data that identifies an object containing multiple characters CH representing the content of the utterance, and feature data that identifies speech features including at least one of the intonation and speech rate during the utterance. The speech reference data RD causes the speech reference image RI, which represents the speech features based on the relative positional relationships between the multiple characters CH, to be displayed on a display device.
[0048] The speech reference data RD with the above configuration is generated, for example, by the following method. The characters CH that make up each performer's lines are identified by extracting audio data and separating speakers from videos of professional comedians performing stand-up comedy. The intonation of the lines is obtained from the above audio data using Fourier transform, etc., and these frequencies are classified into the five levels of intonation (pitching) described above.
[0049] The effect images (EF) and fonts representing the performers' emotional expressions will be identified by performing emotion recognition based on the content of the dialogue, voice, and facial expressions in the aforementioned comedy performance video. The motion panel images (MP) representing the performers' physical movements will be identified by performing motion capture on the aforementioned comedy performance video. For the effect images (EF) and motion panel images (MP), the closest image will be assigned from a pre-prepared set of images. If no close image exists, a new one may be created.
[0050] The execution timing of each object will be determined according to the execution timing of each object in the above-mentioned comedy performance video. For physical movements, for example, the start timing of movements deemed important for the success of the routine will be determined.
[0051] Based on the identified lines (strings of dialogue), emotional expressions, and bodily movements, along with the timing of their execution, speech reference data (RD) representing them is created.
[0052] (Overview of edited image EI) Figure 4 is an explanatory diagram showing an example of an edited image EI in this embodiment. The edited image EI is an image used for editing the speech reference image RI. The edited image EI includes a background image BIe and a musical score image SI, similar to the speech reference image RI described above. The layout of the background image BIe of the edited image EI is substantially the same as the layout of the background image BI of the speech reference image RI. However, the background image BIe of the edited image EI includes several interfaces, which are described below, for editing the speech reference image RI. Note that the musical score image SI of the edited image EI is the same as the musical score image SI of the speech reference image RI, and is therefore omitted from the illustration in Figure 4.
[0053] The background image Bie of the edited image EI includes a dialogue input field E1. When the editor selects either lane L1 or L2 and enters one or more characters CH representing dialogue into the dialogue input field E1, the entered characters CH are displayed in the selected lane. The editor can change the intonation of the dialogue by changing the vertical position of each displayed character CH, for example by dragging with the mouse. The editor can also change the speaking speed and timing by changing the horizontal position of each displayed character CH. Note that changing the horizontal position of a character CH will consequently change the horizontal spacing between that character CH and other adjacent character CHs, and thus change the speaking speed of those character CHs. The editor can also set the font of the character CH, thereby setting the emotional expression of the dialogue. The background image Bie of the edited image EI also includes a furigana input field E2 for entering the phonetic readings of the dialogue.
[0054] The background image Bie of the editing image EI includes the part addition button E5. When the editor selects the part addition button E5, a library screen (not shown) opens displaying candidate objects (effect image EF and motion panel image MP). When the editor selects an object from the library screen and also selects a lane, the selected object is displayed in the selected lane. The editor can change the execution timing of the object by changing its left-right position.
[0055] The background image Bie of the edited image EI includes a right-scroll button E6 and a left-scroll button E7. When the editor selects the right-scroll button E6, each object (score image SI) in the edited image EI scrolls, displaying the image corresponding to a later scene in time. Conversely, when the editor selects the left-scroll button E7, each object in the edited image EI scrolls in the opposite direction, displaying the image corresponding to an earlier scene in time.
[0056] The background image Bie of the edited image EI includes a play button E3 and a pause button E4. When the editor selects the play button E3, the latest image at that time will be played, and when the pause button E4 is selected during playback, playback will be paused.
[0057] The background image BIe of the edited image EI includes the data output button E8. When the editor selects the data output button E8, speech reference data RD is generated, representing the edited speech reference image RI at that time. This enables the updating of the speech reference data RD.
[0058] (Configuration of the information processing device 100) Next, the configuration of the information processing device 100 for performing various processes related to the speech reference image RI will be described. Figure 5 is an explanatory diagram showing the schematic configuration of the information processing device 100. The information processing device 100 is composed of, for example, a computer (PC, server, smartphone, tablet terminal, etc.).
[0059] The information processing device 100 comprises a control unit 110, a storage unit 120, a display unit 130, an operation input unit 140, an interface unit 150, and an audio input / output unit 160. Each of these units is connected to the others via a bus 190 so as to be able to communicate with each other.
[0060] The display unit 130 of the information processing device 100 is configured, for example, as a liquid crystal display, and displays various images and information. The display unit 130 is an example of an output device and display device. The operation input unit 140 is configured, for example, as a keyboard, mouse, buttons, microphone, trackpad, etc., and accepts user operations and instructions. The display unit 130 may also function as the operation input unit 140 by being equipped with a touch panel. The interface unit 150 is configured, for example, as a LAN interface or USB interface, and communicates with other devices by wired or wireless connection. The audio input / output unit 160 is configured, for example, as a microphone or speaker, and inputs and outputs audio.
[0061] The storage unit 120 is composed of, for example, ROM, RAM, a hard disk drive (HDD), and stores various programs and data, and is used as a workspace and temporary storage area for data when executing various programs. For example, the storage unit 120 stores a speech support processing program CP, which is a computer program for executing various processes described later. The speech support processing program CP is provided, for example, stored on a computer-readable recording medium (not shown) such as a CD-ROM, DVD-ROM, or USB memory, or is provided in a state that can be obtained from an external device (a server on a network or other terminal device) via the interface unit 150, and is stored in the storage unit 120 in a state that can be operated on the information processing device 100. The storage unit 120 also stores speech reference data RD, which represents a speech reference image RI.
[0062] The control unit 110 is composed of, for example, a CPU, and controls the operation of the information processing device 100 by executing a computer program read from the storage unit 120. For example, the control unit 110 functions as a speech support processing unit 111 for executing various processes described later by reading and executing a speech support processing program CP from the storage unit 120. The speech support processing unit 111 includes a data acquisition unit 112, an image display unit 113 including a playback speed change unit 114, and an editing unit 115. The functions of each of these units will be explained in accordance with the descriptions of the various processes described later.
[0063] (Speech reference image playback processing) Next, the speech reference image playback process performed by the information processing device 100 of this embodiment will be described. Figure 6 is a flowchart of the speech reference image playback process. The speech reference image playback process is a process that displays (plays) a speech reference image RI, for example, to support a stand-up comedy performance by performers. The speech reference image playback process is started when a user operates the operation input unit 140 of the information processing device 100 and inputs a start command.
[0064] First, the data acquisition unit 112 (Figure 5) of the information processing device 100 acquires speech reference data RD (S110). The data acquisition unit 112 may acquire speech reference data RD from an external device via the interface unit 150, or it may acquire speech reference data RD by generating it itself. The acquired speech reference data RD is stored in the storage unit 120.
[0065] Next, the image display unit 113 (Figure 5) of the information processing device 100 reproduces and displays the speech reference image RI on the display unit 130 based on the speech reference data RD (S120). The performer performs the comedy routine by referring to the reproduced speech reference image RI. The audio during the comedy routine performance may be recorded via the audio input / output unit 160. The physical movements during the comedy routine performance may be recorded via a camera or the like (not shown).
[0066] The speech support processing unit 111 (Figure 5) of the information processing device 100 monitors for any user requests to change the playback speed during playback of the speech reference image RI (S130). These playback speed change requests are input, for example, via the speed adjustment button SB on the speech reference image RI. If a playback speed change request is received (S130: YES), the playback speed change unit 114 changes the playback speed of the speech reference image RI according to the request (S140).
[0067] Although not shown in the diagram, if a user input (pause, fast forward, rewind) is received, for example via the playback-related button PB, during playback of the speech reference image RI, the image display unit 113 will pause, fast forward, or rewind the speech reference image RI according to the input.
[0068] The speech support processing unit 111 (Figure 5) of the information processing device 100 monitors whether the speech reference image RI has been played back to the end (S150). If the speech reference image RI is played back to the end (S150: YES), the speech support processing unit 111 terminates the speech reference image playback process.
[0069] (Editing process) Next, the editing process performed by the information processing device 100 of this embodiment will be described. Figure 7 is a flowchart of the editing process. The editing process is the process of editing and updating the speech reference image RI. The editing process is started when the user operates the operation input unit 140 of the information processing device 100 and inputs a start command.
[0070] First, the image display unit 113 (Figure 5) of the information processing device 100 displays the edited image EI on the display unit 130 based on the speech reference data RD stored in the storage unit 120 (S210).
[0071] Next, the editing unit 115 of the information processing device 100 updates the edited image EI according to the editing instructions from the user (S220). For example, if the user instructs the placement of an object, the editing unit 115 updates the edited image EI so that the object is displayed at the instructed position in the edited image EI. Also, if the user instructs the position of an object to be changed, the editing unit 115 updates the edited image EI so that the position of the object in the edited image EI is changed.
[0072] The speech support processing unit 111 (Figure 5) of the information processing device 100 monitors for the presence or absence of a data output instruction from the user (S230). A data output instruction is input, for example, via the data output button E8 in the edited image EI. When a data output instruction is received (S230: YES), the editing unit 115 updates the speech reference data RD according to the editing results at that time (S240) and terminates the editing process.
[0073] (Effects of this embodiment) As described above, the information processing device 100 of this embodiment comprises a data acquisition unit 112 and an image display unit 113. The data acquisition unit 112 acquires speech reference data RD having object data that identifies an object containing a plurality of characters CH representing the content of an utterance, and feature data that identifies speech features including at least one of the intonation and speech speed at the time of utterance. Based on the speech reference data RD, the image display unit 113 displays a speech reference image RI, which is a moving image representing an object, on the display unit 130. The speech reference image RI is an image that represents the above-mentioned speech features based on the relative positional relationship between the plurality of characters CH.
[0074] Thus, in the information processing device 100 of this embodiment, in the speech reference image RI which shows an object containing multiple characters CH representing the content of the utterance, speech features including at least one of the intonation and speech speed during utterance are represented by the relative positional relationship between the multiple characters CH. Therefore, by referring to the speech reference image RI, the speaker can utter (speak the lines) with appropriate intonation and / or speed without having to memorize the string of characters (lines) representing the content of the utterance. Accordingly, the information processing device 100 of this embodiment can effectively support the speaker's utterance.
[0075] In this embodiment, the speech reference image RI is an image that represents the intonation of speech based on the positional relationship along a direction (e.g., vertical direction) that intersects with the direction (e.g., left-right direction) indicating the pronunciation order of multiple characters CH. Therefore, the speaker can intuitively grasp the appropriate intonation by referring to the speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support the speaker's speech.
[0076] In this embodiment, the speech reference image RI is an image that represents speech speed based on the positional relationship of multiple characters CH along a direction (e.g., left-right direction) indicating the order in which they are spoken. Therefore, the speaker can intuitively grasp the appropriate speech speed by referring to the speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support the speaker's speech.
[0077] In this embodiment, the speech reference data RD includes timing data that specifies the execution timing of each object, and the speech reference image RI is an image that shows the execution timing of each object. Therefore, the speaker can execute objects at the appropriate timing by referring to the speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support the speaker's utterances.
[0078] In this embodiment, the object includes a graphic representing the emotional expression at the time of utterance. Therefore, the speaker can make an utterance with appropriate emotional expression by referring to the utterance reference image RI. Accordingly, the information processing device 100 of this embodiment can effectively support the speaker's utterance.
[0079] In this embodiment, the object includes a graphic representing a body movement during speech. Therefore, the speaker can perform speech with appropriate body movements by referring to the speech reference image RI. Accordingly, the information processing device 100 of this embodiment can effectively support speech by the speaker.
[0080] In this embodiment, the image display unit 113 includes a playback speed changing unit 114 that changes the playback speed of the speech reference image RI. Therefore, the speaker can speak while referring to the speech reference image RI played back at their preferred speed. Accordingly, the information processing device 100 of this embodiment can effectively support the speaker's speech.
[0081] In this embodiment, the image display unit 113 displays an edited image EI, which includes an object, on the display unit 130 based on the speech reference data RD. The information processing device 100 further includes an editing unit 115 that edits the speech reference data RD according to instructions from the user while the edited image EI is being displayed. As a result, the speech reference image RI can be edited to represent more appropriate speech features, thereby more effectively supporting the speaker's utterance. In this embodiment, the user's instructions include instructions to change the position of the character CH in the edited image EI. As a result, the editor can intuitively give editing instructions so that the speech reference image RI becomes an image that represents more appropriate speech features.
[0082] The data structure of the speech reference data RD in this embodiment includes object data that identifies an object containing multiple characters CH, and feature data that identifies speech features including at least one of intonation and speech rate during speech. The data structure is such that a speech reference image RI representing speech features based on the relative positional relationship between the multiple characters CH is displayed on the display unit 130. According to the data structure of the speech reference data RD in this embodiment, a speech reference image RI that can effectively support speech by the speaker can be displayed.
[0083] The information processing device 100 of this embodiment includes a data acquisition unit 112 and an image display unit 113. The data acquisition unit 112 acquires speech reference data RD, which includes object data that identifies an object containing multiple characters CH representing the content of each of multiple speakers' speeches, and timing data that identifies the execution timing of each object by each speaker. Based on the speech reference data RD, the image display unit 113 displays a speech reference image RI, which is a moving image showing an object, on the display unit 130. The speech reference image RI is an image that shows the execution timing of each object by each speaker synchronized with each other.
[0084] Thus, in the information processing device 100 of this embodiment, in the speech reference image RI which shows an object containing multiple characters CH representing the speech content of each of the multiple speakers, the execution timing of each object by each speaker is shown to be synchronized with each other. Therefore, by referring to a single speech reference image RI, each speaker can speak (speak lines) at appropriate, synchronized timings without having to memorize the string of characters (lines) that represent the speech content. Accordingly, the information processing device 100 of this embodiment can effectively support speech by multiple speakers.
[0085] In this embodiment, the speech reference image RI is an image in which an object associated with each speaker moves at a constant speed in a predetermined direction (for example, left or right). Therefore, each speaker can easily grasp the execution timing of each other's objects by referring to a single speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support speech by multiple speakers.
[0086] In this embodiment, the speech reference image RI is an image in which objects associated with each speaker are arranged in a direction that intersects the direction of movement of the objects (e.g., left-right direction). Therefore, each speaker can more easily grasp the execution timing of each other's objects by referring to a single speech reference image RI. Accordingly, the information processing device 100 of this embodiment can more effectively support speech by multiple speakers.
[0087] The data structure of the speech reference data RD in this embodiment comprises object data that identifies an object containing multiple characters CH representing the content of each utterance by multiple speakers, and timing data that identifies the execution timing of each object by each speaker. The data structure is used to display a speech reference image RI on the display unit 130, which shows the execution timing of each object by each speaker synchronized with each other. According to the data structure of the speech reference data RD in this embodiment, it is possible to display a speech reference image RI that can effectively support utterances by multiple speakers.
[0088] (modified version) The technologies disclosed herein are not limited to the embodiments described above and can be modified in various forms without departing from their essence, for example, the following modifications are possible.
[0089] The configuration of the information processing device 100 in the above embodiment is merely an example and can be modified in various ways. Similarly, the content of the various processes in the above embodiment is merely an example and can be modified in various ways. For example, in the above embodiment, the speech reference image RI and the edited image EI are displayed on the display unit 130 of the information processing device 100, but these images may also be displayed on an external display device.
[0090] In the above embodiment, the speaker's body movements are shown using a motion panel image MP in the speech reference image RI. However, as shown in the modified example in Figure 8, the speaker's body movements may also be shown using a body movement video MC. The body movement video MC is, for example, a video generated by motion capture from a performance by a professional comedian.
[0091] In the speech reference image RI of the above embodiment, some or all of the objects associated with multiple speakers may be displayed overlapping each other.
[0092] In the above embodiment, the speech reference image RI is an image showing the content of speech by two speakers, but the speech reference image RI may also be an image showing the content of speech by one speaker, or an image showing the content of speech by three or more speakers.
[0093] In the above embodiment, at least one of the functional units included in the control unit 110 of the information processing device 100 may be included in another device instead of the control unit 110 of the information processing device 100. The various processes in the above embodiment do not necessarily have to be performed by a single device, but may be performed by different devices. Each step of the various processes in the above embodiment does not necessarily have to be performed by a single device, but may be performed by different devices. In the above embodiment, some of the configurations implemented by hardware may be replaced with software, and conversely, some of the configurations implemented by software may be replaced with hardware.
[0094] In the embodiments described above, a stand-up comedy performance was used as an example of one form of speech, but the technology disclosed herein is similarly applicable to support speech in general, including everyday conversation and presentations at international conferences. For example, the technology disclosed herein can be applied to karaoke, allowing users to practice imitating the singing style (intonation, etc.) of artists by referring to speech reference images. Furthermore, the technology disclosed herein can be applied to dubbing movies, allowing users to refer to speech reference images to encourage emotional expression that more closely resembles the atmosphere of the original work, and to make the timing of each performer's lines clearer. [Explanation of Symbols]
[0095] 100: Information processing unit 110: Control unit 111: Speech support processing unit 112: Data acquisition unit 113: Image display unit 114: Playback speed change unit 115: Editing unit 120: Memory unit 130: Display unit 140: Operation input unit 150: Interface unit 160: Audio input / output unit 190: Bus BI: Background image BIe: Background image CH: Text CP: Speech support processing program E1: Dialogue input field E2: Furigana input field E3: Play button E4: Pause button E5: Add part button E6: Screen right scroll button E7: Screen left scroll button E8: Data output button EF: Effect image EI: Editing image L1: First lane L2: Second lane MC: Body movement video MP: Motion panel image P1: Play button P2: Pause button P3: Fast forward button P4: Rewind button PB: Playback related buttons RD: Speech reference data RI: Speech reference image S1: Constant speed button S2: Acceleration button S3: Deceleration button SB: Speed adjustment button SI: Score image TB: Timing bar
Claims
1. An information processing device for supporting speech by multiple speakers who speak while viewing moving images, A data acquisition unit acquires speech reference data having object data that identifies an object containing multiple characters representing the content of each speaker's speech, and timing data that identifies the execution timing of each object by each speaker, which is the execution timing of the action corresponding to the object by the speaker. Based on the aforementioned speech reference data, an image display unit displays the speech reference image, which is a moving image representing the object, in a manner that synchronizes the execution timing of each of the objects by each speaker with each other. Equipped with, The utterance reference image is an image in which the object associated with each of the utterances moves in a predetermined direction at a constant speed, and the image indicates the execution timing of the object when the object reaches a predetermined position, in this information processing device.
2. An information processing apparatus according to claim 1, The utterance reference image is an information processing device in which the objects associated with each utterance are arranged in a direction that intersects the predetermined direction.
3. An information processing apparatus according to claim 1 or claim 2, The aforementioned object is an information processing device that includes a graphic representing the emotional expression at the time of utterance.
4. An information processing apparatus according to claim 1 or claim 2, The aforementioned object is an information processing device that includes a graphic representing a body movement during speech.
5. An information processing apparatus according to claim 1 or claim 2, The image display unit includes a playback speed changing unit that changes the playback speed of the speech reference image, and is an information processing device.
6. An information processing apparatus according to claim 1 or claim 2, The aforementioned speech reference data includes feature data that identifies speech features, which include at least one of the intonation and speech rate during speech. The utterance reference image is an image that represents the utterance features based on the relative positional relationship between the characters, in an information processing device.
7. An information processing apparatus according to claim 6, The aforementioned speech features include intonation during speech, The aforementioned speech reference image is an image that represents intonation based on the positional relationship along a direction that intersects the direction indicating the order in which the multiple characters are spoken, and is an information processing device.
8. An information processing apparatus according to claim 6, The aforementioned speech features include speech rate, The aforementioned speech reference image is an image that represents the speech speed based on the positional relationship along the direction indicating the order in which the multiple characters are spoken, and is an information processing device.
9. An information processing apparatus according to claim 1 or claim 2, The image display unit displays the edited image including the object on the display device based on the speech reference data. The information processing device further comprises an editing unit that edits the speech reference data in accordance with instructions from the user while the edited image is being displayed.
10. An information processing method for supporting speech by multiple speakers who speak while viewing moving images, A step of acquiring speech reference data having object data that identifies an object containing multiple characters representing the content of each speaker's speech, and timing data that identifies the execution timing of each of the objects by each speaker, which is the execution timing of the action corresponding to the object by the speaker. A step of displaying on a display device an utterance reference image, which is an utterance reference image that Equipped with, The utterance reference image is an image in which the object associated with each of the utterances moves in a predetermined direction at a constant speed, and the image indicates the execution timing of the object when the object reaches a predetermined position, in this information processing method.
11. A computer program for supporting speech by multiple speakers who speak while viewing moving images, On the computer, A process for obtaining utterance reference data having object data that identifies an object containing multiple characters representing the content of each utterance, and timing data that identifies the execution timing of each object by each utterance, which is the execution timing of the action corresponding to the object by the utterance; Based on the aforementioned speech reference data, the process involves displaying the speech reference image, which is a moving image representing the object, on a display device, in which the execution timing of each of the objects by each speaker is synchronized with each other. Make it run, A computer program in which the utterance reference image is an image in which the object associated with each of the utterances moves in a predetermined direction at a constant speed, and the timing of the execution of the object is indicated by the timing when the object reaches a predetermined position.
12. Speech reference data for use in a computer comprising a display unit, a control unit and a storage unit, wherein the speech reference image is a moving image that shows an object containing a plurality of characters representing the content of each of the speeches of a plurality of speakers who speak while viewing the moving image, and the speech reference data is stored in the storage unit, the data structure of the speech reference data The system comprises object data that identifies the object, and timing data that identifies the execution timing of each object by each speaker, which is the execution timing of the action corresponding to the object by the speaker. The object data and timing data are data structures used by the control unit to display the utterance reference image on the display unit, which is an image in which the objects associated with each speaker move at a constant speed in a predetermined direction, and the timing of the object reaching a predetermined position indicates the timing of the execution of the object.
Citation Information
Patent Citations
Prompter
JP2016212299A