Processing device, method for processing, and program
The processing device addresses the issue of unintended facial expressions during silence in lip synchronization by adding a second motion to the mouth before silence, ensuring natural transitions and reducing stiffness.
Patent Information
- Application Number
- JP2024010319
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-26
- Publication Date
- 2025-08-07
- Estimated Expiration
- 2044-01-26
AI Technical Summary
Lip synchronization in characters or avatars results in unintended facial expressions, such as a stiff mouth closed during silence, leading to unnatural transitions and prolonged stiff expressions, which can cause various production issues.
A processing device and method that adds a second motion to the mouth of a drawing target from a predetermined time before silence, blending it with the first motion to ensure a smooth transition and avoid stiff expressions.
Reduces the inconvenience of unintended facial expressions during silence by smoothly transitioning mouth movements, maintaining natural facial expressions and preventing unnatural transitions.
Smart Images

Figure 2025115719000001_ABST
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to a processing device, a processing method, and a program. [Background technology]
[0002] Characters, avatars, etc. are drawn on game screens and various UI (user interface) screens, and by moving their mouths in time with the dialogue, they can be drawn more realistically. This technology is called "lip sync." Related technologies are disclosed in Patent Documents 1 and 2.
[0003] The technology disclosed in Patent Document 1 prepares lip movements that utter monosyllables, words, or sentences as video fragments, reads out the video fragments that correspond to the lines, and plays back the video. Then, in the part where two video fragments are joined, the technology synthesizes the two video fragments to create a video, and inserts the created video. The technology disclosed in Patent Document 1 creates a video by synthesizing two video fragments that perform lip movements that correspond to the lines.
[0004] The technology disclosed in Patent Document 2 adjusts the application rate (proportion) of mouth motion in accordance with the lines, thereby causing the character's mouth to perform a desired motion. [Prior art documents] [Patent documents]
[0005] [Patent Document 1] Patent Publication No. 11-226050 [Patent Document 2] Patent Publication No. 2019-57116 Summary of the Invention [Problem to be solved by the invention]
[0006] As a result of examining lip synchronization, the present inventors have newly discovered the following problem.
[0007] When lip sync is used, the mouth of the character or avatar will be in a corresponding state, typically with the mouth tightly closed, during periods of silence (when not speaking). For this reason, the facial expression of the character or avatar will tend to be relatively stiff during periods of silence (when not speaking). If an unintended facial expression appears during a period of silence, various problems can occur.
[0008] In view of the above-mentioned problems, one example of the objective of the present disclosure is to provide a processing device, processing method, and program that alleviate the inconvenience of unintended facial expressions appearing at silent times when lip synchronization is employed. [Means for solving the problem]
[0009] According to the present disclosure, Computer, functioning as a display control means for executing a first control in which the mouth of a drawing target executes a first motion that changes the mouth state corresponding to each sound in accordance with the output timing of each sound included in the spoken lines; The display control means A program is provided that adds a second motion to the mouth of the drawing target from a predetermined time before the start of silence in the dialogue while the mouth of the drawing target is executing the first motion.
[0010] Further, according to the present disclosure, a display control unit that executes a first control to make a mouth of a drawing target execute a first motion that changes the state of the mouth corresponding to each sound in accordance with the output timing of each sound included in the spoken lines; The display control unit A processing device is provided that adds a second motion to the mouth of the drawing target from a predetermined time before a silence in the dialogue starts while the mouth of the drawing target is executing the first motion.
[0011] Further, according to the present disclosure, One or more computers a display control step of executing a first control for causing the mouth of a drawing target to perform a first motion that changes the mouth state corresponding to each sound in accordance with the output timing of each sound included in the spoken lines; In the display control step, A processing method is provided in which a second motion is added to the mouth of the drawing target from a predetermined time before the start of silence in the dialogue while the mouth of the drawing target is executing the first motion. [Effects of the Invention]
[0012] According to one aspect of the present disclosure, a processing device, a processing method, and a program are provided that reduce the inconvenience of unintended facial expressions appearing at silent times when lip synchronization is employed. [Brief explanation of the drawings]
[0013] [Figure 1] FIG. 1 illustrates an example of a functional block diagram of a processing device according to the present disclosure. [Figure 2] FIG. 1 is a diagram illustrating an example of a hardware configuration of a processing device according to the present disclosure. [Figure 3] FIG. 10 is a diagram schematically illustrating an example of a screen displayed by a processing device according to the present disclosure. [Figure 4] FIG. 1 is a diagram schematically illustrating an example of information processed by a processing device according to the present disclosure. [Figure 5] 10 is a flowchart illustrating an example of a processing flow of a processing device according to the present disclosure. [Figure 6] FIG. 2 is a diagram for explaining processing of a processing device according to the present disclosure. [Figure 7] FIG. 10 is another diagram for explaining the processing of the processing device according to the present disclosure. [Figure 8] FIG. 10 is another diagram for explaining the processing of the processing device according to the present disclosure. [Figure 9] FIG. 10 is another diagram for explaining the processing of the processing device according to the present disclosure. [Figure 10] FIG. 10 is another diagram for explaining the processing of the processing device according to the present disclosure. [Figure 11] FIG. 10 is another diagram for explaining the processing of the processing device according to the present disclosure. [Figure 12] FIG. 10 is another diagram for explaining the processing of the processing device according to the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0014] Hereinafter, embodiments of the present disclosure will be described with reference to the drawings. In this disclosure, the drawings relate to one or more embodiments. In all drawings, similar components are designated by similar reference numerals, and descriptions thereof will be omitted as appropriate.
[0015] <<Details of the Problems of the Embodiments of the Present Disclosure>> When lip sync is used, unintended facial expressions may appear during periods of silence. An example of an unintended facial expression is a relatively stiff expression with the mouth closed. When such unintended facial expressions appear during periods of silence, various problems can occur.
[0016] For example, if there is a relatively long period of silence, an unintended facial expression will continue for a long time. This is undesirable from a production perspective. Such a relatively long period of silence may occur after a line has been spoken. For example, in a scene where randomly selected lines are to be spoken, if a line that is short compared to the length of the scene is selected, there will be excess time after the line has been spoken, resulting in a relatively long period of silence.
[0017] In addition, the inclusion of an unintended facial expression during a silence can cause the transition of facial expressions to appear unnatural. For example, when a smiling target is made to speak, the facial expression transitions as follows: (before speaking) "smile" → (while speaking) "expression generated by lip sync" → (after speaking) "smile." If, for example, a relatively stiff expression with the mouth tightly closed is added at the end of this "expression generated by lip sync" (after the line has been spoken), the facial expression will go from "smile" → "expression of a conversation (expression generated by lip sync)" → "relatively stiff expression with the mouth tightly closed (expression generated by lip sync)" → "smile," and the final transition from "relatively stiff expression with the mouth tightly closed (expression generated by lip sync)" to "smile" will appear unnatural.
[0018] Additionally, when a relatively long line is spoken, silence may occur between the lines. For example, when a line contains multiple sentences, such as "Good morning. It's cold today, isn't it?", silence may occur between the sentences. In this case, the facial expression transitions from "a conversational expression (expression generated by lip sync)" to "a relatively stiff expression with the mouth tightly closed (expression generated by lip sync)" to "a conversational expression (expression generated by lip sync)." In this way, if an unintended facial expression appears between utterances, the transition of facial expressions may appear unnatural.
[0019] The embodiments of the present disclosure can reduce the inconvenience of unintended facial expressions appearing at silent times when lip synchronization is employed, thereby alleviating at least one of the various problems described above.
[0020] <<First embodiment>> 1 is a functional block diagram showing an overview of a processing device 10. As shown in FIG.
[0021] The display control unit 11 causes the mouth of the drawing target to execute a first motion (lip sync motion) that changes the mouth state to correspond to each sound in accordance with the output timing of each sound included in the spoken lines.
[0022] Then, the display control unit 11 adds the second motion to the mouth of the drawing target from a predetermined time before the start of silence in the dialogue while the mouth of the drawing target is executing the first motion.
[0023] In this way, by adding the second motion to the mouth of the drawing target a predetermined time before the silence start timing, the second motion can be added to the mouth of the drawing target even at the silence start timing and while the target is silent (while no one is speaking).
[0024] According to such a processing device 10, the inconvenience of an unintended facial expression (for example, a relatively stiff expression with the mouth tightly closed) appearing at the timing of silence can be reduced by adding the intended second motion.
[0025] Incidentally, when adding a second motion to a mouth executing a first motion, as in the processing device 10 of the first embodiment, adding the second motion at a large rate from the beginning can result in unnatural transitions in the mouth motion around the timing at which the addition of the second motion starts. For this reason, it is preferable to add the second motion at a small rate initially and gradually increase the rate of the second motion. However, when this method is adopted, the influence of the second motion is small when the addition of the second motion starts, and the influence of the second motion is hardly obtained. For this reason, if the addition of the second motion starts from the silence start timing, the influence of the second motion cannot be obtained for a predetermined time from the silence start timing, resulting in an unintended facial expression (for example, a relatively stiff expression with the mouth tightly closed).
[0026] By adding the second motion to the mouth of the drawing target from a predetermined time before the silence start timing as in the processing device 10 of the first embodiment, the proportion of the second motion at the silence start timing can be increased to a level where the influence of the second motion can be sufficiently obtained. As a result, it is possible to effectively suppress the inconvenience of unintended facial expressions appearing at the silence timing.
[0027] <<Second embodiment>> <Summary> The processing apparatus 10 of the second embodiment is a specific embodiment of the configuration of the processing apparatus 10 of the first embodiment, which will be described in detail below.
[0028] <Hardware configuration> An example of the hardware configuration of the processing device 10 will be described below. Each functional unit of the processing device 10 is realized by any combination of hardware and software. Those skilled in the art will understand that there are many variations in the realization method and device. The software includes programs that are pre-loaded in the device before shipping, and programs downloaded from recording media such as CDs (Compact Discs) or servers on the Internet.
[0029] FIG. 2 is a block diagram illustrating an example of the hardware configuration of a processing device 10. As shown in FIG. 2, the processing device 10 has a processor 1A, a memory 2A, an input / output interface 3A, a peripheral circuit 4A, and a bus 5A. The peripheral circuit 4A includes various modules. The processing device 10 does not necessarily have to have the peripheral circuit 4A. Note that the processing device 10 may be composed of multiple devices that are physically and / or logically separated. In this case, each of the multiple devices may have the above hardware configuration.
[0030] The bus 5A is a data transmission path for the processor 1A, memory 2A, peripheral circuit 4A, and input / output interface 3A to mutually transmit and receive data. The processor 1A is, for example, a central processing unit (CPU) or a graphics processing unit (GPU). The memory 2A is, for example, a random access memory (RAM) or a read-only memory (ROM). The input / output interface 3A includes interfaces for acquiring information from input devices, external devices, external servers, external sensors, cameras, etc., and interfaces for outputting information to output devices, external devices, external servers, etc. The input / output interface 3A also includes an interface for connecting to a communication network such as the Internet. Examples of input devices include a keyboard, mouse, microphone, physical buttons, touch panel, and dedicated game console controller. Examples of output devices include a display, projection device, speaker, printer, and mailer. The processor 1A can issue commands to each module and perform calculations based on the results of those calculations.
[0031] <Functional configuration> An example of the functional configuration of the processing device 10 will be described. An example of a functional block diagram of the processing device 10 is shown in FIG.
[0032] The display control unit 11 causes the output device to output a screen including the drawing target C as shown in Fig. 3. Then, the display control unit 11 causes the mouth m of the drawing target C to perform the first motion and the second motion described in the first embodiment.
[0033] In the example screen of Figure 3, the drawing object C is shown with a close-up of the face, but the drawing object may be included so that the entire body is visible, or so that only the upper body is visible (i.e., the lower body is not visible), or in other ways.
[0034] The "drawing subject" is a concept that includes at least one of a character and an avatar.
[0035] A "screen" is a screen that includes a drawing target and provides various information to the user.
[0036] An example of a screen is a game screen showing a scene from a game. The game screen can be realized using any well-known technology.
[0037] Another example of a screen is a UI screen that provides various kinds of guidance to the user. Examples of such a UI screen include a UI screen for a chat service. Other examples include a UI screen that outputs various kinds of guidance prepared in advance in response to a user's selection. In one example, on such a UI screen, the drawing target is displayed as a concierge or the like. Such a UI screen may be provided by an application (software), a web page, a dedicated system, or other means.
[0038] The screens shown here are merely examples, and the present invention is not limited to these examples.
[0039] The display control unit 11 may generate a screen and cause the generated screen to be output to an output device. Alternatively, the display control unit 11 may receive a screen from an external device and cause the received screen to be output to an output device.
[0040] In one example, the processing device 10 is an operation terminal operated by a user. In this example, the operation terminal (processing device 10) includes a display control unit 11. The display control unit 11 generates a screen and causes the generated screen to be output to an output device included in the operation terminal (processing device 10) or an output device connected to the operation terminal (processing device 10).
[0041] The "terminal device" may be, but is not limited to, a game device, a smartphone, a tablet terminal, a personal computer, a mobile phone, a terminal dedicated to a service installed in a facility, etc.
[0042] In another example, the processing device 10 is a client terminal operated by a user. In this example, the client terminal (processing device 10) includes a display control unit 11. The display control unit 11 receives a screen from the server and outputs the received screen to an output device included in the client terminal (processing device 10) or an output device connected to the client terminal (processing device 10).
[0043] A "client terminal" is, but is not limited to, a game device, a smartphone, a tablet terminal, a personal computer, a mobile phone, a terminal dedicated to a service installed in a facility, etc. The client terminal operates in conjunction with a server.
[0044] In another example, the processing device 10 is a server. In this example, the server (processing device 10) includes a display control unit 11. The display control unit 11 generates a screen and causes the generated screen to be output to a client terminal (output device).
[0045] Below, the configuration of the processing device 10 will be explained by dividing it into "pre-processing" that is performed before a screen such as that shown in Figure 3 is output to the output device, and "runtime processing" that outputs a screen such as that shown in Figure 3 to the output device and causes the mouth m of the drawing object C to perform a first motion or a second motion.
[0046] "Pre-processing" Pre-processing is processing that is executed before an output device outputs a screen including a rendering target C as shown in Figure 3. Pre-processing is processing that is performed during the creation stage of a game or system.
[0047] The following describes the pre-processing required to control the motion of the mouth of the drawing target. In addition to the pre-processing described below, preparations for controlling the design of the drawing target and the motion of parts other than the mouth of the drawing target can also be made, but these can be realized using any well-known technology.
[0048] Note that the following explanation of the pre-processing assumes that the drawing target will speak in Japanese, but if the drawing target speaks in another language, the same pre-processing will be performed after making appropriate adjustments for that language.
[0049] Pre-processing 1 (preparation of audio data and analysis data) First, at least one line to be spoken by the drawing target is determined in advance, and audio data (audio file) for that line is generated. One piece of audio data is generated corresponding to one line. One line may consist of one sentence, or may consist of multiple sentences. The work of generating the audio data is performed by a worker. The worker determines the line to be spoken by the drawing target based on the content of the game and the content of the guidance provided on a specified UI screen. The worker then generates audio data uttering the determined line by any means.
[0050] The generated voice data is stored in a predetermined storage device. The storage device may be provided within the processing device 10, or may be provided in an external device configured to be able to communicate with the processing device 10. The content, length, and language of the dialogue are not particularly limited. Dialogue and voice data may be prepared for each drawing object.
[0051] Then, by analyzing each piece of voice data, analysis data is generated for each piece of voice data (each line). This operation may be performed by the processing device 10. That is, the processing device 10 may be provided with a means for receiving input of voice data, analyzing the voice data, and generating analysis data. Alternatively, this operation may be performed by a worker. That is, the worker may analyze the voice data through processing using a predetermined voice analysis tool, and generate analysis data.
[0052] The generated analysis data is stored in a predetermined storage device. The storage device may be provided within the processing device 10 or may be provided in an external device configured to be able to communicate with the processing device 10.
[0053] An example of the analysis data is shown in Fig. 4. Fig. 4 shows the analysis data generated by analyzing the voice data with voice data ID (identifier): L00013.
[0054] The illustrated analysis data indicates, for each frame (each elapsed time from the start of the audio data), the content of the sound at each timing, the state of the mouth corresponding to that sound, and the volume of the sound. Note that the analysis data does not necessarily have to include some of the items shown in FIG.
[0055] "Frame" indicates the elapsed time from the start position of the audio data in frame count. "msec" indicates the elapsed time from the start position of the audio data in milliseconds. "Width" indicates the width of the mouth. In the example shown, "width" takes a value between 0 and 1. "height" indicates the height of the mouth. In the example shown, "height" takes a value between 0 and 1. "Tongue" indicates the position of the tongue. In the example shown, "tongue" takes a value between 0 and 1. "A", "I", "U", "E", and "O" indicate the blending amount of each of the sounds "a", "i", "u", "e", and "o". In the example shown, "A", "I", "U", "E", and "O" each take a value between 0 and 1. In the example shown, the total value of "A", "I", "U", "E", and "O" falls within the range of 0 to 1. "Vol (db)" is the decibel value of the volume. In the illustrated example, "Vol (db)" has a value between 0 (maximum volume) and -96 (minimum volume).
[0056] Such analysis of voice data can be realized using widely known analysis techniques. For example, such analysis of voice data can be realized using commercially available voice analysis tools. Note that, depending on the voice analysis tool used, the analysis data may be output in text format. In this case, the analysis data output in text format can be converted into binary data for use in the actual device and stored in the above-mentioned predetermined storage device.
[0057] Pre-processing 2 (preparing data for the first motion (lip sync motion)) Data is generated that indicates the state (shape, etc.) of the mouth corresponding to each sound "a, i, u, e, o," i.e., data that indicates the mouth motion when each sound is uttered (for example, data depicting the mouth motion). Also, data is generated that indicates the state (shape, etc.) of the mouth (base mouth) when none of the sounds is being uttered, i.e., data that indicates the mouth motion when none of the sounds is being uttered. The state of the mouth when none of the sounds is being uttered is, for example, a closed mouth.
[0058] Data for the first motion may be generated for each drawing target. Data for the first motion may also be generated for each situation, such as "when smiling" or "when angry." This work is performed by a worker. The worker can generate the data for the first motion using predetermined software.
[0059] The generated data for the first motion is stored in a predetermined storage device. The storage device may be provided within the processing device 10 or may be provided in an external device configured to be able to communicate with the processing device 10.
[0060] Pre-processing 3 (preparing data for the second motion) As explained in the first embodiment, the second motion is a motion added to the mouth of a drawing target that is executing the first motion, more specifically, a motion added to the mouth of a drawing target while the target is silent (while the target is not speaking). The data for the second motion is data that indicates such a motion (for example, data that depicts the mouth motion).
[0061] The purpose of adding the second motion is to alleviate the inconvenience of the drawn subject's mouth being in an unintended state (for example, a tightly closed state) during periods of silence (when not speaking), causing the drawn subject's face to have an unintended expression (for example, a relatively stiff expression). The second motion may be anything that can achieve this purpose, and various contents may be adopted. For example, the second motion may be a motion in which the mouth is slightly open and maintained in a relaxed state, or may be some other motion.
[0062] The data for the second motion may be generated for each line, so that the second motion appropriate for each line can be added to the mouth of the drawing target.
[0063] Alternatively, the data for the second motion may be generated for each drawing subject, allowing a second motion suited to the personality and characteristics of each drawing subject to be added to the mouth of each drawing subject.
[0064] Alternatively, the data for the second motion may be generated for each game situation (game progress, the scene at that time, the situation of the drawing target (hit points, etc.), etc.). In this way, a second motion that matches the game situation can be added to the mouth of the drawing target.
[0065] Runtime Processing The display control unit 11 executes the runtime processing when providing a game or various services (various kinds of guidance) to a user via a UI screen. The display control unit 11 causes an output device to output a screen including a drawing target C as shown in Fig. 3. Then, the display control unit 11 executes the processing shown in the flowchart of Fig. 5.
[0066] ○ Waiting for dialogue playback (S10) As shown in FIG. 5, the display control unit 11 is waiting for the lines to be reproduced.
[0067] The processing device 10 can determine whether to play back dialogue and which dialogue to play back based on, for example, user input, the current situation (game situation, UI screen display content, etc.), etc. The determination can be realized using any well-known technology.
[0068] ○Reading analysis data (S11) When it is determined that a certain line is to be played back (Yes in S10), the display control unit 11 reads out the analysis data of that line.
[0069] By the above-described pre-processing, analysis data for at least one line is generated in advance and stored in a predetermined storage device. The display control unit 11 reads out the analysis data corresponding to the line that has been determined to be played back from the storage device.
[0070] ○ Execution of the first motion (S12) After reading out the analysis data determined to be played back, the display control unit 11 executes a first control based on the read out analysis data. The display control unit 11 reads out data for a first motion (lip sync motion) that has been generated in advance by the above-described pre-processing and stored in a predetermined storage device. Then, the display control unit 11 causes the mouth of the drawing target to execute the first motion that changes the mouth state to correspond to each sound in accordance with the output timing of each sound included in the spoken lines.
[0071] The analysis data for each line indicates the content of the sound at each timing for each frame (each elapsed time from the start position of the audio data). For example, as shown in FIG. 4, the analysis data for each line indicates the blending amount of each of the sounds "a," "i," "u," "e," and "o." The display control unit 11 adds data indicating the state (shape, etc.) of the mouth corresponding to each sound (data indicating mouth motion) to the base mouth at a ratio according to the blending amount of each sound, and plays the result.
[0072] ○ Identifying the start of silence (S13) Based on the analysis data read out in S11, the display control unit 11 identifies the start timing of silence while the mouth of the drawing target is executing the first motion. The display control unit 11 identifies the start timing of silence in the future based on the elapsed time from the current time. The elapsed time from the current time is expressed in units of frame count, milliseconds, etc.
[0073] The display control unit 11 can identify the silence start timing by, for example, executing the following process 1 or process 2. Note that the display control unit 11 may also identify the silence start timing by other processes.
[0074] Process 1 First, the display control unit 11 identifies the silence start timing in the audio data based on the analysis data as shown in Fig. 4. As shown in Fig. 6, the silence start timing in the audio data is indicated by the elapsed time t1 from the start position of the audio data. The elapsed time t1 is indicated in units such as frame counts or milliseconds.
[0075] The silence start timing is, for example, the timing when "Vol(db)" shown in Fig. 4 becomes equal to or less than a threshold value. The threshold value here may be the minimum value (-96) among the possible values of "Vol(db)" ranging from 0 (maximum volume) to -96 (minimum volume), or may be a value obtained by adding a small correction value (a predetermined value) to this value.
[0076] Then, as shown in FIG. 7, the display control unit 11 specifies the timing t1 after the timing at which the reproduction of the audio data starts (the dialogue reproduction start timing in the figure) as the silence start timing.
[0077] Thereafter, the display control unit 11 manages the elapsed time t2 from the timing at which the reproduction of the audio data starts, and can thereby determine that the silence start timing will come after (t1-t2) from the current time.
[0078] Process 2 First, the display control unit 11 tracks the current playback position of the audio data while the audio data is being played. The current playback position of the audio data is indicated, for example, by the elapsed time from the start position of the audio data. The elapsed time is indicated in units such as frame counts or milliseconds.
[0079] Then, based on the current playback position and the analysis data as shown in Fig. 4, the display control unit 11 determines whether the silence start timing occurs a predetermined time t3 after the current playback position (after a predetermined number of frames, a predetermined number of milliseconds, etc.). The display control unit 11 can make this determination for each frame, for example. Note that the tracking and determination may be performed not based on the "current playback position" but on the "predetermined time after the current playback position (e.g., several seconds)."
[0080] The silence start timing is, for example, the timing when "Vol(db)" shown in Fig. 4 becomes equal to or less than a threshold value. The threshold value here may be the minimum value (-96) among the possible values of "Vol(db)" ranging from 0 (maximum volume) to -96 (minimum volume), or may be a value obtained by adding a small correction value (a predetermined value) to this value.
[0081] After determining that the silence start timing will occur a predetermined time t3 after the current playback position (after a predetermined number of frames, a predetermined number of milliseconds, etc.), the display control unit 11 manages the elapsed time t4 from the time when the determination was made, thereby determining that the silence start timing will occur (t3-t4) after the current time.
[0082] The "predetermined time t3" is a value smaller than the length t0 of the audio data to be played (see FIG. 6). By doing so, it is possible to determine, after starting playback of the audio data, whether the silence start timing will come after the predetermined time t3 (after a predetermined number of frames, a predetermined number of milliseconds, etc.) from the current playback position.
[0083] The predetermined time t3 is, for example, a value equal to or greater than the first predetermined time, and can be a value of about several frames (or about several milliseconds). As explained in the first embodiment, the first predetermined time is a value used to determine the timing for adding the second motion. Details of the first predetermined time will be described later.
[0084] In Process 1, if the timing at which audio data playback starts can be determined, the silence start timing can be determined. Therefore, the silence start timing can be determined relatively early, immediately after audio data playback starts. However, if a problem occurs, such as audio data playback stopping midway, the silence start timing determined based on the timing at which audio data playback started may differ from the actual silence start timing. Process 2, which determines whether the silence start timing occurs a predetermined time t3 (after a predetermined number of frames, a predetermined number of milliseconds, etc.) from the current playback position, can alleviate this problem.
[0085] 5, S13 is executed after S12, but the execution timing of S13 is not limited to this. S13 can be executed at any timing after the analysis data is read in S11. For example, S13 may be executed immediately after the analysis data of the lines to be played back is read in S11.
[0086] Addition of the second motion (S14, S15) The display control unit 11 reads out the data for the second motion that has been generated in advance by the above-mentioned pre-processing and stored in a predetermined storage device. If the data for the second motion has been generated for each condition, such as for each line of dialogue, each drawing object, or each game situation, the display control unit 11 reads out the data for the second motion that matches the conditions at that time.
[0087] Then, after S13, the display control unit 11 adds a second motion to the mouth of the drawing target from a first predetermined time before the identified silence start timing.
[0088] The "process of adding the second motion" adds the second motion at a small rate at first, gradually increasing the rate over time, and making the rate of the second motion 100% at a predetermined timing (switching timing) after the silence start timing. As the second motion is added, the rate of the first motion decreases. As the rate of the second motion gradually increases, the rate of the first motion gradually decreases accordingly. When the rate of the second motion reaches 100%, the rate of the first motion becomes 0%.
[0089] The proportion of the second motion (initial proportion) at the silence start timing may be set to 0% or a small value equivalent thereto. Then, the display control unit 11 may gradually increase the proportion of the second motion from such a value over time. The display control unit 11 may gradually increase the proportion of the second motion linearly, according to another function, or by other means.
[0090] During a time period in which the components of the first motion and the second motion are included, the display control unit 11 blends the first motion and the second motion at the ratio at that time and plays them back.
[0091] This process is illustrated in Figure 7. Figure 7 shows the speech state of the drawing target ("voice" in the figure) and the mouth motion of the drawing target ("mouth motion" in the figure) in time synchronization.
[0092] As shown in FIG. 7, the display control unit 11 causes the mouth of the drawing target to perform a first motion (lip sync motion) from the timing when playback of audio data starts (the timing when playback of lines starts in the figure). Note that the timing when the first motion starts may be slightly earlier or later. Alternatively, the display control unit 11 may cause the mouth of the drawing target to perform the first motion (lip sync motion) from the timing when utterance starts. The timing when utterance starts can be identified based on analysis data such as that shown in FIG. 4. For example, the display control unit 11 can identify the timing when "Vol(db)" shown in FIG. 4 becomes equal to or greater than a threshold as the timing when utterance started.
[0093] Then, the display control unit 11 starts adding the second motion a first predetermined time before the silence start timing. The timing at which the addition of the second motion starts is shown as "addition start timing" in the figure. After this addition start timing, the display control unit 11 causes the mouth of the drawing target to perform a motion that is a blend of the first motion (lip sync motion) and the second motion. The display control unit 11 initially adds the second motion at a small rate, gradually increases the rate of the second motion over time, and can make the rate of the second motion 100% at a predetermined timing (switching timing) after the silence start timing. The timing at which the rate of the second motion is made 100% is shown as "switching timing" in the figure. After the switching timing, the display control unit 11 causes the mouth of the drawing target to perform the second motion. During the "blend" time period in the figure, the display control unit 11 blends the first motion and the second motion at the rate at that time and plays them back.
[0094] The "first predetermined time" is set so that the proportion of the second motion is increased to a degree that the influence of the second motion is sufficiently obtained at the silence start timing. Such a first predetermined time can be determined depending on the way of increasing the proportion of the second motion, the target value of the proportion of the second motion at the silence start timing, etc. In one example, the first predetermined time can be set to about 0.1 to 0.5 seconds.
[0095] The "switching timing" is a timing after the silence start timing. The silence start timing may be the switching timing, or a timing after the silence start timing as shown in FIG. 7 may be the switching timing. In one example, the switching timing may be about 0.03 to 0.07 seconds after the silence start timing. The switching timing may be before the timing at which the playback of the audio data ends (the dialogue playback end timing in the figure), or may be after the timing at which the playback of the audio data ends.
[0096] Up to this point, the processing of the display control unit 11 has been described using an example in which there is a silent portion at the end of the audio data, as shown in Fig. 6. However, as shown in Fig. 8, there may be silent portions in the middle of the audio data. In other words, there may be silent portions between utterance portions. For example, when multiple sentences are uttered in one piece of audio data, such as "Good morning. It's cold today, isn't it?", silent portions will occur between the sentences.
[0097] Even if there is a silent part in the middle of the audio data in this way, the display control unit 11 can identify the start timing of the silence using a process similar to that described above, and add a second motion to the mouth of the drawing target from a first predetermined time before the identified start timing of the silence.
[0098] "Action and effect" According to the processing apparatus 10 of the second embodiment, the same effects as those of the processing apparatus 10 of the first embodiment are achieved.
[0099] Furthermore, the "process of adding the second motion" executed by the processing device 10 of the second embodiment can be a process in which the second motion is added at a small rate at first, the rate of the second motion is gradually increased over time, and the rate of the second motion is made 100% at a predetermined timing (switching timing) after the silence start timing.
[0100] The processing device 10 of the second embodiment can prevent the mouth motion from becoming unnatural around the timing of starting addition due to suddenly adding a large proportion of the second motion. Also, by gradually increasing the proportion of the second motion, it is possible to naturally switch from the first motion to the second motion.
[0101] <<Third embodiment>> <Summary> The processing device 10 of the third embodiment is a specific implementation of the configuration of the processing device 10 of the first embodiment. The processing device 10 of the third embodiment adds the second motion using a method different from that of the processing device 10 of the second embodiment. This will be described in detail below.
[0102] <Functional configuration> An example of the functional configuration of the processing device 10 will be described. An example of a functional block diagram of the processing device 10 is shown in FIG.
[0103] As in the second embodiment, the display control unit 11 causes the output device to output a screen including the drawing target C as shown in Fig. 3. Then, the display control unit 11 causes the mouth m of the drawing target C to execute the first motion and the second motion described in the first embodiment.
[0104] The "drawing target," "screen," "method by which the display control unit 11 causes the output device to output the screen," and the like are as explained in the second embodiment.
[0105] As in the second embodiment, the configuration of the processing device 10 will be described below, divided into "pre-processing" and "runtime processing."
[0106] "Pre-processing" The pre-processing of the third embodiment differs from the pre-processing of the second embodiment in the content of pre-processing 3 (preparation of data for the second motion). The pre-processing of the third embodiment is similar to the pre-processing of the second embodiment except for pre-processing 3 (preparation of data for the second motion).
[0107] Pre-processing 3 (preparing data for the second motion) In the third embodiment, at least one sequence is generated by combining a plurality of mouth motions in a predetermined order. Data representing each mouth motion (for example, data depicting the mouth motion) is also generated. These operations are performed by a worker. The worker can generate this data using predetermined software.
[0108] The generated data is stored in a predetermined storage device. The storage device may be provided within the processing device 10 or may be provided in an external device configured to be able to communicate with the processing device 10.
[0109] As shown in Fig. 9, each sequence shows at least one mouth motion (motions A, C, and D in the figure), the order in which the mouth motions are performed (A → D → C → A in the figure), and the timing for switching between mouth motions. Note that in the example of Fig. 9, the duration of each motion is shown, but instead, conditions for switching between motions may be shown.
[0110] Examples of mouth motions include, but are not limited to, "motions that maintain a smile," "motions that maintain a slightly open mouth," "motions that maintain a pointed mouth," and "motions that maintain a normal mouth."
[0111] In the third embodiment, the mouth motion incorporated into the sequence as described above is used as the second motion.
[0112] Note that a sequence may be generated for each facial feature. Alternatively, a sequence may be generated for each part of the face (for example, the lower half of the face, the upper half of the face). Alternatively, a sequence may be generated for the entire face as a whole. Alternatively, a sequence may be generated for each drawing target.
[0113] Runtime Processing The display control unit 11 causes the output device to output a screen including the drawing target C as shown in Fig. 3. Then, the display control unit 11 executes a second control. That is, the display control unit 11 determines a sequence to be executed in accordance with a user input, the situation at that time (the situation in the game, the display content of the UI screen, etc.), etc. Then, the display control unit 11 causes the mouth of the drawing target to execute a motion according to the determined sequence.
[0114] The display control unit 11 may cause other facial features to perform predetermined motions in accordance with the determined sequence. The sequence to be executed may be determined using any well-known technology.
[0115] The display control unit 11 executes the process shown in the flowchart of FIG. 5 while causing the mouth of the drawing target to perform a predetermined motion in accordance with the determined sequence.
[0116] ○ Waiting for dialogue playback (S10) 5, the display control unit 11 is waiting for the lines to be reproduced. The details of the process are the same as those in the second embodiment.
[0117] ○Reading analysis data (S11) When it is determined that a certain line is to be played back (Yes in S10), the display control unit 11 reads out the analysis data of that line. The details of the process are the same as those in the second embodiment.
[0118] ○ Execution of the first motion (S12) In response to the drawing target uttering a predetermined line, the display control unit 11 switches from a second control in which the drawing target's mouth executes a mouth motion according to a sequence to a first control in which the drawing target's mouth executes a first motion. Details of the first control are the same as those in the second embodiment.
[0119] The display control unit 11 can perform the transition from the mouth motion according to the sequence to the first motion, for example, by the following process. This process can make the transition natural. The following example is merely an example, and the transition may be performed by other methods.
[0120] In this example, instead of suddenly switching from the sequenced mouth motion to the first motion, the mouth performs a motion that blends the sequenced mouth motion and the first motion for a predetermined period of time, and then switches from the sequenced mouth motion to the first motion. In other words, the transition of the mouth motion is "sequenced mouth motion" → "blend" → "first motion."
[0121] The display control unit 11 can initially add the first motion at a small rate, gradually increase the rate of the first motion over time, and make the rate of the first motion 100% at a predetermined timing. As the first motion is added, the rate of the mouth motion according to the sequence decreases. As the rate of the first motion is gradually increased, the rate of the mouth motion according to the sequence gradually decreases accordingly. Then, when the rate of the first motion reaches 100%, the rate of the mouth motion according to the sequence becomes 0%.
[0122] The initial proportion of the first motion may be set to 0% or a small value equivalent thereto. Then, the display control unit 11 may gradually increase the proportion of the first motion from such a value over time. The display control unit 11 may gradually increase the proportion of the first motion linearly, according to another function, or by other means.
[0123] During the blending time period, the display control unit 11 can initially use the mouth motion according to the sequence at that time as the base mouth of the first motion (lip sync motion). In this case, the display control unit 11 adds data (data indicating mouth motion) indicating the state (shape, etc.) of the mouth corresponding to each sound at a ratio according to the blend amount of each sound to the mouth motion (base mouth) according to the sequence at that time, and plays it back.
[0124] Then, when the ratio of the first motion reaches a predetermined threshold or more, the base mouth motion can be switched from the mouth motion according to the sequence at that time to the base mouth motion of the first motion (lip sync motion) generated in advance. After that, the display control unit 11 adds data (data indicating mouth motion) indicating the state (shape, etc.) of the mouth corresponding to each sound at a ratio according to the blend amount of each sound to the base mouth motion of the first motion (lip sync motion) generated in advance, and plays it.
[0125] This process is illustrated in Figure 10. Figure 10 shows the running sequence ("Sequence" in the figure), the speech state of the drawing target ("Audio" in the figure), and the mouth motion of the drawing target ("Mouth Motion" in the figure) in time synchronization.
[0126] As shown in Fig. 10, before playing back the audio data, the display control unit 11 causes the mouth of the drawing target to perform the mouth motion according to the sequence. As shown in Fig. 10, before playing back the audio data, the mouth of the drawing target is performing motion A.
[0127] Then, the display control unit 11 can add a first motion (lip sync motion) to the mouth of the drawing target from the timing when the drawing target starts speaking ("voice start timing (1)" in the figure). Note that the display control unit 11 may add the first motion to the mouth of the drawing target from a predetermined time before the timing when the drawing target starts speaking.
[0128] The display control unit 11 may set the initial proportion of the first motion to 0% or a similarly small value. Then, the display control unit 11 may gradually increase the proportion of the first motion from such a value over time, and set the proportion of the first motion to 100% at a predetermined timing. The display control unit 11 may gradually increase the proportion of the first motion linearly, according to another function, or by other means.
[0129] 10, the time period during which the proportion of the first motion is gradually increased and the mouth of the drawing target executes a motion that is a blend of the mouth motion according to the sequence and the first motion is indicated as "BL." The time period after the point in time when the proportion of the first motion reaches 100% is indicated as "first motion." During the "BL" time period in the figure, the display control unit 11 blends the mouth motion according to the sequence and the first motion at the proportion at that time and plays them.
[0130] In the above example, the "sequenced mouth motion" and the "first motion" are blended so that the total is 100% during blending. However, as another example, the "sequenced mouth motion" and the "base mouth of the first motion" may be blended so that the total is 100%. In this example, the base mouth of the first motion during blending is a blend of the "base mouth of the first motion prepared in advance by pre-processing" and the "sequenced mouth motion." Then, the ratio of the "sequenced mouth motion" and the "base mouth of the first motion prepared in advance by pre-processing" is gradually changed over time.
[0131] The display control unit 11 initially reduces the proportion of the base mouth of the first motion prepared in advance in pre-processing, and gradually increases that proportion over time, so that at a predetermined timing the proportion of the base mouth of the first motion prepared in advance in pre-processing becomes 100%. As the proportion of the base mouth of the first motion prepared in advance in pre-processing gradually increases, the proportion of the mouth motion according to the sequence gradually decreases accordingly. Then, when the proportion of the base mouth of the first motion prepared in advance in pre-processing becomes 100%, the proportion of the mouth motion according to the sequence becomes 0%. This is the timing when blending ends.
[0132] The initial proportion of the base mouth of the first motion prepared in advance in pre-processing may be set to 0% or a small value equivalent thereto. Then, the display control unit 11 may gradually increase the proportion of the base mouth of the first motion prepared in advance in pre-processing from such a value over time. The display control unit 11 may gradually increase the proportion of the base mouth of the first motion prepared in advance in pre-processing linearly, gradually increase it according to another function, or gradually increase it by other means.
[0133] ○ Identifying the start of silence (S13) The display control unit 11 identifies the silence start timing while the mouth of the drawing target is executing the first motion, based on the analysis data read out in S11. Details of the process are the same as those in the second embodiment.
[0134] Addition of the second motion (S14, S15) After S13, the display control unit 11 adds the mouth motion according to the currently executed sequence as a second motion to the mouth of the drawing target from a first predetermined time before the identified silence start timing.
[0135] The "process of adding the second motion" adds the second motion at a small rate at first, gradually increasing the rate over time, and making the rate of the second motion 100% at a predetermined timing (switching timing) after the silence start timing. As the second motion is added, the rate of the first motion decreases. As the rate of the second motion gradually increases, the rate of the first motion gradually decreases accordingly. When the rate of the second motion reaches 100%, the rate of the first motion becomes 0%.
[0136] The proportion of the second motion (initial proportion) at the silence start timing may be set to 0% or a small value equivalent thereto. Then, the display control unit 11 may gradually increase the proportion of the second motion from such a value over time. The display control unit 11 may gradually increase the proportion of the second motion linearly, according to another function, or by other means.
[0137] During a time period in which the components of the first motion and the second motion are included, the display control unit 11 blends the first motion and the second motion at the ratio at that time and plays them back.
[0138] This process is illustrated in Figure 10. Figure 10 shows the running sequence ("Sequence" in the figure), the speech state of the drawing target ("Audio" in the figure), and the mouth motion of the drawing target ("Mouth Motion" in the figure) in time synchronization.
[0139] In the example shown in FIG. 10, the display control unit 11 identifies silence start timing (1) and silence start timing (2).
[0140] The silence start timing (1) indicates the start timing of the silence between utterances. For example, when multiple sentences are uttered in one voice data, such as "Good morning. It's cold today, isn't it?", silence occurs between the sentences.
[0141] The silence start timing (2) indicates the start timing of the silence portion at the end of the voice data, that is, after all utterances have ended.
[0142] The display control unit 11 adds the mouth motion according to the currently running sequence as a second motion to the mouth of the drawing target from a first predetermined time before the silence start timing (1). The "mouth motion according to the currently running sequence" is the motion that is determined to be executed at that time, and in the illustrated example, is motion D.
[0143] The display control unit 11 may set the initial proportion of the second motion (motion D) to 0% or a similarly small value. Then, the display control unit 11 may gradually increase the proportion of the second motion (motion D) from such a value over time, and set the proportion of the second motion (motion D) to 100% at a predetermined timing. The display control unit 11 may gradually increase the proportion of the second motion (motion D) linearly, according to another function, or by other means.
[0144] 10, the time period in which the proportion of the second motion (motion D) is gradually increased and the mouth of the drawing target is caused to execute a motion that is a blend of the first motion and the second motion (motion D) is indicated as "BL." The time period from the point in time when the proportion of the second motion (motion D) reaches 100% is indicated as "motion D." At the point in time when the proportion of the second motion (motion D) reaches 100%, the display control unit 11 switches from the first control in which the mouth of the drawing target executes the first motion to the second control in which the mouth of the drawing target executes a mouth motion according to a sequence.
[0145] In Fig. 10, audio start timing (2) comes after that (see "Audio" in the figure). The display control unit 11 switches from the second control, which causes the mouth of the drawing target to perform the mouth motion according to the sequence, to the first control, which causes the mouth of the drawing target to perform the first motion, by processing similar to "Performing the first motion (S12)" described above. As a result, as shown in the figure, the mouth motion transitions from motion D → BL → first motion.
[0146] In FIG. 10, silence start timing (2) comes after that (see "Audio" in the figure). From a first predetermined time before silence start timing (2), display control unit 11 adds the mouth motion according to the currently running sequence to the mouth of the drawing target as a second motion. The "mouth motion according to the currently running sequence" is the motion that is determined to be executed at that time, which is motion C in the figure.
[0147] The display control unit 11 may set the initial proportion of the second motion (motion C) to 0% or a similarly small value. Then, the display control unit 11 may gradually increase the proportion of the second motion (motion C) from such a value over time, and set the proportion of the second motion (motion C) to 100% at a predetermined timing. The display control unit 11 may gradually increase the proportion of the second motion (motion C) linearly, according to another function, or by other means.
[0148] 10, the time period in which the proportion of the second motion (motion C) is gradually increased and the mouth of the drawing target is made to execute a motion that is a blend of the first motion and the second motion (motion C) is indicated as "BL." The time period from when the proportion of the second motion (motion C) reaches 100% is indicated as "motion C." When the proportion of the second motion (motion C) reaches 100%, the display control unit 11 switches from the first control in which the mouth of the drawing target executes the first motion to the second control in which the mouth of the drawing target executes a mouth motion according to a sequence.
[0149] In the time period "BL" in the figure, the display control unit 11 blends the first motion and the second motion at the ratio at that time and plays them back.
[0150] The "first predetermined time" is the same as in the second embodiment.
[0151] The "timing to switch from the first control to the second control," i.e., the "timing to make the proportion of the second motion 100%," is a timing after the silence start timing. The silence start timing may be the timing in question, or a timing after the silence start timing may be the timing in question.
[0152] In the above example, the "sequenced mouth motion" and the "first motion" are blended so that the total is 100% during blending. However, as another example, the "sequenced mouth motion (second motion)" and the "first motion base mouth" may be blended so that the total is 100%. In this example, the first motion base mouth during blending is a blend of the "first motion base mouth prepared in advance by pre-processing" and the "sequenced mouth motion." Then, the ratio between the "sequenced mouth motion" and the "first motion base mouth prepared in advance by pre-processing" is gradually changed over time.
[0153] Furthermore, when performing processing based on a sequence as in the third embodiment, the display control unit 11 may execute at least one of the following characteristic processes 3 and 4.
[0154] ○Process 3 If the next motion switching timing occurs in the running sequence immediately after the timing of switching from the first control (lip sync control) to the second control (sequence-based control), the mouth motion may switch frequently in a short period of time, which may result in an unnatural appearance. This will be explained using Figure 10.
[0155] As described above, the mouth motion in Figure 10 shows that starting from a first predetermined time before the silence start timing (2), mouth motion C according to the running sequence is added to the mouth of the drawing target as a second motion, and then at a predetermined timing thereafter (the boundary between "BL" and "motion C" in the figure), the first control is switched to the second control.
[0156] If the mouth motion switches according to the running sequence (from motion C to motion A in the example) immediately after switching from the first control to the second control, i.e., immediately after "motion C" in the mouth motion of the figure starts, the mouth motion may switch frequently in a short period of time, which may look unnatural.
[0157] Therefore, as shown in FIG. 11, if there is a next motion switching timing in the currently executing sequence within a second predetermined time from the timing of switching from the first control to the second control ("switching timing" in the figure), the display control unit 11 adds the motion after switching ("motion A" in the figure) to the mouth of the drawing target as the second motion. That is, in the time period BL in the figure, the display control unit 11 adds motion A as the second motion.
[0158] On the other hand, as shown in FIG. 12, if there is no next motion switching timing in the currently executing sequence within a second predetermined time from the timing of switching from the first control to the second control ("switching timing" in the figure), the display control unit 11 adds the motion before switching ("motion C" in the figure) to the mouth of the drawing target as the second motion. That is, in the time period BL in the figure, the display control unit 11 adds motion C as the second motion.
[0159] This reduces the inconvenience of frequent changes in mouth motion occurring in a short period of time, which can be unnatural.
[0160] Process 4 The display control unit 11 may cause the mouth of the drawing target to close firmly once for a predetermined time (for example, about 0.1 to 0.2 seconds) at the end of the line. That is, the display control unit 11 may cause the mouth of the drawing target to execute such a mouth-closing motion.
[0161] In one example, the display control unit 11 may execute control to close the mouth of the drawing target at the timing when the playback of the dialogue (playback of the audio data) ends (the timing when the dialogue playback ends in Figure 10), and cause the mouth of the drawing target to execute a closing motion.
[0162] Additionally, when switching from the first control to the second control during a silent portion, the display control unit 11 may execute control to close the mouth of the drawing target at a predetermined timing between the first control and the second control, causing the mouth of the drawing target to execute a closing motion. The predetermined timing between the first control and the second control is a predetermined timing between the timing at which addition of the second motion starts and the timing at which the proportion of the second motion becomes 100%.
[0163] This control can only be performed on the silent parts at the end of a line of dialogue. For example, when multiple sentences are spoken in one piece of voice data, such as "Good morning. It's cold today, isn't it?", there will be silent parts between the sentences. During the silent parts in the middle of this voice data (dialogue), it is not necessary to execute the mouth closing motion of the drawn object's mouth. In this way, by not closing the mouth during the relatively short silent parts between sentences, and closing the mouth at the silent part at the end of the sentence, it is possible to make the drawn object's mouth move naturally.
[0164] The other configurations of the processing apparatus 10 of the third embodiment are similar to those of the processing apparatus 10 of the first and second embodiments.
[0165] "Action and effect" According to the processing apparatus 10 of the third embodiment, the same effects as those of the processing apparatus 10 of the first and second embodiments are achieved.
[0166] Furthermore, the processing device 10 of the third embodiment uses the mouth motion incorporated in the sequence as the second motion, so that the processing device 10 of the third embodiment does not need to take the trouble of generating the second motion.
[0167] Furthermore, the processing device 10 of the third embodiment can execute at least one of the above-mentioned processes 3 and 4, which are effective when performing processing based on a sequence. By performing such processes, natural drawing can be realized.
[0168] Although the present disclosure has been described above with reference to the embodiments, the present disclosure is not limited to the above-described embodiments. Various modifications that can be understood by those skilled in the art can be made to the configuration and details of the present disclosure within the scope of the present disclosure. Furthermore, each embodiment can be combined with other embodiments as appropriate.
[0169] In addition, in the flowcharts used in the above description, multiple steps (processes) are described in order. However, the order of the steps performed in each embodiment is not limited to the order described. In each embodiment, the order of the steps shown in the drawings can be changed as long as it does not cause any problems in terms of the content.
[0170] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes. 1. Turn your computer functioning as a display control means for executing a first control in which the mouth of a drawing target executes a first motion that sets the mouth in a state corresponding to each sound in accordance with the output timing of each sound included in the spoken lines; The display control means A program that adds a second motion to the mouth of the drawing target from a predetermined time before the start of silence in the dialogue while the mouth of the drawing target is executing the first motion. 2. The display control means further executing a second control for making the mouth of the drawing target execute a motion according to a predetermined sequence in which a plurality of mouth motions are combined in a predetermined order; switching from the second control to the first control in response to the drawing target uttering a predetermined line; adding a mouth motion according to the predetermined sequence as the second motion to the mouth of the drawing target from a first predetermined time before the silence start timing of the predetermined line; 2. The program according to claim 1, wherein the control is switched from the first control to the second control at a predetermined timing after the silence start timing. 3. The display control means If a next motion switching timing exists in the predetermined sequence within a second predetermined time from the timing of switching from the first control to the second control, the motion after switching is added to the mouth of the drawing target as the second motion; A program described in 2, in which if there is no next motion switching timing in the specified sequence within the second specified time from the timing of switching from the first control to the second control, the motion before switching is added to the mouth of the drawing target as the second motion. 4. The display control means The program described in 2, which, when switching from the first control to the second control, executes control to keep the mouth of the drawing target closed between the first control and the second control. 5. The program described in 1, wherein the process of adding the second motion starts a first predetermined time before the silence start timing, gradually increases the proportion of the second motion, and makes the proportion of the second motion 100% at a predetermined timing after the silence start timing. 6. A display control unit that executes a first control to make the mouth of a drawing target execute a first motion that changes the mouth state corresponding to each sound in accordance with the output timing of each sound included in the spoken lines, The display control unit A processing device that adds a second motion to the mouth of the drawing target from a predetermined time before a silence starts during the execution of the first motion by the mouth of the drawing target. 7. One or more computers: a display control step of executing a first control for causing the mouth of a drawing target to perform a first motion that changes the mouth state corresponding to each sound in accordance with the output timing of each sound included in the spoken lines; In the display control step, A processing method for adding a second motion to the mouth of the drawing target from a predetermined time before the start of silence in the dialogue while the mouth of the drawing target is executing the first motion. [Explanation of symbols]
[0171] 10 Processing equipment 11 Display control unit 1A processor 2A Memory 3A input / output I / F 4A peripheral circuit 5A Bus
Claims
1. Computer, functioning as a display control means for executing a first control in which the mouth of a drawing target executes a first motion that sets the mouth in a state corresponding to each sound in accordance with the output timing of each sound included in the spoken lines; The display control means A program that adds a second motion to the mouth of the drawing target from a predetermined time before a silence in the dialogue begins while the mouth of the drawing target is executing the first motion.
2. The display control means a second control for causing the mouth of the drawing target to perform a motion according to a predetermined sequence in which a plurality of mouth motions are combined in a predetermined order; switching from the second control to the first control in response to the drawing target uttering a predetermined line; adding a mouth motion according to the predetermined sequence as the second motion to the mouth of the drawing target from a first predetermined time before the silence start timing of the predetermined line; 2. The program according to claim 1, wherein the control is switched from the first control to the second control at a predetermined timing after the silence start timing.
3. The display control means If a next motion switching timing exists in the predetermined sequence within a second predetermined time from the timing of switching from the first control to the second control, the motion after switching is added to the mouth of the drawing target as the second motion; The program described in claim 2, wherein if there is no next motion switching timing in the specified sequence within the second specified time from the timing of switching from the first control to the second control, the motion before switching is added to the mouth of the drawing target as the second motion.
4. The display control means The program according to claim 2 , wherein when switching from the first control to the second control, a control is executed to keep the mouth of the drawing target in a closed state between the first control and the second control.
5. The program of claim 1, wherein the process of adding the second motion starts a first predetermined time before the silence start timing, gradually increases the proportion of the second motion, and makes the proportion of the second motion 100% at a predetermined timing after the silence start timing.
6. a display control unit that executes a first control to make a mouth of a drawing target execute a first motion that changes the mouth state corresponding to each sound in accordance with the output timing of each sound included in the spoken lines; The display control unit A processing device that adds a second motion to the mouth of the drawing target from a predetermined time before a silence in the dialogue begins while the mouth of the drawing target is executing the first motion.
7. One or more computers a display control step of executing a first control for causing the mouth of the drawing target to perform a first motion that changes the mouth state corresponding to each sound in accordance with the output timing of each sound included in the spoken lines; In the display control step, A processing method for adding a second motion to the mouth of the drawing target from a predetermined time before a silence in the dialogue begins while the mouth of the drawing target is executing the first motion.
Citation Information
Patent Citations
System and method for visually helping hearing and recording medium recording control program for visually helping hearing
JP1999226050A
Lip sink processing program, recording media and lip sink processing method
JP2019057116A