Audio-video providing method and computing device for performing the same
By synthesizing audio-visual state images during the regeneration of waiting state images and utilizing reverse motion imaging technology, the problem of long generation time and large data volume in the existing technology for generating real-time session images is solved, and efficient audio-visual image generation is achieved.
Patent Information
- Application Number
- CN202080038196.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-10-30
- Filing Date
- 2020-12-22
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2040-12-22
AI Technical Summary
Existing technologies are too time-consuming and require a large amount of data to generate real-time conversation-related audio-visual content, making it difficult to achieve efficient generation.
By pre-generating a waiting state image and synthesizing a sound state image during the regeneration process, the waiting state image is returned to the reference frame using a reverse motion image, thus generating a synthesized sound image.
It enables real-time provision of AI-based conversational services, reducing the time and data volume required to generate synthetic audio-visual content.
Smart Images

Figure CN114766036B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to a voice video providing technology. BACKGROUND
[0002] Recently, with the development of technology in the field of artificial intelligence, various types of content have been generated based on artificial intelligence (AI) technology. For example, when there is a voice message to be conveyed, there is a case of attracting people's attention by generating a voice video that seems to be a voice message spoken by a celebrity (e.g., a president, etc.). This is implemented in such a way that a mouth shape, etc. is generated to match a specific message, as if a specific message is spoken by a celebrity in a video of the celebrity.
[0003] Also, technology for having a conversation with a person using artificial intelligence (AI) (e.g., a video call, etc.) is being researched. However, such technology has a problem in that it is difficult to generate a real-time conversation video (or a voice video) because it takes too much time to synthesize a voice video and requires a lot of data. SUMMARY
[0004] TECHNICAL PROBLEM
[0005] An object of embodiments of the present application is to provide a new technology that can provide a voice video based on artificial intelligence in real time.
[0006] TECHNICAL SOLUTION
[0007] A computing device of one disclosed embodiment has at least one processor and a memory for storing at least one program executed by the at least one processor, and includes a waiting state video generation module for generating a waiting state video in which a character in a video is in a waiting state, a voice state video generation module for generating a voice state video in which a character in a video is in a voice state based on a voice content source, and a video reproduction module for reproducing the waiting state video and generating a synthesized voice video by synthesizing the waiting state video being reproduced and the voice state video.
[0008] The video reproduction module can generate the synthesized voice video by synthesizing a preset reference frame of the waiting state video being reproduced and the voice state video.
[0009] The present application is characterized in that the reference frame can be a first frame of the waiting state video.
[0010] The waiting state image generation module generates a waiting state image with a preset reproduction time, generates at least one reverse motion image for at least one frame contained in the waiting state image, and the image reproduction module returns the waiting state image being reproduced to the reference frame based on the reverse motion image, and the synthesized sound image is generated by synthesizing the returned reference frame and the sound state image.
[0011] The reverse motion image is used for image interpolation between a corresponding frame of the waiting state image and the reference frame of the waiting state image.
[0012] The reverse motion image is generated at each preset frame interval in a plurality of frames contained in the waiting state image, the image reproduction module detects a nearest frame with the reverse motion image in a plurality of frames after a current frame of the waiting state image being reproduced, returns the waiting state image to the reference frame based on the reverse motion image of the detected frame, and the synthesized sound image is generated by synthesizing the returned reference frame and the sound state image.
[0013] When the sound state image ends in the process of reproducing the synthesized sound image, the image reproduction module re-reproduces the waiting state image from an ending time point of the sound state image, and when the waiting state image ends, the image reproduction module returns the waiting state image to the reference frame based on the reverse motion image of a last frame of the waiting state image and reproduces.
[0014] The sound state image generation module generates a voice part and an image part of the sound state image based on the sound content source respectively, and the image part can be generated for a face part of a person in the waiting state image.
[0015] The image reproduction module replaces the face part of the waiting state image with the image part of the sound state image, and the synthesized sound image is generated by synthesizing the waiting state image with the replaced face part and the voice part of the sound state image.
[0016] A computing device of another disclosed embodiment has at least one processor and a memory for storing at least one program executed by the at least one processor, and the computing device includes a waiting state image generation module for generating a waiting state image in which a person in an image is in a waiting state, and an image reproduction module for transmitting a sound content source to a server, receiving a sound state image in which a person in an image is in a sound state from the server, reproducing the waiting state image, and generating a synthesized sound image by synthesizing the waiting state image being reproduced and the sound state image.
[0017] The voice video providing method of one disclosed embodiment is performed by a computing device having at least one processor and a memory for storing at least one program executed by the at least one processor, and includes the steps of generating a waiting state video in which a character in a video is in a waiting state, generating a voice state video in which the character in the video is in a voice state based on a voice content source, and reproducing the waiting state video and generating a synthesized voice video by synthesizing the waiting state video being reproduced and the voice state video.
[0018] The voice video providing method of another disclosed embodiment is performed by a computing device having at least one processor and a memory for storing at least one program executed by the at least one processor, and includes the steps of generating a waiting state video in which a character in a video is in a waiting state, transmitting a voice content source to a server, receiving a voice state video in which the character in the video is in a voice state from the server, and reproducing the waiting state video and generating a synthesized voice video by synthesizing the waiting state video being reproduced and the voice state video.
[0019] Technical Effects
[0020] According to the disclosed embodiments, a voice state video can be generated during reproduction of a waiting state video by preparing the waiting state video in advance, and a synthesized voice video can be generated in real time by synthesizing the voice state video and the waiting state video, thereby enabling a conversation-related service based on artificial intelligence to be provided in real time.
[0021] Also, when the voice state video is generated, a video portion is generated for a face portion of the character in the waiting state video, and the face portion of the waiting state video is replaced with the video portion of the voice state video to generate a synthesized voice video, thereby enabling the time and amount of data consumed to generate the synthesized voice video to be reduced.
[0022] Also, by providing a reverse motion image in a frame of the waiting state video and returning the waiting state video being reproduced through the reverse motion image to a first frame, and then synthesizing the first frame of the waiting state video and the voice state video, the synthesized voice video can be easily generated whenever the voice state video is generated during reproduction of the waiting state video, regardless of other variables. BRIEF DESCRIPTION OF DRAWINGS
[0023] Figure 1 A block diagram showing the structure of a voice video providing apparatus according to one embodiment of the present application.
[0024] Figure 2FIG. 1 is a diagram for briefly showing a state in which a frame of a waiting state video is generated into a reverse moving image in an embodiment of the present application.
[0025] Figure 3 FIG. 2 is a diagram for briefly showing a state in which a waiting state video and a sound state video are synthesized in an embodiment of the present application.
[0026] Figure 4 FIG. 3 is a diagram for briefly showing a state in which a video reproduction module returns a waiting state video being reproduced to a first frame in an embodiment of the present application.
[0027] Figure 5 FIG. 4 is a diagram for showing a configuration of a sound video providing system in an embodiment of the present application.
[0028] Figure 6 FIG. 5 is a block diagram for illustrating a computing environment including a computing device suitable for use in the illustrated embodiments. DETAILED DESCRIPTION
[0029] Hereinafter, specific embodiments of the present application will be described with reference to the accompanying drawings. The following detailed description is provided to give a full and complete understanding of the methods, devices, and / or systems described in this specification. However, this is merely illustrative, and the present application is not limited thereto.
[0030] In describing the embodiments of the present application, detailed descriptions of well-known related technologies will be omitted when it is determined that such detailed descriptions of the related technologies can unnecessarily obscure the gist of the present application. Also, the following terms are defined as terms defined after considering the functions in the present application, and can vary according to the intention or custom of the user, operator, etc. Therefore, the terms should be defined according to the entire content of the specification. The terms used in the detailed description are used only to explain embodiments of the present application, and thus, do not have limiting meanings. Unless clearly used in a different manner, the expression of the singular form includes the meaning of the plural form. In the specification, the expressions of "include" or "provide" or the like are used only to specify a certain feature, number, step, action, element, part, or combination thereof, and should be interpreted as not excluding the existence or possibility of one or at least one other feature, number, step, action, element, part, or combination thereof, in addition to the described content.
[0031] In the following description, "transmission", "communication", "sending", "receiving", and terms similar thereto with respect to a signal or information include not only a meaning of directly transmitting the signal or information from one structural element to another structural element, but also a meaning of transmitting through other structural elements. In particular, "transmission" or "sending" of a signal or information to a certain structural element means a final destination of the signal or information thereof, and does not mean a direct destination. This is the same as "receiving" of a signal or information. Also, in the present specification, "related to" two or more data or information means that at least a part of other data (or information) can be obtained based on one data (or information) in a case where the one data (or information) is obtained.
[0032] Also, although the first, second, and the like terms can be used to describe various structural elements, the structural elements are not limited to the terms. The terms can be used to distinguish one structural element from other structural elements. For example, the first structural element can be named as the second structural element without departing from the scope of the invention claimed by the present application, and similarly, the second structural element can be named as the first structural element.
[0033] Figure 1 A block diagram showing a structure of an apparatus for providing a sound image according to an embodiment of the present application.
[0034] Referring to Figure 1 The apparatus for providing a sound image 100 can include a waiting state image generation module 102, a sound state image generation module 104, and an image reproduction module 106.
[0035] According to an embodiment, the waiting state image generation module 102, the sound state image generation module 104, and the image reproduction module 106 can be implemented by at least one apparatus that is physically distinguished, or can be implemented by at least one processor or a combination of at least one processor and software, and can not be explicitly distinguished in a specific action, unlike the example shown in the drawing.
[0036] In the illustrated embodiment, although the apparatus for providing a sound image 100 can be a device for performing an artificial intelligence conversation (AI conversation) or a video call (AI video call), or the like, it is not limited thereto. The apparatus for providing a sound image 100 can generate a sound image (for example, a sound image for a conversation or a video call, or the like) based on artificial intelligence and display the generated sound image on a screen or transmit the generated sound image to the outside (for example, a terminal of a conversation object or a relay server for relaying a terminal of a conversation object and the apparatus for providing a sound image 100, or the like).
[0037] For example, the apparatus for providing a sound image 100 can also be provided in a terminal of a user who wants to have a conversation with artificial intelligence, in a variety of devices or equipment such as a self-service machine without a person, an electronic service desk, an outdoor advertising screen, a robot, or the like.
[0038] Among them, the voiced image as the image synthesized based on artificial intelligence refers to the image in which the specified person speaks. Among them, although the specified person can be a virtual person, or can also be a public figure, but is not limited thereto.
[0039] The waiting state image generation module 102 can generate an image in which the person in the image is in a waiting state (hereinafter referred to as a waiting state image). Among them, the waiting state can be a state before the person in the image speaks (for example, a state of listening to the other party speak, etc.).
[0040] The waiting state image generation module 102 can generate a waiting state image with a preset reproduction time (for example, 5 seconds to 30 seconds, etc.). The waiting state image can display natural actions in the process in which the person in the image is in a waiting state. That is, in the process of listening to the other party speak, the waiting state image can naturally display the expressions, postures and actions of the person in this process (for example, the action of nodding, the action of folding hands and listening, the action of tilting the head, the expression of smiling, etc.).
[0041] The waiting state image has a preset reproduction time and contains multiple frames. Moreover, in order to display natural actions during the period in which the person in the image is in a waiting state, each frame in the waiting state image can include varying images. In an exemplary embodiment, when the waiting state image is reproduced from the first frame to the last frame, it can be repeatedly reproduced from the last frame back to the first frame and then back again.
[0042] In addition to each frame of the waiting state image, the waiting state image generation module 102 can generate a back motion image. The back motion image can be used for image interpolation between any frame of the waiting state image and a preset reference frame of the waiting state image. Hereinafter, the reference frame will be described as the first frame of the waiting state image. However, the reference frame is not limited thereto.
[0043] When returning from any frame of the waiting state image to the first frame (i.e., the reference frame) of the waiting state image, in order to make the any frame and the first frame naturally connected, the waiting state image generation module 102 can generate a back motion image.
[0044] As Figure 2As illustrated, in the exemplary embodiment, the waiting state image generation module 102 can generate the reverse motion image from each frame (from the 2nd frame to the nth frame) in addition to the first frame (1st) of the waiting state image. That is, the waiting state image generation module 102 can generate the reverse motion image for the image interpolation between the corresponding frame and the first frame for each frame in addition to the first frame of the waiting state image. In this case, at least one reverse motion image can be generated for each frame. However, it is not limited thereto, and the reverse motion image can also be generated at every predetermined frame interval in the waiting state image.
[0045] The speaking state image generation module 104 can generate an image in which a person in the image is in a speaking state (hereinafter, referred to as a speaking state image). Herein, the speaking state refers to a state in which a person in the image speaks (for example, a state of speaking with an opposite party such as a conversation or a video call). The speaking state image generation module 104 can generate the speaking state image based on the input speaking content source. Although the speaking content source can be in a text form, it is not limited thereto, and can also be in a voice form.
[0046] Although the speaking content source can be generated by analyzing the speaking content of the opposite party by the speaking image providing apparatus 100 and by artificial intelligence, it is not limited thereto, and can also be input by an external apparatus (not illustrated) (for example, an apparatus that generates the speaking content source by analyzing the speaking content of the opposite party) or a manager. Hereinafter, the speaking content source will be exemplarily described as a text.
[0047] The speaking state image generation module 104 generates a voice portion and an image portion of the speaking state image based on the text of the speaking content (for example, "Hello, I am an artificial intelligence tutor, Danny"), respectively, and can generate the speaking state image by synthesizing the generated voice portion and the image portion. The technology of generating a voice and an image based on a text is a well-known technology, and thus a detailed description thereof will be omitted.
[0048] When the speaking state image generation module 104 generates the image portion based on the text of the speaking content, the image portion can be generated with respect to the face portion of the person in the waiting state image. As described above, by generating the image portion related to the face portion of the corresponding person in the speaking state image, the speaking state image can be further rapidly generated, and the data capacity can be reduced.
[0049] The image reproduction module 106 can reproduce the waiting state image generated by the waiting state image generation module 102. The image reproduction module 106 can reproduce the waiting state image and provide it to a conversation object. In the exemplary embodiment, the image reproduction module 106 can reproduce the waiting state image and display it through a screen provided in the speaking image providing apparatus 100. In this case, the conversation object can have a conversation with the person in the image while gazing at the screen of the speaking image providing apparatus 100.
[0050] Also, the image reproduction module 106 can transmit to an external device (e.g., a terminal of a session object or a relay server, etc.) by reproducing the waiting state image. In this case, the session object can receive the image through his / her own terminal (e.g., a smartphone, a tablet, a notebook, a desktop, etc.), a self-service machine, an electronic service desk, an outdoor advertisement screen, etc. and have a conversation with a person in the image.
[0051] When the voice state image is generated in the process of reproducing the waiting state image, the image reproduction module 106 can generate a synthesized voice image by synthesizing the waiting state image and the voice state image and reproduce the synthesized voice image. The image reproduction module 106 can provide the synthesized voice image to the session object.
[0052] Figure 3 FIG. 1 is a diagram for briefly showing a state of synthesizing a waiting state image and a voice state image according to an embodiment of the present application. Referring to FIG. 1, the image reproduction module 106 can generate a synthesized voice image by synthesizing a waiting state image and a voice state image. Figure 3 , the image reproduction module 106 can generate a synthesized voice image by replacing a face portion of the waiting state image with an image portion of the voice state image (i.e., a face portion of a corresponding person) and synthesizing a voice portion of the voice state image.
[0053] In an exemplary embodiment, when the generation of the voice state image is completed in the process of reproducing the waiting state image, the image reproduction module 106 can generate a synthesized voice image by returning to a first frame of the waiting state image and synthesizing a preset reference frame of the waiting state image and the voice state image. For example, the synthesis of the waiting state image and the voice state image can be performed in the first frame of the waiting state image.
[0054] In this case, in the process of reproducing the waiting state image, whenever the voice state image is generated, regardless of other variables (e.g., a network environment between the voice image providing device 100 and an opposite terminal, etc.), a synthesized voice image can be easily generated by uniformly synthesizing the waiting state image and the voice state image and synthesizing the waiting state image and the voice state image.
[0055] In this case, in order to synthesize the first frame of the waiting state image and the voice state image, after returning the waiting state image being reproduced to the first frame (i.e., the reference frame), the image reproduction module 106 can synthesize the first frame of the waiting state image and the voice state image.
[0056] Figure 4 FIG. 2 is a diagram for briefly showing a state in which the image reproduction module 106 of an embodiment of the present application returns the waiting state image being reproduced to the first frame. Referring to FIG. 2, the image reproduction module 106 can return the waiting state image being reproduced to the first frame by using a return button 201. Figure 4In the reproduction of the waiting state video, when a voice state video is generated at the j-th frame and is synthesized with the waiting state video, the video reproduction module 106 can detect the nearest frame having a reverse motion image in the frames after the j-th frame of the waiting state video being reproduced.
[0057] For example, when the nearest frame having a reverse motion image in the frames after the j-th frame is the k-th frame, the video reproduction module 106 can return the waiting state video to the first frame using the reverse motion image of the k-th frame. That is, the video reproduction module 106 can naturally return the waiting state video to the first frame by reproducing the reverse motion image of the k-th frame. The video reproduction module 106 can generate a synthesized voice video by synthesizing the first frame of the waiting state video and the voice state video.
[0058] If the voice state video ends in the reproduction of the synthesized voice video, the video reproduction module 106 can reproduce the waiting state video again from the end time point of the voice state video. When the waiting state video ends, the video reproduction module 106 can return to the first frame of the waiting state video again using the reverse motion image of the last frame of the waiting state video and reproduce the waiting state video.
[0059] According to the disclosed embodiment, the waiting state video is prepared in advance and the voice state video is generated in the reproduction of the waiting state video to be synthesized with the waiting state video, so that the synthesized voice video can be generated in real time, and thus, the artificial intelligence-based conversation-related service can be provided in real time.
[0060] Also, when the voice state video is generated, an image portion is generated for a face portion of a character in the waiting state video, the face portion of the waiting state video is replaced with the image portion of the voice state video to generate a synthesized voice video, so that the time and the amount of data required to generate the synthesized voice video can be reduced.
[0061] Also, after the reverse motion image is provided in the frames of the waiting state video and the waiting state video being reproduced is returned to the first frame through the reverse motion image, regardless of when the voice state video is generated in the reproduction of the waiting state video, the synthesized voice video can be easily generated by synthesizing the first frame of the waiting state video and the voice state video without considering other variables.
[0062] The module in the specification refers to a combination of functions and structures of hardware for implementing the technical idea of the present invention and software for driving the hardware. For example, the "module" refers to a logical unit of a specified code and a hardware resource for executing the specified code, and is not limited to physically connected codes or a type of hardware.
[0063] Figure 5 A diagram showing the structure of a voice video providing system according to an embodiment of the present invention.
[0064] Referring to Figure 5 The sound image providing system 200 can include a sound image providing apparatus 201, a server 203, and a target terminal 205. The sound image providing apparatus 201 can communicate with the server 203 and the target terminal 205 through a communication network 250.
[0065] In various embodiments, the communication network 250 can include the Internet, at least one local area network (LAN), a wide area network (WAN), a cellular network, a mobile network, other kinds of networks, or a combination of more than one of the above networks.
[0066] The sound image providing apparatus 201 can include a waiting state image generation module 211 and an image reproduction module 213. Among them, since the waiting state image generation module 211 is the same as the waiting state image generation module 102 shown in FIG. 1, a detailed description thereof will be omitted. Figure 1
[0067] When receiving the sound content source, the image reproduction module 213 can transmit the sound content source to the server 203. The server 203 can generate a sound state image based on the sound content source. That is, the server 203 can include a sound state image generation module 221. In an exemplary embodiment, the server 203 can generate a sound state image (i.e., a speech portion and an image portion) from the sound content source based on a machine learning technique. The server 203 can transmit the generated sound state image to the image reproduction module 213.
[0068] The image reproduction module 213 can provide to the target terminal 205 by reproducing the waiting state image. When receiving the sound state image of a preset time component from the server 203 in the process of reproducing the waiting state image, the image reproduction module 213 can generate a synthesized sound image by synthesizing the received sound state image and the waiting state image. The image reproduction module 213 can provide the synthesized sound image to the target terminal 205.
[0069] When the sound state image of the next time component is not received from the server 203, the image reproduction module 213 waits until the sound state image of the next time component is received from the server 203, and then can generate a synthesized sound image by synthesizing the received sound state image and the waiting state image.
[0070] Figure 6 A block diagram of a computing environment 10 including a computing apparatus suitable for use in the illustrated embodiments is shown for purposes of illustration in accordance with one embodiment. In the illustrated embodiment, various components can have different functions and capabilities than those described below and can include additional components.
[0071] The illustrated computing environment 10 includes a computing device 12. In one embodiment, the computing device 12 can be the audio-visual providing device 100, 201. The computing device 12 can be a server 203.
[0072] The computing device 12 includes at least one processor 14, a computer- readable storage medium 16, and a communication bus 18. The processor 14 can cause the computing device 12 to operate in accordance with the previously mentioned exemplary embodiments. For example, the processor 14 can execute at least one program stored in the computer-readable storage medium 16. The at least one program can include at least one computer-executable instruction that, when executed by the processor 14, can cause the computing device 12 to perform operations in the exemplary embodiments.
[0073] The computer-readable storage medium 16 can store computer-executable instructions and program codes, program data, and / or other forms of suitable information. The program 20 stored in the computer-readable storage medium 16 can include a set of processor-executable instructions. In one embodiment, the computer-readable storage medium 16 includes a memory (such as a volatile memory, a non-volatile memory, or a suitable combination thereof), at least one disk storage, an optical disk storage, a flash memory device, among others, or a suitable combination thereof, which is accessible by the computing device 12 and capable of storing information required by the computing device 12.
[0074] The communication bus 18 can include the processor 14, the computer- readable storage medium 16, and can connect a plurality of different components of the computing device 12 to each other.
[0075] The computing device 12 can include at least one input / output interface 22 providing an interface for at least one input / output device 24 and at least one network communication interface 26. The input / output interface 22 and the network communication interface 26 are connected to the communication bus 18. The input / output device 24 can be connected to other components of the computing device 12 through the input / output interface 22. The exemplary input / output device 24 can include an input device and an output device, for example, the input device includes a pointing device (a mouse or a touchpad, etc.), a keyboard, a touch input device (a touchpad or a touchscreen, etc.), a voice or sound input device, various types of sensing devices, and / or a camera device, and the output device can include a display device, a printer, a speaker, and / or a network card. The exemplary input / output device 24, as a part of the components constituting the computing device 12, can be disposed inside the computing device 12, or can be connected to the computing device 12 through an additional device that is different from the computing device 12.
[0076] Although the representative embodiments of the present application have been described in detail above, it should be understood that various modifications can be made to the above embodiments without departing from the scope of the present application. Therefore, the scope of the present application should not be limited to the above embodiments but should be defined by the scope of the appended claims and equivalents thereof.
Claims
1.A computing device having: at least one processor; and a memory storing at least one program for execution by the at least one processor, the computing device comprising: a waiting state video generation module configured to generate a waiting state video in which a character in a video is in a waiting state; a sound state video generation module configured to generate a sound state video in which the character in the video is in a sound state based on a sound content source; and a video reproduction module configured to reproduce the waiting state video, generate a synthesized sound video by synthesizing the waiting state video being reproduced and the sound state video, wherein the video reproduction module detects a nearest frame having a reverse motion image in a plurality of frames after a current frame of the waiting state video being reproduced, returns the waiting state video to a reference frame based on the reverse motion image of the detected frame, and generates the synthesized sound video by synthesizing the returned reference frame and the sound state video, the reverse motion image is used for image interpolation between a corresponding frame of the waiting state video and the reference frame, and is generated at every predetermined frame interval among a plurality of frames included in the waiting state video, and the reference frame is a first frame of the waiting state video. 3.The computing device of claim 1, wherein when the sound state video ends in the process of reproducing the synthesized sound video, the video reproduction module reproduces the waiting state video again from an ending time point of the sound state video, and when the waiting state video ends, the video reproduction module returns the waiting state video to the reference frame based on a reverse motion image of a last frame of the waiting state video and reproduces the waiting state video. 4.The computing device of claim 1, wherein the sound state video generation module generates a voice part and a video part of the sound state video based on the sound content source, respectively, and the video part is generated with respect to a face part of the character in the waiting state video, and the video reproduction module replaces the face part of the waiting state video with the video part of the sound state video, and generates the synthesized sound video by synthesizing the waiting state video in which the face part is replaced and the voice part of the sound state video. 6.A computing device having: at least one processor; and a memory storing at least one program for execution by the at least one processor, the computing device comprising: a waiting state video generation module configured to generate a waiting state video in which a character in a video is in a waiting state; and a video reproduction module configured to transmit a sound content source to a server, receive a sound state video in which the character in the video is in a sound state from the server, reproduce the waiting state video, and generate a synthesized sound video by synthesizing the waiting state video being reproduced and the sound state video, wherein the video reproduction module detects a nearest frame having a reverse motion image in a plurality of frames after a current frame of the waiting state video being reproduced, returns the waiting state video to a reference frame based on the reverse motion image of the detected frame, and generates the synthesized sound video by synthesizing the returned reference frame and the sound state video, the reverse motion image is used for image interpolation between a corresponding frame of the waiting state video and the reference frame, and is generated at every predetermined frame interval among a plurality of frames included in the waiting state video, and the reference frame is a first frame of the waiting state video. 2. The computing device of claim 1, wherein, 5. The computing device of claim 4, wherein, at least one processor; The image reproduction module detects a nearest frame having a reverse motion image among a plurality of frames after a current frame of the waiting state image being reproduced, returns the waiting state image to a reference frame based on the reverse motion image of the detected frame, and generates the synthesized voice image by synthesizing the returned reference frame and the voice state image. The reverse motion image is used for image interpolation between a corresponding frame of the waiting state image and the reference frame, and is generated at every preset frame interval among a plurality of frames included in the waiting state image. 7.A voice image providing method performed by a computing device having: at least one processor; and a memory storing at least one program executed by the at least one processor, the method comprising the steps of: generating a waiting state image in which a character in an image is in a waiting state; generating a voice state image in which the character in the image is in a voice state based on a voice content source; and reproducing the waiting state image and generating a synthesized voice image by synthesizing the waiting state image being reproduced and the voice state image; wherein, the step of generating the synthesized voice image comprises: a step in which the image reproduction module detects a nearest frame having a reverse motion image among a plurality of frames after a current frame of the waiting state image being reproduced, returns the waiting state image to a reference frame based on the reverse motion image of the detected frame, and generates the synthesized voice image by synthesizing the returned reference frame and the voice state image; the reverse motion image is used for image interpolation between a corresponding frame of the waiting state image and the reference frame, and is generated at every preset frame interval among a plurality of frames included in the waiting state image. 8.A voice image providing method performed by a computing device having: at least one processor; and a memory storing at least one program executed by the at least one processor, the method comprising the steps of: generating a waiting state image in which a character in an image is in a waiting state; transmitting a voice content source to a server; receiving a voice state image in which the character in the image is in a voice state from the server; and reproducing the waiting state image and generating a synthesized voice image by synthesizing the waiting state image being reproduced and the voice state image; wherein, the step of generating the synthesized voice image comprises: a step in which the image reproduction module detects a nearest frame having a reverse motion image among a plurality of frames after a current frame of the waiting state image being reproduced, returns the waiting state image to a reference frame based on the reverse motion image of the detected frame, and generates the synthesized voice image by synthesizing the returned reference frame and the voice state image; the reverse motion image is used for image interpolation between a corresponding frame of the waiting state image and the reference frame, and is generated at every preset frame interval among a plurality of frames included in the waiting state image.
Citation Information
Patent Citations
Moving image generating system and moving image display system
WO2017094527A1