Style image generation method, and device, medium and program product

By acquiring image frames and subtitles from video clips, and using a style image generation model to automatically generate style image sets, the problem of long comic production cycles is solved, enabling rapid generation and a personalized viewing experience.

WO2026036302A1PCT designated stage Publication Date: 2026-02-19DOUYIN VISION CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/112177
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-08-14
Publication Date
2026-02-19

AI Technical Summary

Technical Problem

Existing comic production cycles are long, making it difficult to meet users' needs for quick reading. Short dramas and videos are suitable for the comic style, but the production process is too complicated and costly, resulting in low comic output.

Method used

By acquiring image frames and subtitles from video clips, a trained style image generation model is used to generate style images, and subtitles are added to the corresponding positions to automatically generate a set of style images, such as comics.

Benefits of technology

It greatly shortens the generation time and manufacturing cycle of style image collections, providing users with a personalized and diverse viewing experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024112177_19022026_PF_FP_ABST
    Figure CN2024112177_19022026_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present disclosure are a style image generation method, and a device, a storage medium and a computer program product. The method comprises: acquiring a video clip, wherein the video clip comprises a set of image frames and a plurality of subtitles; extracting a plurality of key frames from the set of image frames, wherein each key frame has an associated subtitle; using a trained style image generation model to generate a style image corresponding to each key frame among the plurality of key frames; and correspondingly adding the subtitle associated with each key frame to a predetermined position in the corresponding style image, so as to generate a target style image. On the basis of the method in the embodiments of the present disclosure, artificial intelligence (AI) technology can be used to automatically generate a set of style images for a video clip of a type such as a short drama, thereby greatly shortening the generation time and production cycle of the set of style images, and providing a user with a more diverse viewing experience for content.
Need to check novelty before this filing date? Find Prior Art

Description

Style image generation method, device, medium, and program product TECHNICAL FIELD

[0001] The present disclosure relates generally to the field of computers, and more particularly to a style image generation method, an electronic device, a computer-readable storage medium, and a computer program product. BACKGROUND

[0002] With the continuous development of communication technology, the functions of mobile terminals are increasingly diversified. Users can experience more diversified functions on mobile terminals, such as voice or video calls, image acquisition, video playback, audio playback, browsing various types of files, and the like.

[0003] Among the various functions of mobile terminals, users can browse files in the form of comics and the like through mobile terminals. Through the services provided by mobile terminals, comics can be delivered to users more quickly, allowing users to read at any time and providing users with an experience beyond traditional reading.

[0004] SUMMARY

[0005] According to example embodiments of the present disclosure, a style image generation method, an electronic device, a computer storage medium, and a computer program product are provided.

[0006] In a first aspect of the present disclosure, a style image generation method is provided, including: obtaining a video segment, wherein the video segment includes a set of image frames and a plurality of subtitles; extracting a plurality of key frames from the set of image frames, wherein each key frame has an associated subtitle; generating a style image corresponding to each key frame in the plurality of key frames using a trained style image generation model; and adding the subtitle associated with each key frame to a corresponding style image at a predetermined position corresponding to the subtitle to generate a target style image.

[0007] In a second aspect of the present disclosure, an electronic device is provided, including: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, which when executed by the at least one processing unit causes the electronic device to perform the method described in the first aspect of the present disclosure.

[0008] In a third aspect of the present disclosure, a computer-readable storage medium is provided, having machine executable instructions stored thereon, which when executed by a device cause the device to perform the method described in the first aspect of the present disclosure.

[0009] In a fourth aspect of the present disclosure, a computer program product is provided, comprising computer executable instructions, wherein the computer executable instructions, when executed by a processor, implement the method described according to the first aspect of the present disclosure.

[0010] The summary is provided to introduce a selection of concepts that are further described in the detailed description below. It is not intended to identify key or essential features of the disclosure or to delineate the scope of the disclosure. Other features of the disclosure will be apparent from review of the disclosure below. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other features, aspects and advantages of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings, like reference numerals refer to like elements, wherein:

[0012] FIG. 1 shows a schematic diagram of an example system in which embodiments of the present disclosure can be implemented;

[0013] FIG. 2 shows a flowchart of a method for generating a style image according to embodiments of the present disclosure;

[0014] FIG. 3 shows a flowchart of an implementation method for extracting key frames according to some embodiments of the present disclosure;

[0015] FIG. 4 shows a flowchart of an implementation method for extracting key frames according to other embodiments of the present disclosure;

[0016] FIG. 5 shows a schematic diagram of key frames determined from a set of image frames corresponding to a video clip;

[0017] FIGS. 6A-6C show schematic diagrams of target style images with added subtitles according to embodiments of the present disclosure;

[0018] FIG. 7 shows an example diagram of a style image template according to embodiments of the present disclosure;

[0019] FIG. 8 shows a schematic block diagram of an example apparatus according to some embodiments of the present disclosure; and

[0020] FIG. 9 shows a block diagram of an example device that can be used to implement embodiments of the present disclosure. DETAILED DESCRIPTION

[0021] Embodiments of the present disclosure will be described below in greater detail with reference to the accompanying drawings. While certain embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein, but rather, the embodiments are provided so as to more completely and thoroughly understand the present disclosure. It is understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.

[0022] Mobile terminals provide users with a variety of rich functions. Among the various functions of the mobile terminal, the user can browse files including style images (such as comics, etc.) through the mobile terminal. Taking comics as an example, through the service provided by the mobile terminal, the comic files can be delivered to the user more quickly, allowing the user to read at any time and providing the user with a reading experience beyond the traditional reading experience. However, the comic files on the mobile terminal are usually completed by hand-drawing, which requires a considerable drawing time, and the story plot arrangement and content design of the comic files also require a considerable time. Therefore, the production cycle of the comic files currently provided to the mobile terminal is relatively long, which is difficult to meet the reading needs of the user.

[0023] With the continuous development of short videos, a large number of short drama videos have emerged. A short drama video usually refers to a video with a relatively short time length (for example, from tens of seconds to 15 minutes, etc.), with a certain plot, and two consecutive short drama videos are video clips with continuous plots. Compared with traditional long videos, the plot of a short drama is more compact. Users can quickly understand the plot by watching a short drama video, saving time for watching the video. Due to the compact plot, short dramas are very suitable for the style of comics. However, at the current stage, due to the excessive production links of comics and high cost, the annual output of the comic industry is extremely low. Although short dramas are suitable for the style of comics, in the case of the current bottleneck in the comic industry, it is difficult to process short dramas into comics for users to watch.

[0024] In view of this, embodiments of the present disclosure provide a method for generating style images. The method can include: obtaining a video clip, wherein the video clip includes a set of image frames and a plurality of subtitles; extracting a plurality of key frames from the set of image frames, wherein each key frame has an associated subtitle; generating a style image corresponding to each key frame in the plurality of key frames using a trained style image generation model; and adding the subtitle associated with each key frame to a predetermined position in the corresponding style image correspondingly to generate a target style image.

[0025] By employing the method according to the embodiments of the present disclosure, a collection of style images (such as a comic) can be automatically generated for a video clip of a type such as a short drama by using artificial intelligence (AI) technology, thereby greatly shortening the generation time and manufacturing cycle of the collection of style images (e.g., a comic file). In addition, the method according to the embodiments of the present disclosure further improves the user experience by providing the user with the convenience of inputting prompt information, so that the user can generate a more personalized collection of style images by using the capabilities of AI.

[0026] Embodiments of the present disclosure will be described in detail below with further reference to the accompanying drawings, in which FIG. 1 shows a schematic diagram of an example environment 100 in which embodiments of the present disclosure can be implemented. The example environment 100 includes a computing device 110 and a terminal device 120. The computing device 110 can be deployed with an intelligent model 112 (e.g., a generative intelligent model) that is trained to convert an input image into a style image, such as, but not limited to, a style image in a style of a comic (e.g., a Japanese comic, an American comic, etc.), a cartoon style, an ink painting style, and the like. The terminal device 120 is also shown in FIG. 1. In some embodiments, the terminal device 120 communicates with the computing device 110 through a network 130. The network 130 can include a wired network, a wireless network, or a combination thereof, for providing communication between the terminal device 120 and the computing device 110. In some embodiments, the terminal device 120 can be connected to the computing device 110 through a data line, and the present disclosure does not limit the connection manner between the computing device 110 and the terminal device 120.

[0027] Any of the computing device 110 and the terminal device 120 can include, but is not limited to, a personal computer, a server computer, a handheld or laptop device, a mobile device (such as a mobile phone, a personal digital assistant (PDA), a media player, etc.), a multi-processor system, a consumer electronic product, a wearable electronic device, a smart home device, a minicomputer, a mainframe computer, an edge computing device, a distributed computing system including any of the above systems or devices, and the like.

[0028] In some embodiments, the terminal device 120 can send a request to the computing device 110 through the network 130 for obtaining a collection of style images corresponding to a video clip. The video clip can be a video clip for which the user is interested in browsing the collection of style images (e.g., a comic file) corresponding to the video clip. The computing device 110 can respond to the request, obtain the collection of style images corresponding to the video clip, and send the obtained collection of style images to the terminal device 120 via the network 130, so that the user can browse the collection of style images corresponding to the video clip on the terminal device 120.

[0029] In some embodiments, the intelligent model 112 in the computing device 110 can generate a corresponding set of style images for the video clip. The intelligent model 112 in the computing device 110 can receive prompt information from the terminal device, the prompt information being used to provide constraint information for generating a style image in the set of style images. The intelligent model 112 can further generate the set of style images corresponding to the video clip based on the prompt information.

[0030] It can be understood that although the intelligent model 112 is deployed in the computing device 110 in FIG. 1, the intelligent model 112 can be split into multiple sub-models according to actual needs, and each sub-model can be deployed in a corresponding computing device, achieving distributed deployment such as to support large-scale models. Accordingly, the corresponding computing device can send the set of style images corresponding to the video clip to the terminal device 120 for display in response to a request from the terminal device 120.

[0031] In addition, although the intelligent model 112 is shown in FIG. 1 as being deployed separately from the terminal device 120, it can be understood that as the intelligent model 112 is lightened, the intelligent model 112 can also be deployed at the terminal device 120, thereby being able to respond more quickly to user requests and generate a corresponding set of style images. When the intelligent model 112 is deployed locally at the terminal device 120, the corresponding style image generation method is similar to the process described above in conjunction with FIG. 1, and those skilled in the art can refer to the description above for understanding, and for the sake of brevity, the description will not be repeated here.

[0032] According to the style image generation method of the embodiments of the present disclosure, a set of style images (such as files in the form of comics) for a video clip of a type such as a short drama can be automatically generated using artificial intelligence (AI) technology, thereby greatly shortening the generation time and manufacturing cycle of the set of style images (such as files in the form of comics), and providing users with a more diverse viewing experience of content. In addition, the method according to the embodiments of the present disclosure can provide users with the convenience of inputting prompt information, so that users can take full advantage of the capabilities of AI to generate a more personalized set of style images, thereby further improving the user experience.

[0033] The block diagram of an example environment 100 in which embodiments of the present disclosure can be implemented is described above in connection with FIG. 1. A method for generating style images according to embodiments of the present disclosure is described below in connection with FIG. 2. FIG. 2 illustrates a flowchart of a style image generation method 200 according to embodiments of the present disclosure. The method 200 can be executed at the computing device 110 in FIG. 1 and any suitable computing device. It should be understood that the numbering in the flowchart of the method 200 does not indicate an order of execution of the steps, some or all of the steps can be executed in parallel, or the order of execution can be interchanged, which is not limited in the present disclosure. In addition, the method 200 in FIG. 2 can also include additional steps not shown and / or can omit the steps shown, and the scope of the present disclosure is not limited in this regard.

[0034] In block 202, the computing device 110 can obtain a video clip, where the video clip includes a set of image frames and a plurality of subtitles. In some embodiments, the video clip can include a video that a user is interested in converting into a set of style images, including but not limited to a short drama video.

[0035] In some embodiments, the video clip can include a set of image frames and a plurality of subtitles. The subtitles can include textual information displayed in a plurality of image frame screens in the set of image frames. In some embodiments, for a short drama video, the subtitles can include a voiceover subtitle corresponding to a voiceover and a dialogue subtitle corresponding to a dialogue. The voiceover subtitle generally includes a commentary subtitle that explains the image frames. The dialogue subtitle includes dialogue information initiated by an object in the video, such as a character object or an animal object, etc. In addition, the subtitles can also include other types of prompt subtitles, for example, a text type of prompt subtitle that introduces or prompts the objects in the image frames, which is not limited in the present disclosure.

[0036] In addition, not all of the image frames in the set of image frames necessarily have subtitle information. For example, some image frames have subtitles, while other image frames do not have subtitles. This is not limited in the present disclosure.

[0037] In block 204, the computing device 110 can extract a plurality of key frames from the set of image frames, where each key frame has an associated subtitle. The key frames can be understood as image frames that are relatively key to generating the set of style images for the video clip. These key frames combined together can generate a set of style images that substantially and completely embody the content of the video clip. The extraction process of the key frames will be described in detail below in connection with the accompanying drawings.

[0038] In block 206, the computing device 110 can generate a style image corresponding to each of the plurality of keyframes using the trained style image generation model. The computing device 110 can input each of the plurality of keyframes to the trained style image generation model to generate, by the trained style image generation model, the corresponding style image from the plurality of keyframes, respectively, so that a set of style images can be obtained.

[0039] In block 208, the computing device 110 can add the subtitle associated with each keyframe to a predetermined position in the corresponding style image, respectively, to generate a target style image. In some embodiments, the predetermined position in the style image is associated with the type of the subtitle to be added. In response to the subtitle to be added including a dialogue subtitle corresponding to a dialogue, the predetermined position includes a position in the corresponding style image within a predetermined range of an object initiating the dialogue. In response to the object initiating the dialogue not being in the corresponding style image, the predetermined position includes a position in the corresponding style image opposite to a position of at least one object in the corresponding style image. Further, the computing device 110 can display an indication at the predetermined position, the indication indicating the object initiating the dialogue is in at least one style image out of the corresponding style image. In some embodiments, in response to the subtitle including a voiceover subtitle, the predetermined position for the voiceover subtitle can be adjusted according to a position of an object in the corresponding style image, and the predetermined position is at least partially non-overlapping with the position of the object. For example, the predetermined position for the voiceover can be at a position below or above the corresponding style image, and is at least partially non-overlapping with the position of the object in the image.

[0040] In some embodiments, the computing device 110 can add the subtitle associated with each keyframe to a predetermined position in the corresponding style image, respectively, to generate a target style image. The computing device 110 can obtain each target style image with the added subtitle, and embed each obtained target style image into a style image template to generate a set of style images.

[0041] In some embodiments, the style image template corresponds to a format of the set of style images to be generated. The format of the style image template can match a screen feature of a terminal device. For example, the style image template can be a strip-shaped template to match a screen feature of a mobile terminal. Taking the strip-shaped template as an example, the strip-shaped template can be divided into a plurality of regions in a direction from top to bottom, and each divided region corresponds to one or more style images. The generated set of style images, by being embedded into the style image template, can be displayed in the terminal device in the format of the template, to facilitate a user to browse at the terminal device.

[0042] By using the style image generation method according to the embodiments of the present disclosure, a set of style images (such as files in the form of cartoons) can be automatically generated for video clips of a type such as short plays by using artificial intelligence (AI) technology, so that the generation time and manufacturing cycle of the set of style images (such as files in the form of cartoons) can be greatly shortened, and a more diversified viewing experience of content can be provided for users.

[0043] The specific implementation of the computing device 110 extracting the key frames will be described in detail below with reference to the accompanying drawings. FIG. 3 shows a flowchart of an implementation method 300 of extracting key frames according to some embodiments of the present disclosure. It should be understood that the numbering in the flowchart of the method 300 does not represent the order of execution of these steps, some or all of these steps can be executed in parallel, or the order of execution can be interchanged, and the present disclosure does not limit this. In addition, the method 300 in FIG. 3 can also include additional steps not shown and / or can omit the steps shown, and the scope of the present disclosure is not limited in this regard.

[0044] In block 302, the computing device 110 can extract a plurality of image frames with subtitles from the set of image frames. The computing device 110 can extract the image frames with subtitles from the set of image frames, thereby obtaining a plurality of image frames with subtitles.

[0045] In block 304, the computing device 110 can arrange the extracted plurality of image frames in chronological order to obtain a plurality of arranged image frames.

[0046] In block 306, the computing device 110 can determine the similarity between two adjacent image frames. In some embodiments, the computing device 110 can compare the similarity between two adjacent image frames in the plurality of arranged image frames by using a trained neural network model. The computing device 110 can also compare the similarity between two adjacent image frames in other suitable manners, and the present disclosure does not limit the manner of comparing the similarity between adjacent image frames.

[0047] In block 308, in response to the similarity determined in block 304 being less than a similarity threshold, the computing device 110 can take the latter image frame of the two adjacent image frames as a key frame. According to the requirement of the size (e.g., the number of images) of the target set of style images to be generated, a suitable similarity threshold can be set. For example, if more images are required for the target set of style images to be generated, the similarity threshold can be set to be relatively large; if fewer images are required for the target set of style images to be generated, the similarity threshold can be set to be relatively small.

[0048] By determining the similarity between two adjacent image frames in block 306 and comparing the determined similarity with a similarity threshold, the computing device 110 can determine the latter image frame among the two adjacent image frames as a key frame when the determined similarity is less than the similarity threshold. For example, the computing device 110 can compare the similarity S1 between two adjacent image frames F1 and F2, and if it is determined that S1 is less than the threshold Sth, the computing device can determine the latter image frame F2 as a key frame.

[0049] After performing the operation in block 308, the computing device 110 can continue to select the next set of adjacent image frames from the arranged plurality of image frames and perform the operations in blocks 306 and 308 until the similarity comparison for all adjacent image frames in the arranged plurality of image frames is completed. For the case where the similarity comparison for all adjacent image frames is completed but one image frame is left, the remaining image frame can be determined as a key frame by default, or the remaining image frame can be compared with the previous key frame in similarity and determined as a key frame when the similarity is less than the similarity threshold.

[0050] In some embodiments, the computing device 110 can determine the earliest image frame in the plurality of image frames as a key frame. In other words, the computing device 110 can determine the initial image frame in the plurality of image frames as a key frame.

[0051] FIG. 4 shows a flowchart of an implementation method 400 of extracting key frames according to some other embodiments of the present disclosure. It should be understood that the numbering in the flowchart of the method 400 does not represent the order of execution of the steps, some or all of the steps can be executed in parallel, or the order of execution can be interchanged, and the present disclosure does not limit this. In addition, the method 400 in FIG. 4 can also include additional steps not shown and / or can omit the steps shown, and the scope of the present disclosure is not limited in this respect.

[0052] In block 402, the computing device 110 can extract a plurality of image frames with subtitles from the set of image frames. The computing device 110 can extract the image frames with subtitles from the set of image frames to obtain the plurality of image frames with subtitles. The computing device 110 can arrange the extracted plurality of image frames in chronological order to obtain the arranged plurality of image frames.

[0053] In block 404, the computing device 110 can obtain a first key frame from the plurality of image frames. For example, the computing device 110 can obtain the first key frame from the arranged plurality of image frames. In some embodiments, in an initial stage, the computing device 110 can determine the earliest image frame from the plurality of image frames as the first key frame. In addition, the computing device 110 can employ other criteria to determine the initial first key frame, which are not limited by the present disclosure. In other stages other than the initial stage, the computing device 110 can take the most recently determined key frame as the first key frame.

[0054] In block 406, the computing device 110 can select a to-be-compared image frame from the remaining image frames in the plurality of image frames. In some embodiments, the remaining image frames are the image frames in the plurality of image frames that have not been subjected to the similarity comparison. In other words, the remaining image frames are the image frames in the plurality of image frames that have not been subjected to the key frame determination operation.

[0055] In block 408, the computing device 110 can perform a similarity comparison between the to-be-compared image frame and the first key frame obtained in block 404. The computing device 110 can employ any suitable manner to perform the similarity comparison between the to-be-compared image frame and the obtained first key frame.

[0056] In block 410, the computing device 110 can determine whether the to-be-compared image frame is a key frame according to the similarity comparison result and the frame interval time between the to-be-compared image frame and the first key frame. In response to determining that the to-be-compared image frame is a key frame, the method 400 proceeds to block 412, and in block 412, the computing device 110 can update the first key frame with the to-be-compared image frame determined as a key frame in block 410. In response to determining that the to-be-compared image frame is not a key frame, the computing device 110 can continue to perform the operation in block 406 to select (e.g., sequentially select) the to-be-compared image frame from the remaining image frames in the plurality of image frames that have not been subjected to the similarity comparison, and determine whether the to-be-compared image frame is a key frame by continuing to perform the operations in blocks 408 and 410, until the key frame determination operation is completed for all the image frames in the plurality of image frames.

[0057] In some embodiments, the operation performed by the computing device 110 in block 410 can be implemented by determining that the to-be-compared image frame is a key frame in response to determining that the similarity is less than the similarity threshold value and the frame interval time between the to-be-compared image frame and the first key frame is greater than the predetermined time. In response to the to-be-compared image frame being determined as a key frame, the computing device 110 can update the first key frame with the to-be-compared image frame, thereby obtaining an updated first key frame.

[0058] After the computing device 110 determines that the to-be-compared image frame currently performing the key frame determination operation is a key frame, and updates the first key frame with the to-be-compared image frame to obtain an updated first image frame, the computing device 110 can continue to sequentially obtain the to-be-compared image frame from the remaining image frames, and continue to determine the key frame by performing the operations in blocks 408 and 410 until the key frame determination operation is completed for all the image frames in the plurality of image frames.

[0059] In addition, in block 410, in response to the similarity being greater than the similarity threshold or the frame interval time between the to-be-compared image frame and the first key frame being not greater than the predetermined time, the computing device 110 can determine that the to-be-compared image frame is a non-key frame. The computing device 110 can select a next image frame of the to-be-compared image frame (for example, sequentially selected from the arranged plurality of image frames), and perform a similarity comparison between the next image frame and the first key frame to determine whether the next image frame is a key frame. The computing device 110 can perform the operations in blocks 408 and 410 to determine whether the newly selected to-be-compared image frame is a key frame until the key frame determination operation is completed for all the image frames in the plurality of image frames.

[0060] FIG. 5 shows a schematic diagram of the key frames determined from the set of image frames corresponding to the video clip. The key frames in FIG. 5 are only schematic. A person skilled in the art can set appropriate parameters according to the needs of the set of target style images to be generated, and thus extract suitable key frames by using the method of key frame extraction according to the embodiments of the present disclosure.

[0061] As described above, at least part of the image frames in the set of image frames included in the video clip have corresponding subtitles, and therefore, after the key frames are selected, the computing device 110 also needs to associate the subtitles of the image frames in the set of image frames that are not determined as key frames with the image frames determined as key frames, so that the subtitles in the generated set of style images are complete.

[0062] In some embodiments, in response to there being at least one non-key frame between two adjacent key frames, the computing device 110 can determine the similarity between each of the at least one non-key frame and each of the two adjacent key frames. For example, the computing device 110 can determine that there are two non-key frames A1 and A2 between the key frames F1 and F2. The computing device 110 can perform a similarity comparison between each of the two non-key frames A1 and A2 and each of the key frames F1 and F2 to determine the similarity S11 between the non-key frame A1 and the key frame F1, the similarity S12 between the non-key frame A1 and the key frame F2, the similarity S21 between the non-key frame A2 and the key frame F1, and the similarity S22 between the non-key frame A2 and the key frame F2.

[0063] For each non-key frame, the computing device 110 can associate the subtitle associated with each non-key frame with the key frame of the two adjacent key frames that has a higher similarity with the non-key frame. Continuing with the example in the foregoing text, for non-key frame Al, assuming that the similarity S11 of non-key frame Al with key frame Fl is greater than the similarity S12 of non-key frame Al with key frame F2 (i.e., S11 > S12), the computing device can associate the subtitle associated with non-key frame Al with key frame Fl. For non-key frame A2, assuming that the similarity S21 of non-key frame A2 with key frame Fl is less than the similarity S22 of non-key frame A2 with key frame F2 (i.e., S21 < S22), the computing device can associate the subtitle associated with non-key frame A2 with key frame F2. Through the above processing of the subtitles of the non-key frames between the adjacent key frames, the subtitles in the video clip can be completely embodied in the set of target style images.

[0064] In some embodiments, after obtaining the key frames of the video clip, the computing device 110 can generate a style image corresponding to each of the key frames by using the trained style image generation model. In some embodiments, when generating the style image, the computing device 110 can receive prompt information, and input the received prompt information into the trained style image generation model.

[0065] In some embodiments, the prompt information provides constraint information for generating the style image corresponding to the key frame. For generating the style image of the key frame, the corresponding prompt information can be divided into the following parts: role prompt information, which is used to assign a role type to the trained style image generation model, and by providing the role type to the model, it helps the model to better understand the user's instructions; conversion style prompt information, which is used to indicate the type of style image that the model needs to convert the key frame into (for example, a Japanese manga style, an American manga style, an ink painting style, a cartoon style, etc.); and detail processing prompt information, which is used to indicate the processing rules that the model needs to follow in the conversion process, for example, accurately capturing object features, optimizing local details, or how to process the details, etc.

[0066] Table 1 shows an example of prompt information for converting an input image into a manga image in a Japanese manga style according to an embodiment of the present disclosure.

[0067] Table 1

[0068] The user can input the relevant constraint information in the prompt information according to the processing needs of the target style image, so that the model can process the input image (for example, the key frame) according to the constraints in the prompt information to generate the style image expected by the user.

[0069] It can be understood that the examples in the above Table 1 are only illustrative, and a person skilled in the art can set different prompt information according to needs to generate a desired target style image.

[0070] In some embodiments, after determining the associated subtitles for each key frame and obtaining the style image corresponding to each key frame, the computing device 110 can add the subtitles associated with each key frame to the predetermined position in the style image corresponding to the key frame to generate the target style image. In some embodiments, the predetermined position in the style image is associated with the type of the subtitles to be added.

[0071] In some embodiments, in response to the subtitles to be added including dialogue subtitles corresponding to a dialogue, the predetermined position includes a position in the corresponding style image within a predetermined range of an object initiating the dialogue.

[0072] A schematic diagram of a target style image after adding subtitles according to an embodiment of the present disclosure is shown in FIG. 6A. As shown in FIG. 6A, the converted style image in FIG. 6A is a comic style image, the key frame corresponding to the image is key frame 1 in FIG. 5, and the subtitles associated with the key frame are “Today is a sunny day, and the mood is also good”. The subtitles are dialogue subtitles, and the object initiating the dialogue subtitles is a panda in FIG. 6A. In addition, “Today is a sunny day” can be the subtitles associated with the previous non-key frame of key frame 1, and “the mood is also good” can be the subtitles associated with the current key frame 1. The subtitles associated with the non-key frame can be combined with the subtitles associated with the current key frame, and the combined subtitles are used as the subtitles associated with the current key frame. Accordingly, as shown in FIG. 6A, the computing device 110 can determine a predetermined position within a predetermined range of the object (i.e., the panda) initiating the dialogue in FIG. 6A, as illustrated in block 610, and add the subtitles associated with the current key frame at the predetermined position 610.

[0073] In some embodiments, in response to the object initiating the dialogue not being in the corresponding style image, the predetermined position includes a position in the corresponding style image opposite to a position of at least one object in the corresponding style image. Further, the computing device 110 can display an indication at the predetermined position, the indication indicating that the object initiating the dialogue is in at least one style image other than the corresponding style image.

[0074] A schematic diagram of a target style image with added subtitles according to an embodiment of the present disclosure is shown in FIG. 6B. As shown in FIG. 6B, the converted style image in FIG. 6B is a comic style image, which corresponds to the key frame 2 in FIG. 5, and the subtitles associated with the key frame 2 are "What are you eating?" and "Who is talking?". Also, the subtitle "Who is talking?" is a dialogue subtitle, and the object that initiates the dialogue subtitle is the panda in FIG. 6B. The subtitle "What are you eating?" is a dialogue subtitle, and the object that initiates the dialogue subtitle is not in FIG. 6B. Accordingly, as shown in FIG. 6B, the computing device 110 can determine a predetermined position within a predetermined range of the object (i.e., the panda) that initiates the dialogue in FIG. 6B, as illustrated in block 623, and add the associated subtitle "Who is talking?" at the predetermined position 623. The computing device 110 determines a predetermined position 621 for the subtitle "What are you eating?" at a position opposite to the object (i.e., the panda), and displays an indicator (e.g., an arrow 622 pointing outside of the style image 6B) at the predetermined position, which indicates that the object that initiates the dialogue is in at least one style image outside of the style image 6B.

[0075] In some embodiments, in response to the subtitle comprising a voiceover subtitle, the predetermined position for the voiceover subtitle is adjusted according to a position of an object in a corresponding style image, and the predetermined position is at least partially non-overlapping with the position of the object in the corresponding style image. For example, the predetermined position for the voiceover can be located at a position below or above the corresponding style image, and is at least partially non-overlapping with the position of the object in the image.

[0076] A schematic diagram of a target style image with added subtitles according to an embodiment of the present disclosure is shown in FIG. 6C. As shown in FIG. 6C, the converted style image in FIG. 6C is a comic style image, which corresponds to the key frame 3 in FIG. 5, and the subtitle associated with the key frame 3 is "The panda stopped and carefully distinguished where the sound came from", and the subtitle "The panda stopped and carefully distinguished where the sound came from" is a voiceover subtitle. Accordingly, as shown in FIG. 6C, the computing device 110 can identify a position of an object in FIG. 6C, and set the predetermined position for the voiceover subtitle at a position that is at least partially non-overlapping with the position of the object. As illustrated in FIG. 6C, the computing device 110 can set the predetermined position for the voiceover subtitle at the top right of the image 630, and is not overlapping with the object (i.e., the panda) in the image, so that the completeness of the picture can be presented to the user.

[0077] The computing device 110 can add the subtitles associated with each key frame to the corresponding style image at the predetermined positions in the style image according to the implementation described above, thereby generating the target style image. In some embodiments, the computing device 110 can obtain each target style image with added subtitles, and embed each obtained target style image into a style image template to generate a set of style images.

[0078] In some embodiments, the style image template corresponds to a format of the set of style images to be generated. The format of the style image template can match the screen features of the terminal device. For example, the style image template can be a strip-shaped template to match the screen features of a mobile terminal. Taking the strip-shaped template as an example, the strip-shaped template can be divided into a plurality of regions in a top-to-bottom direction, and each divided region corresponds to one or more style images. The generated set of style images, by being embedded into the style image template, can be displayed in the terminal device in the format of the template to facilitate the user to browse at the terminal device.

[0079] FIG. 7 shows an example diagram 700 of a style image template according to embodiments of the present disclosure. As shown in FIG. 7, the style image template in FIG. 7 includes four regions 710, 720, 730 and 740, which can be understood as being illustrative only. According to the actual needs of the set of target style images, the style image template can include any number of regions for accommodating corresponding target style images. Taking the style image template shown in FIG. 7 as an example, the four regions can accommodate target style image 1 to target style image 4, respectively.

[0080] The computing device 110 can embed target style image 1 to target style image 4 into the template 700 in the order of the style image template (from top to bottom, from left to right) and in the time sequence of target style image 1 to target style image 4, respectively. In this way, the set of style images (e.g., a set of images including target style image 1 to target style image 4) can be displayed in the terminal device in the format of the template 700 by being embedded into the template 700, to facilitate the user to browse at the terminal device. For example, for a mobile terminal, the user can browse at the mobile terminal by swiping up and down.

[0081] FIG. 8 shows a schematic block diagram of an example apparatus 800 according to some embodiments of the present disclosure. The apparatus 800 can be implemented by software, hardware, or a combination of both. As shown in FIG. 7, the apparatus 800 includes an obtaining module 810, an extracting module 820, a style image generation module 830 and an adding module 840.

[0082] In some embodiments, the obtaining module 810 can obtain a video clip, where the video clip includes a set of image frames and a plurality of subtitles. The extracting module 820 can extract a plurality of key frames from the set of image frames, where each key frame has an associated subtitle. The style image generation module 830 can generate a style image corresponding to each key frame of the plurality of key frames using a trained style image generation model. The adding module 840 can correspondingly add the subtitle associated with each key frame to the corresponding style image at a predetermined position in the corresponding style image to generate a target style image.

[0083] The apparatus 800 of FIG. 8 can be used to implement the processes described above in connection with FIGS. 1-7, and thus further description thereof will not be repeated here for brevity.

[0084] The division of modules or units in the embodiments of the present disclosure is illustrative, and is merely a logical function division. When actually implemented, another division manner can also be used. In addition, each functional unit in the disclosed embodiments can be integrated into one unit, or can be physically separated, or two or more units can be integrated into one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.

[0085] FIG. 9 illustrates a block diagram of an example device 900 that can be used to implement embodiments of the present disclosure. It should be understood that the device 900 illustrated in FIG. 9 is merely an example and should not be construed to limit the functionality and scope of the implementations described herein. For example, the device 900 can correspond to the computing device 110 described herein in connection with FIG. 1, and can be used to perform the processes described above in connection with FIGS. 1-7.

[0086] As shown in FIG. 9, the device 900 is in the form of a general-purpose computing device. Components of the computing device 900 can include, but are not limited to, one or more processors or processing units 910, a memory 920, a storage device 930, one or more communication units 940, one or more input devices 950, and one or more output devices 960. The processing unit 910 can be a real or virtual processor and is capable of executing various processing in accordance with programs stored in the memory 920. In a multi-processor system, multiple processing units execute computer-executable instructions in parallel to improve the parallel processing capability of the computing device 900.

[0087] The computing device 900 typically includes a plurality of computer storage media. Such media can be volatile, nonvolatile, removable, and / or non-removable media implemented in any method or technology for storage of information such as computer readable instructions, data structures, program modules or other data. Storage 920 can be volatile (such as RAM), non-volatile (such as ROM, EEPROM, flash memory, etc.), or some combination of the two. Storage 930 can be removable or non-removable and can include machine readable media such as flash drives, magnetic disks, or any other medium that can be used to store information and / or data (e.g., training data for training) and that can be accessed by the computing device 900.

[0088] The computing device 900 can further include additional removable / non-removable, volatile / non-volatile storage media. Although not shown in FIG. 9, a disk drive or other computer readable media drive can provide for reading and writing to a removable, non-removable, and / or non-volatile media such as a floppy disk, a ZIP® disk, a magnetic tape, or a flash drive. In such instances, the disk drive or other computer readable media drive can be connected to the bus by one or more data media interfaces. The memory 920 can include a computer program product 925 having one or more program modules configured to carry out the various methods or actions of the various implementations of the present disclosure.

[0089] The communication unit 940 enables communications with other computing devices over a communication media. Additionally, the functionality of the components of the computing device 900 can be implemented in a single computing cluster or a plurality of computer machines that are capable of communicating over a communication connection. Thus, the computing device 900 can operate in a networked environment using logical connections to one or more other servers, network personal computers (PCs), or another network nodes in the networked environment.

[0090] Input device 950 can be one or more input devices, such as a mouse, a keyboard, a trackball, etc. Output device 960 can be one or more output devices, such as a display, a speaker, a printer, etc. Computing device 900 can also communicate with one or more external devices (not shown) such as a storage device, a display device, etc. through communication unit 940, and with one or more devices that enable a user to interact with computing device 900, or any devices (e.g., a network card, a modem, etc.) that enable computing device 900 to communicate with one or more other computing devices. Such communication can occur via Input / Output (I / O) interface (not shown).

[0091] According to example implementations of the present disclosure, a computer readable storage medium is provided having computer executable instructions stored thereon, where the computer executable instructions are executed by a processor to implement the method described above. According to example implementations of the present disclosure, a computer program product is also provided that is tangibly stored on a non-transitory computer readable medium and includes computer executable instructions, where the computer executable instructions are executed by a processor to implement the method described above. According to example implementations of the present disclosure, a computer program product is provided having a computer program stored thereon, which when executed by a processor implements the method described above.

[0092] Various aspects of the disclosure are now described with reference to the drawings. In general, the drawings described below are diagrammatic and schematic representations of actual or conceptual structures and processes, and are not limiting of the scope of the present disclosure. In the drawings, the same reference numerals are used to represent similar or like items. The drawings are intended to be illustrative, and are not limiting of the scope of the present disclosure. The drawings are not necessarily to scale, and certain features can be exaggerated or omitted for the sake of clarity.

[0093] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions stored therein comprises an article of manufacture including instructions which implement aspects of the function / act specified in the flowchart and / or block diagram block or blocks.

[0094] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0095] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.

[0096] The above-described implementations of this disclosure are illustrative and not exhaustive, and are not limited to the disclosed implementations. Numerous modifications and adaptations will be apparent to those skilled in the art without departing from the scope and spirit of the disclosed implementations. The choice of words in this document is intended to best explain the principles of the implementations, practical application, or improvement over the technology in the market, or to enable other ordinary skilled in the art to understand the various implementations disclosed herein.

Claims

1. A method for generating style images, the method comprising: obtaining a video clip, wherein the video clip comprises a set of image frames and a plurality of subtitles; extracting a plurality of key frames from the set of image frames, wherein each key frame has an associated subtitle; generating, by using a trained style image generation model, a style image corresponding to each of the plurality of key frames; and adding the subtitle associated with the each key frame to the corresponding style image at a predetermined position to generate a target style image. 2.The method of claim 1, wherein the extracting a plurality of key frames comprises: extracting a plurality of image frames with subtitles from the set of image frames; arranging the plurality of image frames in a time sequence; determining a similarity between two adjacent image frames; and in response to the similarity being less than a similarity threshold, determining a latter image frame of the two adjacent image frames as a key frame. 3.The method of claim 2, wherein the extracting a plurality of key frames comprises: determining an earliest image frame of the plurality of image frames as a key frame. 4.The method of claim 1, wherein the extracting a plurality of key frames comprises: extracting a plurality of image frames with subtitles from the set of image frames; obtaining a first key frame from the plurality of image frames; selecting a to-be-compared image frame from remaining image frames of the plurality of image frames, the remaining image frames being image frames of the plurality of image frames that have not been compared for similarity; comparing the to-be-compared image frame with the first key frame for similarity; and determining whether the to-be-compared image frame is a key frame according to a result of the comparison for similarity and an interval time between the to-be-compared image frame and the first key frame. 5.The method of claim 4, wherein the determining whether the to-be-compared image frame is a key frame according to a result of the comparison for similarity and an interval time between the to-be-compared image frame and the first key frame comprises: in response to the similarity being less than a similarity threshold and the interval time between the to-be-compared image frame and the first key frame being greater than a predetermined time, determining the to-be-compared image frame as a key frame, and updating the first key frame with the to-be-compared image frame. 6.The method of claim 4, wherein the determining whether the to-be-compared image frame is a key frame according to a result of the comparison for similarity and an interval time between the to-be-compared image frame and the first key frame comprises: in response to the similarity being greater than a similarity threshold or the interval time between the to-be-compared image frame and the first key frame not being greater than a predetermined time, determining the to-be-compared image frame as a non-key frame. 7.The method of claim 1, further comprising: in response to there being at least one non-key frame between two adjacent key frames, determining a similarity between each non-key frame of the at least one non-key frame and each key frame of the two adjacent key frames; and for each non-key frame, associating the subtitle associated with the each non-key frame with a key frame of the two adjacent key frames that has a higher similarity with the each non-key frame. ​ ​ ​ 8. The method of claim 1, wherein generating, with a trained style image generation model, a style image corresponding to each key frame in the plurality of key frames comprises: inputting a received hint information to the trained style image generation model, wherein the hint information provides constraint information for generating the corresponding style image; and inputting the each key frame to the trained style image generation model to generate, by the trained style image generation model, the corresponding style image based on the hint information.

9. The method of claim 1, wherein the predetermined position is associated with a type of the subtitle.

10. The method of claim 9, wherein in response to the subtitle comprising a dialogue subtitle corresponding to a dialogue, the predetermined position comprises a position in the corresponding style image within a predetermined range of an object initiating the dialogue.

11. The method of claim 10, in response to the object initiating the dialogue not being in the corresponding style image, the predetermined position comprises a position in the corresponding style image opposite to a position of at least one object in the corresponding style image.

12. The method of claim 11, further comprising: displaying an indication marker at the predetermined position, the indication marker indicating the object initiating the dialogue in at least one style image out of the corresponding style image.

13. The method of claim 9, wherein in response to the subtitle comprising a voiceover subtitle or a caption subtitle, the predetermined position is adjusted according to a position of an object in the corresponding style image, and the predetermined position is at least partially non-overlapping with the position of the object.

14. The method of claim 1, further comprising: obtaining each target style image with added subtitle; and embedding the obtained each target style image into a style image template to generate a set of style images.

15. The method of claim 1, wherein the style image comprises a comic style image.

16. The method of claim 1, wherein the video clip comprises a short drama video.

17. An electronic device, comprising: at least one processing unit; at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, the instructions which when executed by the at least one processing unit cause the electronic device to perform the method according to any one of claims 1-16.

18. A computer readable storage medium having stored thereon a computer program which, when executed by a processor, implements the method according to any one of claims 1-16.

19. A computer program product having stored thereon a computer program which, when executed by a processor, implements the method according to any one of claims 1-16. ​ ​ ​ ​ ​ ​

Citation Information

Patent Citations

  • Method and device for video processing

    CN108683924A

  • Image processing method and device, equipment and storage medium

    CN109859298A

  • Video processing and subtitle detection model method and device

    CN113361462A

  • Hypervideo browsing using links generated based on user-specified content features

    US20140040273A1