A method and device for generating a comic video based on interactive labeling

By using an interactive annotation method that combines intelligent recognition and user interaction, comic videos are generated, solving the problems of poor flexibility and low accuracy in existing technologies, and achieving efficient and flexible dynamic video generation.

CN121239924BActive Publication Date: 2026-02-13SHENZHEN YOUYOU INTERNET TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511770743.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-13
Estimated Expiration
2045-11-28

AI Technical Summary

Technical Problem

Existing comic video generation technologies suffer from poor flexibility, low accuracy, and high technical barriers, making it impossible to efficiently generate dynamic videos.

Method used

An interactive annotation-based method is adopted. By generating a canvas covering the comic image, user interactive annotation and intelligent recognition detection are used, combined with partition annotation information, to generate a video frame sequence and process it frame by frame to form a comic video.

Benefits of technology

It enables efficient and accurate generation of dynamic videos from static comics, improving operational efficiency and video generation flexibility while lowering the technical threshold.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121239924B_ABST
    Figure CN121239924B_ABST
Patent Text Reader

Abstract

The application discloses a kind of cartoon video generation method and device based on interactive labeling, and generation method specifically includes: obtaining cartoon image;Response to grid labeling request, generate the canvas covering cartoon image, wherein the canvas includes multiple partitions, each partition respectively corresponds to the different grid area covering cartoon image, and the partition includes initial labeling information;Response to interactive labeling instruction, optimize initial labeling information to form partition labeling information;Combine the grid area of cartoon image and partition labeling information, give video frame sequence;Video frame sequence is processed in frame-by-frame writing mode, and cartoon video is generated.The application balances the flexibility and automation of cartoon grid by fusing user interactive labeling, under the guarantee of detection quality and batch processing, efficiently, high-accuracy static cartoon generates dynamic video.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of comic video generation, and particularly relates to a comic video generation method and device based on interactive labeling. BACKGROUND

[0002] Traditional comic reading is in a static page-turning mode, which cannot highlight key panels, is difficult to control narrative rhythm, has single reading experience and lacks certain immersion.

[0003] With the rise of short videos, the static page-turning mode cannot meet the needs of modern users for dynamic and immersive content, and some comic content creators begin to convert static comics into dynamic videos to improve the spread of comics and user stickiness.

[0004] Currently, the generation methods for comic videos mainly include preset template method, full-automatic recognition method and manual editing method. For example, patent application CN111415399A discloses an image processing method, device, electronic equipment and computer readable storage medium, which belongs to the full-automatic recognition method and includes the following steps: dividing a comic picture according to comic panels to generate a plurality of candidate pictures and a first arrangement order between the plurality of candidate pictures; for each candidate picture, extracting text information from the candidate picture, generating voice information corresponding to the text information, obtaining special effect information based on at least one of the picture content of the candidate picture and the semantics of the text information, and editing at least one of the candidate picture and the voice information based on the special effect information to generate a video segment with a target special effect matched with the candidate picture, wherein the target special effect is a special effect matched with the special effect information; and splicing the video segments respectively matched with the candidate pictures according to the first arrangement order to generate a target video matched with the comic picture.

[0005] However, the fixed camera movement template used in the preset template method has the problem of panel layout adaptability for different comic types; the accuracy of the image segmentation algorithm of the full-automatic recognition method is unstable, and the calculation cost is high; and the manual editing method requires professional video editing software operation skills and takes a long production time.

[0006] Therefore, in view of the problems of poor flexibility, low accuracy and high technical threshold in the existing video generation technology, how to generate dynamic videos from static comics with high efficiency and high accuracy while ensuring detection quality and batch processing is a problem to be solved by those skilled in the art. SUMMARY

[0007] In view of the defects in the prior art, the application provides a cartoon video generation method and device based on interactive labeling, and the generation method specifically comprises the following steps: obtaining a cartoon image; in response to a grid labeling request, generating a canvas covering the cartoon image, wherein the canvas comprises a plurality of partitions, each partition corresponds to a different grid area of the cartoon image, and each partition comprises initial labeling information; in response to an interactive labeling instruction, optimizing the initial labeling information to form partition labeling information; combining the grid areas of the cartoon image and the partition labeling information to give a video frame sequence; and processing the video frame sequence in a frame-by-frame writing manner to generate a cartoon video. The application balances the flexibility and automation of cartoon grid by fusing user interactive labeling, guarantees detection quality and batch processing, and efficiently and accurately generates dynamic video from static cartoon.

[0008] In a first aspect, the application provides a cartoon video generation method based on interactive labeling, which specifically comprises the following steps:

[0009] Obtaining a cartoon image;

[0010] In response to a grid labeling request, generating a canvas covering the cartoon image, wherein the canvas comprises a plurality of partitions, each partition corresponds to a different grid area of the cartoon image, and each partition comprises initial labeling information;

[0011] In response to an interactive labeling instruction, optimizing the initial labeling information to form partition labeling information;

[0012] Combining the grid areas of the cartoon image and the partition labeling information to give a video frame sequence;

[0013] Processing the video frame sequence in a frame-by-frame writing manner to generate a cartoon video.

[0014] Further, the labeling information is divided into initial labeling information and partition labeling information, and the labeling information comprises position information, partition code, entry and exit information, and state information. The entry and exit information comprises an entry and exit mode, the entry and exit mode comprises an entry and exit manner and an entry and exit form, and the state information is a parameter for presenting periodic changes in size of the cartoon video.

[0015] Further, in response to the grid labeling request, the canvas covering the cartoon image is generated, specifically comprising the following steps:

[0016] Receiving an automatic grid labeling instruction containing a type request;

[0017] Based on edge detection, performing contour recognition on the cartoon image to determine each grid area of the cartoon image;

[0018] According to the coordinate information and mapping relationship of the grid areas, giving initial position information of the partitions corresponding to each grid area, labeling partition edge frames and partition center points of each partition;

[0019] Based on the type request, determine the order of each partition, give the partition code, form and mark the partition line linking the center points of each partition based on the partition code;

[0020] Determine the initial in-out information and initial state information of each partition, and generate a canvas covering the comic image.

[0021] Further, it also includes saving and configuring the initial annotation information or partition annotation information to a JSON file to form an annotation information configuration file.

[0022] Further, in response to the grid annotation request, a canvas covering the comic image is generated, which specifically includes the following steps:

[0023] Receive and parse the annotation information configuration file to give the initial position information of each partition corresponding to each grid area, mark the partition edge frame and partition center point of each partition, give the partition code, form and mark the partition line linking the center points of each partition based on the partition code, give the initial in-out information and initial state information of each partition, and form a canvas covering the comic image.

[0024] Further, in response to the interactive annotation instruction, the initial annotation information is optimized to form the partition annotation information, which specifically includes the following steps:

[0025] Select and drag any partition edge frame, and / or select and drag any partition center point, optimize the initial position information of the corresponding partition, and form the partition position information.

[0026] Further, in response to the interactive annotation instruction, the initial annotation information is optimized to form the partition annotation information, which specifically includes the following steps:

[0027] Select and adjust the initial in-out information to form the partition in-out information;

[0028] And / or;

[0029] Select and adjust the initial state information to form the partition state information.

[0030] Further, in combination with the grid area and partition annotation information of the comic image, a video frame sequence is given, which specifically includes the following steps:

[0031] Based on the grid area of the comic image, determine the corresponding key frame;

[0032] Based on the partition corresponding to each grid area, perform sequence processing on each key frame to form a key frame sequence;

[0033] According to the in-out information of each adjacent key frame, and through an interpolation algorithm, intermediate frames are generated between each adjacent key frame, the in-out information of each intermediate frame is determined, and an intermediate frame sequence is formed.

[0034] The key frame sequence and the intermediate frame sequence are fused to give a video frame sequence.

[0035] Further, according to the in-out information of each adjacent key frame, and through an interpolation algorithm, intermediate frames are generated between each adjacent key frame, the in-out information of each intermediate frame is determined, and an intermediate frame sequence is formed, specifically including the following steps:

[0036] The duration of each adjacent key frame is determined by traversing the key frame sequence and combining the in-out information of each adjacent key frame;

[0037] The number of corresponding intermediate frames is given by fusing the duration of each adjacent key frame and the preset frame rate;

[0038] The in-out information of each intermediate frame is determined through linear interpolation and in combination with the in-out information of adjacent key frames;

[0039] The intermediate frames are generated to form an intermediate frame sequence.

[0040] Further, before fusing the key frame sequence and the intermediate frame sequence, the following steps are further included:

[0041] The corresponding relationship between each key frame and the partition determined based on the grid area is obtained, the distance between the center points of each partition is calculated through Euclidean distance, and the positional relationship between each key frame is given;

[0042] The starting key frame is determined and labeled from the key frame sequence, the remaining key frames in the key frame sequence are traversed, and the key frame closest to the starting key frame is labeled as the next key frame;

[0043] The next key frame is updated as the starting key frame, the unmarked key frames in the key frame sequence are traversed, and the key frame closest to the starting key frame is labeled as the next key frame;

[0044] The above labeling process of the next key frame is repeated until the labeling of all key frames in the key frame sequence is completed, and an optimized path of the key frame sequence is formed;

[0045] Based on the optimized path of the key frame sequence, the key frame sequence is optimized, and the intermediate frame sequence is also optimized.

[0046] Further, the video frame sequence is processed in a frame-by-frame writing manner to generate a comic video, specifically including the following steps:

[0047] In a frame-by-frame writing mode, the key frame and corresponding information, state information and intermediate frame and corresponding information are read and analyzed respectively to form a key frame segment and an intermediate frame segment, and a comic video is generated.

[0048] Further, the key frame and corresponding information, state information are read and analyzed to form a key frame segment, specifically including the following steps:

[0049] The key frame is obtained, the fade-in function and the shape parameter are fused, and an initial key frame segment is formed;

[0050] The dwell time period of the initial key frame segment is determined;

[0051] The initial key frame segment is optimized by superimposing a breathing function in the dwell time period to form a key frame segment.

[0052] Further, the intermediate frame and corresponding information are read and analyzed to form an intermediate frame segment, specifically including the following steps:

[0053] The intermediate frame is obtained, the fade-in function and the shape parameter are fused, and an intermediate frame segment is formed.

[0054] Further, the key frame segment and the intermediate frame segment are formed to generate a comic video, specifically including the following steps:

[0055] An audio file is obtained;

[0056] Based on the key frame segment and the intermediate frame segment, an initial comic video is generated;

[0057] The audio file is subjected to format processing suitable for the comic video to form a to-be-combined audio file;

[0058] The initial comic video and the to-be-combined audio file are fused, and the time lengths of the initial comic video and the to-be-combined audio file are compared to determine the time length of the comic video as the shorter one, and the comic video is generated.

[0059] In a second aspect, the present application also provides a comic video generation device based on interactive labeling, which adopts the comic video generation method based on interactive labeling as described above, and specifically includes:

[0060] A collection unit is configured to obtain a comic image, a grid labeling request and an interactive labeling instruction;

[0061] A labeling generation unit is configured to generate a canvas covering the comic image in response to the grid labeling request, wherein the canvas includes a plurality of partitions, each of which corresponds to a different grid area of the comic image, and each partition includes initial labeling information; and optimize the initial labeling information to form partition labeling information in response to the interactive labeling instruction;

[0062] A video generation unit is configured to generate a video frame sequence in combination with a grid area of a comic image and partition label information, and to process the video frame sequence in a frame-by-frame writing manner to generate a comic video.

[0063] The comic video generation method and device based on interactive labeling provided by the application have at least the following beneficial effects:

[0064] (1) The comic video generation scheme of the application balances the flexibility and automation of comic grid by fusing user interactive labeling and intelligent recognition and detection, and efficiently and accurately generates dynamic video from static comic under the guarantee of detection quality and batch processing.

[0065] (2) The application improves the operation efficiency of comic grid by designing a comic labeling method of a covering canvas and a drag-and-drop interaction, realizes real-time adjustment of pixel-level precision, and has high reliability without repeated trial and error. Meanwhile, the application organically combines the optimization of partition position information including a partition edge frame and a partition center point adjustment, and the optimization of partition in-out information and partition state information, adapts to different scenarios, and realizes fast labeling, fine adjustment and batch processing.

[0066] (3) The application labels different in-out information in different grid areas of a comic, embeds different types of easing functions, so that the generated comic video realizes fine effect control of video frame granularity, and enhances visual fluency. Meanwhile, the application labels state information of key frames, fuses a breathing function and an easing function, presents dynamic effects, and improves the expressiveness of key frames. BRIEF DESCRIPTION OF DRAWINGS

[0067] Figure 1 A flowchart of the comic video generation method based on interactive labeling provided by the application is shown.

[0068] Figure 2 An information interaction diagram of an embodiment of the application is shown.

[0069] Figure 3 A canvas flowchart of generating a covering comic image of an embodiment of the application is shown.

[0070] Figure 4 A grid area diagram of an embodiment of the application is shown.

[0071] Figure 5 An interactive labeling interface diagram of an embodiment of the application is shown.

[0072] Figure 6 A flowchart of generating a video frame sequence of an embodiment of the application is shown.

[0073] Figure 7A flowchart of forming an intermediate frame sequence according to an embodiment of the present application is provided.

[0074] Figure 8 A flowchart of video frame path optimization according to an embodiment of the present application is provided.

[0075] Figure 9 A flowchart of forming a key frame segment according to an embodiment of the present application is provided.

[0076] Figure 10 A structure diagram of a cartoon video generation device based on interactive labeling according to an embodiment of the present application is provided. DETAILED DESCRIPTION

[0077] In order to better understand the above technical solutions, the above technical solutions will be described in detail below in combination with the accompanying drawings and specific embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0078] The terms used in the embodiments of the present application are only for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms "a", "said" and "the" used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise. "Multiple" generally includes at least two.

[0079] It should also be noted that the terms "include", "contain" or any other variants thereof are intended to cover non-exclusive inclusion, so that the goods or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or include elements inherent to such goods or devices. Without more limitation, the element defined by the sentence "including a" does not exclude the presence of another identical element in the goods or devices including the element.

[0080] Static cartoons cannot highlight key frames, have poor narrative rhythm, lack immersion, and need to generate dynamic videos. Existing templating tools provide fixed camera movement templates (such as left-right scanning and up-down scrolling), users apply the templates to cartoon images, and the system generates videos according to the preset path, but the number of templates is limited, which cannot adapt to different frame layouts of cartoon images (such as irregular frames, multi-layered frames, etc.), lacks flexibility, cannot adjust the dwell time and zoom ratio for specific frames, and cannot label key details (such as close-up shots of characters, dialogue bubbles, etc.), resulting in poor narrative effect.

[0081] A convolutional neural network is used to train a cell detection model (for example, YOLO, Mask R-CNN, etc.), which can automatically identify the cell boundaries in a comic image and generate a fixed sequence of shot motion video. However, a large amount of labeled data is required to train the model, the development cost is high, and the cell rules of different comic styles (for example, Japanese comics, American comics, and Chinese comics) differ greatly, resulting in unstable model generalization performance, low recognition accuracy (about 60%-75%) for complex cells (such as perspective cells and frameless cells), high computational overhead, the need for GPU acceleration, high deployment cost, and the inability to understand the narrative logic of the comic. Using a manual editing tool (such as Premiere, Final Cut Pro, etc.), the user needs to first import the comic image into the video editing software, then manually set the position, scaling, and rotation of each shot using the keyframe function, and adjust the duration and transition effects on the timeline, and then render the output video. However, this method requires professional video editing skills, has a high learning cost (usually several weeks to several months of learning), takes 30 minutes to 2 hours to produce a video for one page of comic, is inefficient, and is expensive. Also, it cannot be batch processed and requires a lot of repetitive work.

[0082] To solve the problems of poor flexibility, low accuracy, and high technical threshold in the prior art, the present application balances flexibility and automation by combining user interaction labeling and intelligent recognition detection, and designs a flexible and efficient semi-automatic solution.

[0083] In a first aspect, as shown in Figure 1 The present application provides a comic video generation method based on interactive labeling, which specifically includes the following steps:

[0084] Obtain a comic image;

[0085] In response to a cell labeling request, generate a canvas covering the comic image, wherein the canvas includes a plurality of partitions, each partition corresponding to a different cell region of the comic image, and the partition includes initial labeling information;

[0086] In response to an interactive labeling instruction, optimize the initial labeling information to form partition labeling information;

[0087] Combine the cell regions of the comic image and the partition labeling information to give a video frame sequence;

[0088] Process the video frame sequence in a frame-by-frame writing manner to generate a comic video.

[0089] Referring to Figure 2In the example, the comic image refers to a comic picture uploaded by the user to the web interface, which can be in common image formats such as JPG / PNG / BMP. The grid annotation request refers to the identification instruction triggered by the user for the grid area of the comic image, which automatically locates each grid area of the comic through edge detection and contour analysis. The interactive annotation instruction refers to the user's operation on the canvas partition. It is worth noting that the "partition" is for the canvas, and the "grid area" is for the comic image. The partition and the grid area have a corresponding relationship, and the canvas composed of the partition is overlaid on the comic image composed of the grid area. The present application adopts a Canvas multi-layer superposition architecture, the bottom layer displays the original comic image, and the top layer generates a transparent Canvas canvas, which is overlaid on the comic image and used to draw annotation and interactive annotation instructions. Each partition of the canvas corresponds to a grid area of the comic image.

[0090] Further, the annotation information is divided into initial annotation information and partition annotation information. The partition annotation information is obtained by optimizing the initial annotation information according to the user operation. The annotation information includes position information, partition code, entry and exit information, and state information. The entry and exit information includes entry and exit mode, and the entry and exit mode includes entry and exit mode (for example, uniform speed, slow-fast, fast-slow, jump, etc., which can be represented by different types of motion functions), entry and exit form (for example, rotation, scaling, etc., which can be represented by different form parameters), and state information is a parameter that presents periodic changes in size of the comic video, which can be represented by different types of breathing functions.

[0091] In combination with the grid area and the partition annotation information, the video frame sequence is determined, and then each frame is encoded, so that the comic video is generated.

[0092] The existing preset template method pursues automation but sacrifices flexibility, and the manual editing method is flexible but inefficient. Users face the dilemma of "either accepting unsuitable templates or spending a lot of time manually producing". The present application provides a semi-automatic and user-controllable balanced solution by combining intelligent recognition detection and user interactive annotation, solving the contradiction between flexibility and automation.

[0093] For the annotation of the comic image, it is not intuitive to input coordinates by relying on a slider or a text box. Users need to repeatedly try and adjust parameters. It may take 5-10 adjustments to accurately position a single annotation point. There is also a lack of drag-and-drop interactive interface, and real-time preview effect cannot be achieved.

[0094] As shown in Figure 3 In response to the grid annotation request, a canvas covering the comic image is generated, specifically including the following steps:

[0095] Receiving an automatic grid annotation instruction containing a type request;

[0096] Based on edge detection, the outline of the comic image is recognized, and each grid area of the comic image is determined;

[0097] According to the coordinate information of the grid area and the mapping relationship, the initial position information of each subarea corresponding to the grid area is given, and the subarea edge frame and the subarea center point of each subarea are labeled;

[0098] Based on the type request, the order of each subarea is determined, the subarea code is given, and the subarea connection line connecting each subarea center point is formed and labeled based on the subarea code;

[0099] The initial in-out information and the initial state information of each subarea are determined, and the canvas covering the comic image is generated.

[0100] The labeling scheme for the comic image first receives an automatic grid labeling instruction containing a type request. The type request can be a user-selected grid area ordering method. In an embodiment, the type request includes confidence ordering (i.e., preferentially displaying detected accurate grids), vertical ordering (mainly suitable for strip comics), horizontal ordering (mainly suitable for American comics), and Z-shaped ordering (mainly suitable for traditional layouts of Japanese comics, Chinese comics, etc.). After receiving the automatic grid labeling instruction, an edge detection algorithm (such as the Canny edge detection algorithm) is used to perform contour recognition on the comic image, extract the edge frame of each grid area in the comic image, and locate the boundary of each grid area through contour analysis.

[0101] Based on edge detection, the outline of the comic image is recognized, and each grid area of the comic image is determined, specifically including:

[0102] Smooth the comic image through Gaussian filtering to remove noise;

[0103] Calculate the gradient amplitude and direction (which can be done through the Sobel operator);

[0104] Refine the edges through non-maximum suppression;

[0105] Through double-threshold detection (a low threshold of 50 and a high threshold of 150 can be set), strong edges and weak edges are extracted, and each grid area of the comic image is determined.

[0106] Next, each detected grid area (containing a grid area of a comic image as shown in Figure 4 in the original image coordinate system, such as 2000x3000 pixels) is converted to the display coordinate system (such as 800x1200 pixels, controlled by CSS), ensuring that the labeling position is consistent with the picture seen by the user, and the corresponding subarea edge frame and subarea center point are generated for each grid area.

[0107] According to the coordinate relationship and mapping relationship of the sub-grid area, the initial position information of the sub-area corresponding to each sub-grid area is given, specifically including:

[0108] According to the coordinate relationship of each sub-grid area, combined with the preset proportional conversion coefficient, the initial position information of the sub-area corresponding to each sub-grid area is given.

[0109] That is, the original coordinates of each sub-grid area are divided by the scale coefficient to obtain the Canvas drawing coordinates. Of course, when optimizing the initial labeling information under the subsequent interactive labeling instruction, the user clicks the Canvas coordinate, multiplies the scale coefficient, and then corresponds to the original coordinates of each sub-grid area.

[0110] Then, according to the type request (i.e., the user's selected sub-grid sorting method), each sub-area is assigned a serial number, the traversal order is marked, and the center points of each sub-area are connected (which can be presented in the form of a colored dashed line) according to the traversal order, and the moving path of the shot is displayed intuitively, as shown in Figure 5 .

[0111] Finally, the initial entry and exit method of each sub-grid (such as "uniform speed"), the initial entry and exit form (such as "rotation"), and the initial state information (such as "no breathing effect") are determined, and the above information is superimposed on the canvas covering the comic image to form an interactive labeling interface, on which the user can interact. In this example, most of the accurate sub-areas are automatically generated by the intelligent detection algorithm, and the user only needs to fine-tune a small number of irregular sub-areas through interactive labeling, reducing repeated clicking and greatly improving labeling efficiency; multiple sorting modes cover various layouts such as Japanese comics, American comics, Chinese comics, and strip comics, saving sorting time; visual sub-area connection lines and serial number labels allow users to visually see the shot path, predict the narrative effect in advance, and avoid rework after generating the video. In addition, the edge box and center point labeling also provide a data basis for subsequent path optimization, and the preset entry and exit information and state information provide parameters for subsequent video frame generation, improving the efficiency of comic video generation.

[0112] The confidence sorting in the type request refers to the priority sorting of each sub-grid area in the comic image obtained by edge detection according to the confidence score. The confidence sorting usually puts the sub-grid area most likely to be the main sub-grid area in the front row. The confidence score is calculated by the ratio of the area of the sub-grid area to the total area of the comic, and is specifically represented as:

[0113]

[0114] Wherein, confidence is the confidence score, area is the pixel area of the detected sub-grid area, min_area is 1% of the total area of the comic image, and min() represents the minimum.

[0115] In response to the interactive labeling instruction, the initial labeling information is optimized to form the partition labeling information, specifically including the following steps:

[0116] Selecting and dragging any partition edge frame and / or selecting and dragging any partition center point optimizes the initial position information of the corresponding partition to form the partition position information.

[0117] In a specific example, the user can operate the partition edge frame and the partition center point on the Canvas by means of a mouse, touch, etc. (click to select, drag), that is, adjust and optimize the initial position information. It can be understood that after moving the partition center point to the critical position, the mapping of the original partition and the grid area will change to form the correspondence between the partition and the new grid area. The critical position can be pre-set according to the size, contour and edge of each grid area of the comic image. As shown in Figure 5 The #6 partition center point position is moved to the right, and if the partition center point crosses the edge of the grid area corresponding to the #6 partition, the #6 partition needs to be re-corresponded to the grid area corresponding to the original #7 partition.

[0118] Optimizing the initial position information of the corresponding partition specifically includes:

[0119] Clicking the four corners or edges of the partition edge frame to drag the size of the partition;

[0120] Clicking the partition center point to drag the center position of the partition;

[0121] Dragging the partition edge frame and moving the partition center point to realize the synchronous optimization of the center position and size of the partition.

[0122] When the user interacts with the labeling by dragging, the partition edge frame, the partition center point and the connection line (traversal order) are updated in real time, and the user can directly see the adjustment and optimization effect. Through the drag-type interaction and real-time preview, the production time is shortened from several hours to several minutes, without the need for professional skills, greatly reducing the technical threshold and production time. In addition, the drag-type comic image labeling has an operation efficiency 60% higher than that of the traditional slider input, and the real-time adjustment of pixel-level precision also does not require repeated trial and error. Different interactive modes (clicking, dragging) are organically combined to adapt to different scenarios such as fast labeling, fine adjustment and batch processing.

[0123] In response to the interactive labeling instruction, the initial labeling information is optimized to form the partition labeling information, specifically including the following steps:

[0124] Selecting and adjusting the initial entry and exit information to form the partition entry and exit information;

[0125] And / or;

[0126] The initial state information is selected and adjusted to form the partition state information.

[0127] In this example, the user corrects the partition attributes generated by the intelligent detection by adjusting the relevant parameters (entry and exit information, state information), to form entry and exit information and state information that are more in line with the narrative requirements. Specifically, the dwell time is adjusted (e.g., from 5 seconds to 8 seconds) through a slider or input box, the panning duration is adjusted (e.g., from 2 seconds to 1.5 seconds) through a slider or input box, the easing type is selected (e.g., from "linear" to "ease-in-out") through a drop-down menu, the rotation angle is adjusted (e.g., from 0° to -15°, simulating a squint angle) through a slider, and the breathing effect is enabled (e.g., breathe=true, allowing the static partitions to have a slight zoom) through a checkbox. By adjusting the dwell time, the key partitions (e.g., character dialogues, key actions) are highlighted, and the narrative rhythm is optimized; by selecting the easing type (e.g., "ease-in-out"), the panning and rotation movements are more natural, and the movement fluidity is improved; by adjusting the state information (e.g., the breathing effect), the picture is more dynamic, and the visual vividness is enhanced. The user also deeply participates in the video generation process and can adjust the parameters according to the comic style (e.g., humor, etc.) to form a unique video style. Based on the intelligent detection, the user quickly optimizes the position, time, and visual attributes of the partitions, improving the accuracy, rhythm, and vividness of the comic video generation.

[0128] In addition, the initial annotation information or the partition annotation information can be saved after being obtained. That is, the initial annotation information or the partition annotation information is saved and configured to a JSON file to form an annotation information configuration file. Based on the annotation information configuration file, the next comic image annotation can be performed in response to a partition annotation request.

[0129] Therefore, in response to the partition annotation request, the canvas covering the comic image can also specifically include the following steps:

[0130] The annotation information configuration file is received and parsed to give the initial position information of the partitions corresponding to each partition area, annotate the partition edge frame and partition center point of each partition, give the partition code, form and annotate the partition connection line linking the partition center points based on the partition code, give the initial entry and exit information and initial state information of each partition, and form the canvas covering the comic image.

[0131] It can be understood that the prior art lacks a configuration template system and a standardized data format, and the labeling data cannot be saved and reused, and each time the same type of comic image is processed, it needs to be labeled again, resulting in a large amount of repetitive labor and low efficiency when processing series comic images (such as serialized works). The present application provides a mechanism for saving and reusing labeling configuration through a saved labeling information configuration file, supporting saving, loading, importing, and exporting JSON configuration. For serialized comics, after the first episode is labeled and saved, the second to tenth episodes are directly loaded and fine-tuned, and the production efficiency is improved by 90%; after the designer labels, the JSON is exported, and the production personnel import the JSON to batch generate videos, which is suitable for team collaboration scenarios; through version management, multiple configuration versions (quick version, detailed version, music version) are saved, and the selection is made according to the release platform.

[0132] As shown in Figure 6 , in combination with the grid area and partition labeling information of the comic image, a video frame sequence is given, specifically including the following steps:

[0133] Based on the grid area of the comic image, the corresponding key frame is determined;

[0134] Based on the partition of each grid area, each key frame is sequentially processed to form a key frame sequence;

[0135] According to the in-out information of each adjacent key frame, and through an interpolation algorithm, intermediate frames are generated between each adjacent key frame, the in-out information of each intermediate frame is determined, and an intermediate frame sequence is formed;

[0136] Fusion of key frame sequence and intermediate frame sequence, give video frame sequence.

[0137] As can be understood, a keyframe refers to a static frame corresponding to each segmented area, which is the core frame that "stays" in the comic video (such as character dialogue, key actions, etc.). In this example, the coordinates, size, and other information of each segmented area are first obtained through intelligent segmentation detection (edge ​​detection, contour recognition) or interactive annotation (applied to the partitions of the canvas and then mapped to the corresponding segmented areas). The central area of ​​each segmented area (e.g., x=840, y=480, width=400, height=200) is taken as the focus area of ​​the keyframe. Then, the keyframes are arranged according to the user annotation order to form a traversal order, simulating the order in which humans read comics, ensuring the coherence of the narrative logic. Based on the entry and exit information of each adjacent keyframe, dynamic transition frames, i.e., intermediate frames, are formed to connect adjacent keyframes, and the entry and exit information of each intermediate frame is determined to form an intermediate frame sequence. The keyframes and intermediate frames are combined in order to generate a continuous video frame sequence. By using keyframe serialization, jump shots are avoided, enhancing the audience's immersion. Easing function interpolation is used to make the motion of intermediate frames have a process of "acceleration and deceleration" (such as "ease in and ease out" effect), avoiding the stiffness of "uniform motion" in existing technologies. Breathing effects can be superimposed on the pause phase of keyframes to make static images have subtle dynamics (such as "the character's hair swaying slightly"), eliminating the illusion of "stuttering".

[0138] Furthermore, such as Figure 7 As shown, based on the entry and exit information of each adjacent keyframe, and through an interpolation algorithm, intermediate frames are generated between each adjacent keyframe, the entry and exit information of each intermediate frame is determined, and an intermediate frame sequence is formed. Specifically, the steps include:

[0139] Traverse the keyframe sequence and combine the entry and exit information of each adjacent keyframe to determine the duration of each adjacent keyframe.

[0140] By combining the duration of each adjacent keyframe with the preset frame rate, the corresponding number of intermediate frames is given;

[0141] By using linear interpolation and combining the entry and exit information of adjacent keyframes, the entry and exit information of each intermediate frame is determined.

[0142] Generate intermediate frames to form an intermediate frame sequence.

[0143] In this example, the duration refers to the transition time between two adjacent keyframes. The transition time is determined by the entry and exit information. First, the entry and exit information of adjacent keyframes is extracted by traversing the keyframe sequence, and the transition time is determined; then, according to the relationship "number of intermediate frames = duration x preset frame rate", the number of intermediate frames is determined; then, the entry and exit information of each intermediate frame is determined by gradually changing calculation between the entry and exit information (such as position, size, rotation angle) of adjacent keyframes according to normalized time (t∈[0,1]) (which can be combined with the easing function and shape parameters to make the interpolation result more consistent with the natural motion law), and the entry and exit information of each intermediate frame is determined. According to the interpolation result (such as the x, y coordinates of each frame, the width, height size, and rotation angle), each intermediate frame picture is generated using OpenCV affine transformation, that is, a continuous intermediate frame sequence.

[0144] The prior art often uses uniform motion, lacks the easing curve of professional animation software, and makes the lens motion lack acceleration and deceleration, the visual effect is stiff, and the comedy tension cannot be expressed. The present application improves the smoothness and accuracy of the video by precisely controlling the number of intermediate frames and the transition effect.

[0145] The number of intermediate frames is calculated to ensure that the frame rate of the transition effect is adapted, and the stuttering caused by insufficient number of intermediate frames in the prior art is avoided; the entry and exit information of the intermediate frame is gradually changed from the "previous keyframe" to the "next keyframe" through linear interpolation (combined with the easing function), avoiding the jump transition in the prior art; the rotation angle interpolation of the intermediate frame makes the transition effect contain rotation, supports advanced dynamic effects, and improves the three-dimensional expression of the picture; each annotation point can independently select the easing type, rather than globally uniform, to realize fine motion effect control.

[0146] As non-limiting examples, the easing function in the present application includes but is not limited to: linear, ease_in, ease_out, ease_in_out, ease_in_cubic, ease_out_cubic, bounce, elastic, corresponding to uniform, slow→fast, fast→slow, slow→fast→slow, slow→slow→extremely fast, extremely fast→slow, bounce, spring oscillation, etc. Motion type.

[0147] In another embodiment, as shown in Figure 8 Before fusing the keyframe sequence and the intermediate frame sequence, the following steps are further included:

[0148] The corresponding relationship between each keyframe and the partition determined based on the grid area is obtained, the distance between the center points of each partition is calculated by Euclidean distance, and the position relationship between each keyframe is given;

[0149] determining and labeling a starting key frame from the key frame sequence, traversing the remaining key frames in the key frame sequence, and labeling the key frame closest to the starting key frame as the next key frame;

[0150] updating the next key frame as the starting key frame, traversing the unlabeled key frames in the key frame sequence, and labeling the key frame closest to the starting key frame as the next key frame;

[0151] repeating the above labeling process of the next key frame until the labeling of all key frames in the key frame sequence is completed, forming an optimized path of the key frame sequence;

[0152] optimizing the key frame sequence and the intermediate frame sequence based on the optimized path of the key frame sequence.

[0153] In this example, the Euclidean distance between any two partition center points is calculated based on the determined key frame sequence and the corresponding partition center point coordinates of each key frame, and a distance matrix is generated to represent the positional relationship between the key frames. The Euclidean distance is specifically represented as:

[0154]

[0155] where dis(p1, p2) is the Euclidean distance between partition center points p1 and p2, x1 is the horizontal coordinate of partition center point p1, y1 is the vertical coordinate of partition center point p1, x2 is the horizontal coordinate of partition center point p2, and y2 is the vertical coordinate of partition center point p2.

[0156] In a certain example, the optimization of the key frame sequence and the intermediate frame sequence is as follows: a starting key frame is determined from the key frame sequence (e.g., the first key frame #1 labeled by the user); the remaining unvisited key frames are traversed, and the key frame closest to the current starting frame is selected as the next key frame (e.g., the nearest frame of #1 is #3); the next key frame (#3) is updated as the new starting frame, and the above steps are repeated until all key frames are labeled, obtaining an optimized path of the key frame sequence (e.g., the nearest frame of #3 is #2, the nearest frame of #2 is #4, and the final path is #1→#3→#2→#4). According to the optimized path, the original key frame sequence is adjusted to the order of the optimized path (e.g., from #1→#2→#3→#4 to #1→#3→#2→#4); and according to the new key frame order, the intermediate frame sequence is adjusted accordingly (e.g., the translation frames of #2→#3 are changed to the translation frames of #3→#2, adjusting the translation direction and duration). Finally, the optimized key frame sequence and the optimized intermediate frame sequence are output.

[0157] Since the user's manual annotation sequence may not be the optimal path (such as first annotating the upper left corner, then jumping to the lower right corner, and then returning to the left), the lens jumps greatly, and the viewing experience is poor. The application introduces a path optimization algorithm to reduce lens jumping and improve visual coherence and narrative logic. User tests show that the total length of the path is reduced by an average of 35%~45%, long distance jumps (>500px) are reduced by 70%; typical comic scenes (10-20 annotation points) take less than 1ms to optimize, and the response speed of the generated video is fast.

[0158] In another embodiment, the video frame sequence is processed in a frame-by-frame write-in manner to generate a comic video, specifically including the following steps:

[0159] In a frame-by-frame write-in manner, the key frames and corresponding entry and exit information, state information, and intermediate frames and corresponding entry and exit information are read and analyzed respectively to form key frame segments and intermediate frame segments, and a comic video is generated.

[0160] In this example, each frame (including key frames and intermediate frames) in the video frame sequence is traversed, and the frames are written into a video encoding tool (such as OpenCV's VideoWriter) in time sequence to ensure that the order of the frames is consistent with the narrative rhythm; the entry and exit information (entry and exit mode, entry and exit form) and state information of each frame are extracted, the key frame sequence of the stay stage is converted into continuous segments (such as 150 frames of stay frames of #1), and the state information is superimposed (such as when breathe=true, use a sinusoidal function to drive periodic scaling), the intermediate frame sequence of the transition stage is converted into continuous segments (such as 60 frames of transition frames from #1→#2), and the entry and exit information is applied (such as when panDuration=2 seconds, use the ease_in_out interpolation function to interpolate the position and size); the key frame segments and intermediate frame segments are combined in order and encoded into an MP4 video by OpenCV. By converting the user's annotated parameters into continuous and dynamic videos, the unity of narrative rhythm and visual effect is ensured.

[0161] Further, as shown in Figure 9 The key frames and corresponding entry and exit information, state information are read and analyzed to form key frame segments, specifically including the following steps:

[0162] The key frames are obtained, the easing function and form parameters are fused, and the initial key frame segment is formed;

[0163] The stay time period of the initial key frame segment is determined;

[0164] The initial key frame segment is optimized by superimposing a breathing function in the stay time period to form a key frame segment.

[0165] To solve the problem of still pictures in the key frame stay stage and lack of dynamic, the application converts static key frames into rhythmic and dynamic segments through the combination of in-out mode (such as a slow function), shape parameters, and state information (such as a breathing function).

[0166] In this example, the slow function includes eight curves such as linear, slow-in, slow-out, elastic, and bounce, which are extracted from the in-out mode in the in-out information; the shape parameters include a rotation angle (-180°~180°) and a zoom factor (such as 1.0~2.0), which are used to adjust the picture shape of the key frame and are extracted from the in-out shape in the in-out information. First, the key frame is fused with the slow function and the shape parameters to form a dynamic key frame segment (such as the "slow-in slow-out" "zoom" + "-15°" rotation, which simulates the effect of the character gradually entering the picture); according to the stay duration, the static time of the key frame in the video is determined (the stay time period = stay duration * frame rate), which ensures that the narrative focus of the key frame (such as character dialogue and key action) has enough time to be focused on by the audience; finally, during the stay time period, a breathing function (such as the form of a sine function) is used to drive periodic scaling (±2%), which simulates the subtle dynamics of "breathing" to form a key frame segment. By converting static key frames into rhythmic and dynamic segments, the immersion and expressiveness of the video are improved. The slow function allows the dynamic effects (scaling and rotation) in the key frame to have an acceleration and deceleration process, and the shape parameters allow the picture shape of the key frame to be adjusted (such as a rotation of -90° to simulate a character falling down and a zoom of 2.0 to enlarge a close-up), which improves the three-dimensional expressiveness of the picture; the subtle periodic scaling (±2%) during the stay time period makes the static picture have a "breathing" feeling and avoids a sense of lag; dynamic key frames attract more audience attention than static key frames, thereby enhancing the content dissemination power.

[0167] The breathing function is specifically represented as:

[0168]

[0169] In the formula, scale_factor represents the breathing function, frame_idx is the sequence number of the key frame in the stay stage, γ is the amplitude, which is 0.02 in this example, and θ is the angular frequency, which is 0.1 in this example.

[0170] The breathing function ensures periodicity and smoothness, keeps the center of the picture fixed, and slightly expands and contracts the edges, which simulates the natural shaking (micro-ocular movement, blinking, and breathing) when the eyes are fixed. After comparing the videos with "no breathing effect" vs "breathing effect" by multiple test users, the immersion is greatly improved and the sense of lag is greatly reduced.

[0171] In another embodiment, the intermediate frames and corresponding in-out information are read and analyzed to form an intermediate frame segment, specifically including the following steps:

[0172] An intermediate frame is obtained, the easing function and the morph parameter are fused, and an intermediate frame segment is formed.

[0173] It can be understood that the intermediate frame is a transition dynamic frame between adjacent key frames, and is in a dynamic process (such as translation from left to right) by itself. The picture is always changing, and no additional breathing effect is needed to “activate” the picture. The duration of the intermediate frame is relatively short, and the audience will not notice the “stillness” problem, so no breathing function is needed, and smooth motion can be achieved only by easing function and morph parameter interpolation.

[0174] In another embodiment, the key frame segment and the intermediate frame segment are formed, and a comic video is generated, specifically including the following steps:

[0175] An audio file is obtained;

[0176] Based on the key frame segment and the intermediate frame segment, an initial comic video is generated;

[0177] The audio file is processed in a format of the comic video to form a to-be-synthesized audio file;

[0178] The initial comic video and the to-be-synthesized audio file are fused, and the duration of the initial comic video and the to-be-synthesized audio file is compared. The duration of the comic video is determined according to the shorter one, and the comic video is generated.

[0179] In this example, the user uploads an audio file (supports MP3, WAV, M4A, AAC, etc.) through the file selector (id="audioInput") of the Web interface, the server receives and stores the uploaded audio file; based on the key frame segment (dynamic picture in the stay stage) and the intermediate frame segment (translation / rotation animation in the transition stage), an initial comic video without audio is generated using OpenCV VideoWriter; the uploaded audio file is converted in format to adapt to the format requirements of the comic video (such as converting WAV to AAC); then FFmpeg is used to fuse the initial comic video and the to-be-synthesized audio to generate a comic video with background music, and the final duration is determined according to the shorter one of the initial comic video and the to-be-synthesized audio file to prevent black screen or mute. If the format processing fails and the to-be-synthesized audio file does not appear, the initial comic video is directly used as the final comic video. By adding background music, the infectivity and propagation of the comic video are improved, the user experience is improved, and the content demand of the short video era is met.

[0180] In another embodiment, the key frame segment and the intermediate frame segment are formed, and a comic video is generated, further comprising:

[0181] Adjacent frames are obtained, and a transition effect type is determined;

[0182] Determine the alpha value of the adjacent frame and linearly mix.

[0183] It can be understood that the alpha value is used to control the weight ratio of the current frame and the previous frame, and the value range is [0, 1]. In the present example, the fade-in effect (preset transition effect type) is applied to the first 10 frames (about 0.33 seconds) of the translation transition, and the alpha blending is realized by using the addWeighted function of OpenCV to solve the problem of "direct jump of picture when adjacent grid switching, strong visual impact".

[0184] In a second aspect, as shown in the Figure 10 The application also provides a cartoon video generation device based on interactive annotation, which adopts the cartoon video generation method based on interactive annotation as described above, and specifically comprises:

[0185] The acquisition unit is configured to acquire the cartoon image, the grid annotation request, and the interactive annotation instruction.

[0186] The annotation generation unit is configured to generate a canvas covering the cartoon image in response to the grid annotation request, wherein the canvas includes a plurality of partitions, each of which corresponds to a different grid area of the cartoon image, and each partition includes initial annotation information; and optimize the initial annotation information to form partition annotation information in response to the interactive annotation instruction.

[0187] The video generation unit is configured to give a video frame sequence in combination with the grid area of the cartoon image and the partition annotation information, and process the video frame sequence in a frame-by-frame writing manner to generate a cartoon video.

[0188] The cartoon video generation method and device based on interactive annotation provided by the application have at least the following beneficial effects:

[0189] (1) The cartoon video generation scheme of the application fully balances the flexibility and automation of cartoon grid by fusing user interactive annotation and intelligent recognition detection, and efficiently and accurately generates dynamic video from static cartoon while ensuring detection quality and batch processing.

[0190] (2) The application improves the operation efficiency of cartoon grid by designing a cartoon annotation method with a covering canvas and a drag-and-drop interaction, realizes real-time adjustment with pixel-level precision, and has high reliability without repeated trial and error. At the same time, the partition position information optimization including partition edge frame and partition center point adjustment, as well as the partition in-out information and partition state information optimization, are organically combined to adapt to different scenarios, realizing fast annotation, fine adjustment, and batch processing.

[0191] (3) The application realizes the fine control of the video frame granularity of the generated cartoon video by marking different entry and exit information in different grid areas of the cartoon and embedding different types of easing functions, and enhances the visual fluency. Meanwhile, the state information of the key frame is marked, the breathing function and the easing function are fused, the dynamic effect is presented, and the expressiveness of the key frame is improved.

[0192] While the preferred embodiments of the application have been described, additional variations and modifications can be made to the embodiments by those skilled in the art once they learn of the basic inventive concepts. Therefore, the appended claims are intended to encompass within their scope all such variations and modifications as are included within the scope of the application. It should be apparent that the application is not limited to the specific embodiments described herein, but can be practiced with modification and alteration within the scope and spirit of the present application. Accordingly, the specification is to be regarded in an illustrative, rather than a restrictive sense, as the application is open to additional suggested modifications and embodiments.

Claims

1. A method for generating comic videos based on interactive annotation, characterized in that, Specifically, the steps include the following: Get comic book images; In response to a panel annotation request, a canvas covering the comic image is generated. Specifically, this includes: receiving an automatic panel annotation instruction containing a type request; performing contour recognition on the comic image based on edge detection to determine each panel region; providing initial position information for each panel region based on its coordinate relationship and a preset scaling factor, and annotating the panel bounding boxes and center points of each panel; determining the order of each panel based on the type request, providing a panel code, and forming and annotating the panel connection lines linking the center points of each panel based on the panel code; determining the initial entry / exit information and initial state information of each panel, and generating a canvas covering the comic image. The canvas includes multiple panels, each corresponding to a different panel region of the comic image. Each panel includes initial annotation information, which includes position information, panel code, entry / exit information, and state information. The entry / exit information includes entry / exit modes, which include entry / exit methods and entry / exit forms. The state information consists of parameters that periodically change in the size of the comic video. In response to interactive annotation commands, the initial annotation information is optimized to form partition annotation information. Specifically, this includes: selecting and dragging any partition edge box, and / or selecting and dragging any partition center point to optimize the initial position information of the corresponding partition and form partition position information. Here, interactive annotation commands refer to the user's operations on the canvas partitions. Based on the panel divisions and partition annotations of the comic image, a video frame sequence is given; The video frame sequence is processed frame by frame to generate a comic video.

2. The comic video generation method based on interactive annotation as described in claim 1, characterized in that, In response to interactive annotation commands, the initial annotation information is optimized to form partition annotation information, which also includes the following steps: Select and adjust the initial entry and exit information to form the partitioned entry and exit information; and / or; Select and adjust the initial state information to form the partition state information.

3. The comic video generation method based on interactive annotation as described in claim 1, characterized in that, Based on the panel divisions and section annotations of the comic image, a video frame sequence is generated, including the following steps: Based on the grid regions of the comic image, determine the corresponding keyframes; Based on the partitioning of each corresponding grid region, each keyframe is serialized to form a keyframe sequence. Based on the entry and exit information of each adjacent keyframe, and through an interpolation algorithm, intermediate frames are generated between each adjacent keyframe, the entry and exit information of each intermediate frame is determined, and an intermediate frame sequence is formed. By fusing the keyframe sequence and the intermediate frame sequence, a video frame sequence is given.

4. The comic video generation method based on interactive annotation as described in claim 3, characterized in that, Based on the entry and exit information of each adjacent keyframe, and through an interpolation algorithm, intermediate frames are generated between each adjacent keyframe. The entry and exit information of each intermediate frame is determined to form an intermediate frame sequence. The specific steps include the following: Traverse the keyframe sequence and combine the entry and exit information of each adjacent keyframe to determine the duration of each adjacent keyframe. By combining the duration of each adjacent keyframe and the preset frame rate, the corresponding number of intermediate frames is given; By using linear interpolation and combining the entry and exit information of adjacent keyframes, the entry and exit information of each intermediate frame is determined. Generate intermediate frames to form an intermediate frame sequence.

5. The comic video generation method based on interactive annotation as described in claim 3, characterized in that, Before fusing the keyframe sequence and the intermediate frame sequence, the following steps are also included: Obtain the correspondence between each key frame determined based on the grid region and the partition, calculate the distance between the center points of each partition through the Euclidean distance, and give the positional relationship between each key frame; Determine and mark the starting key frame from the key frame sequence, traverse the remaining key frames in the key frame sequence, and mark the key frame closest to the starting key frame as the next key frame; Update the next key frame as the starting key frame, traverse the unmarked key frames in the key frame sequence, and mark the key frame closest to the starting key frame as the next key frame; Repeat the above process of marking the next key frame until all key frames in the key frame sequence are marked, forming an optimized path of the key frame sequence; Based on the optimized path of the key frame sequence, optimize the key frame sequence and optimize the intermediate frame sequence.

6. The comic video generation method based on interactive annotation as described in claim 3, characterized in that, Process the video frame sequence in a frame-by-frame writing manner to generate a comic video, specifically including the following steps: In a frame-by-frame writing manner, respectively read and analyze the key frames and corresponding entry and exit information, status information, and intermediate frames and corresponding entry and exit information, form key frame segments and intermediate frame segments, and generate a comic video.

7. The comic video generation method based on interactive annotation as described in claim 6, characterized in that, Read and analyze the key frames and corresponding entry and exit information, status information, and form key frame segments, specifically including the following steps: Obtain the key frames, fuse the easing function and the morphological parameters to form an initial key frame segment; Determine the停留时间段 of the initial key frame segment; Optimize the initial key frame segment by superimposing the breathing function during the停留时间段 to form a key frame segment.

8. A comic video generation device based on interactive annotation, characterized in that, Adopt the method for generating a comic video based on interactive annotation as described in any one of claims 1-7, specifically including: An acquisition unit for obtaining comic images, grid annotation requests, and interactive annotation instructions; A label generation unit for receiving an automatic grid annotation instruction containing a type request; performing contour recognition on the comic image based on edge detection to determine each grid region of the comic image; giving the initial position information of the partition corresponding to each grid region according to the coordinate relationship of each grid region, combining with a preset proportional conversion coefficient, and marking the partition edge frame and the partition center point of each partition; determining the sorting of each partition based on the type request, giving a partition code, and forming and marking a partition connection line connecting the center points of each partition based on the partition code; determining the initial entry and exit information and initial status information of each partition, and generating a canvas covering the comic image, where the canvas includes multiple partitions, each partition respectively corresponds to a different grid region covering the comic image, and the partition includes initial annotation information; in response to the interactive annotation instruction, optimize the initial annotation information to form partition annotation information, specifically including: selecting and dragging any partition edge frame, and / or, selecting and dragging any partition center point, and optimizing the initial position information of the corresponding partition to form partition position information, where the interactive annotation instruction refers to the user's operation on the canvas partition, and the annotation information includes position information, partition code, entry and exit information, and status information, and the entry and exit information includes an entry and exit mode, and the entry and exit mode includes an entry and exit method and an entry and exit form, and the status information is a parameter for the comic video to present a periodic change in size; It should be noted that there is an unclear expression "停留时间段" in the original text, which may need to be further clarified according to the specific context for a more accurate translation. The video generation unit is used to combine the grid areas and partition annotation information of the comic image to give a video frame sequence; and to process the video frame sequence in a frame-by-frame writing manner to generate a comic video.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and computer readable storage medium

    CN111415399A

  • Cartoon video generation method and device, electronic equipment and storage medium

    CN115811639A