Cover generation method and device

Through the combination of cutout and style guide information, the use of ControlNet and diffusion model to generate covers solves the problem of time-consuming and laborious and lack of personalization in traditional cover generation methods, and achieves efficient, personalized and visually attractive cover generation.

CN120091198APending Publication Date: 2025-06-03SHANGHAI HODE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510265111.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-06
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The traditional cover generation method is time-consuming and laborious. The generated cover lacks visual appeal and personalization, making it difficult to accurately display the first impression of the video content.

Method used

By obtaining the target image frames in the target video, the cutout process is performed to extract the image of the target object; automatically generate style guidance information based on the video content; combine the target image and style guidance information, use ControlNet and diffusion model to generate a cover with embedded target image and a target style.

Benefits of technology

Improves the automation and image quality of cover generation, enhances visual appeal and personalization characteristics, and enables the generated cover to fit the video content closely.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120091198A_ABST
    Figure CN120091198A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a cover generation method and device, and belongs to the technical field of multimedia. The cover generation method comprises the following steps: acquiring a target video, wherein the target video comprises a target image frame with a target object; performing matting on the target image frame to obtain a target image corresponding to a target object; according to the video content of the target video, style guiding information is obtained, and the style guiding information is used for guiding generation of a target style; and according to the target image and the style guide information, generating a cover embedded with the target image and having a target style. According to the technical scheme, the generation efficiency of the cover can be improved. Through the control of the style guide information, the generated cover can not only closely fit the video content, but also enhance the visual appeal and personalized characteristics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of multimedia technologies, and in particular, to a method, apparatus, computer device, computer-readable storage medium, and computer program product for generating a cover page. Background Art

[0002] Today, with the increasing prosperity of digital media and Internet content creation, video content has become an important part of information dissemination and entertainment consumption. Whether in the fields of education, advertising, movies, or short videos, as the first impression of a video, the cover directly affects the click-through rate and viewing experience of the audience.

[0003] Traditional cover generation methods often rely on manually selecting video frames or directly intercepting static images from the video. This method is not only time-consuming and laborious, but also the generated covers lack visual appeal and personalization.

[0004] It should be noted that the above content is not necessarily prior art and does not limit the scope of patent protection of the present application. Summary of the Invention

[0005] Embodiments of the present application provide a method, apparatus, computer device, computer-readable storage medium, and computer program product for generating a cover page to solve or alleviate one or more of the above-mentioned technical problems.

[0006] One aspect of the embodiments of the present application provides a method for generating a cover page, the method including: Obtaining a target video, where the target video includes a target image frame having a target object; Performing matte extraction on the target image frame to obtain a target image corresponding to the target object; Obtaining style guidance information according to the video content of the target video, where the style guidance information is used to guide the generation of a target style; Generating a cover page embedded with the target image and having the target style according to the target image and the style guidance information.

[0007] Optionally, the method for generating a cover page further includes: Extracting a plurality of key frames from the target video; Respectively detecting the plurality of key frames through a pre-trained object detection model to obtain several image frames representing the target image; Wherein, the target object is a person or object having the target image, and the target image frame is selected from the several image frames.

[0008] Optionally, performing matte extraction on the target image frame to obtain a target image corresponding to the target object includes: Separate the background and the foreground corresponding to the target object in the target image frame through the Matting image processing method; Perform edge processing on the separated foreground to obtain the target image.

[0009] Optionally, the cover generation method further includes: Adjust the brightness and / or contrast of the target image.

[0010] Optionally, according to the video content of the target video, obtain style guidance information, including: Determine the video theme of the target video according to the video content; Generate the style guidance information according to the video theme; Wherein, the style guidance information includes prompt words, and the prompt words include visual elements and / or style requirements corresponding to the video theme.

[0011] Optionally, according to the target image and the style guidance information, generate a cover embedded with the target image and having a target style, including: Input the target image into the ControlNet model to generate an intermediate feature map corresponding to the target image; Input the style guidance information and the intermediate feature map into the diffusion model to generate the cover.

[0012] Another aspect of the embodiments of the present application provides a cover generation device, and the device includes: A first acquisition module, configured to acquire a target video, where the target video includes a target image frame having a target object; A matting module, configured to perform matting on the target image frame to obtain a target image corresponding to the target object; A second acquisition module, configured to acquire style guidance information according to the video content of the target video, where the style guidance information is used to guide the generation of a target style; A generation module, configured to generate a cover embedded with the target image and having a target style according to the target image and the style guidance information.

[0013] Another aspect of the embodiments of the present application provides a computer device, including: At least one processor; and A memory communicatively connected to the at least one processor; Wherein: the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method as described above.

[0014] Another aspect of the embodiments of the present application provides a computer-readable storage medium, in which computer instructions are stored, and when the computer instructions are executed by a processor, the above-described method is implemented.

[0015] Another aspect of the embodiments of the present application provides a computer program product, including a computer program, and when the computer program is executed by a processor, the above-described method is implemented.

[0016] The embodiments of the present application adopting the above technical solutions may include the following advantages: First, by performing matte extraction on the target image frame to extract the target image corresponding to the target object (such as a teacher). Then, style guidance information for guiding the generation of the target style is automatically generated according to the content of the target video. Subsequently, combining the target image with the style guidance information, a cover that includes both the target image and has the target style is generated. Through the above automated process, the link of manual participation is eliminated, thereby improving the generation efficiency of the cover. And through the control of the style guidance information, the generated cover can not only closely fit the video content, but also enhance the visual appeal and personalized features. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings exemplarily show the embodiments and form a part of the specification, and are used together with the written description of the specification to explain the exemplary embodiments of the embodiments. The shown embodiments are only for illustrative purposes and do not limit the scope of the claims. In all the drawings, the same reference numerals refer to similar but not necessarily identical elements.

[0018] Figure 1 Schematically shows an operating environment diagram of the cover generation method according to Embodiment 1 of the present application; Figure 2 Schematically shows a flowchart of the cover generation method according to Embodiment 1 of the present application; Figure 3 Schematically shows Figure 2 a sub-step flowchart of step S202 in Figure 4 Schematically shows Figure 2 a sub-step flowchart of step S204 in Figure 5 Schematically shows Figure 2 a sub-step flowchart of step S206 in Figure 6 Schematically shows an additional flowchart of the cover generation method according to Embodiment 1 of the present application; Figure 7 Schematically shows an exemplary application flowchart; Figure 8 Schematically shows a block diagram of the cover generation device according to Embodiment 2 of the present application; and Figure 9 Schematically shows a schematic diagram of the hardware architecture of a computer device according to Embodiment 3 of the present application. Detailed implementation manners

[0019] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts fall within the scope of protection of the present application.

[0020] It should be noted that the descriptions involving "first", "second", etc. in the embodiments of the present application are only for descriptive purposes and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments may be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement them. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present application.

[0021] In the description of the present application, it should be understood that the numerical labels before the steps do not identify the order of execution of the steps, but are only used to facilitate the description of the present application and distinguish each step, and thus cannot be understood as a limitation to the present application.

[0022] First, provide the term explanations involved in the present application: Diffusion model: A generative artificial intelligence model mainly used to generate high-quality data (such as images, audio or text). By simulating the physical diffusion process, it gradually converts random noise into the target data. By processing the input text or image, it generates images that conform to a specific style or theme, and is widely used in fields such as art creation, image generation and visual design.

[0023] ControlNet: A deep learning technology that enhances the image generation control ability. It can use the input additional information (such as key points, edge maps, depth maps, etc.) as constraints to guide the output results of the generation model, making the generated images more in line with user expectations.

[0024] Matting: An image processing technology mainly used to extract target objects (such as people or objects) from complex backgrounds. By precisely separating the foreground and background, it ensures that the edges of the extracted objects have natural transitions and complete details, and is widely used in fields such as matting, video special effects and augmented reality.

[0025] Secondly, to facilitate the understanding of the technical solutions provided in the embodiments of the present application by those skilled in the art, the related technologies will be described below: The present inventor has learned that manually selecting video frames or directly intercepting static images from videos is not only time-consuming and laborious, but also the generated covers lack visual appeal and personalization, making it difficult to accurately display the teacher's image and course theme, thus affecting the display effect of online education content.

[0026] Therefore, the embodiments of the present application provide a cover generation technical solution. In this technical solution, the Matting technology, diffusion model technology, and ControlNet technology are combined to optimize cover generation. According to the Matting technology, the outline of an object (such as a teacher) is accurately extracted from a complex background and separated from other content (i.e., the background) in the video. Subsequently, the extracted object image is stylized using the diffusion model technology and ControlNet technology. The cover image that conforms to the video theme is controlled and generated through prompts. This combination method not only improves the automation degree of cover generation, but also enhances the image quality and visual effect. See the following text for details.

[0027] Finally, for the convenience of understanding, an exemplary operating environment is provided below.

[0028] As Figure 1 shown, the operating environment diagram includes: a service platform 2, clients (4A, 4B,..., 4N).

[0029] The service platform 2 can be connected to the clients (4A, 4B,..., 4N) through a network.

[0030] The service platform 2 can be a single server, a server cluster, or a cloud computing service center.

[0031] The service platform 2 can provide services such as reading services and uploading services to the clients.

[0032] The service platform 2 can be located in a data center such as a single location, or distributed in different geographical locations (for example, in multiple locations). The service platform 2 can provide services via a network. The network includes various network devices such as routers, switches, multiplexers, hubs, modems, bridges, repeaters, firewalls, proxy devices, and / or the like. The network can include physical links such as coaxial cable links, twisted pair cable links, fiber optic links, and their combinations, or wireless links such as cellular links, satellite links, Wi-Fi links, etc.

[0033] Clients (4A, 4B, …, 4N) can be configured to access the content and services of service platform 2. The clients (4A, 4B, …, 4N) can include electronic devices with or external to a display panel, such as mobile devices, tablet devices, laptop computers, workstations, virtual reality devices, gaming devices, digital streaming devices, vehicle terminals, smart TVs, set-top boxes, etc., and can also include virtualized computing instances. The virtualized computing instances can include virtual machines, such as emulations of computer systems, operating systems, servers, etc. The computing device can load a virtual machine based on a virtual image and / or other data defining specific software (e.g., operating system, dedicated application, server) for emulation. As the demand for different types of processing services changes, different virtual machines can be loaded and / or terminated on one or more computing devices.

[0034] Clients (4A, 4B, …, 4N) can be associated with one or more users. A single user can also use one or more of the clients (4A, 4B, …, 4N) to access service platform 2. The clients (4A, 4B, …, 4N) can travel to various locations and use different networks to access service platform 2.

[0035] Clients (4A, 4B, …, 4N) can include an interface. The interface can include a touchpad, a touch screen, a mouse, a keyboard, or other sensing elements. For example, the input element can be configured to receive user instructions, and the user instructions can cause the clients (4A, 4B, …, 4N) to perform various operations, such as uploading a video, selecting a cover, etc.

[0036] It should be noted that the above devices are exemplary, and the number and types of devices can be adjusted in different scenarios or according to different requirements.

[0037] The following takes the client as the execution subject and introduces the technical solutions of this application through multiple embodiments. It should be noted that these embodiments can be implemented in various different forms and should not be construed as being limited only to the embodiments described herein.

[0038] Embodiment 1 Figure 2 A flowchart of a cover generation method according to Embodiment 1 of the present application is schematically shown.

[0039] As Figure 2 shown, the cover generation method can include steps S200 to S206, where: Step S200, obtain a target video, where the target video includes a target image frame with a target object; Step S202, perform matte extraction on the target image frame to obtain a target image corresponding to the target object; Step S204: Obtain style guidance information according to the video content of the target video, where the style guidance information is used to guide the generation of the target style. Step S206: Generate a cover that embeds the target image and has the target style according to the target image and the style guidance information.

[0040] For the cover generation method provided in this embodiment, first, perform matte extraction on the target image frame to extract the target image corresponding to the target object (such as a teacher). Then, automatically generate style guidance information for guiding the generation of the target style according to the content of the target video. Subsequently, combine the target image with the style guidance information to generate a cover that not only includes the target image but also has the target style. Through the above automated process, the manual participation link is omitted, thereby improving the cover generation efficiency. And through the control of the style guidance information, the generated cover can not only closely fit the video content, but also enhance the visual attraction and personalized features.

[0041] The following will Figure 2 elaborate in detail on each step in steps S200 to S206 and other optional steps.

[0042] Step S200 Obtain a target video, where the target video includes a target image frame with a target object.

[0043] The target video is video content that includes target objects (such as people, animals, landmarks, specific items, etc.), such as people interviews, animal documentaries, product demonstrations, movie trailers, etc. The target video can be in various formats such as MP4, AVI, MOV, WMV, FLV, etc.

[0044] The target image frame refers to a video frame selected from the target video that includes the target object and has a relatively simple background and / or high contrast.

[0045] Step S202 Perform matte extraction on the target image frame to obtain the target image corresponding to the target object.

[0046] The target image refers to the target object image that is separated from the target image frame after matte extraction and does not include the background or only includes a simple background (for example, a teacher image).

[0047] The target image frame can be subjected to matte extraction in various ways. For example: (1) Based on a pre-trained deep learning model (such as Mask R-CNN, DeepLab, etc.): The target image frame can be input into the pre-trained deep learning model to identify and locate the target object in the target image frame. Then, a mask that conforms to the shape of the target object is generated. This mask is applied to the original image to obtain the target image corresponding to the target object. Herein, the mask refers to a binary image or layer used to identify the boundary between the target object and the background in the image.

[0048] (2) Based on image segmentation methods such as color space, texture, or edge detection: Perform color space conversion or texture feature extraction on the target image frame to analyze the color or texture information in the image. Use an edge detection algorithm to identify the edge information in the image. Based on this information, an image segmentation algorithm is used to separate the target object from the background. Finally, by generating a binary mask (the target object is white and the background is black) and applying it to the original image, the target image corresponding to the target object is obtained.

[0049] (3) Based on the affinity method: Construct a foreground and background model of the image by calculating the affinity (or similarity) between pixels. Then, input the target image frame into this model for matting to obtain the target image corresponding to the target object. Herein, the affinity can be calculated based on various features such as color, texture, and shape.

[0050] In addition to the above methods of matting the target image frame based on deep learning, image segmentation, and affinity, etc., the target image frame can also be matted through the Matting image processing method: In an alternative embodiment, as Figure 3 shown, step S202 may include: Step S300, separating the background and the foreground corresponding to the target object in the target image frame through the Matting image processing method; Step S302, performing edge processing on the separated foreground to obtain the target image.

[0051] In some embodiments, before applying the Matting algorithm, the target image frame can be preprocessed, such as denoising, enhancing contrast, etc., to improve the performance and accuracy of the algorithm.

[0052] Matting algorithms can include GrabCut, Poisson Matting, KNN Matting, etc. The appropriate Matting algorithm can be selected according to the characteristics of the target image frame (such as color contrast, texture complexity, etc.) for separating the background and foreground. For example, if the target object has a distinct color contrast with the background, the GrabCut algorithm can be adopted. Based on graph cut theory, it can quickly and accurately separate the foreground and background. In some embodiments, the parameters of the Matting algorithm, such as the number of iterations, color space selection (RGB, Lab, etc.), tolerance range, etc., can be adjusted according to the characteristics of the target image frame to obtain the best separation effect.

[0053] A foreground mask (Trimap) can be obtained. The Trimap can be generated by manual annotation or using an automatic algorithm. The target image frame and the obtained Trimap can be input into the selected Matting algorithm to calculate the Alpha value of each pixel and generate an alpha map. Then, the foreground and background in the target image frame can be separated through this alpha map. The foreground image can be extracted by multiplying the alpha map with the target image frame and setting a threshold.

[0054] The edges of the foreground image can be identified through an edge detection algorithm (such as the Canny edge detector). The edge detection algorithm will generate a binary image. Among them, the edge pixels are marked as white, and other pixels are marked as black, thus obtaining the target image. In some embodiments, post-processing can be performed on the target image, such as denoising, smoothing the edges, etc., to improve the quality of the target image.

[0055] In this embodiment, matting is performed on the selected key frame (i.e., the target image frame) through Matting technology to accurately separate the target image (for example, the contour image of the teacher), so as to ensure the subjectivity and consistency of the target object (for example, the teacher image) in the subsequent generated cover as much as possible.

[0056] Step S204 , according to the video content of the target video, obtain style guidance information, where the style guidance information is used to guide the generation of the target style.

[0057] The style guidance information can refer to a set of elements including color, composition, emotion, and theme tags, etc., which is used to guide the style of the generated cover. The style guidance information can be in various formats such as text, pictures, videos, etc.

[0058] The style guidance information can be generated in various ways. For example: (1) Color analysis: The main color, color matching, and color transition mode in the target video can be extracted to generate the color elements of the style guidance information.

[0059] (2) Composition analysis: It is possible to analyze the composition elements of the target video, such as the rule of thirds, symmetry, the relationship between the foreground and the background, etc., to generate the composition elements of the style guidance information.

[0060] (3) Emotion and theme recognition: Utilize natural language processing (NLP) and computer vision technologies to identify the emotional atmosphere (such as joy, sadness, tension, etc.) and themes (such as travel, food, sports, etc.) in the video, and generate the emotion and theme labels of the style guidance information.

[0061] In some embodiments, the historical behavior data of the user and the current popular trends can be combined to generate a cover style that conforms to the user's preferences and the sense of the times.

[0062] In an alternative embodiment, as Figure 4 shown, step S204 may include: Step S400, determining the video theme of the target video according to the video content; Step S402, generating the style guidance information according to the video theme; wherein, the style guidance information includes prompt words, and the prompt words include visual elements and / or style requirements corresponding to the video theme.

[0063] The video content includes visual content, audio content, text content, etc.; the video content can be analyzed in various ways: (1) Visual content analysis: It is possible to analyze the image content in the target video through computer vision technologies, such as object detection, scene recognition, etc. For example, identify elements such as objects (such as teachers), backgrounds, colors, and lights in the target video.

[0064] (2) Audio content analysis: It is possible to use speech recognition and natural language processing technologies to analyze keywords, emotions, etc. in the conversations, voiceovers, or background music in the target video.

[0065] (3) Text content analysis: It is possible to analyze the subtitles, titles, or descriptive texts in the target video.

[0066] It is possible to match the above video content analysis results with predefined themes to further determine the video theme. In some embodiments, it is also possible to dynamically generate a video theme according to the above video content analysis results. The video theme can be course categories (such as science, art, or language), natural scenery, urban scenery, human history, cutting-edge technology, etc.

[0067] Visual elements refer to elements such as images, colors, shapes, compositions, etc. related to the video theme. Style requirements refer to specific requirements for aspects such as the overall style, color matching, animation effects, font styles, etc. of the generated cover. The prompt can include keywords, phrases, or descriptive statements. For example, if the video content is a math course video explaining the problem-solving skills of quadratic functions, covering the image characteristics, vertex coordinates, axis of symmetry of quadratic functions, and their applications in practical problems. Video theme: Problem-solving skills of quadratic functions. Visual elements are: quadratic function images, text information (such as the title "Problem-solving skills of quadratic functions"), color matching (such as the main color being blue or green, reflecting the calmness and rationality of the math course). Style requirements are: "Simple and clear, focusing on the clarity and readability of the content", "Select professional and easy-to-read fonts, such as Arial or Helvetica, to ensure the readability of the title and key information". The generated prompts are: "Math course cover with the problem-solving skills of quadratic functions as the core.", "Blue or green tone, highlighting key information, simple and clear.", "Professional font, clear layout, and animated display of function changes." In this embodiment, by determining the video theme according to the video content and generating a prompt including visual elements and style requirements as style guidance information, the production efficiency and quality of the cover are effectively improved, thereby enhancing the attractiveness of the target video.

[0068] Step S206 , generate a cover embedded with the target image and having a target style according to the target image and the style guidance information.

[0069] Covers can be generated in various ways. For example: (1) The target image can be embedded into a preset background template, and then, according to the style guidance information, the background color, texture, and composition elements are adjusted to generate a cover embedded with the target image and having a target style.

[0070] (2) The style of the target image can be converted to a style matching the style guidance information through a style transfer algorithm (such as AdaIN, CycleGAN, etc.) to generate a cover embedded with the target image and having a target style.

[0071] In addition to the above methods, covers can also be automatically generated through a deep learning model: In an alternative embodiment, as Figure 5 shown, step S206 may include: Step S500, input the target image into the ControlNet model to generate an intermediate feature map corresponding to the target image; Step S502, input the style guidance information and the intermediate feature map into the diffusion model to generate the cover.

[0072] The ControlNet model is used to extract features from the input target image, generating an intermediate feature map including key information. Intermediate features refer to the feature representations that are gradually extracted and encoded during the processing of the target image through each layer of the ControlNet model and can represent the key information of the image. These features include the edges, textures, shapes, objects of the image, and the relationships between them, etc.

[0073] The diffusion model can be a model based on the U-Net structure, a model based on the Transformer structure, etc.

[0074] In this embodiment, by combining the ControlNet model and the diffusion model, a cover that conforms to the video theme is generated, so that the generated cover not only has high resolution and detail retention, but also improves the overall quality of the cover, thereby enhancing the attractiveness of the video content (such as, online course content).

[0075] In an alternative embodiment, as Figure 6 shown, the cover generation method may further include: Step S600, extracting multiple key frames from the target video; Step S602, respectively detecting the multiple key frames through a pre-trained object detection model to obtain several image frames representing the target image; Wherein, the target object is a person or an object having the target image, and the target image frame is selected from the several image frames.

[0076] Multiple algorithms such as shot boundary detection, image content change detection, or video summarization technology can be used to extract key frames from the video. In some embodiments, the number of key frames to be extracted can be determined according to the length of the video, the complexity of the content, and the accuracy of the required cover. In some embodiments, the features of the several image frames representing the target image can be high clarity, the target object is complete, the target object (such as, a person) has a natural expression, etc.

[0077] The object detection model can be various models such as YOLO (You Only Look Once), Faster R-CNN, or SSD (SingleShot MultiBox Detector).

[0078] In this embodiment, by extracting key frames from the target video and using a pre-trained object detection model to identify the image frames representing the target image, the efficiency of screening out the image frames suitable as the cover is improved.

[0079] In an alternative embodiment, after obtaining the target image, the method further includes: adjusting the brightness and / or contrast of the target image.

[0080] For the method of adjusting the brightness of the target image: (1) Linear transformation method: The brightness of each color channel of the target image can be multiplied by a brightness adjustment coefficient (a positive number greater than 1 or less than 1) to adjust (increase or decrease) the brightness of the target image.

[0081] (2) Histogram translation method: The brightness of the target image can be adjusted by shifting the entire histogram of the target image to the right or left by a certain number of gray levels.

[0082] (3) Nonlinear transformation method: A nonlinear function (such as a logarithmic function, an exponential function, etc.) can be used to transform the pixel values of the image to adjust the brightness of the target image.

[0083] For the method of adjusting the contrast of the target image: (1) Linear contrast stretching: The contrast can be adjusted by stretching the gray value range of the image.

[0084] (2) Histogram equalization: The contrast is enhanced by redistributing the gray values of the target image so that the gray histogram is evenly distributed at each gray level to adjust the contrast of the target image.

[0085] (3) Adaptive contrast enhancement: The contrast of the image can be adjusted based on edge detection-based contrast enhancement or content-based contrast optimization of the image.

[0086] In this embodiment, by adjusting the brightness and contrast of the target image, the quality of the target image is improved.

[0087] To make the present application easier to understand, the following is combined with Figure 7 An exemplary application is provided. In this exemplary application, it aims to accurately extract the teacher's image from a video classroom automatically and generate a cover image that matches the course theme through stylization processing, so as to more efficiently display the teacher's image and course content and enhance the learning experience and attraction.

[0088] Step S11, perform frame extraction on the video (for example, a classroom video), sample the frames at a frame rate of 1 frame per second to obtain a plurality of key frames.

[0089] Step S12, perform human detection on the plurality of key frames through a task detection model, identify and filter out the target image frames with the target object (for example, the teacher).

[0090] Step S13A, detect the teacher in the target image frame by using the Matting technology, perform a matting process, separate the teacher's outline from the background, and optimize the edge details of the separated teacher to obtain a target image corresponding to the teacher.

[0091] Step S13B, based on the course theme of the video (such as mathematics, Chinese, etc.), construct prompt words (Prompt) that are highly relevant to the classroom video content.

[0092] In step S14, the target image and the prompt are input into the ControlNet and diffusion model to generate a cover that matches the course theme.

[0093] Matting technology can quickly and accurately extract the teacher's outline from the complex background and separate it from other content in the video. It can effectively remove the background and retain only the main body of the teacher's image, while ensuring the clarity and integrity of the teacher's image.

[0094] Through ControlNet and diffusion model technology, the extracted teacher image is stylized, and the cover that matches the course theme is controlled and generated through prompt words. The powerful ability of diffusion model technology in image generation makes the cover image not only have high resolution and detail retention, but also can give the cover a specific artistic style and theme atmosphere, thereby improving the overall quality of the cover.

[0095] By combining Matting, ControlNet and diffusion model technology to optimize cover generation, the technical defects of existing cover generation methods, such as low efficiency, poor personalization and poor visual effects, are solved, and an efficient, personalized and high-quality cover generation solution is provided, which brings new possibilities and user experience to the display of online educational content. Compared with traditional static screenshots or manual selection methods, this combined method not only improves the automation of cover generation, but also improves image quality and visual effects, and provides a more professional and personalized display method for online educational content.

[0096] Embodiment 2 Figure 8 The block diagram of the cover generation device according to the second embodiment of the present application is schematically shown. The device can be divided into one or more program modules, one or more program modules are stored in a storage medium, and are executed by one or more processors to complete the embodiment of the present application. The program module referred to in the embodiment of the present application refers to a series of computer program instruction segments that can complete specific functions. The following description will specifically introduce the functions of each program module in this embodiment. Figure 8 As shown, the device 1000 may include: a first acquisition module 1100, a cutout module 1200, a second acquisition module 1300, and a generation module 1400, wherein: The first acquisition module 1100 is configured to acquire a target video, where the target video includes a target image frame having a target object; The matting module 1200 is configured to perform matting on the target image frame to obtain a target image corresponding to the target object; The second acquisition module 1300 is configured to acquire style guidance information according to the video content of the target video, where the style guidance information is used to guide the generation of a target style; The generation module 1400 is configured to generate a cover that embeds the target image and has the target style according to the target image and the style guidance information.

[0097] In an optional embodiment, the cover generation device further includes a detection module, configured to: Extract a plurality of key frames from the target video; Detect the plurality of key frames respectively through a pre-trained object detection model to obtain several image frames representing the target image; Wherein, the target object is a person or an object having the target image, and the target image frame is selected from the several image frames.

[0098] In an optional embodiment, the matting module 1200 is further configured to: Separate the background and the foreground corresponding to the target object in the target image frame through the Matting image processing method; Perform edge processing on the separated foreground to obtain the target image.

[0099] In an optional embodiment, the cover generation device further includes an adjustment module, configured to: Adjust the brightness and / or contrast of the target image.

[0100] In an optional embodiment, the second acquisition module 1300 is further configured to: Determine the video theme of the target video according to the video content; Generate the style guidance information according to the video theme; Wherein, the style guidance information includes a prompt, and the prompt includes visual elements and / or style requirements corresponding to the video theme.

[0101] In an optional embodiment, the generation module 1400 is further configured to: Input the target image into the ControlNet model to generate an intermediate feature map corresponding to the target image; Input the style guidance information and the intermediate feature map into the diffusion model to generate the cover.

[0102] Embodiment 3 Figure 9 FIG. schematically shows a hardware architecture diagram of a computer device 10000 suitable for implementing the cover generation method according to Embodiment 3 of the present application. In some embodiments, the computer device 10000 may be a terminal device such as a smart phone, a wearable device, a tablet computer, a personal computer, a vehicle-mounted terminal, a game console, a virtual device, a workbench, a set-top box, a robot, etc. In other embodiments, the computer device 10000 may be a rack server, a blade server, a tower server, or a cabinet server (including an independent server or a server cluster composed of multiple servers), etc. As Figure 9 shown, the computer device 10000 includes, but is not limited to: a memory 10010, a processor 10020, and a network interface 10030 that can communicate with each other through a system bus. Among them: The memory 10010 includes at least one type of computer-readable storage medium. The readable storage medium includes flash memory, a hard disk, a multimedia card, a card-type memory (such as an SD or DX memory), a random access memory (RAM), a static random access memory (SRAM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a programmable read-only memory (PROM), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 10010 may be an internal storage module of the computer device 10000, such as the hard disk or memory of the computer device 10000. In other embodiments, the memory 10010 may also be an external storage device of the computer device 10000, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. equipped on the computer device 10000. Of course, the memory 10010 may also include both the internal storage module and the external storage device of the computer device 10000. In this embodiment, the memory 10010 is generally used to store the operating system and various application software installed on the computer device 10000, such as the program code of the cover generation method. In addition, the memory 10010 may also be used to temporarily store various data that have been output or will be output.

[0103] In some embodiments, the processor 10020 may be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other chips. The processor 10020 is generally used to control the overall operation of the computer device 10000, such as performing control and processing related to data interaction or communication with the computer device 10000. In this embodiment, the processor 10020 is used to run the program code stored in the memory 10010 or process data.

[0104] The network interface 10030 may include a wireless network interface or a wired network interface. The network interface 10030 is generally used to establish a communication link between the computer device 10000 and other computer devices. For example, the network interface 10030 is used to connect the computer device 10000 to an external terminal through a network, and establish a data transmission channel and a communication link between the computer device 10000 and the external terminal. The network may be an enterprise internal network (Intranet), the Internet, the Global System of Mobile communication (GSM for short), Wideband Code Division Multiple Access (WCDMA for short), 4G network, 5G network, Bluetooth, Wi-Fi and other wireless or wired networks.

[0105] It should be noted that Figure 9 Only the computer device with components 10010 - 10030 is shown, but it should be understood that it is not required to implement all the shown components, and more or fewer components may be alternatively implemented.

[0106] In this embodiment, the cover generation method stored in the memory 10010 may also be divided into one or more program modules and executed by one or more processors (such as the processor 10020) to complete the embodiments of the present application.

[0107] Embodiment 4 The embodiments of the present application further provide a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the cover generation method in the embodiments are implemented.

[0108] In this embodiment, the computer-readable storage medium includes flash memory, hard disks, multimedia cards, card-type memories (such as SD or DX memories, etc.), random access memories (RAM), static random access memories (SRAM), read-only memories (ROM), electrically erasable programmable read-only memories (EEPROM), programmable read-only memories (PROM), magnetic memories, magnetic disks, optical disks, etc. In some embodiments, the computer-readable storage medium may be an internal storage unit of a computer device, such as the hard disk or memory of the computer device. In other embodiments, the computer-readable storage medium may also be an external storage device of the computer device, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., equipped on the computer device. Of course, the computer-readable storage medium may also include both the internal storage unit and the external storage device of the computer device. In this embodiment, the computer-readable storage medium is generally used to store the operating system installed on the computer device and various application software, such as the program code of the cover generation method in the embodiment. In addition, the computer-readable storage medium may also be used to temporarily store various data that have been output or will be output.

[0109] Embodiment 5 The embodiment of the present application also provides a computer program product, including a computer program, which implements the method in the above embodiment when executed by a processor.

[0110] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the embodiments of the present application can be implemented by a general computer device. They can be concentrated on a single computer device or distributed on a network composed of multiple computer devices. Optionally, they can be implemented by program codes executable by the computer device. Thus, they can be stored in a storage device and executed by the computer device. And in some cases, the steps shown or described can be executed in a different order from here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module to be implemented. In this way, the embodiments of the present application are not limited to any specific combination of hardware and software.

[0111] It should be noted that the above are only the preferred embodiments of the present application, and do not limit the patent protection scope of the present application. Any equivalent structure or equivalent process transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.

Claims

1. A cover generation method, characterized in that: The method comprises: Acquire a target video, wherein the target video includes a target image frame having a target object; Cutting out the target image frame to obtain a target image corresponding to the target object; Acquire style guidance information according to the video content of the target video, wherein the style guidance information is used to guide the generation of a target style; A cover having a target style and in which the target image is embedded is generated according to the target image and the style guiding information.

2. The method according to claim 1, characterized in that The method further comprises: Extracting multiple key frames from the target video; The plurality of key frames are respectively detected by a pre-trained object detection model to obtain a plurality of image frames representing the target image; The target object is a person or object having the target image, and the target image frame is selected from the plurality of image frames.

3. The method according to claim 1, characterized in that Cutting out the target image frame to obtain a target image corresponding to the target object includes: Separating the background in the target image frame from the foreground corresponding to the target object by using a Matting image processing method; The separated foreground is subjected to edge processing to obtain the target image.

4. The method according to claim 1, characterized in that The method further comprises: The brightness and / or contrast of the target image is adjusted.

5. The method according to claim 1, characterized in that According to the video content of the target video, style guidance information is obtained, including: Determining a video theme of the target video according to the video content; Generating the style guidance information according to the video theme; The style guidance information includes prompt words, and the prompt words include visual elements and / or style requirements corresponding to the video theme.

6. The method according to any one of claims 1 to 5, characterized in that: Generating a cover having a target style and embedded with the target image according to the target image and the style guiding information, comprising: Inputting the target image into the ControlNet model to generate an intermediate feature map corresponding to the target image; The style guidance information and the intermediate feature map are input into a diffusion model to generate the cover.

7. A cover generation device, characterized in that: The device comprises: A first acquisition module is used to acquire a target video, wherein the target video includes a target image frame having a target object; A cutout module, used for cutting out the target image frame to obtain a target image corresponding to the target object; A second acquisition module is used to acquire style guidance information according to the video content of the target video, wherein the style guidance information is used to guide the generation of a target style; A generating module is used to generate a cover with the target image embedded therein and having a target style according to the target image and the style guiding information.

8. A computer device, characterized in that: include: at least one processor; and a memory communicatively connected to the at least one processor; wherein: The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and when the computer instructions are executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to claims 1 to 6 are implemented.