Video card graph generation method and electronic equipment

By obtaining user feedback information of similar videos, common video card pictures are generated, and artificial intelligence editing technology is used to solve the problem of insufficient attractiveness in video publishers' choice of card pictures by themselves, and the video click-through rate and distribution effect are improved.

CN120416624APending Publication Date: 2025-08-01HUAWEI TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202410135133.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-30
Publication Date
2025-08-01

AI Technical Summary

Technical Problem

In the prior art, video card images selected by the video publisher cannot effectively attract video viewers to click, resulting in poor video distribution effects.

Method used

By obtaining multiple videos similar to the current video content, based on user feedback information such as click-through rate and likes of these videos, they filter out similar videos and generate common video card pictures. Combined with artificial intelligence and image processing technology, they edit and generate multiple alternative video card pictures to ensure their relevance and attractiveness with the current video.

Benefits of technology

It improves the attractiveness of video card pictures, increases the click rate of videos, and improves the video distribution effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120416624A_ABST
    Figure CN120416624A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video card graph generation method and electronic equipment. The method is applied to the electronic equipment and comprises the steps that M videos similar to a first video in video content are obtained, and M is an integer larger than or equal to 1; according to the user feedback information of the M videos, K videos are screened from the M videos, K is an integer larger than or equal to 1 and smaller than or equal to M, and the user feedback information comprises the video like number and / or the video click rate; obtaining video card pictures of the K videos; and generating a plurality of alternative video card pictures according to the generality of the video card pictures of the K videos. According to the method provided by the embodiment of the invention, the video card picture of the current video can be generated according to the user feedback information of the similar video by referring to the similar video of the current video, so that the video card picture better meets the application scene requirement of the current video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly relates to a method for generating a video card image and an electronic device. Background Art

[0002] When a video publisher uploads a video for distribution, a beautiful picture (hereinafter collectively referred to as a card image) is generally attached to the video to attract video viewers to click on the video.

[0003] Generally, the video card image is uploaded by the video publisher himself / herself. The video card image can be a picture related to the video selected by the video publisher, or a picture of a certain frame extracted from the video. However, the video publisher often selects the video card image only based on his / her own understanding of the video content, and the video card image selected by the video publisher cannot effectively attract video viewers to click, resulting in poor video distribution effects.

[0004] Therefore, a method for generating a video card image is needed to obtain a video card image that can effectively attract video viewers to click. Summary of the Invention

[0005] In view of the problem of how to obtain a video card image that can effectively attract video viewers to click, this application provides a method for generating a video card image and an electronic device, and this application also provides a computer-readable storage medium.

[0006] The embodiments of this application adopt the following technical solutions:

[0007] In a first aspect, this application provides a method for generating a video card image. The method is applied to an electronic device and includes:

[0008] Obtain M videos whose video content is similar to that of a first video, where M is an integer greater than or equal to 1;

[0009] According to the user feedback information of the M videos, screen K videos from the M videos, where K is an integer greater than or equal to 1 and less than or equal to M, and the user feedback information includes the number of video likes and / or the video click-through rate;

[0010] Obtain the video card images of the K videos;

[0011] Generate multiple alternative video card images according to the commonalities of the video card images of the K videos.

[0012] According to the method of the first aspect, similar videos of the current video can be referred to, and the video card image of the current video can be generated according to the user feedback information of the similar videos, so that the video card image better meets the application scenario requirements of the current video.

[0013] In one implementation of the first aspect, screening K videos from the M videos according to the user feedback information of the M videos includes:

[0014] Screening the K videos with the highest video click-through rates from the M videos.

[0015] In one implementation of the first aspect, generating multiple alternative video card images according to the commonalities of the video card images of the K videos includes:

[0016] Generating a first image template according to the commonalities of the video card images of the K videos;

[0017] Based on the first image template, generating multiple first alternative video card images that match the first image template.

[0018] According to the method of the above implementation, generating an image template based on the commonalities of the video card images of videos with high video click-through rates, and generating alternative video card images according to the image template, so that the alternative video card images are more likely to attract video viewers to click.

[0019] In one implementation of the first aspect, generating multiple first alternative video card images that match the first image template based on the first image template includes:

[0020] Obtaining similar images of the video card images of the K videos;

[0021] Editing the similar images of the video card images of the K videos according to the first image template to generate the first alternative video card images.

[0022] According to the method of the above implementation, using the existing images in the picture library as the source of images, screening and editing the images in the source of images according to the video card images of videos with high video click-through rates to generate alternative video card images, which can effectively limit the source range of the alternative video card images.

[0023] In one implementation of the first aspect, generating multiple first alternative video card images that match the first image template based on the first image template further includes:

[0024] Extracting a first video frame from the first video;

[0025] Editing the first video frame according to the first image template to generate the first alternative video card images.

[0026] According to the method of the above implementation manner, using the video frames of the current video as the source images, screening and editing the images in the source images according to the video card images of the videos with high video click-through rates to generate alternative video card images can ensure the relevance of the alternative video card images to the current video.

[0027] In one implementation manner of the first aspect, the extracting the first video frame from the first video includes:

[0028] Determining the user attention time intervals of the M videos or the K videos according to the user interaction behavior records of the M videos or the K videos;

[0029] Determining the user attention time interval of the first video according to the user attention time intervals of the M videos or the K videos;

[0030] Extracting the first video frame from the video segment of the first video corresponding to the user attention time interval of the first video.

[0031] According to the method of the above implementation manner, extracting video frames according to the user attention time interval can effectively ensure that the extracted video frames are more attractive to video viewers to click.

[0032] In one implementation manner of the first aspect, the determining the user attention time interval of the first video according to the user attention time intervals of the M videos or the K videos includes:

[0033] Retrieving similar video segments in the first video according to the video segments corresponding to the user attention time intervals of the M videos or the K videos, and the time interval corresponding to the similar video segment is the user attention time interval of the first video.

[0034] In one implementation manner of the first aspect, the extracting the first video frame from the first video further includes:

[0035] Determining the user attention time interval of the first video according to the user interaction behavior record of the first video and the user interaction behavior records of the videos, where the user interaction behavior records of the videos include the user interaction behavior records of the M videos or the K videos;

[0036] Extracting the first video frame from the video segment of the first video corresponding to the user attention time interval of the first video.

[0037] In one implementation manner of the first aspect, the generating the first alternative video card image further includes:

[0038] Determine the user's concerned spatial region of the M videos or the K videos according to the user interaction behavior records of the M videos or the K videos;

[0039] Determine the user's concerned spatial region of the first video according to the user's concerned spatial region of the M videos or the K videos;

[0040] Crop the pictures in the second alternative picture set according to the user's concerned spatial region of the first video.

[0041] According to the method of the above implementation manner, cropping the video frames according to the user's concerned spatial region can effectively ensure that the alternative video card pictures are more attractive to video viewers to click.

[0042] In an implementation manner of the first aspect, generating the first alternative video card picture further includes:

[0043] Determine the user's concerned spatial region of the first video according to the user interaction behavior record of the first video and the user interaction behavior records of the videos, where the user interaction behavior records of the videos include the user interaction behavior records of the M videos or the K videos;

[0044] Crop the pictures in the second alternative picture set according to the user's concerned spatial region of the first video.

[0045] In an implementation manner of the first aspect, generating multiple first alternative video card pictures that match the first image template includes:

[0046] Use an artificial intelligence automatic creation and generation content model to generate the first alternative video card picture based on the first image template and the video text description of the first video.

[0047] In an implementation manner of the first aspect, the video text description of the first video includes the alternative video title of the first video and / or the original video title of the first video.

[0048] In an implementation manner of the first aspect, the method further includes:

[0049] Obtain the original video title of the first video;

[0050] Rewrite the original video title of the first video based on the video titles of the M videos or the K videos to generate multiple alternative video titles.

[0051] In an implementation of the first aspect, generating multiple first alternative video card images that match the first image template further includes:

[0052] Using an artificial intelligence automatic creation generation content model, based on the semantic segmentation map and the video text description of the first video, to generate the first alternative video card images, where:

[0053] The semantic segmentation map is a semantic segmentation map obtained by performing semantic segmentation on the images of the video card images of the K videos, and / or the set of first alternative video card images, and / or the set of second alternative video card images;

[0054] The set of first alternative video card images is a set of images generated by editing similar images of the video card images of the K videos according to the first image template;

[0055] The set of second alternative video card images is a set of images generated by editing the video frames extracted from the first video according to the first image template.

[0056] In an implementation of the first aspect, screening K videos from the M videos according to the user feedback information of the M videos includes:

[0057] Screening the K videos with the highest number of video likes from the M videos.

[0058] In an implementation of the first aspect, generating multiple alternative video card images according to the commonalities of the video card images of the K videos includes:

[0059] Determining the video card images selected by the video editor among the video card images of the K videos;

[0060] Generating a second image template according to the commonalities of the video card images selected by the video editor;

[0061] Based on the second image template, generating multiple second alternative video card images that match the second image template.

[0062] According to the method of the above implementation, it is possible to generate the video card image of the video to be edited according to the videos that users hope to imitate among the videos with high video likes, so that the video card image of the video to be edited has a better visual effect.

[0063] In an implementation of the first aspect, generating multiple second alternative video card images that match the second image template based on the second image template includes:

[0064] Extracting second video frames from the first video;

[0065] Edit the second video frame according to the second image template to generate the second alternative video card image.

[0066] In one implementation of the first aspect, the extracting a second video frame from the first video includes:

[0067] Extract a second video frame from the first video whose picture dimension matches the second image template.

[0068] In one implementation of the first aspect, the extracting a second video frame from the first video includes:

[0069] Determine the user attention time interval of the first video according to the video editing information of the M videos or the K videos;

[0070] Extract the second video frame from the video segment of the first video corresponding to the user attention time interval of the first video.

[0071] According to the method of the above implementation, determining the user attention time interval of the current video according to the video with a high number of video likes, and extracting video frames according to the user attention time interval can make the alternative video card image closer to the image effect of the user attention video segment of the video with a high number of video likes.

[0072] In one implementation of the first aspect, the generating the second alternative video card image further includes:

[0073] Determine the user attention spatial region of the first video according to the video editing information of the M videos or the K videos;

[0074] Crop the second video frame according to the user attention spatial region of the first video.

[0075] In one implementation of the first aspect, the extracting a second video frame from the first video includes:

[0076] Determine the user attention time interval of the first video according to the video editing information of the first video;

[0077] Extract the second video frame from the video segment of the first video corresponding to the user attention time interval of the first video.

[0078] According to the method of the above implementation, determining the user attention time interval according to the video editing information of the current video, and extracting video frames according to the user attention time interval can make the extracted video frames more in line with the needs of video editors.

[0079] In an implementation of the first aspect, generating the second alternative video card image further includes:

[0080] Determine the user - focused spatial region of the first video according to the video editing information of the first video;

[0081] Crop the second video frame according to the user - focused spatial region of the first video.

[0082] In an implementation of the first aspect, the method further includes:

[0083] Select one or more alternative video card images from the multiple alternative video card images as the video card image of the first video.

[0084] In an implementation of the first aspect, selecting one or more alternative video card images from the multiple alternative video card images as the video card image of the first video includes:

[0085] Sort the multiple alternative video card images according to a preset rule, and display the multiple alternative video card images to the video publisher according to the sorting result;

[0086] Determine the alternative video card images selected by the video publisher from the multiple alternative video card images, and use the alternative video card images selected by the video publisher as the video card image of the first video.

[0087] In an implementation of the first aspect, sorting the multiple alternative video card images according to a preset rule includes:

[0088] Sort the pictures in the multiple alternative video card images according to the similarity calculation result and / or the relevance calculation result, where:

[0089] The similarity calculation result includes the similarity between the multiple alternative video card images and the video card images of the K videos, or the similarity between the multiple alternative video card images and the selected video card images among the video card images of the K videos;

[0090] The relevance calculation result includes the relevance between the multiple alternative video card images and the video text description of the first video.

[0091] In an implementation of the first aspect, sorting the pictures in the multiple alternative video card images according to the similarity calculation result and the relevance calculation result includes:

[0092] Calculate the similarity between the multiple alternative video card images and the video card images of the K videos to obtain the similarity calculation result;

[0093] Calculate the relevance between the multiple alternative video card images and the video title of the first video, and obtain the relevance calculation result;

[0094] Sort the multiple alternative video card images according to the similarity calculation result and / or the relevance calculation result.

[0095] In an implementation manner of the first aspect, sorting the images in the multiple alternative video card images according to the similarity calculation result and the relevance calculation result includes:

[0096] Display the video text description of the first video, where the video text description includes multiple alternative video titles of the first video;

[0097] Determine the alternative video title selected by the video publisher;

[0098] Calculate the similarity between the multiple alternative video card images and the video card images of the K videos, and obtain the similarity calculation result;

[0099] Calculate the relevance between the multiple alternative video card images and the alternative video title selected by the video publisher, and obtain the relevance calculation result;

[0100] Sort the multiple alternative video card images according to the similarity calculation result and / or the relevance calculation result.

[0101] In a second aspect, the present application provides a method for generating a video card image, where the method is applied to an electronic device, and the method includes:

[0102] Determine the user's attention time interval and the user's attention spatial region of the second video according to the user interaction behavior record of the second video;

[0103] Extract the fifth video frame from the video segment of the second video corresponding to the user's attention time interval of the second video;

[0104] Crop the fifth video frame according to the user's attention spatial region of the second video to generate a third alternative video card image.

[0105] According to the method of the second aspect, generating alternative video card images of a video based on the user interaction behavior of the video can associate the alternative video card images with the user's historical operations on the video, making the alternative video card images better able to assist the user in recalling the video content.

[0106] In a third aspect, the present application provides an electronic device, which includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the electronic device is triggered to execute the method steps described in the first aspect or the second aspect.

[0107] In a fourth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when running on a computer, causes the computer to execute the method described in the first aspect or the second aspect. Description of the Drawings

[0108] Figure 1 Shown is a flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0109] Figure 2 Shown is a flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0110] Figure 3 Shown is a partial flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0111] Figure 4 Shown is a partial flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0112] Figure 5 Shown is a partial flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0113] Figure 6 Shown is a partial flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0114] Figure 7 Shown is a partial flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0115] Figure 8 Shown is a partial flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0116] Figure 9 Shown is a partial flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0117] Figure 10 Shown is a partial flowchart of a method for generating a video card diagram according to an embodiment of the present application;

[0118] Figure 11 Shown is a schematic diagram of the system structure for generating a video card diagram according to an embodiment of the present application;

[0119] Figure 12 FIG2 is a flow chart of a method for generating a video card image according to an embodiment of the present application;

[0120] Figure 13 Shown is a schematic diagram of the interface display effect according to an embodiment of the present application;

[0121] Figure 14 Shown is a schematic diagram of the interface display effect according to an embodiment of the present application;

[0122] Figure 15 FIG2 is a flow chart of a method for generating a video card image according to an embodiment of the present application;

[0123] Figure 16 FIG2 is a flow chart of a method for generating a video card image according to an embodiment of the present application;

[0124] Figure 17 FIG2 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;

[0125] Figure 18 FIG2 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;

[0126] Figure 19 FIG2 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;

[0127] Figure 20 FIG2 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;

[0128] Figure 21 FIG2 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;

[0129] Figure 22 FIG2 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;

[0130] Figure 23 FIG2 is a simplified structural diagram of a video editing system according to an embodiment of the present application;

[0131] Figure 24 FIG2 is a flow chart of a method for generating a video card image according to an embodiment of the present application;

[0132] Figure 25 FIG2 is a flow chart of a method for generating a video card image according to an embodiment of the present application;

[0133] Figure 26 FIG. 1 is a schematic structural diagram of an electronic device according to an embodiment of the present application. Detailed implementation manners

[0134] To make the objectives, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without making creative efforts shall fall within the protection scope of the present application.

[0135] The terms used in the implementation manners part of the present application are only for explaining the specific embodiments of the present application, rather than aiming to limit the present application.

[0136] To obtain a video card image that can effectively attract video viewers to click, an embodiment of the present application provides a method for generating a video card image. The method provided by the embodiment of the present application is applied to an electronic device. In the method provided by the embodiment of the present application, the electronic device obtains similar videos of a first video; based on the current application scenario, filters the similar videos according to the user feedback information of the similar videos (for example, video click-through rate, video like count, etc.), wherein, compared with the videos that are not filtered out, the filtered similar videos better match the application requirements of the current application scenario; generates a video card image of the first video according to the filtered similar videos.

[0137] In the embodiments of this specification, the electronic devices that execute the video card image generation method process include but are not limited to mobile phones, tablet computers, laptop computers, wearable smart devices, desktop electronic computers, servers, etc.

[0138] Optionally, in one implementation, the electronic device that executes the video card image generation method process may be the terminal device of the video publisher, for example, a mobile phone, a tablet computer, a laptop computer, a desktop computer, etc. In another implementation, the electronic device that executes the video card image generation method process may be other electronic devices and / or cloud servers connected to the terminal device of the video publisher. In another implementation, the video card image generation method process may be partially executed by the terminal device of the video publisher and partially executed by other electronic devices and / or cloud servers connected to the terminal device of the video publisher.

[0139] Figure 1 The figure shows a flowchart of a method for generating a video card image according to an embodiment of the present application.

[0140] In one embodiment, the electronic device executes Figure 1 the following process shown to generate a video card image of the first video.

[0141] S101. Obtain M videos whose video content is similar to that of the first video, where M is an integer greater than or equal to 1.

[0142] S102. According to the user feedback information of the M videos, screen out K videos from the M videos, where K is an integer greater than or equal to 1 and less than or equal to M. The user feedback information includes the number of video likes and / or the video click-through rate.

[0143] S103. Obtain the video card images of the K videos, denoted as the first set of card images.

[0144] S104. According to the common features of the video card images in the first set of card images (such as large faces, full-body photos of beautiful women, landscapes, text labels, picture effects, etc.), generate multiple alternative video card images, denoted as the second set of card images.

[0145] S105. Select one or more alternative video card images from the second set of card images as the video card images of the first video.

[0146] According to the method of the embodiment of the present application, similar videos of the current video can be referred to, and the video card images of the current video can be generated according to the user feedback information of the similar videos, so that the video card images better meet the application scenario requirements of the current video.

[0147] The video card image generation method provided by the embodiment of the present application can be applied to the video sharing application scenario.

[0148] Figure 2 Shown is the flowchart of the video card image generation method according to an embodiment of the present application.

[0149] In one embodiment, in the video sharing application scenario, the electronic device executes Figure 2 the following process shown to generate the video card images of the first video.

[0150] In Figure 2 the shown embodiment, the first video is the video that the video publisher expects to publish to the video sharing platform for video sharing. For the first video, the video viewer has not clicked to watch it. For the first video, no user feedback information and user interaction behavior records are recorded.

[0151] S100. Obtain M videos whose video content is similar to that of the first video, where M is an integer greater than or equal to 1.

[0152] Specifically, in S100, screen out M videos from the videos that have been published to the video sharing platform and can be clicked and watched by the video viewer. The user feedback information of the M videos includes the video click-through rate.

[0153] S110. From the M videos obtained in S100, screen out the top K videos with the highest click-through rate (CTR). The K videos are recorded as a similar video set. K is an integer greater than or equal to 1 and less than or equal to M.

[0154] S120. Obtain the video card images of all the videos in the similar video set, which is recorded as the first card image set.

[0155] S130. Generate a first image template according to the common features of the video card images in the first card image set (such as large faces, full-body photos of beautiful women, landscapes, text labels, picture effects, etc.).

[0156] S140. Based on the first image template generated in S130, generate multiple alternative video card images that match the first image template, which is recorded as the second card image set.

[0157] S150. Select one or more alternative video card images from the second card image set as the video card image of the first video.

[0158] According to the method of the embodiment of the present application, an image template is generated based on the common features of the video card images of videos with high video click-through rates, and alternative video card images are generated according to the image template, so that the alternative video card images can attract video viewers to click more.

[0159] In S140, multiple different methods can be used to generate alternative video card images.

[0160] For example, Figure 3 The figure shows a partial flowchart of a method for generating a video card image according to an embodiment of the present application.

[0161] In one embodiment, the electronic device executes Figure 3 The following process shown in the figure to generate alternative video card images.

[0162] S200. Retrieve similar images of the first card image set in the picture library, which is recorded as the first alternative image set.

[0163] Specifically, in one implementation, the picture library is a local picture library. For example, the picture library is stored on the terminal device used by the video publisher to publish videos, or the picture library is stored on other local devices connected to the terminal device (such as a home network storage). In another implementation, the picture library is a cloud picture library stored on the server.

[0164] S210. Edit the images in the first alternative image set according to the first image template to generate alternative video card images, which is recorded as the first alternative video card image set, and use the first alternative video card image set as the second card image set.

[0165] Specifically, in one implementation of S210, editing the pictures in the first alternative picture set includes cropping the pictures (for example, cutting out adult faces, full-body pictures of beautiful women, etc.) and / or applying special effects processing (for example, attaching text labels, picture special effects, etc. to the pictures).

[0166] According to the method of the embodiment of the present application, using the existing pictures in the picture library as the picture source, screening and editing the pictures in the picture source according to the video card pictures of the videos with high video click-through rate to generate alternative video card pictures can effectively limit the source range of the alternative video card pictures.

[0167] For another example, Figure 4 The figure shows a partial method flow chart of the video card picture generation method according to an embodiment of the present application.

[0168] In one embodiment, the electronic device executes Figure 4 the following process shown in the figure to generate alternative video card pictures.

[0169] S300, extract video frames from the first video, and record them as the second alternative picture set.

[0170] Specifically, in S300, various methods can be used to extract video frames from the first video.

[0171] For example, in one embodiment, extract video frames from the first video at a fixed time interval.

[0172] For another example, Figure 5 The figure shows a partial method flow chart of the video card picture generation method according to an embodiment of the present application.

[0173] In one embodiment, the electronic device executes Figure 5 the following process shown in the figure to extract video frames from the first video.

[0174] S400, according to the user interaction behavior records of M videos (the videos obtained in S100) or K videos (the K videos screened in S110), determine the user attention time intervals (the first user attention time intervals) of the M videos or K videos.

[0175] In one embodiment, for the video sharing application scenario, the user attention time interval refers to the time interval during which the video viewer focuses on the video content. Specifically, in one embodiment, during the process of a video viewer watching a video, the time interval corresponding to the video player's video playback control (such as pausing, rewinding, slow playback, etc.) and / or video viewing interaction (such as entering bullet comments, taking screenshots, magnifying images, etc.) is the user attention time interval. In S400, based on the user interaction behavior record of the video, the user attention time interval of the video is determined.

[0176] For example, the video viewer fast-forwards and / or drags the time bar to quickly browse the video content. During the process of quickly browsing the video content, the time period during which the video viewer watches the video at the normal speed in its entirety is the user attention time interval. For another example, during the process of a video viewer watching a video, the video viewer rewinds the video playback progress (such as by dragging the time bar) and repeatedly watches the video, and the time period for this is the user attention time interval. For another example, during the process of a video viewer watching a video, the video viewer pauses the video playback and carefully watches the video content, and the time period for this is the user attention time interval. For another example, during the process of a video viewer watching a video, the time period during which the video viewer enters comment subtitles (bullet comments) is the user attention time interval.

[0177] S410, determine the user attention time interval (the second user attention time interval) of the first video according to the first user attention time interval.

[0178] Specifically, in one implementation manner of S410, according to the video segment of the first user attention time interval, similar video segments are retrieved from the first video (for example, based on the screen similarity retrieval algorithm, similar video frames are retrieved), and the time interval corresponding to the retrieved similar video segments is the second user attention time interval.

[0179] S420, extract video frames from the video segment of the first video corresponding to the second user attention time interval.

[0180] According to the method of the embodiment of the present application, extracting video frames according to the user attention time interval can effectively ensure that the extracted video frames are more likely to attract the video viewer to click.

[0181] S310, edit the pictures in the second alternative picture set according to the first image template to generate alternative video card pictures, denoted as the second alternative video card picture set, and use the second alternative video card picture set as the second card picture set.

[0182] In S310, when editing the pictures in the second alternative picture set, S210 can be referred to.

[0183] According to the method of the embodiments of the present application, using the video frames of the current video as the source images, and screening and editing the images in the source images according to the video card images of the videos with high video click-through rates to generate alternative video card images can ensure the relevance between the alternative video card images and the current video.

[0184] In one implementation manner of S310, a user attention spatial region is also introduced. In S310, the images in the second alternative image set are also cropped according to the user attention spatial region of the first video.

[0185] In one embodiment, for the video sharing application scenario, the user attention spatial region refers to the image region that the video viewer focuses on the video image. Specifically, in one embodiment, during the process of the video viewer watching the video, the image region corresponding to the video viewing interaction (such as partial screenshot, image zooming, etc.) of the video viewer is the user attention spatial region.

[0186] For example, during the process of the video viewer watching the video, if the video viewer zooms in on a certain image region multiple times, this image region is the user attention spatial region. Also for example, during the process of the video viewer watching the video, if the video viewer takes a partial screenshot of a certain image region, this image region is the user attention spatial region.

[0187] Specifically, in one embodiment, determine the user attention spatial region (the first user attention spatial region) of M videos (the videos obtained in S100) or K videos (the K videos screened in S110) in the similar video set. Determine the user attention spatial region (the second user attention spatial region) of the first video according to the first user attention spatial region.

[0188] Specifically, in one implementation manner, according to the video frames of the first user attention spatial region, retrieve similar video frames from the first video, and save the first user attention spatial region as the second user attention spatial region of the retrieved similar video frames.

[0189] According to the method of the embodiments of the present application, cropping the video frames according to the user attention spatial region can effectively ensure that the alternative video card images are more attractive to video viewers to click.

[0190] Also for example, Figure 6 Shown is a partial method flowchart of the video card image generation method according to an embodiment of the present application.

[0191] In one embodiment, the electronic device executes Figure 6 the following process shown to generate alternative video card images.

[0192] S500, call the AI Generated Content (AIGC) model to generate content automatically.

[0193] S510, use the AIGC model to generate alternative video card images, denoted as the third set of alternative video card images, based on the first image template generated in S130 and the video text description of the first video, and use the third set of alternative video card images as the second set of card images.

[0194] In one embodiment, the video text description of the first video includes alternative video titles of the first video and / or the original video title of the first video.

[0195] Specifically, in one embodiment, the original video title of the first video and multiple alternative video titles are input by the video publisher.

[0196] Specifically, in another embodiment, the original video title of the first video is input by the video publisher. Referring to the video titles of K videos (the videos screened in S110), the original video title of the first video is rewritten to generate multiple alternative video titles.

[0197] In S510, the AIGC model generates different alternative video card images by combining different alternative video titles.

[0198] The video text description of the first video can be the video title input by the video publisher or the description information input by the video publisher for the content of the first video.

[0199] For another example, Figure 7 The figure shows a partial flowchart of the method for generating video card images according to an embodiment of the present application.

[0200] In one embodiment, the electronic device executes Figure 7 the following process shown in the figure to generate alternative video card images.

[0201] S600, call the AI Generated Content (AIGC) model to generate content automatically.

[0202] S610, perform semantic segmentation on the images of the first set of card images, and / or the first set of alternative video card images ( Figure 3 embodiment), and / or the second set of alternative video card images ( Figure 4 embodiment) to obtain semantic segmentation maps.

[0203] S620, using the AIGC model, based on the semantic segmentation map generated in S610 and the video text description of the first video, generates an alternative video card map, recorded as the fourth alternative video card map set, and uses the fourth alternative video card map set as the second card map set.

[0204] The video text description of the first video in S620 may refer to S510.

[0205] Furthermore, in another embodiment, a second card image set may be generated by combining multiple alternative video card image sets. For example, any multiple of the first alternative video card image set, the second alternative video card image set, the third alternative video card image set, and the fourth alternative video card image set may be combined to form the second card image set.

[0206] Furthermore, in S150 , a video card image of the first video may be selected from the second card image set in a variety of different ways.

[0207] For example, in one implementation of S150, the pictures in the second card image set are displayed to the video publisher, and the video publisher selects the pictures in the second card image set as the video card images.

[0208] In another implementation of S150, the pictures in the second card picture set are sorted according to a preset rule, and the top N pictures are used as video card pictures, where N is a preset integer greater than or equal to one.

[0209] In another implementation of S150, the images in the second card image set are sorted according to a preset rule, and the images in the second card image set are displayed to the video publisher according to the sorting result (for example, the images in the top ranking are displayed to the video publisher first). The video publisher selects the image as the video card image.

[0210] For example, Figure 8 Shown is a partial method flow chart of a video card image generation method according to an embodiment of the present application.

[0211] In one embodiment, the electronic device executes Figure 8 The following process is shown to achieve 150.

[0212] S700: Calculate the similarity between each picture in the second card image set and the first card image set.

[0213] S710 , sorting the pictures in the second card picture set according to the similarity calculation result in S700 , wherein the pictures with greater similarity are ranked higher.

[0214] S720, according to the sorting result of S710, display the pictures in the second set of card pictures to the video publisher, and give priority to displaying the pictures with higher sorting to the video publisher.

[0215] S730, receive the picture selection operation of the video publisher, and select the picture to be used as the video card picture from the second set of card pictures according to the picture selection operation of the video publisher.

[0216] For another example, Figure 9 The figure shows a partial flowchart of a method for generating a video card picture according to an embodiment of the present application.

[0217] In one embodiment, the electronic device executes Figure 9 the following process shown in the figure to implement 150.

[0218] S800, obtain the video text description of the first video, for example, an alternative video title.

[0219] The video text description of the first video in S800 may refer to S510.

[0220] S810, calculate the relevance between each picture in the second set of card pictures and the video text description.

[0221] S820, sort the pictures in the second set of card pictures according to the relevance calculation result in S810, and the higher the relevance, the higher the sorting.

[0222] S830, according to the sorting result of S820, display the pictures in the second set of card pictures to the video publisher, and give priority to displaying the pictures with higher sorting to the video publisher.

[0223] Specifically, in one embodiment, obtain multiple alternative video titles of the first video in S800. Calculate the relevance between each picture in the second set of card pictures and each alternative video title in S810. Display multiple alternative video titles to the video publisher in S830, and the video publisher selects one video title from the multiple alternative video titles as the video title of the first video. And, after the video publisher selects an alternative video title, sort the pictures in the second set of card pictures according to the relevance between each picture and the selected alternative video title, and display the pictures in the second set of card pictures.

[0224] For example, the second set of card pictures includes picture A, picture B, and picture C. The alternative video titles include title T1 and title T2.

[0225] The relevance ranking of Picture A, Picture B, and Picture C with respect to title T1 from high to low is Picture B, Picture C, Picture A; the relevance ranking of Picture A, Picture B, and Picture C with respect to title T2 from high to low is Picture A, Picture C, Picture B.

[0226] In S830, display title T1 and title T2 to the user. When the user selects title T1, display Picture A, Picture B, and Picture C (display Picture B first) based on the relevance ranking of Picture B, Picture C, Picture A. When the user selects title T2, display Picture A, Picture B, and Picture C (display Picture A first) based on the relevance ranking of Picture A, Picture C, Picture B.

[0227] S840, receive the title selection operation and picture selection operation of the video publisher, and determine the video title of the first video according to the title selection operation and picture selection operation of the video publisher, and select the picture as the video card picture from the second set of card pictures.

[0228] For another example, Figure 10 Shown is a partial method flowchart of a video card picture generation method according to an embodiment of the present application.

[0229] In one embodiment, the electronic device executes Figure 10 the following process shown to implement 150.

[0230] S900, obtain the video text description of the first video.

[0231] The alternative video titles of the first video in S900 may refer to S510.

[0232] S910, calculate the relevance of each picture in the second set of card pictures to the video text description.

[0233] S920, calculate the similarity of each picture in the second set of card pictures to the first set of card pictures.

[0234] S930, sort the pictures in the second set of card pictures according to the relevance calculation result in S910 and the similarity calculation result in S9…

[0235] For example, perform a weighted average calculation on the relevance calculation result and the similarity calculation result, and sort the pictures in the second set of card pictures according to the weighted average calculation result.

[0236] S940, display the pictures in the second set of card pictures to the video publisher according to the sorting result of S930, and preferentially display the pictures with higher sorting to the video publisher. (Refer to S830)

[0237] S950 receives the picture selection operation of the video publisher, and determines the video card picture of the first video according to the picture selection operation of the video publisher.

[0238] Specifically, in one embodiment, in S900, multiple alternative video titles of the first video are obtained. In S910, the relevance between each picture in the second set of card pictures and each alternative video title is calculated. In S940, the multiple alternative video titles are displayed to the video publisher, and the video publisher selects one video title from the multiple alternative video titles as the video title of the first video. Moreover, after the video publisher selects an alternative video title, the pictures in the second set of card pictures are sorted according to the relevance between each picture in the second set of card pictures and the selected alternative video title, and the similarity between each picture in the second set of card pictures and the first set of card pictures, and the pictures in the second set of card pictures are displayed according to the sorting result.

[0239] According to the video card picture generation method proposed by the embodiments of the present application, an embodiment of the present application also proposes a video card picture generation system.

[0240] Optionally, in one implementation, the video card picture generation system can be built in the terminal device used by the video publisher to publish videos, such as mobile phones, tablets, laptops, desktops, etc. In another implementation, the video card picture generation system can be built in other electronic devices and / or cloud servers connected to the terminal device used by the video publisher to publish videos. In another implementation, the video card picture generation system can be partially built in the terminal device used by the video publisher to publish videos and partially built in other electronic devices and / or cloud servers connected to the terminal device used by the video publisher to publish videos.

[0241] Figure 11 Shown is a schematic diagram of the structure of a video card picture generation system according to an embodiment of the present application.

[0242] As Figure 11 shown, the video card picture generation system includes:

[0243] A similarity retrieval module 1010, which is used to perform video similarity and picture similarity retrieval;

[0244] A template extraction module 1020, which is used to extract common features based on a set of pictures and generate a first image template according to the common features;

[0245] An image cropping module 1030, which is used to crop pictures according to specific dimensions, proportions, and image area position information, and edit pictures according to specified rules;

[0246] The AIGC module 1040 is used to generate pictures of a specific size according to any one or a combination of the title / semantic segmentation map / template map.

[0247] The sorting module 1050 is used to calculate the relevance between the picture and the title and the similarity between pictures, and sort multiple pictures according to the calculation results of the relevance and similarity.

[0248] The video frame extraction module 1060 is used to extract video frames.

[0249] Furthermore, the video card image generation method provided by the embodiments of the present application can be applied to the situation where a video publisher publishes a video for the first time in a video sharing application scenario.

[0250] Figure 12 Shown is a flowchart of a video card image generation method according to an embodiment of the present application.

[0251] Figure 13 Shown is a schematic diagram of the interface display effect according to an embodiment of the present application.

[0252] In one embodiment, Figure 11 The system shown executes Figure 12 The steps shown to implement the generation of video card images.

[0253] In one embodiment, Figure 11 The system shown executes Figure 12 During the process of executing the steps shown, the terminal device of the video publisher displays an interface as Figure 13 shown.

[0254] S1100, receive the video to be published (the first video) uploaded by the video publisher and the video title (the original video title) of the video to be published uploaded by the video publisher.

[0255] As Figure 13 shown, the cover / play interface of the received video to be published is displayed at 1200, and the video title uploaded by the video publisher is displayed at 1201.

[0256] Optionally, in one embodiment, in S1100, the video type, the card image layout (such as video large Figure 17 :9 / small square Figure 1 :1), the potential population targeted by the video to be published (for example, male / female), etc. input by the video publisher are also received.

[0257] S1110, the similarity retrieval module 1010 retrieves similar videos (M) from the distributed video library based on the first video and the original video title.

[0258] The video distribution library is used to store videos, as well as the CTR of the videos and user interaction records.

[0259] Specifically, in S1110, when retrieving similar videos, the video type of the first video, the card layout (such as large video Figure 17 :9 / small square Figure 1 :1), the potential population targeted by the first video (for example, male / female), etc. are also matched.

[0260] S1111, the similarity retrieval module 1010 filters out the TOP-K similar videos from the retrieved M similar videos based on the CTR of the videos, and records them as the similar video set.

[0261] S1112, the similarity retrieval module 1010 obtains the video cards of each video in the similar video set, and records them as the first card set.

[0262] As Figure 13 shown, the cover / play interface of the videos in the similar video set and the corresponding video cards are displayed at 1202, and the video publisher can switch to display different videos in the similar video set by dragging the slider 1203.

[0263] S1120, the template extraction module 1020 extracts the common features from the pictures in the first card set to generate the first image template.

[0264] S1113, the similarity retrieval module 1010 retrieves similar pictures in the picture library according to the pictures in the first card set, and records them as the first alternative picture set.

[0265] S1130, the image cropping module 1030 edits the pictures in the first alternative picture set based on the first image template to generate alternative video cards, and records them as the first alternative video card set.

[0266] As Figure 13 shown, the pictures in the first alternative video card set are displayed at 1204.

[0267] S1114, the similarity retrieval module 1010 determines the user attention time interval of the videos in the similar video set according to the user interaction behavior records of the videos in the similar video set.

[0268] As Figure 13 shown, the user attention time interval of the videos in the similar video set is marked in the video progress bar 1205.

[0269] S1115, the similarity retrieval module 1010 retrieves similar video segments from the video to be published according to the video segments corresponding to the user attention time interval of the videos in the similar video set.

[0270] As Figure 13 shown, the video clip corresponding to the user's attention time interval in the video to be released is marked in the video progress bar 1206.

[0271] S1140, the frame extraction module 1060 extracts video frames from the similar video clips, denoted as the second alternative picture set.

[0272] S1116, the similarity retrieval module 1010 determines the user's attention spatial region of the video to be released according to the user interaction behavior records of the videos in the similar video set.

[0273] As Figure 13 shown, the user's attention spatial region of the videos in the similar video set is marked in the playback interface of the videos in the similar video set (for example, the user's attention spatial region of a certain video frame in the videos in the similar video set is 1207). The user's attention spatial region of the video to be released is marked in the playback interface of the video to be released (for example, the user's attention spatial region of a certain video frame of the video to be released is 1208).

[0274] S1131, the image cropping module 1030 edits the pictures in the second alternative picture set based on the first image template and the user's attention spatial region to generate alternative video card pictures, denoted as the second alternative video card picture set.

[0275] As Figure 13 shown, the pictures in the second alternative video card picture set are displayed at 1209.

[0276] S1150, the similarity retrieval module 1010 refers to the video titles of the videos in the similar video set and rewrites the video title uploaded by the video publisher to generate alternative video titles.

[0277] S1160, the AIGC module 1040 generates different alternative video card pictures based on the first image template and in combination with different alternative video titles, denoted as the third alternative video card picture set.

[0278] S1161, the AIGC module 1040 performs semantic segmentation on the pictures in the first alternative video card picture set and the second alternative video card picture set to generate a semantic segmentation map.

[0279] S1162, the AIGC module 1040 generates different alternative video card pictures based on the semantic segmentation map and in combination with different alternative video titles, denoted as the fourth alternative video card picture set.

[0280] As Figure 13 shown, the pictures in the third alternative video card picture set and the fourth alternative video card picture set are displayed at 1211.

[0281] In S1170, the sorting module 1050 calculates the relevance between the pictures and the alternative video titles in the first set of alternative video card pictures, the second set of alternative video card pictures, the third set of alternative video card pictures, and the fourth set of alternative video card pictures, calculates the similarity between the pictures in the first set of alternative video card pictures, the second set of alternative video card pictures, the third set of alternative video card pictures, and the fourth set of alternative video card pictures and the first set of card pictures, and sorts the pictures in the second set of card pictures according to the calculation results of the relevance and similarity.

[0282] Specifically, in one embodiment, the sorting module 1050 calculates the relevance and similarity for the first set of alternative video card pictures, the second set of alternative video card pictures, the third set of alternative video card pictures, and the fourth set of alternative video card pictures respectively, and sorts the first set of alternative video card pictures, the second set of alternative video card pictures, the third set of alternative video card pictures, and the fourth set of alternative video card pictures respectively.

[0283] In S1180, the alternative video titles are displayed, and according to the sorting result, the pictures in the first set of alternative video card pictures, the second set of alternative video card pictures, the third set of alternative video card pictures, and the fourth set of alternative video card pictures are displayed.

[0284] In S1181, according to the user's selection operation, the video title of the first video and the video card picture of the first video are determined.

[0285] As Figure 13 shown, the alternative video titles are displayed at 1210, and the video publisher selects one of the multiple alternative video titles as the video title of the video to be published. After the video publisher selects an alternative video title, according to the sorting result for this video title, pictures are displayed at 1204, 1209, and 1211 (displaying the top two pictures). After the video publisher clicks on one or more of the pictures displayed at 1204, 1209, and 1211, the pictures clicked by the video publisher are used as the video card pictures of the video to be published.

[0286] It should be noted here that in the embodiments of the present application, the interface displayed on the terminal device of the video publisher is not limited to Figure 13 the mode shown. Those skilled in the art can design the interface mode displayed on the terminal device of the video publisher according to actual needs.

[0287] For example, Figure 14 shown is a schematic diagram of the interface display effect according to an embodiment of the present application.

[0288] As Figure 14As shown, the cover / play interface of the video to be released is displayed at 1300, and the video title uploaded by the video publisher is displayed at 1301.

[0289] Alternative video titles are displayed at 1302, and the video publisher selects one alternative video title from multiple alternative video titles as the video title of the video to be released.

[0290] The pictures of the first alternative video card image set, the second alternative video card image set, the third alternative video card image set, and the fourth alternative video card image set are displayed at 1303.

[0291] The sorting module 1050 calculates the relevance between the pictures in the first alternative video card image set, the second alternative video card image set, the third alternative video card image set, and the fourth alternative video card image set and the alternative video titles, calculates the similarity between the pictures in the first alternative video card image set, the second alternative video card image set, the third alternative video card image set, and the fourth alternative video card image set and the first card image set, and sorts the pictures in the second card image set according to the calculation results of the relevance and similarity.

[0292] After the video publisher selects an alternative video title, according to the sorting result for the alternative video title, pictures are displayed at 1303 (displaying the top six pictures). After the video publisher clicks on one or more of the pictures displayed at 1303, the pictures clicked by the video publisher are used as the video card images of the video to be released.

[0293] Furthermore, the video card image generation method provided by the embodiments of the present application can be applied to the update of the current video card images of videos in a video sharing application scenario.

[0294] Figure 15 Shown is a flowchart of a video card image generation method according to an embodiment of the present application.

[0295] In one embodiment, Figure 11 The system shown executes Figure 15 The steps shown to implement the update of the video card image of the first video.

[0296] In Figure 15 The embodiment shown, the first video is a video that has been released to a video sharing platform and has been clicked and viewed by video viewers. For the first video, user feedback information and user interaction behavior records have been recorded.

[0297] S1400, obtain the video (the first video) for which the video card image needs to be updated and the video title (the original video title) of the video.

[0298] Optionally, in one embodiment, in S1400, the video type, card layout, potential population targeted by the video to be released, etc. of the first video are also obtained.

[0299] S1410, the similarity retrieval module 1010 retrieves similar videos (M) of the first video from the distribution video library.

[0300] S1411, the similarity retrieval module 1010 filters out the TOP-K similar videos from the retrieved M similar videos based on CTR, and records them as the similar video set.

[0301] S1412, the similarity retrieval module 1010 obtains the video card images of each video in the similar video set, and records them as the first card image set.

[0302] S1420, the template extraction module 1020 extracts the common features of the images in the first card image set to generate the first image template.

[0303] S1413, the similarity retrieval module 1010 retrieves similar images in the image library according to the images in the first card image set, and records them as the first alternative image set.

[0304] S1430, the image cropping module 1030 edits the images in the first alternative image set based on the first image template to generate alternative video card images, which are recorded as the first alternative video card image set.

[0305] S1414, the similarity retrieval module 1010 determines the user attention time interval of the first video according to the user interaction behavior records of the videos in the similar video set and the user interaction behavior records of the first video.

[0306] S1440, the frame extraction module 1060 extracts video frames from the video segment corresponding to the user attention time interval of the first video, and records them as the second alternative image set.

[0307] S1415, the similarity retrieval module 1010 determines the user attention space area of the first video according to the user interaction behavior records of the videos in the similar video set and the user interaction behavior records of the first video.

[0308] S1431, the image cropping module 1030 edits the images in the second alternative image set based on the first image template and the user attention space area to generate alternative video card images, which are recorded as the second alternative video card image set.

[0309] S1450, the similarity retrieval module 1010 rewrites the video title of the first video with reference to the video titles of the videos in the similar video set to generate alternative video titles.

[0310] S1460. The AIGC module 1040 generates different alternative video card images based on the first image template and in combination with different alternative video titles, which are denoted as the third set of alternative video card images.

[0311] S1461. The AIGC module 1040 performs semantic segmentation on the images in the first set of alternative video card images and the second set of alternative video card images to generate a semantic segmentation map.

[0312] S1462. The AIGC module 1040 generates different alternative video card images based on the semantic segmentation map and in combination with different alternative video titles, which are denoted as the fourth set of alternative video card images.

[0313] S1470. The sorting module 1050 uses the combination of the first set of alternative video card images, the second set of alternative video card images, the third set of alternative video card images, and the fourth set of alternative video card images as the second set of card images, calculates the relevance between the images in the second set of card images and the video title of the first video, calculates the similarity between the images in the second set of card images and the first set of card images, and sorts the images in the second set of card images according to the calculation results of the relevance and similarity.

[0314] S1480. One or more images ranked first in the sorting result are used as the video card images of the first video.

[0315] Furthermore, the video card image generation method provided by the embodiments of the present application can also be applied to the video editing application scenario.

[0316] In some video editing application scenarios, the user hopes to edit the video frames of the video to obtain video card images, and the user hopes that the image effect of the video card images can imitate the image effect of the video card images of some videos.

[0317] Figure 16 The figure shows a flowchart of the video card image generation method according to an embodiment of the present application.

[0318] In one embodiment, in the video editing application scenario, the electronic device executes Figure 16 the following process shown to generate the video card images of the first video.

[0319] In Figure 16 the shown embodiment, the first video is the video to be edited.

[0320] S1500. Obtain M videos similar to the video content of the first video, where M is an integer greater than or equal to 1.

[0321] Specifically, in S1500, M videos are screened out from the sample videos in the editing sample library and / or the shared videos. The user feedback information of the M videos includes the number of video likes.

[0322] S1510. From the M videos obtained in S1500, K (Top-K) videos with the highest video like counts are screened out, where K is an integer greater than or equal to 1 and less than or equal to M.

[0323] S1511. Show the K videos screened out in S1510 and the video card diagrams of the K videos to the user.

[0324] S1520. Receive the user's selection operation, and determine the video card diagrams selected by the user according to the user's selection operation, denoted as the third set of card diagrams.

[0325] In one embodiment, the user selects a video, and the video card diagram corresponding to the video is determined according to the selected video by the user, denoted as the third set of card diagrams. In another embodiment, the user selects a video card diagram, and according to the selected video card diagram by the user, it is denoted as the third set of card diagrams.

[0326] S1530. Generate a second image template according to the commonalities of the video card diagrams in the third set of card diagrams.

[0327] S1540. Generate multiple alternative video card diagrams that match the second image template based on the second image template generated in S, denoted as the fourth set of card diagrams.

[0328] S1550. Select one or more alternative video card diagrams from the fourth set of card diagrams as the video card diagram of the first video.

[0329] According to the method of the embodiment of the present application, the video card diagram of the video to be edited can be generated according to the videos that users hope to imitate among the videos with high video like counts, so that the video card diagram of the video to be edited has a better visual effect.

[0330] Referring to the second set of card diagrams in S140, in S1540, multiple different methods can be used to generate alternative video card diagrams.

[0331] For example, Figure 17 The figure shows a partial flowchart of the video card diagram generation method according to an embodiment of the present application.

[0332] In one embodiment, the electronic device executes Figure 17 The following process shown in the figure to generate alternative video card diagrams.

[0333] S1600. Extract video frames whose picture dimensions match the second image template from the first video, denoted as the third set of alternative pictures.

[0334] For example, the screen dimensions of the second image template include woods, a river, and a human figure. In S1600, video frames that simultaneously include woods, a river, and a human figure are extracted from the first video.

[0335] Another example is that the screen dimensions of the second image template are a large face of a female human figure and a dog's head. In S1600, video frames that simultaneously include a large face of a female human figure and a dog's head are extracted from the first video.

[0336] S1610, according to the second image template, edit the pictures in the third set of alternative pictures to generate alternative video card pictures, denoted as the fifth set of alternative video card pictures, and use the fifth set of alternative video card pictures as the fourth set of card pictures.

[0337] According to the method of the embodiment of the present application, using the video frames of the current video as the source images, and screening and editing the pictures in the source images according to the commonalities of the video card pictures of videos with high video like counts to generate alternative video card pictures can, on the basis of ensuring the relevance between the alternative video card pictures and the current video, make the alternative video card pictures closer to the picture effects of the video card pictures of videos with high video like counts.

[0338] Another example, Figure 18 Shown is a partial method flowchart of a video card picture generation method according to an embodiment of the present application.

[0339] In S1500, M videos are screened out from the sample videos in the editing sample library and / or the videos that have been shared. The M videos are videos that have been edited, and video editing information has been recorded for the M videos.

[0340] In one embodiment, the electronic device executes Figure 18 the following process shown to generate alternative video card pictures.

[0341] S1700, according to the video editing information of the M videos (the videos obtained in S1500) or the K videos (the K videos screened in S1510), determine the user attention time intervals (the third user attention time intervals) of the M videos or the K videos.

[0342] In one embodiment, for the video editing application scenario, the user attention time interval refers to the time interval during which the video editor edits the video content. For example, the time interval corresponding to using special effects such as slow motion / loop playback as a video segment is the user attention time interval.

[0343] S1710, determine the user attention time interval (the fourth user attention time interval) of the first video according to the third user attention time interval.

[0344] Specifically, in one implementation of S1710, according to the video clips within the first user's attention time interval, similar video clips are retrieved from the first video (for example, based on a video frame similarity retrieval algorithm to retrieve similar video frames), and the time interval corresponding to the retrieved similar video clips is used as the fourth user's attention time interval.

[0345] S1720, extract video frames from the video clips of the first video corresponding to the fourth user's attention time interval, and denote them as the fourth set of alternative pictures.

[0346] S1730, according to the second image template, edit the pictures in the fourth set of alternative pictures to generate alternative video card pictures, denoted as the sixth set of alternative video card pictures, and use the sixth set of alternative video card pictures as the fourth set of card pictures.

[0347] According to the method of the embodiments of the present application, determining the user's attention time interval of the current video based on the video with a high number of video likes, and extracting video frames according to the user's attention time interval can make the alternative video card pictures closer to the image effect of the user's attention video clips of the video with a high number of video likes.

[0348] In one implementation of S1730, the user's attention spatial region is also introduced. In S1730, the pictures in the fourth set of alternative pictures are also cropped according to the user's attention spatial region of the first video.

[0349] In one embodiment, for the video editing application scenario, the user's attention spatial region refers to the image region where the video editor edits the video content. For example, the image region where close-up zooming / sticker / image special effects are used is the user's attention spatial region.

[0350] Specifically, in one embodiment, determine the user's attention spatial region (the third user's attention spatial region) of M videos (the videos obtained in S1500) or K videos (the K videos screened in S1510). Determine the user's attention spatial region (the fourth user's attention spatial region) of the first video according to the third user's attention spatial region.

[0351] Specifically, in one implementation, according to the video frames of the second user's attention spatial region, similar video frames are retrieved from the first video, and the second user's attention spatial region is saved as the fourth user's attention spatial region of the retrieved similar video frames.

[0352] According to the method of the embodiments of the present application, using the video frames of the current video as the source images, and screening and editing the images in the source images according to the commonalities of the video card images of the videos with high video like counts, alternative video card images can be generated. On the basis of ensuring the relevance between the alternative video card images and the current video, the alternative video card images can be made more similar to the image effects of the video card images of the videos with high video like counts.

[0353] For another example, Figure 19 The figure shows a partial flowchart of a method for generating video card images according to an embodiment of the present application.

[0354] The first video is a video that has been edited, and video editing information has been recorded for the first video.

[0355] In one embodiment, the electronic device executes Figure 19 the following process shown to generate alternative video card images.

[0356] S1800. Determine the user attention time interval (the fifth user attention time interval) of the first video according to the video editing information of the first video.

[0357] For example, in one embodiment, the time interval corresponding to the video segment edited by the user is the user attention time interval.

[0358] S1810. Extract video frames from the video segments of the first video corresponding to the fifth user attention time interval, and denote them as the fifth alternative image set.

[0359] S1820. Edit the images in the fifth alternative image set according to the second image template to generate alternative video card images, denoted as the seventh alternative video card image set, and use the seventh alternative video card image set as the fourth card image set.

[0360] According to the method of the embodiments of the present application, determining the user attention time interval according to the video editing information of the current video and extracting video frames according to the user attention time interval can make the extracted video frames more in line with the needs of video editors.

[0361] In one implementation manner of S1820, the user attention spatial region is also introduced. In S1820, the images in the fourth alternative image set are also cropped according to the user attention spatial region (the fifth user attention spatial region) of the first video.

[0362] For example, in one embodiment, the image region edited by the user is the user attention spatial region.

[0363] Further, in another embodiment, a combination of multiple alternative video card image sets can be used to generate a fourth card image set. For example, any combination of the fifth alternative video card image set, the sixth alternative video card image set, and the seventh alternative video card image set can be used to form the fourth card image set. Further, the generation process of the second card image set in the video sharing application scenario can be referred to for generating the fourth card image set.

[0364] Further, referring to S150, in S1550, various different methods can be used to select the video card image of the first video from the fourth card image set.

[0365] For example, in one implementation manner of S1550, the images in the fourth card image set are shown to the video editor. The video editor selects the image to be used as the video card image from the fourth card image set.

[0366] In another implementation manner of S1550, according to a preset rule, the images in the fourth card image set are sorted, and the top N images after sorting are used as the video card images, where N is a preset integer greater than or equal to one.

[0367] In another implementation manner of S1550, according to a preset rule, the images in the fourth card image set are sorted, and the images in the fourth card image set are shown to the video editor according to the sorting result (for example, the images ranked in the front are shown to the video editor first). The video editor selects the image to be used as the video card image.

[0368] For example, Figure 20 FIG. shows a partial method flowchart of a video card image generation method according to an embodiment of the present application.

[0369] In one embodiment, the electronic device executes Figure 20 the following process shown to implement 5150.

[0370] S1900, calculate the similarity between each image in the fourth card image set and the third card image set.

[0371] S1910, sort the images in the fourth card image set according to the similarity calculation result in S1900, and the greater the similarity, the higher the ranking.

[0372] S1920, show the images in the fourth card image set to the video editor according to the sorting result of S1910, and give priority to showing the images with higher rankings to the video editor.

[0373] S1930, receive the image selection operation of the video editor, and select the image to be used as the video card image from the fourth card image set according to the image selection operation of the video editor.

[0374] For another example, Figure 21 Shown is a partial method flowchart of a method for generating a video card diagram according to an embodiment of the present application.

[0375] In one embodiment, the electronic device executes Figure 21 the following process shown to implement 1550.

[0376] S2000, calculate the relevance between each picture in the fourth card diagram set and the video title of the first video.

[0377] S2010, sort the pictures in the fourth card diagram set according to the relevance calculation result in S2000, and the higher the relevance, the higher the sorting.

[0378] S2020, according to the sorting result of S2010, display the pictures in the fourth card diagram set to the video editor, and preferentially display the pictures with higher sorting to the video editor.

[0379] S2030, receive the picture selection operation of the video editor, and select the picture to be used as the video card diagram from the fourth card diagram set according to the picture selection operation of the video editor.

[0380] For another example, Figure 22 Shown is a partial method flowchart of a method for generating a video card diagram according to an embodiment of the present application.

[0381] In one embodiment, the electronic device executes Figure 22 the following process shown to implement 1550.

[0382] S2100, calculate the relevance between each picture in the fourth card diagram set and the video title of the first video.

[0383] S2110, calculate the similarity between each picture in the fourth card diagram set and the third card diagram set.

[0384] S2120, sort the pictures in the fourth card diagram set according to the relevance calculation result in S2100 and the similarity calculation result in S2110.

[0385] For example, perform a weighted average calculation on the relevance calculation result and the similarity calculation result, and sort the pictures in the fourth card diagram set according to the weighted average calculation result.

[0386] S2130, according to the sorting result of S2120, display the pictures in the fourth card diagram set and the alternative video titles to the video publisher, and preferentially display the pictures with higher sorting to the video editor.

[0387] S2140, receive the picture selection operation of the video editor, and select a picture as the video card picture from the fourth set of card pictures according to the picture selection operation of the video editor.

[0388] Figure 23 The following shows a schematic diagram of the video editing system structure according to an embodiment of the present application.

[0389] As Figure 23 shown, the video editing system includes:

[0390] A similarity retrieval module 2210, which is used to perform video similarity and picture similarity retrieval;

[0391] A template extraction module 2220, which is used to extract common features based on a set of pictures and generate a second image template according to the common features;

[0392] An image cropping module 2230, which is used to crop pictures according to specific dimensions, ratios, and image area position information, and edit pictures according to specified rules;

[0393] A sorting module 2240, which is used to calculate the relevance between pictures and titles and the similarity between pictures, and sort multiple pictures according to the calculation results of the relevance and similarity;

[0394] A frame extraction module 2250, which is used to extract video frames.

[0395] Figure 24 The following shows a method flow chart of a video card picture generation method according to an embodiment of the present application.

[0396] In one embodiment, Figure 23 the shown video editing system executes Figure 24 the shown steps to implement the generation of video card pictures.

[0397] S2300, obtain the video to be edited (the first video) and the video title of the video to be edited.

[0398] S2310, the similarity retrieval module 2210 retrieves similar videos (M) from the video library / sample library based on the first video and the video title of the first video.

[0399] S2311, the similarity retrieval module 2210 filters out the top-K similar videos with the highest number of likes from the retrieved M similar videos based on the number of likes of the M similar videos.

[0400] S2301, display the K videos and the video card pictures of the K videos screened out in S110 to the user.

[0401] S2302, receive the user's selected operation, determine the N videos selected by the user according to the user's selected operation, and record the video card images of the N videos as the third set of card images.

[0402] S2330, the template extraction module 2220 extracts the common features from the images in the third set of card images to generate a second image template.

[0403] S2340, the frame extraction module 2250 extracts video frames with a picture dimension matching that of the second image template from the first video based on the second image template, and records them as the third set of alternative pictures.

[0404] S2350, the image cropping module 2230 edits the pictures in the third set of alternative pictures according to the second image template to generate alternative video card images, which are recorded as the fifth set of alternative video card images.

[0405] S2312, the similarity retrieval module 2210 determines the user's attention time interval (the fourth user attention time interval) of the first video, and determines the user's attention space region (the fourth user attention space region) of the first video according to the video editing information of the K videos (the K videos screened in S2311).

[0406] S2341, the frame extraction module 2250 extracts video frames from the fourth user attention time interval of the first video and records them as the fourth set of alternative pictures.

[0407] S2351, the image cropping module 2230 edits the pictures in the fourth set of alternative pictures according to the second image template and the fourth user attention space region of the first video to generate alternative video card images, which are recorded as the sixth set of alternative video card images.

[0408] S2313, the similarity retrieval module 2210 determines the user's attention time interval (the fifth user attention time interval) of the first video, and determines the user's attention space region (the fifth user attention space region) of the first video according to the video editing information of the first video.

[0409] S2342, the frame extraction module 2250 extracts video frames from the fifth user attention time interval of the first video and records them as the fifth set of alternative pictures.

[0410] S2352, the image cropping module 2230 edits the pictures in the fifth set of alternative pictures according to the second image template and the fifth user attention space region of the first video to generate alternative video card images, which are recorded as the seventh set of alternative video card images.

[0411] S2360, the sorting module 2240 uses the combination of the fifth alternative video card graph set, the sixth alternative video card graph set, and the seventh alternative video card graph set as the fourth card graph set, calculates the relevance between the pictures in the fourth card graph set and the video title of the first video, calculates the similarity between the pictures in the fourth card graph set and the third card graph set, and sorts the pictures in the second card graph set according to the calculation results of the relevance and similarity.

[0412] S2303, according to the sorting result of the sorting module 2240, display the pictures in the fourth card graph set to the video editor.

[0413] S2304, according to the selection operation of the video editor, determine the picture as the video card graph editing result.

[0414] Further, in an embodiment, in S2303, based on the second image template or the card graph in the third card graph set, superimpose editable graphic attachments (such as, word art, special effects, device components, etc.) on the pictures in the fourth card graph set displayed to the video editor for the video editor to perform secondary editing.

[0415] Further, in some application scenarios, the video is not for public release but for private viewing by the user. For example, the video saved in the local album or cloud album. For such videos, their video card graphs do not need to attract video viewers to click, but are used to assist the video viewers to recall the content of the video.

[0416] For the above application scenarios, an embodiment of the present application provides a method for generating a video card graph. The method provided by the embodiment of the present application is applied to an electronic device. In the method provided by the embodiment of the present application, when the electronic device generates the video card graph of the second video, it does not need to refer to other publicly released videos, but generates the video card graph of the second video according to the user interaction behavior record of the second video.

[0417] Figure 25 Shown is the flowchart of the method for generating a video card graph according to an embodiment of the present application.

[0418] In an embodiment, the electronic device executes Figure 25 the steps shown to implement the generation of the video card graph.

[0419] S2400, record the user interaction behavior of the second video and generate a user interaction behavior record.

[0420] Specifically, in an embodiment, the user interaction behavior includes the user's playing behavior of the video and the user's editing behavior of the video.

[0421] S2410. Determine the user attention time interval and the user attention spatial region of the second video according to the user interaction behavior record.

[0422] Specifically, for the user's video editing behavior, determine the user attention time interval and the user attention spatial region of the second video according to the video segment targeted by the user editing behavior and the image region of the video frame targeted by the user editing behavior.

[0423] S2420. Extract video frames from the video segment corresponding to the user attention time interval, and record them as the candidate picture set.

[0424] S2430. Crop the pictures in the candidate picture set according to the user attention spatial region and the preset video card layout to generate candidate video cards, which are recorded as the candidate video card set.

[0425] S2440. Select one or more candidate video cards from the card set as the video card of the second video.

[0426] According to the method of the embodiment of the present application, generating candidate video cards for a video based on the user interaction behavior of the video can make the candidate video cards associated with the user's historical operations on the video, so that the candidate video cards can better assist the user in recalling the video content.

[0427] Referring to S150, in S2440, various different methods can be used to select the video card of the second video from the card set.

[0428] For example, in one implementation of S2440, display the pictures in the card set to the user. The user selects the picture to be used as the video card from the card set.

[0429] In another implementation of S2440, sort the pictures in the card set according to a preset rule, and use the N pictures ranked at the top as the video cards, where N is a preset integer greater than or equal to one.

[0430] In another implementation of S2440, sort the pictures in the card set according to a preset rule, and display the pictures in the card set to the user according to the sorting result (for example, the pictures ranked in the front are preferentially displayed to the user). The user selects the picture to be used as the video card.

[0431] For example, in one embodiment, perform a comprehensive sorting according to three dimensions: user preference degree (such as the frames with the most frequent user zooming, pausing, and repeated viewing), image quality (such as the picture aesthetics scoring model), and semantic similarity (such as the distance between the picture embedding vector and the video average embedding vector).

[0432] Further, after S2440, continue to record the user interaction behavior of the second video. When the new view count and / or new edit count meet the preset threshold, determine whether the user's attention time interval and / or user's attention spatial region of the second video have changed. If the user's attention time interval and / or user's attention spatial region have changed, repeat the execution of S2420 to S2440 to update the video card diagram of the second video.

[0433] In the description of the embodiments of the present application, for the convenience of description, when describing the device, it is divided into various modules according to functions and described separately. The division of each module is only a logical function division. When implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0434] Specifically, when the device proposed in the embodiments of the present application is actually implemented, it can be fully or partially integrated into a physical entity, or physically separated. And these modules can all be implemented in the form of software called by a processing element; they can also all be implemented in the form of hardware; or some modules can be implemented in the form of software called by a processing element, and some modules can be implemented in the form of hardware. For example, the detection module can be a separately established processing element or can be integrated in a certain chip of the electronic device. The implementation of other modules is similar. In addition, all or part of these modules can be integrated together or can be independently implemented. In the implementation process, each step of the above method or each of the above modules can be completed by the integrated logic circuit in the processor element or the instruction in the form of software.

[0435] For example, the above modules can be one or more integrated circuits configured to implement the above method, such as: one or more Application Specific Integrated Circuits (ASICs), or, one or more Digital Singnal Processors (DSPs), or, one or more Field Programmable Gate Arrays (FPGAs), etc. Again, these modules can be integrated together and implemented in the form of a System-On-a-Chip (SOC).

[0436] An embodiment of the present application also proposes an electronic device.

[0437] Figure 26 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present application.

[0438] Such as Figure 26As shown, the electronic device 2500 includes a memory 2502 for storing computer program instructions and a processor 2501 for executing the program instructions. When the computer program instructions are executed by the processor 2501, the electronic device 2500 is triggered to execute the method steps described in the embodiments of the present application.

[0439] Specifically, in an embodiment of the present application, the above one or more computer programs are stored in the memory 2502. The above one or more computer programs include instructions that, when executed by the electronic device 2500, cause the electronic device 2500 to execute the method steps described in the embodiments of the present application.

[0440] It can be understood that the structural description of the electronic device 2500 in the embodiments of the present application does not constitute a specific limitation on the electronic device 2500. In other embodiments of the present application, the electronic device 2500 may include components other than the processor 2501 and the memory 2502.

[0441] The processor 2501 may be a system-on-chip (SOC). The processor 2501 may include a central processing unit (CPU) and may further include other types of processors.

[0442] The processors related to the processor 2501 may include, for example, a CPU, a DSP, a microcontroller, or a digital signal processor, and may further include a GPU, an embedded neural-network processor (NPU), and an image signal processor (ISP). The processor may also include necessary hardware accelerators or logic processing hardware circuits, such as an ASIC, or one or more integrated circuits for controlling the execution of the technical solution program of the present application. In addition, the processor may have the function of operating one or more software programs, and the software programs may be stored in a storage medium.

[0443] The processor 2501 may include one or more processing units. For example, the processor may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent components or integrated in one or more processors. In some embodiments, the electronic device 2500 may also include one or more processors 2501. Among them, the controller can generate operation control signals according to the instruction operation code and timing signals to complete the control of fetching and executing instructions.

[0444] In some embodiments, the processor 2501 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. Among them, the USB interface is an interface that conforms to the USB standard specification, and specifically may be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge the electronic device or to transfer data between the electronic device and peripheral devices.

[0445] The electronic device 2500 may further include an external memory interface, which is used to connect an external memory card, such as a Micro SD card, to implement the storage capacity expansion of the electronic device. The external memory card communicates with the processor 2501 through the external memory interface to achieve the data storage function. For example, files such as music and videos are saved in the external memory card.

[0446] The memory 2502 may include a code storage area and a data storage area. Among them, the code storage area can store the operating system. The data storage area can store data created during the use of the electronic device 2500, etc. In addition, the memory 2502 may include high-speed random access memory, and may also include non-volatile memory, such as one or more disk storage components, flash memory components, universal flash storage (UFS), etc.

[0447] The memory 2502 can be a read-only memory (ROM), other types of static storage devices that can store static information and instructions, random access memory (RAM), or other types of dynamic storage devices that can store information and instructions. It can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, or other magnetic storage devices. Or it can also be any computer-readable medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer.

[0448] The processor 2501 and the memory 2502 can be integrated into a processing device, and more commonly are independent components of each other.

[0449] Optionally, the devices, apparatuses, and modules described in the embodiments of the present application can be specifically implemented by computer chips or entities, or by products with certain functions.

[0450] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, an apparatus, or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media that contain computer-usable program code.

[0451] In several embodiments provided by the present application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application.

[0452] Specifically, in one embodiment of the present application, a computer-readable storage medium is further provided. A computer program is stored in this computer-readable storage medium. When it runs on a computer, it causes the computer to execute the method provided by the embodiment of the present application.

[0453] One embodiment of the present application further provides a computer program product. This computer program product includes a computer program. When it runs on a computer, it causes the computer to execute the method provided by the embodiment of the present application.

[0454] The embodiments in the present application are described with reference to the flowcharts and / or block diagrams of methods, devices (apparatuses), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for realizing the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0455] These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing devices to work in a specific manner, so that the instructions stored in this computer-readable memory generate a manufactured product including an instruction device, and this instruction device realizes the functions specified in one Figure 1 process or multiple processes and / or blocks Figure 1 or multiple blocks.

[0456] These computer program instructions can also be loaded onto a computer or other programmable data processing devices, so that a series of operation steps are executed on the computer or other programmable devices to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable devices provide for realizing the functions in the process Figure 1One or more processes and / or blocks Figure 1 Steps of functions specified in one or more blocks.

[0457] It should also be noted that in the embodiments of the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent the case where A exists alone, A and B exist simultaneously, or B exists alone. Where A and B may be singular or plural. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "At least one of the following" and its similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c may represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, and c may be single or multiple.

[0458] In the embodiments of the present application, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed, or further includes elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the existence of additional identical elements in the process, method, commodity or device comprising the said element.

[0459] The present application may be described in the general context of computer-executable instructions executable by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media including storage devices.

[0460] The various embodiments in the present application are described in a progressive manner. For the same or similar parts among the various embodiments, reference may be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the device embodiments, since they are basically similar to the method embodiments, they are described relatively simply, and for the relevant parts, reference may be made to the description of the method embodiments.

[0461] Those of ordinary skill in the art can realize that the various units and algorithm steps described in the embodiments of this application can be implemented by a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0462] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, devices, and units can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0463] As described above, the above is only the specific implementation manner of this application. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all should be covered by the protection scope of this application. The protection scope of this application shall be subject to the protection scope of the claimed rights.

Claims

1. A method for generating a video card diagram, characterized in that, The method is applied to an electronic device, and the method includes: Obtaining M videos whose video content is similar to that of a first video, where M is an integer greater than or equal to 1; According to the user feedback information of the M videos, screening K videos from the M videos, where K is an integer greater than or equal to 1 and less than or equal to M, and the user feedback information includes the number of video likes and / or video click-through rates; Obtaining the video card images of the K videos; Generating multiple alternative video card images according to the commonalities of the video card images of the K videos.

2. The method according to claim 1, wherein The screening K videos from the M videos according to the user feedback information of the M videos includes: Screening the K videos with the highest video click-through rates from the M videos.

3. The method according to claim 1 or 2, characterized in that, The generating multiple alternative video card images according to the commonalities of the video card images of the K videos includes: Generating a first image template according to the commonalities of the video card images of the K videos; Based on the first image template, generating multiple first alternative video card images that match the first image template.

4. The method according to claim 3, wherein The generating multiple first alternative video card images that match the first image template based on the first image template includes: Obtaining similar images of the video card images of the K videos; Editing the similar images of the video card images of the K videos according to the first image template to generate the first alternative video card images.

5. The method according to claim 3, characterized in that The generating multiple first alternative video card images that match the first image template based on the first image template further includes: Extracting a first video frame from the first video; Editing the first video frame according to the first image template to generate the first alternative video card images.

6. The method according to claim 5, wherein The extracting a first video frame from the first video includes: Determining the user attention time intervals of the M videos or the K videos according to the user interaction behavior records of the M videos or the K videos; Determining the user attention time interval of the first video according to the user attention time intervals of the M videos or the K videos; Extracting the first video frame from the video segment of the first video corresponding to the user attention time interval of the first video.

7. The method according to claim 6, characterized in that, The determining the user attention time interval of the first video according to the user attention time intervals of the M videos or the K videos includes: Retrieving similar video segments in the first video according to the video segments corresponding to the user attention time intervals of the M videos or the K videos, and the time interval corresponding to the similar video segment is the user attention time interval of the first video.

8. The method according to claim 5, wherein The extracting a first video frame from the first video further includes: Determining the user attention time interval of the first video according to the user interaction behavior record of the first video and the user interaction behavior records of the videos, where the user interaction behavior records of the videos include the user interaction behavior records of the M videos or the K videos; Extracting the first video frame from the video segment of the first video corresponding to the user attention time interval of the first video.

9. The method according to claim 5, characterized in that, Generating the first alternative video card image further includes: Determining the user attention spatial region of the M videos or the K videos according to the user interaction behavior records of the M videos or the K videos; Determining the user attention spatial region of the first video according to the user attention spatial region of the M videos or the K videos; Cropping the first video frame according to the user attention spatial region of the first video.

10. The method according to claim 5, characterized in that Generating the first alternative video card image further includes: Determining the user attention spatial region of the first video according to the user interaction behavior record of the first video and the user interaction behavior records of the videos, where the user interaction behavior records of the videos include the user interaction behavior records of the M videos or the K videos; Cropping the first video frame according to the user attention spatial region of the first video.

11. The method according to claim 3, wherein Generating multiple first alternative video card images that match the first image template includes: Using an artificial intelligence automatic creation and generation content model to generate the first alternative video card image based on the first image template and the video text description of the first video.

12. The method according to claim 11, wherein The video text description of the first video includes the alternative video title of the first video and / or the original video title of the first video.

13. The method according to claim 12, characterized in that, The method further includes: Obtaining the original video title of the first video; Rewriting the original video title of the first video based on the video titles of the M videos or the K videos to generate multiple alternative video titles.

14. The method according to claim 3, characterized in that, Generating multiple first alternative video card images that match the first image template further includes: Using an artificial intelligence automatic creation and generation content model to generate the first alternative video card image based on the semantic segmentation map and the video text description of the first video, where: The semantic segmentation map is a semantic segmentation map obtained by performing semantic segmentation on the video card images of the K videos, and / or the first alternative video card image set, and / or the second alternative video card image set; The first alternative video card image set is a set of images generated by editing similar images of the video card images of the K videos according to the first image template; The second alternative video card image set is a set of images generated by editing the video frames extracted from the first video according to the first image template.

15. The method according to claim 1, wherein Selecting K videos from the M videos according to the user feedback information of the M videos includes: Selecting the K videos with the highest number of video likes from the M videos.

16. The method according to claim 15, wherein Generating multiple alternative video card images according to the commonalities of the video card images of the K videos includes: Determining the video card images selected by the video editor among the video card images of the K videos; Generating a second image template according to the commonalities of the video card images selected by the video editor; Generating multiple second alternative video card images that match the second image template based on the second image template.

17. The method according to claim 16, wherein Generating multiple second alternative video card images that match the second image template, including: Extracting second video frames from the first video; Editing the second video frames according to the second image template to generate the second alternative video card images.

18. The method according to claim 17, wherein The extracting second video frames from the first video includes: Extracting second video frames from the first video whose picture dimensions match the second image template.

19. The method according to claim 17, wherein The extracting second video frames from the first video includes: Determining a user attention time interval of the first video according to the video editing information of the M videos or the K videos; Extracting the second video frames from the video segment of the first video corresponding to the user attention time interval of the first video.

20. The method according to claim 19, wherein The generating the second alternative video card images further includes: Determining a user attention spatial region of the first video according to the video editing information of the M videos or the K videos; Cropping the second video frames according to the user attention spatial region of the first video.

21. The method according to claim 17, characterized in that, The extracting second video frames from the first video includes: Determining a user attention time interval of the first video according to the video editing information of the first video; Extracting the second video frames from the video segment of the first video corresponding to the user attention time interval of the first video.

22. The method according to claim 21, wherein The generating the second alternative video card images further includes: Determining a user attention spatial region of the first video according to the video editing information of the first video; Cropping the second video frames according to the user attention spatial region of the first video.

23. The method according to any one of claims 1-22, characterized in that, The method further includes: Selecting one or more alternative video card images from the multiple alternative video card images as the video card images of the first video.

24. The method according to claim 23, wherein The selecting one or more alternative video card images from the multiple alternative video card images as the video card images of the first video includes: Sorting the multiple alternative video card images according to a preset rule, and presenting the multiple alternative video card images to the video publisher according to the sorting result; Determining the alternative video card images selected by the video publisher from the multiple alternative video card images, and using the alternative video card images selected by the video publisher as the video card images of the first video.

25. The method according to claim 24, characterized in that, Sorting the multiple alternative video card images according to a preset rule includes: Sorting the pictures in the multiple alternative video card images according to the similarity calculation result and / or the relevance calculation result, where: The similarity calculation result includes the similarity between the multiple alternative video card images and the video card images of the K videos, or the similarity between the multiple alternative video card images and the selected video card images among the video card images of the K videos; The relevance calculation result includes the relevance between the multiple alternative video card images and the video text description of the first video.

26. The method according to claim 25, wherein Sorting the pictures in the multiple alternative video card images according to the similarity calculation result and the relevance calculation result includes: Calculate the similarity between the multiple alternative video card graphs and the video card graphs of the K videos, and obtain the similarity calculation result; Calculate the relevance between the multiple alternative video card graphs and the video titles of the first video, and obtain the relevance calculation result; Sort the multiple alternative video card graphs according to the similarity calculation result and / or the relevance calculation result.

27. The method according to claim 25, characterized in that, Sort the pictures in the multiple alternative video card graphs according to the similarity calculation result and the relevance calculation result, including: Display the video text description of the first video, where the video text description includes multiple alternative video titles of the first video; Determine the alternative video title selected by the video publisher; Calculate the similarity between the multiple alternative video card graphs and the video card graphs of the K videos, and obtain the similarity calculation result; Calculate the relevance between the multiple alternative video card graphs and the alternative video title selected by the video publisher, and obtain the relevance calculation result; Sort the multiple alternative video card graphs according to the similarity calculation result and / or the relevance calculation result.

28. A method for generating a video card diagram, characterized in that, The method is applied to an electronic device, and the method includes: Determine the user attention time interval and the user attention spatial region of the second video according to the user interaction behavior record of the second video; Extract the fifth video frame from the video segment of the second video corresponding to the user attention time interval of the second video; Crop the fifth video frame according to the user attention spatial region of the second video to generate a third alternative video card graph.

29. An electronic device, characterized in that, The electronic device includes a memory for storing computer program instructions and a processor for executing the computer program instructions. When the computer program instructions are executed by the processor, the electronic device is triggered to execute the method steps as described in any one of claims 1-27 or claim 28.

30. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When it runs on a computer, the computer is caused to execute the method as described in any one of claims 1-27 or claim 28.

Citation Information

Patent Citations

  • Method and apparatus for determining video file cover, and computer readable storage medium

    CN107707967A

  • Video processing method, storage medium and processor

    CN113382301A

  • Content generation method and device, electronic equipment and storage medium

    CN117221622A

  • Content summarization

    US10090020B1