Video card image generating method and electronic device
By obtaining videos similar to video content, filtering and generating video card pictures based on user feedback information, the problem of insufficient attractiveness of the card pictures selected by the video publisher is solved, and a higher click rate is achieved.
Patent Information
- Application Number
- PCT/CN2024/123630
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2024-10-09
- Publication Date
- 2025-08-07
AI Technical Summary
The video card images selected by the video publisher cannot effectively attract video viewers to click, resulting in poor video distribution results.
By obtaining M videos similar to the first video content, filtering K videos based on user feedback information such as the number of video likes and click-through rates, multiple alternative video card pictures are generated, and content models and semantic segmentation technology are used to automatically create, and video card pictures are generated based on video text description.
The generated video card pictures are more in line with the current application scenario requirements of the video and can attract video viewers to click more effectively.
Smart Images

Figure CN2024123630_07082025_PF_FP_ABST
Abstract
Description
Video card image generation method and electronic device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on January 30, 2024, with application number 2024101351333 and application name “A method and electronic device for generating a video card image”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present application relates to the field of computer technology, and in particular to a method for generating a video card image and an electronic device. Background Art
[0003] When a video publisher uploads a video for distribution, the video is generally accompanied by an exquisite picture (hereinafter referred to as a card picture) to attract video viewers to click on the video.
[0004] Typically, video card images are uploaded by the video publisher themselves. These images can be images selected by the publisher that are relevant to the video, or they can be images of a specific frame within the video. However, publishers often choose video card images based solely on their own understanding of the video content. These selected video card images fail to effectively attract viewers, resulting in poor video distribution.
[0005] Therefore, a video card image generation method is needed to obtain a video card image that effectively attracts video viewers to click.
[0006] Summary of the Invention
[0007] In order to solve the problem of how to obtain a video card image that effectively attracts video viewers to click, the present application provides a video card image generation method and an electronic device, and the present application also provides a computer-readable storage medium.
[0008] The embodiments of this application adopt the following technical solutions:
[0009] In a first aspect, the present application provides a method for generating a video card image, the method being applied to an electronic device, the method comprising:
[0010] Obtain M videos with similar video content to the first video, where M is an integer greater than or equal to 1;
[0011] Filtering K videos from the M videos based on user feedback information of the M videos, where K is an integer greater than or equal to 1 and less than or equal to M, wherein the user feedback information includes the number of video likes and / or the video click-through rate;
[0012] Obtain video card images of the K videos;
[0013] Based on the commonalities of the video card images of the K videos, a plurality of candidate video card images are generated.
[0014] According to the method of the first aspect, similar videos of the current video can be referenced, and a video card image of the current video can be generated based on user feedback information of similar videos, so that the video card image is more in line with the application scenario requirements of the current video.
[0015] In an implementation of the first aspect, screening K videos from the M videos based on user feedback information of the M videos includes:
[0016] From the M videos, select K videos with the highest video click rates.
[0017] In an implementation of the first aspect, generating a plurality of candidate video card images based on commonalities of the video card images of the K videos includes:
[0018] generating a first image template according to the commonalities of the video card images of the K videos;
[0019] Based on the first image template, a plurality of first candidate video card images matching the first image template are generated.
[0020] According to the above-mentioned implementation method, an image template is generated based on the common characteristics of the video card images of videos with high video click-through rates, and an alternative video card image is generated based on the image template, so that the alternative video card image can better attract video viewers to click.
[0021] In an implementation of the first aspect, generating, based on the first image template, a plurality of first candidate video card images that match the first image template includes:
[0022] Obtain similar images of the video card images of the K videos;
[0023] Based on the first image template, similar images of the video card images of the K videos are edited to generate the first candidate video card image.
[0024] According to the above-mentioned implementation method, the existing pictures in the picture library are used as the picture source, and the pictures in the picture source are screened and edited according to the video card pictures of videos with high video click-through rates to generate alternative video card pictures, which can effectively limit the source range of the alternative video card pictures.
[0025] In an implementation of the first aspect, generating, based on the first image template, a plurality of first candidate video card images that match the first image template further includes:
[0026] Extracting a first video frame from the first video;
[0027] The first video frame is edited according to the first image template to generate the first candidate video card image.
[0028] According to the above-mentioned implementation method, the video frame of the current video is used as the image source, and the images in the image source are screened and edited according to the video card images of videos with high video click-through rates to generate alternative video card images, which can ensure the correlation between the alternative video card images and the current video.
[0029] In an implementation of the first aspect, extracting a first video frame from the first video includes:
[0030] Determining a user attention time interval for the M videos or the K videos based on user interaction behavior records for the M videos or the K videos;
[0031] Determining a user attention time interval of the first video according to the user attention time intervals of the M videos or the K videos;
[0032] The first video frame is extracted from a video segment of the first video corresponding to a user attention time interval of the first video.
[0033] According to the above implementation method, video frames are extracted according to the time interval of user attention, which can effectively ensure that the extracted video frames are more attractive to video viewers to click.
[0034] In an implementation of the first aspect, determining the user attention time interval of the first video according to the user attention time intervals of the M videos or the K videos includes:
[0035] According to the video segments corresponding to the user's attention time interval of the M videos or the K videos, similar video segments are retrieved in the first video, and the time interval corresponding to the similar video segments is the user's attention time interval of the first video.
[0036] In an implementation of the first aspect, extracting the first video frame from the first video further includes:
[0037] Determining a user attention time interval of the first video based on the user interaction behavior record of the first video and the user interaction behavior record of the video, wherein the user interaction behavior record of the video includes the user interaction behavior records of the M videos or the K videos;
[0038] The first video frame is extracted from a video segment of the first video corresponding to a user attention time interval of the first video.
[0039] In an implementation of the first aspect, generating the first candidate video card image further includes:
[0040] Determining, based on the user interaction behavior records of the M videos or the K videos, a user attention space area of the M videos or the K videos;
[0041] Determining a user attention space area of the first video according to the user attention space areas of the M videos or the K videos;
[0042] The pictures in the second candidate picture set are cropped according to the user's attention space area of the first video.
[0043] According to the above implementation method, cropping the video frame according to the user's attention space area can effectively ensure that the alternative video card image is more attractive to video viewers to click.
[0044] In an implementation of the first aspect, generating the first candidate video card image further includes:
[0045] Determining a user attention space area of the first video according to the user interaction behavior record of the first video and the user interaction behavior record of the video, wherein the user interaction behavior record of the video includes the user interaction behavior records of the M videos or the K videos;
[0046] The pictures in the second candidate picture set are cropped according to the user's attention space area of the first video.
[0047] In an implementation of the first aspect, generating, based on the first image template, a plurality of first candidate video card images that match the first image template includes:
[0048] Artificial intelligence is used to automatically create a content model, and the first alternative video card image is generated based on the first image template and the video text description of the first video.
[0049] In an implementation of the first aspect, the video text description of the first video includes an alternative video title of the first video and / or an original video title of the first video.
[0050] In an implementation of the first aspect, the method further includes:
[0051] Obtaining the original video title of the first video;
[0052] Based on the video titles of the M videos or the K videos, the original video title of the first video is rewritten to generate a plurality of candidate video titles.
[0053] In an implementation of the first aspect, generating, based on the first image template, a plurality of first candidate video card images that match the first image template further includes:
[0054] Using artificial intelligence to automatically create and generate a content model, based on the semantic segmentation map and the video text description of the first video, a card image of the first candidate video is generated, wherein:
[0055] The semantic segmentation map is a semantic segmentation map obtained by semantically segmenting the video card maps of the K videos, and / or the first candidate video card map set, and / or the second candidate video card map set;
[0056] The first candidate video card image set is a set of images generated by editing similar images of the video card images of the K videos according to the first image template;
[0057] The second candidate video card image set is a set of images generated by editing video frames extracted from the first video according to the first image template.
[0058] In an implementation of the first aspect, screening K videos from the M videos based on user feedback information of the M videos includes:
[0059] From the M videos, select K videos with the highest number of likes.
[0060] In an implementation of the first aspect, generating a plurality of candidate video card images based on commonalities of the video card images of the K videos includes:
[0061] Determine the video card image selected by the video editor among the video card images of the K videos;
[0062] generating a second image template based on the commonalities of the video card images selected by the video editor;
[0063] Based on the second image template, a plurality of second candidate video card images matching the second image template are generated.
[0064] According to the above-mentioned implementation method, a video card image of the video to be edited can be generated based on the video that the user wants to imitate among the videos with high video likes, so that the video card image of the video to be edited has a better visual effect.
[0065] In an implementation of the first aspect, generating, based on the second image template, a plurality of second candidate video card images that match the second image template includes:
[0066] Extracting a second video frame from the first video;
[0067] The second video frame is edited according to the second image template to generate the second candidate video card image.
[0068] In an implementation of the first aspect, extracting the second video frame from the first video includes:
[0069] A second video frame whose picture dimension matches the second image template is extracted from the first video.
[0070] In an implementation of the first aspect, extracting the second video frame from the first video includes:
[0071] Determining a user attention time interval of the first video according to the video editing information of the M videos or the K videos;
[0072] The second video frame is extracted from a video segment of the first video corresponding to a user attention time interval of the first video.
[0073] According to the above-mentioned implementation method, the user attention time interval of the current video is determined based on videos with high video like numbers, and video frames are extracted based on the user attention time interval, so that the alternative video card image can be closer to the image effect of the user attention video clip of the video with high video like numbers.
[0074] In an implementation of the first aspect, generating the second candidate video card image further includes:
[0075] Determining a user focused spatial area of the first video according to the video editing information of the M videos or the K videos;
[0076] The second video frame is cropped according to a user-focused spatial area of the first video.
[0077] In an implementation of the first aspect, extracting the second video frame from the first video includes:
[0078] determining a user attention time interval of the first video according to the video editing information of the first video;
[0079] The second video frame is extracted from a video segment of the first video corresponding to a user attention time interval of the first video.
[0080] According to the above implementation method, the user's attention time interval is determined according to the video editing information of the current video, and video frames are extracted according to the user's attention time interval, so that the extracted video frames can better meet the needs of the video editor.
[0081] In an implementation of the first aspect, generating the second candidate video card image further includes:
[0082] determining a user focused spatial area of the first video according to the video editing information of the first video;
[0083] The second video frame is cropped according to a user-focused spatial area of the first video.
[0084] In an implementation of the first aspect, the method further includes:
[0085] One or more candidate video card images are selected from the multiple candidate video card images as the video card images of the first video.
[0086] In an implementation of the first aspect, selecting one or more candidate video card images from the plurality of candidate video card images as the video card image of the first video includes:
[0087] Sorting the plurality of candidate video card images according to a preset rule, and displaying the plurality of candidate video card images to the video publisher according to the sorting result;
[0088] Determine the alternative video card image selected by the video publisher among the multiple alternative video card images, and use the alternative video card image selected by the video publisher as the video card image of the first video.
[0089] In an implementation of the first aspect, sorting the plurality of candidate video card images according to a preset rule includes:
[0090] The pictures in the plurality of candidate video card images are sorted according to the similarity calculation result and / or the correlation calculation result, wherein:
[0091] The similarity calculation result includes the similarity between the multiple candidate video card images and the video card images of the K videos, or the similarity between the multiple candidate video card images and a selected video card image from the video card images of the K videos;
[0092] The relevance calculation result includes the relevance between the multiple candidate video card images and the video text description of the first video.
[0093] In an implementation of the first aspect, sorting the images in the plurality of candidate video card images according to the similarity calculation result and the correlation calculation result includes:
[0094] Calculating similarities between the plurality of candidate video card images and the video card images of the K videos, and obtaining the similarity calculation result;
[0095] Calculate the relevance between the video title of the first video and the plurality of candidate video card images, and obtain the relevance calculation result;
[0096] The plurality of candidate video card images are sorted according to the similarity calculation result and / or the correlation calculation result.
[0097] In an implementation of the first aspect, sorting the images in the plurality of candidate video card images according to the similarity calculation result and the correlation calculation result includes:
[0098] Displaying a video text description of the first video, wherein the video text description includes a plurality of alternative video titles of the first video;
[0099] Determine the alternative video titles selected by the video publisher;
[0100] Calculating similarities between the multiple candidate video card images and the video card images of the K videos, and obtaining the similarity calculation result;
[0101] Calculating the relevance between the plurality of candidate video card images and the candidate video titles selected by the video publisher, and obtaining the relevance calculation result;
[0102] The plurality of candidate video card images are sorted according to the similarity calculation result and / or the correlation calculation result.
[0103] In a second aspect, the present application provides a method for generating a video card image, which is applied to an electronic device and includes:
[0104] Determining a user attention time interval and a user attention spatial area of the second video based on a user interaction behavior record of the second video;
[0105] extracting a fifth video frame from a video segment of the second video corresponding to a user attention time interval of the second video;
[0106] The fifth video frame is cropped according to the user's attention space area of the second video to generate a third candidate video card image.
[0107] According to the method of the second aspect, an alternative video card image is generated for the video based on the user's interactive behavior with the video, so that the alternative video card image can be associated with the user's historical operations on the video, so that the alternative video card image can better assist the user in recalling the video content.
[0108] In a third aspect, the present application provides an electronic device, comprising a memory for storing computer program instructions and a processor for executing computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to execute the method steps described in the first aspect or the second aspect.
[0109] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method described in the first aspect or the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0110] FIG1 is a flow chart of a method for generating a video card image according to an embodiment of the present application;
[0111] FIG2 is a flow chart of a method for generating a video card image according to an embodiment of the present application;
[0112] FIG3 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0113] FIG4 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0114] FIG5 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0115] FIG6 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0116] FIG7 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0117] FIG8 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0118] FIG9 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0119] FIG10 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0120] FIG11 is a simplified structural diagram of a video card image generation system according to an embodiment of the present application;
[0121] FIG12 is a flow chart of a method for generating a video card image according to an embodiment of the present application;
[0122] FIG13 is a schematic diagram showing an interface display effect according to an embodiment of the present application;
[0123] FIG14 is a schematic diagram showing an interface display effect according to an embodiment of the present application;
[0124] FIG15 is a flow chart of a method for generating a video card image according to an embodiment of the present application;
[0125] FIG16 is a flow chart of a method for generating a video card image according to an embodiment of the present application;
[0126] FIG17 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0127] FIG18 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0128] FIG19 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0129] FIG20 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0130] FIG21 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0131] FIG22 is a partial flow chart of a method for generating a video card image according to an embodiment of the present application;
[0132] FIG23 is a simplified structural diagram of a video editing system according to an embodiment of the present application;
[0133] FIG24 is a flow chart of a method for generating a video card image according to an embodiment of the present application;
[0134] FIG25 is a flow chart of a method for generating a video card image according to an embodiment of the present application;
[0135] FIG26 is a schematic structural diagram of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION
[0136] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.
[0137] The terms used in the implementation section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.
[0138] In order to obtain a video card image that effectively attracts video viewers to click, an embodiment of the present application provides a video card image generation method. The method provided in the embodiment of the present application is applied to an electronic device. In the method provided in the embodiment of the present application, the electronic device obtains similar videos of a first video; based on the current application scenario, the electronic device filters similar videos according to user feedback information of similar videos (for example, video click-through rate, number of video likes, etc.), wherein the filtered similar videos are more in line with the application requirements of the current application scenario than the videos that are not filtered out; and a video card image of the first video is generated based on the filtered similar videos.
[0139] In the embodiments of this specification, electronic devices that execute the video card image generation method process include but are not limited to mobile phones, tablet computers, laptops, wearable smart devices, desktop computers, servers, etc.
[0140] Optionally, in one implementation, the electronic device that executes the video card image generation method process may be a terminal device of the video publisher, such as a mobile phone, tablet computer, laptop computer, desktop computer, etc. In another implementation, the electronic device that executes the video card image generation method process may be other electronic devices and / or cloud servers connected to the terminal device of the video publisher. In another implementation, the video card image generation method process may be partially executed by the terminal device of the video publisher and partially executed by other electronic devices and / or cloud servers connected to the terminal device of the video publisher.
[0141] FIG1 is a flow chart of a method for generating a video card image according to an embodiment of the present application.
[0142] In one embodiment, the electronic device executes the following process shown in FIG. 1 to generate a video card image of the first video.
[0143] S101, obtaining M videos having video contents similar to the first video, where M is an integer greater than or equal to 1.
[0144] S102, based on user feedback information of M videos, select K videos from the M videos, where K is an integer greater than or equal to 1 and less than or equal to M. The user feedback information includes the number of video likes and / or the video click-through rate.
[0145] S103: Obtain video card images of K videos, which are recorded as a first card image set.
[0146] S104: Generate multiple candidate video card images based on the commonalities of the video card images in the first card image set (for example, adult faces, full-body photos of beautiful women, scenery, text labels, picture special effects, etc.), which are recorded as the second card image set.
[0147] S105: Select one or more candidate video card images from the second card image set as the video card images of the first video.
[0148] According to the method of the embodiment of the present application, similar videos of the current video can be referred to, and a video card image of the current video can be generated based on user feedback information of the similar videos, so that the video card image is more in line with the application scenario requirements of the current video.
[0149] The video card image generation method provided in the embodiment of the present application can be applied to video sharing application scenarios.
[0150] FIG2 is a flow chart of a method for generating a video card image according to an embodiment of the present application.
[0151] In one embodiment, in a video sharing application scenario, the electronic device executes the following process shown in FIG. 2 to generate a video card image of the first video.
[0152] In the embodiment shown in FIG2 , the first video is a video that a video publisher desires to publish to a video sharing platform for video sharing. Video viewers have not yet clicked to view the first video. No user feedback or user interaction records have been recorded for the first video.
[0153] S100: Obtain M videos having video content similar to the first video, where M is an integer greater than or equal to 1.
[0154] Specifically, in S100, M videos are selected from videos that have been posted on the video sharing platform and can be clicked and viewed by video viewers. User feedback information of the M videos includes video click-through rates.
[0155] S110, from the M videos obtained in S100, screen the K (Top-K) videos with the highest click-through rate (CTR). The K videos are recorded as a similar video set, where K is an integer greater than or equal to 1 and less than or equal to M.
[0156] S120: Obtain video card images of all videos in the similar video set, which are recorded as a first card image set.
[0157] S130, generating a first image template according to the commonalities of the video card images in the first card image set (for example, adult faces, full-body photos of beautiful women, landscapes, text labels, picture special effects, etc.).
[0158] S140, based on the first image template generated in S130, generate a plurality of candidate video card images that match the first image template, which are recorded as a second card image set.
[0159] S150: Select one or more candidate video card images from the second card image set as the video card images of the first video.
[0160] According to the method of the embodiment of the present application, an image template is generated based on the common characteristics of the video card images of videos with high video click-through rates, and an alternative video card image is generated based on the image template, so that the alternative video card image can better attract video viewers to click.
[0161] In S140 , the candidate video card images may be generated in a variety of different ways.
[0162] For example, FIG3 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0163] In one embodiment, the electronic device executes the following process shown in FIG3 to generate an alternative video card image.
[0164] S200: Retrieve similar pictures of the first card picture set from the picture library and record them as a first candidate picture set.
[0165] Specifically, in one implementation, the image library is a local image library. For example, the image library is stored on the terminal device used by the video publisher to publish the video, or the image library is stored on another local device connected to the terminal device (for example, a home network storage device). In another implementation, the image library is a cloud image library stored on a server.
[0166] S210, based on the first image template, edit the pictures in the first candidate picture set to generate candidate video card pictures, recorded as the first candidate video card picture set, and use the first candidate video card picture set as the second card picture set.
[0167] Specifically, in one implementation of S210, the pictures in the first candidate picture set are edited, including cropping the pictures (for example, cutting out an adult face, a full-body picture of a beautiful woman, etc.) and / or special effects processing (for example, adding text labels, picture special effects, etc. to the pictures).
[0168] According to the method of the embodiment of the present application, the existing pictures in the picture library are used as the picture source, and the pictures in the picture source are screened and edited according to the video card pictures of videos with high video click-through rates to generate alternative video card pictures, which can effectively limit the source range of the alternative video card pictures.
[0169] For another example, FIG4 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0170] In one embodiment, the electronic device executes the following process shown in FIG. 4 to generate an alternative video card image.
[0171] S300: extract video frames from the first video and record them as a second candidate picture set.
[0172] Specifically, in S300 , video frames may be extracted from the first video in a variety of ways.
[0173] For example, in one embodiment, video frames are extracted from the first video at fixed time intervals.
[0174] For another example, FIG5 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0175] In one embodiment, the electronic device executes the following process shown in FIG5 to extract video frames from the first video.
[0176] S400 , determining a user attention time interval (a first user attention time interval) of the M videos or the K videos based on user interaction behavior records of the M videos (the videos obtained in S100 ) or the K videos (the K videos screened in S110 ).
[0177] In one embodiment, for video sharing applications, user attention intervals refer to the time intervals during which video viewers pay close attention to the video content. Specifically, in one embodiment, the time intervals during which video viewers control video playback (e.g., pause, rewind, slow play, etc.) and / or interact with the video (e.g., enter comments, take screenshots, zoom in on an image, etc.) are considered user attention intervals. In S400, the user attention intervals for the video are determined based on the user interaction records for the video.
[0178] For example, a video viewer uses fast forwarding and / or dragging the time bar to quickly browse the video content. During the process of quickly browsing the video content, the time period in which the video viewer watches the video in its entirety at a normal speed is the user's attention time interval. For another example, while watching a video, the video viewer rewinds the video playback progress (for example, by dragging the time bar) and repeatedly watches the video is the user's attention time interval. For another example, while watching a video, the video viewer pauses the video playback and carefully watches the video content is the user's attention time interval. For another example, while watching a video, the time period in which the video viewer enters commentary subtitles (bullet screen) is the user's attention time interval.
[0179] S410: Determine a user attention time interval (a second user attention time interval) of the first video according to the first user attention time interval.
[0180] Specifically, in one implementation of S410, based on the video clips in the time interval that the first user pays attention to, similar video clips are retrieved from the first video (for example, similar video frames are retrieved based on a picture similarity retrieval algorithm), and the time interval corresponding to the retrieved similar video clips is the time interval that the second user pays attention to.
[0181] S420: Extract video frames from the video segment of the first video corresponding to the time interval that the second user pays attention to.
[0182] According to the method of the embodiment of the present application, video frames are extracted according to the time interval of user attention, which can effectively ensure that the extracted video frames are more attractive to video viewers to click.
[0183] S310, based on the first image template, edit the pictures in the second candidate picture set to generate candidate video card pictures, which are recorded as the second candidate video card picture set, and the second candidate video card picture set is used as the second card picture set.
[0184] In S310 , the pictures in the second candidate picture set are edited, and reference may be made to S210 .
[0185] According to the method of the embodiment of the present application, the video frame of the current video is used as the image source, and the images in the image source are screened and edited according to the video card images of videos with high video click-through rates to generate alternative video card images, which can ensure the correlation between the alternative video card images and the current video.
[0186] In one implementation of S310, a user attention space area is also introduced. In S310, the pictures in the second candidate picture set are cropped according to the user attention space area of the first video.
[0187] In one embodiment, for video sharing applications, the user attention space refers to the image area that the video viewer focuses on. Specifically, in one embodiment, the image area corresponding to the video viewer's video viewing interaction (e.g., partial screenshot, image zoom, etc.) during the video viewing process is the user attention space.
[0188] For example, if a video viewer zooms in on a certain image area multiple times while watching a video, the image area becomes the user's focused spatial area. For another example, if a video viewer takes a partial screenshot of a certain image area while watching a video, the image area becomes the user's focused spatial area.
[0189] Specifically, in one embodiment, a user attention space region (first user attention space region) of M videos (the videos obtained in S100) or K videos in a similar video set (the K videos selected in S110) is determined. A user attention space region (second user attention space region) of the first video is determined based on the first user attention space region.
[0190] Specifically, in one implementation, similar video frames are retrieved from the first video based on the video frame of the spatial region of interest to the first user, and the spatial region of interest to the first user is saved as the second spatial region of interest to the retrieved similar video frame.
[0191] According to the method of the embodiment of the present application, the video frame is cropped according to the spatial area of user attention, which can effectively ensure that the alternative video card image is more attractive to video viewers to click.
[0192] For another example, FIG6 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0193] In one embodiment, the electronic device executes the following process shown in FIG6 to generate an alternative video card image.
[0194] S500, calling the AI Generated Content (AIGC) model for automatic creation of generated content.
[0195] S510, using the AIGC model, based on the first image template generated in S130 and the video text description of the first video, generates an alternative video card image, recorded as a third alternative video card image set, and uses the third alternative video card image set as the second card image set.
[0196] In one embodiment, the video text description of the first video includes the alternative video title of the first video and / or the original video title of the first video.
[0197] Specifically, in one embodiment, the video publisher inputs the original video title of the first video and a plurality of alternative video titles.
[0198] Specifically, in another embodiment, the video publisher inputs the original video title of the first video, and the original video title of the first video is rewritten with reference to the video titles of K videos (the videos screened in S110) to generate a plurality of candidate video titles.
[0199] In S510, the AIGC model combines different candidate video titles to generate different candidate video card images.
[0200] The video text description of the first video may be a video title input by a video publisher, or may be description information for the first video content input by the video publisher.
[0201] For another example, FIG7 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0202] In one embodiment, the electronic device executes the following process shown in FIG. 7 to generate an alternative video card image.
[0203] S600, calling the AI Generated Content (AIGC) model for automatic creation of generated content.
[0204] S610, performing semantic segmentation on the pictures of the first card image set, and / or the first alternative video card image set (embodiment of FIG. 3), and / or the second alternative video card image set (embodiment of FIG. 4), to obtain a semantic segmentation map.
[0205] S620, using the AIGC model, based on the semantic segmentation map generated in S610 and the video text description of the first video, generates an alternative video card map, recorded as the fourth alternative video card map set, and uses the fourth alternative video card map set as the second card map set.
[0206] The video text description of the first video in S620 may refer to S510.
[0207] Furthermore, in another embodiment, a second card image set may be generated by combining multiple alternative video card image sets. For example, any multiple of the first alternative video card image set, the second alternative video card image set, the third alternative video card image set, and the fourth alternative video card image set may be combined to form the second card image set.
[0208] Furthermore, in S150 , a video card image of the first video may be selected from the second card image set in a variety of different ways.
[0209] For example, in one implementation of S150, the pictures in the second card image set are displayed to the video publisher, and the video publisher selects the pictures in the second card image set as the video card images.
[0210] In another implementation of S150, the pictures in the second card picture set are sorted according to a preset rule, and the top N pictures are used as video card pictures, where N is a preset integer greater than or equal to one.
[0211] In another implementation of S150, the images in the second card image set are sorted according to a preset rule, and the images in the second card image set are displayed to the video publisher according to the sorting result (for example, the images in the top ranking are displayed to the video publisher first). The video publisher selects the image as the video card image.
[0212] For example, FIG8 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0213] In one embodiment, the electronic device executes the following process shown in FIG. 8 to implement 150 .
[0214] S700: Calculate the similarity between each picture in the second card image set and the first card image set.
[0215] S710 , sorting the pictures in the second card picture set according to the similarity calculation result in S700 , wherein the pictures with greater similarity are ranked higher.
[0216] S720: Display the pictures in the second card picture set to the video publisher according to the sorting result of S710, and give priority to displaying the pictures with the highest sorting to the video publisher.
[0217] S730: Receive a picture selection operation from the video publisher, and select a picture as a video card picture from the second card picture set according to the picture selection operation from the video publisher.
[0218] For another example, FIG9 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0219] In one embodiment, the electronic device executes the following process shown in FIG. 9 to implement 150 .
[0220] S800: Obtain a video text description of a first video, for example, an alternative video title.
[0221] The video text description of the first video in S800 may refer to S510.
[0222] S810: Calculate the relevance between each image in the second card image set and the video text description.
[0223] S820 , sorting the pictures in the second card picture set according to the relevance calculation result in S810 , with pictures having a higher relevance being ranked higher.
[0224] S830: Display the pictures in the second card picture set to the video publisher according to the sorting result of S820, and give priority to displaying the pictures with the highest sorting to the video publisher.
[0225] Specifically, in one embodiment, multiple alternative video titles for the first video are obtained in S800. The relevance between each image in the second card image set and each alternative video title is calculated in S810. The multiple alternative video titles are displayed to the video publisher in S830, and the video publisher selects a video title from the multiple alternative video titles as the video title of the first video. Furthermore, after the video publisher selects an alternative video title, the images in the second card image set are sorted according to the relevance between each image in the second card image set and the selected alternative video title, and the images in the second card image set are displayed.
[0226] For example, the second card image set includes picture A, picture B, and picture C. The candidate video titles include title T1 and title T2.
[0227] The order of relevance of Picture A, Picture B and Picture C to Title T1 from high to low is Picture B, Picture C, Picture A; the order of relevance of Picture A, Picture B and Picture C to Title T2 from high to low is Picture A, Picture C, Picture B.
[0228] In S830, Title T1 and Title T2 are displayed to the user. When the user selects Title T1, Pictures B, C, and A are sorted based on relevance, and Pictures A, B, and C are displayed (with Picture B being displayed first). When the user selects Title T2, Pictures A, C, and B are sorted based on relevance, and Pictures A, B, and C are displayed (with Picture A being displayed first).
[0229] S840, receiving a title selection operation and a picture selection operation from the video publisher, determining the video title of the first video according to the title selection operation and the picture selection operation from the video publisher, and selecting a picture as a video card picture from the second card picture set.
[0230] For another example, FIG10 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0231] In one embodiment, the electronic device executes the following process shown in FIG. 10 to implement 150 .
[0232] S900: Obtain a video text description of the first video.
[0233] The candidate video titles of the first video in S900 may refer to S510.
[0234] S910: Calculate the relevance between each image in the second card image set and the video text description.
[0235] S920: Calculate the similarity between each image in the second card image set and the first card image set.
[0236] S930 , sorting the pictures in the second card picture set according to the correlation calculation result in S910 and the similarity calculation result in S920 .
[0237] For example, a weighted average calculation is performed on the correlation calculation results and the similarity calculation results, and the pictures in the second card image set are sorted according to the weighted average calculation result.
[0238] S940: Display the images in the second card image set to the video publisher according to the ranking result of S930, with priority given to displaying the images with the highest ranking to the video publisher. (See S830)
[0239] S950: Receive a picture selection operation from a video publisher, and determine a video card image for the first video based on the picture selection operation from the video publisher.
[0240] Specifically, in one embodiment, multiple alternative video titles of the first video are obtained in S900. The relevance between each picture in the second card image set and each alternative video title is calculated in S910. The multiple alternative video titles are displayed to the video publisher in S940, and the video publisher selects one video title from the multiple alternative video titles as the video title of the first video. Moreover, after the video publisher selects an alternative video title, the pictures in the second card image set are sorted according to the relevance between each picture in the second card image set and the selected alternative video title, as well as the similarity between each picture in the second card image set and the first card image set, and the pictures in the second card image set are displayed according to the sorting result.
[0241] According to the video card image generation method proposed in an embodiment of the present application, an embodiment of the present application also proposes a video card image generation system.
[0242] Optionally, in one implementation, the video card image generation system can be built into a terminal device used by a video publisher to publish videos, such as a mobile phone, tablet computer, laptop computer, desktop computer, etc. In another implementation, the video card image generation system can be built into other electronic devices and / or cloud servers connected to the terminal device used by the video publisher to publish videos. In another implementation, the video card image generation system can be partially built into the terminal device used by the video publisher to publish videos, and partially built into other electronic devices and / or cloud servers connected to the terminal device used by the video publisher to publish videos.
[0243] FIG11 shows a simplified structural diagram of a video card image generation system according to an embodiment of the present application.
[0244] As shown in FIG11 , the video card image generation system includes:
[0245] Similarity retrieval module 1010, which is used to perform video similarity and image similarity retrieval;
[0246] A template extraction module 1020 is configured to extract common features based on a set of images and generate a first image template based on the common features;
[0247] An image cropping module 1030 is used to crop an image according to a specific size, ratio, and image region location information, and edit the image according to specified rules;
[0248] An AIGC module 1040 is configured to generate an image of a specific size based on a combination of any one of the following: a title, a semantic segmentation map, and a template map;
[0249] A sorting module 1050 is used to calculate the relevance between the image and the title and the similarity between the images, and sort the multiple images according to the calculation results of the relevance and similarity;
[0250] The frame extraction module 1060 is used to extract video frames.
[0251] Furthermore, the video card image generation method provided in the embodiment of the present application can be applied to the situation where a video publisher publishes a video for the first time in a video sharing application scenario.
[0252] FIG12 is a flow chart of a method for generating a video card image according to an embodiment of the present application.
[0253] FIG13 is a schematic diagram showing an interface display effect according to an embodiment of the present application.
[0254] In one embodiment, the system shown in FIG. 11 executes the steps shown in FIG. 12 to achieve the generation of a video card image.
[0255] In one embodiment, when the system shown in FIG. 11 executes the steps shown in FIG. 12 , the terminal device of the video publisher displays an interface shown in FIG. 13 .
[0256] S1100: Receive a video to be published (a first video) uploaded by a video publisher and a video title (original video title) of the video to be published uploaded by the video publisher.
[0257] As shown in FIG13 , the cover / play interface of the received video to be published is displayed at 1200 , and the title of the video uploaded by the video publisher is displayed at 1201 .
[0258] Optionally, in one embodiment, in S1100, the video type, card image layout (such as large video image 17:9 / small square image 1:1), potential audience for the video to be released (for example, male / female), etc. input by the video publisher are also received.
[0259] S1110 , the similarity retrieval module 1010 retrieves similar videos (M) from a distribution video library based on the first video and the original video title.
[0260] The distribution video library is used to save videos, as well as the CTR and user interaction records of the videos.
[0261] Specifically, in S1110, when retrieving similar videos, the video type of the first video, the card image layout (such as 17:9 large video image / 1:1 small square image), the potential population targeted by the first video (for example, male / female), etc. are also matched.
[0262] S1111 , the similarity retrieval module 1010 selects TOP-K similar videos from the retrieved M similar videos based on the CTR of the video, and records them as a similar video set.
[0263] S1112 , the similarity retrieval module 1010 obtains a video card image of each video in the similar video set, which is recorded as a first card image set.
[0264] As shown in FIG13 , the cover / play interface of the video in the similar video collection and the video card image corresponding to the video are displayed at 1202 . The video publisher can switch to display different videos in the similar video collection by dragging the slider 1203 .
[0265] S1120: The template extraction module 1020 extracts commonalities from the images in the first card image set to generate a first image template.
[0266] S1113, the similarity retrieval module 1010 searches for similar pictures in the picture library based on the pictures in the first card picture set, and records them as a first candidate picture set.
[0267] S1130 , the image cropping module 1030 edits the images in the first candidate image set based on the first image template to generate candidate video card images, which are recorded as the first candidate video card image set.
[0268] As shown in FIG. 13 , pictures in the first candidate video card picture set are displayed at 1204 .
[0269] S1114 , the similarity retrieval module 1010 determines the user attention time interval of the videos in the similar video set based on the user interaction behavior records of the videos in the similar video set.
[0270] As shown in FIG. 13 , the user attention time interval of the videos in the similar video set is marked in the video progress bar 1205 .
[0271] S1115 , the similarity retrieval module 1010 retrieves similar video segments from the videos to be published based on the video segments corresponding to the user's attention time interval of the videos in the similar video set.
[0272] As shown in FIG. 13 , the video clips corresponding to the time interval of user attention in the video to be released are marked in the video progress bar 1206 .
[0273] S1140 , the frame extraction module 1060 extracts video frames from similar video clips and records them as a second candidate picture set.
[0274] S1116 , the similarity retrieval module 1010 determines the user attention space area of the video to be released based on the user interaction behavior records of the videos in the similar video set.
[0275] As shown in Figure 13, the user attention space area of the videos in the similar video set is marked on the playback interface of the videos in the similar video set (for example, the user attention space area of a certain video frame of the videos in the similar video set is 1207). The user attention space area of the video to be released is marked on the playback interface of the video to be released (for example, the user attention space area of a certain video frame of the video to be released is 1208).
[0276] S1131, the image cropping module 1030 edits the pictures in the second candidate picture set based on the first image template and the user's attention space area to generate candidate video card images, which are recorded as the second candidate video card image set.
[0277] As shown in FIG. 13 , pictures in the second candidate video card picture set are displayed at 1209 .
[0278] S1150 , the similarity retrieval module 1010 refers to the video titles of the videos in the similar video set, rewrites the video titles uploaded by the video publisher, and generates candidate video titles.
[0279] S1160: The AIGC module 1040 generates different candidate video card images based on the first image template and in combination with different candidate video titles, which are recorded as a third candidate video card image set.
[0280] S1161 , the AIGC module 1040 performs semantic segmentation on the images of the first candidate video card image set and the second candidate video card image set to generate a semantic segmentation map.
[0281] S1162: The AIGC module 1040 generates different candidate video card images based on the semantic segmentation image and in combination with different candidate video titles, which are recorded as a fourth candidate video card image set.
[0282] As shown in FIG. 13 , pictures of the third candidate video card image set and the fourth candidate video card image set are displayed at 1211 .
[0283] S1170, the sorting module 1050 calculates the correlation between the pictures in the first alternative video card image set, the second alternative video card image set, the third alternative video card image set and the fourth alternative video card image set and the alternative video titles, calculates the similarity between the pictures in the first alternative video card image set, the second alternative video card image set, the third alternative video card image set and the fourth alternative video card image set and the first card image set, and sorts the pictures in the second card image set according to the calculation results of the correlation and similarity.
[0284] Specifically, in one embodiment, the sorting module 1050 performs correlation and similarity calculations for the first alternative video card image set, the second alternative video card image set, the third alternative video card image set, and the fourth alternative video card image set, and sorts the first alternative video card image set, the second alternative video card image set, the third alternative video card image set, and the fourth alternative video card image set, respectively.
[0285] S1180, displaying candidate video titles, and displaying pictures of a first candidate video card image set, a second candidate video card image set, a third candidate video card image set, and a fourth candidate video card image set according to the sorting result.
[0286] S1181: Determine the video title of the first video and the video card image of the first video according to the user's selection operation.
[0287] As shown in Figure 13, candidate video titles are displayed at 1210. The video publisher selects one of the candidate video titles as the video title for the video to be published. After the video publisher selects an alternative video title, images are displayed in 1204, 1209, and 1211 according to the ranking results for the video title (the top two images are displayed). After the video publisher selects one or more images displayed in 1204, 1209, and 1211, the image selected by the video publisher becomes the video card image for the video to be published.
[0288] It should be noted that, in the embodiment of the present application, the interface displayed by the terminal device of the video publisher is not limited to the mode shown in Figure 13. Those skilled in the art can design the interface mode displayed by the terminal device of the video publisher according to actual needs.
[0289] For example, FIG14 is a schematic diagram showing an interface display effect according to an embodiment of the present application.
[0290] As shown in FIG14 , the cover / play interface of the video to be published is displayed at 1300 , and the title of the video uploaded by the video publisher is displayed at 1301 .
[0291] The candidate video titles are displayed at 1302 , and the video publisher selects one of the candidate video titles as the video title of the video to be published.
[0292] Pictures of the first alternative video card image set, the second alternative video card image set, the third alternative video card image set, and the fourth alternative video card image set are displayed at 1303 .
[0293] The sorting module 1050 calculates the correlation between the pictures in the first alternative video card image set, the second alternative video card image set, the third alternative video card image set and the fourth alternative video card image set and the alternative video titles, calculates the similarity between the pictures in the first alternative video card image set, the second alternative video card image set, the third alternative video card image set and the fourth alternative video card image set and the first card image set, and sorts the pictures in the second card image set according to the calculation results of the correlation and similarity.
[0294] After the video publisher selects a candidate video title, pictures (the top six pictures) are displayed in 1303 according to the ranking results for the candidate video title. After the video publisher clicks on one or more pictures displayed in 1303, the picture selected by the video publisher becomes the video card picture of the video to be published.
[0295] Furthermore, the video card image generation method provided in the embodiment of the present application can be applied to update the current video card image of the video in the video sharing application scenario.
[0296] FIG15 is a flow chart showing a method for generating a video card image according to an embodiment of the present application.
[0297] In one embodiment, the system shown in FIG. 11 executes the steps shown in FIG. 15 to update the video card image of the first video.
[0298] In the embodiment shown in Figure 15, the first video is a video that has been posted to a video sharing platform and has been clicked by a video viewer. User feedback information and user interaction behavior records have been recorded for the first video.
[0299] S1400: Obtain a video (a first video) whose video card image needs to be updated and the video title of the video (the original video title).
[0300] Optionally, in one embodiment, in S1400, the video type of the first video, the card image format, the potential audience targeted by the video to be released, etc. are also obtained.
[0301] S1410 , the similarity retrieval module 1010 retrieves M videos similar to the first video from the distribution video library.
[0302] S1411 , the similarity retrieval module 1010 selects TOP-K similar videos from the retrieved M similar videos based on CTR, and records them as a similar video set.
[0303] S1412: The similarity retrieval module 1010 obtains a video card image of each video in the similar video set, which is recorded as a first card image set.
[0304] S1420: The template extraction module 1020 extracts commonalities from the images in the first card image set to generate a first image template.
[0305] S1413: The similarity search module 1010 searches for similar pictures in the picture library based on the pictures in the first card picture set, and records them as a first candidate picture set.
[0306] S1430: The image cropping module 1030 edits the images in the first candidate image set based on the first image template to generate candidate video card images, which are recorded as the first candidate video card image set.
[0307] S1414 , the similarity retrieval module 1010 determines a user attention time interval of the first video based on the user interaction behavior records of the videos in the similar video set and the user interaction behavior records of the first video.
[0308] S1440 , the frame extraction module 1060 extracts video frames from the video segment corresponding to the user's attention time interval of the first video, and records them as a second candidate picture set.
[0309] S1415 , the similarity retrieval module 1010 determines the user attention space area of the first video based on the user interaction behavior records of the videos in the similar video set and the user interaction behavior records of the first video.
[0310] S1431, the image cropping module 1030 edits the pictures in the second candidate picture set based on the first image template and the user's attention space area to generate candidate video card images, which are recorded as the second candidate video card image set.
[0311] S1450 , the similarity retrieval module 1010 refers to the video titles of the videos in the similar video set, rewrites the video title of the first video, and generates candidate video titles.
[0312] S1460: The AIGC module 1040 generates different candidate video card images based on the first image template and in combination with different candidate video titles, which are recorded as a third candidate video card image set.
[0313] S1461 , the AIGC module 1040 performs semantic segmentation on the images of the first candidate video card image set and the second candidate video card image set to generate a semantic segmentation map.
[0314] S1462: The AIGC module 1040 generates different candidate video card images based on the semantic segmentation image and in combination with different candidate video titles, which are recorded as a fourth candidate video card image set.
[0315] S1470, the sorting module 1050 uses the combination of the first alternative video card image set, the second alternative video card image set, the third alternative video card image set and the fourth alternative video card image set as the second card image set, calculates the correlation between the pictures in the second card image set and the video title of the first video, calculates the similarity between the pictures in the second card image set and the first card image set, and sorts the pictures in the second card image set according to the calculation results of the correlation and similarity.
[0316] S1480: Using one or more pictures ranked first in the sorting result as video card pictures of the first video.
[0317] Furthermore, the video card image generation method provided in the embodiment of the present application can also be applied to video editing application scenarios.
[0318] In some video editing application scenarios, the user wishes to edit a video frame of a video to obtain a video card image, and the user wishes that the image effect of the video card image can imitate the image effect of a video card image of a certain video.
[0319] FIG16 is a flow chart of a method for generating a video card image according to an embodiment of the present application.
[0320] In one embodiment, in a video editing application scenario, the electronic device executes the following process shown in FIG. 16 to generate a video card image of the first video.
[0321] In the embodiment shown in FIG16 , the first video is the video to be edited.
[0322] S1500: Obtain M videos having video content similar to the first video, where M is an integer greater than or equal to 1.
[0323] Specifically, in S1500, M videos are screened from sample videos in the editing sample library and / or shared videos. User feedback information of the M videos includes the number of likes for the videos.
[0324] S1510 , screening K (Top-K) videos with the highest number of likes from the M videos obtained in S1500 , where K is an integer greater than or equal to 1 and less than or equal to M.
[0325] S1511, displaying the K videos screened out in S1510 and video card images of the K videos to the user.
[0326] S1520, receiving a user selection operation, and determining a video card image selected by the user according to the user selection operation, which is recorded as a third card image set.
[0327] In one embodiment, the user selects a video, and the video card image corresponding to the video is determined based on the video selected by the user, which is recorded as the third card image set. In another embodiment, the user selects a video card image, and the video card image selected by the user is recorded as the third card image set.
[0328] S1530: Generate a second image template based on the commonalities of the video card images in the third card image set.
[0329] S1540: Based on the second image template generated in S1530, generate a plurality of candidate video card images that match the second image template, which are recorded as a fourth card image set.
[0330] S1550: Select one or more candidate video card images from the fourth card image set as the video card images of the first video.
[0331] According to the method of the embodiment of the present application, a video card image of the video to be edited can be generated based on the video that the user wants to imitate among the videos with high video likes, so that the video card image of the video to be edited has a better visual effect.
[0332] Referring to the second card image set in S140 , in S1540 , candidate video card images may be generated in a variety of different ways.
[0333] For example, Figure 17 shows a partial method flow chart of a video card image generation method according to an embodiment of the present application.
[0334] In one embodiment, the electronic device executes the following process shown in FIG. 17 to generate an alternative video card image.
[0335] S1600: Extract, from the first video, video frames whose picture dimensions match the second image template, and record them as a third candidate picture set.
[0336] For example, the second image template has a picture dimension including trees, a river, and a portrait. In S1600, a video frame including trees, a river, and a portrait is extracted from the first video.
[0337] For another example, the second image template has a large female face and a dog head. In S1600, a video frame containing both a large female face and a dog head is extracted from the first video.
[0338] S1610, based on the second image template, edit the pictures in the third candidate picture set to generate candidate video card pictures, recorded as the fifth candidate video card picture set, and use the fifth candidate video card picture set as the fourth card picture set.
[0339] According to the method of the embodiment of the present application, the video frame of the current video is used as the image source, and based on the common characteristics of the video card images of videos with a high number of video likes, the images in the image source are screened and edited to generate alternative video card images. On the basis of ensuring the correlation between the alternative video card images and the current video, the alternative video card images can be made closer to the image effects of the video card images of videos with a high number of video likes.
[0340] For another example, FIG18 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0341] In S1500 , M videos are screened from sample videos in the editing sample library and / or shared videos. The M videos are edited videos, and video editing information is recorded for the M videos.
[0342] In one embodiment, the electronic device executes the following process shown in FIG. 18 to generate an alternative video card image.
[0343] S1700 , determining a user attention time interval (third user attention time interval) of the M videos or K videos based on video editing information of the M videos (the videos obtained in S1500 ) or the K videos (the K videos filtered in S1510 ).
[0344] In one embodiment, for a video editing application scenario, the user-focused time interval refers to the time interval during which the video editor edits the video content. For example, a time interval corresponding to a video clip using special effects such as slow motion or loop playback is considered the user-focused time interval.
[0345] S1710: Determine a user attention time interval (a fourth user attention time interval) for the first video according to the third user attention time interval.
[0346] Specifically, in one implementation of S1710, based on the video clips in the time interval that the first user pays attention to, similar video clips are retrieved from the first video (for example, similar video frames are retrieved based on a picture similarity retrieval algorithm), and the time interval corresponding to the retrieved similar video clips is the time interval that the fourth user pays attention to.
[0347] S1720: Extract video frames from the video segment of the first video corresponding to the fourth user's attention time interval, and record them as a fourth candidate picture set.
[0348] S1730, based on the second image template, edit the pictures in the fourth candidate picture set to generate candidate video card pictures, recorded as the sixth candidate video card picture set, and use the sixth candidate video card picture set as the fourth card picture set.
[0349] According to the method of the embodiment of the present application, the user attention time interval of the current video is determined based on videos with high video like numbers, and video frames are extracted based on the user attention time interval, so that the alternative video card image can be closer to the image effect of the user attention video clip of the video with high video like numbers.
[0350] In one implementation of S1730, a user attention space area is also introduced. In S1730, the pictures in the fourth candidate picture set are cropped according to the user attention space area of the first video.
[0351] In one embodiment, for a video editing application scenario, the user focused spatial area refers to the image area where the video editor edits the video content. For example, the image area using close-up magnification / texture / image effects is the user focused spatial area.
[0352] Specifically, in one embodiment, a user attention spatial region (third user attention spatial region) of M videos (the videos obtained in S1500) or K videos (the K videos filtered in S1510) is determined. A user attention spatial region (fourth user attention spatial region) of the first video is determined based on the third user attention spatial region.
[0353] Specifically, in one implementation, similar video frames are retrieved from the first video based on the video frame of the second user's focused spatial region, and the second user's focused spatial region is saved as the fourth user's focused spatial region of the retrieved similar video frames.
[0354] According to the method of the embodiment of the present application, the video frame of the current video is used as the image source, and based on the common characteristics of the video card images of videos with a high number of video likes, the images in the image source are screened and edited to generate alternative video card images. On the basis of ensuring the correlation between the alternative video card images and the current video, the alternative video card images can be made closer to the image effects of the video card images of videos with a high number of video likes.
[0355] For another example, FIG19 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0356] The first video is a video that has been edited before, and video editing information has been recorded for the first video.
[0357] In one embodiment, the electronic device executes the following process shown in Figure 19 to generate an alternative video card image.
[0358] S1800 : Determine a user attention time interval (fifth user attention time interval) of the first video based on video editing information of the first video.
[0359] For example, in one embodiment, the time interval corresponding to the video clip edited by the user is the time interval that the user focuses on.
[0360] S1810 , extracting video frames from the video segment of the first video corresponding to the fifth user's attention time interval, and recording them as a fifth candidate picture set.
[0361] S1820, based on the second image template, edit the pictures in the fifth candidate picture set to generate candidate video card pictures, recorded as the seventh candidate video card picture set, and use the seventh candidate video card picture set as the fourth card picture set.
[0362] According to the method of the embodiment of the present application, the user's attention time interval is determined according to the video editing information of the current video, and video frames are extracted according to the user's attention time interval, so that the extracted video frames can better meet the needs of the video editor.
[0363] In one implementation of S1820, a user attention space region is also introduced. In S1820, pictures in the fourth candidate picture set are cropped according to the user attention space region (fifth user attention space region) of the first video.
[0364] For example, in one embodiment, the image area that the user edits is the user's focused spatial area.
[0365] Furthermore, in another embodiment, a fourth card image set can be generated by combining multiple alternative video card image sets. For example, any multiple sets of the fifth alternative video card image set, the sixth alternative video card image set, and the seventh alternative video card image set are combined to form the fourth card image set. Furthermore, the fourth card image set can be generated by referring to the generation process of the second card image set in the video sharing application scenario.
[0366] Further, referring to S150, in S1550, a video card image of the first video can be selected from the fourth card image set in a variety of different ways.
[0367] For example, in one implementation of S1550, the pictures in the fourth card image set are displayed to the video editor, and the pictures selected by the video editor as the video card images in the fourth card image set are displayed to the video editor.
[0368] In another implementation of S1550, the pictures in the fourth card picture set are sorted according to a preset rule, and the top N pictures are used as video card pictures, where N is a preset integer greater than or equal to one.
[0369] In another implementation of S1550, the images in the fourth card image set are sorted according to a preset rule, and the images in the fourth card image set are displayed to the video editor according to the sorting result (for example, the images in the top order are displayed to the video editor first). The video editor selects the image as the video card image.
[0370] For example, Figure 20 shows a partial method flow chart of a video card image generation method according to an embodiment of the present application.
[0371] In one embodiment, the electronic device executes the following process shown in FIG. 20 to implement 5150 .
[0372] S1900 , calculating the similarity between each image in the fourth card image set and the third card image set.
[0373] S1910 , sorting the images in the fourth card image set according to the similarity calculation result in S1900 , with the images having a higher similarity being ranked higher.
[0374] S1920 , according to the sorting result of S1910 , display the pictures in the fourth card picture set to the video editor, and give priority to displaying the pictures with the highest sorting to the video editor.
[0375] S1930 , receiving a picture selection operation of a video editor, and selecting a picture as a video card picture from a fourth card picture set according to the picture selection operation of the video editor.
[0376] For another example, FIG21 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0377] In one embodiment, the electronic device executes the following process shown in FIG. 21 to implement 1550 .
[0378] S2000: Calculate the relevance between each picture in the fourth card image set and the video title of the first video.
[0379] S2010 , sorting the pictures in the fourth card picture set according to the correlation calculation result in S2000 , with the pictures having a higher correlation being ranked higher.
[0380] S2020: Display the pictures in the fourth card picture set to the video editor according to the sorting result of S2010, and give priority to displaying the pictures with higher sorting to the video editor.
[0381] S2030: Receive a picture selection operation from a video editor, and select a picture as a video card picture from a fourth card picture set according to the picture selection operation of the video editor.
[0382] For another example, FIG22 shows a partial flow chart of a method for generating a video card image according to an embodiment of the present application.
[0383] In one embodiment, the electronic device executes the following process shown in FIG. 22 to implement 1550 .
[0384] S2100: Calculate the relevance between each picture in the fourth card image set and the video title of the first video.
[0385] S2110 , calculating the similarity between each image in the fourth card image set and the third card image set.
[0386] S2120 , sorting the images in the fourth card image set according to the correlation calculation result in S2100 and the similarity calculation result in S2110 .
[0387] For example, a weighted average calculation is performed on the correlation calculation results and the similarity calculation results, and the pictures in the fourth card image set are sorted according to the weighted average calculation result.
[0388] S2130: Display the pictures and alternative video titles in the fourth card image set to the video publisher according to the sorting result of S2120, and give priority to displaying the pictures with higher sorting to the video editor.
[0389] S2140: Receive a picture selection operation from a video editor, and select a picture as a video card picture from a fourth card picture set according to the picture selection operation of the video editor.
[0390] FIG23 is a simplified structural diagram of a video editing system according to an embodiment of the present application.
[0391] As shown in FIG23 , the video editing system includes:
[0392] Similarity retrieval module 2210, which is used to perform video similarity and image similarity retrieval;
[0393] A template extraction module 2220 is configured to extract common features based on a set of images and generate a second image template based on the common features;
[0394] Image cropping module 2230, which is used to crop the image according to specific size, ratio and image area location information, and edit the image according to specified rules;
[0395] A sorting module 2240 is used to calculate the relevance between the image and the title and the similarity between the images, and sort the multiple images according to the calculation results of the relevance and similarity;
[0396] The frame extraction module 2250 is used to extract video frames.
[0397] Figure 24 is a flow chart of a method for generating a video card image according to an embodiment of the present application.
[0398] In one embodiment, the video editing system shown in FIG. 23 executes the steps shown in FIG. 24 to realize the generation of a video card image.
[0399] S2300: Obtain a video to be edited (a first video) and a video title of the video to be edited.
[0400] S2310 , the similarity retrieval module 2210 retrieves similar videos (M) from a video library / sample library based on the first video and the video title of the first video.
[0401] S2311 , the similarity retrieval module 2210 selects TOP-K similar videos with the highest number of likes from the retrieved M similar videos based on the number of likes of the M similar videos.
[0402] S2301, displaying the K videos screened out in S110 and video card images of the K videos to the user.
[0403] S2302, receiving a user selection operation, determining N videos selected by the user according to the user selection operation, and recording the video card images of the N videos as a third card image set.
[0404] S2330: The template extraction module 2220 extracts commonalities from the images in the third card image set to generate a second image template.
[0405] S2340: The frame extraction module 2250 extracts video frames whose picture dimensions match the second image template from the first video based on the second image template, and records them as a third candidate picture set.
[0406] S2350: The image cropping module 2230 edits the images in the third candidate image set according to the second image template to generate candidate video card images, which are recorded as the fifth candidate video card image set.
[0407] S2312, the similarity retrieval module 2210 determines the user attention time interval of the first video (the fourth user attention time interval) and the user attention spatial area (the fourth user attention spatial area) of the first video based on the video editing information of the K videos (the K videos filtered by S2311).
[0408] S2341: The frame extraction module 2250 extracts video frames from a fourth user attention time interval of the first video and records them as a fourth candidate picture set.
[0409] S2351, the image cropping module 2230 edits the pictures in the fourth candidate picture set according to the second image template and the fourth user attention space area of the first video to generate candidate video card images, which are recorded as the sixth candidate video card image set.
[0410] S2313, the similarity retrieval module 2210 determines the user attention time interval of the first video (the fifth user attention time interval) and the user attention spatial area of the first video (the fifth user attention spatial area) based on the video editing information of the first video.
[0411] S2342: The frame extraction module 2250 extracts video frames from the fifth user attention time interval of the first video and records them as a fifth candidate picture set.
[0412] S2352, the image cropping module 2230 edits the pictures in the fifth candidate picture set according to the second image template and the fifth user attention space area of the first video to generate candidate video card images, which are recorded as the seventh candidate video card image set.
[0413] S2360, the sorting module 2240 uses the combination of the fifth alternative video card image set, the sixth alternative video card image set and the seventh alternative video card image set as the fourth card image set, calculates the correlation between the pictures in the fourth card image set and the video title of the first video, calculates the similarity between the pictures in the fourth card image set and the third card image set, and sorts the pictures in the second card image set according to the calculation results of the correlation and similarity.
[0414] S2303: Display the pictures in the fourth card picture set to the video editor according to the sorting result of the sorting module 2240.
[0415] S2304: Determine the image to be used as the editing result of the video card image according to the selection operation of the video editor.
[0416] Furthermore, in one embodiment, in S2303, based on the second image template or the card image in the third card image set, editable graphic attachments (for example, artistic text, special effects, device components, etc.) are superimposed on the image in the fourth card image set displayed to the video editor for secondary editing by the video editor.
[0417] Furthermore, in some application scenarios, videos are not intended for public release but rather for private viewing. For example, videos saved in local albums or cloud albums. For such videos, the video card image does not need to attract viewers to click, but rather helps them recall the content of the video.
[0418] In view of the above application scenario, an embodiment of the present application provides a method for generating a video card image. The method provided in the embodiment of the present application is applied to an electronic device. In the method provided in the embodiment of the present application, when generating a video card image for a second video, the electronic device does not need to refer to other publicly released videos, but instead generates a video card image for the second video based on user interaction behavior records for the second video.
[0419] Figure 25 is a flow chart of a method for generating a video card image according to an embodiment of the present application.
[0420] In one embodiment, the electronic device executes the steps shown in Figure 25 to achieve the generation of the video card image.
[0421] S2400: Record the user interaction behavior of the second video and generate a user interaction behavior record.
[0422] Specifically, in one embodiment, the user interaction behavior includes the user's video playback behavior and the user's video editing behavior.
[0423] S2410: Determine a user's attention time interval and a user's attention space area for the second video based on the user interaction behavior record.
[0424] Specifically, with respect to the user's editing behavior on the video, the user's focused time interval and the user's focused spatial area of the second video are determined according to the video clip targeted by the user's editing behavior and the image area of the video frame targeted by the user's editing behavior.
[0425] S2420 , extracting video frames from the video clip corresponding to the time interval that the user pays attention to, and recording them as a candidate picture set.
[0426] S2430: Crop the pictures in the candidate picture set according to the user's attention space area and the preset video card picture layout to generate candidate video card pictures, which are recorded as the candidate video card picture set.
[0427] S2440: Select one or more candidate video card images from the card image set as the video card images of the second video.
[0428] According to the method of the embodiment of the present application, an alternative video card image is generated for the video based on the user's interactive behavior with the video, so that the alternative video card image can be associated with the user's historical operations on the video, so that the alternative video card image can better assist the user in recalling the video content.
[0429] Referring to S150 , in S2440 , a video card image of the second video may be selected from the card image set in a variety of different ways.
[0430] For example, in one implementation of S2440, a picture in the card image set is displayed to the user, and the picture selected by the user as the video card image in the card image set.
[0431] In another implementation of S2440, the pictures in the card picture set are sorted according to a preset rule, and the top N pictures are used as video card pictures, where N is a preset integer greater than or equal to one.
[0432] In another implementation of S2440, the pictures in the card image set are sorted according to a preset rule, and the pictures in the card image set are displayed to the user according to the sorting result (for example, the pictures in the top order are displayed to the user first). The user selects the picture to be used as the video card image.
[0433] For example, in one embodiment, a comprehensive ranking is performed according to three dimensions: user preference (such as the frames most frequently zoomed, paused, and repeatedly viewed by users), image quality (such as a picture aesthetic scoring model), and semantic similarity (such as using the distance between the picture embedding vector and the average video embedding vector).
[0434] Furthermore, after S2440, user interaction behavior for the second video continues to be recorded. When the number of new views and / or new edits meets a preset threshold, it is determined whether the user's attention time interval and / or user's attention spatial area of the second video has changed. If the user's attention time interval and / or user's attention spatial area has changed, S2420 to S2440 are repeated to update the video card image of the second video.
[0435] In the description of the embodiments of the present application, for the convenience of description, the device is described as being divided into various modules according to their functions. The division of each module is merely a division of logical functions. When implementing the embodiments of the present application, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0436] Specifically, the device proposed in the embodiment of the present application can be fully or partially integrated into a physical entity during actual implementation, or it can be physically separated. And these modules can all be implemented in the form of software calling through processing elements; or they can all be implemented in the form of hardware; or some modules can be implemented in the form of software calling through processing elements, and some modules can be implemented in the form of hardware. For example, the detection module can be a separately established processing element, or it can be integrated in a chip of an electronic device. The implementation of other modules is similar. In addition, these modules can be fully or partially integrated together, or they can be implemented independently. During the implementation process, each step of the above method or each of the above modules can be completed by the hardware integrated logic circuit in the processor element or the instructions in the form of software.
[0437] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0438] An embodiment of the present application further provides an electronic device.
[0439] FIG26 is a schematic structural diagram of an electronic device according to an embodiment of the present application.
[0440] As shown in Figure 26, the electronic device 2500 includes a memory 2502 for storing computer program instructions and a processor 2501 for executing program instructions, wherein, when the computer program instructions are executed by the processor 2501, the electronic device 2500 is triggered to execute the method steps described in the embodiment of the present application.
[0441] Specifically, in one embodiment of the present application, the above-mentioned one or more computer programs are stored in the above-mentioned memory 2502, and the above-mentioned one or more computer programs include instructions. When the above-mentioned instructions are executed by the above-mentioned electronic device 2500, the above-mentioned electronic device 2500 executes the method steps described in the embodiment of the present application.
[0442] It is understood that the structural description of the electronic device 2500 in the embodiment of the present application does not constitute a specific limitation on the electronic device 2500. In other embodiments of the present application, the electronic device 2500 may include other components besides the processor 2501 and the memory 2502.
[0443] The processor 2501 may be a device on a chip (SOC), and the processor 2501 may include a central processing unit (CPU), and may further include other types of processors.
[0444] The processor involved in processor 2501 may include, for example, a CPU, a DSP, a microcontroller, or a digital signal processor, and may also include a GPU, an embedded neural network processor (NPU), and an image signal processor (ISP). The processor may also include necessary hardware accelerators or logic processing hardware circuits, such as ASICs, or one or more integrated circuits for controlling the execution of the program of the technical solution of this application. In addition, the processor may have the function of operating one or more software programs, and the software programs may be stored in a storage medium.
[0445] The processor 2501 may include one or more processing units. For example, the processor may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Different processing units may be independent components or integrated into one or more processors. In some embodiments, the electronic device 2500 may also include one or more processors 2501. The controller may generate an operation control signal based on the instruction opcode and the timing signal to complete the control of instruction fetching and execution.
[0446] In some embodiments, the processor 2501 may include one or more interfaces. The interfaces may include an inter-integrated circuit (I2C) interface, an integrated circuit sound (I2S) interface, a pulse code modulation (PCM) interface, a universal asynchronous receiver / transmitter (UART) interface, a mobile industry processor interface (MIPI), a general-purpose input / output (GPIO) interface, a SIM card interface, and / or a USB interface, etc. Among them, the USB interface is an interface that complies with the USB standard specification, and specifically can be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc. The USB interface can be used to connect a charger to charge the electronic device, and can also be used to transmit data between the electronic device and peripheral devices.
[0447] Electronic device 2500 may also include an external memory interface for connecting an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device. The external memory card communicates with processor 2501 via the external memory interface to implement data storage. For example, files such as music and videos can be stored on the external memory card.
[0448] Memory 2502 may include a code storage area and a data storage area. The code storage area may store an operating system. The data storage area may store data created during the use of electronic device 2500. Furthermore, memory 2502 may include high-speed random access memory and non-volatile memory, such as one or more disk storage components, flash memory components, and universal flash storage (UFS).
[0449] The memory 2502 may be a read-only memory (ROM), other types of static storage devices that can store static information and instructions, a random access memory (RAM) or other types of dynamic storage devices that can store information and instructions, an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), a magnetic disk storage medium or other magnetic storage device, or any computer-readable medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer.
[0450] The processor 2501 and the memory 2502 may be combined into one processing device, or more commonly, they may be independent components.
[0451] Optionally, the devices, apparatuses, and modules described in the embodiments of the present application may be implemented by computer chips or entities, or by products having certain functions.
[0452] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, apparatus, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0453] In the several embodiments provided in this application, if any function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to the prior art, or part of the technical solution, can be embodied in the form of a software product, which is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to perform all or part of the steps of the method described in each embodiment of this application.
[0454] Specifically, an embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer-readable storage medium is run on a computer, the computer executes the method provided in the embodiment of the present application.
[0455] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program product is run on a computer, it enables the computer to execute the method provided in the embodiment of the present application.
[0456] The embodiment description in this application is described with reference to the flow chart and / or block diagram according to the method, equipment (device) and computer program product of embodiment of the present application.It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions.These computer program instructions can be provided to the processor of general-purpose computer, special-purpose computer, embedded processing machine or other programmable data processing equipment to produce a machine, so that the instruction executed by the processor of computer or other programmable data processing equipment produces the device for realizing the function specified in one flow chart flow chart or multiple flow charts and / or one block or multiple blocks of block diagram.
[0457] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0458] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0459] It should also be noted that, in the embodiments of the present application, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can represent the existence of A alone, the existence of A and B at the same time, and the existence of B alone. Among them, A and B can be singular or plural. The character " / " generally indicates that the previous and subsequent associated objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b and c can represent: a, b, c, a and b, a and c, b and c, or a and b and c, where a, b, c can be single or multiple.
[0460] In the embodiments of the present application, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not preclude the presence of additional identical elements in the process, method, commodity, or apparatus comprising the element.
[0461] The present application may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present application may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communications network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.
[0462] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences from other embodiments. In particular, the device embodiments are generally similar to the method embodiments, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiments.
[0463] Those skilled in the art will appreciate that the various units and algorithm steps described in the embodiments of the present application can be implemented using a combination of electronic hardware, computer software, and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0464] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0465] The above description is merely a specific embodiment of the present application. Any person skilled in the art may easily conceive of variations or substitutions within the technical scope disclosed in this application, and such variations or substitutions shall be within the scope of protection of this application. The scope of protection of this application shall be subject to the scope of protection of the claims.
Claims
1. A method for generating a video card image, characterized in that: The method is applied to an electronic device, and includes: Obtain M videos with similar video content to the first video, where M is an integer greater than or equal to 1; Filtering K videos from the M videos based on user feedback information of the M videos, where K is an integer greater than or equal to 1 and less than or equal to M, wherein the user feedback information includes the number of video likes and / or the video click-through rate; Obtain video card images of the K videos; Based on the commonalities of the video card images of the K videos, a plurality of candidate video card images are generated.
2. The method according to claim 1, characterized in that The step of screening K videos from the M videos based on user feedback information of the M videos includes: From the M videos, select K videos with the highest video click rates.
3. The method according to claim 1 or 2, characterized in that The step of generating a plurality of candidate video card images according to the commonalities of the video card images of the K videos includes: generating a first image template according to the commonalities of the video card images of the K videos; Based on the first image template, a plurality of first candidate video card images matching the first image template are generated.
4. The method according to claim 3, characterized in that The step of generating a plurality of first candidate video card images matching the first image template based on the first image template includes: Obtain similar images of the video card images of the K videos; Based on the first image template, similar images of the video card images of the K videos are edited to generate the first candidate video card image.
5. The method according to claim 3, characterized in that The step of generating a plurality of first candidate video card images matching the first image template based on the first image template further includes: Extracting a first video frame from the first video; The first video frame is edited according to the first image template to generate the first candidate video card image.
6. The method according to claim 5, characterized in that The extracting a first video frame from the first video includes: Determining a user attention time interval for the M videos or the K videos based on user interaction behavior records for the M videos or the K videos; Determining a user attention time interval of the first video according to the user attention time intervals of the M videos or the K videos; The first video frame is extracted from a video segment of the first video corresponding to a user attention time interval of the first video.
7. The method according to claim 6, characterized in that The determining, based on the user attention time intervals of the M videos or the K videos, of the user attention time intervals of the first video includes: According to the video segments corresponding to the user's attention time interval of the M videos or the K videos, similar video segments are retrieved in the first video, and the time interval corresponding to the similar video segments is the user's attention time interval of the first video.
8. The method according to claim 5, characterized in that The extracting the first video frame from the first video further includes: Determining a user attention time interval of the first video based on the user interaction behavior record of the first video and the user interaction behavior record of the video, wherein the user interaction behavior record of the video includes the user interaction behavior records of the M videos or the K videos; The first video frame is extracted from a video segment of the first video corresponding to a user attention time interval of the first video.
9. The method according to claim 5, characterized in that The generating of the first candidate video card image further includes: Determining, based on the user interaction behavior records of the M videos or the K videos, a user attention space area of the M videos or the K videos; Determining a user attention space area of the first video according to the user attention space areas of the M videos or the K videos; The first video frame is cropped according to a user-focused spatial area of the first video.
10. The method according to claim 5, characterized in that The generating of the first candidate video card image further includes: Determining a user attention space area of the first video according to the user interaction behavior record of the first video and the user interaction behavior record of the video, wherein the user interaction behavior record of the video includes the user interaction behavior records of the M videos or the K videos; The first video frame is cropped according to a user-focused spatial area of the first video.
11. The method according to claim 3, characterized in that The step of generating a plurality of first candidate video card images matching the first image template based on the first image template includes: Artificial intelligence is used to automatically create a content model, and the first alternative video card image is generated based on the first image template and the video text description of the first video.
12. The method according to claim 11, characterized in that The video text description of the first video includes an alternative video title of the first video and / or an original video title of the first video.
13. The method according to claim 12, characterized in that The method further comprises: Obtaining the original video title of the first video; Based on the video titles of the M videos or the K videos, the original video title of the first video is rewritten to generate a plurality of candidate video titles.
14. The method according to claim 3, characterized in that The step of generating a plurality of first candidate video card images matching the first image template based on the first image template further includes: Using artificial intelligence to automatically create and generate a content model, based on the semantic segmentation map and the video text description of the first video, a card image of the first candidate video is generated, wherein: The semantic segmentation map is a semantic segmentation map obtained by semantically segmenting the video card maps of the K videos, and / or the first candidate video card map set, and / or the second candidate video card map set; The first candidate video card image set is a set of images generated by editing similar images of the video card images of the K videos according to the first image template; The second candidate video card image set is a set of images generated by editing video frames extracted from the first video according to the first image template.
15. The method according to claim 1, wherein The step of screening K videos from the M videos based on user feedback information of the M videos includes: From the M videos, select K videos with the highest number of likes.
16. The method according to claim 15, characterized in that The step of generating a plurality of candidate video card images according to the commonalities of the video card images of the K videos includes: Determine the video card image selected by the video editor among the video card images of the K videos; generating a second image template based on the commonalities of the video card images selected by the video editor; Based on the second image template, a plurality of second candidate video card images matching the second image template are generated.
17. The method according to claim 16, characterized in that The step of generating a plurality of second candidate video card images matching the second image template based on the second image template includes: Extracting a second video frame from the first video; The second video frame is edited according to the second image template to generate the second candidate video card image.
18. The method according to claim 17, characterized in that Extracting a second video frame from the first video includes: A second video frame whose picture dimension matches the second image template is extracted from the first video.
19. The method according to claim 17, wherein Extracting a second video frame from the first video includes: Determining a user attention time interval of the first video according to the video editing information of the M videos or the K videos; The second video frame is extracted from a video segment of the first video corresponding to a user attention time interval of the first video.
20. The method according to claim 19, characterized in that The generating of the second candidate video card image further includes: Determining a user focused spatial area of the first video according to the video editing information of the M videos or the K videos; The second video frame is cropped according to the user's attention space area of the first video.
21. The method according to claim 17, wherein Extracting a second video frame from the first video includes: determining a user attention time interval of the first video according to the video editing information of the first video; The second video frame is extracted from a video segment of the first video corresponding to a user attention time interval of the first video.
22. The method according to claim 21, characterized in that The generating of the second candidate video card image further includes: determining a user focused spatial area of the first video according to the video editing information of the first video; The second video frame is cropped according to the user's attention space area of the first video.
23. The method according to any one of claims 1 to 22, characterized in that The method further comprises: One or more candidate video card images are selected from the multiple candidate video card images as the video card images of the first video.
24. The method according to claim 23, wherein The step of selecting one or more candidate video card images from the plurality of candidate video card images as the video card image of the first video includes: Sorting the plurality of candidate video card images according to a preset rule, and displaying the plurality of candidate video card images to the video publisher according to the sorting result; Determine the alternative video card image selected by the video publisher among the multiple alternative video card images, and use the alternative video card image selected by the video publisher as the video card image of the first video.
25. The method according to claim 24, characterized in that Sorting the plurality of candidate video card images according to a preset rule includes: The pictures in the plurality of candidate video card images are sorted according to the similarity calculation result and / or the correlation calculation result, wherein: The similarity calculation result includes the similarity between the multiple candidate video card images and the video card images of the K videos, or the similarity between the multiple candidate video card images and a selected video card image from the video card images of the K videos; The relevance calculation result includes the relevance between the multiple candidate video card images and the video text description of the first video.
26. The method according to claim 25, characterized in that Sorting the pictures in the plurality of candidate video card pictures according to the similarity calculation result and the correlation calculation result, including: Calculating similarities between the plurality of candidate video card images and the video card images of the K videos, and obtaining the similarity calculation result; Calculate the relevance between the video title of the first video and the plurality of candidate video card images, and obtain the relevance calculation result; The plurality of candidate video card images are sorted according to the similarity calculation result and / or the correlation calculation result.
27. The method according to claim 25, characterized in that Sorting the pictures in the plurality of candidate video card pictures according to the similarity calculation result and the correlation calculation result, including: Displaying a video text description of the first video, wherein the video text description includes a plurality of alternative video titles of the first video; Determine the alternative video titles selected by the video publisher; Calculating similarities between the multiple candidate video card images and the video card images of the K videos, and obtaining the similarity calculation result; Calculating the relevance between the plurality of candidate video card images and the candidate video titles selected by the video publisher, and obtaining the relevance calculation result; The plurality of candidate video card images are sorted according to the similarity calculation result and / or the correlation calculation result.
28. A method for generating a video card image, characterized in that: The method is applied to an electronic device, and includes: Determining a user attention time interval and a user attention spatial area of the second video based on a user interaction behavior record of the second video; extracting a fifth video frame from a video segment of the second video corresponding to a user attention time interval of the second video; The fifth video frame is cropped according to the user's attention space area of the second video to generate a third candidate video card image.
29. An electronic device, characterized in that: The electronic device comprises a memory for storing computer program instructions and a processor for executing the computer program instructions, wherein when the computer program instructions are executed by the processor, the electronic device is triggered to perform the method steps of any one of claims 1 to 27 or claim 28.
30. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed on a computer, enables the computer to execute the method according to any one of claims 1 to 27 or claim 28.
Citation Information
Patent Citations
Video publishing method, device and equipment and storage medium
CN111491202A
Video cover determination method and device, medium and equipment
CN112800276A
Video cover recommendation method, device and equipment and computer readable storage medium
CN115706836A
Content generation method and device, electronic equipment and storage medium
CN117221622A
Systems and methods for content curation in video based communications
US20180359530A1