Dynamic cover generation method and device, electronic equipment and storage medium
By segmenting the target video and using intelligent capture network matching, dynamic covers that match the user's intent are generated, solving the problem of the limited variety of dynamic cover generation methods and improving the video's attractiveness and interactive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MIGU CO LTD
- Filing Date
- 2025-12-01
- Publication Date
- 2026-04-14
AI Technical Summary
The existing methods for generating dynamic covers are rather simplistic, resulting in poor video appeal.
By acquiring the target video and the cover description information input by the target object, the video is segmented, and an intelligent capture network is used to match the cover description information and video segments to generate a dynamic cover that matches the target object's intent.
It improves the intelligence of dynamic cover generation, enhances the appeal of videos, meets users' real intentions, and improves the interactive experience and reach of videos.
Smart Images

Figure CN121865068A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to a method, apparatus, electronic device, and computer storage medium for generating dynamic covers. Background Technology
[0002] Currently, the dynamic cover of a video is generated by the platform selecting several highlight video clips as cover candidates based on the video content. Users then choose a dynamic cover from these candidates. However, the generation method of this dynamic cover is relatively simple, resulting in poor video appeal. Summary of the Invention
[0003] This application provides a method, apparatus, electronic device, and computer storage medium for generating dynamic covers.
[0004] The technical solution of this application is implemented as follows: This application provides a method for generating a dynamic cover, the method comprising: Obtain the target video and the cover description information input by the target object for the target video; The target video is segmented to obtain multiple video segments; The cover description information and the multiple video clips are matched using an intelligent capture network to obtain matching results; Based on the matching results, a dynamic cover image corresponding to the target video is generated.
[0005] This application also provides a dynamic cover generation apparatus, the apparatus comprising: The acquisition module is used to acquire the target video and the cover description information input by the target object for the target video; The segmentation module is used to segment the target video to obtain multiple video segments; The matching module is used to match the cover description information and the multiple video clips using an intelligent capture network to obtain matching results; The generation module is used to generate a dynamic cover corresponding to the target video based on the matching results.
[0006] This application provides an electronic device, the device including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the dynamic cover generation method provided by one or more of the foregoing technical solutions.
[0007] This application provides a computer storage medium storing a computer program; when the computer program is executed, it can implement the dynamic cover generation method provided by one or more of the aforementioned technical solutions.
[0008] This application provides a computer program product, including a computer program that, when executed by a processor, implements the dynamic cover generation method provided by one or more of the aforementioned technical solutions.
[0009] This application provides a method, apparatus, electronic device, and computer storage medium for generating dynamic cover images. The method includes: acquiring a target video and cover description information input by a target object for the target video; segmenting the target video to obtain multiple video segments; using an intelligent capture network to match the cover description information and the multiple video segments to obtain a matching result; and generating a dynamic cover image corresponding to the target video based on the matching result.
[0010] As can be seen, in this embodiment of the application, since the cover description information is input by the target object itself for the target video, the cover description information can reflect the true intention of the target object. By matching the cover description information with multiple segmented video segments, a video segment that conforms to the true intention of the target object can be obtained. Using this video segment to generate a dynamic cover can improve the intelligence level of dynamic cover generation, enhance the video's attractiveness, and thus solve the problem of the current single dynamic cover generation method, which leads to poor video attractiveness. Attached Figure Description
[0011] Figure 1 A flowchart illustrating a dynamic cover generation method provided in this application embodiment; Figure 2 This application provides a schematic diagram of a scenario for inputting cover description information. Figure 3 This is a schematic diagram illustrating another scenario for inputting cover description information, provided in an embodiment of this application. Figure 4 A flowchart illustrating another dynamic cover generation method provided in this application embodiment; Figure 5 A flowchart for obtaining matching results is provided in this application embodiment; Figure 6 Another flowchart for obtaining a matching result provided in this application embodiment; Figure 7A A network architecture diagram of a spatial-temporal attention layer provided in this application embodiment; Figure 7B Another network architecture diagram of the spatial-temporal attention layer provided in this application embodiment; Figure 7C A network architecture diagram of another spatial-temporal attention layer provided in this application embodiment; Figure 8 Another flowchart for obtaining a matching result provided in this application embodiment; Figure 9 A network architecture diagram of a multi-scale self-attention layer provided in this application embodiment; Figure 10 This is a schematic diagram of the composition structure of the dynamic cover generation device according to an embodiment of this application; Figure 11 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions in this application will now be clearly and completely described with reference to the accompanying drawings.
[0013] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are merely illustrative of the present application and are not intended to limit the present application. Furthermore, the embodiments provided below are some embodiments for implementing the present application, and not all embodiments for implementing the present application. Unless otherwise specified, the technical solutions described in the present application can be implemented in any combination.
[0014] It should be noted that, in this application, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method or apparatus that includes a list of elements includes not only the elements expressly described, but also other elements not expressly listed, or elements inherent to implementing the method or apparatus. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other related elements (e.g., steps in the method or units in the apparatus, such as portions of a processor, program, or software, etc.) in the method or apparatus that includes that element.
[0015] For example, the dynamic cover generation method provided in this application includes a series of steps, but the dynamic cover generation method provided in this application is not limited to the steps described. Similarly, the dynamic cover generation device provided in this application includes a series of modules, but the dynamic cover generation device provided in this application is not limited to the modules explicitly described, but may also include modules that need to be set for obtaining relevant information or processing based on information.
[0016] In some embodiments of this application, the dynamic cover generation method can be implemented using a processor in the dynamic cover generation device. The processor can be at least one of the following: Application Specific Integrated Circuit (ASIC), Digital Signal Processor (DSP), Digital Signal Processing Device (DSPD), Programmable Logic Device (PLD), Field Programmable Gate Array (FPGA), CPU, controller, microcontroller, and microprocessor.
[0017] It should be noted that the collection, use, storage, sharing and transfer of user personal information involved in the technical solution of this application all comply with the provisions of relevant laws and regulations, and require notification to users and obtaining their consent or authorization. Where applicable, user personal information has been subjected to de-identification and / or anonymization and / or encryption technical processing.
[0018] Figure 1 A flowchart of a dynamic cover generation method provided in this application embodiment is shown below. Figure 1 As shown, the method includes the following steps: Step 100: Obtain the target video and the cover description information input by the target object for the target video.
[0019] In this embodiment of the application, the executing entity of the dynamic cover generation method can be a terminal device; the terminal device can include, but is not limited to, electronic devices such as mobile phones, tablets, personal computers (PCs), and vehicle terminals; the terminal device has a video application capable of playing videos.
[0020] For example, the target video refers to the video for which a dynamic cover needs to be generated; the target video can be an image sequence consisting of multiple image frames.
[0021] For example, the target audience can be users, operators, or other relevant personnel with a need for dynamic cover generation; the cover description information refers to the requirements of the target audience for the visual presentation of the dynamic cover of the target video; for example, the cover description information can be "use the goal scored by athlete A as the cover, with afterimage effects."
[0022] In this embodiment of the application, in order to improve the convenience and flexibility of target object interaction, various input methods can be supported for the cover description information, namely text input, voice input, text and image combination input, voice and image combination input, etc., without specific limitations here.
[0023] For example, the above input methods can be implemented through corresponding input controls, wherein text input corresponds to a text input control, and voice input corresponds to a voice input control; that is, text input controls and / or voice input controls can be displayed on the relevant interface of the target video.
[0024] To facilitate understanding, this section combines... Figure 2 and Figure 3 For further explanation, please refer to [link / reference]. Figure 2 The ballpoint pen icon represents a text input control, and the microphone icon represents a voice input control. In response to the target object's trigger operation on the text input control, the cover description information input by text input can be obtained. This cover description information is in text form and can be something like "Find a spectacular vaulting athlete's move and add a ghosting effect." Similarly, in response to the target object's trigger operation on the voice input control, the cover description information input by voice input can be obtained. This cover description information is in voice form.
[0025] For example, the input method for the cover description information can also be a combination of text and image input or a combination of voice and image input; for example, see Figure 3 In response to the target object's selection of an image frame in the target video and the triggering of a text input control, cover description information input using a combination of text and image input can be obtained. This cover description information is in text form and can be something like "Select the main shot of this character in this episode". Similarly, in response to the target object's selection of an image frame in the target video and the triggering of a voice input control, cover description information input using a combination of voice and image input can be obtained. This cover description information is in voice form.
[0026] It should be noted that when obtaining the cover description information in voice form, technologies such as automatic speech recognition can be used to automatically convert the voice-based cover description information into text-based cover description information for easier subsequent processing; that is, the final cover description information obtained is in text form.
[0027] Step 101: Segment the target video to obtain multiple video segments.
[0028] In this embodiment of the application, after obtaining the target video, the target video can be segmented to obtain multiple video segments for subsequent matching.
[0029] Here, the segmentation method for the target video is not specifically limited; in one embodiment, if the target video includes multiple shots, the target video can be segmented according to the shots to obtain multiple video segments.
[0030] For example, if the target video includes multiple shots, it means that the target video is not a fixed perspective, that is, there are shot transitions. In this case, the shot transition points in the target video can be detected, and the target video can be segmented according to the shot transition points to obtain multiple video segments.
[0031] Each shot corresponds to a video segment, and each video segment has a corresponding time interval; that is, each shot has a corresponding time interval. For example, the video segments corresponding to each shot in the target video can be represented as follows: , where [t i ,t i+1 ) represents the video segment corresponding to the i-th shot, which includes multiple image frames corresponding to the i-th shot.
[0032] In another embodiment, if the target video includes a single shot, the target video can be segmented according to a set number of frames to obtain multiple video segments.
[0033] Here, the set frame number is also called the fixed frame number, and its value can be set according to the actual situation. This application embodiment does not make specific limitations on this.
[0034] For example, if the target video consists of a single shot, it means that the target video is a fixed perspective, that is, there is no shot switching. In this case, the target video can be segmented according to a set number of frames to obtain multiple video segments.
[0035] Step 102: Use the intelligent capture network to match the cover description information with multiple video clips to obtain the matching results.
[0036] In this embodiment of the application, after obtaining the cover description information and multiple segmented video segments, an intelligent capture network can be used to match the cover description information and multiple video segments to obtain the matching result.
[0037] For example, the matching result may include at least one target video segment that matches the cover description information, wherein the at least one target video segment includes one or more target video segments. It should be noted that the target video segment may be one of the multiple video segments, or it may be a portion of the video segments.
[0038] For example, assuming each video segment includes 10 image frames, if at least one target video segment includes two target video segments, then these two target video segments can each include 10 image frames, or each include 5 image frames, or one includes 10 image frames and the other includes 5 image frames, without specific limitations here.
[0039] For example, the matching methods of the intelligent capture network may include hard matching, soft matching, and cross-scale matching; the following are exemplary descriptions of these three matching methods.
[0040] In some embodiments, the above method may further include: using a language model to identify the cover description information to obtain a first identification result; identifying different modal content in each video segment to obtain a second identification result; and correspondingly, using an intelligent capture network to match the cover description information and multiple video segments may include: using an intelligent capture network to match the first identification result and the second identification result.
[0041] For example, the language model can be an open-source model or a closed-source model, and it can be trained without training or fine-tuned. Here, there is no specific limitation on the type of language model. For example, it can be a Bidirectional Encoder Representations from Transformers (BERT) model or a Generative Pre-trained Transformer (GPT) model, etc.
[0042] In this embodiment of the application, a language model can be used to identify the cover description information to obtain a first identification result. The first identification result may include at least one key piece of information in the cover description information; that is, one or more key pieces of information; wherein, the key information may include, but is not limited to, time information, action information, character information and motion effect information.
[0043] For example, see Figure 4 The process on the left involves obtaining the text-based cover description from the user. This text is then input into a language model, which processes it to obtain time, action, character, and animation information. It's worth noting that user input can also include selected images. As you can see, the language model breaks down the user-input cover description into clearly defined key pieces of information.
[0044] For example, suppose the cover description is "Use player A's goal around 16 minutes into the cover, with afterimage effect added". After language model recognition, the corresponding time information can be obtained: 15:00-17:00, action information: goal, character information: player A, and motion effect information: afterimage.
[0045] In this embodiment of the application, after obtaining multiple video segments after segmenting the target video, different modal contents in each video segment can be identified to obtain a second identification result.
[0046] Each video segment may include at least one modal content, i.e., one or more modal contents; here, at least one modal content may include at least one of image content, audio content and text content; the second recognition result may include recognition results corresponding to different modal contents.
[0047] For example, after obtaining each video segment, the modal content of each video segment can be identified and processed; see [link to relevant documentation]. Figure 4 The process on the right, assuming the target video is segmented into various video clips based on camera angles, allows for character and motion recognition of the image content within each clip, speech recognition of the audio content, and optical character recognition (OCR) of the text content. The recognition results are then input into an intelligent capture network for hard matching. Here, the text content can include subtitles, on-screen text information, and, for example, the large amounts of text found in educational videos.
[0048] For example, suppose I ti,ti+1 S represents the image content of the i-th video segment. ti,ti+1 T represents the speech content of the i-th video segment. ti,ti+1 The text representing the i-th video segment is obtained by analyzing the image content I. ti,ti+1 Performing person recognition and motion recognition yields person recognition results and motion recognition results, denoted as IFR, respectively. ti,ti+1 and IAR ti,ti+1 ; By analyzing the voice content S ti,ti+1 Speech recognition can be performed to obtain the speech recognition result, which is represented as SR. ti,ti+1 ; By analyzing the text content T ti,ti+1 OCR recognition can be performed to obtain the OCR recognition result, which is represented as TR. ti,ti+1 In the recognition result, the character R represents the result, the character F represents the face, and the character A represents the action.
[0049] Furthermore, the recognition results of each video segment can be assembled according to the time sequence relationship to obtain the following: Result_IFR:{…,IFR ti,ti+1 :[Athlete A,…],…}; Result_IAR:{…,IAR ti,ti+1 :[Goal,…],…}; Result_SR:{…,SR ti,ti+1 :[Goal!,...],...}; Result_TR:{…,TR ti,ti+1 :[Goal!,...],...}
[0050] It should be noted that when the user input includes a selected image, the selected image can also be used for person recognition and OCR recognition to obtain the recognition result Result_BR. The recognition result Result_BR includes the person recognition result IFR and the OCR recognition result TR. In the recognition result, the character R represents the result, the character B represents the box, the character F represents the face, and the character T represents the text in the selected image.
[0051] For example, by assembling the above recognition results, a second recognition result can be obtained, which can be represented as Result_AI. When the user input includes a selected image, Result_AI can be represented as: [Result_IFR, Result_IAR, Result_SR, Result_TR, Result_BR]; when the user input does not include a selected image, Result_AI can be represented as: [Result_IFR, Result_IAR, Result_SR, Result_TR].
[0052] In this embodiment of the application, after obtaining the first identification result and the second identification result according to the above steps, an intelligent capture network can be used to match the first identification result and the second identification result.
[0053] For example, the above matching method can be as follows: First, use an intelligent capture network to match the time information in the first recognition result with the time interval corresponding to each video segment in multiple video clips to obtain one or more target video clips that match the time information; then, match the first recognition result with the second recognition result corresponding to each target video clip. For ease of understanding, the following is combined with... Figure 5 An example is provided.
[0054] For example, see Figure 5 After obtaining one or more target video segments that conform to time information from multiple video segments after shot segmentation, the person information and action information in the first recognition result can be matched with the second recognition result Result_AI corresponding to each target video segment. This result can include person recognition result, action recognition result, speech recognition result, and OCR recognition result to obtain the matching result, which can be represented as Result_Hard. For example, the matching result Result_Hard can be {…,t i ,t i+1 : 2,…};where the value 2 represents the target video segment [t i ,t i+1 The number of matches between the second recognition result Result_AI and the key information in the cover description information is 2. For example, the second recognition result Result_AI matches the person information and action information in the cover description information. It should be noted that the larger this value is, the higher the matching degree of the corresponding target video segment. The matching result can include the number of key information matches for each target video segment.
[0055] For example, the above matching method is hard matching. As can be seen from the above, the matching method of the intelligent capture network can also be soft matching. The soft matching process is described below as an example.
[0056] In some embodiments, using an intelligent capture network to match cover description information and multiple video segments to obtain matching results may include: using a text processing network to perform feature processing on the cover description information to obtain text features; using a video processing network to perform feature processing on each video segment and text features to obtain cross-attention features; using a classifier to classify the cross-attention features of each video segment to obtain a matching score for each video segment; the matching results include the matching score for each video segment.
[0057] In this embodiment, the intelligent capture network may include a text processing network, a video processing network, and a classifier. After obtaining cover description information in text form and multiple segmented video clips, the cover description information can be input into the text processing network, and each video clip can be input into the video processing network. For ease of understanding, the following describes the process in conjunction with... Figure 6 The above soft matching process will be explained.
[0058] For example, see Figure 6 Text processing networks may include a text encoding layer (EmbeddingText), a bidirectional self-attention layer, and a feedforward network layer; video processing networks may include a video encoding layer (EmbeddingVideo), a spatial-temporal attention layer, a cross-attention layer, and a feedforward network layer.
[0059] For example, using a text processing network to extract features from the cover description information to obtain text features may include: first, encoding the cover description information using a text encoding layer to obtain text encoding features, and simultaneously obtaining the positional encoding corresponding to the cover description information; then, performing N feature calculations through a bidirectional self-attention layer and a feedforward network to finally output the text features, which can be represented as F. text N is an integer greater than 1.
[0060] Here, positional encoding can be obtained by using positional encoding methods such as Rotary Positional Embeddings (RoPE). By adding positional encoding to text encoding features, the model can understand the positional information of elements in the text encoding features. Positional encoding is usually a vector that is added to the embedding of each element in the text encoding features.
[0061] Specifically, after each pass through the bidirectional self-attention layer and the feedforward network, the output dimension is L. text ×d text intermediate layer features, where L text d represents the length of the text describing the cover information, i.e., the input dimension; d represents the feature dimension; during feature calculation, the network consisting of the bidirectional self-attention layer and the feedforward network layer is repeated N times to obtain the final text feature F. text .
[0062] For example, a video processing network is used to perform feature processing on each video segment and text feature to obtain cross-attention features, including: encoding each video segment using a video coding layer to obtain video coding features; performing spatial and temporal attention calculations on the video coding features of each video segment using a spatial-temporal attention layer to obtain spatial-temporal features; and performing cross-attention calculations on the spatial-temporal features and text features of each video segment using a cross-attention layer to obtain cross-attention features.
[0063] For example, see Figure 6For each video segment, a video coding layer can be used to encode each video segment to obtain video coding features. At the same time, the positional code corresponding to the video segment can be obtained, and the method of obtaining the positional code corresponding to the cover description information is similar to that of obtaining the positional code. It will not be repeated here. Here, the video segment input to the video coding layer can be represented as T×H×W, where T is the image frame in a certain video segment, and H and W are the height and width of the image frame, respectively. In the spatial dimension, the image frames in the video segment can be divided into nh×nw spatial units with h×w as the basic unit. In the temporal dimension, the frame interval can be set to t', that is, sampling is performed at an interval of t', compressing T image frames into nti time units. Here, nti=T / / t', nh=H / / h, nw=W / / w, where the symbol " / / " represents integer division. As can be seen, the video coding features output by the video coding layer can be represented as nti×nh×nw×d; where i represents the i-th video segment and nti represents the number of image frames in the i-th compressed video segment.
[0064] Depend on Figure 6 It can be seen that the processing flow for video clips differs from that for cover description information in that the core network of the video processing network is a spatial-temporal attention layer. It should be noted that the spatial-temporal attention layer can employ various network architectures, such as... Figures 7A to 7C As shown.
[0065] For example, see Figure 7A For the video coding features of each video segment, attention calculations at multiple spatial scales can be performed uniformly first, followed by attention calculations at multiple temporal scales to obtain the spatial-temporal features; see [link to relevant documentation]. Figure 7B Alternatively, attention calculations at both spatial and temporal scales can be performed together. Both attention calculation methods use (nti)×(nh•nw•d) as the input for the spatial scale and (nh•nw)×(nti•d) as the input for the temporal scale.
[0066] For example, see Figure 7C Furthermore, different multi-head attention mechanisms can be used to calculate the attention weights of each video coding feature at both spatial and temporal scales, constructing... , That is, the keys and values corresponding to these scales, for half of the attention heads, the spatial scale is through To calculate, the time dimension is through To calculate, where Q, K, and V are all inputs, through different transformation matrices W q W k W vMultiplying them together, we get the character 's' representing the spatial dimension input and the character 't' representing the temporal dimension input; finally, we combine the outputs of multiple attention heads. Concat is used for matrix concatenation. It is also a transformation matrix.
[0067] It should be noted that the spatial-temporal attention layers of the three different network architectures mentioned above all have a final output spatial-temporal feature dimension of (nti•nh•nw)×d.
[0068] Furthermore, after obtaining text features (L) according to the above steps text The spatial-temporal features of each video segment ((nti•nh•nw)×d) and can be denoted as L video After (×d), these two features can be input into the cross-attention layer; that is, the cross-attention layer is used to calculate the cross-attention of the spatial-temporal features and text features of each video segment to obtain the cross-attention features of each video segment; correspondingly, the cross-attention layer can be used to calculate the cross-attention (Q) of the text features and spatial-temporal features. video ,K text V text ), where text represents text-dimensional input and video represents video-dimensional input.
[0069] For example, see Figure 6 The network module, consisting of a spatial-temporal attention layer, a cross-attention layer, and a feedforward network layer, is repeated N times to obtain the final cross-attention feature F. video Then, for the cross-attention feature F video Perform feature fusion to make the cross-attention feature F video The dimension is changed from (nti•nh•nw)×d to 1×d, and finally the fused cross-attention feature F is... video The data is fed into a classifier, which classifies the fused cross-attention features of each video segment to obtain a matching score for each segment. The matching result can include the matching score for each video segment, and can be represented as Result_Soft:{…,t i ,t i+1 :0.8,…}, where the value 0.8 represents the i-th video segment [t i ,t i+1 The higher the matching score, the better the corresponding video clip matches the requirements of the cover description information.
[0070] For example, as can be seen from the above, the matching method of the intelligent capture network can also be cross-scale matching. The cross-scale matching process is described below as an example.
[0071] In some embodiments, the above method may further include: after obtaining the cross-attention features of each video segment, using a multi-scale self-attention layer to perform feature processing on the cross-attention features of each sub-video segment in each video segment to obtain video features; each video segment includes one or more sub-video segments; using a classifier to classify the video features of each sub-video segment to obtain a matching score for each sub-video segment; the matching result includes the matching score of each sub-video segment.
[0072] Understandably, the hard matching process and soft matching process described above judge each video segment as a benchmark to determine whether it can be used to generate a dynamic cover; however, in practice, it may be necessary to combine segments from multiple video segments to generate a dynamic cover, which can be achieved through cross-scale matching of intelligent capture networks.
[0073] Among them, the network structure corresponding to cross-scale matching is as follows: Figure 8 As shown: For ease of understanding, the following is combined with... Figure 8 The above cross-scale matching process will be explained.
[0074] For example, see Figure 8 The intelligent capture network can also include multi-scale self-attention layers, which obtain the cross-attention features F of each video segment according to the soft matching process described above. video Then, the cross-attention feature F can be... video The input is a multi-scale self-attention layer, that is, the multi-scale self-attention layer is used to process the cross-attention features of each sub-video segment in each video segment to obtain video features; wherein, each video segment includes one or more sub-video segments; at the same time, the position code corresponding to each video segment is obtained, and the method of obtaining the position code corresponding to the cover description information is similar to that of obtaining the position code, which will not be repeated here.
[0075] Depend on Figure 6 and Figure 8 It can be seen that the difference between the soft matching process and the cross-scale matching process lies in the multi-scale self-attention layer; the network architecture of the multi-scale self-attention layer can be as follows: Figure 9 As shown.
[0076] For example, Figure 9 A four-layer multi-scale self-attention layer is shown, where the input to this multi-scale self-attention layer is no longer a single video segment, but rather the entire target video, i.e., all video segments, as input; for example... Figure 9As shown in the figure, one square represents a sub-video segment composed of multiple frames. Each video segment can include one or more sub-video segments. Different video segments are distinguished by squares of different patterns. For example, squares 0 to 1 are sub-video segments in a video segment, squares 2 to 3 are sub-video segments in a video segment, squares 4 to 5 are sub-video segments in a video segment, and so on.
[0077] For example, tt' can be set as the number of frames in a sub-video segment, and ntall / / tt' represents the number of sub-video segments, where ntall represents the total number of frames in the target video. If tt'=1, it means that one frame is considered as a unit, that is, a sub-video segment includes one image frame.
[0078] It should be noted that before using the multi-scale self-attention layer to process the cross-attention features of each sub-video segment in each video clip, the cross-attention features of each sub-video segment can be fused first by using a transformation matrix, or they can be fused after passing through the final feedforward network layer to obtain the video features of each sub-video segment. The dimension of the video features is (ntall / / tt')×d.
[0079] For example, the first layer in a multi-scale self-attention layer is no different from a regular attention layer, except that it selects a window size L. window Calculate attention for each image frame instead of all image frames. Figure 9 L window The value of is 5 because the value of (ntall / / tt') is relatively large, and the total computational load is extremely large. Using a window can greatly reduce the computational load. Starting from the second layer, the sampling interval r can be set. The sampling interval r increases exponentially by 2^n, where n is the layer number 0, 1, ..., N-1, forming an image frame interval sequence of 1, 2, 4, 8...; r is the attention selection element for the interval. Only the elements in the window are selected for each layer to calculate the attention color to represent the relative position, and so on.
[0080] Correspondingly, a multi-scale self-attention layer can be used to perform multi-scale attention calculation on the cross-attention features of each sub-video segment in each video segment, resulting in multi-scale features MultiScaleAttention(Q). i ,K i+rk V i+rk ), where i+rk refers to Q i K achievable with a sampling interval of r 1:Lvideo The position in the sequence; when i+rk exceeds the sequence range [1,L] videoThe process is repeated N times, using -inf for padding. The module, composed of multi-scale self-attention layers and feedforward network layers, yields the final video feature Fscale. Finally, the video feature Fscale is fed into a classifier, which classifies the video features of each sub-video segment to obtain a matching score for each segment. The matching result includes the matching score for each sub-video segment. The matching result can then be represented as Result_Scale:{…,tti,tti+1:0.8,…}, where the value 0.8 represents the matching score for the i-th sub-video segment [tti,tti+1). A higher matching score indicates that the corresponding sub-video segment better meets the requirements of the cover description information.
[0081] It should be noted that the network architecture used in the intelligent capture network to achieve soft matching and cross-scale matching needs to be trained. When training the network architecture for soft matching, the input can be each video segment and a label indicating whether the video segment can be used as a dynamic cover, which can be represented by 0 and 1. When training the network architecture for soft matching, the input can be the cross-attention features of all video segments and a label indicating whether the corresponding video segment can be used as a dynamic cover, which can be represented by 0 or 1. The model parameters will be adjusted during training, but once training is complete, the model parameters will no longer be adjusted during inference.
[0082] Step 103: Generate a dynamic cover corresponding to the target video based on the matching results.
[0083] In this embodiment of the application, after obtaining the matching result output by the intelligent capture network, a dynamic cover corresponding to the target video can be generated based on the matching result.
[0084] For example, as can be seen from the above, the matching results output by the intelligent capture network may include one or more of the following: the matching result Result_Hard corresponding to the hard matching method, the matching result Result_Soft corresponding to the soft matching method, and the matching result Result_Scale corresponding to the cross-scale matching method.
[0085] It should be noted that the above three matching results can be used individually, or they can be combined by taking the intersection of the time dimension, or the three matching results can be further fused to obtain the final result.
[0086] For example, in the case of using the matching result Result_Hard, see Figure 4If the cover description information includes motion effect information, after obtaining the matching result Result_Hard output by the intelligent capture network, the target video segment with the most matching key information in the matching result Result_Hard and the motion effect information can be input into the dynamic cover generation module to generate a dynamic cover. The dynamic cover generation module can find the corresponding motion effect matching template based on the motion effect information, such as motion afterimage; it can track the person or object in the current j-th frame, extract the person or object, and copy the corresponding object in frames jk,j-(k-1),...,j-1 according to the position of each frame in the target video segment, and copy it to the j-th frame to obtain the afterimage effect. At this time, a dynamic cover that matches the intention of the target object can be obtained.
[0087] For example, when using the matching result Result_Soft, if the cover description information includes motion effect information, the video segment with the highest matching score in the matching result Result_Soft and the motion effect information can be input into the dynamic cover generation module to generate a dynamic cover. The generation process is similar to the generation process corresponding to the matching result Result_Hard mentioned above, and will not be repeated here to avoid repetition.
[0088] For example, when using the matching result Result_Scale, generating a dynamic cover corresponding to the target video based on the matching result may include: determining multiple target sub-video segments from the target video for generating the dynamic cover based on the matching score of each sub-video segment in the target video; combining the multiple target sub-video segments, and determining the combined result as the dynamic cover corresponding to the target video.
[0089] For example, as can be seen from the above, the matching result Result_Scale includes the matching score of each sub-video segment in the target video, and the higher the matching score, the better the effect as a dynamic cover; therefore, when using the matching result Result_Scale to generate a dynamic cover, multiple target sub-video segments for generating the dynamic cover can be determined from the target video based on the matching score of each sub-video segment in the target video; here, multiple target sub-video segments can come from the same video segment or from different video segments.
[0090] Understandably, multiple target sub-video segments include each sub-video segment in the matching result Result_Scale whose matching score exceeds a set threshold, or each sub-video segment whose matching score is ranked among the top few after sorting from largest to smallest.
[0091] Further, see Figure 9The final result is that after obtaining multiple target sub-video segments for generating dynamic covers, these multiple target sub-video segments can be combined to generate dynamic covers corresponding to the target videos based on the combination results.
[0092] For example, if the cover description information includes animation information, the combined result and animation information can be input into the dynamic cover generation module to generate a dynamic cover. The generation process is similar to the generation process corresponding to the matching result Result_Hard mentioned above, and will not be repeated here to avoid repetition.
[0093] As can be seen, the embodiments of this application can use multimodal information fusion to trim suitable video segments according to the needs of operators and users. These segments can be shot-by-shot or a combination of multiple segments. The video segments can also be processed to add motion effects, providing operators and users with lightweight, fast, and convenient cover animation effects, thereby increasing video reach and clicks. Furthermore, since the cover description information is input by the target object itself for the target video, it reflects the target object's true intent. Therefore, the dynamic video cover generation method provided by the embodiments of this application can directly address user needs. By matching the cover description information with multiple segmented video segments, a video segment that matches the target object's true intent can be obtained. Using this video segment to generate a dynamic cover improves the intelligence of dynamic cover generation, enhances video appeal, and solves the problem of poor video appeal caused by the current singular dynamic cover generation method. In addition, the input of the cover description information supports multimodal user input, such as voice, text, and selection, which can improve the video interaction experience.
[0094] Figure 10 This is a schematic diagram of the composition structure of the dynamic cover generation device according to an embodiment of this application, as shown below. Figure 10 As shown, the device includes: an acquisition module 300, a segmentation module 301, a matching module 302, and a generation module 303, wherein: The acquisition module 300 is used to acquire the target video and the cover description information input by the target object for the target video; The segmentation module 301 is used to segment the target video to obtain multiple video segments; The matching module 302 is used to match the cover description information and multiple video clips using an intelligent capture network to obtain the matching results; The generation module 303 is used to generate a dynamic cover corresponding to the target video based on the matching results.
[0095] In some embodiments, the intelligent capture network includes a text processing network, a video processing network, and a classifier; the matching module 302 is further configured to: A text processing network is used to perform feature processing on the cover description information to obtain text features; A video processing network is used to perform feature processing on each video segment and text feature to obtain cross-attention features; A classifier is used to classify the cross-attention features of each video segment, and a matching score is obtained for each video segment; the matching results include the matching score of each video segment.
[0096] In some embodiments, the video processing network includes a video coding layer, a spatial-temporal attention layer, and a cross-attention layer. The matching module 302 is further configured to: Each video segment is encoded using a video coding layer to obtain video coding features; Spatial-temporal attention layers are used to perform spatial and temporal attention calculations on the video coding features of each video segment to obtain spatial-temporal features; Cross-attention layers are used to perform cross-attention calculations on the spatial-temporal features and textual features of each video segment to obtain cross-attention features.
[0097] In some embodiments, the intelligent capture network further includes a multi-scale self-attention layer, and the matching module 302 is further configured to: After obtaining the cross-attention features of each video segment, a multi-scale self-attention layer is used to process the cross-attention features of each sub-video segment in each video segment to obtain video features; each video segment includes one or more sub-video segments. A classifier is used to classify the video features of each sub-video segment, and a matching score is obtained for each sub-video segment; the matching result includes the matching score of each sub-video segment.
[0098] In some embodiments, the generation module 303 is further configured to: Based on the matching score of each sub-video segment in the target video, multiple target sub-video segments are determined from the target video to generate the dynamic cover. Multiple target sub-video segments are combined, and a dynamic cover corresponding to the target video is generated based on the combination result.
[0099] In some embodiments, the above-described apparatus further includes an identification module, which is configured to: The language model is used to identify the cover description information to obtain the first identification result, which includes one or more key pieces of information. Different modalities in each video segment are identified to obtain a second identification result; the second identification result includes the identification results corresponding to different modalities. Matching module 302 is also used for: The first and second identification results are matched using an intelligent capture network.
[0100] In some embodiments, the segmentation module 301 is further configured to: The target video is segmented according to the camera angles to obtain multiple video clips; or... The target video is segmented according to a set number of frames to obtain multiple video segments.
[0101] In practical applications, the acquisition module 300, segmentation module 301, matching module 302, generation module 303 and identification module can all be implemented by a processor located in an electronic device. The processor can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller and microprocessor.
[0102] Furthermore, in this embodiment, the functional modules can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional module.
[0103] If the integrated unit is implemented as a software functional module and is not sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this embodiment, in essence, or the part that contributes to related technologies, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the method of this embodiment. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0104] Specifically, the computer program instructions corresponding to a dynamic cover generation method in this embodiment can be stored on storage media such as optical discs, hard disks, and USB flash drives. When the computer program instructions corresponding to a dynamic cover generation method in the storage media are read or executed by an electronic device, any of the dynamic cover generation methods in the aforementioned embodiments are implemented.
[0105] Based on the same technical concept as the foregoing embodiments, see Figure 11 It illustrates an electronic device 500 provided in an embodiment of this application, which may include: a memory 501 and a processor 502; wherein, Memory 501 is used to store computer programs and data; The processor 502 is configured to execute a computer program stored in the memory to implement any of the dynamic cover generation methods described in the foregoing embodiments.
[0106] In practical applications, the aforementioned memory 501 can be volatile memory, such as RAM; or non-volatile memory, such as ROM, flash memory, hard disk drive (HDD), or solid-state drive (SSD); or a combination of the above types of memory, and provide instructions and data to the processor 502.
[0107] The processor 502 described above can be at least one of ASIC, DSP, DSPD, PLD, FPGA, CPU, controller, microcontroller, and microprocessor. It is understood that, for different model training devices, the electronic device used to implement the above processor function can also be other types, and this application embodiment does not specifically limit the specific types.
[0108] This application provides a computer program product, including a computer program that, when executed by a processor, implements any of the dynamic cover generation methods described in the foregoing embodiments.
[0109] In some embodiments, the functions or modules of the apparatus provided in this application can be used to perform the methods described in the above method embodiments. The specific implementation can be referred to the description of the above method embodiments, and for the sake of brevity, it will not be repeated here.
[0110] The description of the various embodiments above tends to emphasize the differences between the various embodiments. The similarities or similarities between them can be referred to, and for the sake of brevity, they will not be repeated here.
[0111] The methods disclosed in the various method embodiments provided in this application can be arbitrarily combined to obtain new method embodiments without conflict.
[0112] The features disclosed in the various product embodiments provided in this application can be arbitrarily combined without conflict to obtain new product embodiments.
[0113] The features disclosed in the various method or device embodiments provided in this application can be arbitrarily combined without conflict to obtain new method or device embodiments.
[0114] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of hardware embodiments, software embodiments, or embodiments combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage and optical storage) containing computer-usable program code.
[0115] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0116] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0117] The above are merely preferred embodiments of this application and are not intended to limit the scope of protection of this application.
Claims
1. A method for generating a dynamic cover, characterized in that, The method includes: Obtain the target video and the cover description information input by the target object for the target video; The target video is segmented to obtain multiple video segments; The cover description information and the multiple video clips are matched using an intelligent capture network to obtain matching results; Based on the matching results, a dynamic cover image corresponding to the target video is generated.
2. The method according to claim 1, characterized in that, The intelligent capture network includes a text processing network, a video processing network, and a classifier. The process of using the intelligent capture network to match the cover description information and the multiple video clips to obtain matching results includes: The text processing network is used to perform feature processing on the cover description information to obtain text features; The video processing network is used to perform feature processing on each video segment and the text features to obtain cross-attention features; The classifier is used to classify the cross-attention features of each video segment to obtain a matching score for each video segment; the matching result includes the matching score of each video segment.
3. The method according to claim 2, characterized in that, The video processing network includes a video coding layer, a spatial-temporal attention layer, and a cross-attention layer. The process of using the video processing network to perform feature processing on each video segment and the text features to obtain cross-attention features includes: Each video segment is encoded using the video coding layer to obtain video coding features; The spatial-temporal attention layer is used to perform spatial and temporal attention calculations on the video coding features of each video segment to obtain spatial-temporal features; The cross-attention layer is used to perform cross-attention calculations on the spatial-temporal features and the text features of each video segment to obtain the cross-attention features.
4. The method according to claim 2, characterized in that, The intelligent capture network further includes a multi-scale self-attention layer, and the method further includes: After obtaining the cross-attention features of each video segment, the multi-scale self-attention layer is used to perform feature processing on the cross-attention features of each sub-video segment in each video segment to obtain video features; each video segment includes one or more sub-video segments; The classifier is used to classify the video features of each sub-video segment to obtain a matching score for each sub-video segment; the matching result includes the matching score of each sub-video segment.
5. The method according to claim 4, characterized in that, The step of generating a dynamic cover image corresponding to the target video based on the matching result includes: Based on the matching score of each sub-video segment in the target video, a plurality of target sub-video segments are determined from the target video for generating the dynamic cover; The multiple target sub-video segments are combined, and a dynamic cover corresponding to the target video is generated based on the combination result.
6. The method according to claim 1, characterized in that, The method further includes: The cover description information is identified using a language model to obtain a first identification result, which includes one or more key pieces of information. Different modalities in each video segment are identified to obtain a second identification result; the second identification result includes the identification results corresponding to different modalities. The step of matching the cover description information and the multiple video clips using a smart capture network includes: The intelligent capture network is used to match the first identification result and the second identification result.
7. The method according to claim 1, characterized in that, The target video is segmented to obtain multiple video segments, including: The target video is segmented according to the camera angles to obtain multiple video clips; or... The target video is segmented according to a set number of frames to obtain multiple video segments.
8. A dynamic cover generation device, characterized in that, The device includes: The acquisition module is used to acquire the target video and the cover description information input by the target object for the target video; The segmentation module is used to segment the target video to obtain multiple video segments; The matching module is used to match the cover description information and the multiple video clips using an intelligent capture network to obtain matching results; The generation module is used to generate a dynamic cover corresponding to the target video based on the matching results.
9. An electronic device, characterized in that, The device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the method according to any one of claims 1 to 7.
10. A computer storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the computer program implements the method described in any one of claims 1 to 7.