Method and device for generating video with edge-pasted picture
By extracting content text from the video and generating side images associated with the video, the problem of poor video quality in the existing technology is solved, personalized and rich video generation is achieved, and the video generation effect is improved.
Patent Information
- Application Number
- CN202410282374.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-12
- Publication Date
- 2025-09-12
AI Technical Summary
The video effect generated by the existing technology is poor, the side-mounted pictures are single in form and lack personalization, resulting in poor video generation effect.
The content text is extracted from the first video, and a side-posting image associated with the video content is generated based on the text. The side-posting image is synthesized with the video, and the richness and personalization of the side-posting image are improved through the image recognition model and the image generation model.
It improves the video generation effect, enhances the matching degree between the side-posted pictures and the video content, realizes personalized and rich video generation, and reduces the user operation complexity and labor costs.
Smart Images

Figure CN120640088A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of computer vision technology, and more particularly to a method and device for generating a video with side-mounted images. Background Art
[0002] In the field of computer vision, content is displayed through images, videos, and other means. For example, advertisements can be displayed through videos. Videos offer a more intuitive user experience and are widely used in various scenarios.
[0003] In the prior art, videos can be further processed, for example, by adding text, images, audio, etc. to the video to enrich the content of the processed video. The text can be input through the keyboard, and the images and audio can be selected by the user from a preset library. Both images and text can include text.
[0004] However, the above solution has the problem that the generated video effect is poor. Summary of the Invention
[0005] The embodiments of the present disclosure provide a method and device for generating a video with side-mounted images, which can improve the generation effect of the video.
[0006] In a first aspect, an embodiment of the present disclosure provides a method for generating a video with a side-mounted image, comprising:
[0007] Get the first video;
[0008] Extracting text describing content in the first video from the first video;
[0009] generating a side-posting image based on the text, wherein the text is used to determine the attributes and / or content of the side-posting image;
[0010] The first video and the side-mounted image are synthesized to obtain a second video.
[0011] In a second aspect, an embodiment of the present disclosure provides a video generation device with a side-mounted image, comprising:
[0012] A first video acquisition module, configured to acquire a first video;
[0013] A text extraction module, configured to extract text describing the content of the first video from the first video;
[0014] A side-posting picture generation module, configured to generate a side-posting picture based on the text, wherein the text is used to determine the attributes and / or content of the side-posting picture;
[0015] The second video generation module is used to synthesize the first video and the side-mounted image to obtain a second video.
[0016] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: at least one processor and a memory;
[0017] The memory stores computer-executable instructions;
[0018] The at least one processor executes the computer-executable instructions stored in the memory, so that the electronic device implements the method according to the first aspect.
[0019] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the computing device implements the method described in the first aspect.
[0020] In a fifth aspect, an embodiment of the present disclosure provides a computer program, which is used to implement the method described in the first aspect.
[0021] The disclosed embodiments provide a method and device for generating a video with a side-image. The method and device can obtain a first video; extract text describing the content of the first video from the first video; generate a side-image based on the text, where the text is used to determine the attributes and / or content of the side-image; and synthesize the first video and the side-image to generate a second video. The disclosed method can generate a side-image based on the text describing the content of the first video, thereby improving the matching of the attributes and / or content of the side-image with the first video and thus enhancing the quality of the second video generation. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0023] Figure 1 This is a schematic diagram of the relationship between a first video, a side image, and a second video provided by the present disclosure;
[0024] Figure 2 This is a schematic diagram of a video generation process provided by the present disclosure;
[0025] Figure 3 This is a flowchart of the steps of a method for generating a video with side-mounted images provided by an embodiment of the present disclosure;
[0026] Figure 4 This is a schematic diagram of a video generation process provided by an embodiment of the present application;
[0027] Figure 5 This is a flowchart of another method for generating a video with side-mounted images provided by an embodiment of the present disclosure;
[0028] Figure 6 This is a structural block diagram of a video generation device with side-mounted pictures provided by an embodiment of the present disclosure;
[0029] Figure 7 This is a structural block diagram of an electronic device provided by an embodiment of the present disclosure. DETAILED DESCRIPTION
[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.
[0031] The present disclosure is used to add a sidebar image to a first video to form a new second video. The sidebar image is positioned around the periphery of the first video, so that each video frame of the second video includes the sidebar image and a video frame of the first video located within the sidebar image. The sidebar image can be understood as a border around the first video, similar to a photo frame.
[0032] It is understood that the first video is any video without adding a side-posting picture, and the second video is a video obtained by adding a side-posting picture to the first video. The first video can be any video, for example, it can be a UGC (User Generated Content) video.
[0033] Figure 1 FIG1 is a schematic diagram of the relationship between a first video, a side-posted image, and a second video provided by the present disclosure. Referring to FIG1 , the video frame of the first video is located inside the side-posted image, and the two together form the video frame of the second video.
[0034] In some scenarios, a user may select a side-posting picture from a side-posting picture library and add it to the periphery of the first video to form a second video. Figure 2 This is a schematic diagram of a video generation process provided by the present disclosure. As shown in Figure 2, the video generation process includes two phases: preparation and generation. The preparation phase includes two sub-phases: image design and image upload. The generation phase includes two sub-phases: image selection and video generation.
[0035] In the above-mentioned side-mounted image design sub-stage, the designer can design the side-mounted image through AE (Adobe After Effects, a graphics and video processing software) software.
[0036] In the above-mentioned side-sticker picture uploading sub-stage, the designer uploads the designed side-sticker picture to the side-sticker picture library for storage.
[0037] In the above-mentioned side-posting picture selection sub-stage, the user logs in to the client to select the side-posting picture to be used from the above-mentioned side-posting picture library, and selects the first video on the client.
[0038] In the above-mentioned video generation sub-stage, the user operates on the client to confirm the use of the side-mounted image and the first video to generate the second video based on the two.
[0039] However, the side-pasted images in the above solution have the problem of being single in form and lacking in personalization, resulting in poor quality of the generated video.
[0040] To address the above technical issues, the present invention extracts text from the first video to generate a sidebar image based on the text. The generated sidebar image is an image associated with the content of the first video. This helps to increase the richness and personalization of the sidebar image, effectively improving the quality of video generation.
[0041] The following detailed description of the technical solutions of the embodiments of the present disclosure and how the technical solutions of the present disclosure solve the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.
[0042] Figure 3 This is a flowchart of a method for generating a video with side-mounted images provided by an embodiment of the present disclosure. Figure 3 As shown, the method for generating a video with a side-mounted picture includes:
[0043] S101: Obtain a first video.
[0044] The first video may be any video obtained in any manner.
[0045] S102: Extracting text for describing content in the first video from the first video.
[0046] The text here is any text used to represent the content of the first video, and can be obtained by performing image recognition from any video frame of the first video. The text used to describe the content of the first video can be semantically continuous natural language text, or it can be keyword text including multiple keywords. It is understood that when the text used to describe the content of the first video includes keywords, it can reduce the information redundancy input into the image generation model when generating the side-posted image, thereby reducing the generation complexity of the side-posted image.
[0047] Specifically, all or part of the video frames of the first video can be input into an image recognition model to obtain descriptive text corresponding to each video frame, and the text describing the content of the first video can be generated based on the descriptive text corresponding to each video frame. Specifically, the descriptive text of each video frame can be aggregated, and duplicate content in the descriptive text can be deleted to obtain the text describing the content of the first video.
[0048] It is understandable that the above-mentioned image recognition model can be any pre-trained neural network model. For example, a commonly used image recognition model can be CLIP (Contrastive Language Image Pre Training, a pre-trained neural network model for matching images and text), which is a large open source model.
[0049] S103: Generate a side-posting image based on the text, where the text is used to determine the attributes and / or content of the side-posting image.
[0050] Among them, the side-posting picture is any picture used to be added to the periphery of the first video and associated with the text used to describe the content in the first video. The attributes of the side-posting picture include at least one of the following: the target position, color, size, style, and shape of the side-posting picture for adding the video. Different texts used to describe the content in the first video can correspond to side-posting pictures with different attributes and / or contents. It can be seen that the above-mentioned text used to describe the content in the first video is used to determine the target position of the first video in the side-posting picture, the color of the side-posting picture, the size of the side-posting picture, the style of the side-posting picture, the shape of the side-posting picture, and the content of the side-posting picture.
[0051] The target location for adding a video in the above-mentioned side-posted image can be a portion of the area within the side-posted image. For example, it can be the central sub-area within the side-posted image, so that the video frame of the first video is added to the central sub-area of the side-posted image. In the embodiment of the present disclosure, the target location is not fixed, but is associated with the content of the first video. Different content in different first videos corresponds to different target locations. In this way, the location of first videos with different content in the generated second video can be different, which can improve the personalization of the second video.
[0052] The above-mentioned side-posting image can be generated by the following process: inputting text used to describe the content of the first video into an image generation model to obtain the side-posting image. The image generation model can be any neural network model with image generation function.
[0053] Of course, the above-mentioned image generation model can also be generated through pre-training. The second training set required for the pre-training process may include a large number of sample texts and sample images corresponding to each sample text. These sample texts are associated with the attributes and / or content of the sample images. Thus, this association relationship between the two can be identified through the pre-training process, so that the image generation model can generate a side-stick image associated with the text used to describe the content in the first video when applied.
[0054] In some embodiments, the text describing the content of the first video is extracted from some or all video frames in the first video. Taking one of the video frames as an example, a corresponding sidebar image can be generated based on the text describing the content of the first video. In this way, each of the some or all video frames receives a sidebar image, allowing different video frames in the first video to correspond to different sidebar images, further improving the richness of the sidebar images and the resulting video quality.
[0055] S104: Synthesize the first video and the side-mounted image to obtain a second video.
[0056] It can be understood that the process of generating the second video includes adding the first video to the target position in the side-mounted image to obtain the second video. The specific implementation method is: adding each video frame of the first video to the target position in the corresponding side-mounted image to obtain each video frame of the second video.
[0057] The target position may be any position within the border image, and usually the target position may be the center position of the border image.
[0058] To sum up, the embodiment of the present disclosure can generate a side-posting picture associated with the content of the first video to form a second video. In this way, the first video and the side-posting picture in the second video can be associated, ensuring the matching between the two, thereby improving the generation effect of the second video.
[0059] In some embodiments, when extracting text for describing the content in the first video, the above S102 may specifically include: first, extracting content description text from the first video; then, splitting the content description text according to the subject and the scene to obtain subject description text and scene description text; finally, generating text for describing the content in the first video based on the subject description text and the scene description text.
[0060] The content description text is used to describe the content of the first video and can be recognized by the image recognition model. The content description text output by the image recognition model is a continuous natural language text that can be directly used as the text used to describe the content of the first video. However, in order to more accurately describe the content of the first video and improve the accuracy of the text used to describe the content of the first video, the content description text can be processed and used as the text used to describe the content of the first video.
[0061] As can be seen from the above description, the text used to describe the content in the first video includes a subject description text and a scene description text. The subject description text is used to describe the subject in the first video, and the scene description text is used to describe the scene in the first video. Therefore, when generating a side-posted picture based on the text used to describe the content in the first video, it can be generated based on the subject and the scene respectively, and the degree of influence of the two on the side-posted picture can be flexibly adjusted. For example, the weight of the subject description text can be set higher, and the weight of the scene description text can be set lower, so that the generated side-posted picture is more in line with the subject. Of course, the weights here can be flexibly set according to actual needs.
[0062] In order to further reduce the redundancy of the text used to describe the content in the first video, the subject description text and the scene description text can also be segmented to obtain multiple segmentations, so as to determine the text used to describe the content in the first video based on the multiple segmentations. Thus, the text used to describe the content in the first video includes keywords in two dimensions: keywords for describing the subject and keywords for describing the scene. In this way, not only can the content of the first video be described from two dimensions, but the redundancy of the text used to describe the content in the first video can also be reduced as much as possible, reducing the complexity of the picture generation model in generating pictures based on text. In addition, the interference of redundant information can be avoided, which helps to improve the accuracy of the generated pictures.
[0063] In some embodiments, before determining the text used to describe the content of the first video based on the multiple segmented words, the first importance of each segmented word corresponding to the main description text in the main description text and the second importance of each segmented word corresponding to the scene description text in the scene description text can be determined. Then, the target segmented word is deleted from the multiple segmented words.
[0064] The target segmented words include one or more segmented words with the lowest first importance and one or more segmented words with the lowest second importance. The first importance level indicates the importance of each segmented word in the subject description text to the subject, while the second importance level indicates the importance of each segmented word in the scene description text to the scene. This allows us to delete some unimportant segmented words, further reducing the redundancy of the target segmented words.
[0065] The following examples illustrate the relationship between the content description text, the main description text, the scene description text, and the text used to describe the content in the first video. For example, the content description text can be "The parrot is walking on the balcony, with the blue sky and woods behind it", the split main description text can be "Parrot, walking", and the scene description text can be "Balcony, blue sky, woods". After word segmentation of the main description text, two participles "parrot" and "walk" are obtained, and after word segmentation of the scene description text, three participles "balcony", "blue sky", and "woods" are obtained. After deleting some participles, the text used to describe the content in the first video can include "parrot" and "balcony".
[0066] In some implementations, to further improve processing complexity and the accuracy of the content description text, the content description text can be extracted from some video frames of the first video. Specifically, first, second quality information corresponding to each video frame in the first video is determined. The second quality information includes at least one of the following: content richness, video frame clarity, exposure, and color contrast. Then, a target video frame is selected from each video frame based on the second quality information. Finally, the content description text is extracted from the target video frame.
[0067] The target video frame is one or more video frames with the highest quality in the first video. The process of extracting the content description text from the target video frame can refer to the aforementioned process of extracting the content description text from any video frame.
[0068] The above-mentioned content richness can be represented by the number of objects and the number of object types included in the video frame. The larger the number of objects and / or the number of object types, the higher the content richness, and thus the higher the quality of the video frame. For example, when there is no clear object in a video frame, the content richness of the video frame is low. For another example, for a video frame B1 having four types of objects such as human faces, animals, grass, blue sky, and woods, and a video frame B2 having one type of object, human faces, the number of object types of video frame B1 is 4, and the number of object types of video frame B2 is 1, so that the content richness of video frame B1 is higher than that of video frame B2. It can be understood that the content description text extracted from a video frame with higher content richness can include richer content, thereby generating richer side-posting pictures, which helps to further improve the relevance of the side-posting pictures with the first video, as well as the video effect.
[0069] Regarding the video frame clarity, a larger value indicates higher quality. For video frames with higher video clarity, the recognized text describing the content of the first video is more accurate, which can improve the relevance of the side-link image to the first video and the video quality.
[0070] A reasonable exposure range indicates higher video frame quality. For video frames with reasonable exposure, the recognized text describing the content of the first video is more accurate, improving the relevance of the side-link image to the first video and enhancing the video quality.
[0071] Regarding the aforementioned color contrast, a higher value indicates higher quality of the video frame. For video frames with higher color contrast, the recognized text describing the content of the first video is more accurate, which can improve the relevance of the side-image to the first video and the video quality.
[0072] As can be seen, the disclosed embodiment can select target video frames to extract text describing the content of the first video based on quality information from multiple dimensions. This can maximize the accuracy and richness of the text describing the content of the first video, thereby improving the video quality.
[0073] In practical applications, keyframes can be used as target video frames to avoid quality analysis of the first video and reduce the complexity and time required to extract the text describing the content in the first video. Furthermore, keyframes are generally of higher quality, ensuring the accuracy of the text describing the content in the first video.
[0074] It is understood that the series of processes used in the process of generating the text describing the content of the first video based on the content description text can be implemented by a natural language processing model. For example, a commonly used natural language processing model can be a large language model (LLM).
[0075] After obtaining the text for describing the content in the first video through the above process, S103 can also be executed to generate a side-posting picture. From the description of the aforementioned S103, it can be known that the text for describing the content in the first video can be input into the picture generation model to use the picture output by the picture generation model as a side-posting picture. In some embodiments, when generating a side-posting picture, the above S203 may include: first, inputting the text for describing the content in the first video into the picture generation model to obtain at least two candidate pictures; then, determining the degree of relevance between each candidate picture and the vertical category of the video content; finally, selecting a side-posting picture from the candidate pictures based on the degree of relevance. In this way, the picture generation model can be controlled to generate multiple candidate pictures, from which side-posting pictures related to the vertical category of the first video can be selected, which can improve the matching of the side-posting picture with the first video and improve the quality of the second video.
[0076] The vertical category is the vertical category of the first video, which is used to indicate the content type of the first video. For example, the vertical category of the first video is a food video, an interview video, a live broadcast video, etc.
[0077] In an embodiment of the present disclosure, the vertical categories of the candidate image and the first video can be input into a correlation model to obtain the degree of correlation between the two in terms of vertical categories. The correlation model can be any neural network model, for example, a chatGPT model.
[0078] In addition to filtering the images for side-posting based on the vertical categories above, you can also select images for side-posting based on quality. Specifically, first, obtain the first quality information of each candidate image; then, select the image for side-posting from the candidate images based on the first quality information and the degree of relevance.
[0079] The side-posting picture may be a picture with higher quality and higher relevance among multiple candidate pictures. Before selecting the side-posting picture, the first quality information is numerically processed to obtain a quality parameter, and the side-posting picture is selected based on the quality parameter and the relevance.
[0080] In the first example of selecting a side-posting picture, a candidate picture whose quality parameter is greater than or equal to a preset quality threshold and whose correlation degree is greater than or equal to a preset correlation threshold is selected as a side-posting picture, or the candidate pictures are comprehensively sorted by quality parameter and correlation degree, so that the candidate picture with the highest ranking is selected as the side-posting picture.
[0081] In the second example of selecting a side-posting picture, the comprehensive quality parameter of each candidate picture is determined by the quality parameter and the degree of correlation, so that the candidate picture with the comprehensive quality parameter greater than or equal to the preset comprehensive quality threshold is used as the side-posting picture, or the candidate picture with the largest comprehensive quality parameter is used as the side-posting picture.
[0082] It can be seen that the embodiment of the present disclosure can improve the quality of the side-posted pictures as much as possible while ensuring the association with the first video vertical category, which helps to improve the quality of the second video.
[0083] It is understood that the first quality information can describe the quality of the candidate image from at least one dimension. Thus, the first quality information includes at least one of the following: clarity, exposure, noise, color, composition, subject emphasis, whether a subject is missing, and whether a scene is missing.
[0084] Among them, when the clarity is higher, the exposure is within a reasonable range, the noise is smaller, the color is brighter, the composition is better, the subject is more emphasized, the subject is not missing, and the scene is not missing, the quality of the candidate image is higher.
[0085] When the clarity is higher, the area in the candidate image that is blurred, out of focus, jagged, mosaic, etc. is smaller.
[0086] Figure 4This is a schematic diagram of a video generation process provided by an embodiment of the present application. After the user selects the first video, the following steps can be performed: Figure 4 The process shown.
[0087] Reference Figure 4 As shown, after the user selects the first video, the first video is firstly extracted to obtain a target video frame. The target video frame here can be a video frame with better quality in the first video, for example, a key frame.
[0088] After extracting the target video frame, the target video frame can be input into the image recognition model to extract content description text from the target video frame.
[0089] After obtaining the content description text, the content description text can be processed through the LLM model to obtain text used to describe the content of the first video. This processing can include multiple steps: splitting the content description text into a main description text and a scene description text, segmenting the main description text and the scene description text, and determining the first importance of the segmented words in the main description text and the second importance of the segmented words in the scene description text. These multiple processes can each correspond to an LLM model, that is, there can be multiple LLM models.
[0090] After obtaining the text for describing the content in the first video, the text for describing the content in the first video is input into an image generation model to obtain a plurality of candidate images.
[0091] After obtaining multiple candidate images, you can filter out the edge-pasting images.
[0092] After obtaining the side-posting picture, the first video and the side-posting picture can be combined to obtain a second video. The second video is a video combined with the first video and the side-posting picture.
[0093] In summary, this application can build an AIGC (Artificial Intelligence Generated Content) platform to intelligently generate personalized, diverse side-by-side images for a first video, improving the quality of video generation. Furthermore, it can eliminate the need for users to select side-by-side images, reducing operational complexity, and reduce the number of steps required for designers to design side-by-side images, reducing labor costs and enabling mass production of side-by-side images.
[0094] Figure 5 This is a flowchart of another method for generating a video with side-mounted pictures provided by an embodiment of the present disclosure. Figure 5 As shown, the above-mentioned method for generating a video with a side-mounted picture may include:
[0095] S301: Obtain a first video.
[0096] S302: Determine second quality information corresponding to each video frame in the first video.
[0097] The second quality information includes at least one of: content richness, video frame clarity, exposure level, and color contrast.
[0098] S303: Select a target video frame from each video frame of the first video according to the second quality information.
[0099] S304: Extracting content description text from the target video frame.
[0100] S305: Split the content description text according to subject and scene to obtain subject description text and scene description text.
[0101] S306: Perform word segmentation processing on the subject description text and the scene description text to obtain multiple word segments.
[0102] S307: Determine the first importance of each segmentation corresponding to the main description text in the main description text, and determine the second importance of each segmentation corresponding to the scene description text in the scene description text.
[0103] S308: Delete the target segmented word from the multiple segmented words according to the first importance and the second importance.
[0104] The target segmented words include: one or more segmented words with the least first importance, and one or more segmented words with the least second importance.
[0105] S309: Determine text for describing the content in the first video based on the multiple segmented words.
[0106] S310: Inputting text used to describe the content of the first video into an image generation model to obtain at least two candidate images.
[0107] S311: Determine the degree of relevance between each candidate image and the vertical category of the video content of the first video.
[0108] S312: Obtain first quality information of each candidate picture.
[0109] S313: Selecting a side-posting picture from the candidate pictures according to the first quality information and the correlation degree.
[0110] S314: Synthesize the first video and the side-mounted image to obtain a second video.
[0111] It should be noted that the above Figure 5The steps shown can be adjusted in sequence independently of each other. Figure 3 The corresponding description of the video generation method with side-mounted pictures is shown and will not be repeated here.
[0112] Corresponding to the video generation method with side-mounted pictures in the above embodiment, Figure 6 This is a structural block diagram of a video generation device with side-mounted pictures provided by an embodiment of the present disclosure. For ease of explanation, only the parts related to the embodiment of the present disclosure are shown. Figure 6 The above-mentioned video generation device 400 with side-mounted pictures includes:
[0113] The first video acquisition module 401 is configured to acquire a first video.
[0114] The text extraction module 402 is configured to extract text describing the content of the first video from the first video.
[0115] The side-stick image generation module 403 is configured to generate a side-stick image based on the text, where the text is used to determine the attributes and / or content of the side-stick image.
[0116] The second video generation module 404 is configured to synthesize the first video and the side-mounted image to obtain a second video.
[0117] Optionally, the side image generation module 403 is further configured to:
[0118] The text used to describe the content of the first video is input into the image generation model to obtain at least two candidate images; the degree of correlation between each of the candidate images and the vertical category of the first video is determined; and the side-posting image is selected from the candidate images based on the degree of correlation.
[0119] Optionally, the side image generation module 403 is further configured to:
[0120] Acquire first quality information of each candidate image; and select the side-posting image from the candidate images based on the first quality information and the correlation degree.
[0121] Optionally, the first quality information includes at least one of the following: clarity, exposure, noise, color, composition, subject emphasis, whether a subject is missing, and whether a scene is missing.
[0122] Optionally, the text extraction module 402 is further configured to:
[0123] Extract content description text from the first video; split the content description text according to subject and scene to obtain the subject description text and the scene description text; generate the text for describing the content in the first video based on the subject description text and the scene description text.
[0124] Optionally, the text extraction module 402 is further configured to:
[0125] The subject description text and the scene description text are segmented to obtain a plurality of segmented words; and the text used to describe the content in the first video is determined based on the plurality of segmented words.
[0126] Optionally, it also includes:
[0127] The importance determination module is used to determine the first importance of each segmentation corresponding to the main description text in the main description text before determining the text used to describe the content in the first video based on the multiple segmentations, and to determine the second importance of each segmentation corresponding to the scene description text in the scene description text.
[0128] The segmentation deletion module is used to delete a target segmentation from the multiple segmentations, where the target segmentation includes: one or more segmentations with the least first importance and one or more segmentations with the least second importance.
[0129] Optionally, the text extraction module 402 is further configured to:
[0130] Determine second quality information corresponding to each video frame in the first video, where the second quality information includes at least one item: content richness, video frame clarity, exposure level, and color contrast; select a target video frame from the video frames based on the second quality information; and extract the content description text from the target video frame.
[0131] Optionally, the attributes of the side-posted image include at least one of the following: a target position, color, size, style, and shape for adding a video in the side-posted image.
[0132] The video generation device with side-mounted pictures provided in this embodiment can be used to implement the technical solution of the above-mentioned video generation method with side-mounted pictures embodiment. Its implementation principle and technical effects are similar and will not be repeated here in this embodiment.
[0133] Figure 7 6 is a block diagram of an electronic device provided by an embodiment of the present disclosure. The electronic device 600 includes a memory 602 and at least one processor 601.
[0134] The memory 602 stores computer-executable instructions.
[0135] At least one processor 601 executes the computer-executable instructions stored in the memory 602 , so that the electronic device 600 implements the aforementioned method for generating a video with side-mounted pictures.
[0136] In addition, the electronic device may further include a receiver 603 and a transmitter 604. The receiver 603 is configured to receive information from other devices or equipment and forward it to the processor 601. The transmitter 604 is configured to send the information to the other devices or equipment.
[0137] In a first example of the first aspect, an embodiment of the present disclosure provides a method for generating a video with a side-mounted image, including:
[0138] Get the first video.
[0139] Text describing content in the first video is extracted from the first video.
[0140] A sidebar image is generated according to the text, and the text is used to determine the attributes and / or content of the sidebar image.
[0141] The first video and the side-mounted image are combined to obtain a second video.
[0142] Based on the first example of the first aspect, in the second example of the first aspect, generating a sidebar image according to the text includes:
[0143] The text is input into an image generation model to obtain at least two candidate images; the degree of correlation between each candidate image and the vertical category of the first video is determined; and the side-posting image is selected from the candidate images based on the degree of correlation.
[0144] Based on the second example of the first aspect, in the third example of the first aspect, selecting the side-posting picture from the candidate pictures according to the correlation degree includes:
[0145] Acquire first quality information of each candidate image; and select the side-posting image from the candidate images based on the first quality information and the correlation degree.
[0146] Based on the third example of the first aspect, in the fourth example of the first aspect, the first quality information includes at least one of the following: clarity, exposure, noise, color, composition, subject emphasis, whether the subject is missing, and whether the scene is missing.
[0147] Based on the first to fourth examples of the first aspect, in a fifth example of the first aspect, extracting text describing content in the first video from the first video includes:
[0148] Extract content description text from the first video; split the content description text according to subject and scene to obtain the subject description text and the scene description text; generate the text for describing the content in the first video based on the subject description text and the scene description text.
[0149] Based on the fifth example of the first aspect, in the sixth example of the first aspect, generating the text according to the subject description text and the scene description text includes:
[0150] The subject description text and the scene description text are segmented to obtain a plurality of segmented words; and the text used to describe the content in the first video is determined based on the plurality of segmented words.
[0151] Based on the sixth example of the first aspect, in the seventh example of the first aspect, before determining the text used to describe the content in the first video according to the multiple word segmentations, the method further includes:
[0152] Determine the first importance of each segmentation word corresponding to the main description text in the main description text, and determine the second importance of each segmentation word corresponding to the scene description text in the scene description text; delete target segmentations from the multiple segmentations, and the target segmentations include: one or more segmentations with the smallest first importance and one or more segmentations with the smallest second importance.
[0153] Based on the sixth example of the first aspect, in an eighth example of the first aspect, extracting content description text from the first video includes:
[0154] Determine second quality information corresponding to each video frame in the first video, where the second quality information includes at least one item: content richness, video frame clarity, exposure level, and color contrast; select a target video frame from the video frames based on the second quality information; and extract the content description text from the target video frame.
[0155] Based on the first to fourth examples of the first aspect, in the ninth example of the first aspect, the attributes of the side-posted image include at least one of the following: a target position, color, size, style, and shape for adding a video in the side-posted image.
[0156] In a first example of the second aspect, a video generating apparatus with a side-mounted image is provided, including:
[0157] The first video acquisition module is used to acquire the first video.
[0158] A text extraction module is used to extract text describing the content in the first video from the first video.
[0159] The side-posting picture generation module is used to generate a side-posting picture according to the text, and the text is used to determine the attributes and / or content of the side-posting picture.
[0160] The second video generation module is used to synthesize the first video and the side-mounted image to obtain a second video.
[0161] Based on the first example of the second aspect, in the second example of the second aspect, the edge sticker image generation module is further configured to:
[0162] The text is input into an image generation model to obtain at least two candidate images; the degree of correlation between each candidate image and the vertical category of the first video is determined; and the side-posting image is selected from the candidate images based on the degree of correlation.
[0163] Based on the second example of the second aspect, in the third example of the second aspect, the edge sticker image generation module is further configured to:
[0164] Acquire first quality information of each candidate image; and select the side-posting image from the candidate images based on the first quality information and the correlation degree.
[0165] Based on the third example of the second aspect, in the fourth example of the second aspect, the first quality information includes at least one of the following: clarity, exposure, noise, color, composition, subject emphasis, whether the subject is missing, and whether the scene is missing.
[0166] Based on the first to fourth examples of the second aspect, in a fifth example of the second aspect, the text extraction module is further configured to:
[0167] Extract content description text from the first video; split the content description text according to subject and scene to obtain the subject description text and the scene description text; generate the text for describing the content in the first video based on the subject description text and the scene description text.
[0168] Based on the fifth example of the second aspect, in a sixth example of the second aspect, the text extraction module is further configured to:
[0169] The subject description text and the scene description text are segmented to obtain a plurality of segmented words; and the text used to describe the content in the first video is determined based on the plurality of segmented words.
[0170] Based on the sixth example of the second aspect, in the seventh example of the second aspect, further comprising:
[0171] The importance determination module is used to determine the first importance of each segmentation corresponding to the main description text in the main description text before determining the text used to describe the content in the first video based on the multiple segmentations, and to determine the second importance of each segmentation corresponding to the scene description text in the scene description text.
[0172] The segmentation deletion module is used to delete a target segmentation from the multiple segmentations, where the target segmentation includes: one or more segmentations with the least first importance and one or more segmentations with the least second importance.
[0173] Based on the sixth example of the second aspect, in an eighth example of the second aspect, the text extraction module is further configured to:
[0174] Determine second quality information corresponding to each video frame in the first video, where the second quality information includes at least one item: content richness, video frame clarity, exposure level, and color contrast; select a target video frame from the video frames based on the second quality information; and extract the content description text from the target video frame.
[0175] Based on the first to fifth examples of the second aspect, in the ninth example of the second aspect, the attributes of the side-posted image include at least one of the following: a target position, color, size, style, and shape for adding a video in the side-posted image.
[0176] In a third aspect, according to one or more embodiments of the present disclosure, an electronic device is provided, comprising: at least one processor and a memory.
[0177] The memory stores computer-executable instructions.
[0178] The at least one processor executes the computer-executable instructions stored in the memory, so that the electronic device implements any method of generating a video with side-mounted pictures in the first aspect.
[0179] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the computing device implements any one of the methods for generating a video with side-mounted images in the first aspect.
[0180] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program is provided, wherein the computer program is used to implement the method for generating a video with side-mounted pictures according to any one of the first aspects.
[0181] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.
[0182] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.
[0183] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A method for generating a video with side-mounted pictures, characterized in that: The method comprises: Get the first video; Extracting text describing content in the first video from the first video; generating a side-posting image based on the text, wherein the text is used to determine the attributes and / or content of the side-posting image; The first video and the side-mounted image are combined to obtain a second video.
2. The method according to claim 1, characterized in that Generating a side image according to the text includes: Inputting the text into an image generation model to obtain at least two candidate images; Determining a degree of relevance between each of the candidate images and the vertical category of the first video; The side-posting picture is selected from the candidate pictures according to the correlation degree.
3. The method according to claim 2, characterized in that Selecting the side-posting picture from the candidate pictures according to the correlation degree includes: Obtaining first quality information of each candidate image; The side-posting picture is selected from the candidate pictures according to the first quality information and the correlation degree.
4. The method according to claim 3, characterized in that The first quality information includes at least one of the following: clarity, exposure, noise, color, composition, subject emphasis, whether a subject is missing, and whether a scene is missing.
5. The method according to any one of claims 1 to 4, characterized in that Extracting text describing content in the first video from the first video includes: extracting content description text from the first video; Splitting the content description text according to subject and scene to obtain subject description text and scene description text; A text for describing the content in the first video is generated based on the main description text and the scene description text.
6. The method according to claim 5, characterized in that Generating text for describing the content of the first video according to the subject description text and the scene description text includes: Performing word segmentation processing on the subject description text and the scene description text to obtain multiple word segments; Determine text for describing content in the first video based on the multiple word segments.
7. The method according to claim 6, characterized in that Before determining text for describing the content in the first video according to the multiple word segmentations, the method further includes: Determining a first importance of each segmentation corresponding to the subject description text in the subject description text, and determining a second importance of each segmentation corresponding to the scene description text in the scene description text; Delete a target participle from the multiple participles, where the target participle includes: one or more participles with the least first importance and one or more participles with the least second importance.
8. The method according to claim 6, characterized in that Extracting content description text from the first video includes: Determining second quality information corresponding to each video frame in the first video, where the second quality information includes at least one of: content richness, video frame clarity, exposure, and color contrast; selecting a target video frame from the video frames according to the second quality information; The content description text is extracted from the target video frame.
9. The method according to any one of claims 1 to 4, characterized in that The attributes of the side-posted image include at least one of the following: a target position, color, size, style, and shape for adding a video in the side-posted image.
10. A video generation device with side-mounted pictures, characterized in that: include: A first video acquisition module, configured to acquire a first video; A text extraction module, configured to extract text describing the content of the first video from the first video; A side-posting picture generation module, configured to generate a side-posting picture based on the text, wherein the text is used to determine the attributes and / or content of the side-posting picture; The second video generation module is used to synthesize the first video and the side-mounted image to obtain a second video.
11. An electronic device, characterized in that: include: at least one processor and memory; The memory stores computer-executable instructions; The at least one processor executes the computer-executable instructions stored in the memory, so that the electronic device implements the method according to any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, and when a processor executes the computer-executable instructions, the computing device implements the method according to any one of claims 1 to 9.
13. A computer program, characterized in that The computer program is used to implement the method according to any one of claims 1 to 9.