Video processing method and device, electronic equipment and computer readable storage medium
By acquiring audience information and tags for the video, suitable special effects images are selected and added to the video, solving the problem of insufficient accuracy in existing video processing technologies and achieving higher video processing accuracy.
Patent Information
- Application Number
- CN202111367476.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-18
- Publication Date
- 2025-11-21
- Estimated Expiration
- 2041-11-18
AI Technical Summary
Existing video processing methods only consider the intrinsic information of the video in specific business scenarios, resulting in low accuracy in selecting special effects images and affecting the accuracy of video processing.
By obtaining audience information for the video to be processed and combining it with video tags, special effects images are selected to enhance the video's performance, added to the video, and then edited.
It improves the accuracy of video processing, especially in business scenarios that require dissemination, by fully considering the impact on the target audience and improving the accuracy of special effects image selection.
Smart Images

Figure CN116137648B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of communication technology, and more specifically to a video processing method, apparatus, and computer-readable storage medium. Background Technology
[0002] In recent years, with the rapid development of internet technology, video processing methods have become increasingly diversified. For example, adding personalized stickers and other special effects images to videos can enhance their visual appeal. Existing video processing methods often select target stickers from a sticker library and process them based on video facial recognition results, speech recognition results, or the attribute information of objects in the image.
[0003] In the process of researching and practicing existing technologies, the inventors of this invention discovered that in some specific business scenarios, such as in scenarios where the video to be processed needs to be disseminated, existing video processing methods only consider the intrinsic information of the video, and the factors for selecting special effects images are relatively singular, which reduces the accuracy of the selected special effects images, thus resulting in insufficient accuracy in video processing. Summary of the Invention
[0004] The present invention provides a video processing method, apparatus, electronic device, and computer-readable storage medium, which can improve the accuracy of video processing.
[0005] A video processing method, comprising:
[0006] Obtain the video to be processed and the audience information corresponding to the video to be processed, wherein the audience information is used to indicate the audience information targeted by the video to be processed;
[0007] The content of the video to be processed is identified so as to select at least one video tag from a preset video tag set;
[0008] Based on the audience information and video tags, a special effects image corresponding to the video to be processed is determined, and the special effects image is used to enhance the video performance of the video to be processed;
[0009] The special effects image is added to the video to be processed to obtain the processed video, and the processed video is sent to the terminal so that the terminal can edit the special effects image in the processed video.
[0010] Optionally, embodiments of the present invention may also provide another video processing method, including:
[0011] Send the video to be processed to the server;
[0012] A preview page is displayed showing the processed video after adding special effects images to the video to be processed. The preview page includes editing controls for the special effects images.
[0013] In response to an editing operation on the editing control, a target video is generated.
[0014] Accordingly, embodiments of the present invention provide a video processing apparatus, including:
[0015] The acquisition unit is used to acquire the video to be processed and the audience information corresponding to the video to be processed, wherein the audience information is used to indicate the information of the audience to which the video to be processed is targeted;
[0016] The identification unit is used to identify the content of the video to be processed, so as to filter out at least one video tag from a preset video tag set;
[0017] The determining unit is used to determine the special effects image corresponding to the video to be processed based on the audience information and video tags, wherein the special effects image is used to enhance the video performance of the video to be processed;
[0018] An adding unit is used to add the special effects image to the video to be processed to obtain a processed video, and send the processed video to a terminal so that the terminal can edit the special effects image in the processed video.
[0019] Optionally, embodiments of the present invention may also provide another video processing apparatus, including:
[0020] The sending unit is used to send the video to be processed to the server;
[0021] The display unit is used to display a preview page of the processed video after adding special effects images to the video to be processed, and the preview page includes editing controls for the special effects images;
[0022] The generation unit is used to generate a target video in response to an editing operation on the editing control.
[0023] Optionally, in some embodiments, the identification unit may be specifically used to extract video streams and audio streams from the video to be processed; perform content recognition on the video streams and audio streams respectively to obtain the screen content and text content of the video to be processed; and based on the screen content and text content, filter out at least one video tag from a preset video tag set.
[0024] Optionally, in some embodiments, the identification unit may be specifically used to extract video frames from the video stream to obtain a set of video frames; identify the picture information of each video frame in the set of video frames to obtain the picture content of the video to be processed; and identify text information in the picture content and audio stream to obtain the text content of the video to be processed.
[0025] Optionally, in some embodiments, the recognition unit may be specifically used to recognize text information in the screen content to obtain video stream text; convert the audio stream into text information to obtain audio stream text, and use the video stream text and audio stream text as the text content of the video to be processed.
[0026] Optionally, in some embodiments, the identification unit may be specifically used to: filter at least one key video frame from the video frame set based on the image content; extract keywords from the audio stream text and video stream text, and fuse the extracted keywords to obtain at least one target keyword of the video to be processed; and filter at least one video tag from a preset video tag set based on the key video frame and the target keyword.
[0027] Optionally, in some embodiments, the identification unit may be specifically used to filter out style tags and item tags from a preset video tag set to obtain a style tag set and an item tag set; based on the key video frame, filter out at least one style tag corresponding to the video to be processed from the style tag set; based on the target keyword, filter out at least one item tag corresponding to the video to be processed from the item tag set, wherein the item tag is used to indicate information about the items contained in the video to be processed.
[0028] Optionally, in some embodiments, the identification unit may be specifically used to extract style features from the key video frames to obtain the style features of the video to be processed; and to filter at least one style tag corresponding to the style features from the style tag set to obtain the style tag corresponding to the video to be processed.
[0029] Optionally, in some embodiments, the identification unit may be specifically used to extract features from the target keywords to obtain the item features of the video to be processed; based on the item features, determine at least one item information in the video to be processed; and filter out the item tags corresponding to the item information from the item tag set to obtain the item tags of the video to be processed.
[0030] Optionally, in some embodiments, the determining unit may be specifically used to obtain a preset set of special effects images, the preset set of special effects images including at least one preset special effects image and attribute information of the preset special effects image; match the attribute information of the special effects image with the audience information, style tags and item tags of the video to be processed respectively; and based on the matching results, filter out the special effects image corresponding to the image to be processed from the preset set of special effects images.
[0031] Optionally, in some embodiments, the determining unit may be specifically used to identify the target style tag set, target item tag set, and audience information set corresponding to the preset special effects image in the attribute information; match the audience information with the audience information set, match the style tags with the target style tag set, and match the item tags with the target item tag set; the step of filtering the special effects image corresponding to the image to be processed from the preset special effects image set based on the matching results includes: filtering the preset special effects image from the preset special effects image set that matches all the audience information, style tags, and item tags of the video to be processed, so as to obtain the special effects image corresponding to the video to be processed.
[0032] Optionally, in some embodiments, the determining unit may be specifically used to filter out preset special effects images from the preset special effects image set that completely match the audience information, style tags, and item tags of the video to be processed, to obtain candidate special effects images; when the number of candidate special effects images is one, the candidate special effects image is used as the special effects image corresponding to the video to be processed; when the number of candidate special effects images is multiple, the special effects image corresponding to the video to be processed is filtered out from the candidate special effects images according to the attribute information of the candidate special effects images.
[0033] Optionally, in some embodiments, the adding unit may be specifically used to obtain the position information of the special effects image; based on the position information, identify the addition position of the special effects image in each video frame of the video to be processed; and add the special effects image to each video frame according to the addition position to obtain the processed video.
[0034] Optionally, in some embodiments, the display unit may specifically be used to include, on the preview page, special effects evaluation data for the processed video, the special effects data being used to evaluate the effect data after adding special effects images to the video to be processed.
[0035] Optionally, in some embodiments, the generation unit may specifically be used to adjust the special effects image in the processed video in response to an editing operation on the editing control; update the preview page of the processed video based on the adjusted special effects image and display the updated preview page, the updated preview page including a subtitle generation control and an application control; add subtitle information to the video in the updated preview page in response to the generation operation of the subtitle generation control to obtain a candidate video; and use the candidate video as the target video in response to the application operation of the application control.
[0036] Furthermore, embodiments of the present invention also provide an electronic device, including a processor and a memory, wherein the memory stores an application program, and the processor is used to run the application program in the memory to implement the video processing method provided in embodiments of the present invention.
[0037] Furthermore, embodiments of the present invention also provide a computer-readable storage medium storing a plurality of instructions adapted for loading by a processor to execute steps in any of the video processing methods provided in embodiments of the present invention.
[0038] In this embodiment of the invention, after obtaining the video to be processed and the corresponding audience information, the content of the video to be processed is identified to filter at least one video tag from a preset video tag set. Then, based on the audience information and the video tag, the special effects image corresponding to the video to be processed is determined. The special effects image is then added to the video to be processed to obtain the processed video, which is then sent to the terminal so that the terminal can edit the special effects image in the processed video. Since this scheme filters the special effects image by using two different factors, audience information and video tags, in the process of determining the special effects image of the video to be processed, it fully considers the influence of the audience of the video to be processed in the business scenario that needs to be disseminated, thereby increasing the accuracy of the selected special effects image. Therefore, it can improve the accuracy of the video processing process. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0040] Figure 1 This is a schematic diagram of a scene using the video processing method provided in an embodiment of the present invention;
[0041] Figure 2This is a flowchart illustrating the video processing method provided in an embodiment of the present invention;
[0042] Figure 3 This is a schematic diagram illustrating the tagging of the original video provided in an embodiment of the present invention;
[0043] Figure 4 This is a schematic diagram of the process for training a label recognition model according to an embodiment of the present invention;
[0044] Figure 5 This is a schematic diagram of the process for obtaining a preset set of special effects images provided in an embodiment of the present invention;
[0045] Figure 6 This is a schematic diagram of the preview page of the processed video provided in an embodiment of the present invention;
[0046] Figure 7 This is another flowchart illustrating the video processing method provided in this embodiment of the invention;
[0047] Figure 8 This is a schematic diagram of the overall process of the video processing method provided in the embodiments of the present invention;
[0048] Figure 9 This is a schematic diagram of the advertising video processing flow provided in an embodiment of the present invention;
[0049] Figure 10 This is a schematic diagram illustrating the effect of adding stickers to an advertising video according to an embodiment of the present invention;
[0050] Figure 11 This is a schematic diagram of the structure of the first video processing device provided in an embodiment of the present invention;
[0051] Figure 12 This is a schematic diagram of the structure of the second video processing device provided in an embodiment of the present invention;
[0052] Figure 13 This is a schematic diagram of the structure of the electronic device provided in an embodiment of the present invention. Detailed Implementation
[0053] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0054] This invention provides a video processing method, apparatus, electronic device, and computer-readable storage medium. The video processing apparatus can be integrated into an electronic device, which may be a server or a terminal, etc. Specifically, this invention provides a video processing apparatus suitable for a first electronic device (which may be referred to as a first video processing apparatus for distinction), and a video processing apparatus suitable for a second electronic device (which may be referred to as a second video processing apparatus for distinction).
[0055] The first electronic device can be a network-side device such as a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. The second electronic device can be a terminal, such as a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, etc., but is not limited to these. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited herein.
[0056] For example, see Figure 1 Taking the integration of a video processing device into a first electronic device as an example, after the electronic device obtains the video to be processed and the audience information corresponding to the video to be processed, it identifies the content of the video to be processed, filters out at least one video tag from a preset video tag set, and then determines the special effects image corresponding to the video to be processed based on the audience information and the video tag. Then, the special effects image is added to the video to be processed to obtain the processed video, and the processed video is sent to the terminal so that the terminal can edit the special effects image in the processed video, thereby improving the accuracy of video processing.
[0057] There are various ways to process video. For example, special effects images can be added to the video to enhance its performance.
[0058] Taking the integration of a video processing device into a second device as an example, the electronic device sends the video to be processed and the corresponding audience information to the server, receives the processed video generated based on the audience information returned by the server, the processed video includes special effects images, and then displays a preview page of the processed video, which includes editing controls for the special effects images. In response to editing operations on the editing controls, the target video is generated.
[0059] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0060] This embodiment will be described from the perspective of a first video processing device, which can be integrated into an electronic device. This electronic device can be a server, which can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms. A video processing method includes:
[0061] The system acquires the video to be processed and its corresponding audience information, which indicates the target audience of the video. It then identifies the content of the video to be processed, selecting at least one video tag from a preset video tag set. Based on the audience information and video tags, it determines the special effects image corresponding to the video to be processed. This special effects image enhances the video's visual appeal. The special effects image is added to the video to be processed, resulting in a processed video. The processed video is then sent to a terminal for editing of the special effects image.
[0062] like Figure 2 As shown, the specific process of this video processing method is as follows:
[0063] 101. Obtain the video to be processed and the corresponding audience information.
[0064] The audience information is used to indicate the target audience of the video to be processed. The audience can be understood as the target group of the video. For example, if the video to be processed is a video advertisement, the audience would be the target audience of that video advertisement. Audience information can include the scope of the audience, the criteria for the audience, the location of the audience, etc. For example, if the audience refers to a group of people, the audience information could include the user range, the user's age, education level, location, identity, and occupation.
[0065] There are several ways to obtain the video to be processed and the corresponding audience information, as follows:
[0066] For example, the system can receive the video to be processed and the corresponding audience information sent by the terminal; or, it can receive the video to be processed and the page information of the audience selection page for the video to be processed sent by the terminal. Based on the page information of the audience selection page, it can filter at least one audience tag corresponding to the video to be processed from a preset set of audience tags, and then merge the audience tags to obtain the audience information. Alternatively, it can filter the original video from the network or video database, select the target type video as the video to be processed from the original video, and then identify the audience of the video to be processed to obtain the audience information. Or, when the memory of the video to be processed is large or the number is large, it can also receive a video processing request sent by the terminal, which carries the storage address of the video to be processed and the corresponding audience information. Based on the storage address, it can retrieve the video to be processed and the corresponding audience information from the terminal's memory, cache, or third-party database.
[0067] There are several ways to filter out videos of the target type from the original videos as videos to be processed. For example, the video type of the original videos can be identified, and videos that are advertisements or other types that need to be disseminated or displayed can be selected as videos to be processed. After selecting the videos to be processed, the target audience for the videos can be identified. There are several ways to identify the target audience. For example, multi-dimensional feature extraction can be performed on the videos to be processed, and the extracted audience features can be fused to obtain global audience features. Based on the global audience features, the target audience for the videos to be processed can be selected from a preset set of target audiences, thus obtaining the audience information for the videos to be processed.
[0068] 102. Identify the content of the video to be processed, and filter out at least one video tag from the preset tag set.
[0069] Video tags are used to describe information about the video to be processed from specific dimensions. These tags can include multiple types. Taking a video advertisement as an example, the video tags could include style tags and item tags. Item tags indicate information about the items contained in the video. Again, using a video advertisement as an example, item tags primarily indicate various information about the products contained in the video advertisement. Item tags can also include multiple types, such as product content tags, selling point tags, and holiday tags. Style tags primarily indicate the style of the video to be processed. These styles can be of various types, such as lighthearted, energetic, sad, funny, serious, and humorous.
[0070] There are several methods for recognizing the content of the video to be processed, as follows:
[0071] For example, extract the video stream and audio stream from the video to be processed, perform content recognition on the video stream and audio stream respectively to obtain the visual content and text content of the video to be processed, and based on the visual content and text content, select at least one video tag from a preset video tag set, as follows:
[0072] S1. Extract the video stream and audio stream from the video to be processed.
[0073] The video stream is a data stream composed of each video frame in the video to be processed in the playback order, and the audio stream is a data stream composed of each audio frame in the video to be processed in the playback order.
[0074] There are several ways to extract video and audio streams, including the following:
[0075] For example, video and audio information in the video to be processed can be separated to obtain target video and target audio information. Based on playback parameters, the target video information can be converted into a video stream and the target audio information can be converted into an audio stream. Alternatively, video frames and audio frames can be directly extracted from the video to be processed, and the video frames can be merged in the playback order to obtain a video stream, and the audio frames can be merged to obtain an audio stream.
[0076] S2. Perform content recognition on the video stream and audio stream respectively to obtain the image content and text content of the video to be processed.
[0077] The image content can be the image information contained within the video frames in the video stream, and the text content can be the text information contained within the audio frames and video frames.
[0078] There are several ways to perform content recognition on video and audio streams, including the following:
[0079] For example, video frames are extracted from the video stream to obtain a set of video frames. The image information of each video frame in the set of video frames is identified to obtain the image content of the video to be processed. Text information is identified from the image content and audio stream to obtain the text content of the video to be processed.
[0080] There are several ways to identify text information in video content and audio stream. For example, text information can be identified in video content to obtain video stream text, audio stream can be converted into text information to obtain audio stream text, and video stream text and audio stream text can be used as the text content of the video to be processed.
[0081] There are several ways to identify text information from video content. For example, a video text recognition model can be used to identify text information appearing in the content of each video frame, which can include the text appearing in the video frame, thus obtaining the video stream text. Similarly, there are several ways to convert audio streams into text information. For instance, a speech-to-text model can be used to convert audio frames in the audio stream into subtitle information for the video to be processed, and then extract text information from the subtitle information to obtain the audio stream text.
[0082] S3. Based on the image content and text content, select at least one video tag from the preset video tag set.
[0083] For example, based on the content of the video, at least one key video frame can be selected from the set of video frames, keywords can be extracted from the audio stream text and video stream text, and the extracted keywords can be merged to obtain at least one target keyword for the video to be processed. Based on the key video frame and the target keyword, at least one video tag can be selected from the preset tag set.
[0084] The key video frame can be a video frame containing preset key features. These key features can be of various types, such as frames containing multiple or specific objects, frames containing a specific scene, or frames containing specific actions or elements. There are several ways to select at least one key video frame from the set of video frames based on the content. For example, features can be extracted from the content of each video frame, and then the feature similarity between the extracted content features and the preset key features can be calculated. Video frames with a feature similarity exceeding a preset similarity threshold can then be selected as key video frames.
[0085] The keywords can be pre-defined words that indicate information about items contained in the video to be processed. There are various ways to merge the extracted keywords, such as deduplicating and filtering the audio stream keywords extracted from the audio stream text and the video stream keywords extracted from the video stream text, so as to obtain at least one target keyword of the video to be processed.
[0086] Optionally, for the extraction of target keywords, the video stream text and audio stream text can be merged, and then keywords can be extracted from the merged text to obtain at least one target keyword of the video to be processed.
[0087] After extracting the key video frames and target keywords, at least one video tag can be selected from the preset video tag set. There are multiple ways to select tags. For example, style tags and item tags can be selected from the preset video tag set to obtain style tag sets and item tag sets. Based on the key video frames, at least one style tag corresponding to the video to be processed can be selected from the style tag set. Based on the target keywords, at least one item tag corresponding to the video to be processed can be selected from the item tag set. The item tags and style tags can then be used as the video tags for the video to be processed.
[0088] There are several ways to select at least one style feature corresponding to the video to be processed from the style feature set based on key video frames. For example, style features can be extracted from key video frames to obtain the style features of the video to be processed, and at least one style tag corresponding to the style feature can be selected from the style tag set to obtain the style tag corresponding to the video to be processed.
[0089] There are several ways to select at least one item tag corresponding to the video to be processed from the item tag set based on the target keywords. For example, feature extraction can be performed on the target keywords to obtain the item features of the video to be processed. Based on the item features, at least one item information in the video to be processed can be determined. Then, the item tags corresponding to the item information can be selected from the item tag set to obtain the item tags of the video to be processed.
[0090] The process of identifying video tags in the video to be processed can be performed by a trained tag recognition model. Specifically, this involves extracting the video and audio streams from the video, using the trained tag recognition model to perform content recognition on both streams to obtain the visual and textual content of the video. Based on the visual content, at least one key video frame is selected from the video frames of the video stream, and at least one target keyword is extracted from both the visual content and the audio stream. Based on the key video frame, at least one style tag corresponding to the video to be processed is selected from a pre-defined style tag set. Based on the target keyword, at least one item tag corresponding to the video to be processed is selected from a pre-defined item tag set.
[0091] The trained label recognition model can be configured according to the needs of actual applications. Furthermore, it should be noted that the trained label recognition model can be pre-configured by maintenance personnel or trained automatically by the video processing device. That is, before the step "using the trained label recognition model to perform content recognition on the video stream and audio stream respectively to obtain the image content and text content of the video to be processed," the video processing method may further include:
[0092] Obtain video samples, including target videos with labeled video tags. Use a pre-defined label recognition model to predict the video tags of the video samples to obtain predicted video tags. Based on the predicted video tags and labeled video tags, converge the pre-defined label recognition model to obtain the trained label recognition model. Specifically, it can be done as follows:
[0093] (1) Obtain video samples.
[0094] The video samples include target videos that have been labeled with video tags. The labeled video tags can include style tags and item tags. The item tags can also include multiple types or dimensions of item sub-tags.
[0095] There are several ways to obtain video samples, including the following:
[0096] For example, at least one original video can be obtained and sent to a tagging server, which then tags the original video with video tags. The original video with the tagged video tags returned by the tagging server is then received, thus obtaining a video sample.
[0097] The tagging server primarily uses manual selection of corresponding video tags from different tag sets to annotate the original video. For example, taking a video advertisement as an example, it can be tagged with style tags and product tags. Product tags can include video tags such as product content tags, selling point tags, and holiday tags. The original video with the tagged video can then be used as a video sample and input into a preset tag recognition model for recognition processing. Specifically, it can be done as follows: Figure 3 As shown.
[0098] (2) The video labels of the video samples are predicted by using a preset label recognition model to obtain the predicted video labels.
[0099] For example, video stream samples and audio stream samples are extracted from video samples. A pre-defined label recognition model is used to perform content recognition on the video stream samples and audio stream samples respectively, obtaining image content samples and text content samples for the video samples. Based on the image content samples, at least one key video frame sample is selected from the video frames of the video stream samples. At least one image keyword sample is extracted from the image content samples. The audio stream is converted into text information, and at least one audio keyword sample is extracted from the converted text information. The image keyword samples and audio keyword samples are merged to obtain the target keyword samples for the video samples. Using the style label recognition sub-model in the pre-defined label recognition model, at least one style label is predicted from a pre-defined style label set based on the key video frame samples, obtaining the predicted style label for the video samples. Using the target keyword samples in the pre-defined label recognition model, at least one item label is predicted from a pre-defined item label set, obtaining the predicted item label for the video samples. The predicted item label and predicted style label are used as the predicted video labels for the video samples.
[0100] (3) The preset label recognition model is converged based on the predicted video labels and the labeled video labels to obtain the trained label recognition model.
[0101] For example, the predicted style labels are compared with the labeled style labels to obtain style loss information, and the predicted item labels are compared with the labeled item labels to obtain item loss information. The style loss information and item loss information are then fused to obtain the label loss information for the video samples. Based on this label loss information, the network parameters of the preset label recognition model are updated to converge the preset label recognition model, thereby obtaining the trained label recognition model.
[0102] The training process of the pre-defined label recognition model mainly involves predicting style labels and item labels. Specifically, the video is divided into video stream samples and audio stream samples. Key video frame samples are extracted from the video stream samples. Keywords are identified in each video frame. The audio stream samples are converted into text information, and at least one keyword is identified from the text information. The identified keywords are then merged to obtain the target keyword samples of the video samples. The style label recognition sub-model in the pre-defined label recognition model identifies the key video frame samples, thus outputting the predicted style labels. Similarly, the item label recognition sub-model in the pre-defined label recognition model identifies the target keyword samples, thus outputting the predicted item labels. The predicted style labels and predicted item labels are then compared with the labeled style labels and labeled item labels, respectively, to converge the pre-defined label recognition model and obtain the trained label recognition model. This can be further described as follows: Figure 4 As shown.
[0103] 103. Based on the audience information and video tags, determine the special effects image corresponding to the video to be processed.
[0104] Special effects images are used to enhance the visual appeal of the video being processed. Taking an advertising video as an example, the special effects images for an advertising video primarily enhance the video's visual appeal by showcasing at least one of its key benefits. These special effects images can take various forms, such as animated or static images, or video stickers. Video stickers, on the other hand, can be static or animated images that cover a specific area of the video and are unrelated to the video content.
[0105] There are several ways to determine the special effects image corresponding to the video to be processed, as follows:
[0106] For example, a set of preset special effects images is obtained, which includes at least one preset special effects image and attribute information of the preset special effects image. The attribute information of the special effects image is matched with the audience information, style tags, and item tags of the video to be processed, respectively. Based on the matching results, the special effects image corresponding to the image to be processed is selected from the set of preset special effects images.
[0107] There are several ways to obtain a set of preset special effects images. For example, one can receive preset special effects images and their tag selection information from the terminal. Based on this tag selection information, audience tags, style tags, and item tags are selected from a tag library. These tags are then merged to generate attribute information for the preset special effects images. This attribute information indicates the applicable set of style tags, item tags, and audience information for the preset image, thus obtaining the set of preset special effects images. Alternatively, the terminal can manually select audience tags, style tags, and item tags for each preset special effects image from a tag library. These tags are then merged to generate attribute information for the preset special effects images, thus obtaining the set of preset special effects images. Specifically, this can be done as follows: Figure 5 As shown.
[0108] There are several ways to match the attribute information of the special effects image with the audience information, style tags, and item tags of the video to be processed. For example, the target style tag set, target item tag set, and audience information set corresponding to the preset special effects image can be identified from the attribute information. The audience information is then matched with the audience information set, the style tags are matched with the target style tag set, and the item tags are matched with the target item tag set. Finally, preset special effects images that completely match the audience information, style tags, and item tags of the video to be processed are selected from the preset special effects image set to obtain the special effects image corresponding to the video to be processed.
[0109] There are several ways to select preset effect images from the preset effect image set that completely match the audience information, style tags, and item tags of the video to be processed. For example, you can select preset effect images from the preset effect image set that completely match the audience information, style tags, and item tags of the video to be processed to obtain candidate effect images. When there is only one candidate effect image, it is used as the effect image corresponding to the video to be processed. When there are multiple candidate effect images, the effect image corresponding to the video to be processed is selected from the candidate effect images based on the attribute information of the candidate effect images.
[0110] In this context, a successful match between the target audience information, style tags, and item tags of the video to be processed can be understood as a successful match between the target audience information set, a successful match between the style tags and the target style tag set, and a successful match between the item tags and the target item tag set. Here, a successful match can be understood as the target audience information, style tags, and item tags of the video to be processed being subsets or proper subsets of the target audience information set, the target style tag set, and the target item tag set, respectively. For example, a successful match in the target audience information set means that the target audience information set of the preset effects image contains the target audience information of the video to be processed; a successful match in the style tags set means that the target style tag set of the preset effects image contains the style tags of the video to be processed; and a successful match in the item tags set means that the target item tag set of the preset effects image contains the item tags of the video to be processed.
[0111] When there are multiple candidate effect images, there are several ways to select the effect image corresponding to the video to be processed from the candidate effect images based on their attribute information. For example, feature extraction can be performed on the attribute information of the candidate effect images to obtain the effect features of each candidate effect image. Effect features can also be extracted from the audience information, style tags, and item tags of the video to be processed, and the extracted basic effect features can be fused to obtain the target effect features of the video to be processed. The effect similarity between the target effect features and the effect features of each candidate effect image can be calculated, and the effect image corresponding to the video to be processed can be selected from the candidate effect images based on the effect similarity.
[0112] 104. Add special effects images to the video to be processed to obtain the processed video, and send the processed video to the terminal so that the terminal can edit the special effects images in the processed video.
[0113] There are several ways to add special effects images to the video to be processed, as follows:
[0114] For example, the location information of the special effects image can be obtained. Based on the location information, the addition position of the special effects image can be identified in each video frame of the video to be processed. According to the addition position, the special effects image is added to each video frame respectively to obtain the processed video.
[0115] The positional information of the special effects image can be used to indicate the location of the special effects image within a video frame, or it can be the required size information of the special effects image within the video frame. The addition position can be the location information for adding the special effects image within the video frame. Based on the positional information of the special effects image, there are several ways to identify the addition position of the special effects image in each frame of the video to be processed. For example, when the positional information of the special effects image indicates the location of the special effects image within the video frame, the corresponding coordinates and other positional information can be directly identified in each video frame of the video to be processed, and the identified coordinates and other positional information can be used as the addition position of the special effects image. When the positional information of the special effects image is the required size information within the video frame, a blank area containing that size information can be identified in each video frame of the video to be processed, and the positional information of that blank area can be obtained, thereby obtaining the addition position of the special effects image in the video frame.
[0116] After adding special effects images to the video to be processed, the processed video can be sent to the terminal so that the terminal can edit the special effects images in the processed video. There are several ways for the terminal to edit the special effects images. For example, the terminal can send the video to be processed and the corresponding audience information to the server, receive the processed video generated based on the audience information returned by the server, the processed video includes special effects images, display a preview page of the processed video, the preview page includes editing controls for the special effects images, and generate the target video in response to the editing operation on the editing controls.
[0117] Wherein, "responding to" is used to indicate the conditions or states on which the operation performed depends. When the conditions or states on which it depends are met, one or more operations performed can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0118] After receiving the processed video from the server, a preview page can be displayed. This preview page may include the processed video with special effects images, and may also include editing and application controls for the special effects images, specifically as follows: Figure 6As shown, there are various types of editing controls, primarily used to edit the text content of special effects images, modify the position and size of special effects images, replace or delete special effects images on the preview page. Each time an editing operation is performed on the special effects image, the preview page is updated to display the updated preview. Once the user has finished editing the special effects image, the terminal generates the target video in response to the application operation on the application controls.
[0119] As can be seen from the above, after obtaining the video to be processed and the corresponding audience information, this embodiment of the application identifies the content of the video to be processed to filter out at least one video tag from a preset video tag set. Then, based on the audience information and video tags, it determines the special effects image corresponding to the video to be processed. The special effects image is then added to the video to be processed to obtain the processed video, and the processed video is sent to the terminal so that the terminal can edit the special effects image in the processed video. Since this scheme filters out the special effects image by using two different factors, audience information and video tags, in the process of determining the special effects image of the video to be processed, it fully considers the influence of the audience of the video to be processed in the business scenario that needs to be disseminated, thereby increasing the accuracy of filtering out the special effects image. Therefore, it can improve the accuracy of the video processing process.
[0120] This embodiment will be described from the perspective of a second video processing device, which can be integrated into an electronic device, such as a terminal; wherein the terminal may include a tablet computer, a laptop computer, a personal computer (PC), a wearable device, a virtual reality device, or other smart devices capable of video processing.
[0121] A video processing method, comprising:
[0122] The system sends the video to be processed to the server and displays a preview page of the processed video after adding special effects images. The preview page includes editing controls for the special effects images. In response to editing operations on the editing controls, the system generates the target video.
[0123] like Figure 7 The specific process of this video processing method is as follows:
[0124] 201. Send the video to be processed to the server.
[0125] For example, you can directly send the video to be processed to the server, or you can store the video to be processed, add the storage address to the video processing request, and send the video processing request with the added storage address to the server, so that the server can obtain the video to be processed based on the storage address carried in the video processing request.
[0126] Optionally, the audience information of the video to be processed can also be obtained and sent to the server. There are several ways to send the audience information to the server, such as sending the audience information directly to the server, or packaging the audience information and the video to be processed together and sending the packaged data to the server. Alternatively, the audience information of the video to be processed can be stored, the storage address can be added to the video processing request, and the video request with the added storage address can be sent to the server so that the server can obtain the audience information of the video to be processed based on the storage address.
[0127] 202. Displays a preview page of the processed video after adding special effects images to the video to be processed.
[0128] The processed video is the video returned by the server after adding special effects images to the video to be processed.
[0129] The preview page includes editing controls for special effects images. These controls can be of various types; for example, they may include controls for replacing logos in the special effects image, or controls for adjusting decorative elements, and so on. The editing controls are primarily used on the preview page to edit the text content of the special effects image, modify the position and size of the special effects image, replace the special effects image, or delete the special effects image.
[0130] There are several ways to display the preview page, as follows:
[0131] For example, the server receives the processed video after adding special effects images to the video to be processed based on the audience information, adds the processed video to a preset page, plays the processed video on the preset page, and obtains a preview page, which may also include editing controls for the special effects images.
[0132] The server can add special effects images to the video to be processed based on audience information in several ways. For example, the server can extract the video stream and audio stream from the video to be processed, and use a trained label recognition model to perform content recognition on the video stream and audio stream respectively to obtain the picture content and text content of the video to be processed. Based on the picture content, at least one key video frame is selected from the video frames of the video stream, and at least one target keyword is extracted from the picture content and audio stream. Based on the key video frame, at least one style tag corresponding to the video to be processed is selected from a preset style tag set. Based on the target keyword, at least one item tag corresponding to the video to be processed is selected from a preset item tag set. The item tag and style tag can be used as the video tags of the video to be processed. A preset special effects image set is obtained, which includes at least one preset special effects image and the attribute information of the preset special effects image. The attribute information of the special effects image is matched with the audience information, style tag and item tag of the video to be processed respectively. Based on the matching results, the special effects image corresponding to the image to be processed is selected from the preset special effects image set. The position information of the special effects image is obtained. Based on the position information, the addition position of the special effects image is identified in each video frame of the video to be processed. According to the addition position, the special effects image is added to each video frame respectively to obtain the processed video.
[0133] Optionally, the preview page may also include special effects evaluation data for the processed video. This special effects evaluation data is used to evaluate the effect data after adding special effects images to the video to be processed. The special effects evaluation data can be calculated by the server and returned to the terminal, or it can be calculated by the terminal itself and displayed on the preview page. There can be various types of special effects evaluation data, such as the rate of change or improvement of conversion data such as click-through rate (CTR), click conversion rate, conversion rate, or interaction rate relative to the video without special effects images. The special effects evaluation data displayed on the preview page can be predicted by the terminal based on historical data. There are several ways to predict the special effects evaluation data. For example, you can obtain a set of video samples with added special effects images, select multiple target video samples with the same special effects images added to the processed video from the video sample set, select multiple original video samples from the set of original video samples without added special effects images, play or push the target video samples and original video samples using the same video platform, and then count the basic evaluation data of the target video samples and original video samples within a preset time period. Then, compare the basic evaluation data of the target video samples and original video samples to obtain the rate of change of the basic evaluation data of the target video samples relative to the original video samples, and use this rate of change as the special effects evaluation data. Alternatively, you can obtain historical business data from the video platform, count the basic evaluation data between the target video samples with added special effects images and the original videos without added special effects images within a preset time period from the historical business data, and then compare the basic evaluation data of the target video samples and original videos to obtain the rate of change of the basic evaluation data of the target video samples relative to the original video samples, and use this rate of change as the special effects evaluation data.
[0134] The basic evaluation data can be conversion data such as CTR, click-through rate, conversion rate, or interaction rate.
[0135] 203. In response to editing operations on the editing controls, generate the target video.
[0136] For example, in response to an editing operation on the editing control, the special effects image in the processed video can be adjusted. Based on the adjusted special effects image, the preview page of the processed video can be updated and the updated preview image can be displayed. The updated preview page includes a subtitle generation control and an application control. In response to the subtitle generation operation of the subtitle generation control, subtitle information can be added to the video in the updated preview page to obtain candidate videos. In response to the application operation of the application control, the candidate videos can be used as the target videos.
[0137] There are several ways to adjust special effects images in the processed video. For example, you can edit the text of the special effects image, modify the position and size of the special effects image, replace or delete the special effects image in the preview page, and you can also adjust the decorative elements or logos in the special effects image.
[0138] The subtitle information can be understood as the letter text that needs to be added to the processed video. It can be added through a subtitle generation control. There are several ways to add subtitle information. For example, the audio information in the processed video can be directly converted into subtitle information and added directly to the corresponding position of the audio frame of the processed video. Alternatively, a subtitle generation page can be displayed. This subtitle generation page can include a subtitle input control. In response to the input operation of the subtitle input control, the subtitle information input by the user can be received and added to the video frame of the processed video to obtain candidate videos.
[0139] Once the user has finished editing the special effects image and added subtitles, the application control can be triggered. In response to the application operation on the application control, the candidate video is selected as the target video.
[0140] As can be seen from the above, after sending the video to be processed to the server, this embodiment of the application displays a preview page of the processed video. The preview page includes editing controls for special effects images. In response to editing operations on the editing controls, the target video is generated. Since this solution filters out special effects images through audience information and can edit the special effects images to generate the target video, the accuracy of video processing can be improved.
[0141] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.
[0142] In this embodiment, the video processing device will be specifically integrated into an electronic device, with the electronic device serving as a server, the video to be processed being an advertising video, the item label being a product label, and the special effects image being a video sticker, as an example for illustration.
[0143] like Figure 8 As shown, a video processing method has the following specific steps:
[0144] 301. The server obtains the advertising video and the corresponding audience of the advertising video.
[0145] For example, the server receives an advertising video and its corresponding audience from a terminal. Alternatively, it can receive the advertising video and the page information of the audience selection page for the advertising video from the terminal. Based on the page information of the audience selection page, it filters at least one audience tag corresponding to the advertising video from a preset set of audience tags, and merges the audience tags to obtain the audience. Alternatively, it can filter the original video from the network or video database, select the target type video as the advertising video from the original video, and then identify the audience of the advertising video to obtain the audience. Or, when the advertising video has a large memory or a large number of videos, it can also receive a video processing request from the terminal. This video processing request carries the storage address of the advertising video and its corresponding audience. Based on this storage address, it retrieves the advertising video and its corresponding audience from the terminal's memory, cache, or a third-party database.
[0146] 302. The server extracts the video and audio streams from the advertisement video.
[0147] For example, the server can separate the video and audio information in the advertisement video to obtain the target video and target audio information. Based on the playback parameters, the target video information can be converted into a video stream and the target audio information can be converted into an audio stream. Alternatively, the server can directly extract the video and audio frames from the advertisement video, merge the video frames in the playback order to obtain a video stream, and merge the audio frames to obtain an audio stream.
[0148] 303. The server performs content recognition on the video stream and audio stream respectively to obtain the visual content and text content of the advertising video.
[0149] For example, the server extracts video frames from the video stream, obtaining a set of video frames. Within this set, it identifies the image information of each video frame to obtain the content of the advertising video. A text recognition model identifies text information appearing in the image content of each video frame; this text information can include the text appearing in the video frames, thus obtaining the video stream text. A speech-to-text model converts audio frames in the audio stream into subtitle information for the advertising video, extracting text information from the subtitle information to obtain the audio stream text. The video stream text and audio stream text are then used as the text content of the advertising video.
[0150] 304. The server selects at least one video tag from a preset set of video tags based on the screen content and text content.
[0151] For example, the server extracts features from the content of each video frame, then calculates the feature similarity between the extracted features and preset key features. Video frames with feature similarity exceeding a preset similarity threshold are selected as key video frames. Keyword extraction is performed on the audio stream text and video stream text. The audio stream keywords extracted from the audio stream text are compared with the video stream keywords extracted from the video stream text, and deduplication and filtering are performed to obtain at least one target keyword for the video to be processed.
[0152] The server filters style tags and product tags from a preset video tag set, obtaining style tag sets and product tag sets. Style features are extracted from key video frames to obtain the style features of the video to be processed. At least one style tag corresponding to the style features is then selected from the style tag sets to obtain the style tags for the video to be processed. Feature extraction is performed on target keywords to obtain the item features of the video to be processed. Based on these item features, at least one item information in the video to be processed is identified. Product tags corresponding to the item information are then selected from the product tag sets to obtain the product tags for the video to be processed.
[0153] The server identifies video tags in the video to be processed using a trained tag recognition model. The specific process involves extracting the video and audio streams from the video, performing content recognition on both streams using the trained model to obtain the video and text content. Based on the video content, at least one key video frame is selected from the video frames, and at least one target keyword is extracted from both the video and audio streams. Based on the key video frame, at least one style tag corresponding to the video to be processed is selected from a pre-defined style tag set. Based on the target keyword, at least one product tag corresponding to the video to be processed is selected from a pre-defined product tag set.
[0154] The trained label recognition model can be configured according to the needs of actual applications. It should also be noted that the trained label recognition model can be pre-configured by maintenance personnel or trained automatically by the video processing device. The training process can be as follows:
[0155] (1) The server obtains advertising video samples.
[0156] For example, the server can obtain at least one original video, send the original video to the tagging server, so that the tagging server can tag the original video with video tags, and receive the original video with tagged video tags returned by the tagging server, thereby obtaining an advertising video sample.
[0157] (2) The server uses a preset tag recognition model to predict the video tags of the advertising video samples and obtain the predicted video tags.
[0158] For example, the server extracts video stream samples and audio stream samples from the advertising video sample. A pre-defined tag recognition model is used to perform content recognition on the video stream samples and audio stream samples respectively, obtaining the image content samples and text content samples of the advertising video sample. Based on the image content samples, at least one key video frame sample is selected from the video frames of the video stream sample. At least one image keyword sample is extracted from the image content samples. The audio stream is converted into text information, and at least one audio keyword sample is extracted from the converted text information. The image keyword samples and audio keyword samples are merged to obtain the target keyword samples of the advertising video sample. Using the style tag recognition sub-model in the pre-defined tag recognition model, at least one style tag is predicted from the pre-defined style tag set based on the key video frame samples, obtaining the predicted style tag of the advertising video sample. Using the target keyword samples in the pre-defined tag recognition model, at least one product tag is predicted from the pre-defined product tag set, obtaining the predicted product tag of the advertising video sample. The predicted product tag and predicted style tag are used as the predicted video tag of the advertising video sample.
[0159] (3) The server converges the preset label recognition model based on the predicted video labels and the labeled video labels to obtain the trained label recognition model.
[0160] For example, the server compares the predicted style labels with the labeled style labels to obtain style loss information, and compares the predicted product labels with the labeled product labels to obtain item loss information. The style loss information and item loss information are then fused to obtain the label loss information for the advertising video sample. Based on this label loss information, the network parameters of the preset label recognition model are updated to converge the preset label recognition model, thus obtaining the trained label recognition model.
[0161] 305. The server determines the video sticker corresponding to the advertisement video based on the target audience and video tags.
[0162] For example, the server can receive preset video stickers and their tag selection information from the terminal. Based on this tag selection information, it can filter audience tags, style tags, and product tags from the tag library, and then merge these tags to generate attribute information for the preset video stickers. This attribute information indicates the set of style tags, product tags, and target audiences to which the preset image is applicable, thus obtaining a set of preset video stickers. Alternatively, the server can receive audience tags, style tags, and product tags from the tag library for each preset video sticker through manual selection by the terminal, and then merge these tags to generate attribute information for the preset video stickers, thus obtaining a set of preset video stickers.
[0163] The server can identify the target style tag set, target product tag set, and audience set corresponding to the preset video stickers from the attribute information. It then matches the audiences with their respective sets, matches the style tags with the target style tag set, and matches the product tags with the target product tag set. Finally, it filters out preset video stickers from the preset video sticker set that completely match the audience, style tags, and product tags of the advertisement video, thus obtaining candidate video stickers.
[0164] In this context, a successful audience match means that the target audience set of the preset video sticker includes the target audience of the advertisement video. A successful style tag match means that the target style tag set of the preset video sticker includes the style tag of the advertisement video. A successful product tag match means that the target product tag set of the preset video sticker includes the product tag of the advertisement video.
[0165] When there is only one candidate video sticker, the server uses it as the video sticker corresponding to the advertisement video. When there are multiple candidate video stickers, the server extracts features from the attribute information of the candidate video stickers to obtain the special effect features of each candidate video sticker. The server then extracts special effect features from the audience, style tags, and product tags of the advertisement video, and merges the extracted basic special effect features to obtain the target special effect features of the advertisement video. The server calculates the special effect similarity between the target special effect features and the special effect features of each candidate video sticker, and selects the video sticker corresponding to the advertisement video from the candidate video stickers based on the special effect similarity.
[0166] 306. The server adds the video sticker to the advertisement video, obtains the processed advertisement video, and sends the processed advertisement video to the terminal.
[0167] For example, the server obtains the location information of a video sticker. When the location information indicates the position where the video sticker will appear in a video frame, the server can directly identify the coordinates and other location information corresponding to that position in each video frame of the advertisement video. This identified coordinates and other location information is then used as the addition position for the video sticker. When the location information indicates the required size information within a video frame, the server can identify the blank area containing that size information in each video frame of the advertisement video and obtain the location information of that blank area, thus determining the addition position for the video sticker in the video frame. Based on the addition position, the video sticker is added to each video frame, resulting in a processed video. This processed video is then sent to the terminal, allowing the terminal to edit the video stickers within the processed video to obtain the target video.
[0168] 307. The terminal edits the special effects images in the processed advertising video to obtain the target advertising video.
[0169] For example, a preview page is displayed showing the processed video after adding special effects images. This preview page includes editing controls for the special effects images. In response to editing operations on these controls, users can edit the text content of the special effects images, modify their position and size, replace or delete them. Each time an editing operation is performed on the special effects images, the preview page is updated to display the updated preview. Once the user has finished editing the special effects images, the terminal generates the target advertising video in response to application operations on the application controls.
[0170] In the advertising business, video ads, due to their rich content attributes, can convey more advertising information to users compared to text and image ads. However, their production costs are higher, making it difficult for small and medium-sized advertisers to create videos that fully showcase the characteristics of their ads and achieve good advertising results. The video processing in this solution can be as follows: Figure 9 As shown, advertisers fill in the target audience information and upload the advertising video. The advertising video is then input into the tag recognition model for tag recognition, which identifies the video's style tags and category tags. The sticker recommendation system then filters out a benefit sticker that matches the target audience, style, and product tags. The video preview page displays the video with the sticker. Advertisers can edit the sticker text, position, and size. After editing, the advertiser generates a finished video with the sticker for the advertisement.
[0171] As can be seen from the above, in this embodiment, after the server obtains the advertising video and the corresponding audience, it identifies the content of the advertising video to filter out at least one video tag from a preset video tag set. Then, based on the audience and video tag, it determines the video sticker corresponding to the advertising video, adds the video sticker to the advertising video, obtains the processed video, and sends the processed video to the terminal so that the terminal can edit the video sticker in the processed video. Since this solution filters out special effects images through two different factors, audience and video tag, in the process of determining the video sticker of the advertising video, it fully considers the influence of the audience of the advertising video in the business scenario that needs to be disseminated, thereby increasing the accuracy of the video sticker selection. Therefore, it can improve the accuracy of the video processing process.
[0172] Based on the method described in the above embodiments, the following examples will provide further detailed explanations.
[0173] In this embodiment, the video processing device is specifically integrated into an electronic device, which is a server. The video to be processed is an advertising video, the item tag is a product tag, the special effects image is a video sticker, and a certain video sticker in the sticker library is a Double Eleven tag with the default text "Discounted Products Big Promotion". Its target audience tag is all people, the style tag is relaxed, the product content tag is unspecified, the selling point tag is holiday activities, and the holiday tag is Double Eleven.
[0174] For example, in the ad creative library, select N ad videos and manually label them with style and product information tags from a tag library. Input these videos into a deep learning model and train them using the tags as the target values. Then, select M different ad videos, input them into the deep learning model, and obtain the output storyline, style, and product information tags. For instance, an ad video might feature a live streamer promoting a beauty product during a Double 11 promotion, with upbeat background music. In this case, during manual labeling, the video's style tag could be set to "upbeat," the product content tag to "beauty," the selling point tag to "holiday promotion," and the holiday tag to "Double 11."
[0175] When an advertiser uploads a video with a width of 720 pixels and a height of 1280 pixels, and the algorithm analyzes the video's style tag as "light and cheerful," the product content tag as "beauty," the selling point tag as "holiday event," and the holiday tag as "Double 11," and the advertiser selects a female target audience, the sticker recommendation system may recommend Double 11 stickers with the default text "Big Sale on Discounted Products." On the preview page, the sticker may appear in the upper left corner of the video, covering an area 150 pixels wide and 50 pixels high, and will appear throughout the video playback. The advertiser can move the sticker to the upper right corner or change it to another sticker. Based on the final preview adjustments, a 720-pixel wide and 1280-pixel high video with the sticker will be generated for the advertiser. The sticker will be a static Double 11 image sticker with the text "Big Sale on Discounted Products," appearing in the upper right corner of the video, covering an area 150 pixels wide and 50 pixels high, and will appear throughout the video playback. See details below. Figure 10 As shown.
[0176] To better implement the above methods, embodiments of the present invention also provide a video processing device (i.e., a first video processing device), which can be integrated into an electronic device. The electronic device can be a server, which can be a single server or a server cluster composed of multiple servers.
[0177] For example, such as Figure 11 As shown, the video processing device may include an acquisition unit 401, an identification unit 402, a determination unit 403, and an addition unit 404, as follows:
[0178] (1) Obtain unit 401;
[0179] The acquisition unit 401 is used to acquire the video to be processed and the audience information corresponding to the video to be processed. The audience information is used to indicate the audience information targeted by the video to be processed.
[0180] For example, the acquisition unit 401 can be specifically used to receive the video to be processed and the audience information corresponding to the video to be processed sent by the terminal; or, it can receive the video to be processed and the page information of the audience selection page of the video to be processed sent by the terminal, and based on the page information of the audience selection page, filter at least one audience tag corresponding to the video to be processed from a preset audience tag set, and merge the audience tags to obtain the audience information; or, it can filter the original video from the network or video database, filter the target type video from the original video as the video to be processed, and then identify the audience of the video to be processed to obtain the audience information; or, when the memory of the video to be processed is large or the number is large, it can also receive the video processing request sent by the terminal, which carries the storage address of the video to be processed and the audience information corresponding to the video to be processed, and based on the storage address, obtain the video to be processed and the audience information corresponding to the video to be processed from the terminal's memory, cache or third-party database.
[0181] (2) Identification unit 402;
[0182] The recognition unit 402 is used to recognize the content of the video to be processed, so as to filter out at least one video tag from the preset video tag set.
[0183] For example, the recognition unit 402 can specifically be used to extract video streams and audio streams from the video to be processed, extract video frames from the video streams to obtain a set of video frames, identify the image information of each video frame in the set of video frames to obtain the image content of the video to be processed, identify text information in the image content to obtain video stream text, convert the audio stream into text information to obtain audio stream text, and use the video stream text and audio stream text as the text content of the video to be processed. Based on the image content, at least one key video frame is selected from the set of video frames, keywords are extracted from the audio stream text and video stream text, and the extracted keywords are fused to obtain at least one target keyword for the video to be processed. Based on the key video frame and the target keyword, at least one video tag is selected from a preset tag set.
[0184] (3) Determine unit 403;
[0185] The determining unit 403 is used to determine the special effects image corresponding to the video to be processed based on the audience information and video tags. The special effects image is used to enhance the video performance of the video to be processed.
[0186] For example, the determining unit 303 can be used to obtain a preset special effects image set, which includes at least one preset special effects image and attribute information of the preset special effects image. The attribute information of the special effects image is matched with the audience information, style tags and item tags of the video to be processed, respectively. Based on the matching results, the special effects image corresponding to the image to be processed is selected from the preset special effects image set.
[0187] (4) Add unit 404;
[0188] The adding unit 404 is used to add special effects images to the video to be processed to obtain a processed video, and send the processed video to the terminal so that the terminal can edit the special effects images in the processed video.
[0189] For example, the adding unit 404 can be used to obtain the position information of the special effects image. Based on the position information, the adding position of the special effects image is identified in each video frame of the video to be processed. According to the adding position, the special effects image is added to each video frame respectively to obtain the processed video. The processed video is sent to the terminal so that the terminal can edit the special effects image in the processed video.
[0190] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.
[0191] As can be seen from the above, in this embodiment, after the acquisition unit 401 acquires the video to be processed and the audience information corresponding to the video to be processed, the identification unit 402 identifies the content of the video to be processed to filter out at least one video tag from the preset video tag set. Then, the determination unit 403 determines the special effects image corresponding to the video to be processed based on the audience information and the video tag. Then, the addition unit 404 adds the special effects image to the video to be processed to obtain the processed video, and sends the processed video to the terminal so that the terminal can edit the special effects image in the processed video. Since this scheme filters out the special effects image by two different factors, audience information and video tag, in the process of determining the special effects image of the video to be processed, it fully considers the influence of the audience of the video to be processed in the business scenario that needs to be disseminated, thereby increasing the accuracy of filtering out the special effects image. Therefore, it can improve the accuracy of the video processing process.
[0192] To better implement the above methods, embodiments of the present invention also provide a video processing device (i.e., a second video processing device), which can be integrated into a terminal, including a tablet computer, a laptop computer, and / or a personal computer.
[0193] For example, such as Figure 12 As shown, the first data processing device may include a sending unit 501, a display unit 502, and a generating unit 503, as follows:
[0194] (1) Transmitting unit 501;
[0195] The sending unit 501 is used to send the video to be processed and the audience information corresponding to the video to be processed to the server.
[0196] For example, the sending unit 501 can be used to directly send the video to be processed and the audience information corresponding to the video to be processed to the server. Alternatively, it can store the video to be processed and the audience information corresponding to the video to be processed, add the storage address to the video processing request, and send the video processing request with the added storage address to the server, so that the server can obtain the video to be processed and the audience information corresponding to the video to be processed according to the storage address carried in the video processing request.
[0197] (2) Display unit 502;
[0198] Display unit 502 is used to display a preview page of the processed video after adding special effects images to the video to be processed. The preview page includes editing controls for the special effects images.
[0199] For example, the display unit 502 can be used to add the processed video to a preset page and play the processed video on the preset page to obtain a preview page, which may also include editing controls for special effects images.
[0200] (3) Generation unit 503;
[0201] The generation unit 503 is used to generate the target video in response to an editing operation on the editing control.
[0202] For example, the generation unit 503 can be specifically used to adjust the special effects image in the processed video in response to the editing operation of the editing control, update the preview page of the processed video based on the adjusted special effects image, and display the updated preview image. The updated preview page includes a subtitle generation control and an application control. In response to the subtitle generation operation of the subtitle generation control, subtitle information is added to the video in the updated preview page to obtain candidate videos. In response to the application operation of the application control, the candidate videos are used as target videos.
[0203] In practice, each of the above units can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units, please refer to the previous method embodiments, which will not be repeated here.
[0204] As can be seen from the above, in this embodiment of the application, after the sending unit 501 sends the video to be processed and the audience information corresponding to the video to be processed to the server, the display unit 502 displays a preview page of the processed video after adding special effects images to the video to be processed. The preview page includes editing controls for the special effects images. The generation unit 503 generates the target video in response to the editing operation on the editing controls. Since this solution filters out special effects images through audience information and can edit the special effects images to generate the target video, the accuracy of video processing can be improved.
[0205] This invention also provides an electronic device, such as... Figure 13 As shown, it illustrates a structural schematic diagram of the electronic device involved in an embodiment of the present invention, specifically:
[0206] The electronic device may include components such as a processor 601 with one or more processing cores, a memory 602 with one or more computer-readable storage media, a power supply 603, and an input unit 604. Those skilled in the art will understand that... Figure 13 The electronic device structure shown does not constitute a limitation on the electronic device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:
[0207] The processor 601 is the control center of the electronic device. It connects various parts of the electronic device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 602, and by calling data stored in the memory 602, it performs various functions and processes data, thereby managing the electronic device as a whole. Optionally, the processor 601 may include one or more processing cores; preferably, the processor 601 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 601.
[0208] The memory 602 can be used to store software programs and modules. The processor 601 executes various functional applications and data processing by running the software programs and modules stored in the memory 602. The memory 602 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 602 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 602 may also include a memory controller to provide the processor 601 with access to the memory 602.
[0209] The electronic device also includes a power supply 603 that supplies power to the various components. Preferably, the power supply 603 can be logically connected to the processor 601 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 603 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0210] The electronic device may also include an input unit 604, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0211] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 601 in the electronic device loads the executable files corresponding to the processes of one or more applications into the memory 602 according to the following instructions, and the processor 601 runs the applications stored in the memory 602 to realize various functions, as follows:
[0212] The system acquires the video to be processed and its corresponding audience information, which indicates the target audience of the video. It then identifies the content of the video to be processed, selecting at least one video tag from a preset video tag set. Based on the audience information and video tags, it determines the special effects image corresponding to the video to be processed. This special effects image enhances the video's visual appeal. The special effects image is added to the video to be processed, resulting in a processed video. The processed video is then sent to a terminal for editing of the special effects image.
[0213] or
[0214] The system sends the video to be processed to the server and displays a preview page of the processed video after adding special effects images. The processed video is the video to be processed with added special effects images. The preview page includes editing controls for the special effects images. In response to editing operations on the editing controls, the target video is generated.
[0215] For example, the system can receive the video to be processed and the corresponding audience information sent by the terminal; or, it can receive the video to be processed and the page information of the audience selection page for the video to be processed sent by the terminal, and based on the page information of the audience selection page, filter at least one audience tag corresponding to the video to be processed from a preset set of audience tags, and merge the audience tags to obtain the audience information; or, it can filter the original video from the network or video database, filter the target type video from the original video as the video to be processed, and then identify the audience of the video to be processed to obtain the audience information; or, when the memory of the video to be processed is large or the number is large, it can also receive a video processing request sent by the terminal, which carries the storage address of the video to be processed and the corresponding audience information of the video to be processed, and based on the storage address, retrieve the video to be processed and the corresponding audience information of the video to be processed from the terminal's memory, cache, or third-party database.
[0216] The process involves extracting video and audio streams from the video to be processed. Video frames are extracted from the video streams to obtain a set of video frames. The image information of each video frame in the set is identified to obtain the image content of the video to be processed. Text information is then identified from the image content to obtain the video stream text. The audio stream is converted into text information to obtain the audio stream text. Both the video stream text and the audio stream text are used as the text content of the video to be processed. Based on the image content, at least one key video frame is selected from the set of video frames. Keywords are extracted from the audio stream text and the video stream text, and the extracted keywords are fused to obtain at least one target keyword for the video to be processed. Based on the key video frame and the target keyword, at least one video tag is selected from a preset tag set.
[0217] A set of preset special effects images is obtained, which includes at least one preset special effects image and its attribute information. The attribute information of the special effects images is matched with the audience information, style tags, and item tags of the video to be processed. Based on the matching results, the special effects images corresponding to the images to be processed are selected from the set of preset special effects images. The position information of the special effects images is obtained. Based on the position information, the addition position of the special effects image is identified in each video frame of the video to be processed. According to the addition position, the special effects image is added to each video frame respectively to obtain the processed video. The processed video is sent to the terminal. The terminal receives the processed video generated based on the audience information returned by the server. The processed video includes special effects images. A preview page of the processed video is displayed. The preview page includes editing controls for the special effects images. In response to editing operations on the editing controls, the special effects images are adjusted in the processed video. Based on the adjusted special effects images, the preview page of the processed video is updated, and the updated preview image is displayed. The updated preview page includes a subtitle generation control and an application control. In response to the subtitle generation operation of the subtitle generation control, subtitle information is added to the video in the updated preview page to obtain candidate videos. In response to the application operation of the application control, the candidate videos are used as the target videos.
[0218] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0219] As can be seen from the above, after obtaining the video to be processed and the corresponding audience information, this embodiment of the application identifies the content of the video to be processed to filter out at least one video tag from a preset video tag set. Then, based on the audience information and video tags, it determines the special effects image corresponding to the video to be processed. The special effects image is then added to the video to be processed to obtain the processed video, and the processed video is sent to the terminal so that the terminal can edit the special effects image in the processed video. Since this scheme filters out the special effects image by using two different factors, audience information and video tags, in the process of determining the special effects image of the video to be processed, it fully considers the influence of the audience of the video to be processed in the business scenario that needs to be disseminated, thereby increasing the accuracy of filtering out the special effects image. Therefore, it can improve the accuracy of the video processing process.
[0220] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0221] Therefore, embodiments of the present invention provide a computer-readable storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the video processing methods provided in the embodiments of the present invention. For example, the instructions can execute the following steps:
[0222] The system acquires the video to be processed and its corresponding audience information, which indicates the target audience of the video. It then identifies the content of the video to be processed, selecting at least one video tag from a preset video tag set. Based on the audience information and video tags, it determines the special effects image corresponding to the video to be processed. This special effects image enhances the video's visual appeal. The special effects image is added to the video to be processed, resulting in a processed video. The processed video is then sent to a terminal for editing of the special effects image.
[0223] or
[0224] The system sends the video to be processed to the server and displays a preview page of the processed video after adding special effects images. The processed video is the video to be processed with added special effects images. The preview page includes editing controls for the special effects images. In response to editing operations on the editing controls, the target video is generated.
[0225] For example, the system can receive the video to be processed and the corresponding audience information sent by the terminal; or, it can receive the video to be processed and the page information of the audience selection page for the video to be processed sent by the terminal, and based on the page information of the audience selection page, filter at least one audience tag corresponding to the video to be processed from a preset set of audience tags, and merge the audience tags to obtain the audience information; or, it can filter the original video from the network or video database, filter the target type video from the original video as the video to be processed, and then identify the audience of the video to be processed to obtain the audience information; or, when the memory of the video to be processed is large or the number is large, it can also receive a video processing request sent by the terminal, which carries the storage address of the video to be processed and the corresponding audience information of the video to be processed, and based on the storage address, retrieve the video to be processed and the corresponding audience information of the video to be processed from the terminal's memory, cache, or third-party database. The process involves extracting video and audio streams from the video to be processed. Video frames are extracted from the video streams to obtain a set of video frames. The image information of each video frame in the set is identified to obtain the image content of the video to be processed. Text information is then identified from the image content to obtain the video stream text. The audio stream is converted into text information to obtain the audio stream text. Both the video stream text and the audio stream text are used as the text content of the video to be processed. Based on the image content, at least one key video frame is selected from the set of video frames. Keywords are extracted from the audio stream text and the video stream text, and the extracted keywords are fused to obtain at least one target keyword for the video to be processed. Based on the key video frame and the target keyword, at least one video tag is selected from a preset tag set.
[0226] A set of preset special effects images is obtained, which includes at least one preset special effects image and its attribute information. The attribute information of the special effects images is matched with the audience information, style tags, and item tags of the video to be processed. Based on the matching results, the special effects images corresponding to the images to be processed are selected from the set of preset special effects images. The position information of the special effects images is obtained. Based on the position information, the addition position of the special effects image is identified in each video frame of the video to be processed. According to the addition position, the special effects image is added to each video frame respectively to obtain the processed video. The processed video is sent to the terminal. The terminal receives the processed video generated based on the audience information returned by the server. The processed video includes special effects images. A preview page of the processed video is displayed. The preview page includes editing controls for the special effects images. In response to editing operations on the editing controls, the special effects images are adjusted in the processed video. Based on the adjusted special effects images, the preview page of the processed video is updated, and the updated preview image is displayed. The updated preview page includes a subtitle generation control and an application control. In response to the subtitle generation operation of the subtitle generation control, subtitle information is added to the video in the updated preview page to obtain candidate videos. In response to the application operation of the application control, the candidate videos are used as the target videos.
[0227] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0228] The computer-readable storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0229] Since the instructions stored in the computer-readable storage medium can execute the steps of any of the video processing methods provided in the embodiments of the present invention, the beneficial effects that any of the video processing methods provided in the embodiments of the present invention can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0230] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations of the video processing or advertising video adding video stickers described above.
[0231] The foregoing has provided a detailed description of a video processing method, apparatus, electronic device, and computer-readable storage medium provided by embodiments of the present invention. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, those skilled in the art will recognize that, based on the ideas of the present invention, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A video processing method, characterized in that, Applied to servers, including: Obtaining the video to be processed and the audience information corresponding to the video to be processed, wherein the audience information is used to indicate the audience target of the video to be processed, the video to be processed is a video advertisement, and the audience information of the video to be processed is the audience of the video advertisement; identifying the audience corresponding to the video to be processed based on the video to be processed includes: extracting multi-dimensional features from the video to be processed, fusing the extracted audience features to obtain global audience features, and filtering the audience corresponding to the video to be processed from a preset audience set based on the global audience features, thereby obtaining the audience information of the video to be processed; The content of the video to be processed is identified so as to filter out at least one video tag from a preset video tag set, the video tag including style tag and item tag; Based on the audience information and video tags, a special effects image corresponding to the video to be processed is determined, and the special effects image is used to enhance the video performance of the video to be processed; The special effects image is added to the video to be processed to obtain the processed video, and the processed video is sent to the electronic device terminal so that the electronic device terminal can edit the special effects image in the processed video; The step of identifying the content of the video to be processed, in order to filter out at least one video tag from a preset video tag set, includes: Extract the video stream and audio stream from the video to be processed; Content recognition is performed on the video stream and audio stream respectively to obtain the screen content and text content of the video to be processed; Based on the video content and text content, at least one video tag is selected from the preset video tag set; The step of performing content recognition on the video stream and audio stream respectively to obtain the image content and text content of the video to be processed includes: Extract video frames from the video stream to obtain a video frame set; The image information of each video frame is identified in the set of video frames to obtain the image content of the video to be processed. Text information is identified from the video content and audio stream to obtain the text content of the video to be processed; The step of identifying text information in the video content and audio stream to obtain the text content of the video to be processed includes: Text information is identified from the video content to obtain the video stream text. The audio stream is converted into text information to obtain audio stream text, and the video stream text and audio stream text are used as the text content of the video to be processed.
2. The video processing method according to claim 1, characterized in that, The step of selecting at least one video tag from a preset video tag set based on the image content and text content includes: Based on the content of the image, at least one key video frame is selected from the set of video frames. Keyword extraction is performed on the audio stream text and video stream text, and the extracted keywords are fused to obtain at least one target keyword for the video to be processed; Based on the key video frames and target keywords, at least one video tag is selected from a preset set of video tags.
3. The video processing method according to claim 2, characterized in that, The step of selecting at least one video tag from a preset video tag set based on the key video frames and target keywords includes: Filter out style tags and item tags from the preset video tag set to obtain style tag set and item tag set; Based on the key video frames, at least one style tag corresponding to the video to be processed is selected from the style tag set; Based on the target keywords, at least one item tag corresponding to the video to be processed is selected from the item tag set. The item tag is used to indicate information about the items contained in the video to be processed.
4. The video processing method according to claim 3, characterized in that, The step of filtering at least one style tag corresponding to the video to be processed from the style tag set based on the key video frames includes: Style features are extracted from the key video frames to obtain the style features of the video to be processed; At least one style tag corresponding to the style feature is selected from the style tag set to obtain the style tag corresponding to the video to be processed.
5. The video processing method according to claim 3, characterized in that, The step of filtering at least one item tag corresponding to the video to be processed from the item tag set based on the target keyword includes: Feature extraction is performed on the target keywords to obtain the item features of the video to be processed; Based on the characteristics of the items, at least one item information in the video to be processed is determined; The item tags corresponding to the item information are filtered out from the item tag set to obtain the item tags of the video to be processed.
6. The video processing method according to claim 3, characterized in that, The step of determining the special effects image corresponding to the video to be processed based on the audience information and video tags includes: Obtain a set of preset special effects images, the set of preset special effects images including at least one preset special effects image and the attribute information of the preset special effects image; The attribute information of the special effects image is matched with the audience information, style tags, and item tags of the video to be processed; Based on the matching results, the special effects images corresponding to the image to be processed are selected from the preset set of special effects images.
7. The video processing method according to claim 6, characterized in that, The step of matching the attribute information of the special effects image with the audience information, style tags, and item tags of the video to be processed includes: The target style tag set, target item tag set, and audience information set corresponding to the preset special effects image are identified from the attribute information. The audience information is matched with the audience information set, the style tag is matched with the target style tag set, and the item tag is matched with the target item tag set. The step of filtering out the special effects image corresponding to the image to be processed from the preset special effects image set based on the matching results includes: filtering out preset special effects images that completely match the audience information, style tags, and item tags of the video to be processed from the preset special effects image set, so as to obtain the special effects image corresponding to the video to be processed.
8. The video processing method according to claim 7, characterized in that, The step of selecting preset special effects images from the preset special effects image set that completely match the audience information, style tags, and item tags of the video to be processed, in order to obtain the special effects image corresponding to the video to be processed, includes: Candidate effect images are obtained by filtering out preset effect images from the preset effect image set that match all the audience object information, style tags, and item tags of the video to be processed. When there is only one candidate effect image, the candidate effect image is used as the effect image corresponding to the video to be processed. When there are multiple candidate effect images, the effect image corresponding to the video to be processed is selected from the candidate effect images according to the attribute information of the candidate effect images.
9. The video processing method according to any one of claims 1 to 8, characterized in that, The step of adding the special effects image to the video to be processed to obtain the processed video includes: Obtain the location information of the special effects image; Based on the location information, the location where the special effects image was added is identified in each video frame of the video to be processed; According to the addition position, the special effects image is added to each video frame to obtain the processed video.
10. A video processing method, characterized in that, Applied to electronic device terminals, including: Send the video to be processed to the server; Receive the processed video sent by the server, wherein the processed video is a video obtained by adding special effects images to the video to be processed; A preview page is displayed showing the processed video after adding special effects images to the video to be processed. The preview page includes editing controls for the special effects images. In response to an editing operation on the editing control, a target video is generated; The processed video is obtained using the following video processing method: The system retrieves the video to be processed from the electronic device terminal and the audience information corresponding to the video to be processed. The audience information is used to indicate the information of the audience to which the video to be processed is targeted. The video to be processed is a video advertisement, and the audience of the video to be processed is the audience of the video advertisement. The audience information includes, but is not limited to, the user range of the audience, the user's age, the user's education level, the user's location, the user's identity, and the user's job type. The content of the video to be processed is identified so as to filter out at least one video tag from a preset video tag set, the video tag including style tag and item tag; Based on the audience information and video tags, a special effects image corresponding to the video to be processed is determined, and the special effects image is used to enhance the video performance of the video to be processed; The special effects image is added to the video to be processed to obtain the processed video.
11. The video processing method according to claim 10, characterized in that, The preview page also includes special effects evaluation data for the processed video, which is used to evaluate the effect data after adding special effects images to the video to be processed.
12. The video processing method according to claim 10, characterized in that, The step of generating a target video in response to an editing operation on the editing control includes: In response to an editing operation on the editing control, the special effects image is adjusted in the processed video; Based on the adjusted special effects image, the preview page of the processed video is updated and the updated preview page is displayed. The updated preview page includes a subtitle generation control and an application control. In response to the subtitle generation operation of the subtitle generation control, subtitle information is added to the video in the updated preview page to obtain candidate videos; In response to the application operation of the application control, the candidate video is selected as the target video.
13. A video processing apparatus, characterized in that, include: The acquisition unit is used to acquire the video to be processed and the audience information corresponding to the video to be processed. The audience information is used to indicate the information of the audience to which the video to be processed is targeted. The video to be processed is a video advertisement, and the audience information of the video to be processed is the audience of the video advertisement. The acquisition unit is specifically used for: extracting multi-dimensional features from the video to be processed, fusing the extracted audience features to obtain global audience features, and filtering out the audience objects corresponding to the video to be processed from a preset audience object set based on the global audience features, thereby obtaining the audience object information of the video to be processed. The identification unit is used to identify the content of the video to be processed, so as to filter out at least one video tag from a preset video tag set, the video tag including style tag and item tag; The determining unit is used to determine the special effects image corresponding to the video to be processed based on the audience information and video tags, wherein the special effects image is used to enhance the video performance of the video to be processed; An adding unit is used to add the special effects image to the video to be processed to obtain a processed video, and send the processed video to an electronic device terminal so that the electronic device terminal can edit the special effects image in the processed video; The identification unit is specifically used to: extract the video stream and audio stream from the video to be processed; Content recognition is performed on the video stream and audio stream respectively to obtain the screen content and text content of the video to be processed; Based on the video content and text content, at least one video tag is selected from the preset video tag set; The identification unit is specifically used to: extract video frames from the video stream to obtain a set of video frames; The image information of each video frame is identified in the set of video frames to obtain the image content of the video to be processed. Text information is identified from the video content and audio stream to obtain the text content of the video to be processed; The recognition unit is specifically used to: identify text information in the screen content to obtain video stream text; The audio stream is converted into text information to obtain audio stream text, and the video stream text and audio stream text are used as the text content of the video to be processed.
14. A video processing apparatus, characterized in that, include: The sending unit is used to send the video to be processed to the server; The sending unit is used to receive the processed video sent by the server, wherein the processed video is a video obtained by adding special effects images to the video to be processed; The display unit is used to display a preview page of the processed video after adding special effects images to the video to be processed, and the preview page includes editing controls for the special effects images; A generation unit is configured to generate a target video in response to an editing operation on the editing control; The processed video was obtained using the following video processing method: The system retrieves the video to be processed from the electronic device terminal and the audience information corresponding to the video to be processed. The audience information is used to indicate the information of the audience to which the video to be processed is targeted. The video to be processed is a video advertisement, and the audience of the video to be processed is the audience of the video advertisement. The audience information includes, but is not limited to, the user range of the audience, the user's age, the user's education level, the user's location, the user's identity, and the user's job type. The content of the video to be processed is identified so as to filter out at least one video tag from a preset video tag set, the video tag including style tag and item tag; Based on the audience information and video tags, a special effects image corresponding to the video to be processed is determined, and the special effects image is used to enhance the video performance of the video to be processed; The special effects image is added to the video to be processed to obtain the processed video.
15. An electronic device, characterized in that, It includes a processor and a memory, the memory storing an application program, and the processor running the application program within the memory to perform the steps of the video processing method according to any one of claims 1 to 12.
16. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the steps of the video processing method according to any one of claims 1 to 12.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a plurality of instructions adapted for loading by a processor to perform the steps of the video processing method according to any one of claims 1 to 12.
Citation Information
Patent Citations
Information recommending method based on broadcast program, electronic equipment and server
CN106231362A
Business object recommendation method and device, memory medium and electronic device
CN108076353A
Information processing method and device, electronic equipment and storage medium
CN110688498A
Video and image processing method and device, electronic equipment and storage medium
CN111541936A
Music recommendation method and device and readable storage medium
CN113569088A