Video processing method and device, electronic equipment, computer readable storage medium and computer program product
Patent Information
- Application Number
- CN202510370701.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2026-09-29
AI Technical Summary
相关技术中,在对广告视频进行自动化修改的过程中缺乏品牌元素保护机制,容易导致品牌方的LOGO变形、或者产品展示不完整等问题
[0019]通过将视频划分成可修改区域和禁止修改区域,并在对视频进行修改的过程中,保持禁止修改区域不变,仅对可修改区域进行更新,如此,一方面能够对视频进行修改,另一方面也可以保证禁止修改区域显示的目标元素的完整性,避免了在对视频进行修改的过程中造成目标元素的变形或者显示不完整等问题。
Smart Images

Figure CN122845844A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of Internet technology, and in particular to a video processing method, apparatus, electronic device, computer-readable storage medium, and computer program product. Background Technology
[0002] With the rise of live streaming and short video platforms, more and more brands are using these platforms to place advertising videos. However, the automated modification of these advertising videos often lacks a mechanism to protect brand elements, easily leading to issues such as distorted brand logos or incomplete product displays. Summary of the Invention
[0003] This application provides a video processing method, apparatus, electronic device, computer-readable storage medium, and computer program product that can ensure the integrity of target elements during video editing.
[0004] The technical solution of this application embodiment is implemented as follows:
[0005] This application provides a video processing method, including:
[0006] The video interface includes a video, which includes an editable area and an uneditable area.
[0007] In response to a video editing trigger, display the editing controls;
[0008] In response to an editing operation on the modifiable region based on the editing control, the modifiable region is updated according to the editing operation.
[0009] This application provides a video processing apparatus, including:
[0010] A display module is used to display a video interface, wherein the video interface includes a video, and the video includes an editable area and an uneditable area;
[0011] The display module is also used to display editing controls in response to a video editing trigger operation;
[0012] An update module is configured to update the modifiable region in response to an editing operation performed on the modifiable region based on the editing control.
[0013] This application provides an electronic device, including:
[0014] Memory is used to store executable instructions for a computer;
[0015] The processor, when executing computer-executable instructions stored in the memory, implements the video processing method provided in the embodiments of this application.
[0016] This application provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the video processing method provided in this application.
[0017] This application provides a computer program product, including a computer program or computer executable instructions, which, when executed by a processor, implements the video processing method provided in this application.
[0018] The embodiments of this application have the following beneficial effects:
[0019] By dividing the video into modifiable and non-modifiable areas, and keeping the non-modifiable areas unchanged while updating only the modifiable areas during the video modification process, the video can be modified while ensuring the integrity of the target elements displayed in the non-modifiable areas. This avoids problems such as deformation or incomplete display of target elements during the video modification process. Attached Figure Description
[0020] Figure 1 This is a schematic diagram of the architecture of the video processing system 100 provided in an embodiment of this application;
[0021] Figure 2 This is a schematic diagram of the structure of the electronic device 500 provided in the embodiments of this application;
[0022] Figure 3 This is a first flowchart illustrating the video processing method provided in this application embodiment;
[0023] Figure 4 This is a second flowchart illustrating the video processing method provided in the embodiments of this application;
[0024] Figure 5 This is a schematic diagram of the third process of the video processing method provided in the embodiments of this application;
[0025] Figure 6A This is a schematic diagram illustrating the annotation of unmodifiable areas in a video frame, provided in an embodiment of this application.
[0026] Figure 6B This is a schematic diagram illustrating the annotation of modifiable areas in a video frame, provided in an embodiment of this application.
[0027] Figure 7 This is a schematic diagram illustrating an application scenario of the video processing method provided in the embodiments of this application;
[0028] Figure 8 This is a schematic diagram of the architecture of the video processing system provided in the embodiments of this application;
[0029] Figure 9 This is a schematic diagram of the video processing method provided in the embodiments of this application. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this application clearer, the application will be further described in detail below with reference to the accompanying drawings. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0031] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0032] It is understood that in the embodiments of this application, data such as user information are involved. When the embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of related data must comply with relevant laws, regulations and standards.
[0033] In the following description, the terms “first, second, ...” are used merely to distinguish similar objects and do not represent a specific ordering of objects. It is understood that “first, second, ...” may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0034] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0035] Before providing a further detailed description of the embodiments of this application, the nouns and terms involved in the embodiments of this application will be explained, and the nouns and terms involved in the embodiments of this application shall be interpreted as follows.
[0036] 1) Responding to: used to indicate the conditions or states on which the operation is performed depends. When the conditions or states on which it depends are met, one or more operations can be performed in real time or with a set delay. Unless otherwise specified, there is no restriction on the order in which the multiple operations are performed.
[0037] 2) Feature data: also known as feature variables, is a set of variables used in data science to describe the attributes of data points. These variables provide information about the data points, and feature data can be numerical, Boolean, or other types of data.
[0038] 3) Image recognition models: These are machine learning models used to automatically identify and classify objects, scenes, or concepts in images. These models are typically based on deep learning, especially convolutional neural networks, including YOLO, Mask R-CNN, ResNet, and VGGNet. YOLO is a real-time object detection model that achieves fast and accurate detection by dividing the image into a grid and predicting the object category and bounding box in each grid. Mask R-CNN is an instance segmentation model that generates an accurate segmentation mask for each detected object, building upon object detection. ResNet solves the vanishing gradient problem in deep neural networks by introducing residual connections, making it possible to train very deep networks. VGGNet uses multiple 3×3 convolutional and pooling layers, resulting in a very simple structure.
[0039] Taking video advertising scenarios as an example, intelligent video generation in the digital advertising field currently generally adopts static templates or limited dynamic solutions, which presents a dual contradiction of low user interest matching and insufficient protection of brand elements. Specifically, in the digital advertising field, dynamic video generation technology is gradually replacing traditional static template ads, and the current mainstream solutions are mainly based on the following two technical routes:
[0040] 1. Template-based replacement technology: Advertisers provide pre-set templates (such as text placeholders and image replacement areas), and the system replaces partial content (such as user names, product images, etc.) based on user tags. Typical applications include replacing product images and price information in e-commerce promotional advertisements.
[0041] 2. Collaborative filtering recommendation technology: Based on user behavior data (such as clicks, purchase records, etc.), it matches similar content in a pre-stored video library and selects the optimal version for delivery through A / B testing. Typical applications include: short video platforms recommending different styles of ad clips based on user interests.
[0042] However, the applicant discovered the following problems with the above solution during the implementation of the embodiments of this application:
[0043] 1. Lack of dynamic modification capabilities: The solutions provided by related technologies only support the replacement of text / image placeholders and cannot perform pixel-level redrawing of video images (such as replacing background items, adjusting visual styles, etc.), resulting in serious homogenization of advertising content and easy user fatigue.
[0044] 2. Inefficient use of user feature data: The solutions provided by related technologies rely only on basic tags (such as gender, age, etc.) for coarse-grained matching, without exploring the deep relationship between users' implicit behaviors (such as page dwell patterns, scrolling speed, etc.) and visual elements. This results in insufficient matching between advertising content and user interests, leading to low click-through rates.
[0045] 3. Brand security risks: The solutions provided by related technologies lack brand element protection mechanisms during the automated modification process, which can easily lead to problems such as deformation of the advertiser's logo and incomplete product display, thereby damaging the advertiser's brand image and reducing the advertiser's trust.
[0046] In view of this, embodiments of this application provide a video processing method to solve the above-mentioned technical problems through the following means:
[0047] 1. Intelligent material deconstruction technology: Based on computer vision (CV) segmentation networks (such as deep learning visual models), video frames are automatically divided into brand protection areas (i.e., areas that cannot be modified) and dynamic variable areas (i.e., areas that can be modified). It can also support advertisers to bind heat map annotations to policy rules. For example, the surface of a lunch box can be replaced with stickers, but the logo area cannot be modified.
[0048] 2. Cross-modal feature translation: Transform users' implicit behaviors (such as preferences for a certain type of music, local weather information, etc.) into visual parameters (such as color saturation offset, coordinates of dynamic element insertion, etc.), auditory parameters (such as the type of background music), and narrative logic (such as the weight of plot branches, scene changes, etc.) of advertising videos.
[0049] In other words, the embodiments of this application provide a video processing method, apparatus, electronic device, computer-readable storage medium, and computer program product, which, on the one hand, ensures the integrity of the target elements displayed in the prohibited modification area during the video modification process, and on the other hand, makes the modified video conform to the user's interests. The electronic device provided in the embodiments of this application will be described below. The electronic device provided in the embodiments of this application can be implemented as a terminal device.
[0050] For example, see Figure 1 , Figure 1 This is a schematic diagram of the architecture of the video processing system 100 provided in the embodiments of this application, as shown below. Figure 1As shown, the video processing system 100 provided in this application embodiment includes: a server 200, a network 300, and a terminal device 400. The network 300 can be a local area network or a wide area network, or a combination of the two. The terminal device 400 is a terminal device associated with a user (e.g., an advertiser). A client 410 runs on the terminal device 400. The client 410 can be various types of video editing software or a browser, etc.
[0051] In some embodiments, a video interface may be displayed on the client 410, wherein the video interface includes a video (e.g., an advertising video generated by the server 200 based on advertising materials uploaded by the advertiser), the video includes an editable area (e.g., an area other than the area where the advertiser's logo or product is located, including the area where the background environment is located in the video frame, the area where wall decorations are located, etc.) and a non-editable area (e.g., the area where the advertiser's logo or product is located), when the client 410 receives a video editing trigger operation triggered by the advertiser in the video interface (e.g., receiving a click operation by the advertiser on the video editing entry displayed in the video interface), When the video is modified, a video editing interface can be displayed (for example, the video editing interface can be displayed floating in the video interface, or the video interface can be jumped to the video editing interface). The video editing interface includes editing controls. Subsequently, the client 410 can respond to the advertiser's editing operation on the modifiable area based on the editing controls, and update the modifiable area according to the editing operation (for example, adding special effects in the modifiable area, adjusting the scene displayed in the modifiable area, adjusting the items displayed in the modifiable area, etc.). In this way, the integrity of the target elements (such as the advertiser's logo or product) in the prohibited modification area can be guaranteed during the modification of the video.
[0052] It should be noted that, Figure 1 The server 200 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms. The terminal device 400 can be a smartphone, tablet, laptop, desktop computer, smart speaker, smartwatch, in-vehicle terminal, etc., but is not limited to these. The terminal device 400 and the server 200 can be directly or indirectly connected via wired or wireless communication, which is not limited in this embodiment.
[0053] In some embodiments, the terminal device can also implement the video processing method provided in this application embodiment by running various computer-executable instructions or computer programs. For example, computer-executable instructions can be microprogram-level commands, machine instructions, or software instructions. Computer programs can be native programs or software modules in an operating system; they can be native applications (APPs), i.e., programs that need to be installed in the operating system to run, such as video editing APPs; or they can be applets that can be embedded in any APP, i.e., programs that only need to be downloaded to a browser environment to run. In summary, the aforementioned computer-executable instructions can be any form of instruction, and the aforementioned computer programs can be any form of application, module, or plugin.
[0054] The structure of the electronic device provided in the embodiments of this application will be further described below. Taking the electronic device as a terminal device as an example, see... Figure 2 , Figure 2 This is a schematic diagram of the structure of the electronic device 500 provided in the embodiments of this application. Figure 2 The illustrated electronic device 500 includes at least one processor 510, a memory 550, at least one network interface 520, and a user interface 530. The various components in the electronic device 500 are coupled together via a bus system 540. It is understood that the bus system 540 is used to implement communication between these components. In addition to a data bus, the bus system 540 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 2 The general labeled all buses as Bus System 540.
[0055] The processor 510 can be an integrated circuit chip with signal processing capabilities, such as a general-purpose processor, a digital signal processor (DSP), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0056] User interface 530 includes one or more output devices 531 that enable the presentation of media content, including one or more speakers and / or one or more visual displays. User interface 530 also includes one or more input devices 532, including user interface components that facilitate user input, such as a keyboard, mouse, microphone, touch screen display, camera, other input buttons and controls.
[0057] The memory 550 may be removable, non-removable, or a combination thereof. Exemplary hardware devices include solid-state storage, hard disk drives, optical disk drives, etc. The memory 550 may optionally include one or more storage devices physically located away from the processor 510.
[0058] The memory 550 may include volatile memory or non-volatile memory, or both. The non-volatile memory may be read-only memory (ROM), and the volatile memory may be random access memory (RAM). The memory 550 described in this application embodiment is intended to include any suitable type of memory.
[0059] In some embodiments, memory 550 is capable of storing data to support various operations, examples of which include programs, modules, and data structures or subsets or supersets thereof, as illustrated below.
[0060] Operating system 551 includes system programs for handling various basic system services and performing hardware-related tasks, such as the framework layer, core library layer, driver layer, etc., for implementing various basic business functions and handling hardware-based tasks;
[0061] The network communication module 552 is used to reach other computing devices via one or more (wired or wireless) network interfaces 520, exemplary network interfaces 520 including: Bluetooth, WiFi, and Universal Serial Bus (USB), etc.
[0062] Presentation module 553 is configured to enable the presentation of information (e.g., a user interface for operating peripheral devices and displaying content and information) via one or more output devices 531 (e.g., a display screen, a speaker, etc.) associated with user interface 530;
[0063] The input processing module 554 is used to detect and translate one or more user inputs or interactions from one or more input devices 532.
[0064] In some embodiments, the apparatus provided in this application can be implemented in software. Figure 2A video processing device 555 stored in memory 550 is shown. This device can be software in the form of programs and plugins, and includes the following software modules: a display module 5551, an update module 5552, a determination module 5553, an acquisition module 5554, a jump module 5555, a generation module 5556, and a recognition module 5557. These modules are logically connected and can therefore be arbitrarily combined or further separated according to the functions implemented. It should be noted that... Figure 2 For ease of explanation, all the above modules are shown at once, but this should not be construed as excluding the implementation of the video processing apparatus 555 which may only include the display module 5551 and the update module 5552. The functions of each module will be described below.
[0065] The video processing method provided in this application will be specifically described below with reference to the exemplary application and implementation of the terminal device provided in the embodiments of this application.
[0066] For example, see Figure 3 , Figure 3 This is a first flowchart illustrating the video processing method provided in this application embodiment, which will be combined with... Figure 3 The steps shown are explained.
[0067] It should be noted that, Figure 3 The method illustrated can be executed by various forms of computer programs running on the terminal device, and is not limited to a client. For example, it can also be the operating system, software module, script, and applet mentioned above. Therefore, the client-side examples used below should not be considered as limiting the embodiments of this application. Furthermore, for ease of description, no specific distinction will be made between the terminal device and the client running on the terminal device below.
[0068] In step 101, the video interface is displayed.
[0069] Here, the video interface may include a video (e.g., a video that is playing or a video that is paused), wherein the video may include modifiable areas (e.g., the area where the background environment is located) and non-modifiable areas (e.g., the area where the advertiser's logo or product is located).
[0070] In some embodiments, during execution Figure 3 Before step 101 shown, the following processing can be performed on each frame of the video: the image is identified by a pre-trained image recognition model; the region where the target element is located in the image is designated as a prohibited modification region, and the other regions in the image other than the region where the target element is located are designated as modifiable regions. The image recognition model can be trained based on sample images, which may include the target element and labels labeled for the target element.
[0071] For example, taking the Mask R-CNN model as an image recognition model, Mask R-CNN is a deep learning model for object detection and instance segmentation. It is developed based on the Faster R-CNN model. Mask R-CNN achieves instance segmentation by adding an extra branch to the Faster R-CNN model to predict a pixel-level mask for each detected object. The Mask R-CNN model mainly includes components such as a backbone network, a region proposal network, a ROIAlign layer, and a Mask branch. The backbone network can be a convolutional neural network (such as ResNet or ResNeXt) used to extract feature representations of the image; the region proposal network generates region proposals that may contain objects, which are then fed into classification and regression branches for further processing; the ROIAlign layer solves the feature alignment problem caused by quantization through bilinear interpolation, improving the accuracy of the segmentation mask; the Mask branch can be a fully convolutional network branch used to generate a binary mask for each detected object, representing the pixel-level segmentation of the object. For example, the Mask R-CNN model can be trained first using sample images containing the advertiser's product logo or packaging. This allows the trained Mask R-CNN model to identify the advertiser's product logo or packaging from the images. Then, the trained Mask R-CNN model can be used to analyze the advertiser's uploaded video frame by frame. For example, for each frame of the advertiser's uploaded video, the following processing can be performed: the trained Mask R-CNN model identifies the image. When the advertiser's product logo or packaging is identified, the area containing the advertiser's product logo or packaging is designated as a prohibited modification area, and the remaining areas in the image (i.e., areas other than the area containing the advertiser's product logo or packaging) are designated as modifiable areas. This allows for the rapid identification of the modifiable and prohibited modification areas in each frame.
[0072] It should be noted that if no target element is identified in an image, then all regions included in that image can be considered as modifiable regions. For example, if the 10th frame of an advertisement video uploaded by an advertiser is not identified in the 10th frame using the trained Mask R-CNN model, then all regions included in the 10th frame can be considered as modifiable regions.
[0073] In addition, it should be noted that the target element mentioned above can be the advertiser's logo or product packaging, or it can be a person or animal in the video footage. For example, when the video is a nature documentary, the target element can be the animal that the documentary mainly records. This application embodiment does not specifically limit this.
[0074] In other embodiments, after displaying the video interface, at least one of the following processes may be performed: marking the modifiable area with a transparent layer; marking the non-modifiable area with a mask; and displaying a prompt message in the video interface, wherein the prompt message may be used to indicate the modifiable area and the non-modifiable area.
[0075] For example, continuing from the above example, to facilitate advertisers in distinguishing between modifiable and non-modifiable areas in a video frame, after using the trained Mask R-CNN model to identify each frame of the advertiser's uploaded video, the area where the advertiser's product logo or packaging is located (i.e., the non-modifiable area) can be marked with a red mask; while the modifiable areas identified in the image (such as the background area) can be marked with a blue semi-transparent layer; of course, advertisers can also be informed of which areas in the video frame are non-modifiable and which are modifiable through text, for example, by displaying corresponding prompt text in the upper left corner of the video interface. This application embodiment does not specifically limit this.
[0076] In some embodiments, the video displayed in the video interface may be generated using a text-based video model, then during execution... Figure 3 Before step 101 shown, the following processes can also be performed: obtaining video footage and prompts; and generating a video using a text-based video model based on the video footage and prompts to obtain the video.
[0077] For example, taking the scenario of pushing advertising videos as an example, the technical solution provided in this application embodiment may include a complete video mode and a material package mode. That is to say, advertisers can not only upload complete videos, but also upload only video materials (such as product images, product copy, etc.) and prompts (such as storyboards), and generate corresponding videos based on video materials and prompts through the Wensheng video model. For example, the Wensheng video model can generate multiple corresponding videos based on the video materials and prompts uploaded by advertisers for advertisers to choose from.
[0078] It should be noted that for the material package mode, advertisers can also upload priority tags (such as "must-retain scenes" and "replaceable elements"). In this way, the Wensheng video model can combine the priority tags uploaded by the advertiser to generate the video, so that the final generated video is more in line with the advertiser's needs.
[0079] In step 102, in response to the video editing trigger operation, the editing controls are displayed.
[0080] In some embodiments, a video editing entry can also be displayed in the video interface described above. In this case, a trigger operation targeting the video editing entry can be identified as a video editing trigger operation. For example, when a user (e.g., an advertiser) clicks on the video editing entry displayed in the video interface, the video editing interface can be displayed floating in the video interface. The size of the video editing interface can be smaller than the size of the video interface. For example, the video editing interface can be displayed floating in the lower middle area of the video interface. Editing controls can be displayed in the video editing interface.
[0081] It should be noted that the video interface and the video editing interface can also be displayed independently in a split-screen manner. For example, when an advertiser clicks on the video editing entry displayed in the video interface (e.g., a full-screen video interface), the video interface can be switched from full-screen to non-full-screen display, and the video editing interface can be displayed in an area independent of the video interface. Of course, when an advertiser clicks on the video editing entry displayed in the video interface, one can also jump from the video interface to the video editing interface. This application embodiment does not specifically limit this.
[0082] In step 103, in response to an editing operation on the modifiable area based on the editing control, the modifiable area is updated according to the editing operation.
[0083] Here, when updating the editable area based on editing operations, a mask can be used to cover the uneditable area, thereby ensuring the integrity of the target element displayed in the uneditable area (such as the advertiser's logo or product packaging) and preventing the target element displayed in the uneditable area from being deformed during the video modification process.
[0084] In some embodiments, the above-mentioned editing control may include an effect addition control, and the editing operation may include an effect addition operation. Then, step 103 can be implemented in the following way: the triggering operation for the effect addition control is identified as an effect addition operation; in response to the effect addition operation, the object feature data of the target object is obtained, wherein the target object is the object to be pushed in the video; the effect matching the object feature data is queried from the effect library, and the matching effect is added to the modifiable area.
[0085] For example, taking user A as the target audience, when pushing an ad video to user A, upon receiving a click action from the advertiser to add controls for special effects, the system can obtain user A's real-time feature data (e.g., user A's search history for the past week). Then, it can query a pre-built effects library to find effects that match user A's real-time feature data. For instance, assuming user A has searched for "cherry blossoms" multiple times in the past week, the effect of cherry blossoms falling from the effects library can be used as the matching effect. Subsequently, the effect of cherry blossoms falling can be added to the editable area included in the video frame. In this way, while protecting the integrity of the advertiser's brand information (e.g., the advertiser's logo or product), it achieves precise video targeting tailored to each individual, saving advertisers the cost of video production and delivery, and effectively improving user conversion rates.
[0086] In other embodiments, following the above, the above-mentioned addition of matching effects to the modifiable area can be achieved in the following way: for at least one frame of the video, the following processing is performed: adding matching effects to the modifiable area included in the image, wherein the at least one frame of the video includes: multiple frames of the video, at least one frame of the multiple frames of the video that corresponds to the set time period, and at least one selected frame of the multiple frames of the video.
[0087] For example, continuing from the previous example and using an advertising video scenario, assuming the advertiser's uploaded video consists of 100 frames, a cherry blossom falling effect can be added to each of the editable areas within those 100 frames. Of course, the advertiser can also set the effect's activation time. For instance, the effect addition controls mentioned above can be displayed in the video editing interface, which also displays the playback timeline of the advertiser's uploaded video. The advertiser can set the effect's activation time based on the playback timeline displayed in the video editing interface. For example, assuming the advertiser sets... If the effective time is set to the 5th to the 10th second, then the cherry blossom falling effect can be added only to the modifiable area of the multiple frames (e.g., 10 frames) corresponding to the 5th to the 10th second in the above 100 frames; or, after clicking the effect addition control, the advertiser can also set the image frames to which the effect should be added. For example, if the advertiser sets the 10th to the 20th frame, then the cherry blossom falling effect can be added only to the modifiable area of the 10 images corresponding to the 10th to the 20th frame. In this way, the advertiser's needs for more refined modification of the advertising video can be met.
[0088] In some embodiments, following the above, the above-mentioned querying of effects from the effects library that match the object feature data can be achieved in the following way: for each effect in the effects library, determine the similarity between the effect feature data of the effect and the object feature data of the target object; and take the effects with a similarity greater than the similarity threshold as the effects that match the object feature data.
[0089] For example, taking user A as the target audience, assuming there are 100 effects pre-stored in the effects library, numbered from effect 1 to effect 100, after obtaining user A's real-time feature data, the similarity (e.g., cosine similarity) between the effect feature data of each effect in the effects library and user A's real-time feature data can be calculated. Then, effects with a similarity greater than a similarity threshold (e.g., 90%) can be considered as effects that match user A's real-time feature data (i.e., effects that highly match user A's interests). For example, assuming the similarity between effect 10 and user A's real-time feature data is greater than 90%, effect 10 can be considered as an effect that matches user A's real-time feature data, meaning effect 10 is an effect that highly matches user A's interests. Thus, by adding effect 10 to the editable area of the video, the advertising video can be made more in line with user A's interests, thereby improving user conversion rates.
[0090] In some embodiments, the above-mentioned editing control may further include a scene switching control, and the above-mentioned editing operation may further include a scene switching operation. Then, the above-mentioned step 103 may also be implemented in the following way: the trigger operation for the scene switching control is identified as a scene switching operation; in response to the scene switching operation, the regional feature data of the current location of the target object (e.g., the temperature of the user's current location) is obtained, wherein the target object is the object to be pushed by the video; the scene displayed in the modifiable area is switched to a scene that conforms to the regional feature data.
[0091] For example, taking user A as the target audience, when pushing an ad video to user A, upon receiving a click from the advertiser on the scene switching control (e.g., the "location" button) displayed in the video editing interface, the advertiser can obtain user A's current location (e.g., the city user A is currently in). This can be determined, for example, by the location uploaded by user A's mobile device (e.g., a mobile phone or tablet). When the temperature in user A's city is higher than a temperature threshold (e.g., 25℃), the scene displayed in the editable area of the video (e.g., a camping scene) can be switched to a beach vacation scene. This makes the modified ad video more closely match the user's current real-world environment, thus better matching the user's interests and effectively improving conversion rates.
[0092] It should be noted that the aforementioned regional feature data can be not only the temperature of the area where the target object is currently located, but also attractions, food, etc. in the area where the target object is currently located. For example, the scene displayed in the modifiable area can be switched to the attractions in the area where the target object is currently located, or the scene displayed in the modifiable area can be switched to a scene that includes the food in the area where the target object is currently located. This application embodiment does not specifically limit this.
[0093] In other embodiments, following the above, the above-mentioned switching of the scene displaying the modifiable area to the scene matching the regional feature data can be achieved in the following way: querying a pre-built mapping table based on the regional feature data of the current area of the target object to obtain the scene matching the regional feature data, wherein the mapping table may include the mapping relationship between different regional feature data and different scenes; and switching the scene displaying the modifiable area to the scene matching the regional feature data.
[0094] For example, taking user A as the target, after obtaining the temperature of user A's current location (e.g., 30℃), a pre-created mapping table can be queried based on the temperature of user A's current location to obtain the scene corresponding to 30℃ (e.g., a beach vacation scene). The mapping table can include the mapping relationship between different temperatures and different scenes. Then, the scene displayed in the modifiable area of the video frame (e.g., a camping scene on the grass scene) can be switched to the beach vacation scene. In this way, by pre-creating a mapping table that includes the mapping relationship between different temperatures and different scenes, after obtaining the temperature of the user's current location, the scene matching the temperature of the user's current location can be quickly obtained by querying the mapping table, improving the scene search efficiency.
[0095] In some embodiments, following the above, the above-described switching of the scene displaying the modifiable area to a scene conforming to the regional feature data can be achieved in the following manner: for at least one frame of the video, the following processing is performed: the scene displaying the modifiable area included in the image is switched to a scene conforming to the regional feature data, wherein the at least one frame of the image includes: multiple frames of the video, at least one frame of the multiple frames of the video that corresponds to the set time period, and at least one selected frame of the multiple frames of the video.
[0096] For example, taking the ad video push scenario as an example, assuming the ad video uploaded by the advertiser (e.g., a video of a camping scene) consists of 100 frames, after querying a pre-created mapping table based on the temperature (e.g., 30℃) of the target object's (e.g., user A's) current location, and obtaining the scene corresponding to 30℃ (e.g., a beach vacation scene), the camping scene displayed in the modifiable area of these 100 frames can be switched to the beach vacation scene. Of course, the advertiser can also set the time period for scene switching. For example, the scene switching control mentioned above (e.g., the "geographic location" button) can be displayed in the video editing interface. The video editing interface can also display the playback timeline of the ad video uploaded by the advertiser. The advertiser can set the time period for scene switching based on the playback timeline displayed in the video editing interface. For example, assuming the advertiser sets the time period to the 10th to the 20th second, only the camping scene displayed in the modifiable area of the images corresponding to the 10th to the 20th second of these 100 frames (e.g., the 10th to the 20th frames) can be switched to the beach vacation scene. Of course, after clicking the "Geographic Location" button displayed in the video editing interface, advertisers can also set the video frames for which the scene needs to be switched. For example, if the advertiser sets the 10th to 20th frames, the scene displayed in the modifiable area included in the 100 frames can be switched to a beach vacation scene. This further refines the granularity of the modification, thereby meeting the advertiser's more refined modification needs for the advertising video.
[0097] In some embodiments, the above-mentioned editing control may further include a scene switching control, and the above-mentioned editing operation may further include a scene switching operation. Then, the above-mentioned step 103 may be implemented in the following way: the trigger operation for the scene switching control is identified as a scene switching operation; in response to the scene switching operation, the object feature data of the target object is obtained, wherein the target object is the object to be pushed in the video; and the scene displayed in the modifiable area is switched to the scene that conforms to the object feature data.
[0098] For example, taking user A as the target audience, when an ad video needs to be pushed to user A, upon receiving a click operation from the advertiser on the scene switching control displayed in the video editing interface, real-time feature data of user A can be obtained (e.g., user A's feature data within the past month). Then, the scene displayed in the modifiable area of the video screen (e.g., a camping scene) can be switched to a scene that matches user A's real-time feature data. For example, assuming that user A is a middle-aged father with children based on real-time feature data, the camping scene displayed in the modifiable area can be switched to an amusement park scene suitable for parent-child activities. At the same time, the couple displayed in the camping scene can be switched to a father and son playing in the amusement park. In this way, the modified ad video can better match user A's interests, achieving precise video delivery with "personalized recommendations," saving advertisers the cost of video production and delivery, and effectively improving user conversion rates.
[0099] In other embodiments, following the above, the above-mentioned switching of the scene displaying the modifiable area to the scene matching the object feature data can be achieved in the following way: for each scene in the scene library, determine the similarity between the scene feature data and the object feature data of the scene; take the scene with a similarity greater than the similarity threshold as the scene matching the object feature data; and switch the scene displaying the modifiable area to the scene matching the object feature data.
[0100] For example, taking user A as the target, assuming there are 100 scenes pre-stored in the scene library, numbered from scene 1 to scene 100, after obtaining user A's real-time feature data, the similarity (e.g., cosine similarity) between the scene feature data of the 100 scenes in the scene library and user A's real-time feature data can be calculated. Then, scenes with a similarity greater than a similarity threshold (e.g., 90%) (e.g., scene 3) can be used as the scenes that match user A's real-time feature data. Subsequently, the scene displayed in the modifiable area can be switched to scene 3. Since the similarity between scene 3 and user A's real-time feature data is greater than the similarity threshold, scene 3 is a scene that user A may be interested in. Thus, switching the scene displayed in the modifiable area of the video to scene 3 can better match user A's interests, thereby achieving accurate video push.
[0101] In some embodiments, the above-mentioned editing control may further include an item adjustment control, and the above-mentioned editing operation may further include an item adjustment operation. Then, step 103 can be implemented as follows: the triggering operation on the item adjustment control is identified as an item adjustment operation; in response to the item adjustment operation, object feature data of the target object is obtained, wherein the target object is the object to be pushed in the video; the items displayed in the modifiable area are adjusted to items that conform to the object feature data. For example, after querying a matching item from the item library based on the object feature data of the target object, a priority tag can be assigned to the queried matching item. Thus, when updating the modifiable area of the video subsequently, the corresponding matching item can be obtained based on the priority tag, and the items displayed in the modifiable area can be adjusted to the matching item.
[0102] For example, taking user A as the target audience, when an ad video needs to be pushed to user A, upon receiving a click operation from the advertiser regarding the adjustment controls for items displayed in the video editing interface, real-time characteristic data of user A can be obtained (e.g., user A's characteristic data for the past month). Then, the items displayed in the editable area of the video screen (e.g., a lunchbox) can be adjusted to items that match user A's real-time characteristic data. For example, assuming that user A is determined to be a music lover based on real-time characteristic data, the lunchbox displayed in the editable area can be changed to a vinyl record. In this way, by adjusting the items displayed in the editable area to items that match the user's interests, the modified ad video can be more in line with the user's interests, thereby effectively improving the user's conversion rate.
[0103] It should be noted that when the modifiable area includes multiple items, all items can be adjusted. For example, multiple items can be adjusted to match the object characteristic data. Alternatively, after clicking the item adjustment control, the advertiser can also set which items need to be adjusted. For example, if the advertiser selects item 1 among multiple items, only item 1 can be adjusted to match the object characteristic data, while the other items remain unchanged.
[0104] In other embodiments, when the number of items included in the modifiable region is multiple, this application embodiment can also generate an attention heatmap corresponding to the modifiable region based on an attention mechanism. Then, items with heat values greater than a heat value threshold in the attention heatmap (i.e., items in the modifiable region that the user is most likely to pay attention to) can be adjusted to items that match the user's interests. For example, a U-Net can be used to generate an attention heatmap corresponding to the modifiable region. The specific process is as follows: First, the U-Net needs to be modified (e.g., adding an extra layer to the decoder or modifying existing layers) so that it can output an attention heatmap. Then, the U-Net is trained using a dataset with attention annotations (e.g., the original image and the corresponding attention heatmap), so that the trained U-Net can learn how to generate an attention heatmap from the original image. After training, the modifiable region can be input into the trained U-shaped network to generate a corresponding attention heatmap. Finally, the generated attention heatmap can be overlaid on the modifiable region to visually display the heat value of each item in the modifiable region. In this way, items with heat values greater than the heat value threshold (i.e., items that users are most likely to pay attention to) can be adjusted to match the user's interests.
[0105] In other embodiments, following the above, the above-mentioned adjustment of the items displayed in the modifiable area to items that match the object feature data can be achieved in the following way: for each item in the item library, determine the similarity between the item feature data and the object feature data; take items with a similarity greater than a similarity threshold as items that match the object feature data; and adjust the items displayed in the modifiable area to items that match the object feature data.
[0106] For example, taking user A as the target, assuming there are 500 items pre-stored in the item library, numbered from item 1 to item 500, after obtaining user A's real-time feature data (e.g., user A's feature data for the most recent month), the similarity between the feature data of these 500 items in the item library and user A's real-time feature data can be calculated (e.g., cosine similarity, Euclidean distance, or Manhattan distance). Then, items with a similarity greater than a similarity threshold (e.g., 90%) (e.g., item 80) can be considered as items matching user A's real-time feature data, and the items displayed in the editable area can be adjusted to item 80 (e.g., vinyl records). In this way, since item 80 is an item that user A might be interested in based on similarity calculations, by adjusting the items displayed in the editable area to item 80, the modified video can be made more in line with user A's interests.
[0107] It should be noted that, in addition to adjusting the items displayed in the modifiable area to conform to the object characteristic data, the style of the items displayed in the modifiable area can also be adjusted to conform to the object characteristic data. In other words, the embodiments of this application can not only adjust the items, but also only adjust the style of the items. The embodiments of this application do not make specific limitations in this regard.
[0108] In some embodiments, the modifiable region described above may include multiple modifiable sub-regions, see [link to relevant documentation]. Figure 4 , Figure 4 This is a schematic diagram of the second process of the video processing method provided in the embodiments of this application, as shown below. Figure 4 As shown, Figure 3 Step 103 shown can be achieved through Figure 4 The implementation of steps 1031 and 1032 shown will be combined with Figure 4 The steps shown are explained.
[0109] In step 1031, in response to a selection operation for multiple modifiable sub-regions, the selected modifiable sub-region is taken as the target modifiable sub-region.
[0110] In some embodiments, the above-mentioned editing control may be displayed in a video editing interface. The video editing interface may also include multiple sub-area controls that correspond one-to-one with multiple modifiable sub-areas. Step 1031 can be implemented in the following way: In response to the triggering operation for multiple sub-area controls, the triggered sub-area control is highlighted (for example, the sub-area control clicked by the advertiser can be highlighted to indicate that the sub-area control is currently selected), and the modifiable sub-area corresponding to the triggered sub-area control is taken as the target modifiable sub-area.
[0111] For example, taking three editable sub-regions as an example, let's say they are editable sub-region 1 (e.g., the area containing the background sky), editable sub-region 2 (e.g., the area containing wall decorations), and editable sub-region 3 (e.g., the area containing desktop decorations). Then, in the video editing interface, three corresponding sub-region controls can be displayed: sub-region control 1, sub-region control 2, and sub-region control 3. Sub-region control 1 can display the text "Background Sky" to indicate that it corresponds to editable sub-region 1; sub-region control 2 can display the text "Wall" to indicate that it corresponds to editable sub-region 1. Area 2; The text "Desktop" can be displayed on sub-area control 3 to indicate that sub-area control 3 corresponds to editable sub-area 3. Advertisers can select these 3 sub-area controls in the video editing interface. For example, if an advertiser clicks on sub-area control 2, sub-area control 2 can be highlighted to indicate that it is currently selected. Subsequently, when an edit operation is triggered by the advertiser, the editable sub-area 2 corresponding to sub-area control 2 can be updated according to the edit operation. That is, only the area where the wall decorations are located in the video frame can be updated. In this way, more granular modification is provided, allowing advertisers to modify the advertising video more precisely.
[0112] In step 1032, in response to an editing operation on the target modifiable sub-region based on the editing control, the target modifiable sub-region is updated according to the editing operation.
[0113] In some embodiments, following the above example, taking the target modifiable sub-area (e.g., the modifiable sub-area corresponding to the sub-area control clicked by the advertiser) as modifiable sub-area 2 as an example, when receiving the edit operation triggered by the advertiser, only modifiable sub-area 2 can be updated according to the edit operation. For example, only effects that match the user's interests can be added to modifiable sub-area 2, or only the items displayed in modifiable sub-area 2 can be adjusted to items that match the user's interests.
[0114] In some embodiments, the aforementioned editing controls may be displayed in a video editing interface. The video editing interface may also include a video playback timeline. After displaying the video editing interface in response to a video editing trigger operation, the following processing may also be performed: In response to a time period setting operation based on the playback timeline, the set time period is obtained. At this time, step 103 above may be implemented in the following way: In response to an editing operation based on the editing controls for a modifiable area, the modifiable area in the multi-frame image corresponding to the target time period of the video is updated according to the editing operation. The target time period is the time period other than the set time period in the total playback duration of the video.
[0115] For example, taking the scenario of pushing advertising videos as an example, in addition to displaying editing controls, the video editing interface can also display the playback timeline corresponding to the advertising video uploaded by the advertiser. The advertiser can set a protected time period (i.e., the time period in the advertising video that is prohibited from being modified) based on the playback timeline displayed in the video editing interface. For example, assuming that the total duration of the advertising video uploaded by the advertiser is 30 seconds, and assuming that the protected time period set by the advertiser is from the 10th second to the 15th second (for example, assuming that it is the time period corresponding to the product close-up shot), then when the advertiser's editing operation on the editable area is received based on the editing controls, only the editable area included in the multi-frame images corresponding to the 1st second to the 9th second and the 16th second to the 30th second of the advertising video uploaded by the advertiser can be updated. In this way, it can avoid inserting other content in the product close-up shot, thereby affecting the product's push effect.
[0116] In some embodiments, the above-mentioned editing controls may be displayed in a video editing interface. The video editing interface may also include style adjustment controls (e.g., a "Screen Style" button). After displaying the video editing interface in response to a video editing trigger operation, the following processing may also be performed: in response to a trigger operation on the style adjustment control, obtain the object feature data of the target object, wherein the target object is the object to be pushed in the video; adjust the style of the video to match the style of the object feature data.
[0117] For example, taking user A as the target audience, when an ad video needs to be pushed to user A, upon receiving a click operation from the advertiser on the style adjustment control (such as the "Style" button) displayed in the video editing interface, real-time feature data of user A can be obtained (for example, feature data of user A over the past month). Then, the style of the ad video uploaded by the advertiser can be adjusted to match the style of user A's real-time feature data. For example, assuming that user A is determined to be an artistic youth based on real-time feature data, the style of the ad video can be adjusted to match the style of an artistic youth. For example, corresponding style filters can be superimposed on the ad video uploaded by the advertiser, so that the modified ad video can better match user A's interests, achieving precise video delivery with "personalized recommendations," saving advertisers the cost of video production and delivery, and effectively improving user conversion rates.
[0118] In other embodiments, following the above, the above-mentioned adjustment of the video style to match the object feature data can be achieved in the following way: for each style in the style library, determine the similarity between the style feature data of the style and the object feature data of the target object; take the style with a similarity greater than the similarity threshold as the style that matches the object feature data of the target object; adjust the style of the video to match the style that matches the object feature data of the target object.
[0119] For example, continuing with user A as the target audience, suppose there are 100 styles pre-stored in the style library, numbered from style 1 to style 100. After obtaining user A's feature data for the most recent month, the similarity between the style feature data of each of the 100 styles in the style library and user A's feature data for the most recent month can be calculated. Then, styles with a similarity greater than a similarity threshold (e.g., 90%) (e.g., style 10) can be used as the style matching user A. Subsequently, the style of the ad video uploaded by the advertiser can be adjusted to style 10. Since style 10 is calculated based on user A's feature data for the most recent month, it matches user A's interests. Therefore, by adjusting the style of the ad video uploaded by the advertiser to style 10, the style-adjusted ad video can better match user A's interests, thereby effectively improving the user's conversion rate.
[0120] In some embodiments, the aforementioned editing controls may be displayed in a video editing interface, which may further include a preview control (e.g., a "Preview" button). After execution... Figure 3 Following step 103, the following processing can also be performed: in response to a trigger operation on the preview control, jump from the video editing interface to the preview interface, wherein the preview interface may include multiple object templates (e.g., user templates), and different object templates correspond to different object feature data; in response to a selection operation on multiple object templates, generate a target video that matches the selected object template based on the video after the modifiable area is updated.
[0121] For example, taking the ad video push scenario as an example, a preview button can be displayed in the video editing interface. When the advertiser clicks the preview button, the video editing interface can be redirected to the preview interface. The preview interface can display multiple user templates (such as "XX music lover", "white-collar worker who often works overtime", "travel lover", etc.) for the advertiser to choose from. Assuming the advertiser selects the "travel lover" user template in the preview interface, a target ad video matching "travel lover" can be generated based on the updated ad video in the editable area (for example, an ad video with cherry blossom falling effects added to the editable area). For example, the scene displayed in the editable area of the ad video uploaded by the advertiser can be adjusted to a natural scene (such as a scene of a mountain). At the same time, the background music of the ad video can be adjusted to music that matches "travel lover", thus obtaining the final target ad video. In this way, the advertiser can preview the final target ad video in the preview interface. If the advertiser is not satisfied with the final generated target ad video, they can return to the video editing interface to modify the video again.
[0122] In some embodiments, during execution Figure 3 Before step 102 shown, the following processing can also be performed: In response to the range adjustment operation for the modifiable area, the adjusted modifiable area is obtained, and step 103 above can also be achieved in the following way: In response to the editing operation based on the editing control for the adjusted modifiable area, the adjusted modifiable area is updated according to the editing operation.
[0123] For example, taking an ad video push scenario, after an advertiser uploads an ad video, the trained Mask R-CNN model can identify the uploaded ad video frame by frame. The identified modifiable areas can be marked with a blue semi-transparent layer. If the advertiser is not satisfied with the range of the modifiable areas output by the trained Mask R-CNN model, they can manually adjust it. For example, the advertiser can adjust the range of the modifiable areas by clicking or dragging a selection box to obtain the adjusted modifiable areas. Thus, when an edit operation triggered by the advertiser based on the edit controls is received subsequently, the adjusted modifiable areas can be updated according to the edit operation. In other words, this embodiment of the application also supports advertisers manually adjusting the range of modifiable areas, further enhancing the dimensionality of the advertiser's modifications to the ad video.
[0124] In some embodiments, the editing controls described above may be displayed in a video editing interface, which may further include audio adjustment controls (e.g., a "background music" button), see [link to relevant documentation]. Figure 5 , Figure 5This is a schematic diagram of the third process of the video processing method provided in the embodiments of this application, as shown below. Figure 5 As shown, after execution Figure 3 After step 102 shown, you can also execute... Figure 5 Steps 104 to 106 shown will combine Figure 5 The steps shown are explained.
[0125] In step 104, in response to a trigger operation on the audio adjustment control, object feature data of the target object is obtained.
[0126] Here, the target object is the object to which the video is to be pushed.
[0127] In some embodiments, taking user A as an example, when it is necessary to push an advertising video to user A, when the advertiser clicks the "background music" button displayed in the video editing interface, the real-time feature data of user A can be obtained (for example, the feature data of user A in the past month).
[0128] In step 105, audio matching the object feature data is queried from the audio library.
[0129] In some embodiments, step 105 can be implemented as follows: for each audio in the audio library, perform the following processing: determine the similarity between the audio feature data of the audio and the object feature data of the target object; and use audio with a similarity greater than a similarity threshold as the audio that matches the object feature data of the target object.
[0130] For example, continuing from the previous example, let's take user A as the target object. Suppose that the audio library has 100 audio files stored in advance, numbered from audio 1 to audio 100. After obtaining user A's real-time feature data, we can calculate the similarity between the audio feature data of these 100 audio files in the audio library and user A's real-time feature data. Then, the audio files with a similarity greater than the similarity threshold (e.g., 90%) (e.g., audio 10) can be used as the audio files that match user A's real-time feature data. In other words, audio 10 is the audio file that user A might be interested in. For example, if we assume that user A is a jazz enthusiast based on the real-time feature data, then audio 10, which belongs to jazz, can be used as the audio file that matches user A.
[0131] In step 106, the audio associated with the video is adjusted to match the audio of the object feature data.
[0132] In some embodiments, following the above, after obtaining the audio (i.e., audio 10) that matches the real-time feature data of user A, the original audio associated with the advertiser's uploaded ad video can be adjusted to audio 10. This makes the modified ad video more aligned with user A's interests. Furthermore, it should be noted that if the advertiser's uploaded ad video originally did not have any associated audio, audio 10 can be directly added to the advertiser's uploaded ad video.
[0133] The following example, using the scenario of pushing advertising videos, illustrates an exemplary application of the embodiments of this application in a real-world application scenario.
[0134] This application provides a video processing method that automatically identifies brand protection zones (i.e., areas where modification is prohibited, such as the advertiser's logo or product location) and dynamic variable zones (i.e., areas where modification is allowed, such as the background sky, wall decorations, or tabletop items) in the video frame. This allows for dynamic modification and optimization of the video uploaded by the advertiser while ensuring the integrity of brand information. Furthermore, this application can also incorporate real-time user characteristic data to adjust the content, style, and even background music effects. In addition, advertisers can control the degree and direction of modification using preset or custom rules, making the final generated advertising video more aligned with user interests and thus effectively improving advertising conversion rates.
[0135] The video processing method provided in the embodiments of this application will be described in detail below.
[0136] In some embodiments, advertisers can upload videos (i.e., full video mode) or related creative materials (i.e., creative package mode). For full video mode, after the advertiser uploads the original video (e.g., a 30-second camping scene video), the uploaded original video can be analyzed frame by frame based on a deep learning visual model (e.g., an improved Mask R-CNN model) to obtain the brand protection zone and dynamically variable zones. For example, as... Figure 6A As shown, for region 601 (i.e., brand protection zone) where fixed elements (such as advertiser's logo, product packaging, etc.) identified from the video frame based on a deep learning visual model are located, a red mask can be used to mark it to indicate that the region is unmodifiable (or prohibited from modification); or, as... Figure 6B As shown, the dynamically variable region 602 identified from the video frame based on the deep learning visual model (e.g., the area where the background sky, wall decorations, tabletop items, etc. are located in the video frame) can be marked in blue to indicate that the region is a modifiable region.
[0137] Furthermore, for the material package mode, this application embodiment supports uploading graphic and textual materials (such as product images, product copy, artistic fonts, etc.), video clips (such as material clips from different scenes), and priority tags (such as "must-retain scenes," "replaceable elements," etc.). Advertisers can also input storyboard scripts as prompts to initially generate video content. For example, product-related information can be extracted using a large language model, and advertisers can supplement or modify the automatically extracted information. The system can provide advertisers with multiple generated videos to choose from. After selecting one video, similar to the full video mode, the selected video can be analyzed frame-by-frame using a deep learning visual model to obtain the brand protection zone and the dynamically variable zone, which will not be elaborated further in this application embodiment.
[0138] In other embodiments, a video editing interface can also be provided to advertisers, allowing them to select different strategies for display. For example, the system can pre-set different rule strategy templates, from which advertisers can choose to combine or modify based on the templates. The types of rule strategies can include: 1) Spatial strategy: modifying the range of a dynamically variable area through clicks or drag-and-drop selection (e.g., only allowing modification of the lunchbox sticker); 2) Time strategy: setting the display period for core elements (e.g., prohibiting the insertion of other content in product close-up shots); 3) Logical strategy: binding trigger conditions based on the user's real-time characteristics, where the trigger conditions are shown in Table 1.
[0139] Table 1. Triggering Condition Diagram
[0140]
[0141] In some embodiments, following the above, after setting the rules, advertisers can also use editing tools to perform dynamic simulation previews and strategy adjustments to ensure the quality of the final generated advertising video. For example, embodiments of this application can also provide the following auxiliary functions: 1) Multi-profile simulator: It can provide multiple user templates (e.g., including "XXX music lovers", "family users", etc.) for advertisers to choose from. After selecting any user template, advertisers can generate corresponding variant advertising videos in real time; 2) Brand security detection: Threshold alarms can be set, for example, when the integrity of the logo is <95%, rendering can be forcibly stopped, and modification suggestions can be provided, such as reducing the range of the modifiable area.
[0142] For example, see Figure 7 , Figure 7 This is a schematic diagram illustrating an application scenario of the video processing method provided in the embodiments of this application, such as... Figure 7As shown, the video editing interface 701 displays the original video uploaded by the advertiser 702, a "background environment" control 703, a "desktop decoration" control 704, a "style filter" control 705, an "enable" control for "product close-up shots to prevent the insertion of other content" 706, a "screen style" control 707, a "background music" control 708, a "geographic location" control 709, a "confirm settings" control 710, and a "preview" control 711. For example, when an advertiser clicks on the "Background Environment" control 703 displayed in the video editing interface 701, only the background environment of the original video 702 uploaded by the advertiser can be modified; when an advertiser clicks on the "Desktop Decoration" control 704 displayed in the video editing interface 701, only the desktop decorations of the original video 702 uploaded by the advertiser can be modified; when an advertiser clicks on the "Style Filter" control 705 displayed in the video editing interface 701, the style filter of the original video 702 uploaded by the advertiser can be modified; when the advertiser checks the "Prevent Insertion of Other Content for Product Close-up" enable control 706 displayed in the video editing interface 701 (i.e., the "Prevent Insertion of Other Content for Product Close-up" enable control 706 is selected), other content can be prevented from being inserted into the video clips showing product close-up shots included in the original video 702 uploaded by the advertiser (e.g., the video clip from the 10th to the 18th second).
[0143] See also Figure 7When the system receives a click from the advertiser on the "Screen Style" 707 displayed in the video editing interface 701, it can adjust the screen style of the original video 702 uploaded by the advertiser to a style that matches the user's preferences. When the system receives a click from the advertiser on the "Background Music" control 708 displayed in the video editing interface 701, it can adjust the background music of the original video 702 uploaded by the advertiser to music that matches the user's preferences (for example, if the user likes jazz, the background music of the original video 702 uploaded by the advertiser can be switched to jazz). When the system receives a click from the advertiser on the "Geographic Location" control 709 displayed in the video editing interface 701, it can obtain the temperature of the user's current location and switch the background of the original video 702 uploaded by the advertiser to a background that matches the temperature of the user's current location. For example, if the temperature of the user's current location is >25℃, the background of the original video 702 uploaded by the advertiser can be switched from a camping scene to a beach scene. After completing the settings, advertisers can click the "Confirm Settings" button 710. Subsequently, advertisers can also preview the video by clicking the "Preview" button 711 displayed in the video editing interface 701. For example, when an advertiser clicks the "Preview" button 711 displayed in the video editing interface 701, the system will jump from the video editing interface 701 to the preview interface 712. The preview interface 712 can display multiple user templates (such as "XX Music Lover", "White-collar Worker Who Often Works Overtime", "Travel Lover", etc.) for advertisers to choose from. When an advertiser selects a user template, a corresponding advertisement video can be generated in real time for the advertiser to preview.
[0144] In some embodiments, this application can also fuse implicit user characteristics (such as user location, page scrolling speed, etc.) and explicit behavioral data (such as user search history, click hotspots, etc.) as strategy inputs to control the effect of the final advertising video presented to the user. That is, the effect of the final advertising video will be different for users with different characteristics and behaviors. For example, taking user A as an example, assuming that user A's search history (i.e. explicit behavioral data) includes a jazz playlist, it can be determined that user A is a jazz enthusiast. At this time, the lunchbox displayed in the video frame of the original video uploaded by the advertiser can be adjusted to the shape of a retro vinyl record box, and the background music of the original video can also be switched to jazz. Similarly, taking user B as an example, assuming that user B is a middle-aged father with a child, the scene displayed in the video frame of the original video uploaded by the advertiser can be switched from a camping scene to an amusement park scene, the character interaction can be changed to parents and children, and the background music of the original video can also be switched to piano music that matches the parent-child relationship.
[0145] In other words, the technical solution provided in this application reconstructs the production chain of advertising content, while protecting the integrity of brand information, achieving precise "personalized" advertising video delivery, saving advertisers the cost of material production and delivery, and effectively improving user conversion rates.
[0146] The following will continue to combine Figure 8 The video processing method provided in the embodiments of this application will be described.
[0147] For example, see Figure 8 , Figure 8 This is a schematic diagram of the architecture of the video processing system provided in the embodiments of this application, such as... Figure 8 As shown, the video processing system provided in this application embodiment can adopt a modular layered architecture, wherein the core components include a material parsing engine, a feature hub, a rule decision engine, and a video synthesis module. First, advertisers can upload materials (including videos / material packages) to the material parsing engine through the advertiser console. The material parsing engine can analyze the uploaded video frame by frame based on a deep learning visual model (such as an improved Mask R-CNN model, which needs to be fine-tuned to learn from the product logo, product packaging, etc. uploaded by the advertiser), and output pixel-level unmodifiable areas (i.e., brand protection areas) and modifiable areas (i.e., dynamic variable areas). For the dynamic variable areas, the rule decision engine can combine visual saliency analysis and semantic correlation calculation. For example, it can first use a U-Net to generate a screen attention heatmap, then use a graph database (such as Neo4j) to calculate the substitutability score of items in the modifiable areas (e.g., calculate the correlation between a lunchbox and a vinyl record), and finally output suggested editable areas (e.g., labeled with a blue semi-transparent layer) and priority labels for items with high scores.
[0148] See also Figure 8 The feature hub can collect explicit user behavior data (such as search keywords and click records) and implicit feature data (such as the user's current location and page scrolling speed) in real time, generating structured vectors (i.e., the user's real-time feature vectors). The rule decision engine can parse the spatial, temporal, and logical strategies shown in Table 2 based on the user's real-time feature vectors output by the feature hub and the metadata of the dynamically variable areas output by the content parsing engine, and generate operation instructions for the modifiable areas. The rule decision engine can send the generated rendering instruction set to the video compositing module, enabling the video compositing module to adjust the original video uploaded by the advertiser based on the rendering instruction set sent by the rule decision engine, and store the final synthesized video in the Content Delivery Network (CDN).
[0149] Table 2. Schematic diagram of the execution logic of the rule decision engine
[0150]
[0151] For example, for the spatial strategy mentioned above, advertisers can manually select the editable area (for example, assuming the coordinate range is: x∈[120, 300], y∈[80, 200]), which will generate the corresponding constraint instructions (for example, constraint instructions in JSON format). Then ControlNet can generate the corresponding mask based on the constraint instructions and limit the redrawing area based on the mask.
[0152] For example, for the above logical strategy, the rule definition can be: if the temperature at the user's location is >25℃, the background is switched to a beach scene. The rule decision engine can trigger the corresponding rendering instruction and send it to the video compositing mode. The video compositing module can call the pre-stored beach scene for replacement and compositing according to the rendering instruction generated by the rule decision engine.
[0153] The following will continue to combine Figure 9 The process of video synthesis is explained.
[0154] For example, see Figure 9 , Figure 9 This is a schematic diagram of the video processing method provided in the embodiments of this application, such as... Figure 9 As shown, after an advertiser uploads the original video to the content analysis engine through the advertiser console, the engine can identify the brand protection zone and dynamic variable zones in the video frame. Simultaneously, the feature hub can collect and analyze user feature data in real time, sending the resulting real-time user feature vector to the rule decision engine. Upon receiving the real-time user feature vector from the feature hub and the metadata for the brand protection zone and dynamic variable zones from the content analysis engine, the rule decision engine can parse the spatial, temporal, and logical strategies, generating rendering instructions for the dynamic variable zones. Subsequently, the rule decision engine can send the generated rendering instructions to the video compositing module, enabling the module to modify the dynamic variable zones frame-by-frame according to the rendering instructions, while maintaining the brand protection zone unchanged during the automated modification process.
[0155] In summary, the video processing method provided in this application has the following beneficial effects:
[0156] 1. Balancing Brand Security and Flexible Creative: Improved CV segmentation models (e.g., Mask R-CNN) lock in pixel-level brand information (e.g., advertiser's logo, product packaging), ensuring the integrity and consistency of brand information during automated modification. This mitigates the risk of accidental brand information distortion inherent in related technologies. Furthermore, based on visual saliency analysis and semantic association calculation, high-value, dynamically variable areas (e.g., background elements, prop textures) are intelligently recommended. Advertisers can flexibly configure modification strategies through visualization tools, significantly improving the reusability of creative materials.
[0157] 2. Precise Targeting Based on Real-Time User Characteristics: Explicit user behaviors (e.g., search history, click hotspots) and implicit characteristics (e.g., user's geographical location, page scrolling speed) are transformed into visual parameters (e.g., color shifts, element placement coordinates), auditory parameters (e.g., background music matching), and narrative logic (e.g., scene transition weighting) to achieve atomic-level matching between ad content and user interests. Furthermore, through flexible combinations of spatial, temporal, and logical strategies (e.g., enabling movie filters when the user's artistic index is >0.8), highly customized ad content can be generated while ensuring brand safety, thereby effectively improving conversion rates.
[0158] 3. High-efficiency production chain reconstruction: After advertisers upload the original video, the division of the brand protection area and the dynamic variable area can be completed automatically, reducing the cost of manual labeling. In addition, by providing visual rule configuration tools (such as drag-and-drop selection of modifiable areas, setting timeline protection segments, etc.), it supports non-technical personnel to quickly complete the deployment of complex strategies, lowering the operation threshold.
[0159] The following description continues to illustrate the exemplary structure of the video processing apparatus 555 provided in the embodiments of this application as a software module. In some embodiments, such as Figure 2 As shown, the software modules stored in the video processing device 555 in the memory 550 may include: a display module 5551 and an update module 5552.
[0160] Display module 5551 is used to display a video interface, wherein the video interface includes a video, and the video includes an editable area and an uneditable area; display module 5551 is also used to display editing controls in response to video editing trigger operations; update module 5552 is used to update the editable area according to the editing operation based on the editing controls for editing operations on the editable area.
[0161] In some embodiments, the editing control includes an effect addition control, and the editing operation includes an effect addition operation; the update module 5552 is further configured to identify the triggering operation for the effect addition control as an effect addition operation; in response to the effect addition operation, obtain object feature data of the target object, wherein the target object is the object to be pushed in the video; query the effect library for an effect that matches the object feature data, and add the matching effect in the modifiable area.
[0162] In some embodiments, the update module 5552 is further configured to perform the following processing on at least one frame of the video: adding a matching effect to a modifiable area included in the image, wherein the at least one frame of the video includes: multiple frames of the video, at least one frame of the multiple frames of the video corresponding to a set time period, and at least one selected frame of the multiple frames of the video.
[0163] In some embodiments, the update module 5552 is further configured to, for each effect in the effects library, determine the similarity between the effect feature data of the effect and the object feature data; and to use effects with a similarity greater than a similarity threshold as effects that match the object feature data.
[0164] In some embodiments, the editing control includes a scene switching control, and the editing operation includes a scene switching operation; the update module 5552 is further configured to identify the triggering operation for the scene switching control as a scene switching operation; in response to the scene switching operation, obtain the regional feature data of the area where the target object is currently located, wherein the target object is the object to be pushed in the video; and switch the scene displayed in the modifiable area to a scene that conforms to the regional feature data.
[0165] In some embodiments, the update module 5552 is further configured to query a pre-built mapping table based on the regional feature data to obtain a scene that matches the regional feature data, wherein the mapping table includes the mapping relationship between different regional feature data and different scenes; and switch the scene displayed in the modifiable region to the scene that matches the regional feature data.
[0166] In some embodiments, the update module 5552 is further configured to perform the following processing on at least one frame of image included in the video: switch the scene displayed in the modifiable area included in the image to a scene that conforms to the area feature data, wherein the at least one frame of image includes: multiple frames of image included in the video, at least one frame of image included in the multiple frames of image included in the video that corresponds to the set time period, and at least one selected frame of image included in the multiple frames of image included in the video.
[0167] In some embodiments, the editing control includes a scene switching control, and the editing operation includes a scene switching operation; the display module 5552 is further configured to identify the triggering operation for the scene switching control as a scene switching operation; in response to the scene switching operation, obtain object feature data of the target object, wherein the target object is the object to be pushed in the video; and switch the scene displayed in the modifiable area to a scene that conforms to the object feature data.
[0168] In some embodiments, the update module 5552 is further configured to, for each scene in the scene library, determine the similarity between scene feature data and object feature data; identify scenes with a similarity greater than a similarity threshold as scenes that match the object feature data; and switch the scenes displayed in the modifiable area to scenes that match the object feature data.
[0169] In some embodiments, the editing control includes an item adjustment control, and the editing operation includes an item adjustment operation; the update module 5552 is further configured to identify the triggering operation for the item adjustment control as an item adjustment operation; in response to the item adjustment operation, obtain the object feature data of the target object, wherein the target object is the object to be pushed in the video; and adjust the items displayed in the modifiable area to items that conform to the object feature data.
[0170] In some embodiments, the update module 5552 is further configured to, for each item in the item library, determine the similarity between the item feature data and the object feature data; identify items with a similarity greater than a similarity threshold as items that match the object feature data; and adjust the items displayed in the modifiable area to match the items.
[0171] In some embodiments, the modifiable region includes a plurality of modifiable sub-regions; the video processing apparatus 555 further includes a determining module 5553, configured to, in response to a selection operation for the plurality of modifiable sub-regions, designate the selected modifiable sub-region as the target modifiable sub-region; and an updating module 5552, configured to, in response to an editing operation for the target modifiable sub-region based on an editing control, update the target modifiable sub-region according to the editing operation.
[0172] In some embodiments, the editing control is displayed in a video editing interface, which also includes multiple sub-region controls corresponding to multiple modifiable sub-regions; the display module 5551 is further configured to highlight the triggered sub-region control in response to a trigger operation on the multiple sub-region controls; the determination module 5553 is further configured to use the modifiable sub-region corresponding to the triggered sub-region control as the target modifiable sub-region.
[0173] In some embodiments, the editing control is displayed in a video editing interface, which also includes an audio adjustment control; the video processing device 555 further includes an acquisition module 5554, configured to acquire object feature data of a target object in response to a trigger operation on the audio adjustment control, wherein the target object is the object to be pushed in the video; the update module 5552 is further configured to query an audio from an audio library that matches the object feature data, and adjust the audio associated with the video to the matching audio.
[0174] In some embodiments, the update module 5552 is further configured to determine, for each audio in the audio library, the similarity between the audio feature data of the audio and the object feature data; and to use audio with a similarity greater than a similarity threshold as audio that matches the object feature data.
[0175] In some embodiments, before the display module 5551 displays the editing control in response to a video editing trigger operation, the acquisition module 5554 is further configured to acquire the adjusted modifiable area in response to a range adjustment operation for the modifiable area; the update module 5552 is further configured to update the adjusted modifiable area according to the editing operation based on the editing control for the adjusted modifiable area.
[0176] In some embodiments, the editing control is displayed in a video editing interface, which also includes a video playback timeline; the acquisition module 5554 is further configured to acquire the set time period in response to a time period setting operation based on the playback timeline; the update module 5552 is further configured to update the modifiable area in the multi-frame image corresponding to the target time period of the video according to the editing operation based on the editing control for the modifiable area, wherein the target time period is the time period other than the set time period in the total playback duration of the video.
[0177] In some embodiments, the editing control is displayed in the video editing interface, which also includes a style adjustment control; the acquisition module 5554 is further configured to acquire object feature data of the target object in response to a trigger operation on the style adjustment control, wherein the target object is the object to be pushed in the video; the update module 5552 is further configured to adjust the style of the video to conform to the style of the object feature data.
[0178] In some embodiments, the editing control is displayed in a video editing interface, which also includes a preview control; the video processing device 555 further includes a jump module 5555 and a generation module 5556, wherein the jump module 5555 is used to jump from the video editing interface to the preview interface in response to a trigger operation on the preview control, wherein the preview interface includes multiple object templates, and different object templates correspond to different object feature data; the generation module 5556 is used to generate a target video that conforms to the selected object template based on the video after the modifiable area is updated in response to a selection operation on multiple object templates.
[0179] In some embodiments, the video processing apparatus 555 further includes a recognition module 5557, configured to perform the following processing on each frame of the video before the display module 5551 displays the video interface: recognizing the image using a pre-trained image recognition model; the determination module 5553 is further configured to designate the region where the target element is located in the image as a prohibited modification region, and designate other regions in the image besides the region where the target element is located as modifiable regions; wherein the image recognition model is trained based on sample images, and the sample images include the target element and labels annotated for the target element.
[0180] In some embodiments, after displaying the video interface, the display module 5551 is further configured to perform at least one of the following processes: marking the modifiable area with a transparent layer; marking the unmodifiable area with a mask; and displaying a prompt message in the video interface, wherein the prompt message is used to indicate the modifiable area and the unmodifiable area.
[0181] It should be noted that the description of the apparatus in this application embodiment is similar to the description of the method embodiment above, and has similar beneficial effects as the method embodiment; therefore, it will not be repeated. For any technical details not covered in the video processing apparatus provided in this application embodiment, please refer to... Figure 3 , Figure 4 ,or Figure 5 The meaning is understood in accordance with the description of any of the accompanying drawings.
[0182] This application provides a computer program product, which includes a computer program or computer-executable instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer-executable instructions from the computer-readable storage medium and executes the computer-executable instructions, causing the computer device to perform the video processing method described in this application.
[0183] This application provides a computer-readable storage medium storing computer-executable instructions. When these computer-executable instructions are executed by a processor, they cause the processor to execute the video processing method provided in this application. For example, ... Figure 3 , Figure 4 ,or Figure 5 The video processing method shown.
[0184] In some embodiments, the computer-readable storage medium may be a memory such as FRAM, ROM, PROM, EPROM, EEPROM, flash memory, magnetic surface memory, optical disk, or CD-ROM; or it may be a variety of devices including one or any combination of the above-mentioned memories.
[0185] In some embodiments, executable instructions may take the form of a program, software, software module, script, or code, written in any form of programming language (including compiled or interpreted languages, or declarative or procedural languages), and may be deployed in any form, including as a standalone program or as a module, component, subroutine, or other unit suitable for use in a computing environment.
[0186] As an example, executable instructions can be deployed to execute on a single electronic device, or on multiple electronic devices located in one location, or on multiple electronic devices distributed across multiple locations and interconnected via a communication network.
[0187] The above description is merely an embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, and improvements made within the spirit and scope of this application are included within the scope of protection of this application.
Claims
1. A video processing method, characterized in that, The method includes: The video interface includes a video, which includes an editable area and an uneditable area. In response to a video editing trigger, display the editing controls; In response to an editing operation on the modifiable region based on the editing control, the modifiable region is updated according to the editing operation.
2. The method according to claim 1, characterized in that, The editing control includes a special effects addition control, and the editing operation includes a special effects addition operation; The step of responding to an editing operation on the modifiable region based on the editing control and updating the modifiable region according to the editing operation includes: The trigger operation for adding controls to the special effect is identified as the special effect addition operation; In response to the special effects addition operation, the object feature data of the target object is obtained, wherein the target object is the object to be pushed in the video; The system queries the effects library for effects that match the object's feature data and adds the matching effects to the modifiable area.
3. The method according to claim 2, characterized in that, Adding the matching effect to the modifiable area includes: For at least one frame of the video, the following processing is performed: the matching effect is added to the modifiable area included in the image, wherein the at least one frame includes: multiple frames of the video, at least one frame of the multiple frames of the video corresponding to a set time period, and at least one selected frame of the multiple frames of the video.
4. The method according to claim 2, characterized in that, The step of querying the special effects library for special effects that match the object's feature data includes: For each effect in the effects library, determine the similarity between the effect's feature data and the object's feature data; The effects with a similarity greater than a similarity threshold are considered as effects that match the object feature data.
5. The method according to claim 1, characterized in that, The editing control includes a scene switching control, and the editing operation includes a scene switching operation; The step of responding to an editing operation on the modifiable region based on the editing control and updating the modifiable region according to the editing operation includes: The trigger operation for the scene switching control is identified as the scene switching operation; In response to the scene switching operation, the region feature data of the area where the target object is currently located is obtained, wherein the target object is the object to be pushed in the video; Switch the scene displayed in the modifiable area to a scene that matches the area's feature data.
6. The method according to claim 5, characterized in that, The step of switching the scene displayed in the modifiable area to a scene that conforms to the area feature data includes: Based on the regional feature data, a pre-built mapping table is queried to obtain the scene that matches the regional feature data. The mapping table includes the mapping relationship between different regional feature data and different scenes. Switch the scene displayed in the modifiable area to a scene that matches the feature data of the area.
7. The method according to claim 5, characterized in that, The step of switching the scene displayed in the modifiable area to a scene that conforms to the area feature data includes: For at least one frame of the video, the following processing is performed: the scene displayed in the modifiable area of the image is switched to a scene that conforms to the feature data of the area, wherein the at least one frame of the video includes: multiple frames of the video, at least one frame of the multiple frames of the video that corresponds to the set time period, and at least one selected frame of the multiple frames of the video.
8. The method according to claim 1, characterized in that, The editing control includes a scene switching control, and the editing operation includes a scene switching operation; The step of responding to an editing operation on the modifiable region based on the editing control and updating the modifiable region according to the editing operation includes: The trigger operation for the scene switching control is identified as the scene switching operation; In response to the scene switching operation, object feature data of the target object is obtained, wherein the target object is the object to be pushed in the video; Switch the scene displayed in the modifiable area to a scene that matches the object's characteristic data.
9. The method according to claim 8, characterized in that, The step of switching the scene displayed in the modifiable area to a scene that conforms to the object characteristic data includes: For each scene in the scene library, determine the similarity between the scene feature data of the scene and the object feature data; The scenarios with a similarity greater than a similarity threshold are considered as scenarios that match the object feature data; Switch the scene displayed in the modifiable area to the matching scene.
10. The method according to claim 1, characterized in that, The editing control includes an item adjustment control, and the editing operation includes an item adjustment operation; The step of responding to an editing operation on the modifiable region based on the editing control and updating the modifiable region according to the editing operation includes: The trigger operation for adjusting the control of the item is identified as the item adjustment operation; In response to the item adjustment operation, the object feature data of the target object is obtained, wherein the target object is the object to be pushed in the video; Adjust the items displayed in the modifiable area to match the object's characteristic data.
11. The method according to claim 10, characterized in that, The step of adjusting the items displayed in the modifiable area to conform to the object characteristic data includes: For each item in the item database, determine the similarity between the item feature data and the object feature data; Items with a similarity greater than a similarity threshold are considered as items that match the object feature data; Adjust the items displayed in the modifiable area to match the items.
12. The method according to any one of claims 1 to 11, characterized in that, The modifiable region includes multiple modifiable sub-regions; The step of responding to an editing operation on the modifiable region based on the editing control and updating the modifiable region according to the editing operation includes: In response to a selection operation for the plurality of modifiable sub-regions, the selected modifiable sub-region is taken as the target modifiable sub-region; In response to an editing operation on the target modifiable sub-region based on the editing control, the target modifiable sub-region is updated according to the editing operation.
13. The method according to claim 12, characterized in that, The editing controls are displayed in the video editing interface, which also includes multiple sub-area controls that correspond one-to-one with the multiple modifiable sub-areas. The step of responding to a selection operation for the plurality of modifiable sub-regions by taking the selected modifiable sub-region as the target modifiable sub-region includes: In response to a trigger operation on the plurality of sub-region controls, the triggered sub-region control is highlighted, and the modifiable sub-region corresponding to the triggered sub-region control is taken as the target modifiable sub-region.
14. The method according to any one of claims 1 to 11, characterized in that, The editing controls are displayed in the video editing interface, which also includes audio adjustment controls; The method further includes: In response to a trigger operation on the audio adjustment control, object feature data of the target object is obtained, wherein the target object is the object to be pushed in the video; The audio is retrieved from the audio library to match the object's feature data, and the audio associated with the video is adjusted to match the audio.
15. The method according to claim 14, characterized in that, The step of querying the audio library for audio that matches the object feature data includes: For each audio file in the audio library, determine the similarity between the audio feature data of the audio file and the object feature data; The audio with a similarity greater than the similarity threshold is considered as the audio that matches the object feature data.
16. The method according to any one of claims 1 to 11, characterized in that, Before displaying the editing controls in response to a video editing trigger operation, the method further includes: In response to a range adjustment operation on the modifiable region, the adjusted modifiable region is obtained; The step of responding to an editing operation on the modifiable region based on the editing control and updating the modifiable region according to the editing operation includes: In response to an editing operation on the adjusted modifiable area based on the editing control, the adjusted modifiable area is updated according to the editing operation.
17. The method according to any one of claims 1 to 11, characterized in that, The editing controls are displayed in the video editing interface, which also includes the video playback timeline; The method further includes: In response to a time period setting operation based on the playback timeline, the set time period is obtained; The step of responding to an editing operation on the modifiable region based on the editing control and updating the modifiable region according to the editing operation includes: In response to an editing operation on the modifiable region based on the editing control, the modifiable region in the multi-frame image corresponding to the target time period of the video is updated according to the editing operation, wherein the target time period is the time period in the total playback duration of the video excluding the set time period.
18. The method according to any one of claims 1 to 11, characterized in that, The editing controls are displayed in the video editing interface, which also includes style adjustment controls; The method further includes: In response to a trigger operation on the style adjustment control, object feature data of the target object is obtained, wherein the target object is the object to be pushed in the video; Adjust the style of the video to match the style of the object feature data.
19. The method according to any one of claims 1 to 11, characterized in that, The editing controls are displayed in the video editing interface, which also includes a preview control; The method further includes: In response to a trigger operation on the preview control, the user is redirected from the video editing interface to the preview interface, wherein the preview interface includes multiple object templates and different object feature data corresponding to different object templates. In response to the selection operation for the plurality of object templates, a target video matching the selected object template is generated based on the video after the modification region is updated.
20. The method according to any one of claims 1 to 11, characterized in that, Prior to displaying the video interface, the method further includes: For each frame of the video, the following processing is performed: The image is recognized using a pre-trained image recognition model; The region containing the target element in the image is designated as the region that cannot be modified, and the other regions in the image, excluding the region containing the target element, are designated as the regions that can be modified. The image recognition model is trained based on sample images, which include the target element and labels annotated for the target element.
21. The method according to any one of claims 1 to 11, characterized in that, After displaying the video interface, the method further includes: Perform at least one of the following processes: Use a transparent layer to mark the modifiable area; Use a mask to mark the areas that are prohibited from modification; The video interface displays a prompt message, which indicates the modifiable area and the prohibited area.
22. A video processing apparatus, characterized in that, The device includes: A display module is used to display a video interface, wherein the video interface includes a video, and the video includes an editable area and an uneditable area; The display module is also used to display editing controls in response to a video editing trigger operation; An update module is configured to update the modifiable region in response to an editing operation performed on the modifiable region based on the editing control.
23. An electronic device, characterized in that, include: Memory is used to store executable instructions for a computer; A processor, when executing computer-executable instructions stored in the memory, implements the video processing method according to any one of claims 1 to 21.
24. A computer-readable storage medium storing computer-executable instructions, characterized in that, When the computer-executable instructions are executed by the processor, they implement the video processing method according to any one of claims 1 to 21.
25. A computer program product comprising a computer program or computer-executable instructions, characterized in that, When the computer program or computer-executable instructions are executed by the processor, the video processing method according to any one of claims 1 to 21 is implemented.