Video generation method and device, equipment, medium and product
By acquiring demand information and user preferences, and using the AIGC model to generate videos with embedded ads, the problem of the separation between ads and content in the traditional ad overlay mode is solved, and the natural integration of ads and video content is achieved, improving user experience and ad conversion rate.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-05
AI Technical Summary
Traditional ad overlay models are difficult to match with user experience in the commercialization of video content, resulting in a disconnect between ads and content, which affects the viewing immersion and ad conversion rate.
By acquiring generation demand information, we determine embedding decision information, including embedding carrier, quantity, and display information. We use the AIGC model to generate video content that naturally integrates advertisements and videos, and optimize the advertisement embedding strategy by combining user preferences and historical data.
Achieve seamless integration between advertising and video content, reduce user resistance, and improve ad acceptance and conversion rates.
Smart Images

Figure CN121985187A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of video processing technology, and more specifically, to a video generation method, apparatus, device, medium, and product. Background Technology
[0002] In the process of commercializing video content, traditional ad overlay models are no longer suitable for today's users' high demands for content experience. These models often use simple overlays and hard-sell placements without considering the scene atmosphere, visual style, and narrative rhythm of the video itself. This results in a clear disconnect between the advertisement and the content, which not only destroys the viewing immersion but also easily triggers strong user resistance, even leading to users skipping ads or uninstalling the app, seriously affecting ad conversion rates. Summary of the Invention
[0003] The purpose of this disclosure is to provide a video generation method, apparatus, device, medium, and product.
[0004] To achieve the above objectives, in a first aspect, this disclosure provides a video generation method, the method comprising:
[0005] Obtain the generation requirement information; Based on the generated demand information, the embedding decision information of the target advertisement is determined, wherein the embedding decision information includes at least one of the following: embedding carrier information, advertisement quantity information, and advertisement display information; Based on the generated demand information and the embedded decision information, a video with the target advertisement embedded is generated.
[0006] Optionally, determining the embedding decision information of the target advertisement based on the generated demand information includes: Based on the generated demand information and user preference information, the embedding decision information of the target advertisement is determined. The user preference information includes the acceptance information of different users in historical videos regarding embedding carrier information, number of advertisements, and advertisement display information.
[0007] Optionally, determining the embedding decision information of the target advertisement based on the generated demand information and user preference information includes: Based on the generated demand information, the extended information of the target advertisement is determined, and the extended information includes at least one of scene-related extended information, person-related extended information, and object-related extended information; Based on the extended information and the user preference information, the embedding decision information for the target advertisement is determined.
[0008] Optionally, generating a video embedded with the target advertisement based on the generated demand information and the embedded decision information includes: Based on the generated requirement information, determine the supplementary information for the generated requirement information; Based on the generation requirement information, the supplementary information, and the embedding decision information, the video generation instruction is determined; The video containing the target advertisement is generated based on the video generation instructions and the AIGC model.
[0009] Optionally, generating the video embedded with the target advertisement according to the video generation instructions and the AIGC model includes: Based on the video generation instructions and the AIGC model, multiple video generation schemes with different ad display levels are determined; In response to the instruction to determine the target video generation scheme, a video embedded with the target advertisement is generated, and the display degree of the target advertisement is consistent with the advertisement display degree corresponding to the target video generation scheme.
[0010] Optionally, each display resolution corresponds to a different incentive benefit. After generating the video embedding the target advertisement according to the video generation instructions and the AIGC model, the method further includes: In response to the generation of the video, the corresponding incentive benefits for the video are distributed to the account.
[0011] Optionally, determining supplementary information for the generated requirement information based on the generated requirement information includes: Obtain the dimension information corresponding to the generated requirement information, wherein the dimension information includes at least one of scene dimension information, character dimension information and object dimension information; Based on the dimensional information, the generated requirement information is supplemented to obtain the supplementary information.
[0012] Secondly, this disclosure provides a video generation apparatus, the apparatus comprising: The information acquisition module is configured to acquire the generation requirement information; The information determination module is configured to determine the embedding decision information of the target advertisement based on the generated demand information, wherein the embedding decision information includes at least one of embedding carrier information, advertisement quantity information, and advertisement display information; The video generation module is configured to generate a video with the target advertisement embedded in it, based on the generation requirement information and the embedding decision information.
[0013] Thirdly, this disclosure provides an electronic device, including: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method of any one of the first aspects.
[0014] Fourthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon that, when executed by a processor, implements the steps of the method described in any of the first aspects.
[0015] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in any one of the first aspects, or the steps of the method described in any one of the second aspects.
[0016] The above technical solution determines the embedding decision information of advertisements based on the video generation requirements information. During the video generation process, the advertisements are generated together based on the generation requirements information and the embedding decision information, so as to achieve a natural integration of advertisements and video content, reduce users' resistance to advertisements in the generated video, and improve the acceptance of advertisements.
[0017] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0018] The accompanying drawings are provided to further illustrate the present disclosure and form part of the specification. They are used together with the following detailed description to explain the present disclosure, but do not constitute a limitation thereof. In the drawings: Figure 1 This is a flowchart illustrating a video generation method according to an exemplary embodiment.
[0019] Figure 2 This is a block diagram of a video generation apparatus according to an exemplary embodiment.
[0020] Figure 3 This is a block diagram illustrating an electronic device according to an exemplary embodiment. Detailed Implementation
[0021] The specific embodiments of this disclosure will be described in detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are for illustration and explanation only and are not intended to limit this disclosure.
[0022] It should be noted that all actions involving the acquisition of signals, information, or data in this disclosure are carried out in compliance with the relevant data protection laws and policies of the country where the location is situated, and with authorization from the owner of the relevant device.
[0023] To address the technical problems raised in the background, this application provides a video generation method. Figure 2 This is a flowchart illustrating a video generation method according to an exemplary embodiment, the method comprising the following steps.
[0024] In step S11, the generation requirement information is obtained.
[0025] The generation requirement information refers to the specific parameters or conditions used to guide video generation, including key information such as target ad content, video duration, resolution, style template, and placement scenario. This generation requirement information can be obtained by extracting it from user-input text and / or voice information, by calling a pre-configured interface from the backend system, or by intelligently recommending and generating it based on user historical behavior data.
[0026] In this embodiment, to facilitate interaction with the user, a client interface is provided for the client, which is used not only to receive text and / or voice information input by the user, but also to provide feedback to the user on the video generated based on the generation requirement information.
[0027] After obtaining user-inputted text or voice information through the client interface, the system first converts the voice information to text, then cleans the text, and finally standardizes the information format to obtain unified, structured data for generating requirements. Voice-to-text conversion can be performed using a model, such as the Whisper model, which supports multilingual recognition and dialect adaptation. Text cleaning removes special symbols and redundant interjections, and performs word segmentation, part-of-speech tagging, and stop word detection. Information format standardization uses a predefined rule engine to map the cleaned text into structured fields containing ad content category, target duration range, resolution level, style tags, and scenario usage, ensuring that subsequent modules can accurately parse and execute generation instructions. The entire process is completed in milliseconds after the user triggers the generation request, ensuring real-time interaction and a smooth user experience.
[0028] In step S12, based on the generated demand information, the embedding decision information of the target advertisement is determined. The embedding decision information includes at least one of the embedding carrier information, advertisement quantity information, and advertisement display information.
[0029] Embedding decision information refers to the key parameters that guide the specific embedding method of advertisements within video content. Embedding carrier information specifies the video segment or scene location where the advertisement will be inserted. Ad quantity information specifies the total number of advertisements inserted in a single video. Ad display information includes details such as the ad presentation format, duration, playback timing, and interactive functions. By analyzing the placement scenario and target audience characteristics in the generated demand information, the optimal embedding strategy is determined to ensure that the ad content blends naturally into the video narrative, improving user experience and conversion rates. The generation of embedding decision information relies on an understanding of the target ad content and a precise grasp of the video's contextual semantics.
[0030] Furthermore, the method for determining the embedded decision information can be a multi-dimensional understanding of the generated demand information. This can be based on user preference information or on the user's historical data; no specific limitations are imposed here. Multi-dimensional understanding includes, but is not limited to, dimensions such as person, scene, and object. Information related to the person dimension includes, but is not limited to, the number of people, their identities (office worker, student, athlete, driver, etc.), core actions (walking, drinking, working, driving, exercising, etc.), clothing characteristics (clothing type: T-shirt / jacket / sportswear; accessories: backpack / hat / watch, etc.), and posture (standing / sitting / dynamic movement). Information related to the scene dimension includes, but is not limited to, macro-scenes (outdoor / indoor), further refined to micro-scenes (outdoor: street, park, sports field, highway, etc.; indoor: office, family living room, restaurant, coffee shop, etc.), simultaneously extracting scene attributes (lighting conditions: strong light / backlight / weak light; time: day / night; style tone: realistic / cartoon / science fiction, etc.). The object dimension includes basic object dimensions and transportation dimensions. The basic object dimension includes, but is not limited to, everyday items (water cups, computers, mobile phones, backpacks, tableware), decorative elements (walls, desktops, ornaments), and consumer products (drinks, food, cosmetics), recording the object's shape (round / square) and usage status (held / placed / in use). The transportation dimension includes, but is not limited to, the presence of a transportation vehicle (car, bicycle, subway, bus, etc.), recording the transportation type, usage scenario (moving / stopped), and appearance characteristics (color, model, blank painted areas).
[0031] Specifically, taking the determination of embedding decision information based on user's historical data as an example, we can first determine the correspondence between the video scenes corresponding to different generation requirement information and the embedding decision information based on the user's historical data. When the user inputs the generation requirement information, the embedding decision information will be determined based on this correspondence, thereby improving the accuracy and personalization of the embedded decision information determination.
[0032] In some implementations, determining the embedding decision information of the target advertisement based on the generated demand information includes: determining the embedding decision information of the target advertisement based on the generated demand information and user preference information, wherein the user preference information includes the acceptance information of different users in historical videos regarding embedding carrier information, number of advertisements, and advertisement display information.
[0033] User preference information refers to personalized preference features derived from data such as users' historical data (e.g., types of ad carriers previously accepted, modified ad embedding positions), feedback records (e.g., ad styles rejected, number of preferred ads), and interactive behaviors (e.g., clicks, dwell time, swipe patterns). User preference information can take the form of a structured tag set or a preference model. Preferably, it is in the form of a preference model. The preference model is trained on user behavior data using machine learning algorithms to label users' acceptance ratings for different embedding carrier information, ad quantity information, and ad display information.
[0034] Furthermore, the selection of embedding carriers, the setting of the number of ads, and the configuration of display methods in the embedded decision information are all dynamically optimized based on the acceptance scores output by the preference model, ensuring that ad embedding conforms to user visual habits and scene logic. The selection of embedding carriers can prioritize carriers with high historical user acceptance, combined with the rationality of the current video scene. The setting of the number of ads can be based on video length (e.g., a maximum of 2 ads in a 15-second video, a maximum of 4 ads in a 60-second video) and user preferences to determine a reasonable number of ads, avoiding excessive insertion. The acceptance information of ad display includes the selection and adjustment of display level. The default display level is medium clarity; if the user has a historical preference for high display level, the level is increased; if the user has a historical aversion to ads, a blurred level is selected. Display level is controlled by adjusting the size, position, and dwell time of the ads on the screen (e.g., high display level: center of the screen, 10%-15% of the screen area, dwell time of more than 3 seconds; low display level: edge of the screen, 5% or less of the screen area, dwell time of 1-2 seconds).
[0035] In some specific implementations, determining the embedding decision information of the target advertisement based on the generated demand information and user preference information includes: determining the extended information of the target advertisement based on the generated demand information, wherein the extended information includes at least one of scene-related extended information, person-related extended information, and object-related extended information; and determining the embedding decision information of the target advertisement based on the extended information and the user preference information.
[0036] In this implementation, there are instances where the generated requirement information is ambiguous. To better understand the generation intent, the generated requirement information needs semantic expansion and contextual completion. Expanded information includes, but is not limited to, scene-related expanded information, person-related expanded information, and object-related expanded information. Scene-related expanded information is refined based on the environmental background suitable for the target advertisement, supplementing lighting conditions, spatial layout, and time-of-day characteristics. For example, an outdoor street scene can be expanded to include building wall advertisements, roadside light boxes, and bus stop advertisements. An office scene can be expanded to include advertisements on office equipment, desktop ornaments, and beverage cups. Person-related expanded information is refined based on the interaction possibilities between the person and the advertisement, supplementing the person's gaze direction, movement trajectory, and social attributes. For example, a person's clothing can be expanded to include clothing brand logos and clothing print advertisements. A person holding a water cup can be expanded to include beverage brand and cup packaging advertisements. A person wearing a watch can be expanded to include wristwatch brand advertisements. Object-related expanded information is extended based on the functional attributes and usage scenarios of objects in the environment, supplementing the object's material, shape characteristics, and interaction frequency. For example, a desktop computer can be expanded to include computer body sticker advertisements and screen saver advertisements. Cars in motion can be expanded into body paint advertising, window sticker advertising, and in-vehicle device interface advertising. Park benches can be expanded into brand lettering on the bench surface and side advertising signs. By using extended information and user preference information to jointly determine the embedding decision information, the accuracy and naturalness of ad embedding can be effectively improved.
[0037] In step S13, a video with the target advertisement embedded is generated based on the generated demand information and the embedded decision information.
[0038] In this embodiment, based on the content framework in the generated demand information and the specific parameters in the embedded decision information, a pre-trained video generation model is invoked to perform multimodal fusion rendering. This process integrates advertising elements into the original scene in a semantically consistent and spatially coordinated manner, ensuring lighting matching, perspective uniformity, and dynamic coherence.
[0039] In some implementations, generating a video embedded with the target advertisement based on the generation requirement information and the embedding decision information includes: determining supplementary information of the generation requirement information based on the generation requirement information; determining a video generation instruction based on the generation requirement information, the supplementary information, and the embedding decision information; and generating the video embedded with the target advertisement based on the video generation instruction and an AIGC model.
[0040] The supplementary information for generating demand information refers to the details missing from the generated demand information. The supplementary rules can be: original generated demand information + multi-dimensional element details + advertising carrier constraints (position / size / display) + style consistency requirements + natural integration conditions (light and shadow adaptation, perspective consistency). Supplementary information can include environmental parameters, interaction logic, and temporal relationships. Environmental parameters supplement information such as light intensity, weather conditions, and spatial dimensions. Interaction logic clarifies the interaction methods between the advertisement and the main subject, such as touch, gaze, or usage behavior. Temporal relationships define the start and end times of the advertisement's appearance, its duration, and its order with other elements.
[0041] For example, let's take the original input for generating demand information as "an office worker drinking coffee outdoors". The supplementary information is "an office worker wearing a business casual T-shirt is drinking coffee outside an outdoor street cafe. The white water cup he is holding has a minimalist beverage brand logo printed on its surface (the logo occupies no more than 10% of the cup and is consistent with the direction of light and shadow on the cup). There are low-saturation lightbox advertisements for daily necessities brands on both sides of the street in the background (located at the edge of the image, occupying less than 5%). The overall lighting of the image is natural daylight, and the advertising elements are consistent with the style of the scene."
[0042] Furthermore, AI-Generated Content (AIGC) is a type of deep learning model specifically designed for automatically generating text, images, audio, video, code, and other content. This model, trained on massive amounts of multimodal data, can understand the relationship between semantic context and visual structure, thereby automatically generating semantically consistent visual content. In this embodiment, the AIGC model generates a sequence of video frames containing target advertising elements based on video generation instructions, ensuring that the ad embedding position, timing, and interaction logic conform to the instructions. The AIGC model parses the spatial coordinates and environmental parameters in the embedded decision information, dynamically adjusting the perspective angle, lighting matching, and motion trajectory of the advertising materials, allowing brand elements to naturally blend into the scene.
[0043] Understandably, video generation commands need to pay attention to details related to ad blending, including but not limited to lighting and shadow adaptation, perspective adaptation, and style consistency. Lighting and shadow adaptation refers to adjusting the lighting and shadow effects of the ad creative based on the direction and intensity of light in the video scene (e.g., adding shadows to the ad creative in backlit scenes, reducing the brightness of the ad creative in bright light scenes). Perspective adaptation refers to adjusting the perspective distortion of the ad creative based on the spatial angle of the ad carrier. For example, a tilted water glass or curved clothing fabric, ensuring it fits the carrier perfectly. Style consistency refers to adjusting the style of the ad creative to be consistent with the overall style of the video. For example, in a cartoon-style video, realistic ad creatives are converted to a cartoon style.
[0044] During video generation, to improve the quality of videos generated by the AIGC model, a fusion quality check is performed after video generation. This check uses an AI detection model to assess the naturalness of the merged ads, checking for issues such as mismatched lighting, perspective errors, or jarring styles. If the check fails, the fusion parameters are readjusted and a new video is generated.
[0045] Furthermore, to better cater to user needs, we will respond to user feedback and modification requests, and then generate corresponding videos based on these requests. Modification requests include, but are not limited to, changes to the placement, number, and visibility of advertisements within the video. The video generation process is only completed after quality testing and user confirmation. The entire process follows a closed-loop logic of "generation—testing—optimization," ensuring that the advertising content blends naturally with the scene, improving visual consistency and user experience.
[0046] In some specific implementations, generating the video embedded with the target advertisement according to the video generation instructions and the AIGC model includes: determining multiple video generation schemes with different ad display levels according to the video generation instructions and the AIGC model; and generating a video embedded with the target advertisement in response to the instruction to determine the target video generation scheme, wherein the display level of the target advertisement is consistent with the ad display level corresponding to the target video generation scheme.
[0047] Ad visibility refers to the clarity of an ad's display within a video. Ad visibility can be categorized into three levels: blurry, medium clarity, and clear. Blurry means only the core visual features and style are retained, such as "blue-toned, minimalist beverage packaging." Medium clarity includes added product type and some details, such as "blue-toned, minimalist bottled mineral water packaging with a circular logo area." Clear means the brand and product details are fully preserved, such as "First Brand blue bottled mineral water with a circular First Brand logo on the bottle and a simple white label."
[0048] In this embodiment, the AIGC model's generation parameters are dynamically adjusted based on the user-selected ad display level to control the level of detail in the ad elements within the video. When a blurred display is selected, only the color tone and outline are retained; when medium clarity is selected, product category and structural features are overlaid; and when clear display is selected, brand logos, packaging details, and text content are fully reproduced. The rendering parameters of the ad creatives are intelligently matched to the scene's lighting, shadows, and perspective to ensure visual realism in different scenarios. For example, in a scene with strong light, clear ads will simultaneously generate highlights and edge reflections, while blurred ads will only slightly reflect light and shadow tendencies, avoiding information overload. This tiered control strategy satisfies both brand exposure needs and visual harmony. The system supports real-time preview of multiple display options, allowing users to choose the optimal path based on visual effects, ultimately generating high-quality video content that aligns with creative intent and communication goals.
[0049] The video generation scheme can include at least 2-3 embedding schemes with different ad counts and display levels. These schemes can be sorted based on user preferences or historical data to better present multiple schemes to the user. The system automatically sorts and recommends schemes based on the user's historical adoption preferences, such as a preference for high-definition or low-interference embedding. Manual filtering and adjustment are also supported to ensure flexibility and control. In this implementation, a hybrid strategy of "collaborative filtering + reinforcement learning" can also be used: based on the collaborative filtering algorithm, matching materials are recommended using the ad preferences of similar users; based on the reinforcement learning algorithm, embedding parameter decisions are dynamically optimized using "user acceptance rate + ad exposure effect + user experience score" as the reward function; and the gradient boosting tree (GBDT) model is used to predict the probability of user acceptance of different schemes, improving the accuracy of scheme recommendations.
[0050] Specifically, each display resolution corresponds to a different incentive benefit. After generating the video with the target advertisement embedded according to the video generation instructions and the AIGC model, the method further includes: in response to the generation of the video, distributing the incentive benefits corresponding to the video to the account.
[0051] Here, "account" refers to the user account. Users need to log in to their relevant accounts to perform operations such as inputting requirement information and generating videos. Ad visibility and incentive benefits are positively correlated; the higher the clarity, the more incentive benefits users can obtain. For example, when the visibility is blurry and there are few ads, the incentive benefits are basic points. When the visibility is moderately clear and there are a reasonable number of ads, the incentive benefits include, but are not limited to, advanced points and free video generation attempts. When the visibility is clear and there are a compliant number of ads, the incentive benefits include, but are not limited to, high points and cash subsidies. After completing video generation, users can view and claim the corresponding incentive benefits through their personal accounts. The system automatically calculates rewards based on the ad visibility level and the number of ads embedded. Basic points are credited instantly and used for subsequent service deductions; advanced points can be accumulated and redeemed for membership benefits or exclusive material library access; cash subsidies are reviewed and issued periodically, subject to platform compliance requirements. All incentive records are searchable within the account, ensuring transparency and fairness. Clarity grading not only affects visual expression but also deepens the linkage mechanism between user behavior and commercial value. The high incentives brought by high clarity encourage users to actively optimize ad presentation quality, thereby improving brand communication effectiveness.
[0052] In some other specific embodiments, determining supplementary information for the generated requirement information based on the generated requirement information includes: obtaining dimensional information corresponding to the generated requirement information, wherein the dimensional information includes at least one of scene dimensional information, character dimensional information, and object dimensional information; and supplementing the generated requirement information based on the dimensional information to obtain the supplementary information.
[0053] In this implementation, the dimensional information can be obtained by analyzing the generated requirement information from three dimensions: visual, semantic, and style. The visual dimension involves extracting color features (dominant color tone, secondary color), shape features (logo outline, product shape), and texture features (such as clothing fabric texture, product surface texture). The semantic dimension involves extracting core information (brand name, product type, core selling points) and applicable scenarios (such as sports brands for sports scenarios, and office equipment brands for office scenarios). The style dimension involves matching common AIGC video styles (such as realistic, cartoon, Chinese style, and science fiction) to generate style-appropriate descriptions. By generating requirements through multi-dimensional analysis, the system can intelligently recommend suitable advertising element libraries, ensuring a high degree of integration between advertising content and scenes, people, and objects.
[0054] Furthermore, the acquisition of dimensional information can be further refined by analyzing scene, character, and object dimensions. For scene acquisition, we can first identify the macro-level scene (outdoor / indoor) in the generation requirements information, then refine it to identify the micro-level scene (outdoor: streets, parks, sports fields, highways, etc.; indoor: offices, family living rooms, restaurants, cafes, etc.), and simultaneously extract scene attributes (lighting conditions: strong light / backlight / weak light; time: day / night; style tone: realistic / cartoon / science fiction, etc.). For character acquisition, we can first identify the macro-level scene (outdoor / indoor) in the generation requirements information, then refine it to identify the micro-level scene (outdoor: streets, parks, sports fields, highways, etc.; indoor: offices, family living rooms, restaurants, cafes, etc.), and simultaneously extract scene attributes (lighting conditions: strong light / backlight / weak light; time: day / night; style tone: realistic / cartoon / science fiction, etc.). The acquisition of object dimensions includes the acquisition of basic object dimensions and the acquisition of vehicles. Acquiring the basic object dimension allows for the classification and identification of core objects within the scene generated from the requirement information. These include everyday items (water cups, computers, mobile phones, backpacks, tableware), decorative elements (walls, desktops, ornaments), and consumer products (drinks, food, cosmetics), recording the object's shape (round / square) and usage status (held / placed / in use). Acquiring the transportation dimension allows for the identification of whether transportation (cars, bicycles, subways, buses, etc.) exists in the generated requirement information, recording the transportation type, usage scenario (moving / parked), and appearance characteristics (color, model, blank painted areas).
[0055] Understandably, dimensional information can be obtained through a semantic understanding model that deeply analyzes the generated requirement information, extracting multi-dimensional semantic features such as scenes, people, and objects, and combining this with the context to achieve accurate recognition. The semantic understanding model can be a fusion model of BERT and CLIP, or other semantic understanding models. By using a BERT and CLIP fusion model, alignment between textual and visual semantic spaces can be achieved, accurately capturing the implicit information in the generated requirements.
[0056] In its implementation, this application provides a client-side interface and an advertising interface. The client-side interface receives text / voice input from user-generated videos, displays recommended advertising embedding schemes, receives user feedback on modifications, and displays incentive benefits. The advertising interface receives uploaded advertising materials, feedback on material review results, and queries advertising delivery data statistics. To ensure interface security, prevent malicious requests, and control concurrent access, this application also provides authentication and traffic control modules. Users log in to their user accounts and submit generation requirements using the client-side interface. Upon receiving the generation requirements, the system performs multi-dimensional semantic analysis to extract scene, character, object, and advertising-related features. The advertising interface is then used to obtain relevant advertising material information. Combined with scene adaptability, user preference tags, and historical delivery data, intelligent matching is performed to generate multiple candidate embedding schemes. Finally, based on user feedback and selection, the final generated video content and advertising embedding scheme are determined, and dynamic rendering output is completed.
[0057] In this solution, the embedding decision information for advertisements is determined based on the video generation requirements. During the video generation process, the advertisements are generated together based on both the generation requirements and the embedding decision information, achieving a natural integration of advertisements and video content. This reduces user resistance to advertisements in the generated video and increases ad acceptance.
[0058] Figure 2 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment. (Refer to...) Figure 2 The video generation device 300 includes an information acquisition module 310, an information determination module 320, and a video generation module 330.
[0059] The information acquisition module 310 is configured to acquire generation requirement information; The information determination module 320 is configured to determine the embedding decision information of the target advertisement based on the generated demand information, wherein the embedding decision information includes at least one of embedding carrier information, advertisement quantity information, and advertisement display information; The video generation module 330 is configured to generate a video with the target advertisement embedded in it based on the generation requirement information and the embedding decision information.
[0060] In one possible implementation, the information determination module 320 is further configured to determine the embedding decision information of the target advertisement based on the generated demand information and user preference information, wherein the user preference information includes the acceptance information of different users in historical videos regarding embedding carrier information, number of advertisements, and advertisement display information.
[0061] In one possible implementation, the information determination module 320 is further configured to determine extended information of the target advertisement based on the generated demand information, the extended information including at least one of scene-related extended information, person-related extended information, and object-related extended information; Based on the extended information and the user preference information, the embedding decision information for the target advertisement is determined.
[0062] In one possible implementation, the video generation module 330 is further configured to determine supplementary information to the generation requirement information based on the generation requirement information. Based on the generation requirement information, the supplementary information, and the embedding decision information, the video generation instruction is determined; The video containing the target advertisement is generated based on the video generation instructions and the AIGC model.
[0063] In one possible implementation, the video generation module 330 is further configured to determine multiple video generation schemes with different ad display levels based on the video generation instructions and the AIGC model. In response to the instruction to determine the target video generation scheme, a video embedded with the target advertisement is generated, and the display degree of the target advertisement is consistent with the advertisement display degree corresponding to the target video generation scheme.
[0064] In one possible implementation, the video generation apparatus further includes a benefit distribution module configured to distribute the incentive benefits corresponding to the video to an account in response to the generation of the video.
[0065] In one possible implementation, the video generation module 330 is further configured to obtain dimensional information corresponding to the generation requirement information, the dimensional information including at least one of scene dimensional information, character dimensional information and object dimensional information; Based on the dimensional information, the generated requirement information is supplemented to obtain the supplementary information.
[0066] Figure 3 This is a block diagram illustrating an electronic device 400 according to an exemplary embodiment. Figure 3 As shown, the electronic device 400 may include a processor 401 and a memory 402. The electronic device 400 may also include one or more of a multimedia component 403, an input / output (I / O) interface 404, and a communication component 405.
[0067] The processor 401 controls the overall operation of the electronic device 400 to complete all or part of the steps in the video generation method described above. The memory 402 stores various types of data to support the operation of the electronic device 400. This data may include, for example, instructions for any application or method operating on the electronic device 400, and application-related data such as contact data, sent and received messages, pictures, audio, video, etc. The memory 402 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The multimedia component 403 may include a screen and audio components. The screen may be, for example, a touchscreen, and the audio component is used to output and / or input audio signals. For example, the audio component may include a microphone for receiving external audio signals. The received audio signals may be further stored in memory 402 or transmitted via communication component 405. The audio component also includes at least one speaker for outputting audio signals. I / O interface 404 provides an interface between processor 401 and other interface modules, such as a keyboard, mouse, buttons, etc. These buttons may be virtual or physical buttons. Communication component 405 is used for wired or wireless communication between the electronic device 400 and other devices. Wireless communication may include Wi-Fi, Bluetooth, Near Field Communication (NFC), 2G, 3G, or 4G, or a combination thereof; therefore, the corresponding communication component 405 may include a Wi-Fi module, a Bluetooth module, or an NFC module.
[0068] In an exemplary embodiment, the electronic device 400 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the video generation method described above.
[0069] In another exemplary embodiment, a computer-readable storage medium including program instructions is also provided, which, when executed by a processor, implement the steps of the video generation method described above. For example, the computer-readable storage medium may be the memory 402 including the program instructions described above, which may be executed by the processor 401 of the electronic device 400 to complete the video generation method described above.
[0070] In another exemplary embodiment, a computer program product is also provided, which includes a computer program executable by a processor, which, when executed by the processor, implements the steps of the video generation method described above.
[0071] The preferred embodiments of this disclosure have been described in detail above with reference to the accompanying drawings. However, this disclosure is not limited to the specific details of the above embodiments. Within the scope of the technical concept of this disclosure, various simple modifications can be made to the technical solutions of this disclosure, and these simple modifications all fall within the protection scope of this disclosure.
[0072] It should also be noted that the various specific technical features described in the above embodiments can be combined in any suitable manner without contradiction. To avoid unnecessary repetition, this disclosure will not describe the various possible combinations separately.
[0073] Furthermore, various different embodiments of this disclosure can be combined in any way, as long as they do not violate the spirit of this disclosure, they should also be regarded as the content disclosed in this disclosure.
Claims
1. A video generation method, characterized in that, The method includes: Obtain the generation requirement information; Based on the generated demand information, the embedding decision information of the target advertisement is determined, wherein the embedding decision information includes at least one of the following: embedding carrier information, advertisement quantity information, and advertisement display information; Based on the generated demand information and the embedded decision information, a video with the target advertisement embedded is generated.
2. The video generation method according to claim 1, characterized in that, The step of determining the embedding decision information of the target advertisement based on the generated demand information includes: Based on the generated demand information and user preference information, the embedding decision information of the target advertisement is determined. The user preference information includes the acceptance information of different users in historical videos regarding embedding carrier information, number of advertisements, and advertisement display information.
3. The video generation method according to claim 2, characterized in that, The step of determining the embedding decision information of the target advertisement based on the generated demand information and user preference information includes: Based on the generated demand information, the extended information of the target advertisement is determined, and the extended information includes at least one of scene-related extended information, person-related extended information, and object-related extended information; Based on the extended information and the user preference information, the embedding decision information for the target advertisement is determined.
4. The video generation method according to claim 1, characterized in that, The step of generating a video embedded with the target advertisement based on the generated demand information and the embedded decision information includes: Based on the generated requirement information, determine the supplementary information for the generated requirement information; Based on the generation requirement information, the supplementary information, and the embedding decision information, the video generation instruction is determined; The video containing the target advertisement is generated based on the video generation instructions and the AIGC model.
5. The video generation method according to claim 4, characterized in that, The step of generating the video embedded with the target advertisement according to the video generation instructions and the AIGC model includes: Based on the video generation instructions and the AIGC model, multiple video generation schemes with different ad display levels are determined; In response to the instruction to determine the target video generation scheme, a video embedded with the target advertisement is generated, and the display degree of the target advertisement is consistent with the advertisement display degree corresponding to the target video generation scheme.
6. The video generation method according to claim 5, characterized in that, Each display resolution corresponds to a different incentive benefit. After generating the video embedding the target advertisement according to the video generation instructions and the AIGC model, the method further includes: In response to the generation of the video, the corresponding incentive benefits for the video are distributed to the account.
7. The video generation method according to claim 4, characterized in that, The step of determining supplementary information for the generated requirement information based on the generated requirement information includes: Obtain the dimension information corresponding to the generated requirement information, wherein the dimension information includes at least one of scene dimension information, character dimension information and object dimension information; Based on the dimensional information, the generated requirement information is supplemented to obtain the supplementary information.
8. A video generation apparatus, characterized in that, The device includes: The information acquisition module is configured to acquire the generation requirement information; The information determination module is configured to determine the embedding decision information of the target advertisement based on the generated demand information, wherein the embedding decision information includes at least one of embedding carrier information, advertisement quantity information, and advertisement display information; The video generation module is configured to generate a video with the target advertisement embedded in it, based on the generation requirement information and the embedding decision information.
9. An electronic device, characterized in that, include: A memory on which computer programs are stored; A processor for executing the computer program in the memory to implement the steps of the method according to any one of claims 1-7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the method described in any one of claims 1-7.
11. A computer program product, comprising a computer program, characterized in that, When executed by a processor, the computer program implements the steps of the method described in any one of claims 1-7.