Advertisement generation method, device and equipment

By using an artificial intelligence model to identify preset events in the video stream in real time, extracting visual element data and integrating it with advertising information, the problem of the disconnect between advertising and video content in existing technologies is solved, achieving seamless integration of advertising and video content, and improving the advertising generation effect and audience experience.

CN121728323APending Publication Date: 2026-03-24MIGU VIDEO TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-24
Publication Date
2026-03-24

AI Technical Summary

Technical Problem

Existing technologies for generating ad animations lack the ability to deeply perceive and analyze video content in real time, making it impossible to dynamically generate animated ads that are highly relevant to video events. This results in ads becoming disconnected from video content, affecting ad conversion rates and viewer experience.

Method used

By using an artificial intelligence model to identify preset event types in the video stream in real time, extracting visual element data such as motion trajectories and human action outlines, and integrating them with preset advertising information, the target advertising content is generated and rendered in real time to the set position in the video stream, achieving seamless integration of advertising and video content.

Benefits of technology

It improved the relevance of advertising information to video content, enhanced the fit and visual appeal of advertisements with event content, and improved the effectiveness of ad generation and the viewing experience for viewers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121728323A_ABST
    Figure CN121728323A_ABST
Patent Text Reader

Abstract

The invention provides an advertisement generation method, device and equipment, and relates to the technical field of computers, and the method comprises the steps: receiving a real-time video stream; identifying a preset event type in the real-time video stream through an artificial intelligence model; wherein the preset event type comprises a preset action type; extracting visual element data of the preset event type, and performing fusion processing on the visual element data and preset advertisement information through a preset fusion mode to generate target advertisement content; rendering the target advertisement content to a set position of the real-time video stream in real time, and generating an advertisement real-time video stream; and outputting the advertisement real-time video stream. The association degree of advertisement information and video streams can be improved, and the advertisement generation effect and quality are improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of computer, and particularly relates to an advertisement generation method, device and equipment. BACKGROUND

[0002] In the related art, the generation of advertisement motion effects is usually achieved by preset template mode or adaptation mode based on simple rule driving. For example, it can be achieved by static picture superposition, fixed animation playing, time axis triggering, etc. This mode is usually based on program arrangement nodes, operation personnel delivery or video playing timeline as triggering conditions, and the generated motion effect has poor effect. SUMMARY

[0003] The present disclosure provides an advertisement generation method, device and equipment.

[0004] According to a first aspect of the present disclosure, an advertisement generation method is provided, comprising: receiving a real-time video stream; identifying a preset event type in the real-time video stream by an artificial intelligence model; wherein the preset event type comprises a preset action type; extracting visual element data of the preset event type, and performing fusion processing on the visual element data and preset advertisement information by a preset fusion mode to generate target advertisement content; real-time rendering the target advertisement content to a set position of the real-time video stream to generate an advertisement real-time video stream; outputting the advertisement real-time video stream.

[0005] According to a second aspect of the present disclosure, an advertisement generation device is provided, comprising: a receiving module configured to receive a real-time video stream; an identifying module configured to identify a preset event type in the real-time video stream by an artificial intelligence model; wherein the preset event type comprises a preset action type; a fusion module configured to extract visual element data of the preset event type, and perform fusion processing on the visual element data and preset advertisement information by a preset fusion mode to generate target advertisement content; a rendering module configured to real-time render the target advertisement content to a set position of the real-time video stream to generate an advertisement real-time video stream; an outputting module configured to output the advertisement real-time video stream.

[0006] According to a third aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory connected with the at least one processor in communication; wherein, The memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method of the first aspect.

[0007] According to a fourth aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to perform the method of the first aspect.

[0008] According to a fifth aspect of the present disclosure, there is provided a computer program product comprising a computer program which, when executed by a processor, implements the method of the first aspect.

[0009] In an embodiment of the present disclosure, by receiving a real-time video stream; identifying a preset event type in the real-time video stream through an artificial intelligence model; wherein the preset event type includes a preset action type; extracting visual element data of the preset event type, and performing fusion processing on the visual element data and preset advertising information through a preset fusion manner to generate target advertising content; real-time rendering the target advertising content to a set position of the real-time video stream to generate an advertising real-time video stream; and outputting the advertising real-time video stream. In this way, by receiving a real-time video stream and utilizing an artificial intelligence model to identify a preset event type, visual element data related to the event can be accurately extracted; then by a preset fusion manner, advertising information and visual element data are fused to generate target advertising content that is naturally integrated into video content; finally, the target advertising content is real-time rendered to a set position of the video stream and output. In this way, seamless combination of advertising and video content can be achieved, thereby not only improving the correlation between advertising information and video content, achieving high matching and high visual attraction of dynamic advertising fusion effect of advertising and event content, improving the correlation between advertising information and video stream, thereby improving the advertising generation effect and quality, improving the display effect of the advertisement, and enhancing the viewing experience of the audience; and the display mode of the advertisement can be dynamically adjusted according to the video content, thereby improving the targeting and attractiveness of the advertisement.

[0010] It should be understood that the content described in this part is not intended to identify key or important features of the embodiments of the present disclosure, nor to limit the scope of the present disclosure. Other features of the present disclosure will become apparent from the following description. BRIEF DESCRIPTION OF DRAWINGS

[0011] The accompanying drawings are used to better understand the present scheme and do not limit the present disclosure. Among them: Figure 1 A flowchart of an advertising generation method according to an embodiment of the present disclosure is shown; Figure 2A schematic diagram of cross resonance provided for an embodiment of the present disclosure; Figure 3 A schematic diagram of event trajectory conversion into a structured path feature tree provided for an embodiment of the present disclosure; Figure 4 A structural schematic diagram of an advertisement generation device provided for an embodiment of the present disclosure. DETAILED DESCRIPTION

[0012] Exemplary embodiments of the present disclosure are described below with reference to the accompanying drawings, which include various details of the embodiments of the present disclosure to assist in understanding, and should be considered as merely exemplary. Thus, those of ordinary skill in the art will realize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Also, for the sake of brevity and clarity, descriptions of well-known functions and constructions are omitted from the following description.

[0013] As known from the background, advertisement motion effect generation usually adopts a preset template mode or an adaptive mode based on simple rule driving, for example, advertisement display is realized through static picture superposition, fixed animation playing, time axis triggering, etc. This mode usually takes program arrangement nodes, operation manual delivery or video playing timeline as the triggering basis, lacks deep perception and real-time analysis capability for video event (for example, live event such as sports live event) content, and cannot dynamically generate motion effect advertisements with strong correlation, thereby improving advertisement conversion effect. Moreover, in the process of sports event live broadcast, advertisement delivery mainly takes the form of edge display, video patch or intermission push, etc. This mode often causes emotional fragmentation to the audience at key competition nodes, and cannot realize natural integration of advertisement information and event emotion. Furthermore, part of the advertisement delivery system usually relies on large language models or semantic vector matching technology to realize the correspondence between on-site event semantics and brand advertisement semantics, but this method at least has the following questions: it depends on pre-training semantic matching and does not have immediate response to on-site emotional rhythm and audience momentum, that is, the pre-delivery mode of advertisement cannot respond to sudden key events (such as goals, celebrations, conflicts, etc.) in sports live broadcast and match appropriate advertisement motion effects, which will cause event response lag; advertisement presentation still mainly takes the form of content splicing, lacks the sense of fusion with event action trajectory and dynamic rhythm, that is, lacks multi-dimensional understanding of sports live event (such as character appearance, event rewards and punishments, celebration action, etc.), the content correlation between advertisement content and video content is weak, the advertisement effect is poor, which leads to lack of emotional extension of advertisement motion effect generation, poor interactivity, and stiff advertisement pop-up window easily causes audience aversion; it cannot realize natural emergence or emotional resonance of brand image in the motion trajectory, and has a sense of abruptness; moreover, the motion effect generation lacks intelligence, that is, the design of motion effect parameters relies on manual or rule pre-configuration, and lacks intelligent generation mechanism.

[0014] Based on this, the embodiment of the disclosure provides an advertisement generation method, device and equipment, which can recognize a key event (i.e., a preset event) of a competition in real time by using an artificial intelligence model, extract a motion trajectory (a motion object trajectory), a person action contour, a video segment and other multi-modal features, and dynamically match the multi-modal features with advertisement information in a brand style envelope space, drive a brand particle unit to generate a visualized advertisement motion effect, and embed the advertisement motion effect in a live video picture in real time in a plurality of fusion modes, so that a dynamic advertisement fusion effect with high compatibility and high visual attraction of the advertisement and the competition content can be achieved, the association degree of the advertisement information and the video stream is improved, and the advertisement generation effect and quality are improved.

[0015] The advertisement generation method, device and equipment of the embodiment of the disclosure are described below with reference to the accompanying drawings.

[0016] Figure 1 A flowchart of an advertisement generation method provided by the embodiment of the disclosure is shown in FIG. 1. As shown in FIG. 1, the advertisement generation method includes the following steps: Figure 1 Step 101, receiving a real-time video stream.

[0017] In the embodiment of the disclosure, the real-time video stream can be obtained by a video acquisition device (such as a camera, a broadcast vehicle, etc.). For example, the real-time video stream can be a live broadcast of a sports competition or a video signal of a concert live broadcast. The reception of the real-time video stream is the basis of the entire advertisement generation method, and subsequent steps can be based on the real-time video stream for analysis and processing.

[0018] Step 102, identifying a preset event type in the real-time video stream by using an artificial intelligence model.

[0019] The preset event type includes a preset action type.

[0020] In the embodiment of the disclosure, after receiving the real-time video stream, the real-time video stream can be analyzed by using a pre-trained artificial intelligence model to identify the preset event type in the real-time video stream. For example, the preset event type can include but is not limited to a preset action type, such as a shot of an athlete, a goal, a foul, etc. The artificial intelligence model can be a deep learning model, such as a CNN (Convolutional Neural Network), which can be trained by a large amount of labeled data to accurately identify a specific event type in a video.

[0021] Step 103, extracting visual element data of the preset event type, and fusing the visual element data and preset advertisement information by using a preset fusion mode to generate target advertisement content.

[0022] ​In the embodiments of the present disclosure, after identifying the preset event type, visual element data related to the event type can be further extracted. For example, the visual element data can include a trajectory of a moving object, a motion contour of a person, a video clip of a person, etc. Then, the visual element data can be fused with the preset advertisement information through a preset fusion manner. The preset fusion manner can be deforming the trajectory of the moving object into a preset advertisement logo, integrating the preset advertisement logo into the motion contour of the person, superimposing the preset advertisement information associated with the semantic information of the video clip in the video clip of the person, or generating a virtual image of the person and making the motion trajectory of the virtual image evolve into the preset advertisement logo, etc. Through the fusion processing, the advertisement information can be naturally integrated into the video stream, so that the advertisement information is coordinated with the video content, and the acceptance and effect of the advertisement are improved.

[0023] In step 104, the target advertisement content is rendered to a set position of the real-time video stream in real time to generate an advertisement real-time video stream.

[0024] In the embodiments of the present disclosure, the target advertisement content generated through the fusion processing can be rendered to the set position of the real-time video stream in real time. For example, the set position can be a specific area in the video picture, such as a corner, a frame, etc., or a dynamically determined position according to the video content. Through real-time rendering, the target advertisement content can be synchronously displayed with the video stream, and there is no delay or misalignment.

[0025] In step 105, the advertisement real-time video stream is output.

[0026] In the embodiments of the present disclosure, after generating the advertisement real-time video stream, the generated advertisement real-time video stream can be output, for example, output to a display device or transmitted to other platforms for the audience to watch. It can be understood that the output advertisement real-time video stream can be a video signal transmitted through television broadcasting, network live streaming, etc., or a video file stored in a storage medium. In this way, the audience can receive the advertisement information integrated with the video content while watching the real-time video content.

[0027] In the embodiments of the present disclosure, a real-time video stream is received; a preset event type in the real-time video stream is identified by an artificial intelligence model; wherein the preset event type includes a preset action type; visual element data of the preset event type is extracted, and the visual element data and preset advertisement information are fused by a preset fusion manner to generate target advertisement content; the target advertisement content is rendered to a set position of the real-time video stream in real time to generate an advertisement real-time video stream; and the advertisement real-time video stream is output. In this way, the real-time video stream is received, and the preset event type is identified by the artificial intelligence model to accurately extract visual element data related to the event; then the advertisement information and the visual element data are fused by the preset fusion manner to generate target advertisement content naturally integrated into the video content; finally, the target advertisement content is rendered to the set position of the video stream in real time and output. In this way, seamless combination of the advertisement and the video content can be realized, so as to not only improve the correlation between the advertisement information and the video content, achieve high matching and high visual attraction of the dynamic advertisement fusion effect of the advertisement and the event content, improve the correlation between the advertisement information and the video stream, thereby improving the advertisement generation effect and quality, improving the display effect of the advertisement, and enhancing the viewing experience of the audience; but also dynamically adjust the display mode of the advertisement according to the video content, improve the pertinence and attractiveness of the advertisement.

[0028] In some possible implementations, the visual element data of the preset event type is extracted, including: extracting at least one of a moving object trajectory, a motion contour of a person, and a video clip of the person of the preset event type; The preset fusion manner includes at least one of: deforming the moving object trajectory into a preset advertisement logo; integrating the preset advertisement logo into the motion contour of the person; superimposing preset advertisement information associated with semantic information of the video clip in the video clip of the person; wherein the preset advertisement information includes preset advertisement text content; generating a virtual image of the person, and evolving a motion trajectory of the virtual image into the preset advertisement logo.

[0029] In embodiments of the present disclosure, the extracted visual element data includes: extracting a moving object trajectory, for example, the moving trajectory of a moving object (such as a football, a basketball, etc.) in a video stream can be extracted by computer vision technology, such as object detection and tracking algorithms, for example, the moving trajectory can be represented as a series of coordinate points for subsequent fusion processing; extracting the action contour of the person, for example, the key points (such as head, shoulder, hand, etc.) of the person can be detected by using a deep learning model, and the action contour of the person is generated based on this, which can be used to integrate the advertisement logo based on the action of the person; extracting a video clip of a person, for example, a video clip containing a person can be cut from a real-time video stream, which can be used for subsequent semantic analysis and advertisement information superimposition. After the visual element data is extracted, the data can be fused with the preset advertisement information through a preset fusion manner to generate target advertisement content.

[0030] In further possible implementations, deforming the moving object trajectory into a preset advertisement logo includes: extracting a moving trajectory coordinate sequence of the moving object; controlling a preset tail element to generate a dynamic tail along the moving trajectory coordinate sequence according to the moving trajectory coordinate sequence; adjusting the morphological parameters of the dynamic tail to gradually change the dynamic tail into the preset advertisement logo.

[0031] In an embodiment of the present disclosure, when deforming the trajectory of a moving object into a preset advertisement logo, the coordinate sequence of the trajectory of the moving object (such as a football, a basketball, etc.) in the real-time video stream can be extracted first. For example, a target detection algorithm (such as YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), etc.) can be used to analyze each frame in the real-time video stream to detect the position of the moving object, and a target tracking algorithm (such as SORT (Simple Online and Realtime Tracking), DeepSORT (Deep Simple Online and Realtime Tracking), etc.) can be used to track the object to generate the trajectory. Then, the dynamic tailing effect can be generated according to the coordinate sequence of the trajectory. For example, the initial shape and attributes of the preset tailing element can be defined, which can be a particle system, a flame effect, a halo, etc., and the specific shape can be set according to the design requirements of the advertisement. According to the coordinate sequence of the trajectory, the dynamic effect of the tailing element along the coordinate sequence of the trajectory can be controlled, for example, in each frame, the position and shape of the tailing element can be dynamically adjusted according to the position and speed of the object, for example, if the object moves faster, the tailing can be longer and sparser; if the object moves slower, the tailing can be shorter and denser. After the dynamic tailing is generated, the shape parameters of the tailing can be further adjusted to gradually deform into the preset advertisement logo. For example, the shape parameters of the tailing, such as length, width, transparency, color, etc., can be defined, which can be dynamically adjusted by a mathematical function (such as a Bezier curve) to achieve smooth shape changes; a gradual change algorithm (such as an interpolation algorithm) can be used to gradually adjust the shape of the dynamic tailing to the outline of the target advertisement logo, for example, by adjusting the shape and color of the tailing to gradually form a certain brand logo. It can be understood that during the gradual change process, the similarity between the tailing and the target advertisement logo can be monitored in real time, and the shape parameters can be adjusted according to the similarity feedback to ensure the accuracy and naturalness of the gradual change effect. For example, if the similarity between the tailing and the target logo is lower than a preset threshold, the shape parameters can be dynamically adjusted to speed up or optimize the gradual change process. In this way, the trajectory of the moving object can be dynamically deformed into the preset advertisement logo, the natural fusion of the advertisement content and the video content can be achieved, the advertisement generation effect and the visual appeal can be further improved, and the audience experience can be improved.

[0032] In a further possible implementation, the action contour of the character is deformed into the preset advertisement logo, including: performing UV unwrapping processing on the action contour of the virtual character to generate a mapping relationship between the action contour of the character and the UV coordinates; The preset advertisement logo is taken as a texture map and mapped to a preset region of the motion contour of the character based on a mapping relationship; the preset region is located in the motion contour of the character or on the surface of the motion contour of the character.

[0033] In the embodiments of the present disclosure, when the preset advertisement logo is integrated into the motion contour of the character, the motion contour of the character can be subjected to UV unfolding processing first. For example, a three-dimensional model of the character can be extracted from a real-time video stream by using a three-dimensional modeling technology, including appearance features and motion features of the character, to accurately reflect the posture and motion of the character in the video; then the three-dimensional model of the character extracted is subjected to UV unfolding processing (UV unfolding is a technology of mapping the surface of a three-dimensional model to a two-dimensional plane, which can generate a mapping relationship between the surface of the model and UV coordinates by calculating the texture coordinates of the surface of the model), and a mapping relationship between the motion contour of the character and UV coordinates can be generated through the UV unfolding processing. Then, the preset advertisement logo can be taken as a texture map and mapped to a preset region of the motion contour of the character. For example, the preset advertisement logo (such as a brand logo, an advertisement slogan, etc.) can be designed as a texture map, wherein the texture map can be a two-dimensional image file (such as a PNG (Portable Network Graphics) or JPEG (Joint Photographic Experts Group) format), and the resolution and size of the texture map can be adjusted according to the motion contour of the character and advertisement design requirements; then, according to the advertisement design requirements, a preset region of the motion contour of the character is selected, which can be located in the motion contour of the character (such as the back, the chest, etc.) or on the surface of the motion contour of the character (such as the arms, the legs, etc.); then, based on the mapping relationship between the motion contour of the character and UV coordinates, the preset advertisement logo is taken as a texture map and mapped to the preset region of the motion contour of the character, for example, the texture coordinates of the advertisement logo can be aligned with the UV coordinates of the motion contour of the character by using a texture mapping technology, to ensure that the advertisement logo can be naturally integrated into the motion contour of the character. It can be understood that in the texture mapping process, the transparency, color, size, etc. of the advertisement logo can be adjusted to achieve better visual effects. In this way, the preset advertisement logo can be naturally integrated into the motion contour of the character through UV unfolding and texture mapping, and seamless combination of the advertisement content and the video content can be achieved, so as to further enhance the visual appeal and display effect of the advertisement and improve the pertinence and appeal of the advertisement.

[0034] In further possible embodiments, preset advertisement information associated with semantic information of a video segment is superimposed in the video segment, including: extracting semantic information of the video segment; determining, according to the semantic information, preset advertisement information associated with the semantic information from a preset advertisement information set; The preset advertisement information associated with the semantic information is superimposed into the video clip by a preset special effect.

[0035] In the embodiments of the present disclosure, when superimposing the preset advertisement information associated with the semantic information of the video clip in the video clip of the person, the semantic information can be extracted from the video clip of the person. For example, the video clip containing the person can be intercepted from the real-time video stream, and the video clip can be a fixed-length clip or a dynamically generated clip according to a specific event; the video clip is analyzed semantically, for example, the video clip can be analyzed semantically by using NLP (Natural Language Processing) and computer vision technology, and the semantic information of the video clip is obtained, for example, the key objects (such as people, balls, goals, etc.) in the video clip can be identified by using a target detection algorithm (such as YOLO, SSD, etc.), the action (such as shooting, celebration, foul, etc.) of the person can be analyzed by using an action recognition algorithm (such as OpenPose, 3D action recognition model), and then the emotional tendency (such as excitement, tension, disappointment, etc.) of the video clip can be judged by analyzing the expression and action of the person, combining with the audio information (such as the cheers of the audience, the tone of the commentator, etc.), and further understanding the semantic information of the video clip by combining the scene background (such as the stadium, the audience stand, etc.) in the video clip.

[0036] Then, after extracting the semantic information of the video clip, a preset advertisement information associated with the semantic information can be selected from a preset advertisement information set. For example, the preset advertisement information set is a preset advertisement information set containing various advertisement text contents, images, video clips, etc., which can be customized according to brand requirements and advertising strategies. According to the extracted semantic information, a semantic matching algorithm (such as cosine similarity, vector similarity, etc.) can be used to select the preset advertisement information with the highest association degree with the semantic information of the video clip from the preset advertisement information set. For example, if the person in the video clip is celebrating a goal, an advertisement text content related to victory and celebration can be selected. After determining the preset advertisement information associated with the semantic information, the advertisement information can be superimposed into the video clip through a preset special effect. For example, according to the type of advertisement information and the semantic information of the video clip, a suitable special effect can be selected or designed, which can include text animation, image fusion, particle effect, etc., to enhance the visual effect of the advertisement information; and according to the semantic information and visual focus of the video clip, a suitable superimposition position can be selected, for example, if the person in the video clip is celebrating a goal, the advertisement information can be superimposed around the person or in the background to enhance the visual effect. Finally, the designed advertisement information can be rendered into the video clip in real time through the preset special effect, ensuring the natural integration of the advertisement information and the video content. During the rendering process, the transparency, color, size, etc. of the advertisement information can be adjusted to achieve better visual effects. In this way, the preset advertisement information associated with the semantic information of the video clip can be naturally superimposed into the video clip, realizing the seamless combination of advertisement content and video content; and through semantic analysis and matching algorithm, the semantic consistency of the advertisement information and the video clip can be ensured, enhancing the targeting and attractiveness of the advertisement; and through the design of special effects and superimposition position, the visual attractiveness of the advertisement information can be further enhanced, further improving the advertising effect.

[0037] It can be understood that after selecting the preset advertisement information associated with the semantic information from the preset advertisement information set, the association degree of the selected advertisement information with the video clip can be evaluated, and if the association degree is lower than a preset original threshold, the selection strategy can be further adjusted to reselect the preset advertisement information associated with the semantic information, to ensure the semantic consistency of the advertisement information and the video content and optimize the matching effect of the advertisement information.

[0038] In a further possible implementation, a virtual image of the person is generated, and a motion trajectory of the virtual image is evolved into a preset advertisement logo, comprising: obtaining a person feature of the person in the real-time video stream; wherein the person feature includes an appearance feature and a motion feature; generating a virtual image of the person based on the appearance feature; According to the preset advertisement identifier and the action feature, the action track of the virtual image is evolved into the preset advertisement identifier.

[0039] In the embodiments of the present disclosure, when generating a virtual image of a person and evolving the action track of the virtual image into a preset advertisement identifier, the appearance features and action features of the person can be obtained from the real-time video stream first. For example, the real-time video stream can be analyzed by using computer vision technology to extract the relevant features of the person. For example, the appearance features of the person, such as facial features, clothing color, hairstyle, etc., can be extracted by using image processing and deep learning algorithms. The action features of the person, such as joint positions, action tracks, speed, acceleration, etc., can be extracted by using motion capture technology. Then, the virtual image of the person can be generated based on the appearance features. For example, a three-dimensional model of the person can be generated according to the extracted appearance features by using three-dimensional modeling software or algorithms. The virtual image of the person can be generated by mapping the extracted appearance features (such as facial texture, clothing color, etc.) to the surface of the three-dimensional model. Alternatively, the generated three-dimensional model can be optimized, including smoothing processing, detail enhancement, etc., to improve the realism and visual effect of the virtual image of the person. After the virtual image of the person is generated, the action track of the virtual image can be evolved into the preset advertisement identifier according to the preset advertisement identifier and the action feature. For example, the preset advertisement identifier (such as a brand logo, an advertisement slogan, etc.) can be designed as a two-dimensional or three-dimensional model. The action track of the virtual image can be planned according to the extracted action features. For example, the action track can be the movement path of the person or a specific action sequence (such as hand gestures, postures, etc.). The action track of the virtual image can be gradually evolved into the shape of the preset advertisement identifier by using algorithms. For example, if the preset advertisement identifier is a brand logo, the hand gestures or postures of the virtual image can be adjusted to gradually form the shape of the logo. Finally, the evolved virtual image can be rendered in real time into the video stream to ensure that the action track of the virtual image is naturally integrated with the shape of the preset advertisement identifier. It can be understood that during the rendering process, the transparency, color, size, etc. of the virtual image can be adjusted to achieve the best visual effect. In this way, the virtual image of the person can be generated, and the action track of the virtual image can be evolved into the preset advertisement identifier, achieving the natural integration of the advertisement content and the video content, thereby further improving the advertisement generation effect and quality.

[0040] In further possible implementations, visual element data of the preset event type is extracted, and the visual element data and the preset advertisement information are fused by a preset fusion manner to generate target advertisement content, including: extracting key action features of the preset event type; wherein the key action features include action path, speed, acceleration, duration, and visual focus of the screen; determine the preset event type mood curve according to the shot switching frequency, slow motion mark, audio track intensity, and color fluctuation; combine the key action features and the mood curve into an event vector trajectory; In the case where the event vector trajectory and the preset advertisement information envelope space exist a cross resonance point, extract the visual element data of the preset event type, and perform fusion processing on the visual element data and the preset advertisement information through a preset fusion manner to generate the target advertisement content.

[0041] In an embodiment of the present disclosure, in order to realize accurate advertisement fusion, the key action features of the preset event type can be extracted, including action path, speed, acceleration, duration, and picture visual focus, etc. After extracting the key action features, the mood curve of the preset event type can also be determined according to other information in the video stream. For example, the rhythm and dynamic change of the video can be determined by analyzing the shot switching frequency in the video stream; whether there is a slow motion mark in the video stream can be detected, and slow motion is usually used to emphasize key events such as goals, shots, etc.; the mood intensity of the video can be determined by analyzing the audio track intensity change; the mood atmosphere of the video can be determined by analyzing the color fluctuation of the video frame. By comprehensively analyzing the shot switching frequency, slow motion mark, audio track intensity, and color fluctuation, the mood curve of the preset event type can be generated, which can be represented as a time sequence, for example, to reflect the mood intensity and change trend of the event. Then, after determining the key action features and the mood curve, these features can be combined into an event vector trajectory. For example, the key action features (such as action path, speed, acceleration, duration, and picture visual focus) can be fused with the mood curve to generate an event vector trajectory, which can be represented as a multi-dimensional vector, each dimension corresponding to a feature. After generating the event vector trajectory, it can be judged whether the trajectory and the preset advertisement information envelope space exist a cross resonance point. For example, the preset advertisement information envelope space can be a multi-dimensional space, which can define the style, mood, and dynamic features of the advertisement, and the envelope space can be defined by features such as brand logo, advertisement text, color scheme, etc.; whether there is a cross resonance point can be detected by calculating the intersection of the event vector trajectory and the preset advertisement information envelope space, wherein the cross resonance point is used to indicate that the matching degree of the event vector trajectory and the brand advertisement features is high, which is the opportunity for advertisement fusion. In the case where the cross resonance point is detected, the visual element data of the preset event type can be extracted, and the visual element data and the preset advertisement information can be fused through a preset fusion manner to generate the target advertisement content. In this way, only when the cross resonance point exists, the advertisement generation processing is performed, which can further improve the advertisement generation efficiency.

[0042] In a further possible implementation, the target advertisement content is rendered to a set position of the real-time video stream in real time to generate an advertisement real-time video stream, including: In the case that there is no cross resonance point between the event vector trajectory and the preset advertising information envelope space, the preset advertising identifier is rendered by a preset rendering manner or projected on a non-interference area of the real-time video stream; the preset rendering manner includes transparency gradual rendering or boundary suspension manner rendering; the non-interference area includes a virtual lens frame, a real-time score floating layer or a corner particle scattering area.

[0043] In the embodiments of the present disclosure, when the target advertising content is rendered to the set position of the real-time video stream to generate the advertising real-time video stream, in the case that there is no cross resonance point between the event vector trajectory and the preset advertising information envelope space, that is, in the case that the matching degree between the current video content and the preset advertising information is low and the deep fusion is not suitable, the preset advertising identifier can be rendered to the non-interference area of the real-time video stream by the preset rendering manner. Exemplarily, the preset rendering manner can include transparency gradual rendering or boundary suspension manner rendering, wherein the transparency gradual rendering can be that the advertising identifier gradually changes from completely transparent to completely opaque, and the boundary suspension manner rendering can be that the advertising identifier is displayed in a suspended manner at the edge or a specific area of the video picture. The non-interference area can include a virtual lens frame, a real-time score floating layer or a corner particle scattering area, and the non-interference area can be an area that does not interfere with the viewing experience of the audience, which is usually located at the edge or corner of the video picture. When rendering, the preset rendering manner can be selected to render the preset advertising identifier to the non-interference area. During the rendering process, the transparency, position and size of the advertising identifier can be adjusted to ensure that the advertising identifier is clearly visible in the non-interference area and does not interfere with the viewing of the video content. In this way, the rendering manner and position can be flexibly selected according to the matching condition of the event vector trajectory and the preset advertising information envelope space, the display effect of the advertising content is ensured, the advertising identifier is rendered to the non-interference area in the case that there is no cross resonance point, the interference with the viewing experience of the audience is reduced, and the advertising generation effect and quality are further improved.

[0044] To make the advertising generation method provided by the embodiments of the present disclosure clearer, the following examples are combined for illustration.

[0045] The embodiment of the present disclosure aims at the problem that the advertising presentation mode of sports live broadcast is single, disconnected with the content of the event, and lacks real-time and emotional linkage, and provides an advertising generation method for converting key events in sports live broadcast into advertising motion effect driving signals. The advertising generation method can use an artificial intelligence model to identify key events of an event in real time, extract multi-modal features such as motion trajectories and emotional rhythms, and dynamically match them with a brand style envelope space to drive brand particle units to generate visualized advertising motion effects, and embed the advertising motion effects in a live broadcast picture in a multi-layer fusion manner to achieve a dynamic advertising fusion effect with high consistency between the advertising and the content of the event, emotional rendering capability, and visual attraction. For example, the advertising generation method includes the following steps: 1. constructing a brand envelope space, and finding a cross resonance point by fitting a live event event trajectory to form a personalized driving signal for advertising generation; 2. proposing a structured path feature tree to dynamically divide the event trajectory into visual style segments to drive the animation splicing of the brand particle module; 3. using a multi-level visual fusion strategy (action main path layer, background residual image layer, and rhythm response layer) to embed the advertising to ensure visual naturalness and low interference; and 4. providing a bottom mechanism to ensure consistency and rhythm of brand presentation when there is insufficient event signal. Based on the above, the advertising generation method provided by the embodiment of the present disclosure has the following technical advantages: real-time: the advertising and the key event are synchronously responsive, which can improve the attention capturing capability; personalization: the motion effect can be customized based on event details and brand semantics to achieve precise matching; fusion feeling: the brand motion effect is embedded in the event scene and rhythm, which can avoid a jarring visual effect; the brand exposure quality and interaction efficiency can be improved, and the user impression can be enhanced; and the advertising generation method is suitable for multiple scene applications such as sports live broadcast, general entertainment, and interactive short video.

[0046] As a specific example, the advertising generation method provided by the embodiment of the present disclosure includes the following steps: 1. receiving a live event video stream (the live event video stream is a real-time video stream).

[0047] 2. using an artificial intelligence model to analyze the live event video stream in real time, detecting and identifying event types (i.e., preset motion types) of the event, the event types including shooting, entering the field, scoring, fouling, substitution, specific difficult motion, player celebration motion, or slow-motion replay trigger point.

[0048] The specific difficult motion includes hook shot, fish jump, over-the-head shot, acrobatic dribble, Marseille spin, scorpion tail, etc.; and the player celebration motion includes kneeling and sliding, sliding and jumping, running and pointing to the sky, sliding and kneeling, kneeling and praying, and rolling and celebrating, etc.

[0049] 3. generating emotional advertising content associated with the specific event type and fused with brand elements according to the identified specific event type.

[0050] The generation mode of the emotional advertising content can be to fuse the detected visual elements or emotional connotations of the event with brand elements, the visual elements including at least one of the following: trajectory of a moving object, action contour of an athlete (i.e., a person), historical image segment of the athlete (i.e., a video segment of the person). The fusion mode includes at least one of the following: dynamically deforming the moving object trajectory into a brand logo, integrating the brand logo in or on the action contour of the athlete, superimposing a brand slogan associated with the image semantics on the historical image segment of the athlete, generating a virtual image of the athlete and evolving the action trajectory of the virtual image into the brand logo. Examples can be seen in Table 1.

[0051] Table 1

[0052] The implementation idea based on the fusion of the moving object trajectory + the action contour of the athlete + the brand logo is described.

[0053] The specific implementation steps of dynamically deforming the moving object trajectory into the brand logo can include: extracting the moving trajectory coordinate sequence of the moving object; and controlling an advertising motion effect system to simulate a flame tailing special effect according to the moving trajectory coordinate sequence, and dynamically adjusting advertising motion effect attributes to gradually change the tailing shape into the contour of the brand logo.

[0054] The specific implementation steps of integrating the brand logo in or on the action contour of the athlete can include: performing UV unfolding on the action contour model of the athlete, and mapping the brand logo as a texture map to a preset region on the model surface.

[0055] (1) Extract the key action features of the event, including: Action path (such as linear, rotation, diffusion); speed / acceleration; duration; visual focus of the picture (such as medium shot, close-up), etc.

[0056] (2) Determine the emotional curve position of the event (such as: low segment, burst segment, climax segment, contrast segment) according to the information such as the frequency of lens switching in the live event, whether the slow motion is started, the peak value of the audience audio track intensity, the color saturation / contrast fluctuation, the player skeleton change or the trajectory key point sequence, etc.

[0057] (3) Determine the visual angle tension factor of the event (such as close-up, picture focus aggregation degree).

[0058] (4) Extract the brand style features, including: advertising motion tone (such as linear, arc, vortex, etc.); brand emotion target (such as burst, warmth, stability, inspiration); visual style density (such as simple / cool / deep).

[0059] (5) The features of (1), (2), and (3) above are combined into an event vector trajectory according to a time frame / event sequence, as shown in the following example: For example, Event_Vector_t = [ Action type (one-hot), Speed value, Acceleration value, Duration, Focus feature, Lens switching frequency, Slow motion mark, Audio track intensity, Color fluctuation, Emotion stage label (one-hot), Angle of view focus, Close-up intensity ] (6) Project the event vector trajectory into the brand envelope space to find the cross resonance point.

[0060] The event vector trajectory can be fitted with a spline curve; the brand envelope is projected in segments according to the style broken line; if the event trajectory and the brand envelope have a return overlap in any dimension, it is considered to have resonance interference, and the cross resonance point is found, which is used as a driving signal for subsequent advertisement generation. That is, if there is a cross resonance point, the visual element data and the preset advertisement information are fused by the preset fusion method to generate target advertisement content; the target advertisement content is rendered in real time to the set position of the real-time video stream to generate an advertisement real-time video stream; and the process of outputting the advertisement real-time video stream is performed. See Figure 2 , Figure 2 It is shown how the "event trajectory (such as the shooting path)" is fitted with a spline curve and compared with the "brand envelope space (such as the Swoosh shape of brand A)", and the area (red shadow) of "cross resonance" in visual style is generated, providing a trigger point and style guide for advertisement motion effect generation.

[0061] Among them, the brand envelope space is a potential boundary and expression style range of brand style in a multi-dimensional space, which can reflect the tonality and distinguishability of brand advertising in visual, emotion, and motion.

[0062] (7) Convert the event trajectory into a structured path feature tree, dynamically arrange brand particle elements based on its structure, and generate symbolic animation effects in real time to realize the binding of brand and event physical trajectories.

[0063] As can be extracted, the trajectory point set T = {P0, P1,..., Pn} of the event; the path feature tree is generated using the rhythm of the direction change in the trajectory; each node of the path feature tree defines a visual style segment, and the entire path is divided into several visual style segments, each of which is labeled as: Straight smooth segment: suitable for LOGO emergence; Corner burst segment: suitable for particle burst; Rhythm echo segment: suitable for flowing ripple animation.

[0064] Referring to Figure 3 , Figure 3 The process of converting the event trajectory (point set T) into a structured path feature tree is shown, and three visual style segments are marked with different colors. Among them, the first color: trajectory smooth segment, suitable for the emergence animation of brand LOGO; the second color: direction mutation segment, suitable for triggering particle burst and other high-energy dynamic effects; the third color: rhythm reciprocating segment, suitable for generating ripples or flowing visual feedback.

[0065] Each trajectory point P0-Pn is a potential particle emission or dynamic effect control node, and the system can dynamically splice brand visual modules based on the path tree to realize the effect of natural emergence of brand image along the event trajectory.

[0066] (8) Call brand style particle element library: each brand can provide edge flow unit, color dispersion point array unit, and LOGO flow form basic shape; one or more particle units are bound to each path segment to realize real-time animation splicing.

[0067] The trajectory tree can construct animation flow in the form of plug-in special effect splicing during runtime; for example, a sliding path with 3 angle change points generates 3 different morphed LOGO emergence segments and tail scattering segments.

[0068] The edge flow unit function provides point line animation flowing along the path boundary, such as energy drink: red fire flow particles; footwear brand: speed wind stripe; The color dispersion point array unit function is to burst or aggregate into LOGO particles at a specific position, such as cola brand: bubble particles; Nike: interlaced point group to spell Swoosh shape.

[0069] The LOGO flow form basic shape function is a semi-transparent, vector, and residual image form LOGO contour structure; for example, contour line LOGO, feathering appearance, streamline appearance, etc.

[0070] In the embodiments of the present disclosure, the advertisement special effect is not a pre-set template, but an image structure dynamically compiled and generated by calling brand style particle base elements in real time driven by the event trajectory.

[0071] 4. The generated emotional advertising content (i.e., target advertising content) is rendered in real time and superimposed on the corresponding position (i.e., set position) of the live event video stream, forming an enhanced live video (i.e., advertising real-time video stream) output.

[0072] In this embodiment, instead of hard inserting the advertisement, multi-layer special effect nesting rendering is performed, and the advertisement is presented as a background stream form or visual afterimage, etc., for example: (1) Use low-order visual features (edge density, motion intensity, composition center, etc.) to generate foreground layers; map the resonance points identified in the front to the appearance nodes on the time axis; each node controls the start, gradual appearance, expansion or end of the advertisement animation; ensure that the advertisement starts with the action and hides with the action.

[0073] (2) Implement a multi-level visual fusion strategy: Main path layer (action track): LOGO grows or emerges along the track; Background afterimage layer: LOGO can slowly appear in the edge, grassland or audience seat blur area; Auxiliary rhythm layer: music beat synchronous particle jumps, forming sound and picture resonance; All animation transparency, color brightness, interframe speed, etc. are adjusted according to the emotion curve.

[0074] (3) Bottom mechanism handles special cases: If the path or resonance point is not enough to drive the complete animation: Trigger the default lightweight animation of the brand (such as LOGO floating appearance, boundary floating text, etc.); Or project the advertising information on the virtual lens frame, real-time score floating layer, corner particle scattering, etc. non-interference area.

[0075] Through the description of the above implementation mode, those skilled in the art can clearly understand that the method according to the above embodiment can be realized by means of software and necessary general hardware platform, of course, it can also be realized by hardware, but in many cases the former is a better implementation mode.

[0076] According to the embodiments of the present disclosure, the present disclosure also provides an advertisement generation device. Exemplarily, Figure 4 A structural schematic diagram of an advertisement generation device provided by an embodiment of the present disclosure. The advertisement generation device 400 comprises: The receiving module 410 is configured to receive a real-time video stream; The identification module 420 is configured to identify a preset event type in the real-time video stream by using an artificial intelligence model; wherein the preset event type comprises a preset action type; The fusion module 430 is configured to extract visual element data of the preset event type, and perform fusion processing on the visual element data and preset advertisement information in a preset fusion manner to generate target advertisement content. The rendering module 440 is configured to render the target advertisement content to a set position of the real-time video stream in real time to generate an advertisement real-time video stream. The output module 450 is configured to output the advertisement real-time video stream.

[0077] Further, the fusion module 430 is configured to: extract at least one of a motion object trajectory, a motion contour of a person, and a video clip of a person of the preset event type; The preset fusion manner includes at least one of the following: deforming the motion object trajectory into a preset advertisement logo; integrating the preset advertisement logo into the motion contour of the person; superimposing preset advertisement information associated with semantic information of the video clip in the video clip of the person; wherein the preset advertisement information includes preset advertisement text content; generating a virtual image of the person, and evolving a motion trajectory of the virtual image into the preset advertisement logo.

[0078] Further, the fusion module 430 is configured to: extract a motion trajectory coordinate sequence of the motion object; control a preset tailing element to generate a dynamic tailing along the motion trajectory coordinate sequence according to the motion trajectory coordinate sequence; adjust a form parameter of the dynamic tailing, so that the dynamic tailing gradually changes into the preset advertisement logo.

[0079] Further, the fusion module 430 is configured to: perform UV unwrapping processing on the motion contour of the virtual person to generate a mapping relationship between the motion contour of the person and a UV coordinate; map the preset advertisement logo as a texture map to a preset region of the motion contour of the person based on the mapping relationship; wherein the preset region is located in the motion contour of the person or on the surface of the motion contour of the person.

[0080] Further, the fusion module 430 is configured to: extract semantic information of the video clip; determine preset advertisement information associated with the semantic information in a preset advertisement information set according to the semantic information; The preset advertisement information associated with the semantic information is superimposed into the video segment by a preset special effect.

[0081] Further, the fusion module 430 is configured to: obtain a character feature of a character in the real-time video stream; wherein the character feature comprises an appearance feature and a motion feature; generate a virtual image of the character based on the appearance feature; evolve a motion trajectory of the virtual image into the preset advertisement identifier according to the preset advertisement identifier and the motion feature.

[0082] Further, the fusion module 430 is configured to: extract a key motion feature of the preset event type; wherein the key motion feature comprises a motion path, a speed, an acceleration, a duration, and a visual focus; determine an emotion curve of the preset event type according to a shot switching frequency, a slow motion mark, a sound track intensity, and a color fluctuation; combine the key motion feature and the emotion curve into an event vector trajectory; in a case where the event vector trajectory and a preset advertisement information envelope space have a cross resonance point, extract visual element data of the preset event type, and perform fusion processing on the visual element data and the preset advertisement information by a preset fusion manner to generate target advertisement content.

[0083] Further, the rendering module 440 is configured to: in a case where the event vector trajectory and a preset advertisement information envelope space do not have a cross resonance point, render the preset advertisement identifier or project the preset advertisement identifier to a non-interference area of the real-time video stream by a preset rendering manner; wherein the preset rendering manner comprises a transparency gradual rendering or a boundary suspension manner rendering; and the non-interference area comprises a virtual lens frame, a real-time score floating layer, or a corner particle scattering area.

[0084] It should be noted that the features of the advertisement generation device corresponding to the embodiments can refer to the related descriptions of the method corresponding embodiments, which will not be repeated here.

[0085] Embodiments of the present disclosure also provide an electronic device, comprising a memory and a processor, the memory stores a computer program, and the processor is configured to run the computer program to execute the steps in any one of the above method embodiments.

[0086] The embodiment of the disclosure further provides a computer readable storage medium, and the computer readable storage medium stores a computer program. When the computer program is executed, the steps in any of the method embodiments described above are performed.

[0087] In an example embodiment, the computer readable storage medium described above can include, but is not limited to, a U disk, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk, and various media that can store computer programs.

[0088] The embodiment of the disclosure further provides a computer program product, and the computer program product includes a computer program. When the computer program is executed by a processor, the steps in any of the method embodiments described above are implemented.

[0089] The embodiment of the disclosure further provides another computer program product, and the computer program product includes a non-volatile computer readable storage medium. The non-volatile computer readable storage medium stores a computer program. When the computer program is executed by a processor, the steps in any of the method embodiments described above are implemented.

[0090] The skilled person can further realize that the units and algorithm steps of the examples described in conjunction with the embodiments disclosed herein can be realized in electronic hardware, computer software or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of the examples have been described in the above description in general terms. Whether the functions are realized in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to realize the described functions for each specific application, but such implementation should not be considered beyond the scope of the disclosure.

[0091] The above describes in detail the advertisement generation method provided by the disclosure. The principles and implementation manners of the disclosure are described by applying specific examples. The above description of the embodiments is only used to help understand the method of the disclosure and its core idea. It should be noted that, for those skilled in the art, without departing from the principles of the disclosure, some improvements and modifications can be made to the disclosure. These improvements and modifications also fall within the protection scope of the claims of the disclosure.

Claims

1. An advertisement generation method, characterized in that, include: Receive real-time video streams; The preset event types in the real-time video stream are identified using an artificial intelligence model; wherein, the preset event types include preset action types; Extract visual element data of the preset event type, and fuse the visual element data and preset advertising information through a preset fusion method to generate target advertising content; The target advertisement content is rendered in real time to a set position in the real-time video stream to generate an advertisement real-time video stream; Output the real-time video stream of the advertisement.

2. The advertisement generation method according to claim 1, characterized in that, The extraction of visual element data for the preset event type includes: Extract at least one of the following: the trajectory of a moving object, the silhouette of a person's movement, and a video clip of a person, based on the preset event type; The preset fusion method includes at least one of the following: The trajectory of the moving object is transformed into a preset advertising logo; The preset advertising logo is integrated into the silhouette of the character's movement. Preset advertising information associated with the semantic information of the video clips is superimposed onto the video clips of the person; wherein, the preset advertising information includes preset advertising text content; A virtual image of the character is generated, and the movement trajectory of the virtual image is transformed into the preset advertising logo.

3. The advertisement generation method according to claim 2, characterized in that, The step of transforming the trajectory of the moving object into a preset advertising logo includes: Extract the motion trajectory coordinate sequence of the moving object; Based on the motion trajectory coordinate sequence, control the preset trailing element to generate a dynamic trailing along the motion trajectory coordinate sequence; Adjust the shape parameters of the dynamic trail so that the dynamic trail gradually transforms into the preset advertising logo.

4. The advertisement generation method according to claim 2, characterized in that, The process of integrating the preset advertising logo into the character's movement silhouette includes: The character's motion contour is subjected to UV unwrapping to generate a mapping relationship between the character's motion contour and UV coordinates; The preset advertising logo is used as a texture map and mapped to a preset area of ​​the character's action outline based on the mapping relationship; wherein the preset area is located in the character's action outline or on the surface of the character's action outline.

5. The advertisement generation method according to claim 2, characterized in that, The method of overlaying preset advertising information associated with the semantic information of the video clip onto the video clip of the person includes: Extract the semantic information of the video segment; Based on the semantic information, determine the preset advertising information associated with the semantic information in the preset advertising information set; Preset advertising information associated with the semantic information is superimposed onto the video clip using preset special effects.

6. The advertisement generation method according to claim 2, characterized in that, The process of generating a virtual avatar of the character and transforming the motion trajectory of the virtual avatar into the preset advertising logo includes: Obtain the character features of the people in the real-time video stream; wherein, the character features include appearance features and action features; A virtual image of the character is generated based on the aforementioned appearance features; Based on the preset advertising identifier and the action characteristics, the movement trajectory of the virtual character is controlled to evolve into the preset advertising identifier.

7. The advertisement generation method according to any one of claims 3-6, characterized in that, The step of extracting visual element data of the preset event type and fusing the visual element data and preset advertising information through a preset fusion method to generate target advertising content includes: Extract key action features of the preset event type; wherein, the key action features include action path, speed, acceleration, duration, and visual focus of the screen; The emotional curve of the preset event type is determined based on the camera switching frequency, slow-motion markers, audio track intensity, and color fluctuations. The key action features are combined with the emotion curve to form an event vector trajectory; When there is a point of intersection and resonance between the event vector trajectory and the preset advertising information envelope space, the visual element data of the preset event type is extracted, and the visual element data and the preset advertising information are fused through a preset fusion method to generate target advertising content.

8. The advertisement generation method according to claim 7, characterized in that, The step of rendering the target advertisement content to a predetermined position in the real-time video stream to generate the advertisement real-time video stream includes: When there is no intersection or resonance point between the event vector trajectory and the preset advertising information envelope space, the preset advertising logo is rendered using a preset rendering method or the preset advertising logo is projected onto a non-interference area of ​​the real-time video stream; wherein, the preset rendering method includes transparency gradient rendering or boundary floating rendering; the non-interference area includes virtual lens borders, real-time score overlays, or corner particle scattering areas.

9. An advertisement generation device, characterized in that, include: The receiving module is used to receive real-time video streams; The recognition module is used to identify preset event types in the real-time video stream through an artificial intelligence model; wherein, the preset event types include preset action types; The fusion module is used to extract visual element data of the preset event type, and to fuse the visual element data and preset advertising information through a preset fusion method to generate target advertising content; The rendering module is used to render the target advertisement content to a set position in the real-time video stream to generate the advertisement real-time video stream; The output module is used to output the real-time video stream of the advertisement.

10. An electronic device, characterized in that, include: At least one processor; as well as, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1-8.