Video poster display method based on artificial intelligence
By constructing a dual-channel feature extraction network and a lightweight network model, combined with real-time user profiling and lighting adjustments, the problems of low efficiency and insufficient content representativeness in video poster generation were solved, achieving efficient and intelligent video poster display.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- STATE GRID HUBEI ELECTRIC POWER CO LTD WUHAN DONGHU NEW TECH DEV ZONE POWER SUPPLY CO
- Filing Date
- 2025-06-11
- Publication Date
- 2026-04-21
AI Technical Summary
Existing video poster generation technologies rely on manual editing of keyframes, which is inefficient and lacks content representativeness. Furthermore, they lack the ability to intelligently adapt to user preferences and scene characteristics, resulting in stiff video poster effects.
An AI-based approach is employed, which involves constructing a dual-channel feature extraction network (DCFN) to detect dynamic semantic anchors, using a differentiable rendering engine to generate video poster elements, and combining a lightweight network model to predict the duration of user gaze, thereby updating user profiles and adjusting display lighting in real time and presenting video poster content in stages.
It improved the production efficiency and content representativeness of video posters, enhanced user interaction features and visual appeal, and optimized the display and dissemination effects.
Smart Images

Figure CN120676202B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimedia digital data processing technology, specifically to a video poster display method based on artificial intelligence. Background Technology
[0002] Because static posters struggle to capture the dynamic features and core narrative logic of videos, dynamic video posters, as a visual communication tool, are now widely used in advertisements, promotional videos, or short clips of video content. Video posters aim to quickly grab the viewer's attention with visual and concise information. They combine images, text, and possibly audio elements to present core information in a short duration (usually 10 seconds or less). Suitable for various platforms and media, such as social media, websites, television commercials, movie trailers, product promotions, and event announcements, video posters often significantly enhance the appeal and dissemination of content due to their concise and powerful message delivery.
[0003] However, existing video poster generation technologies have the following shortcomings: First, they rely on manual editing of keyframes, which is inefficient and lacks content representativeness; second, they lack the ability to intelligently adapt to user preferences and scene characteristics, resulting in stiff video poster effects. Summary of the Invention
[0004] The purpose of this invention is to provide a video poster display method based on artificial intelligence.
[0005] To achieve the above objectives, the present invention provides the following technical solution:
[0006] An AI-based video poster display method includes inputting user commands to obtain at least one video and / or poster related to the target display content, constructing a dual-channel feature extraction network (DCFN), and simultaneously generating a multi-dimensional semantic heatmap to achieve the detection of dynamic semantic anchor points;
[0007] Differentiable rendering engine (DRE) technology is used to process the detection results of the above dynamic semantic anchor points to generate several video poster elements. By detecting dynamic semantic anchor points, a dual-channel feature extraction network can be constructed to extract feature information of videos and / or posters related to the target display content. By using a differentiable rendering engine to generate several video poster elements, the content of the video poster can be more representative, avoiding the lack of content representativeness caused by relying on manual editing of keyframes, and at the same time, it is beneficial to improve the production efficiency of video posters.
[0008] The various video poster elements generated above are simultaneously displayed on a display terminal, and the user profile is updated in real time during the display process. By updating the user profile in real time during the display process, it is beneficial to improve the interactive characteristics between the video poster and the user, enhance the visual appeal of the video poster to the user, and improve the display and dissemination effect of the video poster. Furthermore, by sensing and adjusting the display light brightness of the video poster in the display environment, it is possible to adjust the light contrast and brightness of the video poster on the display terminal in a timely manner, which is beneficial to improving the display effect of the video poster and ensuring the full presentation of the video poster elements.
[0009] By deploying a lightweight (LSTM) network model to predict the duration of user gaze on the displayed video poster, it is beneficial to collect the time users spend looking at different elements of the video poster. This allows for the optimization of the displayed elements, analysis of the attractiveness of different elements to users, and adaptive presentation of the complete content of the video poster in three stages based on the length of time users' gaze stays. Presenting different elements of the video poster in stages helps users focus on observing the elements presented at different stages, while also enriching the presentation of the video poster.
[0010] As a further aspect of the present invention: the method for constructing a dual-channel feature extraction network (DCFN) includes:
[0011] A three-dimensional convolutional architecture (3D-CNN) is used to construct the visual channel and extract spatiotemporal features.
[0012] As a further aspect of the present invention: the method for constructing a dual-channel feature extraction network (DCFN) further includes:
[0013] Audio text features are extracted using a pre-trained model (BERT).
[0014] As a further aspect of the present invention: the method for constructing a dual-channel feature extraction network (DCFN) further includes:
[0015] Select key segments with narrative integrity from the extracted audio text features.
[0016] As a further aspect of the present invention: the method for generating a multidimensional semantic heatmap includes:
[0017] The seaborn library in Python is used to describe the size of different data by varying the intensity of color, and the aggregation state of the data is shown by displaying the intensity of color changes.
[0018] As a further aspect of the present invention: the method for generating several video poster elements includes:
[0019] A high-resolution image generation model is used to generate 4K-level resolution visual images, and then image elements are synthesized.
[0020] 3D dynamic effect layers are generated using Neural Radiation Field (NeRF) technology to synthesize dynamic elements in video posters;
[0021] Large Language Models (LLMs) are used to generate slogans that adapt to image elements.
[0022] As a further aspect of the present invention: the method for generating several video poster elements further includes:
[0023] The composition of the resolution visual image and dynamic elements is optimized, and the color harmony and visual focus distribution of the resolution visual image and dynamic elements are evaluated using the Aesthetic Evaluation Model (AESNet).
[0024] As a further aspect of the present invention: the method for real-time updating of user profiles includes:
[0025] Track user eye-tracking data and determine the visual weight of elements on video posters based on the eye-tracking data, and train a recommendation model for video posters based on historical interaction data.
[0026] As a further aspect of the present invention: the method for sensing and adjusting the display light brightness of the video poster includes:
[0027] The resolution of the video poster display terminal is detected, and the brightness and contrast of the display terminal are dynamically adjusted based on the ambient light data of the video poster display collected by the light sensor.
[0028] As a further aspect of the present invention: the method for adaptively presenting the complete content of the video poster in three stages based on the length of time the user's gaze lingers includes:
[0029] The presentation of the video poster is divided into three stages: the first stage, the second stage, and the third stage. The core visual symbols are presented on the terminal during the first stage; audio content with text narration appears during the second stage; and micro-dynamic effects are activated on the core visual symbols during the third stage.
[0030] Compared with the prior art, the beneficial effects of the present invention are:
[0031] 1. In this invention, by detecting dynamic semantic anchor points, a dual-channel feature extraction network can be constructed to extract feature information of videos and / or posters related to the target display content. By using a differentiable rendering engine to generate several video poster elements, the content of the video poster can be made more representative, avoiding the lack of content representativeness caused by relying on manual editing of key frames, and at the same time, it is beneficial to improve the production efficiency of video posters.
[0032] 2. In this invention, by updating the user profile in real time during the display process, it is beneficial to improve the interactive characteristics between the video poster and the user, enhance the visual appeal of the video poster to the user, and improve the display and dissemination effect of the video poster.
[0033] 3. In this invention, by sensing the brightness of the display light in the display environment of the video poster, it is possible to adjust the light contrast and brightness of the video poster on the display terminal in a timely manner, which is conducive to improving the display effect of the video poster and fully presenting the elements of the video poster.
[0034] 4. In this invention, by deploying a lightweight (LSTM) network model to predict the duration of user gaze on the displayed video poster, it is beneficial to collect the time users spend looking at different elements of the video poster, which is helpful to optimize the display elements of the video poster and analyze the attractiveness of different elements of the video poster to users.
[0035] 5. In this invention, by presenting different elements of the video poster in stages, it is beneficial for users to focus on observing the elements presented at different stages, while also enriching the display richness of the video poster. Attached Figure Description
[0036] Figure 1 This is a flowchart of the method steps of the present invention. Detailed Implementation
[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] Example
[0039] Please see Figure 1 In this embodiment of the invention, a video poster display method based on artificial intelligence includes:
[0040] S1: Input user instructions to obtain at least one video and / or poster related to the target content, construct a dual-channel feature extraction network (DCFN), and simultaneously generate a multi-dimensional semantic heatmap to achieve the detection of dynamic semantic anchor points;
[0041] S2: Differentiable rendering engine (DRE) technology is used to process the detection results of the above dynamic semantic anchor points to generate several video poster elements; by detecting dynamic semantic anchor points, a dual-channel feature extraction network can be constructed to extract feature information of videos and / or posters related to the target display content. By using a differentiable rendering engine to generate several video poster elements, the content of the video poster can be more representative, avoiding the lack of content representativeness caused by relying on manual editing of keyframes, and at the same time, it is conducive to improving the production efficiency of video posters.
[0042] S3: The generated video poster elements are displayed synchronously through the display terminal, and the user profile is updated in real time during the display process. By updating the user profile in real time during the display process, it is beneficial to improve the interactive characteristics between the video poster and the user, enhance the visual appeal of the video poster to the user, and improve the display and dissemination effect of the video poster. In addition, the display light brightness of the video poster can be perceived and adjusted. By perceiving the display light brightness in the display environment of the video poster, it is convenient to adjust the light contrast and brightness of the video poster on the display terminal in a timely manner, which is beneficial to improve the display effect of the video poster and the full presentation of the video poster elements.
[0043] S4: By deploying a lightweight (LSTM) network model to predict the duration of user gaze on the displayed video poster, it is beneficial to collect the time users spend looking at different elements of the video poster. This helps to optimize the display elements of the video poster, analyze the attractiveness of different elements to users, and adaptively present the complete content of the video poster in three stages based on the length of time users' gaze stays. By presenting different elements of the video poster in stages, it is beneficial for users to focus on observing the elements presented at different stages, while also enriching the presentation of the video poster.
[0044] Preferably, the method for constructing the dual-channel feature extraction network (DCFN) includes: constructing a visual channel using a three-dimensional convolutional architecture (3D-CNN) and extracting spatiotemporal features.
[0045] Preferably, the method for constructing a dual-channel feature extraction network (DCFN) further includes: using a pre-trained model (BERT) to extract audio text features.
[0046] Preferably, the method for constructing a dual-channel feature extraction network (DCFN) further includes selecting key segments with narrative integrity from the extracted audio text features.
[0047] Preferably, the method for generating a multidimensional semantic heatmap includes: using the seaborn library in Python to describe the size of different data by varying the intensity of colors, and displaying the aggregation state of the data by showing the intensity of color variations.
[0048] Preferably, the method for generating several video poster elements includes: generating a 4K resolution visual image using a high-resolution image generation model, and then synthesizing the image elements;
[0049] 3D dynamic effect layers are generated using Neural Radiation Field (NeRF) technology to synthesize dynamic elements in video posters;
[0050] Large Language Models (LLMs) are used to generate slogans that adapt to image elements.
[0051] Preferably, the method for generating several video poster elements further includes: optimizing the composition of the resolution visual image and dynamic elements, and using an aesthetic evaluation model (AESNet) to evaluate the color harmony and visual focus distribution of the resolution visual image and dynamic elements.
[0052] Preferably, the method for real-time updating of user profiles includes: tracking user eye movement data and determining the visual weight of elements on the video poster based on the eye movement data, and training a recommendation model for the video poster based on historical interaction data.
[0053] Preferably, the method for sensing and adjusting the display light brightness of the video poster includes: detecting the resolution of the video poster display terminal, and dynamically adjusting the brightness and contrast of the display terminal based on the ambient light data of the video poster display collected by the light sensor.
[0054] Preferably, the method of adaptively presenting the complete content of the video poster in three stages according to the length of time the user's gaze lingers includes: dividing the presentation of the video poster into a first presentation stage, a second presentation stage, and a third presentation stage; presenting the core visual symbol on the display terminal in the first presentation stage; displaying audio content of text narration in the second presentation stage; and activating micro-dynamic effects on the core visual symbol in the third presentation stage.
[0055] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A video poster display method based on artificial intelligence, characterized in that, include: By inputting user commands, at least one video and / or poster related to the target content is obtained. A dual-channel feature extraction network is constructed, and a multi-dimensional semantic heatmap is generated simultaneously to detect dynamic semantic anchor points. Differentiable rendering engine technology is used to process the detection results of the above dynamic semantic anchor points to generate several video poster elements; The generated video poster elements are displayed synchronously through the display terminal, and the user profile is updated in real time during the display process, as well as the display light brightness of the video poster is sensed and adjusted. By deploying a lightweight network model to predict the duration of user gaze on the displayed video poster, the full content of the video poster is adaptively presented in three stages based on the length of time the user's gaze lingers. The method for generating multidimensional semantic heatmaps includes: The seaborn library in Python is used to describe the size of different data by varying the intensity of color, and the aggregation state of the data is shown by displaying the intensity of color changes. The method for generating several types of video poster elements includes: A high-resolution image generation model is used to generate 4K-level resolution visual images, and then image elements are synthesized. 3D dynamic effect layers are generated using neural radiation field technology to synthesize dynamic elements in video posters; Large-scale language models are used to generate slogans that adapt to image elements; The method for real-time updating of user profiles includes: Track user eye-tracking data and determine the visual weight of elements on video posters based on the eye-tracking data, and train a recommendation model for video posters based on historical interaction data; The method of adaptively presenting the complete content of the video poster in three stages based on the length of time the user's gaze lingers includes: The presentation of the video poster is divided into three stages: the first stage, the second stage, and the third stage. The core visual symbols are presented on the terminal during the first stage; audio content with text narration appears during the second stage; and micro-dynamic effects are activated on the core visual symbols during the third stage.
2. The video poster display method based on artificial intelligence according to claim 1, characterized in that: The method for constructing a dual-channel feature extraction network includes: A three-dimensional convolutional architecture is used to construct a sensory channel and extract spatiotemporal features.
3. The video poster display method based on artificial intelligence according to claim 2, characterized in that: The method for constructing a dual-channel feature extraction network also includes: A pre-trained model is used to extract audio text features.
4. The video poster display method based on artificial intelligence according to claim 3, characterized in that: The method for constructing a dual-channel feature extraction network also includes: Select key segments with narrative integrity from the extracted audio text features.
5. The video poster display method based on artificial intelligence according to claim 1, characterized in that: The method for generating several types of video poster elements also includes: The composition of high-resolution visual images and dynamic elements is optimized, and an aesthetic evaluation model is used to evaluate the color harmony and visual focus distribution of high-resolution visual images and dynamic elements.
6. The video poster display method based on artificial intelligence according to claim 1, characterized in that: The method for sensing and adjusting the display light brightness of the video poster includes: The resolution of the video poster display terminal is detected, and the brightness and contrast of the display terminal are dynamically adjusted based on the ambient light data of the video poster display collected by the light sensor.
Citation Information
Patent Citations
Multi-feature learning online video popularity prediction method based on time information perception
CN119364121A
Poster generation method and device based on RAG poster design and electronic equipment
CN119886053A