Video poster display method based on artificial intelligence
By constructing a dual-channel feature extraction network and a lightweight network model, the problems of low efficiency and insufficient content representativeness in video poster generation technology are solved, intelligent video poster generation is realized, and user interaction and display effects are improved.
Patent Information
- Application Number
- CN202510774578.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-11
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-06-11
AI Technical Summary
Existing video poster generation technology relies on manual editing of key frames, which is inefficient and lacks content representativeness. It also lacks the ability to intelligently adapt to user preferences and scene characteristics, resulting in a stiff effect.
An AI-based approach is used to construct a dual-channel feature extraction network (DCFN) to detect dynamic semantic anchors, generate multi-dimensional semantic heat maps, and use a differentiable rendering engine to generate video poster elements. Combined with a lightweight network model, it predicts the duration of user gaze retention, updates user portraits in real time, adjusts display light brightness, and presents content in stages.
It improves the production efficiency and content representativeness of video posters, enhances user interaction features and visual appeal, and optimizes display effects and richness.
Smart Images

Figure CN120676202A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of multimedia digital data processing, and in particular to a video poster display method based on artificial intelligence. Background Art
[0002] Because static posters struggle to capture the dynamic nature and core narrative logic of a video, dynamic video posters, as a visual communication tool, are currently widely used for advertisements, promotional videos, or short clips of video content. Video posters aim to quickly capture viewers' attention through visuals and concise messaging. Combining images, text, and possibly audio elements, they present core information in a short clip (usually 10 seconds or less). They are suitable for a variety of platforms and media, including social media, websites, television commercials, film trailers, product promotions, and event announcements. Because they convey information concisely and powerfully, video posters often enhance the appeal and effectiveness of content.
[0003] However, the existing video poster generation technology has the following shortcomings: first, it relies on manual editing of key frames, which is inefficient and lacks content representativeness; second, it lacks the ability to intelligently adapt to user preferences and scene characteristics, resulting in a stiff video poster effect. Summary of the Invention
[0004] The purpose of the present invention is to provide a video poster display method based on artificial intelligence.
[0005] To achieve the above object, the present invention provides the following technical solutions: An AI-based video poster display method includes inputting user instructions to obtain at least one video and / or poster related to the target display content, constructing a dual-channel feature extraction network (DCFN), and simultaneously generating a multi-dimensional semantic heat map to detect dynamic semantic anchor points; The detection results of the dynamic semantic anchors are processed using differentiable rendering engine (DRE) technology to generate several video poster elements. By detecting dynamic semantic anchors, a dual-channel feature extraction network can be constructed to extract feature information of the video and / or poster related to the target display content. By using the differentiable rendering engine to generate several video poster elements, the content of the video altitude can be made more representative, avoiding the lack of content representativeness caused by relying on manual editing of keyframes, and at the same time helping to improve the efficiency of video poster production. The several types of video poster elements generated above are synchronously displayed through the display terminal, and the user portrait is updated in real time during the display process. By updating the user portrait in real time during the display process, it is beneficial to improve the interactive characteristics of the video poster and the user, enhance the visual appeal of the video poster to the user, and facilitate the display and dissemination effect of the video poster, as well as perceive and adjust the display light brightness of the video poster. By perceiving the display light brightness in the display environment of the video poster, it is possible to facilitate timely adjustment of the light contrast and brightness of the video poster on the display terminal, which is beneficial to improve the display effect of the video poster and facilitate the full presentation of the video poster elements; By deploying a lightweight (LSTM) network model to predict the length of time a user's gaze stays on a displayed video poster, this method is conducive to collecting the time a user's gaze stays on different elements of a video poster, optimizing the display elements of a video poster, analyzing the degree of attraction of different elements of a video poster to users, and adaptively presenting the complete content of the video poster in three stages according to the length of time the user's gaze stays. By presenting different elements of a video poster in stages, it is conducive to users' concentrated observation of the elements presented in different stages, and at the same time enriching the display richness of the video poster.
[0006] As a further solution of the present invention: the method for constructing a dual-channel feature extraction network (DCFN) includes: A three-dimensional convolutional network (3D-CNN) is used to construct the visual channel and extract spatiotemporal features.
[0007] As a further solution of the present invention: the method for constructing a dual-channel feature extraction network (DCFN) further includes: A pre-trained model (BERT) is used to extract audio text features.
[0008] As a further solution of the present invention: the method for constructing a dual-channel feature extraction network (DCFN) further includes: Select key segments with narrative integrity from the extracted audio text features.
[0009] As a further solution of the present invention: the method for generating a multidimensional semantic heat map includes: The seaborn library in Python is used to describe the size of different data by changing the depth of color, and the aggregation state of the data is displayed by showing the change of color depth.
[0010] As a further solution of the present invention: the method for generating several video poster elements includes: Use a high-resolution image generation model to generate 4K-level resolution visual images and synthesize image elements; Generate 3D dynamic effect layers through Neural Radiance Field (NeRF) technology to synthesize dynamic elements in video posters; Large language models (LLMs) are used to generate slogans adapted to image elements.
[0011] As a further solution of the present invention: the method for generating several video poster elements further includes: The composition of the high-resolution visual image and the dynamic elements is optimized, and the aesthetic evaluation model (AESNet) is used to evaluate the color harmony and the distribution of visual focus of the high-resolution visual image and the dynamic elements.
[0012] As a further solution of the present invention: the method for updating the user portrait in real time includes: Track user eye movement data and determine the visual weight of elements on video posters based on the eye movement data, and train a video poster recommendation model based on historical interaction data.
[0013] As a further solution of the present invention: the method for sensing and adjusting the display light brightness of a video poster includes: Detect the resolution of the video poster display terminal and dynamically adjust the brightness and contrast of the display terminal through the ambient light data of the video poster display collected by the light sensor.
[0014] As a further solution of the present invention: the method of adaptively presenting the complete content of the video poster in three stages according to the length of time the user's gaze remains includes: The video poster presentation stages are divided into the first presentation stage, the second presentation stage and the third presentation stage. In the first presentation stage, the core visual symbols are presented on the display terminal; in the second presentation stage, the audio content of the text commentary appears; and in the third presentation stage, the micro-dynamic effects are activated on the core visual symbols.
[0015] Compared with the prior art, the present invention has the following beneficial effects: 1. In the present invention, by detecting dynamic semantic anchor points, a dual-channel feature extraction network can be constructed to extract feature information of videos and / or posters related to the target display content. By adopting a differentiable rendering engine to generate several video poster elements, the content of the video altitude can be made more representative, avoiding the lack of content representativeness caused by relying on manual editing of key frames, and at the same time helping to improve the production efficiency of video posters.
[0016] 2. In the present invention, by updating the user portrait in real time during the display process, it is beneficial to improve the interactive characteristics of the video poster and the user, enhance the visual appeal of the video poster to the user, and facilitate the display and dissemination effect of the video poster.
[0017] 3. In the present invention, by sensing the brightness of the display light in the display environment of the video poster, it is possible to timely adjust the light contrast and brightness of the video poster on the display terminal, which is beneficial to improving the display effect of the video poster and facilitating the full presentation of the video poster elements.
[0018] 4. In the present invention, a lightweight (LSTM) network model is deployed to predict the length of time a user's gaze stays on a displayed video poster, which is beneficial for collecting the length of time a user's gaze stays on different elements of a video poster, optimizing the display elements of a video poster, and analyzing the degree of attraction of different elements of a video poster to users.
[0019] 5. In the present invention, by presenting different elements of the video poster in stages, it is beneficial for users to concentrate on observing the elements presented in different stages, and at the same time it can enrich the display richness of the video poster. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 The figure is a flow chart of the method steps of the present invention. DETAILED DESCRIPTION
[0021] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0022] Example
[0023] See also Figure 1 In an embodiment of the present invention, a video poster display method based on artificial intelligence includes: S1: Input user instructions to obtain at least one video and / or poster related to the target display content, build a dual-channel feature extraction network (DCFN), and simultaneously generate a multi-dimensional semantic heat map to detect dynamic semantic anchors; S2: Using the Differentiable Rendering Engine (DRE) technology to process the detection results of the above-mentioned dynamic semantic anchors, several video poster elements are generated. By detecting the dynamic semantic anchors, a dual-channel feature extraction network can be constructed to extract feature information of the video and / or poster related to the target display content. By using the differentiable rendering engine to generate several video poster elements, the content of the video altitude can be made more representative, avoiding the lack of content representativeness caused by relying on manual editing of keyframes, and at the same time helping to improve the efficiency of video poster production. S3: Synchronously displaying the several types of video poster elements generated above through a display terminal, and updating the user portrait in real time during the display process. By updating the user portrait in real time during the display process, it is beneficial to improve the interactive characteristics of the video poster and the user, enhance the visual appeal of the video poster to the user, and facilitate the display and dissemination effect of the video poster, and perceive and adjust the display light brightness of the video poster. By perceiving the display light brightness in the display environment of the video poster, it is possible to facilitate timely adjustment of the light contrast and brightness of the video poster on the display terminal, which is beneficial to improving the display effect of the video poster and facilitating the full presentation of the video poster elements; S4: By deploying a lightweight (LSTM) network model to predict the length of time the user's gaze stays on the displayed video poster, deploying a lightweight (LSTM) network model to predict the length of time the user's gaze stays on the displayed video poster is conducive to collecting the time the user's gaze stays on different elements of the video poster, which is conducive to optimizing the display elements of the video poster, analyzing the degree of attraction of different elements of the video poster to users, and adaptively presenting the complete content of the video poster in three stages according to the length of time the user's gaze stays. By presenting different elements of the video poster in stages, it is conducive to users' concentrated observation of the elements presented in different stages, and at the same time can enrich the display richness of the video poster.
[0024] Preferably, the method for constructing a dual-channel feature extraction network (DCFN) includes: using a three-dimensional convolutional neural network (3D-CNN) to construct a visual channel and extract spatiotemporal features.
[0025] Preferably, the method for constructing a dual-channel feature extraction network (DCFN) further includes: extracting audio text features using a pre-trained model (BERT).
[0026] Preferably, the method for constructing a dual-channel feature extraction network (DCFN) further includes: selecting key segments with narrative integrity from the extracted audio text features.
[0027] Preferably, the method for generating a multidimensional semantic heat map includes: using the seaborn library in Python to describe the size of different data by changing the depth of color, and displaying the aggregation state of the data by displaying the depth change state of color.
[0028] Preferably, the method for generating several video poster elements includes: using a high-resolution image generation model to generate a 4K-level resolution visual image and synthesizing the image elements; Generate 3D dynamic effect layers through Neural Radiance Field (NeRF) technology to synthesize dynamic elements in video posters; Large language models (LLMs) are used to generate slogans adapted to image elements.
[0029] Preferably, the method for generating several video poster elements further includes: optimizing the composition of the high-resolution visual image and dynamic elements, and using an aesthetic evaluation model (AESNet) to evaluate the color harmony and distribution of visual focus of the high-resolution visual image and dynamic elements.
[0030] Preferably, the method for updating the user portrait in real time includes: tracking the user's eye movement data and determining the visual weight of the elements on the video poster based on the eye movement data, and training a recommendation model for the video poster based on historical interaction data.
[0031] Preferably, the method for sensing and adjusting the brightness of the display light of the video poster includes: detecting the resolution of the video poster display terminal, and dynamically adjusting the brightness and contrast of the display terminal through the ambient light data of the video poster display collected by the light sensor.
[0032] Preferably, the method of adaptively presenting the complete content of the video poster in three stages according to the length of time the user's gaze stays includes: dividing the stages of video poster presentation into a first presentation stage, a second presentation stage, and a third presentation stage, and presenting core visual symbols on the display terminal in the first presentation stage; presenting audio content of text commentary in the second presentation stage; and activating micro-dynamic effects on the core visual symbols in the third presentation stage.
[0033] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A video poster display method based on artificial intelligence, characterized in that: include: Input user instructions to obtain at least one video and / or poster related to the target display content, build a dual-channel feature extraction network, and simultaneously generate a multi-dimensional semantic heat map to detect dynamic semantic anchors; The detection results of the dynamic semantic anchor points are processed using differentiable rendering engine technology to generate several video poster elements. The generated video poster elements are synchronously displayed through the display terminal, and the user portrait is updated in real time during the display process, and the display light brightness of the video poster is sensed and adjusted; By deploying a lightweight network model to predict how long the user's gaze stays on the displayed video poster, the full content of the video poster is adaptively presented in three stages based on the length of time the user's gaze stays.
2. The method for displaying video posters based on artificial intelligence according to claim 1, characterized in that: The method for constructing a dual-channel feature extraction network includes: A three-dimensional convolutional architecture is used to construct the sensory channel and extract spatiotemporal features.
3. The method for displaying video posters based on artificial intelligence according to claim 2, characterized in that: The method for constructing a dual-channel feature extraction network also includes: Use pre-trained models to extract audio text features.
4. The method for displaying video posters based on artificial intelligence according to claim 3, characterized in that: The method for constructing a dual-channel feature extraction network also includes: Select key segments with narrative integrity from the extracted audio text features.
5. The artificial intelligence-based video poster display method according to claim 4, characterized in that: The method for generating a multidimensional semantic heat map includes: The seaborn library in Python is used to describe the size of different data by changing the depth of color, and the aggregation state of the data is displayed by showing the change of color depth.
6. The method for displaying video posters based on artificial intelligence according to claim 1, characterized in that: The method for generating several video poster elements includes: Use a high-resolution image generation model to generate 4K-level resolution visual images and synthesize image elements; Generate 3D dynamic effect layers through neural radiation field technology to synthesize dynamic elements in video posters; Use large language models to generate slogans adapted to image elements.
7. The method for displaying video posters based on artificial intelligence according to claim 1, characterized in that: The method for generating several video poster elements further includes: The composition of the high-resolution visual image and the dynamic elements is optimized, and an aesthetic evaluation model is used to evaluate the color harmony and distribution of visual focus of the high-resolution visual image and the dynamic elements.
8. The method for displaying video posters based on artificial intelligence according to claim 1, characterized in that: The method for updating the user portrait in real time includes: Track user eye movement data and determine the visual weight of elements on video posters based on the eye movement data, and train a video poster recommendation model based on historical interaction data.
9. The method for displaying video posters based on artificial intelligence according to claim 1, characterized in that: The method for sensing and adjusting the display light brightness of a video poster includes: Detect the resolution of the video poster display terminal and dynamically adjust the brightness and contrast of the display terminal through the ambient light data of the video poster display collected by the light sensor.
10. The method for displaying video posters based on artificial intelligence according to claim 1, characterized in that: The method of adaptively presenting the complete content of the video poster in three stages according to the length of time the user's gaze remains includes: The video poster presentation stages are divided into the first presentation stage, the second presentation stage and the third presentation stage. In the first presentation stage, the core visual symbols are presented on the display terminal; in the second presentation stage, the audio content of the text commentary appears; and in the third presentation stage, the micro-dynamic effects are activated on the core visual symbols.
Citation Information
Patent Citations
Multi-feature learning online video popularity prediction method based on time information perception
CN119364121A
Poster generation method and device based on RAG poster design and electronic equipment
CN119886053A
Poster generation method and device and storage medium
CN120107413A
Animatable Neural Radiance Fields from Monocular RGB-D Inputs
US20240104828A1