Methods, devices, media, and computer program products for generating video stories for photo stories
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2026-08-14
AI Technical Summary
然而,固定的模板结构不仅无法适应多样化的叙事需求,还缺乏对图片语义的深度理解,也忽视了用户的个性化偏好,导致产出的图片故事视频千篇一律,缺乏创意
[0013]如以下将详细描述的,根据本公开实施例的一种图片故事的视频生成方法,通过获取用户选定的图片集,并提取图片集的多模态特征;采用预构的图片故事生成算法,基于多模态特征,生成图片集对应的目标故事文案;图片故事生成算法包括复合型奖励函数,复合型奖励函数用于评估故事文案的视觉相关性、流畅性、信息量和用户偏好相关性;采用预构的图片故事布局算法,基于多模态特征和目标故事文案,生成图片集对应的目标故事布局;图片故事布局算法包括布局奖励函数,布局奖励函数用于评估故事布局的视觉表现力和语义相关性;基于图片集、目标故事文案和目标故事布局,生成图片故事视频。
Smart Images

Figure CN121099083B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of artificial intelligence technology, and in particular to methods, devices, media, and computer program products for generating video stories from images. Background Technology
[0002] With the booming development of social networks and the increasing popularity of video content sharing, users are keen to record their daily lives and showcase their personal creativity through photo stories and videos.
[0003] In related technologies, the generation of photo story videos often relies on template generation methods, requiring users to select a template and manually choose photos to fill in the template. However, the fixed template structure not only fails to adapt to diverse narrative needs but also lacks a deep understanding of the semantics of the images and ignores users' personalized preferences, resulting in photo story videos that are all the same and lack creativity. Summary of the Invention
[0004] In view of this, exemplary embodiments of the present disclosure provide a method, apparatus, medium, and computer program product for generating video stories from pictures to address the problems existing in the related art.
[0005] One aspect of an exemplary embodiment of this disclosure provides a method for generating a video of a picture story, the method comprising:
[0006] Obtain the image set selected by the user and extract the multimodal features of the image set;
[0007] A pre-constructed image story generation algorithm is used to generate target story texts corresponding to the image set based on the multimodal features. The image story generation algorithm includes a composite reward function, which is used to evaluate the visual relevance, fluency, information content, and user preference relevance of the story text.
[0008] A pre-constructed image story layout algorithm is used to generate the target story layout corresponding to the image set based on the multimodal features and the target story text. The image story layout algorithm includes a layout reward function, which is used to evaluate the visual expressiveness and semantic relevance of the story layout.
[0009] Based on the image set, the target story text, and the target story layout, an image story video is generated.
[0010] In another aspect of exemplary embodiments of this disclosure, a computer device is provided, including a memory, a processor, and a computer program stored in the memory, the processor executing the computer program to implement the methods described in exemplary embodiments of this disclosure.
[0011] In another aspect of exemplary embodiments of this disclosure, a computer-readable storage medium is provided having a computer program / instructions stored thereon that, when executed by a processor, implements the methods described in exemplary embodiments of this disclosure.
[0012] In another aspect of exemplary embodiments of this disclosure, a computer program product is provided, including a computer program / instructions that, when executed by a processor, implement the methods described in exemplary embodiments of this disclosure.
[0013] As will be described in detail below, a method for generating video of an image story according to an embodiment of this disclosure includes: acquiring a user-selected image set and extracting multimodal features from the image set; employing a pre-constructed image story generation algorithm to generate a target story text corresponding to the image set based on the multimodal features; the image story generation algorithm includes a composite reward function used to evaluate the visual relevance, fluency, information content, and user preference relevance of the story text; employing a pre-constructed image story layout algorithm to generate a target story layout corresponding to the image set based on the multimodal features and the target story text; the image story layout algorithm includes a layout reward function used to evaluate the visual expressiveness and semantic relevance of the story layout; and generating an image story video based on the image set, the target story text, and the target story layout.
[0014] Therefore, the image story video generation method provided in this disclosure, through the synergistic optimization of multimodal feature extraction and composite reward functions, ensures that the generated target story text accurately conveys image information and meets user expectations, thereby solving technical problems such as insufficient multimodal fusion, lack of innovation in generated stories, and domain adaptability. Secondly, the image story layout algorithm achieves dual-objective optimization through a layout reward function. Visual expressiveness evaluation covers compositional integrity, clarity, and dynamic coordination, while semantic relevance focuses on image-text focus matching, emotional consistency, and temporal logic, thus addressing the technical problem of monotonous and unattractive video layouts. Finally, through multimodal features, the image story generation algorithm, and the image story layout algorithm, the generated video content maintains a high degree of stylistic consistency with the user-selected image set and automatically optimizes video parameters to adapt to the publishing requirements of different social media platforms, significantly lowering the production threshold for high-quality video content. Attached Figure Description
[0015] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0016] Figure 1 A flowchart illustrating the method for generating video stories from images provided in this embodiment of the disclosure;
[0017] Figure 2 A flowchart illustrating the video generation method for image stories provided in this embodiment of the disclosure;
[0018] Figure 3 A schematic diagram illustrating the process of generating target story text provided in this embodiment of the disclosure;
[0019] Figure 4 This is a schematic diagram of the structure of the composite reward function provided in the embodiments of this disclosure;
[0020] Figure 5 A schematic block diagram of the functional modules of the video generation apparatus for picture stories provided in the embodiments of this disclosure;
[0021] Figure 6 A structural block diagram of an electronic device provided in an embodiment of this disclosure;
[0022] Figure 7 A schematic diagram of a computer program product provided in an embodiment of this disclosure. Detailed Implementation
[0023] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0024] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0025] The term "comprising" and its variations as used herein are open-ended, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below. It should be noted that the concepts of "first", "second", etc., used in this disclosure are only used to distinguish different devices, modules, or units, and are not intended to limit the order of functions performed by these devices, modules, or units or their interdependencies.
[0026] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0027] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0028] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0029] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0030] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device. It is understood that the above notification and user authorization process is merely illustrative and does not constitute a limitation on the implementation of this disclosure; other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0031] With the booming development of social networks and the increasing popularity of video content sharing, users are keen to record their daily lives and showcase their personal creativity through photo stories and videos.
[0032] In related technologies, the generation of photo story videos often relies on template generation methods, requiring users to select a template and manually choose photos to fill in the template. However, the fixed template structure not only fails to adapt to diverse narrative needs but also lacks a deep understanding of the semantics of the images and ignores users' personalized preferences, resulting in photo story videos that are all the same and lack creativity.
[0033] Therefore, in order to solve the above problems, this exemplary embodiment provides a method for generating video of picture stories. Figure 1This is a flowchart illustrating the method for generating video stories from images provided in this embodiment. The process begins with a user uploading an image to a cloud storage system. The image is first processed by an image preprocessing module to generate a preprocessed image. Subsequently, a feature extraction module extracts visual and textual features from the image. These features are then passed to a multimodal feature fusion module to generate initial multimodal features. Next, a knowledge graph embedding module combines the knowledge graph features with the initial multimodal features, weighting and adjusting feature importance through an attention mechanism to generate knowledge-enhanced multimodal features. The MMRL-PSG (Multimodal Reinforcement Learning for Photo Story Generation) algorithm module generates a target story text based on the multimodal features. The RL-APSL (Reinforcement Learning for Adaptive Photo Story Layout) algorithm matches the target story layout to the target story text and the user-uploaded image. The story sharing and recommendation module generates a video story based on the target story text, the user-uploaded image, and the target story layout, and implements intelligent recommendation, thus completing the generation and recommendation of the video story.
[0034] For example, Figure 2 A flowchart illustrating the video generation method for image stories provided in this embodiment of the disclosure is shown below. Figure 2 As shown, the specific steps may include:
[0035] Step S210: Obtain the image set selected by the user and extract the multimodal features of the image set.
[0036] In this embodiment, the cloud drive client can filter authorized and usable image data from the user's personal cloud storage space. The image data may include: images actively uploaded by the user, automatically backed-up cached images from applications such as social platforms or instant messaging tools, and material images from the cloud drive's proprietary material library.
[0037] During the data collection process, the cloud drive client verifies the access permissions and file integrity of the images. After filtering, the image set composed of the image data is submitted to the image preprocessing module for preprocessing.
[0038] For example, preprocessing may include: using nonlocal means (NLM) or deep learning models (such as DnCNN) to eliminate sensor noise or compress artifacts, thereby achieving image denoising.
[0039] Preprocessing may also include: adjusting the resolution based on adaptive scaling according to content importance.
[0040] Preprocessing may also include: unifying white balance and exposure to avoid color deviations caused by differences across devices.
[0041] Preprocessing may also include metadata extraction. Metadata is used to record background information, attributes, or characteristics of a file. In image processing, metadata extraction refers to the process of automatically parsing this structured information from image files.
[0042] The high-quality image set output after preprocessing can provide well-structured input data for subsequent multimodal feature extraction.
[0043] The preprocessed image is sent to the feature extraction module in the cloud for feature extraction. The extracted features are then input into the multimodal feature fusion module for feature fusion processing to generate initial multimodal features.
[0044] Then, the knowledge graph features are combined with the initial multimodal features through the knowledge graph embedding module to generate multimodal features. Specifically, a pre-built visual-semantic knowledge graph can be used to map the labels to the semantic space and supplement implicit associations (such as "beach" can be associated with "vacation" and "summer").
[0045] The multimodal feature fusion module uses an attention mechanism to weight and adjust the importance of multimodal features, and outputs multimodal features that have been enhanced with knowledge.
[0046] Step S220: Using a pre-built image story generation algorithm, based on multimodal features, generate target story text corresponding to the image set. The image story generation algorithm includes a composite reward function, which is used to evaluate the visual relevance, fluency, information content, and user preference relevance of the story text.
[0047] For example, a high-dimensional wavelet manifold multimodal reinforcement learning-based image story generation algorithm (MMRL-PSG) can be used to address issues such as insufficient multimodal fusion, inadequate knowledge utilization, and a lack of innovation and domain adaptability in generated stories. This algorithm evaluates the visual relevance, fluency, information content, and user preference relevance of the story text through a composite reward function, prompting the image story generation algorithm to transform user-uploaded image sets into coherent and engaging story text.
[0048] Step S230: Using a pre-constructed image story layout algorithm, based on multimodal features and target story text, generate the target story layout corresponding to the image set. The image story layout algorithm includes a layout reward function, which is used to evaluate the visual expressiveness and semantic relevance of the story layout.
[0049] In this embodiment, in order to address the problems of excessive reliance on fixed templates, lack of semantic understanding and insufficient personalization in existing image and video layout schemes, a reinforcement learning adaptive image story layout algorithm (RL-APSL) is proposed. This algorithm achieves intelligent story video generation by deeply integrating multimodal semantic understanding and personalized layout optimization.
[0050] The RL-APSL algorithm automatically matches a story video template based on the target story text and the client's selected image set. It also automatically adjusts the layout and placement of each image within the video template, adding the target story text. The entire process requires minimal user intervention, achieving adaptive and automated generation of image story videos. This addresses the shortcomings of existing technologies in personalization and semantic relevance, enabling intelligent and adaptive image story layout generation, and providing users with a more aesthetically pleasing, personalized, and engaging visual experience.
[0051] Step S240: Generate a photo story video based on the photo set, target story text, and target story layout.
[0052] After completing the target story copy and layout, the story sharing and recommendation module can generate image story videos based on the user-uploaded image sets, the target story copy generated by the MMRL-PSG algorithm, and the target story layout generated by the RL-APSL algorithm. Then, based on the user's social networks, interest tags, and other information, the generated image stories are recommended to potentially interested users, completing the generation and recommendation of image story videos.
[0053] Specifically, when generating photo story videos, the story sharing and recommendation module can automatically match theme templates, adjust fonts, colors, and transition animations, and arrange the photo story videos according to the optimal image placement and playback order based on the edited text information.
[0054] In one alternative approach, the playback order can also be adjusted based on semantic coherence (e.g., from early to late according to timeline) or emotional rhythm (from calm to climax).
[0055] In an alternative approach, clickable hotspot tags can be embedded, allowing users to view details.
[0056] In one alternative approach, background music can be added based on the emotional context of the scene.
[0057] For example, the story sharing and recommendation module can use a hybrid recommendation strategy to accurately push the generated image story videos to potentially interested users, achieving efficient dissemination and interaction.
[0058] Specifically, based on social graph analysis, content can be prioritized for recommendations to users strongly associated with the publisher, and members of social circles with common tags or co-occurring locations can be identified to enhance the relevance and intimacy of the content.
[0059] Interest modeling can also be used to predict interest matching by combining users’ historical behavior with real-time feedback (such as likes or dwell time), ensuring that recommended content matches personal preferences.
[0060] Finally, a multi-scenario reach mechanism can be adopted to support private sharing and public recommendations, thereby covering different sharing needs and improving content exposure and user engagement.
[0061] Therefore, by dynamically optimizing the recommendation strategy, we can ensure that photo stories and videos can accurately reach the target users and promote social interaction and dissemination.
[0062] Based on this, by collaboratively optimizing multimodal feature extraction and a composite reward function, the generated target story text accurately conveys the image information while meeting user expectations, thus solving technical problems such as insufficient multimodal fusion, lack of innovation in generated stories, and domain adaptability. Secondly, the image story layout algorithm achieves dual-objective optimization through a layout reward function. Visual expressiveness evaluation covers compositional integrity, clarity, and dynamic coordination, while semantic relevance focuses on image-text focus matching, emotional consistency, and temporal logic, thereby addressing the technical problem of monotonous and unattractive video layouts. Finally, through multimodal features, the image story generation algorithm, and the image story layout algorithm, the generated video content maintains a high degree of stylistic consistency with the user-selected image set and automatically optimizes video parameters to adapt to the posting requirements of different social media platforms, significantly lowering the production threshold for high-quality video content.
[0063] Based on the above embodiments, in another embodiment provided in this disclosure, a specific process for generating target story text is also provided.
[0064] For example, Figure 3 This is a flowchart illustrating the process of generating target story text provided in an embodiment of this disclosure, such as... Figure 3 As shown, the specific steps may include:
[0065] Step S301: Fuse the image set and its corresponding multimodal features to obtain f i And form the initial multimodal features M.
[0066] For example, each image in the image set contains multiple multimodal features, such as the names of objects in the image, the number of people in the image, the location of the background, text information in the image, and image file name information. These different modalities of data usually contain complementary information that is valuable to the target task. By extracting different multimodal features, the input data can be more comprehensively characterized, improving the model's understanding and generation capabilities.
[0067] Based on the above embodiments, in another embodiment provided in this disclosure, the extraction of multimodal features from the image set may include:
[0068] Extract visual and textual features from the image set;
[0069] Visual and textual features are fused to obtain initial multimodal features;
[0070] The initial multimodal features are enhanced by knowledge distillation to obtain more multimodal features.
[0071] For example, deep convolutional neural networks (CNNs) or visual transformers (ViTs) can be used to extract visual features, and the output can be image recognition labels. Visual features can include: object classification, number of people, scene classification, and sentiment classification.
[0072] It can semantically encode embedded and related text in images and generate dense vectors using a pre-trained large language model. By capturing keywords, sentiment, and contextual relationships, it extracts text features. Text features can include: background text of the image and image filename.
[0073] Then, the extracted visual features and text features are input into the multimodal feature fusion module for feature fusion processing to generate initial multimodal features.
[0074] For example, feature fusion processing may include directly concatenating visual feature vectors with text feature vectors to generate initial multimodal features.
[0075] Specifically, the extracted visual and textual features are first combined into a feature sequence f. i ∈R di .
[0076] Furthermore, the fused feature sequence f i Making high-dimensional ψ w The mapping representation yields a high-dimensional representation h of the features. i , can be represented as:
[0077]
[0078] in, and For mapping matrix and bias terms; σ is the scale parameter of the wavelet basis; σ is the nonlinear activation function; N is the number of modes.
[0079] Where, ψ w Formulaic mapping can fully capture the multi-scale information of features in the time and frequency domains, enhancing the representational power of features. The mapping function ψ w The formula can be expressed as:
[0080]
[0081] in, For two d x The input vector is d; |y| is the magnitude of vector y; L is a d x ×d x A diagonal matrix, whose diagonal elements are the elements of y; Z is a d x ×d x The symmetric positive definite matrix is the Mahalanobis distance matrix used to measure the distance between sample points x and y in the feature space. The Mahalanobis distance matrix can be obtained by taking the inverse of the covariance matrix of y. For a group of d x An orthogonal basis vector of dimension vj Tv k = δjk, where δjk is the Kronecker delta function; <·,·> are the inner product operations of vectors; ψ(·) is a one-dimensional mother wavelet function used to extract multi-scale features of the input signal; exp(·) is the exponential function.
[0082] Finally, based on the high-dimensional mapping ψ w and high-dimensional representation of features h i Generate a multimodal representation M;
[0083] For the calculated high-dimensional mapping ψ w feature An attention fusion mechanism was designed to obtain the final multimodal representation M:
[0084]
[0085] Where, α ij For query vector W q h i For the key vector W k h j Attention weights; h i ,h j ,h l The i-th, j-th, and l-th high-dimensional vectors are respectively calculated from the aforementioned steps; These are the query matrix and the key matrix, used to store the input vector h. i and h j Mapped to a common attention space; (W q h i ) T (W k h j ) is the query vector W q h i and key vector W k h j The inner product in the attention space is used to measure the similarity between two vectors. is the scaling factor for the inner product result; exp(·) is the exponential function that transforms the inner product result into a non-negative similarity score; |W q h i -W k h j |2 represents the query vector W q h i and key vector W k h j Distance in the Euclidean sense, i.e., the l2 norm distance between two vectors; tanh(·) is the hyperbolic tangent function, mapping real numbers to the interval [-1, 1]; 1-tanh 2 (·) is used to measure the similarity between the query vector and the key vector in the Euclidean distance sense; ⊙ is the Hadamard product, i.e., the element-wise multiplication of two matrices or vectors; N is the total number of input vectors; W v This is the value vector obtained after mapping.
[0086] Step S302: Perform knowledge distillation enhancement processing on the initial multimodal features to obtain multimodal features.
[0087] Based on the above embodiments, in another embodiment provided in this disclosure, the above-mentioned knowledge distillation enhancement processing of the initial multimodal features to obtain multimodal features may include:
[0088] Acquire a pre-built domain knowledge base and multi-head attention mechanism;
[0089] Extract entity nodes related to the image set from the domain knowledge base;
[0090] A gated graph convolutional network is used to learn entity nodes and generate knowledge graph features;
[0091] By using weight allocation in the multi-head attention mechanism and knowledge graph features, the initial multimodal features are corrected to obtain multimodal features.
[0092] For example, in the knowledge distillation enhancement stage, feature correction of multimodal features M can be performed by accessing a pre-built domain knowledge base. First, a structured knowledge graph containing information on tourist attractions, architectural knowledge, and cultural common sense is maintained. When a specific domain image is detected, relevant knowledge nodes are automatically retrieved, and domain features are injected into the base features through a graph attention network. This significantly improves the accuracy of professional domain descriptions and avoids common-sense errors.
[0093] For example, constructing a multimodal representation M enhanced by knowledge distillation. e , can be represented as:
[0094]
[0095] Where M represents multimodal representation; Let be the mapping matrix from the attention key vectors to the distance matrix augmentation space, where d k Let be the dimension of the attention key vector; The bias term is used to enhance the distance matrix, thereby increasing the expressiveness and flexibility of the model.
[0096] Where, γ i ∈R represents the importance weight of the i-th attention head, reflecting its contribution to the enhancement of the distance matrix. Its calculation formula can be expressed as:
[0097]
[0098] Where sim(·,·) is the cosine similarity function; τ is the temperature parameter; To sum the enhancement terms for K attention heads.
[0099] k i This refers to extracting entity nodes related to images from external knowledge graphs (such as scenic spot graphs, city landmark graphs, and cultural knowledge graphs). The feature representation obtained by learning the knowledge representation through a gated graph convolutional network (GGCN) is calculated using the following formula:
[0100]
[0101] g i =σ(W g [k i ||x i ]+b g )
[0102] Among them, g i Let W be the gate vector at the i-th time step; σ(·) is the Sigmoid activation function; W gHere is the weight matrix for the gating mechanism; x i b is the i-th input vector in M for multimodal operation; g For the bias term of the gating mechanism; W k b is the attention key transformation matrix; k This is a bias term for the attention key transformation, used to increase the model's expressiveness and flexibility.
[0103] For example, this disclosure also continuously optimizes model performance during the knowledge distillation process, employing a dynamic evaluation mechanism to determine training termination conditions. During the distillation phase, multiple text generation quality metrics, such as BLEU (Bilingual Evaluation Understudy) and ROUGE (Recall-Oriented Understudy for Gisting Evaluation), are monitored in real time. These text generation quality metrics can measure the validation of generated M... e Similarity between the story and the reference samples accompanying the external knowledge graph. When M e Training can be stopped if the performance metrics do not show significant improvement in several consecutive evaluations (e.g., if they are below a certain preset threshold δ).
[0104] Based on this, through knowledge distillation, the model can establish an accurate visual-semantic mapping and an enhanced multimodal representation M. e By using external knowledge graphs for correction, potential recognition biases during model recognition are mitigated. This not only significantly reduces semantic bias in visual recognition but also ensures that the subsequently generated stories have the accuracy of descriptions within their respective professional domains.
[0105] Step S303: Generate the initial story text.
[0106] Based on the above embodiments, in another embodiment provided in this disclosure, the above-mentioned pre-constructed image story generation algorithm, based on multimodal features, generates target story text corresponding to the image set, which may include:
[0107] Obtain the descriptive text input by the user;
[0108] A composite reward function is used to generate initial story text based on multimodal features, image sets, and descriptive text.
[0109] The initial generated copy is fine-tuned using a domain-adaptive reward function to obtain the target story copy.
[0110] In this embodiment, a manifold regularized reinforcement learning framework can be used to achieve high-quality story generation. A multimodal enhanced decision process (MEDP) model is established to learn the story text generation strategy π. θ The model inputs include: a user-selected image set I as the visual narrative basis for generating story videos, and a knowledge-distilled multimodal enhanced representation M. e Provides optimized features, as well as descriptive text S provided when a user initiates the generation of a story video. u As a style guide, the model introduces a manifold regularization term during policy gradient optimization to ensure that the generated story text maintains a smooth transition and continuity in the semantic space, allowing adjacent concepts to connect naturally, the story development to conform to realistic logic, and the emotional changes to be gradual.
[0111] For example, maximizing the expected reward θ in the MEDP model * It can be represented as:
[0112]
[0113] Where S = (w1, w2, ..., w T R(·) represents the generated story text; R(·) represents the reward function; S ~ π θ To be based on strategy π θ The story sequence S is sampled; E[·] is the expectation operator, used to calculate the average value of the random variable.
[0114] Based on the above embodiments, in another embodiment provided in this disclosure, the composite reward function includes a visual relevance reward subfunction, a fluency reward subfunction, an information content reward subfunction, and a user preference reward subfunction.
[0115] In this embodiment, to ensure that the generated story text is more consistent, fluent, informative, and personalized, a composite reward function R(S) based on a manifold structure is also constructed. Figure 4 This is a schematic diagram of the structure of the composite reward function provided in an embodiment of this disclosure.
[0116] The composite reward function R(S) can be expressed as:
[0117] R(S)=λ1R co (S,I)+λ2R fl (S)+λ3R in (S)+λ4R pr (S,S u )
[0118] Where: R co(S,I) represents the visual relevance reward between the story text S and the image set I; R fl (S) is a bonus for the fluency of the story copywriting S, R fl The (S) function measures the fluency and coherence of the generated story in terms of syntax, logic, and semantics, thereby encouraging the model to generate coherent and logically sound stories; R in (S) is the information content reward for the story copy S, R in The (S) function measures the amount and richness of information contained in the generated story, encouraging the model to generate stories with rich content and high information content; Rpr(S,S) u (S) represents the story copy S relative to the user-provided description of needs S. u The preferred reward; λ1, λ2, λ3, and λ4 are the weight coefficients of the four sub-reward functions, used to control the importance of different quality dimensions in the total reward.
[0119] Specifically, R co (S,I) is used to ensure the consistency between the generated story and the input image. The calculation formula can be expressed as:
[0120]
[0121] in, M is the embedding vector of the t-th word; m Let be the multimodal representation of the m-th image; T be the length of the story sequence S, i.e., the number of words in the story; M be the number of regions extracted from each image in the image set I; TM be the total number of matching pairs between the story sequence and the image set; sim(·,·) is the similarity function.
[0122] R fl (S) Based on the fluidity of the story manifold, it encourages the generation of stories with smooth transitions on the manifold structure. The calculation formula can be expressed as:
[0123]
[0124] Where Δ is the Laplace-Beltrami operator of the manifold; λ L To control the intensity of punishment; Let be the embedding vector of the t-th word.
[0125] R in (S) Based on the information content output by the policy network, the model is encouraged to generate stories that are rich in information and highly unexpected. The calculation formula can be expressed as:
[0126]
[0127] Where, π θ Generate story strategies for the target; w tLet S be the t-th word in the story sequence S; 1:t-1 S is the subsequence of the story sequence S up to the (t-1)th word. 1:t-1 =(w1,w2,…,w t-1 M e This is a multimodal representation enhanced by knowledge distillation.
[0128] Rpr(S,S u This is used to measure the difference between the generated story S and the user-provided requirement description S. u The correlation between them can be expressed by the following formula:
[0129]
[0130] Where ||·|| is the manifold distance metric, which characterizes the distance between the story S and the user requirement description S. u The similarity between them. If S u If the value is empty, meaning the user selected an image without entering any text description, then R... pr (S,S u The value is 0; σ is the bandwidth or scaling parameter of the Gaussian kernel function.
[0131] Based on the above reward function, the parameters θ of the policy network are continuously optimized through reinforcement learning and other methods, ultimately yielding a high-quality story generation policy, forming the generated image story text S=(w1,w2,…,w T ).
[0132] Based on this, a composite reward function was constructed, which significantly improved the quality of story generation through multi-dimensional synergy. Among them, R... co (S,I) Strictly ensure text and image matching, and avoid vague descriptions; R fl (S) Maintain natural narrative logic and avoid abrupt transitions; R in (S) Ensure content density; R pr (S,S u It deeply adapts to users' individual characteristics. Furthermore, the weight coefficients of the four sub-reward functions allow users to adjust the story copy generation mode in real time. Through this refined multi-objective optimization mechanism, it can flexibly adapt to different scenarios such as travelogues, family photo albums, and workplace reports, ultimately outputting high-quality story content that combines accuracy, fluency, and personalization.
[0133] Step S304: Is fine-tuning required for a specific area?
[0134] If yes, proceed to step S305. If no, proceed to step S306.
[0135] Step S305: Adapt and fine-tune the initial story text for a specific domain.
[0136] To meet the specialized needs of different specific fields (such as travel, food, or Xiaohongshu style), the MEDP model can also include a domain adaptation mechanism in the story copy generation stage, by constructing a domain adaptation reward function R. (d) The model is guided to generate story copy that fits the characteristics of the field.
[0137] Domain Adaptive Reward Function R (d) It can be represented as:
[0138]
[0139] Where R(S) is the result of the composite reward function; To assess the relevance and professionalism of stories to the domain, based on domain corpora and knowledge graphs; λ d Weighting coefficients for domain-specific rewards.
[0140] In calculation When, it can be calculated using the following metric function:
[0141]
[0142] Where, μ S and μ d denoted as S and d, respectively, represent the probability distributions of story S and domain d in the semantic space; W2 represents the 2-Wasserstein distance.
[0143] By minimizing the W2 distance between the story and the domain, the model can generate content that is more in line with the characteristics of the target domain.
[0144] Step S306: Generate the target story text.
[0145] If the user chooses not to fine-tune the content for a specific area, the initial story text will be used as the target story text. If the user chooses to fine-tune the content for a specific area, the fine-tuned story text will be used as the target story text.
[0146] Based on the above embodiments, in another embodiment provided in this disclosure, the above-mentioned pre-constructed image story layout algorithm, which generates the target story layout corresponding to the image set based on multimodal features and target story text, may include:
[0147] A layout reward function is used to match the target layout template from a preset template library based on multimodal features and target story text.
[0148] Adjust the layout of the target layout template to obtain the target story layout.
[0149] In this embodiment, the Image Story Layout Algorithm (RL-APSL) models the video layout optimization problem as a multimodal augmented decision process (MEDP) and designs a dual reward mechanism of layout template matching and visual semantics. Through two-stage reinforcement learning, it achieves global layout selection and local fine-tuning, and trains the matching strategy with the highest matching degree among the image set, target story text and video layout template.
[0150] The input to the MEDP model includes: a predefined video layout template T = {T1, ..., T...} m The target story text S and the set of images I selected by the user to be generated as a picture story video.
[0151] Among them, the predefined video layout template library T = {T1,…,T m In the template, each template defines the general layout of the images, which can include grid layout and golden ratio layout, etc.
[0152] The decision-making process of the MEDP model may include the following elements:
[0153] Status: The currently selected layout template.
[0154] Action: Select another layout template from the template library.
[0155] Rewards: Quantify the matching degree between the new template and the current text and image content, and consider: visual adaptability, that is, the compatibility between the image composition and the template; semantic consistency, that is, whether the story text fragments and the narrative logic of the template are consistent; user preferences, that is, the degree of user preference for a certain type of template in historical data.
[0156] The decision-making process of the MEDP model can be represented as: "Let the spatial state s be the currently selected template T". i Action a is to select another template T. j Reward function R(T) i C) Measuring template T i Image-story text pair The matching degree. Where I k For the k-th image, S k "For the corresponding story fragment copy."
[0157] For example, the layout reward function R(T) based on visual-semantic consistency. i C) can be represented as:
[0158]
[0159] Where sim(·,·) is a measure of visual features φ(I) and semantic features. Similarity measures the semantic consistency between images and text;k For the k-th image; S k The corresponding story segment; while p(I,S|T) is the probability that the image-text pair (I,S) is well laid out in the template T, reflecting the compatibility between the template and the content.
[0160] Reward function R(T) i C) In the decision-making process of the MEDP model, the MEDP model is encouraged to choose a layout that can simultaneously improve visual-semantic consistency and template compatibility.
[0161] Because layout assessment involves a degree of subjectivity and complexity, in order to accurately calculate p(I) k ,S k |T i ), can be approximated using a strategy, p(I) k ,S k |T i It is modeled as a parameterized function, and the parameters of the function are learned from the training data through deep learning.
[0162] Specifically, a layout evaluation network f can be designed. θ (I,S,T), where θ represents the network parameters. This network takes an image set I, a target story text S, and a layout template T as input, and outputs a matching score within the range [0,1]. Let f θ (I,S,T) is considered an approximation of p(I,S|T): p(I,S|T)≈f θ (I,S,T).
[0163] First, image I, target story text S, and layout template T are mapped to a common semantic space, and then their compatibility is calculated in that space. Let φ(I)∈R d , and ψ(T)∈R d Let φ(·) represent the d-dimensional semantic vector representations of the image, text, and layout template, respectively, where φ(·) ψ(·) can be a pre-trained deep neural network, such as CNN or RNN.
[0164] Furthermore, a compatibility estimation formula based on an attention mechanism is designed, which can be specifically expressed as:
[0165]
[0166] Here, sigmoid(·) is the Sigmoid function, which compresses the compatibility score to the [0,1] interval; attn(·,·) is the attention scoring function, which is used to measure the relevance between two semantic vectors.
[0167] Specifically, the standardized similarity of g1(x) and g2(y) can be calculated first through the inner product, and then a trilinear function <·,·,·> can be introduced to model the ternary interaction between x, y and template T to improve the fusion capability between modalities. g1(·), g2(·), h1(·), h2(·), and h3(·) are all learnable mapping functions used to enhance the feature representation capability.
[0168] Reward function R(T) i C) Reinforcement learning of optimal layout selection strategy using Q-learning π * (s), this step can be expressed as:
[0169]
[0170] Where α∈(0,1] is used to control the degree of influence of each update on the Q-value estimation; the discount factor γ∈[0,1] is the importance of future rewards relative to current rewards.
[0171] Finally, based on the optimal layout selection strategy π * (s), matching the target layout template of the current image-text pair. It automatically selects the most suitable photo story video layout template for the user's current photo set.
[0172] Furthermore, based on visual expressiveness and semantic relevance, the target layout template can be optimized. Make minor adjustments to the layout. In the selected template... Based on this, the size and placement of each image in the layout template are automatically adjusted according to the visual expressiveness of the images and the semantic relevance of the images and text, forming the final image story video layout.
[0173] Specifically, first calculate I for each image. k Visual expressiveness score v k Image-text semantic relevance score r k Among them, the visual expressiveness score v k The scoring dimensions can include aesthetic quality scores such as whether the image is placed upside down, whether the image is clear when scaled, and whether it is obscured.
[0174] For example, the visual expressiveness score v k Image-text semantic relevance score r k It can be represented as:
[0175] v k =f v (φ(I k ))
[0176]
[0177] Among them, f v (·) and f r (·) represent the visual expressiveness assessment function and the image-text relevance assessment function, respectively, and their specific calculation formulas are as follows:
[0178]
[0179] Wherein, φ(I k )=φ1,φ2,…,φ N For image I k N local visual features, which can include the surface texture, shape, or color of the main scene, as well as facial expressions or clothing textures; α i The attention weight for the i-th local feature is calculated using the attention mechanism. is a layout relationship graph constructed based on image content; GNN(·,·) is a graph neural network applied to the layout relationship graph to learn the representation of nodes; MLPv(·), MLPr(·), MLPφ(·) and MLPφ(·) are multilayer perceptron networks used for feature transformation and score calculation.
[0180] Based on the above embodiments, in another embodiment provided in this disclosure, the above-mentioned layout adjustment processing of the target layout template to obtain the target story layout may include:
[0181] The image set is obtained in the target layout template, and the first visual expressiveness score and the first image-text semantic relevance score are obtained before adjustment, and the second visual expressiveness score and the second image-text semantic relevance score are obtained after adjustment.
[0182] Obtain the difference in visual performance scores between the first and second visual performance scores;
[0183] Obtain the difference in semantic relevance scores between the first and second text-image semantic relevance scores.
[0184] The target story layout is determined based on the difference in visual expressiveness score and the difference in semantic relevance score between text and images.
[0185] In this embodiment, the local video layout fine-tuning problem is modeled as another MEDP model. The decision-making process of the MEDP model can be represented as: "State s is based on template..." The initial layout is defined by action 'a', which involves swapping the positions of two images based on the initial layout (s) or adjusting their sizes. The reward function R(s,a) comprehensively considers both the improvement in visual appeal and semantic relevance before and after the layout adjustment.
[0186] The reward function R(s,a) can be expressed as:
[0187]
[0188] in, and For image I k Visual expressiveness and semantic relevance scores under layout s; and λ represents the corresponding score under layout a; λ is the balance factor.
[0189] Local fine-tuning strategy using Q-learning reinforcement learning Set a termination condition for layout iteration optimization, and output the target story layout L after the termination condition is met. * Termination conditions may include: an iteration count threshold or a reward increase threshold.
[0190] Therefore, by adopting the reward function R(s,a), the overall visual expressiveness and semantic relevance of the picture story video can be significantly improved after fine-tuning the local video layout.
[0191] Based on this, the RL-APSL algorithm learns the optimal layout selection strategy π. * (s) and local fine-tuning strategy The entire process, from global template matching to local fine-tuning, has been automated. Furthermore, a matching algorithm integrating visual features, textual semantics, and user preferences has been developed, ensuring that the matched templates not only match the content of the user's selected image set but also reflect the user's personalized characteristics.
[0192] One or more technical solutions provided in the exemplary embodiments of this disclosure, through the synergistic optimization of multimodal feature extraction and composite reward functions, enable the generated target story text to accurately convey image information and meet user expectations, thereby solving technical problems such as insufficient multimodal fusion, lack of innovation in generated stories, and domain adaptability. Secondly, the image story layout algorithm achieves dual-objective optimization through a layout reward function. Visual expressiveness evaluation covers compositional completeness, clarity, and dynamic coordination, while semantic relevance focuses on image-text focus matching, emotional consistency, and temporal logic, thereby addressing the technical problem of monotonous and unattractive video layouts.
[0193] Therefore, the video generation method for picture stories provided in the exemplary embodiments of this disclosure can, through multimodal features, picture story generation algorithms, and picture story layout algorithms, enable the generated video content to maintain a high degree of consistency in overall style with the picture set selected by the user, and can automatically optimize video parameters to adapt to the posting requirements of different social media platforms, thereby significantly reducing the production threshold for high-quality video content.
[0194] The foregoing primarily describes the solutions provided by exemplary embodiments of this disclosure. It is understood that, in order to achieve the above functions, the electronic device includes corresponding hardware structures and / or software modules for performing each function. Those skilled in the art should readily recognize that, based on the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, this disclosure can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.
[0195] The exemplary embodiments of this disclosure can divide the electronic device into functional units according to the above method examples. For example, each function can be divided into a separate functional module, or two or more functions can be integrated into a single processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in the exemplary embodiments of this disclosure is illustrative and only represents one logical functional division; other division methods may be used in actual implementation.
[0196] By dividing each functional module according to its corresponding function, an exemplary embodiment of this disclosure provides a video generation apparatus for picture stories, which can be a server or a chip applied to a server. Figure 5 This is a schematic block diagram illustrating the functional modules of the video generation apparatus for picture stories provided in embodiments of this disclosure. Figure 5 As shown, the video generation device 500 for this photo story includes:
[0197] The data acquisition module 510 is used to acquire the image set selected by the user and extract the multimodal features of the image set;
[0198] The data processing module 520 is used to generate target story text corresponding to the image set based on the multimodal features using a pre-constructed image story generation algorithm; the image story generation algorithm includes a composite reward function, which is used to evaluate the visual relevance, fluency, information content and user preference relevance of the story text;
[0199] The data processing module 520 is further configured to use a pre-constructed image story layout algorithm to generate a target story layout corresponding to the image set based on the multimodal features and the target story text; the image story layout algorithm includes a layout reward function, which is used to evaluate the visual expressiveness and semantic relevance of the story layout.
[0200] The data processing module 520 is also used to generate a picture story video based on the picture set, the target story text, and the target story layout.
[0201] In another embodiment provided in this disclosure, the image story generation algorithm further includes a domain adaptation reward function; the data processing module 520 is further configured to obtain descriptive text input by the user; generate initial story text based on the multimodal features, the image set, and the descriptive text using the composite reward function; and perform domain fine-tuning processing on the initial generated text using the domain adaptation reward function to obtain the target story text.
[0202] In another embodiment provided in this disclosure, the data processing module 520 further includes: the composite reward function includes a visual relevance reward subfunction, a fluency reward subfunction, an information content reward subfunction, and a user preference reward subfunction.
[0203] In another embodiment provided in this disclosure, the data processing module 520 is further configured to use the layout reward function to match a target layout template in a preset template library based on the multimodal features and the target story text; and to perform layout adjustment processing on the target layout template to obtain the target story layout.
[0204] In another embodiment provided in this disclosure, the data processing module 520 is further configured to obtain, respectively, the first visual expressiveness score and the first image-text semantic relevance score corresponding to the image set before adjustment in the target layout template, and the second visual expressiveness score and the second image-text semantic relevance score corresponding to the image set after adjustment; obtain the visual expressiveness score difference between the first visual expressiveness score and the second visual expressiveness score; obtain the image-text semantic relevance score difference between the first image-text semantic relevance score and the second image-text semantic relevance score; and determine the target story layout based on the visual expressiveness score difference and the image-text semantic relevance score difference.
[0205] In another embodiment provided in this disclosure, the data processing module 520 is further configured to extract visual features and text features of the image set; fuse the visual features and the text features to obtain initial multimodal features; and perform knowledge distillation enhancement processing on the initial multimodal features to obtain multimodal features.
[0206] In another embodiment provided in this disclosure, the data processing module 520 is further configured to acquire a pre-built domain knowledge base and a multi-head attention mechanism; extract entity nodes related to the image set from the domain knowledge base; learn the entity nodes using a gated graph convolutional network to generate knowledge graph features; and perform correction processing on the initial multimodal features through the weight allocation in the multi-head attention mechanism and the knowledge graph features to obtain the multimodal features.
[0207] Exemplary embodiments of this disclosure also provide an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor. The memory stores a computer program executable by the at least one processor, the computer program being executed by the at least one processor to cause the electronic device to perform a method according to an embodiment of this disclosure.
[0208] Exemplary embodiments of this disclosure also provide a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when executed by a computer's processor, is used to cause the computer to perform a method according to embodiments of this disclosure.
[0209] Figure 6 The structural block diagram of the electronic device provided in the embodiments of this disclosure will now be described as follows: An electronic device 600 that can serve as a server or client of this disclosure is an example of a hardware device that can be applied to various aspects of this disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the disclosure described and / or claimed herein.
[0210] like Figure 6 As shown, the electronic device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) 602 or a computer program loaded from a storage unit 608 into a random access memory (RAM) 603. The RAM 603 may also store various programs and data required for the operation of the electronic device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.
[0211] Multiple components in electronic device 600 are connected to I / O interface 605, including: input unit 606, output unit 607, storage unit 608, and communication unit 609. Input unit 606 can be any type of device capable of inputting information to electronic device 600. Input unit 606 can receive input digital or character information and generate key signal inputs related to user settings and / or function control of electronic device. Output unit 607 can be any type of device capable of presenting information and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. Storage unit 608 may include, but is not limited to, disks and optical discs. Communication unit 609 allows electronic device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth™ devices, WiFi devices, WiMax devices, cellular communication devices, and / or the like.
[0212] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above. The various methods described above can all be implemented as computer software programs, which are tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 600 via ROM 602 and / or communication unit 609.
[0213] Figure 7 The diagram illustrates a computer program product provided in an embodiment of this disclosure. An exemplary embodiment of this disclosure also provides a computer program product 700, including a computer program 701, wherein the computer program 701, when executed by a computer's processor, is used to cause the computer to perform a method according to an embodiment of this disclosure.
[0214] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable data processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.
[0215] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0216] As used in this disclosure, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) for providing machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal for providing machine instructions and / or data to a programmable processor.
[0217] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0218] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0219] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0220] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of this disclosure are performed, in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, a terminal, a user equipment, or other programmable device. The computer program or instructions can be stored in a computer-readable storage medium or transferred from one computer-readable storage medium to another. For example, the computer program or instructions can be transferred from one website, computer, server, or data center to another website, computer, server, or data center via wired or wireless means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center integrating one or more available media. The available medium can be a magnetic medium, such as a floppy disk, hard disk, or magnetic tape; it can also be an optical medium, such as a digital video disc (DVD); or it can be a semiconductor medium, such as a solid-state drive (SSD).
[0221] Although this disclosure has been described in conjunction with specific features and embodiments, it will be apparent that various modifications and combinations can be made therein without departing from the spirit and scope of this disclosure. Accordingly, this specification and drawings are merely exemplary illustrations of the disclosure as defined by the appended claims and are to be considered as covering any and all modifications, variations, combinations, or equivalents within the scope of this disclosure. It is obvious that those skilled in the art can make various alterations and modifications to this disclosure without departing from its spirit and scope. Thus, this disclosure is also intended to include any such modifications and modifications that fall within the scope of the claims of this disclosure and their equivalents.
Claims
1. A method for generating video from a picture story, characterized in that, The method includes: Obtain the image set selected by the user and extract the multimodal features of the image set; A pre-constructed image story generation algorithm is used to generate target story text corresponding to the image set based on the multimodal features. The image story generation algorithm includes a composite reward function, which is used to evaluate the visual relevance, fluency, information content, and user preference relevance of the story text. The composite reward function includes a visual relevance reward sub-function, a fluency reward sub-function, an information content reward sub-function, and a user preference reward sub-function. A pre-constructed image story layout algorithm is used to generate the target story layout corresponding to the image set based on the multimodal features and the target story text. The image story layout algorithm includes a layout reward function, which is used to evaluate the visual expressiveness and semantic relevance of the story layout. Based on the image set, the target story text, and the target story layout, generate an image story video; The pre-constructed image story layout algorithm, based on the multimodal features and the target story text, generates a target story layout corresponding to the image set, including: using the layout reward function, matching a target layout template in a preset template library based on the multimodal features and the target story text; and performing layout adjustment processing on the target layout template to obtain the target story layout. The step of adjusting the target layout template to obtain the target story layout includes: obtaining the first visual expressiveness score and the first image-text semantic relevance score of the image set in the target layout template before adjustment, and the second visual expressiveness score and the second image-text semantic relevance score after adjustment; obtaining the visual expressiveness score difference between the first visual expressiveness score and the second visual expressiveness score; obtaining the image-text semantic relevance score difference between the first image-text semantic relevance score and the second image-text semantic relevance score; and determining the target story layout based on the visual expressiveness score difference and the image-text semantic relevance score difference.
2. The method according to claim 1, characterized in that, The image story generation algorithm further includes a domain-adaptive reward function; the pre-constructed image story generation algorithm, based on the multimodal features, generates target story text corresponding to the image set, including: Obtain the descriptive text input by the user; Using the composite reward function, an initial story text is generated based on the multimodal features, the image set, and the descriptive text. The initial story text is fine-tuned using the domain-adaptive reward function to obtain the target story text.
3. The method according to claim 1, characterized in that, The extraction of multimodal features from the image set includes: Extract the visual and textual features of the image set; The visual features and the text features are fused to obtain the initial multimodal features; The initial multimodal features are subjected to knowledge distillation enhancement processing to obtain multimodal features.
4. The method according to claim 3, characterized in that, The step of performing knowledge distillation enhancement on the initial multimodal features to obtain multimodal features includes: Acquire a pre-built domain knowledge base and multi-head attention mechanism; Extract entity nodes related to the image set from the domain knowledge base; The entity nodes are learned using a gated graph convolutional network to generate knowledge graph features; The initial multimodal features are corrected using the weight allocation in the multi-head attention mechanism and the knowledge graph features to obtain the multimodal features.
5. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the method according to any one of claims 1-4.
6. A computer-readable storage medium having a computer program / instructions stored thereon, characterized in that, When the computer program / instructions are executed by the processor, they implement the method of any one of claims 1-4.
7. A computer program product comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, they implement the method of any one of claims 1-4.
Citation Information
Patent Citations
Text-driven automatic image generation method and system
CN119444935A
Photo album generation method, device and equipment
CN119850791A