Environmentally Aware Video Generation Method, System, Server, and Media
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-03
- Publication Date
- 2026-08-11
AI Technical Summary
[0002]在视频生成领域,现有的视频大都为基于采集得到的环境数据生成局限于采集数据的相关推广视频或者相关介绍视频,但是由于生成的视频为基于采集到的数据进行随机生成,因此常常无法得到采集人员想要体现的内容与信息,导致推广视频无法满足采集人员的需求
本申请通过获取目标人员的眼动信息并基于该信息从第一影像信息中提取多个第二影像信息,进而生成对应的目标描述文本。由于眼动信息能够反映目标人员对影像中特定内容的关注,所以基于眼动信息提取的影像及其描述文本更能体现采集人员所关注和想要体现的内容。将上述信息输入视频生成大模型后,生成的视频能够满足采集人员的需求,解决了现有技术中视频无法满足需求的问题。本申请不仅获取了第一影像信息并生成其影像描述文本,还基于目标眼动信息提取多个第二影像信息并生成对应的影像描述文本,最后将两者以及目标描述文本共同输入视频生成大模型。丰富了视频所包含的内容与信息,提升了视频生成的准确性,提升用户体验。
Smart Images

Figure CN120711257B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of video generation technology, and more specifically, relates to a video generation method and system, server, and medium based on environment awareness. Background Technology
[0002] In the field of video generation, most existing videos are promotional or introductory videos based on collected environmental data. However, since the generated videos are randomly generated based on the collected data, they often fail to convey the content and information that the data collectors want to present, resulting in promotional videos that cannot meet the needs of the data collectors.
[0003] Therefore, an environment-aware video generation method is needed. Summary of the Invention
[0004] The purpose of this application is to provide an environment-aware video generation method, system, server, and medium to improve the accuracy of video generation and enhance user experience.
[0005] A first aspect of this application provides an environment-aware video generation method, applied to a server corresponding to a wearable device, including: Acquire first image information and target eye movement information. The first image information is the image information of the target area collected by the wearable device, and the target eye movement information is the eye movement information of the target person collected by the wearable device. The target person is the person wearing the wearable device. Generate image description text corresponding to the first image information; Based on the target eye-tracking information, extract multiple second image information from the first image information and generate image description text corresponding to the multiple second image information; Based on preset text conversion rules, the image description text of multiple second image information is converted into target description text; At least the image description text and target description text of the first image information are input into the large video generation model to obtain the video corresponding to the first image information.
[0006] A second aspect of this application provides an environment-aware video generation system applied to a server corresponding to a wearable device. The system includes: The information acquisition module is used to acquire first image information and target eye movement information. The first image information is the image information of the target area collected by the wearable device, and the target eye movement information is the eye movement information of the target person collected by the wearable device. The target person is the person wearing the wearable device. The first text generation module is used to generate image description text corresponding to the first image information; The second text generation module is used to extract multiple second image information from the first image information based on the target eye movement information, and generate image description text corresponding to the multiple second image information. The text conversion module is used to convert the image description text of multiple second image information into target description text based on preset text conversion rules; The video generation module is used to input at least the image description text and target description text of the first image information into the large video generation model to obtain the video corresponding to the first image information.
[0007] A third aspect of this application provides a server, including a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that the processor executes the computer program to implement the steps of the method as described in any one of claims 1 to 7.
[0008] A fourth aspect of this application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the above-described environment-aware video generation method.
[0009] The beneficial effects of the environment-aware video generation method, system, server, and medium provided in this application embodiment are as follows: This application acquires the eye-tracking information of a target person and extracts multiple second-image information from a first-image dataset based on this information, thereby generating corresponding target description text. Since eye-tracking information reflects the target person's attention to specific content in the image, the images and their description text extracted based on eye-tracking information better reflect the content that the data acquisition personnel are interested in and want to convey. After inputting the above information into a large-scale video generation model, the generated video meets the needs of the data acquisition personnel, solving the problem that existing technologies cannot meet the required video requirements. This application not only acquires the first-image information and generates its image description text, but also extracts multiple second-image information based on the target's eye-tracking information and generates corresponding image description text. Finally, both of these, along with the target description text, are input into a large-scale video generation model. This enriches the content and information contained in the video, improves the accuracy of video generation, and enhances the user experience. Attached Figure Description
[0010] To more clearly illustrate the technical solutions in the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0011] Figure 1A flowchart illustrating an embodiment of the video generation method based on environment awareness provided in this application; Figure 2 This is a structural block diagram of an environment-aware video generation system provided in an embodiment of this application; Figure 3 This is a schematic block diagram of a server provided in one embodiment of this application. Detailed Implementation
[0012] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of this application. However, those skilled in the art will understand that this application may also be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods have been omitted so as not to obscure the description of this application with unnecessary detail.
[0013] To make the objectives, technical solutions, and advantages of this application clearer, the following description will be provided in conjunction with the accompanying drawings and specific embodiments.
[0014] Please refer to Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of a video generation method based on environment awareness provided in this application. The method is executed by a server corresponding to a wearable device. The server can be a distributed server, a cloud server, or an edge server, etc. The method includes: S101: Acquire first image information and target eye movement information. The first image information is image information of the target area collected by the wearable device, and the target eye movement information is eye movement information of the target person collected by the wearable device. The target person is the person wearing the wearable device.
[0015] In this embodiment, the first image information is the raw image data of the target area collected by a wearable device (such as smart glasses, helmet cameras, etc.). For example, when a user wears smart glasses to shop in a supermarket, the glasses' camera captures a continuous image of the shelves, goods, and surrounding environment. The target eye movement information is the eye movement data of the user (target person) wearing the device, which may include the fixation point position, pupil movement, blink frequency, etc. The target area can be the spatial range of the image collected by the wearable device, and can be any area. The influence information can be an image or video; in this embodiment, video information is preferred.
[0016] In this embodiment, the first image information can be acquired by the image acquisition device of the wearable device, such as a camera. The target eye movement information can be acquired by the eye tracking device of the wearable device. The server can essentially be a computer that can perform video generation tasks, and the wearable device can perform information acquisition tasks.
[0017] S102: Generate image description text corresponding to the first image information.
[0018] In this embodiment, the first image information can be input into the semantic extraction model. The semantic extraction model can perform operations such as target detection, semantic segmentation and instance recognition on the input image information, and finally obtain the corresponding text information. The specific principles and processes will not be repeated in this embodiment.
[0019] For example, the first input image information is a scene of a user selecting goods in front of a supermarket shelf. The corresponding image description text could be: object: shelf, beverage bottle, fruit, shopping basket; action: reaching out to pick up, looking; scene: supermarket food section.
[0020] S103: Extract multiple second image information from the first image information based on the target eye movement information, and generate image description text corresponding to the multiple second image information.
[0021] In this embodiment, the second image information can be understood as the image information corresponding to the position where the target person's attention is in the first image information. That is, the image information corresponding to the position where the target person's gaze lingers for a long time during the acquisition of the first image information. Long-term lingering can refer to exceeding a preset lingering time, which can be set based on personal preference. The process of generating image description text corresponding to multiple second image information based on multiple second image information can be the same as S102, and will not be described again in this embodiment.
[0022] In this embodiment, multiple pieces of second image information can be extracted from the first image information in the following way: extracting multiple pieces of second image information from the first image information based on target eye movement information, including: The eye focus area of the target person is determined based on the target eye movement information in the first time period; the first time period is the time period during which the wearable device collects the eye movement information of the target person, and the start and end times of the first image information collection are consistent with the start and end times of the target eye movement information collection. Multiple second image information is extracted from the first image information based on the eye focus area of the target person in the first time period.
[0023] In this embodiment, the target eye movement information contains the target person's eye movement information during the data acquisition period, which may include fixation point coordinates, fixation duration, pupil changes, etc. It is important to note that the start and end times of the first image information acquisition are consistent with the start and end times of the target eye movement information acquisition. Therefore, the eye movement information and image information need to be aligned using timestamps. The calculation of the focal region can be performed as follows: Spatial dimension: converting eye movement coordinates (such as normalized (x, y) values) into image pixel coordinates to determine the specific location of the user's gaze. Temporal dimension: statistically analyzing all fixation points within the first time period, and identifying high-frequency focal regions using a clustering algorithm (such as DBSCAN). For example, if a user repeatedly gazes at a Coca-Cola bottle on a supermarket shelf within 10 seconds, this area is marked as the "eye focus region."
[0024] In this embodiment, the pixel coordinates of the focused area can be mapped to each frame of the first image to determine the rectangular area to be cropped, such as a 200×200 pixel window centered on the gaze point. All image frames within the first time period are traversed, and the focused area at the same position is extracted from each frame to form multiple second image information. Alternatively, in this embodiment, the focused areas corresponding to more than a first number of influence frames representing the same focused area in all image frames can be used as the second image information. That is, when the user wears the wearable device (during the acquisition of the first image information), the position where their eyes focus for a long time is determined as the second image information. In other words, the second image information can be understood as the image information that the user pays attention to for a long time in the first image information.
[0025] S104: Based on preset text conversion rules, convert the image description text of multiple second image information into target description text; the target description text is used to instruct the server to generate video.
[0026] In this embodiment, the preset text conversion rule can be a pre-defined algorithm or logic used to convert the original description text into a more structured format that is more suitable for video generation, or it can be understood as converting multiple second image description texts into one description text, which is the image description text that is relatively important for video generation among the multiple second image description texts.
[0027] In this embodiment, the image description text of multiple second image information can be converted into target description text in the following way: Based on preset text conversion rules, the image description text of multiple second image information is converted into target description text, including: For each image description text of the second image information, obtain the visual part-of-speech words, visual entity words, structural words, and descriptive words in the image description text of the second image information; Repeated words in the visual part-of-speech vocabulary, visual entity vocabulary, structural vocabulary, and descriptive vocabulary corresponding to each second image information are converted into target description text.
[0028] In this embodiment, visual part-of-speech words can be words describing visual actions or states, such as seeing, touching, moving, and remaining still. Visual entity words can be words describing specific objects or scenes, such as apples, tables, supermarkets, and cars. Structural words can be the position and action of the main subject in the image, background, shot size, angle, style reference, generation time, etc. Adjective words can be modifiers describing the attributes of objects, such as red, round, smooth, and tall.
[0029] In this embodiment, converting repeated words into target descriptive text is considered because if a large number of repeated words appear in the descriptive text corresponding to multiple second image information, it indicates that the target personnel are focusing on or paying close attention to this area, and that the target personnel want this area to be highlighted in the generated video. Therefore, repeated words can be converted into target descriptive text. The conversion process can be simple concatenation or classification based on the aforementioned word types, ultimately resulting in target descriptive text. The target descriptive text contains the content that the target personnel want to be reflected in the generated video, and therefore can be used to guide the server in generating the video.
[0030] S105: Input at least the image description text and target description text of the first image information into the large video generation model to obtain the video corresponding to the first image information.
[0031] In this embodiment, the image description text of the first image information can be understood as a natural language description of the original panoramic image. It can also be understood as the first image information serving as the background in the generated video. The target description text can be understood as a structured instruction generated based on keywords extracted from the user's eye movement focus, highlighting the key content that the user is interested in. The large-scale video generation model can be a deep learning-based text-to-video generation model (such as Pika Labs or Stable VideoDiffusion), generating video content through the text prompt. The video corresponding to the first image information is a video generated based on the input text, integrating the panoramic scene and the user's focus, which can reflect the scene and the key points the target audience wants to highlight in the video.
[0032] As can be seen from the above, this application obtains the eye-tracking information of the target person and extracts multiple second-image information from the first image information based on this information, thereby generating corresponding target description text. Since eye-tracking information can reflect the target person's attention to specific content in the image, the image and its description text extracted based on eye-tracking information better reflect the content that the acquisition personnel are concerned with and want to convey. After inputting the above information into the video generation model, the generated video can meet the needs of the acquisition personnel, solving the problem that the video in the prior art cannot meet the requirements. This application not only obtains the first image information and generates its image description text, but also extracts multiple second-image information based on the target eye-tracking information and generates corresponding image description text. Finally, both of these, along with the target description text, are input into the video generation model. This enriches the content and information contained in the video, improves the accuracy of video generation, and enhances the user experience.
[0033] In one embodiment of this application, at least the image description text and target description text of the first image information are input into a large video generation model to obtain the video corresponding to the first image information, including: Determine the richness score of the image description text of the first image information; If the richness score of the image description text of the first image information is greater than the specified score, the image description text of the first image information and the target description text are input into the video generation model to obtain the video corresponding to the first image information.
[0034] In this embodiment, the richness score of the image description text of the first image information can be determined based on multiple dimensions. For example, it can be calculated based on information density, structural integrity, and detail granularity, as detailed below. Alternatively, the richness score can be determined solely based on information density or the length of the image description text of the first image information. This embodiment does not impose any limitations on this approach. The assigned score can be set based on experience or personal preference.
[0035] In this embodiment, considering that when the richness score of the image description text of the first image information is greater than a specified score, it indicates that the richness of the image description text of the first image information is relatively large, that is, the richness of the image description text of the first image information is sufficient to generate the desired video. Therefore, a video can be generated based on the image description text of the first image information and the target description text.
[0036] Correspondingly, in one embodiment of this application, at least the image description text and target description text of the first image information are input into the video generation model to obtain the video corresponding to the first image information, and the method further includes: If the richness score of the image description text of the first image information is less than or equal to the specified score, the scene to which the first image information belongs is determined based on the image description text, and the storyboard description text corresponding to the scene to which it belongs is determined as the storyboard description text of the first image information. The image description text, target description text, and storyboard description text of the first image information are input into the video generation model to obtain the video corresponding to the first image information.
[0037] In this embodiment, when the richness score of the image description text of the first image information is less than or equal to a specified score, it indicates that the richness of the image description text of the first image information is low, the text description is too brief, and it cannot directly support video generation. At this time, the scene to which the image belongs can be determined based on the image description text. For example, keywords (such as "park", "supermarket", "kitchen") or semantic features in the text can be analyzed by semantic model to determine the actual scene type corresponding to the first image. Different scenes have different storyboard description texts pre-stored, which include the common shot order and content in that scene. For example, the storyboard description text corresponding to the supermarket scene may include "entrance panoramic view → shelf medium shot → product close-up". The storyboard description text is used to supplement the missing structured storyboard information in the original description to ensure the coherence of video generation. It does not depend on the details of the original description and provides a general storyboard framework based on the commonality of the scene.
[0038] In this embodiment, after determining the storyboard description text of the first image information, the image description text, target description text, and storyboard description text of the first image information can be input into the video generation model to provide more information input to the video generation model, so as to generate a higher quality video that better meets the user's needs.
[0039] As can be seen from the above, this application determines different video generation strategies by judging the richness score of the image description text of the first image information. When the richness score is greater than a specified score, it indicates that the richness of the image description text is sufficient to provide enough accurate information for video generation. In this case, only the image description text of the first image information and the target description text need to be input into the large-scale video generation model to generate a video that meets the requirements, ensuring the basic quality of video generation, avoiding unnecessary additional processing steps, simplifying the video generation process, and improving video generation efficiency. However, when the richness score is less than or equal to the specified score, it indicates that the image description text is relatively brief, and direct use may lead to poor video generation quality. In this case, by determining the scene to which the first image information belongs and introducing the storyboard description text corresponding to that scene, structured storyboard information is added to the large-scale video generation model, making up for the deficiencies of the original description, ensuring the coherence and completeness of video generation, and thus guaranteeing the quality of the final generated video.
[0040] In one embodiment of this application, determining the richness score of the image description text of the first image information includes: Input the image description text of the first image information into the first semantic model to obtain the visual part-of-speech words and visual entity words in the image description text of the first image information. The visual part-of-speech words are words in the image description text of the first image information that have a greater correlation with the preset visual part-of-speech words than the preset correlation. The visual entity words are words in the image description text of the first image information that have a greater correlation with the preset visual entity words than the preset correlation. The information density score is determined based on the proportion of the total number of visual part-of-speech words and visual entity words to the total number of words in the image description text of the first image information. Input the image description text of the first image information into the second semantic model to obtain the number of structural words in the image description text of the first image information. Structural words are words in the image description text of the first image information whose relevance to preset structural words is greater than preset relevance. The structural integrity score is determined based on the number of structural vocabulary categories. Extract detailed words from the image description text of the first image information. Detailed words are descriptive words in the image description text of the first image information. The detail granularity score is determined based on the proportion of detail words in the image description text of the first image information to the total number of words in the image description text of the first image information. The richness score is obtained by weighting the information density score, structural integrity score, and detail granularity score.
[0041] In this embodiment, both the first semantic model and the second semantic model are trained models. The training dataset for the first semantic model consists of a large number of input words and whether each word is a visual part-of-speech word or a visual entity word. The training dataset for the second semantic model consists of a large number of input words and whether each word is a structured word.
[0042] In this embodiment, the first semantic model can identify visual parts-of-speech words and visual entity words in the text. For example, given the input description "the user picks up a red apple," the first semantic model identifies "pick up" as a visual action word and "user" and "apple" as visual entity words. Visual parts-of-speech words can be words describing visual actions or states (such as "look," "take," "move"), and their relevance to a preset vocabulary (such as {gazing, picking up, placing}) must exceed a threshold, such as a cosine similarity > 0.8. The second semantic model can identify structural words in the text and evaluate the completeness of the description. Visual entity words can be words describing specific objects or scenes, and their relevance to a preset entity vocabulary must exceed a threshold. The preset relevance can be set based on experience.
[0043] In this embodiment, the method can be consistent with the aforementioned method for determining visual parts of speech, visual entity words, or structural words. For example, it can be based on a third semantic model, whose training dataset consists of a large number of input words and whether each word is a detail word. It should be noted that the first semantic model, the second semantic model, and the third semantic model can be the same model or the same model. That is, the training dataset should include a large number of words and whether each word is a visual part of speech, visual entity word, structural word, or detail word. In this embodiment, the first semantic model, the second semantic model, and the third semantic model are not limited.
[0044] In this embodiment, the information density score measures the proportion of "effective visual information" in the description; a higher value indicates richer content. The structural integrity score assesses whether the description contains important information about video generation, such as the main subject, background, style references, and logical relationships such as space and time, ensuring content coherence. The detail granularity score measures the level of detail in the description; more modifiers result in a stronger visual impact.
[0045] In this embodiment, the information density score can be calculated as (number of visual part-of-speech words + number of visual entity words) / total number of words in the text. The types of structural integrity score are the aforementioned main subject, background, style reference, etc. The more types there are, the more complete the structure, which can be determined based on a preset mapping relationship. The detail granularity score can be calculated as: number of detail words / total number of words in the text.
[0046] In this embodiment, before performing weighted calculations, in order to ensure consistency of evaluation dimensions, the scores are normalized. The weights for weighted calculations can be set based on actual application scenarios or based on experience to set fixed weight allocations.
[0047] As can be seen from the above, this application quantifies the richness of the image description text of the first image information through multiple dimensions. It utilizes a first semantic model to identify visual part-of-speech and visual entity words, and calculates an information density score based on their proportion, which can measure the content of effective visual information in the description text. It uses a second semantic model to identify structural words, determines a structural integrity score based on the number of categories, and assesses whether the description contains the important information and logical relationships required for video generation. It extracts detail words and calculates their proportion to obtain a detail granularity score, clearly measuring the level of detail in the description. This multi-dimensional and refined evaluation method can accurately quantify the richness of the description text, contributing to improved video generation quality.
[0048] In one embodiment of this application, considering that the application scenarios of the generated videos are different, the core expression requirements of the videos are also different. Therefore, this embodiment of the application adjusts the calculation weight of the richness score of the image description text of the aforementioned first image information based on different application scenarios. For example, in response to the application scenario of the video corresponding to the first image information being news, the weight corresponding to the information density score is increased based on the first information step size, the weight corresponding to the structural integrity score is decreased based on the first structure step size, and the weight corresponding to the detail granularity score is decreased based on the first detail step size; the first information step size is equal to the sum of the first structure step size and the first detail step size. In response to the fact that the application scenario of the video corresponding to the first image information is a narrative type, the weight corresponding to the structural integrity score is increased based on the second structural step size, the weight corresponding to the information density score is decreased based on the second information step size, and the weight corresponding to the detail granularity score is decreased based on the second detail step size; the second structural step size is equal to the sum of the second information step size and the first detail step size. In response to the application scenario of the video corresponding to the first image information being promotional, the weight corresponding to the detail granularity score is increased based on the third detail step length, the weight corresponding to the information density score is decreased based on the third information step length, and the weight corresponding to the structural integrity score is decreased based on the third structure step length. The third detail step length is equal to the sum of the third information step length and the third structure step length.
[0049] In this embodiment, the application scenario of the video corresponding to the first image information can be that the target person sends it to the server through a wearable device based on their own needs. The application scenarios include, but are not limited to, news, drama and promotion in this application. When the scenario set by the target person is not one of the above three types, the richness score can be calculated based on a preset fixed weight.
[0050] In this embodiment, considering that the core requirement of news videos is to quickly and accurately convey event information, the weight of the information density score should be increased. High information density ensures concise content with sufficient information. The core requirement of narrative videos is coherent narration and logically consistent plot. The plot needs to rely on structural terms such as timelines and causal relationships to construct the story framework. High structural integrity avoids narrative gaps; therefore, the weight of the structural integrity score should be increased. The core requirement of promotional videos is to enhance visual appeal and the impact of details. Promotional content needs to stimulate the user's senses through detailed descriptions. High detail granularity enhances the persuasiveness of product selling points; therefore, the weight of the detail granularity score should be increased. In this embodiment, the specific values of the first information step, the second structural step, and the third detail step can be set based on experience. The specific data of the first structural step, the first detail step, the second information step, the second detail step, the third information step, and the third structural step can be set based on preferences, with the aim of ensuring that the sum of the weights is 1.
[0051] As can be seen from the above, this application considers the differences in the core expressive needs of generated videos under different application scenarios. By adjusting the calculation weight of the richness score of the image description text of the first image information based on different application scenarios, it achieves precise adaptation of video generation to specific application scenarios. For example, news videos emphasize the rapid and accurate delivery of event information; increasing the information density score weight ensures that the video content is concise and information-rich, meeting the timeliness and accuracy requirements of news dissemination. Drama videos emphasize coherent narrative and logically consistent plots; increasing the structural integrity score weight helps avoid narrative gaps and build a complete story framework. Promotional videos aim to enhance visual appeal and the impact of details; increasing the detail granularity score weight enhances the persuasiveness of product selling points and attracts audience attention. The aforementioned targeted weight adjustment method enables the generated videos to better meet the specific needs of different application scenarios, improve the accuracy of video generation, and enhance the user experience.
[0052] Corresponding to the environment-aware video generation method in the above embodiments, Figure 2 This is a structural block diagram of an environment-aware video generation system according to an embodiment of this application. For ease of explanation, only the parts relevant to the embodiment of this application are shown. References Figure 2 The environment-aware video generation system 20 is applied to the server corresponding to the wearable device. The environment-aware video generation system 20 includes: an information acquisition module 21, a first text generation module 22, a second text generation module 23, a text conversion module 24, and a video generation module 25.
[0053] Among them, the information acquisition module 21 is used to acquire first image information and target eye movement information. The first image information is the image information of the target area collected by the wearable device, and the target eye movement information is the eye movement information of the target person collected by the wearable device. The target person is the person wearing the wearable device. The first text generation module 22 is used to generate image description text corresponding to the first image information; The second text generation module 23 is used to extract multiple second image information from the first image information based on the target eye movement information, and generate image description text corresponding to the multiple second image information. The text conversion module 24 is used to convert the image description text of multiple second image information into target description text based on preset text conversion rules; the target description text is used to instruct the server to generate video. The video generation module 25 is used to input at least the image description text and target description text of the first image information into the large video generation model to obtain the video corresponding to the first image information.
[0054] In one embodiment of this application, the video generation module 25 is specifically used to determine the richness score of the image description text of the first image information; If the richness score of the image description text of the first image information is greater than the specified score, the image description text of the first image information and the target description text are input into the video generation model to obtain the video corresponding to the first image information.
[0055] In one embodiment of this application, the video generation module 25 is further configured to determine the scene to which the first image information belongs based on the image description text when the richness score of the image description text of the first image information is less than or equal to a specified score, and to determine the storyboard description text corresponding to the scene to which the first image information belongs as the storyboard description text of the first image information. The image description text, target description text, and storyboard description text of the first image information are input into the video generation model to obtain the video corresponding to the first image information.
[0056] In one embodiment of this application, the video generation module 25 is specifically used to input the image description text of the first image information into the first semantic model to obtain visual part-of-speech words and visual entity words in the image description text of the first image information. The visual part-of-speech words are words in the image description text of the first image information that have a greater correlation with the preset visual part-of-speech words than the preset correlation, and the visual entity words are words in the image description text of the first image information that have a greater correlation with the preset visual entity words than the preset correlation. The information density score is determined based on the proportion of the total number of visual part-of-speech words and visual entity words to the total number of words in the image description text of the first image information. Input the image description text of the first image information into the second semantic model to obtain the number of structural words in the image description text of the first image information. Structural words are words in the image description text of the first image information whose relevance to preset structural words is greater than preset relevance. The structural integrity score is determined based on the number of structural vocabulary categories. Extract detailed words from the image description text of the first image information. Detailed words are descriptive words in the image description text of the first image information. The detail granularity score is determined based on the proportion of detail words in the image description text of the first image information to the total number of words in the image description text of the first image information. The richness score is obtained by weighting the information density score, structural integrity score, and detail granularity score.
[0057] In one embodiment of this application, the environment-aware video generation system 20 further includes: a weight adjustment module, configured to, in response to the application scenario of the video corresponding to the first image information being news-related, increase the weight corresponding to the information density score based on a first information step size, decrease the weight corresponding to the structural integrity score based on a first structure step size, and decrease the weight corresponding to the detail granularity score based on a first detail step size; the first information step size is equal to the sum of the first structure step size and the first detail step size. In response to the fact that the application scenario of the video corresponding to the first image information is a narrative type, the weight corresponding to the structural integrity score is increased based on the second structural step size, the weight corresponding to the information density score is decreased based on the second information step size, and the weight corresponding to the detail granularity score is decreased based on the second detail step size; the second structural step size is equal to the sum of the second information step size and the first detail step size. In response to the application scenario of the video corresponding to the first image information being promotional, the weight corresponding to the detail granularity score is increased based on the third detail step length, the weight corresponding to the information density score is decreased based on the third information step length, and the weight corresponding to the structural integrity score is decreased based on the third structure step length. The third detail step length is equal to the sum of the third information step length and the third structure step length.
[0058] In one embodiment of this application, the second text generation module 23 is specifically used to determine the eye focus area of the target person within a first time period based on the target eye movement information; the first time period is the time period during which the wearable device collects the eye movement information of the target person, and the start and end times of the first image information collection are consistent with the start and end times of the target eye movement information collection; Multiple second image information is extracted from the first image information based on the eye focus area of the target person in the first time period.
[0059] In one embodiment of this application, the text conversion module 24 is specifically used to obtain visual part-of-speech words, visual entity words, structural words and descriptive words in the image description text of each second image information. Repeated words in the visual part-of-speech vocabulary, visual entity vocabulary, structural vocabulary, and descriptive vocabulary corresponding to each second image information are converted into target description text.
[0060] See Figure 3 , Figure 3 This is a schematic block diagram of a server provided in one embodiment of this application. Figure 3 The server 300 in this embodiment may include one or more processors 301, one or more input devices 302, one or more output devices 303, and one or more memories 304. The processors 301, input devices 302, output devices 303, and memories 304 communicate with each other via a communication bus 305. The memories 304 store computer programs, including program instructions. The processors 301 execute the program instructions stored in the memories 304. Specifically, the processors 301 are configured to invoke the program instructions to perform the functions of the modules in the various system embodiments described above, for example... Figure 2 The functions of the information acquisition module 21, the first text generation module 22, the second text generation module 23, the text conversion module 24, and the video generation module 25 are shown.
[0061] It should be understood that, in the embodiments of this application, the processor 301 may be a central processing unit (CPU), but it may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor.
[0062] Input device 302 may include a touchpad, a fingerprint sensor (for collecting the user's fingerprint information and fingerprint orientation information), a microphone, etc., and output device 303 may include a display (LCD, etc.), a speaker, etc.
[0063] The memory 304 may include read-only memory and random access memory, and provides instructions and data to the processor 301. A portion of the memory 304 may also include non-volatile random access memory. For example, the memory 304 may also store a first information step size.
[0064] In specific implementations, the processor 301, input device 302, and output device 303 described in the embodiments of this application can execute the implementation methods described in the embodiments of the environment-aware video generation method provided in the embodiments of this application, or they can execute the server implementation methods described in the embodiments of this application, which will not be repeated here.
[0065] In another embodiment of this application, a computer-readable storage medium is provided. This computer-readable storage medium stores a computer program, which includes program instructions. When executed by a processor, the program instructions implement all or part of the processes in the methods described above. Alternatively, the computer program can instruct related hardware to complete the process. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include any entity or device capable of carrying computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.
[0066] The computer-readable storage medium can be an internal storage unit of the server in any of the foregoing embodiments, such as the server's hard drive or memory. The computer-readable storage medium can also be an external storage device of the server, such as a plug-in hard drive, smart media card (SMC), secure digital (SD) card, flash card, etc., installed on the server. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the server. The computer-readable storage medium is used to store computer programs and other programs and data required by the server. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.
[0067] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this application.
[0068] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working process of the server and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0069] In the several embodiments provided in this application, it should be understood that the disclosed servers and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be an indirect coupling or communication connection through some interfaces or units, or it may be an electrical, mechanical, or other form of connection.
[0070] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the embodiments of this application, depending on actual needs.
[0071] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0072] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A video generation method based on environment awareness, characterized in that, Servers used in wearable devices include: Acquire first image information and target eye movement information, wherein the first image information is image information of the target area collected by the wearable device, and the target eye movement information is eye movement information of the target person collected by the wearable device, wherein the target person is the person wearing the wearable device; Generate image description text corresponding to the first image information; Based on the target eye-tracking information, extract multiple second image information from the first image information and generate image description text corresponding to the multiple second image information; Based on preset text conversion rules, the image description text of the multiple second image information is converted into target description text; Determine the richness score of the image description text of the first image information; If the richness score of the image description text of the first image information is greater than the specified score, the image description text of the first image information and the target description text are input into the video generation model to obtain the video corresponding to the first image information. The determination of the richness score of the image description text of the first image information includes: The image description text of the first image information is input into the first semantic model to obtain visual part-of-speech words and visual entity words in the image description text of the first image information. The visual part-of-speech words are words in the image description text of the first image information that have a greater correlation with the preset visual part-of-speech words than the preset correlation. The visual entity words are words in the image description text of the first image information that have a greater correlation with the preset visual entity words than the preset correlation. The information density score is determined based on the ratio of the total number of visual part-of-speech words and visual entity words to the total number of words in the image description text of the first image information. The image description text of the first image information is input into the second semantic model to obtain the number of structural words in the image description text of the first image information. The structural words are words in the image description text of the first image information that have a greater correlation with preset structural words than preset correlation. The structural integrity score is determined based on the number of structural vocabulary categories mentioned above; Extract detailed words from the image description text of the first image information, wherein the detailed words are descriptive words in the image description text of the first image information; The detail granularity score is determined based on the ratio of the number of detail words in the image description text of the first image information to the total number of words in the image description text of the first image information. The richness score is obtained by weighting the information density score, the structural integrity score, and the detail granularity score. The weight adjustment process for the information density score, the structural integrity score, and the detail granularity score includes: In response to the application scenario of the video corresponding to the first image information being news, the weight corresponding to the information density score is increased based on the first information step size, the weight corresponding to the structural integrity score is decreased based on the first structure step size, and the weight corresponding to the detail granularity score is decreased based on the first detail step size; the first information step size is equal to the sum of the first structure step size and the first detail step size. In response to the fact that the application scenario of the video corresponding to the first image information is a narrative type, the weight corresponding to the structural integrity score is increased based on the second structural step size, the weight corresponding to the information density score is decreased based on the second information step size, and the weight corresponding to the detail granularity score is decreased based on the second detail step size; the second structural step size is equal to the sum of the second information step size and the first detail step size. In response to the application scenario of the video corresponding to the first image information being promotional, the weight corresponding to the detail granularity score is increased based on the third detail step length, the weight corresponding to the information density score is decreased based on the third information step length, and the weight corresponding to the structural integrity score is decreased based on the third structure step length. The third detail step length is equal to the sum of the third information step length and the third structure step length.
2. The video generation method based on environment awareness as described in claim 1, characterized in that, Also includes: If the richness score of the image description text of the first image information is less than or equal to the specified score, the scene to which the first image information belongs is determined based on the image description text, and the storyboard description text corresponding to the scene to which it belongs is determined as the storyboard description text of the first image information. The image description text of the first image information, the target description text, and the storyboard description text of the first image information are input into the video generation model to obtain the video corresponding to the first image information.
3. The video generation method based on environment perception as described in claim 1, characterized in that, The step of extracting multiple second image information from the first image information based on the target eye movement information includes: Based on the target eye movement information, the eye focus area of the target person is determined in the first time period; the first time period is the time period during which the wearable device collects the eye movement information of the target person, and the start and end times of the first image information collection are consistent with the start and end times of the target eye movement information collection; Multiple second image information is extracted from the first image information based on the eye focus area of the target person during the first time period.
4. The video generation method based on environment awareness as described in claim 1, characterized in that, The method of converting the image description text of the plurality of second image information into target description text based on preset text conversion rules includes: For each image description text of the second image information, obtain the visual part-of-speech words, visual entity words, structural words, and descriptive words in the image description text of the second image information; Repeated words in the visual part-of-speech vocabulary, visual entity vocabulary, structural vocabulary, and descriptive vocabulary corresponding to each second image information are converted into target description text.
5. A video generation system based on environmental awareness, characterized in that, For implementing the method as described in any one of claims 1-4, the system is applied to a server corresponding to a wearable device, the system comprising: The information acquisition module is used to acquire first image information and target eye movement information. The first image information is image information of the target area collected by the wearable device, and the target eye movement information is eye movement information of the target person collected by the wearable device. The target person is the person wearing the wearable device. The first text generation module is used to generate image description text corresponding to the first image information; The second text generation module is used to extract multiple second image information from the first image information based on the target eye movement information, and generate image description text corresponding to the multiple second image information. The text conversion module is used to convert the image description text of the plurality of second image information into target description text based on preset text conversion rules; The video generation module is used to input at least the image description text of the first image information and the target description text into the large video generation model to obtain the video corresponding to the first image information.
6. A server comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method as described in any one of claims 1 to 4.
7. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Method for generating text description information of image
CN117037177A
Video generation method and device
CN118400478A
Video generation method and device, electronic equipment and storage medium
CN119788936A
Video end-to-end generation method and device, equipment and medium
CN120186427A
Generation of imagery from descriptive text
US10074200B1