Interactive video and model generation method and device, electronic equipment and storage medium
By generating interactive feature vectors and inputting them into the interactive video model, the problems of interactive object recognition and action matching in the digital human generation system are solved, realizing accurate interaction between digital humans and interactive objects and efficient video generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- ZHENGZHOU APUS DIGITAL CLOUD INFORMATION TECH CO LTD
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-05
AI Technical Summary
Existing digital human generation systems cannot automatically identify interactive objects in scene text descriptions, resulting in digital humans lacking clear interaction targets, leading to problems such as remote interaction and mismatched actions, and the operation is complex and inefficient.
By generating interactive feature vectors, including digital human interaction capabilities, scene attributes, interactive object features, and interactive action sub-vectors, and inputting them into a trained interactive video model, real-time interaction between the digital human and the interactive object is achieved.
It enables precise interaction between digital humans and interactive objects, improves video generation efficiency, simplifies the operation process, requires no manual intervention, and generates videos that meet the user's personalized needs.
Smart Images

Figure CN121985196A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to a method, apparatus, electronic device and storage medium for generating interactive videos and models. Background Technology
[0002] With the development of artificial intelligence (AI) technology, digital human generation technology can quickly generate high-precision digital human images and voice videos from a single photo or video. Video generation technology can generate virtual scene videos that conform to scene settings based on scene text descriptions.
[0003] In related technologies, digital humans are overlaid on the overall scene to generate digital human scene videos. However, the above solutions have the following drawbacks: 1) They cannot automatically identify core interactive objects (such as "Model A icon" or "product device") from the scene text description, resulting in digital humans having no clear interactive targets; 2) The lack of clear interactive targets (pointing, touching, displaying, etc.) for digital humans leads to a lack of spatial calibration and logical binding with interactive objects, often resulting in awkward effects such as "air interaction" and "mismatch between action and object size," for example, the digital human's finger does not accurately point to the icon position; 3) The scene text description, scene, and interaction are disconnected, making it impossible to automatically transform the scene text description into a closed-loop process of "scene generation + object creation + digital human interaction," requiring manual configuration of object parameters and action instructions, which is complex and inefficient. Summary of the Invention
[0004] The purpose of this application is to provide a method, apparatus, electronic device, and storage medium for generating interactive videos and models, so as to solve the problems in related technologies where digital humans have no clear interaction target, cannot accurately point to the interaction object, and are complex and inefficient to operate.
[0005] To achieve the above objectives, the embodiments of this application adopt the following technical solutions: In a first aspect, embodiments of this application provide a method for generating interactive videos, comprising: acquiring a digital human image and scene text; generating corresponding interactive feature vectors based on the digital human image and the scene text, the interactive feature vectors including: digital human interactive ability sub-vectors, scene attribute sub-vectors, interactive object feature sub-vectors in the scene, and interactive action sub-vectors; and inputting the scene text and the interactive feature vectors into a trained interactive video model to obtain a target interactive video.
[0006] Secondly, embodiments of this application provide a method for generating an interactive video model, comprising: acquiring a training dataset, wherein the training dataset includes multiple sets of training data, each set of training data including a training scene text, a training interaction feature vector, and a corresponding training target interactive video, wherein the training interaction feature vector includes: a training digital human interaction ability sub-vector, a training scene attribute sub-vector, a training interaction object feature sub-vector in the training scene, and a training interaction action sub-vector; inputting the training scene text and the training interaction feature vector into an interactive video model to be trained to obtain a training candidate interactive video; and training the interactive video to be trained based on the training candidate interactive video and the training target interactive video to obtain a trained interactive video model. Thirdly, embodiments of this application provide an interactive video generation apparatus, comprising: a first acquisition module for acquiring a digital human image and scene text; a generation module for generating a corresponding interactive feature vector based on the digital human image and the scene text, the interactive feature vector including: a digital human interactive ability sub-vector, a scene attribute sub-vector, an interactive object feature sub-vector in the scene, and an interactive action sub-vector; and a first input module for inputting the scene text and the interactive feature vector into a trained interactive video model to obtain a target interactive video.
[0007] Fourthly, embodiments of this application provide an apparatus for generating an interactive video model, comprising: a second acquisition module, configured to acquire a training dataset, the training dataset including multiple sets of training data, each set of training data including a training scene text, a training interaction feature vector, and a corresponding training target interactive video, the training interaction feature vector including: a training digital human interaction ability sub-vector, a training scene attribute sub-vector, a training interaction object feature sub-vector in the training scene, and a training interaction action sub-vector; a second input module, configured to input the training scene text and the training interaction feature vector into an interactive video model to be trained, to obtain a training candidate interactive video; and a training module, configured to train the interactive video to be trained based on the training candidate interactive video and the training target interactive video, to obtain a trained interactive video model.
[0008] Fifthly, embodiments of this application provide an electronic device, including: a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein when the program or instructions are executed by the processor, they implement the steps of the method described in the first aspect of this application, or implement the steps of the method described in the second aspect of this application.
[0009] In a sixth aspect, embodiments of this application provide a readable storage medium on which a program or instructions are stored. When executed by a processor, the program or instructions implement the steps of the method described in the first aspect of this application, or implement the steps of the method described in the second aspect of this application.
[0010] The above-described technical solutions adopted in the embodiments of this application can achieve the following beneficial effects: In this embodiment, when generating interactive videos, an interaction feature vector is generated based on the input digital human image and scene text. This vector includes: digital human interaction capability sub-vectors, scene attribute sub-vectors, interaction object feature sub-vectors in the scene, and interaction action sub-vectors. The scene text and interaction feature vectors are then input into a trained interactive video model to obtain a target interactive video of real-time interaction between the digital human and the interaction object. The generated interaction feature vector includes sub-vectors in four dimensions, enabling the identification of interaction objects from the scene text, giving the digital human a clear interaction target. It also enables spatial calibration and logical binding of the digital human's interaction actions with the interaction object, allowing the digital human to accurately point to the interaction object. Furthermore, it achieves a complete process of "input digital human image + scene text → automatic generation of a scene containing interaction objects + precise interaction between the digital human and the interaction object → output of a complete fused video," requiring no manual intervention and improving video generation efficiency. Attached Figure Description
[0011] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings: Figure 1 A flowchart illustrating an interactive video generation method provided in one embodiment of this application; Figure 2 A flowchart illustrating an interactive video generation method provided in another embodiment of this application; Figure 3 A flowchart illustrating a method for generating an interactive video model, provided as an embodiment of this application; Figure 4 A schematic diagram of the structure of an interactive video generation apparatus provided in one embodiment of this application; Figure 5 A schematic diagram of the structure of an interactive video model generation device provided for another embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided in one embodiment of this application. Detailed Implementation
[0012] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application will be clearly and completely described below in conjunction with specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0013] The terms "first," "second," etc., used in this application are used to distinguish similar objects and not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of this application can be implemented in orders other than those illustrated or described herein. Furthermore, "and / or" in this application indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship. It should be noted that all data involved in this application was obtained with the user's authorization.
[0014] The technical solutions provided by the various embodiments of this application are described in detail below with reference to the accompanying drawings.
[0015] Figure 1 This is a flowchart illustrating an interactive video generation method according to an embodiment of this application. Figure 1 As shown, the interactive video generation method of this application embodiment may specifically include the following steps: S101, obtain digital human image and scene text.
[0016] In this embodiment of the application, the executing entity of the interactive video generation method is an interactive video generation device, which can be located in an electronic device. This electronic device can be a terminal device or a server. The terminal device can be a mobile phone, tablet computer, desktop computer, laptop, in-vehicle device, etc.; the server can be a standalone server or a server cluster composed of multiple servers.
[0017] This application describes the application process of an interactive video model. By deploying the trained (i.e., fine-tuned) interactive video model as an online service, a real-time closed loop of "input-generation-output" is achieved, meeting users' needs for quickly generating interactive videos.
[0018] The trained interactive video model can be loaded into the server environment and provided to the outside world through the Application Programming Interface (API). To ensure high throughput and low latency, techniques such as model quantization (INT8 quantization) and optimization of the inference engine (such as TensorRT-LLM, vLLM) can be used to ensure that the latency of a single generation is ≤1.5 seconds / 10 seconds of video, and support 100+ concurrent requests per second.
[0019] In client applications (such as web-based tools and mobile apps), users input a digital avatar (which can be a photo or video) and scene text T (containing an open-ended description of the interactive object, such as "Company A released Model A" or "Demonstrating the noise-canceling function of the new wireless headphones"). The digital avatar can be generated based on ZEGO Digital Human API, D-ID Image to Avatar, iFlytek Smart Creation, etc.
[0020] S102, Based on the digital human image and scene text, generate corresponding interaction feature vectors. The interaction feature vectors include: digital human interaction ability sub-vectors, scene attribute sub-vectors, interaction object feature sub-vectors in the scene, and interaction action sub-vectors.
[0021] In this embodiment, based on the digital human image and scene text T obtained in step S101, an interaction feature vector V is generated, containing complete features of the digital human, scene, interactive object, and interactive action. The interaction feature vector V includes sub-vectors in the following four dimensions: digital human interaction capability sub-vector human_interact_cap, scene attribute sub-vector scene_attr, interactive object feature sub-vector object_feat in the scene, and interactive action sub-vector interact_action.
[0022] The dimensions of the interaction feature vector V are defined as follows: Digital human interaction capability dimensions include: support for fine hand movements (0-1.0, 1.0 being full support), range of motion adaptation (small, medium, large), and facial expression coordination (0-1.0).
[0023] Scene attribute dimensions: including scene type (technology, life, business), light source parameters (direction 0°-360°, intensity 0.0-1.0), and spatial layout (open, compact).
[0024] Interactive object feature dimensions include object type (icon, physical product, virtual prop), size parameters (length × width × height, unit cm), shape (rectangle, circle, irregular), and interactive hotspot coordinates (X, Y, Z).
[0025] Interactive action dimensions: including action type (pointing, touching, circling, holding), action speed (slow, medium, fast), and action duration (0.1-5.0 seconds).
[0026] Specific quantization rules are set for each dimension to ensure that the feature vectors are computable and comparable. Specifically: Digital human interaction capability dimension: fine hand movements are quantified as [0.0, 1.0] (1.0 supports fingertip touch, 0.0 only supports arm pointing); the range of motion is quantified by enumeration values (small = 0.3, medium = 0.6, large = 1.0).
[0027] Scene attribute dimensions: Light source intensity is quantized as [0.0, 1.0] (0.0 completely dark, 1.0 bright light), and the direction of the light source is calibrated by angle with the center point of the scene as the origin.
[0028] Interactive object feature dimensions: Size parameters are directly calibrated with actual values (e.g., the size of the A model icon is 5cm×5cm), and the coordinates of interactive hotspots are output with the scene ground as the Z=0 plane, and the three-dimensional coordinates are output (e.g., X=0.5, Y=1.2, Z=0.05).
[0029] Interactive action dimension: Action speed is quantified using enumerated values (slow = 0.4, medium = 0.7, fast = 1.0), and action duration is specified in seconds.
[0030] Furthermore, step S102, "generating corresponding interaction feature vectors based on the digital human image and scene text," may specifically include the following steps: generating digital human interaction capability sub-vectors based on the digital human image; generating scene attribute sub-vectors based on the scene text; using semantic parsing technologies such as GPT-5, Claude 3, and Tongyi Qianwen to extract information about interactive objects (such as "A model icon") from the scene text, and generating interactive object feature sub-vectors based on the interactive object information; generating interactive action sub-vectors based on preset recommended interactive actions; and generating the interaction feature vector V based on the digital human interaction capability sub-vectors, scene attribute sub-vectors, interactive object feature sub-vectors, and interactive action sub-vectors.
[0031] It should be noted here that if users have personalized needs, they can supplement the input through probe questions to update the interaction feature vector V.
[0032] Specifically, for undefined feature dimensions, closed-ended multiple-choice questions ("probe questions") are designed to transform users' qualitative, personalized needs into quantitative vectors. For example: Probe question: What kind of interaction style do you want the digital human to have with interactive objects? A) Fine and gentle (mapped to the interaction action subvector interaction_action speed = 0.4, action type = touch).
[0033] B) Concise and clear (mapped to the interaction action subvector, speed = 0.7, action type = target).
[0034] C) Vivid and exaggerated (mapped to the interactive action subvector interaction_action velocity = 1.0, action type = surround).
[0035] Using the above rules, the qualitative descriptions of digital humans, scenes, interactive objects, and interaction requirements can be converted into a quantified N-dimensional interaction feature vector V. For example, the interaction feature vector V can be represented as follows: Interaction feature vector V={human_interact_cap:{Fine action support: 1.0, Action amplitude: small}, scene_attr:{Type: Technology, Light source direction: 90°, Intensity: 0.8}, object_feat:{Type: Icon, Size: 5×5cm, Interaction hotspot (0.5, 1.2, 0.05)}, interaction_action:{Type: Touch, Speed: Slow, Duration: 0.5 seconds}}.
[0036] S103, input the scene text and interaction feature vector into the trained interactive video model to obtain the target interactive video.
[0037] In this embodiment, the scene text obtained in step S101 and the interaction feature vector generated in step S102 can be sent to the model server via API. The interactive video model trained on the server performs inference and generates a fused video (MP4 format, 720P-4K resolution) showing the precise interaction between the digital human and the interactive object. This video is then returned to the client via API and displayed to the user. The interactive video model generates a scene video containing the interactive object based on the input scene text and interaction feature vector. Through spatial positioning, action calibration, and logical binding, the digital human performs dynamic interactions matching the interactive object (such as pointing to the A model icon, tapping the icon to display details, etc.). Finally, a completely fused interactive video is output, denoted as the target interactive video, or target V_out.
[0038] Interactive video models can be based on large models that support video generation and conditional input, such as the open-source Pika Labs base model and the Runway Gen-2 series models, which require spatial understanding and motion generation capabilities.
[0039] It should be noted here that if the basic model does not support direct video generation, a "staged generation + fusion" strategy can be adopted. Specifically, firstly, based on the interaction feature vector V, digital human interactive video (with alpha channel) and scene video containing interactive objects are generated respectively. Then, the trained fusion model is used to achieve spatial calibration, light and shadow matching and motion synchronization, and finally output the complete target interactive video.
[0040] The trained interactive video model can be obtained by training the training dataset. For details, please refer to the relevant descriptions in the following embodiments, which will not be repeated here.
[0041] Furthermore, before inputting the interaction feature vector V into the interactive video model, the structured interaction feature vector V needs to be converted into a format readable by the interactive model. Correspondingly, such as... Figure 2 As shown, step S103, "inputting the scene text and interaction feature vector into the trained interactive video model to obtain the target interactive video," may specifically include the following steps: S201 converts the interaction feature vector into a predefined format string and wraps the predefined format string with a target tag to obtain the serialized interaction feature vector string.
[0042] Specifically, the quantized interaction feature vector V is converted into a predefined format string and wrapped with a special marker (i.e., target marker). For example, the interaction feature vector V={human_interact_cap:{fine action support: 1.0, action amplitude: small}, scene_attr:{type: technology, light source direction: 90°, intensity: 0.8}, object_feat:{type: icon, size: 5×5cm, interaction hotspot (0.5, 1.2, 0.05)}, interaction_action:{type: touch, speed: slow, duration: 0.5 seconds}} can be converted into the following interaction feature vector string. String: <<|interact_vec|>human_interact_cap:{"Fine-grained motion support":1.0,"Motion amplitude":0.3},scene_attr:{"Type":"Technology","Light source direction":90°,"Intensity":0.8},object_feat:{"Type":"Icon","Size":"5×5cm","Interaction hotspot":"(0.5,1.2,0.05)"},interact_action:{"Type":"Touch","Speed":0.4,"Duration":0.5}<<| / interact_vec|>.
[0043] S202, the serialized interactive feature vector string, scene text and video generation prompt are concatenated to obtain a single text string.
[0044] Specifically, the concatenated single text string can be as follows: <<|interact_vec|>...<<| / interact_vec|>\n Scene text: Company A released Model A\n Interactive video generation requirements: The digital human interacts precisely with the Model A icon, with consistent lighting and shadows, and natural movements\n Video output:
[0045] S203, input a single text string into the trained interactive video model to obtain the target interactive video.
[0046] Furthermore, if the user is not satisfied with the generated target interactive video, the video generation process can be iteratively optimized. Correspondingly, after step S103 above, "inputting the scene text and interactive feature vector into the trained interactive video model to obtain the target interactive video," the interactive video generation method of this application embodiment may further include the following steps: based on the user's modification of at least one of the following information based on the target interactive video: digital human image, scene text T, and recommended interactive actions, regenerate the corresponding interactive feature vector V and the target interactive video until the user is satisfied.
[0047] This application also includes a feedback collection mechanism to obtain explicit user feedback (satisfaction rating of 1-5 stars) and implicit feedback (video dwell time, download rate, modification frequency). After filtering and labeling high-quality interaction logs (including interaction feature vector V, scene text T, and target interaction video V_out), the training dataset used for training the interaction model is expanded, and the interaction model is retrained regularly to achieve continuous iterative improvement of interaction capabilities.
[0048] In summary, the interactive video generation method of this application generates an interactive feature vector based on the input digital human image and scene text. This vector includes: a digital human interaction capability sub-vector, a scene attribute sub-vector, an interactive object feature sub-vector in the scene, and an interactive action sub-vector. The scene text and interactive feature vector are then input into a trained interactive video model to obtain a target interactive video of real-time interaction between the digital human and the interactive object. The generated interactive feature vector includes four-dimensional sub-vectors, enabling the identification of interactive objects from the scene text, giving the digital human a clear interaction target. It also enables spatial calibration and logical binding of the digital human's interactive actions with the interactive object, allowing the digital human to accurately point to the interactive object. Furthermore, it achieves a complete process of "input digital human image + scene text → automatic generation of a scene containing interactive objects + precise interaction between the digital human and the interactive object → output of a complete fused video," requiring no manual intervention and improving video generation efficiency. Updating the interactive feature vector with user-personalized requirements through probe questions makes the generated video more tailored to individual user needs.
[0049] Figure 3 This is a flowchart illustrating a method for generating an interactive video model, provided as an embodiment of this application. Figure 3 As shown, the method for generating an interactive video model according to an embodiment of this application may specifically include the following steps: S301, Obtain the training dataset. The training dataset includes multiple sets of training data. Each set of training data includes training scene text, training interaction feature vectors, and corresponding training target interaction videos. The training interaction feature vectors include: training digital human interaction ability sub-vectors, training scene attribute sub-vectors, training interaction object feature sub-vectors in the training scene, and training interaction action sub-vectors.
[0050] In this embodiment of the application, the execution entity of the interactive video model generation method is an interactive video model generation device, which can be located in an electronic device. This electronic device can be a terminal device or a server. The terminal device can be a mobile phone, tablet computer, desktop computer, laptop, in-vehicle device, etc.; the server can be a standalone server or a server cluster composed of multiple servers.
[0051] The training dataset in this application embodiment is the dataset used to train (or fine-tune) the interactive video model during the model training (i.e., conditional fine-tuning) phase. The training dataset includes multiple sets of training data, each set of training data including: training scene text, training interactive feature vectors, and the corresponding training target interactive video.
[0052] The training target interactive video V_out is a video that is manually produced by experts and / or synthesized with AI assistance. It conforms to the training interactive feature vector V and the training scene text T. It requires that the digital human and the interactive object are spatially adapted, the movements are precise, and the lighting and shadows are consistent.
[0053] Expert-driven production: Organizing digital human animators, scene designers, and interaction engineers to collaborate and create interactive videos that meet the requirements based on the trained interaction feature vector V and the trained scene text T.
[0054] AI-assisted synthesis: Utilizing a more powerful general-purpose large model (teacher model) for automated generation, and designing a "meta-prompt" that includes role-playing, feature injection, and task instructions.
[0055] An example of a meta-prompt could be as follows: "You are a professional interactive video production expert. You need to generate a video based on the following parameters: Interaction feature vector V={human_interact_cap: {Fine motion support: 1.0, Motion amplitude: small}, scene_attr: {Type: Technology, Light source direction: 90°, Intensity: 0.8}, object_feat: {Type: Icon, Size: 5×5cm, Interaction hotspot (0.5, 1.2, 0.05)}, interaction_action: {Type: Touch, Speed: Slow, Duration: 0.5 seconds}}; Scene text T='Company A has released Model A'. Video requirements: In a technological scene, the Model A icon floats in the center of the booth. A digital human stands in the scene, slowly raises his right hand, and lightly touches the icon's interaction hotspot with his index finger. The movement is natural, his face is smiling, the lighting is consistent with the scene's light source, and the video duration is 3 seconds."
[0056] All generated training data are compiled into a structured file (such as JSONL format), with each line representing a complete set of training data, which serves as input for fine-tuning the interactive video model.
[0057] In addition, to improve the generalization ability of the interactive video model, when generating the training dataset, multiple different training interactive feature vectors V can be generated for the same training scene text T and multiple different training interactive videos V_out (i.e., covering different regions of the feature space).
[0058] S302, input the training scene text and training interaction feature vector into the interactive video model to be trained to obtain the training candidate interactive video.
[0059] In this embodiment, step S302 is similar to step S103 in the above embodiment, and will not be described again here.
[0060] The interactive video model to be trained can be a pre-trained large model that supports video generation and conditional input as the base model, such as the open-source Pika Labs base model and the Runway Gen-2 series models, which are required to have spatial understanding and action generation capabilities.
[0061] S303, Based on the training candidate interactive video and the training target interactive video, train the interactive video to be trained to obtain the trained interactive video model.
[0062] In this embodiment, standard supervised fine-tuning techniques are used to train the base model. The model receives a concatenated single text string as input, and its learning objective is to generate a video consistent with the "training target interactive video V_out" in the training data. During training, the difference between the generated training candidate interactive videos and the training target interactive video is measured using a video similarity loss function (PSNR+SSIM+action accuracy score). The model's weights (i.e., the model's parameters) are adjusted using backpropagation and gradient descent algorithms. After multiple epochs of training, the model will learn to understand the feature information within the <<|interact_vec|> tags and generate a qualified interactive video based on this information.
[0063] After training, the model's performance can be evaluated using an independent validation dataset (i.e., a test dataset). Key performance metrics for evaluating the model include: interactive action accuracy (≥90%), lighting and shadow matching accuracy (≥85%), and scene-object-digital human coordination accuracy (≥88%). The fine-tuned model weights that meet the performance requirements are then saved, completing the training of the interactive video model.
[0064] In summary, the interactive video model generation method of this application, when fine-tuning the interactive video model, trains the interactive video model based on multiple sets of training data in the training dataset. Each set of training data includes a training interactive feature vector comprising four sub-vectors, enabling the identification of interactive objects from scene text, giving the digital human a clear interactive target, and achieving spatial calibration and logical binding between the digital human's interactive actions and the interactive object. This allows the digital human to accurately point to the interactive object, thereby improving the performance of the trained interactive video model. Furthermore, it achieves a complete process of "inputting a digital human image + scene text → automatically generating a scene containing interactive objects + precise interaction between the digital human and the interactive object → outputting a complete fused video," requiring no manual intervention, thus improving video generation efficiency and consequently improving the training efficiency of the interactive video model. When generating the training dataset, multiple different training interactive feature vectors V are matched for the same training scene text T, generating multiple different training target interactive videos V_out, improving the generalization ability of the interactive video model.
[0065] Figure 4 This is a schematic diagram of the structure of an interactive video generation apparatus provided in one embodiment of this application. Figure 4 As shown, the interactive video generation device 400 of this application embodiment may specifically include: a first acquisition module 401, a generation module 402, and a first input module 403. Wherein: The first acquisition module 401 is used to acquire the digital human image and scene text.
[0066] The generation module 402 is used to generate corresponding interaction feature vectors based on the digital human image and scene text. The interaction feature vectors include: digital human interaction ability sub-vectors, scene attribute sub-vectors, interaction object feature sub-vectors in the scene, and interaction action sub-vectors.
[0067] The first input module 403 is used to input the scene text and interaction feature vector into the trained interactive video model to obtain the target interactive video.
[0068] In the embodiments of this application, the specific process by which each module and unit in the interactive video generation device of this application implements its function can be found in the relevant descriptions in the embodiments of the interactive video generation method described above, and will not be repeated here.
[0069] In summary, the interactive video generation device of this application, when generating interactive videos, generates an interactive feature vector based on the input digital human image and scene text, including: digital human interactive ability sub-vectors, scene attribute sub-vectors, interactive object feature sub-vectors in the scene, and interactive action sub-vectors. The scene text and interactive feature vectors are then input into a trained interactive video model to obtain a target interactive video of real-time interaction between the digital human and interactive objects. The generated interactive feature vector includes four-dimensional sub-vectors, enabling the identification of interactive objects from the scene text, giving the digital human a clear interactive target, and achieving spatial calibration and logical binding between the digital human's interactive actions and interactive objects, allowing the digital human to accurately point to the interactive object. It also achieves a complete process of "input digital human image + scene text → automatic generation of a scene containing interactive objects + precise interaction between the digital human and interactive objects → output of a complete fused video," without manual intervention, thus improving video generation efficiency. Updating the interactive feature vector through a probe question based on user personalized needs makes the generated video more aligned with those needs.
[0070] Figure 5 This is a schematic diagram of the structure of an interactive video model generation device provided in one embodiment of this application. Figure 5 As shown, the interactive video model generation device 500 of this application embodiment may specifically include: a second acquisition module 501, a second input module 502, and a training module 503. Wherein: The second acquisition module 501 is used to acquire the training dataset. The training dataset includes multiple sets of training data. Each set of training data includes training scene text, training interaction feature vectors and corresponding training target interaction videos. The training interaction feature vectors include: training digital human interaction ability sub-vectors, training scene attribute sub-vectors, training interaction object feature sub-vectors in the training scene and training interaction action sub-vectors.
[0071] The second input module 502 is used to input the training scene text and training interaction feature vector into the interactive video model to be trained, so as to obtain the training candidate interactive video.
[0072] The training module 503 is used to train the interactive video to be trained based on the training candidate interactive video and the training target interactive video, so as to obtain the trained interactive video model.
[0073] In the embodiments of this application, the specific process by which each module and unit in the interactive video model generation device of this application implements its function can be found in the relevant descriptions in the embodiments of the interactive video model generation method described above, and will not be repeated here.
[0074] In summary, the interactive video model generation device of this application, when fine-tuning the interactive video model, trains the interactive video model based on multiple sets of training data in the training dataset. Each set of training data includes a training interactive feature vector comprising four sub-vectors, enabling the identification of interactive objects from scene text, giving the digital human a clear interactive target, and achieving spatial calibration and logical binding between the digital human's interactive actions and the interactive object. This allows the digital human to accurately point to the interactive object, thereby improving the performance of the trained interactive video model. Furthermore, it achieves a complete process of "inputting a digital human image + scene text → automatically generating a scene containing interactive objects + precise interaction between the digital human and the interactive object → outputting a complete fused video," requiring no manual intervention, thus improving video generation efficiency and consequently improving the training efficiency of the interactive video model. When generating the training dataset, multiple different training interactive feature vectors V are matched for the same training scene text T, generating multiple different training target interactive videos V_out, improving the generalization ability of the interactive video model.
[0075] This application also provides an electronic device. For example... Figure 6 As shown, the electronic device 600 can vary considerably due to differences in configuration or performance. It may include one or more processors 601 and memory 602, with memory 602 storing one or more programs or instructions. Memory 602 may be temporary or persistent storage. The application program stored in memory 602 may include one or more modules (not shown), each module including a series of computer-executable instructions for the electronic device 600. Furthermore, processor 601 may be configured to communicate with memory 602 and execute the series of computer-executable instructions stored in memory 602 on the electronic device 600. The electronic device 600 may also include one or more power supplies 603, one or more wired or wireless network interfaces 604, one or more input / output interfaces 605, and one or more keyboards 606.
[0076] Specifically, in the embodiments of this application, the electronic device includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor. When the program or instructions are executed by the processor, they implement the steps of any of the above-described interactive video generation method embodiments, or implement the steps of any of the above-described interactive video model generation method embodiments.
[0077] The electronic device in this application, when generating interactive videos, generates an interaction feature vector based on the input digital human image and scene text. This vector includes: a digital human interaction capability sub-vector, a scene attribute sub-vector, an interaction object feature sub-vector in the scene, and an interaction action sub-vector. The scene text and interaction feature vector are then input into a trained interactive video model to obtain a target interactive video of real-time interaction between the digital human and the interaction object. The generated interaction feature vector includes four-dimensional sub-vectors, enabling the identification of interaction objects from the scene text, giving the digital human a clear interaction target. It also enables spatial calibration and logical binding of the digital human's interaction actions with the interaction object, allowing the digital human to accurately point to the interaction object. Furthermore, it achieves a complete process of "input digital human image + scene text → automatic generation of a scene containing interaction objects + precise interaction between the digital human and the interaction object → output of a complete fused video," without manual intervention, thus improving video generation efficiency. Updating the interaction feature vector with user-personalized requirements through probe questions makes the generated video more consistent with user-personalized needs. Fine-tuning the interactive video model improves its training efficiency. When generating the training dataset, multiple different training interaction feature vectors V are matched for the same training scenario text T, generating multiple different training target interaction videos V_out, which improves the generalization ability of the interaction video model.
[0078] This application also proposes a readable storage medium storing one or more computer programs or instructions that, when executed by a processor in an electronic device, enable the processor in the electronic device to perform the steps of any of the above-described interactive video generation method embodiments, or to implement the steps of any of the above-described interactive video model generation method embodiments.
[0079] The readable storage medium of this application, when generating interactive videos, generates an interaction feature vector based on the input digital human image and scene text. This vector includes: a digital human interaction capability sub-vector, a scene attribute sub-vector, an interaction object feature sub-vector in the scene, and an interaction action sub-vector. The scene text and interaction feature vector are then input into a trained interactive video model to obtain a target interactive video of real-time interaction between the digital human and the interaction object. The generated interaction feature vector includes four-dimensional sub-vectors, enabling the identification of interaction objects from the scene text, giving the digital human a clear interaction target. It also enables spatial calibration and logical binding of the digital human's interaction actions with the interaction object, allowing the digital human to accurately point to the interaction object. Furthermore, it achieves a complete process of "input digital human image + scene text → automatic generation of a scene containing interaction objects + precise interaction between the digital human and the interaction object → output of a complete fused video," without manual intervention, thus improving video generation efficiency. Updating the interaction feature vector with user-personalized requirements through probe questions makes the generated video more aligned with user-personalized needs. This also improves the training efficiency of the interactive video model when fine-tuning it. When generating the training dataset, multiple different training interaction feature vectors V are matched for the same training scenario text T, generating multiple different training target interaction videos V_out, which improves the generalization ability of the interaction video model.
[0080] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions. A typical implementation device is a computer. Specifically, a computer can be, for example, a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email device, game console, tablet computer, wearable device, or any combination of these devices.
[0081] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.
[0082] Those skilled in the art will understand that embodiments of this application can be provided as methods, systems, or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0083] This application is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart... Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0084] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0085] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0086] In a typical configuration, a computing device includes one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0087] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0088] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0089] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0090] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.
[0091] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to interchangeably. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0092] The above are merely embodiments of this application and are not intended to limit the scope of this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the scope of the claims of this application.
Claims
1. An interactive video generation method, characterized in that, include: Acquire digital human avatars and scene descriptions; Based on the digital human image and the scene text, a corresponding interaction feature vector is generated. The interaction feature vector includes: digital human interaction ability sub-vector, scene attribute sub-vector, interaction object feature sub-vector in the scene, and interaction action sub-vector. The scene text and the interaction feature vector are input into the trained interactive video model to obtain the target interactive video.
2. The method according to claim 1, characterized in that, The step of generating corresponding interaction feature vectors based on the digital human image and the scene text includes: Generate the digital human interaction capability sub-vector based on the digital human image; Generate the scene attribute sub-vector based on the scene text; Extract information about interactive objects from the scene text, and generate feature sub-vectors of the interactive objects based on the information of the interactive objects; The interaction action sub-vector is generated based on the preset recommended interaction actions; The interaction feature vector is generated based on the digital human interaction capability sub-vector, the scene attribute sub-vector, the interaction object feature sub-vector, and the interaction action sub-vector.
3. The method according to claim 1, characterized in that, The step of inputting the scene text and the interaction feature vector into the trained interactive video model to obtain the target interactive video includes: The interaction feature vector is converted into a string in a predefined format, and the string in the predefined format is wrapped with a target tag to obtain the serialized interaction feature vector string; The serialized interactive feature vector string, the scene text, and the video generation prompt are concatenated to obtain a single text string; The single text string is input into the trained interactive video model to obtain the target interactive video.
4. The method according to claim 2, characterized in that, After inputting the scene text and the interaction feature vector into the trained interactive video model to obtain the target interactive video, the process further includes: Based on the user's modifications to at least one of the following information based on the target interactive video: the digital human image, the scene text, and the recommended interactive actions, the corresponding interactive feature vector and the target interactive video are regenerated.
5. A method for generating an interactive video model, characterized in that, include: Obtain a training dataset, which includes multiple sets of training data. Each set of training data includes training scene text, training interaction feature vectors, and corresponding training target interaction videos. The training interaction feature vectors include: training digital human interaction ability sub-vectors, training scene attribute sub-vectors, training interaction object feature sub-vectors in the training scene, and training interaction action sub-vectors. The training scene text and the training interaction feature vector are input into the interactive video model to be trained to obtain training candidate interactive videos; The interactive video to be trained is trained based on the candidate interactive video and the target interactive video to obtain the trained interactive video model.
6. The method according to claim 5, characterized in that, The acquisition of the training dataset includes: For the same training scenario text and multiple different training interaction feature vectors, multiple different training target interaction videos are generated.
7. An interactive video generation apparatus, characterized in that, include: The first acquisition module is used to acquire the digital human image and scene text; The generation module is used to generate corresponding interaction feature vectors based on the digital human image and the scene text. The interaction feature vectors include: digital human interaction ability sub-vectors, scene attribute sub-vectors, interaction object feature sub-vectors in the scene, and interaction action sub-vectors. The first input module is used to input the scene text and the interaction feature vector into the trained interactive video model to obtain the target interactive video.
8. An apparatus for generating an interactive video model, characterized in that, include: The second acquisition module is used to acquire a training dataset, which includes multiple sets of training data. Each set of training data includes training scene text, training interaction feature vectors, and corresponding training target interaction videos. The training interaction feature vectors include: training digital human interaction ability sub-vectors, training scene attribute sub-vectors, training interaction object feature sub-vectors in the training scene, and training interaction action sub-vectors. The second input module is used to input the training scene text and the training interaction feature vector into the interactive video model to be trained, so as to obtain the training candidate interactive video. The training module is used to train the interactive video to be trained based on the candidate interactive video and the target interactive video to obtain the trained interactive video model.
9. An electronic device, characterized in that, It includes a processor, a memory, and a program or instructions stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the method as claimed in any one of claims 1-4, or implement the steps of the method as claimed in any one of claims 5-6.
10. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the steps of the method as described in any one of claims 1-4, or the steps of the method as described in any one of claims 5-6.