Method, system, device and computer storage medium for generating interactive digital actor video

By constructing video scenes and digital actor databases and generating digital actor videos, the problem of poor interaction between digital people in multimodal interaction is solved, and higher degrees of freedom and flexible user interaction is achieved.

CN120125718BActive Publication Date: 2025-08-29GUANGDONG-HONG KONG-MACAO GREATER BAY AREA DIGITAL ECONOMY RESEARCH INSTITUTE (INTERNATIONAL ADVANCED TECHNOLOGY APPLICATION PROMOTION CENTER (SHENZHEN) +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510603049.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-12
Publication Date
2025-08-29
Estimated Expiration
2045-05-12

Smart Images

  • Figure CN120125718B_ABST
    Figure CN120125718B_ABST
Patent Text Reader

Abstract

This application discloses a method, system, device, and computer storage medium for generating interactive digital actor videos. The method includes constructing a video scene database and a digital actor database based on interactive videos, wherein the video scene database includes video description information of the interactive videos; the digital actor database includes feature data of character objects in the interactive videos; and generating digital actor videos based on interactive speech for the interactive videos, the video scene database, and the digital actor database. This application processes interactive videos into a video scene database and a digital actor database, and uses the video scene database and the digital actor database to provide information support for digital actor generation. This method can quickly generate digital actor voices based on interactive speech, thereby improving the freedom and interactive effect of video interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of human-computer interaction technology, and in particular to a method, system, device and computer storage medium for generating interactive digital actor videos. Background Art

[0002] Digital humans can simulate real-life expressions and movements, enabling user interaction in a variety of fields, including entertainment, education, and customer service. However, when interacting with users through videos, digital humans generally use pre-designed character tones and speaking habits to communicate with users around the video content. This makes it difficult to flexibly apply digital humans to multimodal real-time interactive portrait-driven applications, especially for video images with rich emotions and personalities, which affects the effectiveness of video interactions.

[0003] Therefore existing technology still needs to be improved and improved. Summary of the Invention

[0004] The technical problem to be solved by this application is to provide a method, system, device and computer storage medium for generating interactive digital actor videos in response to the shortcomings of the existing technology.

[0005] In order to solve the above technical problems, the first aspect of the present application provides a method for generating an interactive digital actor video, wherein the method for generating an interactive digital actor video specifically includes:

[0006] Constructing a video scene database and a digital actor database based on the interactive video, wherein the video scene database includes video description information of the interactive video; the digital actor database includes feature data of character objects in the interactive video;

[0007] A digital actor video is generated based on the interactive voice for the interactive video, the video scene database, and the digital actor database.

[0008] The method for generating interactive digital actor videos, wherein the step of constructing a video scene database and a digital actor database based on the interactive videos specifically includes:

[0009] Segmenting the interactive video into a plurality of scene video segments, and adding a segment identifier to each scene video segment;

[0010] Obtaining video description information of each scene segment and feature data of a character object in each scene video segment;

[0011] Constructing a video scene database according to the video description information of each scene video clip;

[0012] A digital actor database is constructed according to the feature data of the role objects in each scene video clip.

[0013] The method for generating an interactive digital actor video, wherein the feature data of the character object in each scene video clip is obtained, specifically includes:

[0014] The audio data corresponding to each scene video clip is divided into several audio clips based on the character object, and a serial number label is added to each audio clip;

[0015] Extract the speech style of each audio clip to obtain the voice representation data of the character object corresponding to each audio clip;

[0016] Perform face detection on each scene video clip to obtain the facial cropped video corresponding to each character object;

[0017] Extracting a three-dimensional facial representation from the facial cropped video to obtain a three-dimensional facial representation of each character object;

[0018] The voice representation data, the facial cropping video and the three-dimensional representation of the human face of each character object are used as feature data of each character object.

[0019] The method for generating an interactive digital actor video, wherein the step of constructing a video scene database based on the video description information of each scene video clip specifically includes:

[0020] determining a video summary of each scene video clip based on the video description information of each scene video clip;

[0021] A video scene database is constructed according to the video description information and video summary of each scene video clip.

[0022] The method for generating an interactive digital actor video, wherein determining the video summary of each scene video segment based on the video description information of each scene video segment specifically includes:

[0023] Perform speech recognition on each audio clip corresponding to each scene video clip to obtain the dialogue text of each audio clip;

[0024] The video description information of each scene video clip is updated based on all the dialogue texts corresponding to each scene video clip, and the updated video description information is summarized by a large language model to obtain a video summary of each scene video clip.

[0025] The method for generating an interactive digital actor video, wherein the step of constructing a video scene database based on the video description information and video summary of each scene video clip specifically includes:

[0026] Using the segment identifier of each scene video segment as the key and the video description information and summary information of each scene video segment as the value, construct the video scene data corresponding to each scene video segment;

[0027] A video scene database is formed based on the video scene data corresponding to all scene video clips.

[0028] The method for generating interactive digital actor videos, wherein the step of constructing a digital actor database based on the feature data of the character objects in each scene video clip, specifically comprises:

[0029] Use each character object as the key and the characteristic data of each character object as the value to construct digital actor data;

[0030] A digital actor database is formed based on the digital actor data corresponding to all role objects.

[0031] The method for generating an interactive digital actor video, wherein the step of generating the digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database, specifically comprises:

[0032] generating a digital actor voice based on the interactive voice, the video scene database, and the digital actor database;

[0033] generating a digital actor video stream according to the digital actor voice and the digital actor database;

[0034] A digital actor video is generated based on the digital actor voice and the digital actor video stream.

[0035] The method for generating an interactive digital actor video, wherein the step of generating the digital actor voice based on the interactive voice, the video scene database, and the digital actor database, specifically comprises:

[0036] Encoding the interactive speech to obtain a speech representation of the interactive speech;

[0037] Selecting video description information corresponding to the interactive voice in the video scene database;

[0038] Based on the video description information and the speech representation, obtaining target speech data for driving a digital actor constructed based on a target character object through a large language model;

[0039] Selecting target feature data of a target role object from the digital actor database;

[0040] The digital actor voice is generated according to the target feature data and the target voice data.

[0041] The method for generating an interactive digital actor video, wherein the step of generating a digital actor video stream based on the digital actor's voice and the digital actor database, specifically comprises:

[0042] Selecting target feature data of a target role object from the digital actor database;

[0043] Generate facial movement features for driving a digital actor constructed based on a target character object based on the target feature data and the digital actor voice;

[0044] A digital actor video stream is synthesized based on the facial motion features and the target feature data.

[0045] The method for generating an interactive digital actor video, wherein the step of synthesizing a digital actor video stream based on the facial motion features and the target feature data specifically includes:

[0046] Receive control information from user interaction;

[0047] A digital actor video stream is synthesized based on the control information, the facial motion features and the target feature data.

[0048] The method for generating an interactive digital actor video, wherein, after generating the digital actor video based on the digital actor's voice and the digital actor video stream, the method further comprises:

[0049] A playback area, an image selection area, and an image interaction area are configured for the digital actor video. The playback area is used to play the digital actor video, the image selection area is used to determine the digital actor constructed by the target role object, and the interaction area is used for interaction between the user and the digital actor.

[0050] A second aspect of the present application provides a system for generating an interactive digital actor video, wherein the system for generating an interactive digital actor video specifically includes:

[0051] A construction module is used to construct a video scene database and a digital actor database based on the interactive video, wherein the video scene database includes video description information of the interactive video; the digital actor database includes feature data of role objects in the interactive video;

[0052] A generation module is used to generate a digital actor video based on the interactive voice for the interactive video, the video scene database and the digital actor database.

[0053] A third aspect of the present application provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps in any of the above-described methods for generating interactive digital actor videos.

[0054] A fourth aspect of the present application provides a terminal device, comprising: a processor and a memory;

[0055] The memory stores a computer-readable program executable by the processor;

[0056] When the processor executes the computer-readable program, the processor implements the steps in any of the above-described methods for generating an interactive digital actor video.

[0057] Beneficial effects:

[0058] 1. This application can process interactive videos into a video scene database and a digital actor database, and use the video scene database and the digital actor database to provide information support for digital actor generation, so as to generate a digital actor video stream for interacting with the interactive video, thereby improving the degree of freedom of video interaction and the interactive effect.

[0059] 2. This application can receive user interaction voice and, based on the interactive voice video scene database and the digital actor database, determine and generate a digital actor voice that closely matches the character object the user needs to interact with, thereby improving the interactive effect of the video. At the same time, this application can directly clone the timbre of the character object in the interactive video based on the audio clip of the character object through the large language model, without the need for pre-training or fine-tuning, further reducing the interaction cost.

[0060] 3. This application can accept user control information and use the control information to control the generated digital actor video so that the digital actor video can better meet user needs. BRIEF DESCRIPTION OF THE DRAWINGS

[0061] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0062] Figure 1 A schematic diagram of an application scenario of the method for generating an interactive digital actor video provided in an embodiment of the present application.

[0063] Figure 2 A flowchart of a method for generating an interactive digital actor video provided in an embodiment of the present application.

[0064] Figure 3 A flowchart illustrating an example of a method for generating an interactive digital actor video provided in an embodiment of the present application.

[0065] Figure 4 The flowchart of an example of the process of generating a video scene database and a digital actor database is shown in FIG.

[0066] Figure 5 A flowchart illustrating an example of the process of generating music for digital actors.

[0067] Figure 6 A flowchart illustrating an example of a process for generating a digital actor video stream.

[0068] Figure 7 The following is an example of the display interface.

[0069] Figure 8 This is a functional block diagram of the system for generating interactive digital actor videos provided in an embodiment of the present application.

[0070] Figure 9 This is a block diagram of the principles of the terminal device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0071] The embodiments of the present application provide a method, system, device, and computer storage medium for generating interactive digital actor videos. To clarify the objectives, technical solutions, and effects of the present application, the present application is further described below with reference to the accompanying drawings and examples. It should be understood that the specific embodiments described herein are intended only to illustrate the present application and are not intended to limit the present application.

[0072] It will be understood by those skilled in the art that, unless expressly stated otherwise, the singular forms "a", "an", "said" and "the" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application refers to the presence of the features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it may be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connections or wireless couplings. The term "and / or" used herein includes all or any units and all combinations of one or more associated listed items.

[0073] It will be understood by those skilled in the art that, unless otherwise defined, all terms (including technical and scientific terms) used herein have the same meaning as commonly understood by those skilled in the art to which this application belongs. It should also be understood that terms such as those defined in common dictionaries should be understood to have meanings consistent with their meanings in the context of the prior art and will not be interpreted in an idealized or overly formal sense unless specifically defined as herein.

[0074] It should be understood that the sequence numbers and sizes of the steps in this embodiment do not imply the order of execution. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiment of this application.

[0075] After research, it was found that current video platforms provide users with interactive interfaces through which users interact with videos. Among them, the interactive interface is mainly divided into two methods: barrage or chat box and pre-recorded video clips with multiple branches. Barrage or chat box uses pop-up windows to count user preferences during video playback, or edit barrage to express user attitudes. This method can only watch the communication between users, and is often a one-sided and low-frequency information output by users, making it difficult to form effective interaction on a large scale. Pre-recorded video clips with multiple branches interact by detecting different user choices to play corresponding clips. Users can choose different pre-designed branches in the video to view the story development after the corresponding branch. This method requires manual preparation of videos with different branches. In particular, the story development of different branches often requires a lot of time and effort to design.

[0076] With the rapid development of technologies such as artificial intelligence, digital humans are being used to interact with users through videos. However, when interacting with users through digital humans, digital humans generally use pre-designed character tones and speaking habits to communicate with users around video content. This makes it difficult for digital humans to be flexibly applied to multimodal real-time interactive portrait-driven applications, especially video images with rich emotions and personalities.

[0077] To address the above issues, in an embodiment of the present application, a video scene database and a digital actor database are constructed based on the interactive video; a digital actor video is generated based on the interactive speech in the interactive video, the video scene database, and the digital actor database, wherein the video scene database includes video description information of the interactive video; and the digital actor database includes feature data of the character objects in the interactive video. This embodiment of the present application processes the interactive video into a video scene database and a digital actor database, and uses these video scene databases and digital actor databases to provide information support for digital actor generation. This allows for rapid generation of digital actor voices based on the interactive speech, improving the freedom and effectiveness of video interaction.

[0078] An application environment diagram of the method for generating interactive digital actor videos provided in the embodiment of the present application can be as follows: Figure 1 As shown. Figure 1 The system for generating an interactive digital actor video includes a user terminal 110 and a server 120. The user terminal 110 and the server 120 are connected via a network. The user terminal 110 can specifically be a desktop user terminal 110 or a mobile user terminal 110, and the mobile user terminal 110 can specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented as an independent server 120 or a server 120 cluster consisting of multiple servers 120. The user terminal can play the interactive video and receive the user's interactive voice for the interactive video. The server 120 obtains the interactive voice from the user terminal 110, and the server 120 processes the interactive video to form a video scene database and a digital actor database, and generates a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database. The user terminal can play the digital actor video.

[0079] The application content will be further explained below through description of embodiments in conjunction with the accompanying drawings.

[0080] like Figure 2 and Figure 3 As shown, the method for generating an interactive digital actor video provided in this embodiment specifically includes:

[0081] S10. Construct a video scene database and a digital actor database based on the interactive video.

[0082] Specifically, the interactive video can be a video uploaded by the user, a video obtained from the server according to the user's playback instruction, a video stored locally on the user terminal, or a video sent by an external device, etc. In addition, the video format of the interactive video and the language used by the video are not limited here. For example, the video format of the interactive video can be p4, avi, m4v, flv, etc., and the language used by the video can be Chinese, English, French, etc. In other words, the embodiment of the present application can interact with videos of any format and any language. In addition, the interactive video can include a number of character objects, which can be human characters, animal characters, etc. Each of the several character objects can be used to form a digital actor with a consistent character appearance, speech tone and style, and the digital actor is used to interact with the user for the interactive video. For example, if the interactive video is a video centered on a character (such as a short play, variety show, interview, talk show, TV series, etc.), a digital actor will be created for the character. The digital actor and its corresponding character have the same appearance, speaking tone and style, and the digital actor can interact with the user in response to the interactive video. The interaction method can be text interaction, voice interaction, etc., and each interaction method can use any language, for example, Chinese, English, French, etc.

[0083] The video scene database includes video description information of scene video clips in the interactive video; the digital actor database includes feature data of each character object in the scene video clip. The video description information is used to describe the video scene in the interactive video, and the feature data is feature information of each character object, which is used to reflect the character object's appearance, speech timbre, speech style, and other characteristics. The video scene database and the digital actor database are obtained by processing the interactive video, and the interactive video can be processed when the interactive video is acquired, when the interactive video is played through the video playback platform, or when the interactive voice for the interactive video is first received. In addition, the processing time of the interactive video can also be determined based on the source of the interactive video. For example, when the interactive video is a user-uploaded video, the time point when the interactive video is acquired can be used as the processing time; when the interactive video is a locally stored video, the time point when the interactive video is played through the video playback platform can be used as the processing time; when the interactive video is a video that is currently being played, the time point when the interactive voice based on the interactive video is first received can be used as the processing time, etc. Of course, the processing timing of the interactive video can also be pre-set. For example, the acquisition of the interactive video can be used as the processing timing so that the video scene database and the digital actor database can be acquired in advance, so that when the user's interactive voice is received, the video scene database and the digital actor database can be directly read, avoiding the interaction delay caused by acquiring the video scene database and the digital actor database.

[0084] In one implementation, Figure 4 As shown, the construction of the video scene database and the digital actor database based on the interactive video specifically includes:

[0085] S11, dividing the interactive video into a plurality of scene video segments, and adding a segment identifier to each scene video segment;

[0086] S12, obtaining video description information of each scene segment and feature data of the character object in each scene video segment;

[0087] S13, constructing a video scene database according to the video description information of each scene video clip;

[0088] S14. Construct a digital actor database based on the feature data of the role objects in each scene video clip.

[0089] In step S21, several scene video segments are obtained by segmenting the interactive video. The segmentation of the interactive video can be based on scene changes, shot changes, etc. That is, after acquiring the interactive video, a temporal analysis of the video content of the interactive video is performed to obtain transition time points in the interactive video that meet the segmentation criteria. The several scene videos are then segmented using each transition time point as a segmentation point to obtain several scene video segments. At the same time, a segment identifier is assigned to each segmented scene video segment. Each scene video segment is assigned a different segment identifier, so that the scene video segment can be quickly located based on the segment identifier.

[0090] In one implementation, after acquiring an interactive video, a shot cut detection algorithm is used to perform temporal analysis of the video content based on scene and / or shot cuts in the interactive video to identify transition points in the video segments. The interactive video is then segmented into several video segments based on the transition points. Each segment is considered a scene video segment, and a segment identifier is assigned to each scene video segment, which uniquely identifies the scene video segment. The shot cut detection algorithm can utilize techniques such as pyscenedetect, and the segment identifiers can be assigned based on the temporal order of the scene video segments in the interactive video. In other implementations, the interactive video can also be segmented based on character objects. Alternatively, after segmenting to obtain several video segments, the segmented videos can be filtered (e.g., retaining segments containing character objects) and used as the segmented scene video segments.

[0091] In step S12, video description information is obtained by performing content analysis on the scene video clip, so that the video description information can reflect the content information of the scene video clip. The video description information may include one or more of the following: a text description of the scene, dialogue content, a timestamp of the scene video clip, and identification of character objects (character objects contained in the scene video clip). For example, the scene video clip is input into a scene understanding module equipped with a video understanding algorithm. The scene understanding module identifies the scene structure and key elements in the scene video clip and outputs video description information for the scene video clip. The video description information includes the timestamp of the scene video clip, identification of character objects, dialogue content, and text description of the scene video clip. The video understanding algorithm may use a video-captioning or HumanOmni algorithm.

[0092] Furthermore, after obtaining the video description information of each scene video clip, the feature data of the character object in each scene video clip can be obtained, wherein the feature data may include voice representation data and / or face representation data, etc., and the face representation data may include one or more of a three-dimensional face representation, a face cropped video, and a face expression. In a specific implementation, the feature data of the character object includes voice representation data and face representation data, and the face representation data includes a face cropped video and a three-dimensional face representation. Accordingly, if Figure 4 As shown, the process of acquiring the feature data of the character object in each scene video clip includes:

[0093] S121, dividing the audio data corresponding to each scene video clip into a plurality of audio clips based on the character object, and adding a sequence number label to each audio clip;

[0094] S122, extracting the speech style of each audio clip to obtain voice representation data of the character object corresponding to each audio clip;

[0095] S123, performing face detection on each scene video clip to obtain a facial cropping video corresponding to each character object;

[0096] S124, extracting a three-dimensional facial representation from the face cropping video to obtain a three-dimensional facial representation of each character object;

[0097] S125 , using the voice representation data, the facial cropping video, and the three-dimensional facial representation of each character object as feature data of each character object.

[0098] Specifically, each scene video clip may include one character object (e.g., a scene video clip is a monologue scene for one character object) or multiple character objects (e.g., a scene video clip is a dialogue scene for multiple character objects). To obtain voice representation data for each character object in a scene video clip, the audio data corresponding to the scene video clip may be first extracted. Then, based on the character object, the vocal time intervals in the audio data corresponding to the scene video clip may be detected. Finally, the audio data corresponding to the scene video clip may be segmented into several audio segments based on the detected vocal time intervals. The vocal time intervals may be detected using an active vocal detection algorithm, such as silero_vad.

[0099] Furthermore, each audio clip includes all the audio data of a character object. For example, if the scene video clip contains two character objects, then the scene video clip includes the speaking video clip A1 of character object A, the speaking video clip B1 of character object B, the speaking video clip A2 of character object A, and the speaking video clip B2 of character object B. Then, four vocal time intervals can be detected in the scene video clip, and the corresponding audio data are the vocal time interval of the speaking video clip A1, the vocal time interval of the speaking video clip B1, the vocal time interval of the speaking video clip A2, and the vocal time interval of the speaking video clip B2. Accordingly, the audio data of the scene video clip will be divided into two audio clips, which are the audio clips corresponding to the vocal time interval of the speaking video clip A1 and the vocal time interval of the speaking video clip A2 of character object A; and the audio clips corresponding to the vocal time interval of the speaking video clip B1 and the vocal time interval of the speaking video clip B2 of character object B. Of course, in actual applications, other methods can also be used to divide the scene video clip into several audio clips. Taking the above-mentioned scene video clip as an example, each audio clip includes a continuous audio data of a character object. Then, the audio data corresponding to the scene video clip will be divided into 4 audio clips, namely the audio clip corresponding to the vocal time interval of the speaking video clip A1, the audio clip corresponding to the vocal time interval of the speaking video clip B1, the audio clip corresponding to the vocal time interval of the speaking video clip A2, and the audio clip corresponding to the vocal time interval of the speaking video clip B2.

[0100] After segmenting to obtain several audio clips, speech style extraction is performed on each audio clip to obtain speech representation data for each character object. The speech representation data is used to reflect the emotional changes and / or language style of the digital actor, and may include intonation, speaking speed, emotion, etc. The speech representation data can be obtained by extracting the speech style of the character object in each audio clip using a pre-trained speech style extraction network (such as wav2vec, hubert, etc. speech encoder).

[0101] A facial cropping video is obtained by cropping the facial region of a human face in a scene video clip, and each facial cropping video corresponds to a character object. Specifically, facial recognition (for example, using facial recognition algorithms such as dlib, YOLO, or face-alignment) can be used to detect the facial images of the character objects in the scene video clip, and an image sequence consisting of all facial images corresponding to the character objects can be used as a facial cropping video to obtain a facial cropping video corresponding to each character object. In addition, after each facial cropping video corresponding to each character object is obtained, a 3D facial representation of each character object can be extracted using 3D reconstruction technology to obtain a 3D facial representation of each character object. The 3D facial representation can be an implicit expression using FLAME representation, BFM representation, or warping-vae.

[0102] After acquiring the voice representation data, the face cropping video, and the three-dimensional face representation of each character object, the voice representation data, the face cropping video, and the three-dimensional face representation can be used as feature data of a character object.

[0103] Further, in step S13, when constructing a video scene database based on the video description information of each scene video clip, the video description information of each scene video clip can be directly used as the video scene data of the scene video clip, and the video scene data of the scene video clip can be associated with each scene video clip, and then the database composed of the video scene data of all scene video clips can be used as the video scene database.

[0104] In one implementation, to enrich the video scene database, when constructing the video scene database based on the video description information of each scene video clip, a video summary of the video clip is first determined based on the video description information, and then the video scene database is constructed based on the video summary and the video description information. Based on this, constructing the video scene database based on the video description information of each scene video clip specifically includes:

[0105] determining a video summary of each scene video clip based on the video description information of each scene video clip;

[0106] A video scene database is constructed according to the video description information and video summary of each scene video clip.

[0107] Specifically, the video summary is a summary description of the scene video clip, which is determined based on the video description information. That is, the video clip can be summarized based on the video description information to obtain a video summary, and the video summary is added to the video description letter, so that the video description letter also includes the video summary.

[0108] Exemplarily, determining the video summary of each scene video clip based on the video description information of each scene video clip specifically includes:

[0109] Perform speech recognition on each audio clip corresponding to each scene video clip to obtain the dialogue text of each audio clip;

[0110] The video description information of each scene video clip is updated based on all the dialogue texts corresponding to each scene video clip, and the updated video description information is summarized by a large language model to obtain a video summary of each scene video clip.

[0111] Specifically, the audio segment is obtained by segmenting the audio data of the scene video segment in step S12. That is, after the audio data corresponding to the scene video segment is segmented into several audio segments, the dialogue content in each audio segment can be converted into dialogue text, and then the dialogue text is used to update the video description information. When the video description information includes dialogue text, the dialogue text is used to update the dialogue text in the video description information to update the video description information. When the video description information does not include dialogue text, the dialogue text is added to the video description information to update the video description information.

[0112] After obtaining the updated video description information, it is fed into a large language model (such as OpenAI's ChatGpt or the DeepSeek series of models). This large language model summarizes the video description information for each scene video clip to generate a video summary for each scene video clip. This removes redundant information from the scene understanding output, facilitating quick browsing and subsequent analysis. Here, the video description information is updated using the dialogue content of each audio clip, ensuring that the video description only includes the dialogue content of the characters, filtering out other distracting dialogue content, thereby improving the accuracy of the video summary.

[0113] Of course, in practical applications, it is also possible not to extract the dialogue content of each audio clip, but to directly summarize the video description information through a large language model to generate a video summary of the scene video clip.

[0114] Exemplarily, constructing a video scene database according to the video description information of each scene video clip specifically includes:

[0115] Using the segment identifier of each scene video segment as the key and the video description information and summary information of each scene video segment as the value, construct the video scene data corresponding to each scene video segment;

[0116] A video scene database is formed based on the video scene data corresponding to all scene video clips.

[0117] Specifically, in the video scene database, video description information is stored in a key-value format, where the key is the segment identifier and the value is the video description information and video summary, so that the video scene data corresponding to each scene video segment can be quickly determined.

[0118] Further, in step S14, after acquiring the feature data of each character object in each scene video clip, as shown in FIG. Figure 4 As shown, a database consisting of the feature data of each character object in all scene video clips is used as a digital actor database to provide information support for user voice interaction and digital actor generation. The digital actor database is constructed based on the feature data of the character object in each scene video clip, which specifically includes:

[0119] Use each character object as the key and the characteristic data of each character object as the value to construct digital actor data;

[0120] A digital actor database is formed based on the digital actor data corresponding to all role objects.

[0121] Specifically, in the digital actor database, feature data is also stored in key-value format, where the key is the character object identifier and the value is the feature data. The feature data stored in key-value format is also associated with the segment identifier, so that the feature data of the character object in each video segment can be quickly determined.

[0122] Of course, in actual applications, the video scene database and digital actor database corresponding to the interactive video can be stored together with the video scene database and digital actor database of other interactive videos, and the video scene database and digital actor database corresponding to the interactive video can be associated with the interactive video.

[0123] The embodiment of the present application converts the interactive video into a video with which the character objects can interact by obtaining the video scene database and the digital actor database of the interactive video. The user can interact with the character object based on the video scene database and the digital actor database, and supports real-time video interaction between the user and the user character, thereby improving the existing video expression capabilities.

[0124] S20: Generate a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database.

[0125] Specifically, interactive voice can be the voice generated by a user interacting with an interactive video. This interactive voice can be picked up by a terminal device with a sound pickup function, transmitted by other devices, or uploaded by the user. For example, the display used to play the interactive video is equipped with a sound pickup function (such as a microphone), and the sound pickup function can be used to pick up the user's voice in response to the interactive video to obtain the interactive voice. Of course, in actual applications, in order to accurately capture the user's interactive voice in response to the interactive video, a trigger button can be configured on the interactive video playback interface. When the trigger button is triggered, it is determined that the user has begun to input interactive voice in response to the interactive video, and the user voice is picked up to obtain the interactive voice. Alternatively, the user voice can be picked up in real time, and the correlation between the user voice and the interactive video is detected. When the detected correlation meets the requirements (for example, including the video title, actor names, video lines, etc. of the interactive video), the picked up user voice is used as the interactive voice.

[0126] For example: a user plays an interactive video through a user terminal equipped with a speaker and a microphone. While watching the interactive video, the user speaks a voice commenting on a character object in the interactive video. The user terminal will then pick up the comment voice through the microphone and use the comment voice as the interactive voice for the interactive video.

[0127] The digital actor's voice is the response voice corresponding to the interactive voice. Its voice style is cloned from the character's speaking style, and its content is the response content about the interactive video, generated based on the video description information. The digital actor's video stream is the sequence of video frames corresponding to the response voice. The digital actor in the digital actor video stream is cloned from the character in the interactive video, using the facial features of the corresponding character. The following describes the acquisition process of the digital actor's voice and the digital actor video stream.

[0128] Exemplarily, generating the digital actor voice based on the interactive voice, the video scene database, and the digital actor database specifically includes:

[0129] H10. Encode the interactive speech to obtain a speech representation of the interactive speech;

[0130] H20, selecting video description information corresponding to the interactive voice in the video scene database;

[0131] H30, based on the video description information and the voice representation, obtaining target voice data for driving a digital actor constructed based on a target role object through a large language model;

[0132] H40. Select target feature data of a target role object from the digital actor database, and generate the digital actor voice according to the target feature data and the target voice data.

[0133] Specifically, in step H10, the target character object can be determined based on user selection. For example, after receiving the interactive voice, the scene video clip corresponding to the interactive voice is obtained, and then the character object identifier included in the scene video clip is read in the video scene database, and the obtained character object identifier is fed back to the user, and the target character object is determined based on the user's selection operation. Of course, in actual applications, other methods can also be used to determine the target character object, for example, directly identifying the target character object from the interactive voice, or using the character object displayed in the interactive video when the interactive voice is received as the target character object, or first allowing the user to select the target character object. If the target character object selected by the user is not obtained, the character object displayed in the interactive video when the interactive voice is received is used as the target character object, etc.

[0134] The speech representation can be obtained by extracting the speech representation of the interactive speech, such as Figure 5 As shown, after the interactive voice is detected by active human voice detection (such as by a human voice detection algorithm, etc.), the interactive voice is audio-encoded by a voice representation encoder (such as an ASR algorithm) to obtain a voice representation of the interactive voice.

[0135] In step H20, the video description information is the video description information of the scene video segment with which the interactive voice needs to interact, which is selected from the video scene database. The process of obtaining the video description information can first obtain the reception time of the interactive voice, and then determine the scene video segment with which the interactive voice needs to interact based on the reception time, and then read the corresponding video description information in the video scene database based on the segment identifier of the video segment; or, the video description information in the interactive voice can be first identified, and then the segment identifier corresponding to the interactive voice can be determined based on the video description information, and then the scene video segment can be determined based on the segment identifier, etc.

[0136] In step H30, if Figure 5 As shown, after obtaining the speech representation and video description information, the speech representation and video description information can be first input into the large language model. The large language model uses the video description information as context information to generate text information for the speech representation to drive the digital actor built based on the target role object, and converts the text information into speech data. For example, each text block is converted into speech data through a text-to-speech algorithm (such as the TTS algorithm).

[0137] In step H40, after acquiring the voice data, the speaking style of the voice data is cloned into the timbre of the target character object. Specifically, after acquiring the voice data, target feature data of the target character object is selected from the digital actor database, and then the voice data is subjected to timbre cloning based on the target feature data. The timbre cloning process can be performed using a sound cloning algorithm (such as voice-craft) or a large language model. For example, the target feature data and a speech example of the target character object are used as prior knowledge, and this prior knowledge and the voice data are input into the large language model. The timbre of the target character object is cloned using the large language model to generate a digital actor voice with the timbre of the target character object.

[0138] This embodiment of the application uses video description information as contextual information and generates a digital actor voice corresponding to the interactive speech based on feature data. This digital actor voice can be matched to the context of the scene video clip, ensuring that the digital actor voice closely matches the character's personality. In addition, the target feature data of the target character object is pre-extracted. When determining the digital actor voice, only reference audio of the target character object (e.g., 5-10 seconds) is required to generate the digital actor voice corresponding to the target character object. No pre-training or fine-tuning is required, reducing the cost of generating the digital actor voice.

[0139] The above completes the description of the process of generating the digital actor's voice. The following is a description of the process of generating the digital actor's video stream. The digital actor's video stream is generated based on the digital actor's voice and the digital actor database. The process of generating the digital actor's video stream based on the digital actor's voice and the digital actor database specifically includes:

[0140] Selecting target feature data of a target role object from the digital actor database;

[0141] Generate facial movement features for driving a digital actor constructed based on a target character object based on the target feature data and the digital actor voice;

[0142] A digital actor video stream is synthesized based on the facial motion features and the target feature data.

[0143] Specifically, the digital actor video stream is a sequence of video frames corresponding to the digital actor's voice, wherein each video frame in the digital actor video stream contains a digital actor constructed based on the target character object, and the digital actor's facial features are constructed based on the facial features of the target character object. In other words, a digital actor will be created based on the target character object, and the digital actor will perform the digital actor's voice with the facial expressions and emotions of the target character object. Based on this, Figure 6As shown, after acquiring the target feature data, facial motion features are first generated from the digital actor's voice and feature data (e.g., using a seq2seq model). These facial motion features are a set of feature vectors that record the movement of the target character's head and facial expressions. Portrait video synthesis is then performed based on these facial motion features to produce a digital actor video stream featuring the target character as the digital actor. This portrait video synthesis can be performed using a pre-configured portrait video synthesis network (e.g., a U-Net network or Gaussian splattering).

[0144] It should be noted that when generating a digital actor video stream, the user can also control the digital actor in the digital actor video stream. Based on this, in one implementation, Figure 6 As shown, the synthesizing of the digital actor video stream based on the facial motion features and the target feature data specifically includes:

[0145] Receive control information from user interaction;

[0146] A digital actor video stream is synthesized based on the control information, the facial motion features and the target feature data.

[0147] Specifically, the control information is user-interactive and is used to control the digital actor constructed based on the target character object, for example, to control the digital actor's emotions, facial perspective, and real-time dialogue content. The control information can be one or more of text, audio, image, and video information. For example, the control information can be a segment of dialogue, or a segment of dialogue and a character image. After obtaining the control information, a digital actor video stream can be synthesized based on the control information and facial motion feature input. During the digital actor video stream synthesis process, the facial motion features are adjusted according to the control information to generate a digital actor video stream that conforms to the control information.

[0148] After acquiring the digital actor's voice and digital actor video streams, a digital actor video can be generated based on the digital actor's voice and the digital actor video streams. The digital actor's voice is the audio data of the digital actor video, and the digital actor video stream is the image data of the digital actor video. This digital actor video is used to interact with the user. The interaction can be synchronized with the interactive video, played while the interactive video is paused, or played on a user-specified device.

[0149] In one implementation, the interactive method configures a playback area for the digital actor video, and then plays the digital actor video in the playback area. Accordingly, after generating the digital actor video based on the digital actor's voice and the digital actor video stream, the method further includes:

[0150] A playback area, an image selection area, and an image interaction area are configured for the digital actor video. The playback area is used to play the digital actor video, the image selection area is used to determine the digital actor constructed by the target role object, and the interaction area is used for interaction between the user and the digital actor.

[0151] Specifically, the playback area corresponding to the digital actor video and the playback area corresponding to the interactive video are located in the same display interface, and the user can control the interactive video through the playback area corresponding to the interactive video, for example, play, pause or adjust the video playback progress. Figure 7 As shown, the display interface is configured with an image selection area and an image interaction area. The image selection area is used to select an area to determine a character object, and the image interaction area is used to input control information. In other words, the target character object can be determined through the image selection area, and control information can be exchanged through the image interaction area.

[0152] In addition, the display interface may also include a chat history record area, which displays each user's interaction records for the interactive video, such as chat records, digital actor video playback records, etc. Users can review past conversation history and view historical digital actor videos through the chat history record area, and quickly jump to a historical conversation or historical digital actor video, etc.

[0153] In summary, this embodiment provides a method for generating an interactive digital actor video, comprising processing the interactive video to form a video scene database and a digital actor database; generating a digital actor voice based on the interactive voice for the interactive video, the video scene database, and the digital actor database, and generating a digital actor video stream based on the digital actor voice and the digital actor database; and generating a digital actor video based on the digital actor voice and the digital actor video stream. This embodiment of the application processes the interactive video into a video scene database and a digital actor database, and uses these video scene database and the digital actor database to provide information support for digital actor generation, thereby improving the degree of freedom and interactive effect of video interaction.

[0154] Based on the above-mentioned method for generating interactive digital actor videos, this embodiment provides a system for generating interactive digital actor videos, such as Figure 8 As shown, the interactive digital actor video generation system specifically includes:

[0155] The receiving module 100 is configured to construct a video scene database and a digital actor database based on the interactive video, wherein the video scene database includes video description information of the interactive video; and the digital actor database includes feature data of character objects in the interactive video;

[0156] The generating module 200 is configured to generate a digital actor video based on the interactive speech for the interactive video, the video scene database, and the digital actor database.

[0157] Based on the above-mentioned method for generating interactive digital actor videos, this embodiment provides a computer-readable storage medium, which stores one or more programs. The one or more programs can be executed by one or more processors to implement the steps in the method for generating interactive digital actor videos as described in the above-mentioned embodiment.

[0158] Based on the above-mentioned method for generating interactive digital actor videos, the present application also provides a terminal device, such as Figure 9 As shown, it includes at least one processor 20; a display screen 21; and a memory 22. It may also include a communications interface 23 and a bus 24. The processor 20, display screen 21, memory 22, and communications interface 23 can communicate with each other via bus 24. The display screen 21 is configured to display a preset user guidance interface in the initial setup mode. The communications interface 23 can transmit information. The processor 20 can invoke logic instructions in the memory 22 to execute the method described in the above embodiment.

[0159] In addition, the logic instructions in the memory 22 can be implemented in the form of software functional units and can be stored in a computer-readable storage medium when sold or used as an independent product.

[0160] The memory 22, as a computer-readable storage medium, can be configured to store software programs or computer-executable programs, such as program instructions or modules corresponding to the methods in the embodiments of the present disclosure. The processor 20 executes the software programs, instructions, or modules stored in the memory 22 to perform functional applications and data processing, thereby implementing the methods in the above embodiments.

[0161] The memory 22 may include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function; the data storage area may store data created based on the use of the terminal device. In addition, the memory 22 may include high-speed random access memory and non-volatile memory. For example, various media that can store program code, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, may also be transient storage media.

[0162] In addition, the specific process of loading and executing the multiple instructions in the storage medium and the processor in the terminal device has been described in detail in the above method and will not be described here one by one.

[0163] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for generating an interactive digital actor video, characterized in that: The method for generating an interactive digital actor video specifically includes: A video scene database and a digital actor database are constructed based on the interactive video, wherein the video scene database and the digital actor database are obtained by processing the interactive video, and the interactive video is processed when the interactive video is acquired; the video scene database includes video description information of the interactive video, and the video description information reflects the segment content information of the scene video segment; the digital actor database includes feature data of the character objects in the interactive video, and the feature data is used to reflect the character objects' appearance, speech timbre, and speech style; generating a digital actor video based on the interactive speech for the interactive video, the video scene database, and the digital actor database; The step of constructing a video scene database based on interactive videos specifically includes: Segmenting the interactive video into a plurality of scene video segments, and adding a segment identifier to each scene video segment; Get the video description information of each scene segment; determining a video summary of each scene video clip based on the video description information of each scene video clip; Constructing a video scene database based on the video description information and video summary of each scene video clip; The process of acquiring the characteristic data specifically includes: The audio data corresponding to each scene video clip is divided into several audio clips based on the character object, and a serial number label is added to each audio clip; Extract the speech style of each audio clip to obtain the voice representation data of the character object corresponding to each audio clip; Perform face detection on each scene video clip to obtain the facial cropped video corresponding to each character object; Extracting a three-dimensional facial representation from the facial cropped video to obtain a three-dimensional facial representation of each character object; The voice representation data, the face cropped video and the three-dimensional representation of the face of each character object are used as feature data of each character object; The generating of the digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database specifically includes: generating a digital actor voice based on the interactive voice, the video scene database, and the digital actor database; generating a digital actor video stream according to the digital actor voice and the digital actor database; A digital actor video is generated based on the digital actor voice and the digital actor video stream.

2. The method for generating an interactive digital actor video according to claim 1, wherein: The digital actor database is constructed based on the interactive video, specifically including: Obtain feature data of character objects in each scene video clip; A digital actor database is constructed according to the feature data of the role objects in each scene video clip.

3. The method for generating an interactive digital actor video according to claim 1, wherein: The determining of the video summary of each scene video clip based on the video description information of each scene video clip specifically includes: Perform speech recognition on each audio clip corresponding to each scene video clip to obtain the dialogue text of each audio clip; The video description information of each scene video clip is updated based on all the dialogue texts corresponding to each scene video clip, and the updated video description information is summarized by a large language model to obtain a video summary of each scene video clip.

4. The method for generating an interactive digital actor video according to claim 1, wherein: The step of constructing a video scene database based on the video description information and the video summary of each scene video clip specifically includes: Using the segment identifier of each scene video segment as the key and the video description information and summary information of each scene video segment as the value, construct the video scene data corresponding to each scene video segment; A video scene database is formed based on the video scene data corresponding to all scene video clips.

5. The method for generating an interactive digital actor video according to claim 1, wherein: The step of constructing a digital actor database based on the feature data of the character objects in each scene video clip specifically includes: Use each character object as the key and the characteristic data of each character object as the value to construct digital actor data; A digital actor database is formed based on the digital actor data corresponding to all role objects.

6. The method for generating an interactive digital actor video according to claim 1, wherein: Generating the digital actor voice based on the interactive voice, the video scene database, and the digital actor database specifically includes: Encoding the interactive speech to obtain a speech representation of the interactive speech; Selecting video description information corresponding to the interactive voice in the video scene database; Based on the video description information and the speech representation, obtaining target speech data for driving a digital actor constructed based on a target character object through a large language model; Selecting target feature data of a target role object from the digital actor database; The digital actor voice is generated according to the target feature data and the target voice data.

7. The method for generating an interactive digital actor video according to claim 1, wherein: Generating a digital actor video stream according to the digital actor voice and the digital actor database specifically includes: Selecting target feature data of a target role object from the digital actor database; Generate facial movement features for driving a digital actor constructed based on a target character object based on the target feature data and the digital actor voice; A digital actor video stream is synthesized based on the facial motion features and the target feature data.

8. The method for generating an interactive digital actor video according to claim 7, wherein: The synthesizing of the digital actor video stream based on the facial motion features and the target feature data specifically includes: Receive control information from user interaction; A digital actor video stream is synthesized based on the control information, the facial motion features and the target feature data.

9. The method for generating an interactive digital actor video according to claim 7, wherein: After generating the digital actor video based on the digital actor voice and the digital actor video stream, the method further includes: A playback area, an image selection area, and an image interaction area are configured for the digital actor video. The playback area is used to play the digital actor video, the image selection area is used to determine the digital actor constructed by the target role object, and the interaction area is used for interaction between the user and the digital actor.

10. A system for generating interactive digital actor videos, characterized in that: The interactive digital actor video generation system specifically includes: A construction module is configured to construct a video scene database and a digital actor database based on the interactive video, wherein the video scene database and the digital actor database are obtained by processing the interactive video, and the interactive video is processed when the interactive video is acquired; the video scene database includes video description information of the interactive video, and the video description information reflects the content information of the scene video segment; the digital actor database includes feature data of the character objects in the interactive video, and the feature data is used to reflect the character objects' appearance, speech timbre, and speech style; A generating module, configured to generate a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database; The step of constructing a video scene database based on interactive videos specifically includes: Segmenting the interactive video into a plurality of scene video segments, and adding a segment identifier to each scene video segment; Get the video description information of each scene segment; determining a video summary of each scene video clip based on the video description information of each scene video clip; Constructing a video scene database based on the video description information and video summary of each scene video clip; The process of acquiring the characteristic data specifically includes: The audio data corresponding to each scene video clip is divided into several audio clips based on the character object, and a serial number label is added to each audio clip; Extract the speech style of each audio clip to obtain the voice representation data of the character object corresponding to each audio clip; Perform face detection on each scene video clip to obtain the facial cropped video corresponding to each character object; Extracting a three-dimensional facial representation from the facial cropped video to obtain a three-dimensional facial representation of each character object; The voice representation data, the face cropped video and the three-dimensional representation of the face of each character object are used as feature data of each character object; The generating of the digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database specifically includes: generating a digital actor voice based on the interactive voice, the video scene database, and the digital actor database; generating a digital actor video stream according to the digital actor voice and the digital actor database; A digital actor video is generated based on the digital actor voice and the digital actor video stream.

11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for generating an interactive digital actor video as described in any one of claims 1 to 9.

12. A terminal device, characterized in that: include: processor and memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the processor implements the steps in the method for generating an interactive digital actor video as described in any one of claims 1 to 9.

Citation Information

Patent Citations

  • Visual interaction system based on holographic projection

    CN113821104A

  • Visual guiding method and system of scene video superimposed digital human, and storage medium

    CN116740311A

  • Simulation role-based information interaction method and device and storage medium

    CN117560340A