Generation method, system and equipment of interactive digital actor video and computer storage medium
By constructing a video scene database and a digital actor database, digital actor videos are generated, and the problem of difficulty in applying digital people in multimodal real-time interaction is solved, and the freedom and effect of video interaction is improved.
Patent Information
- Application Number
- CN202510603049.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-12
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-12
AI Technical Summary
In the prior art, digital people are difficult to flexibly apply to portrait drivers with multimodal real-time interaction when interacting with users, especially video images with rich emotions and personality, which affects the video interaction effect.
By constructing a video scene database and a digital actor database, the video description information of the interactive video and the feature data of the character object are used to generate digital actor videos to improve the freedom and effect of video interaction.
It realizes efficient interaction between digital actors and users, improves the freedom and effect of video interaction, and reduces the cost of interaction.
Smart Images

Figure CN120125718A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of human-computer interaction technologies, and particularly relates to a method, system, device, and computer storage medium for generating an interactive digital actor video. Background Art
[0002] Digital humans can simulate the expressions, actions, etc. of real people to achieve interaction with users in multiple fields such as entertainment, education, and customer service. However, when interacting with users about videos through digital humans, digital humans generally use pre-designed character tones and speaking habits to exchange with users around the video content. This makes it difficult for digital humans to be flexibly applied to portrait-driven multi-modal real-time interactions, especially in video images with rich emotions and personalities, affecting the video interaction effect.
[0003] Therefore, the existing technologies still need to be improved. Summary of the Invention
[0004] The technical problem to be solved by the present application is to provide a method, system, device, and computer storage medium for generating an interactive digital actor video in view of the deficiencies of the existing technologies.
[0005] To solve the above technical problem, a first aspect of the present application provides a method for generating an interactive digital actor video. Specifically, the method for generating an interactive digital actor video includes: Constructing a video scene database and a digital actor database according to an interactive video, where the video scene database includes video description information of the interactive video; the digital actor database includes feature data of a character object in the interactive video; Generating a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database.
[0006] In the method for generating an interactive digital actor video, the constructing of the video scene database and the digital actor database according to the interactive video specifically includes: Dividing the interactive video into several scene video segments and adding segment identifiers to each scene video segment; Obtaining video description information of each scene segment and feature data of a character object in each scene video segment; Constructing a video scene database according to the video description information of each scene video segment; Constructing a digital actor database according to the feature data of a character object in each scene video segment.
[0007] In the method for generating an interactive digital actor video, obtaining the feature data of a character object in each scene video segment specifically includes: Based on the role object, the audio data corresponding to each scene video segment is segmented into several audio segments, and sequence number tags are added to each audio segment; Extract the speech style of each audio segment to obtain the speech representation data of the role object corresponding to each audio segment; Perform face detection on each scene video segment to obtain the facial cropped video corresponding to each role object; Extract the three-dimensional facial representation of each role object from the facial cropped video to obtain the three-dimensional facial representation of each role object; Take the speech representation data, the facial cropped video, and the three-dimensional facial representation of each role object as the feature data of each role object.
[0008] The method for generating an interactive digital actor video, wherein, the construction of the video scene database according to the video description information of each scene video segment specifically includes: Determine the video summary of each scene video segment based on the video description information of each scene video segment; Construct a video scene database according to the video description information and video summary of each scene video segment.
[0009] The method for generating an interactive digital actor video, wherein, the determination of the video summary of each scene video segment based on the video description information of each scene video segment specifically includes: Perform speech recognition on each audio segment corresponding to each scene video segment to obtain the dialogue text of each audio segment; Update the video description information of each scene video segment based on all the dialogue texts corresponding to each scene video segment, and summarize the updated video description information through a large language model to obtain the video summary of each scene video segment.
[0010] The method for generating an interactive digital actor video, wherein, the construction of the video scene database according to the video description information and video summary of each scene video segment specifically includes: Construct the video scene data corresponding to each scene video segment with the segment identifier of each scene video segment as the key and the video description information and summary information of each scene video segment as the value; Form a video scene database based on the video scene data corresponding to all scene video segments.
[0011] The method for generating an interactive digital actor video, wherein, the construction of the digital actor database according to the feature data of the role object in each scene video segment specifically includes: Construct digital actor data with each character object as the key and the feature data of each character object as the value. Form a digital actor database based on the digital actor data corresponding to all character objects.
[0012] The method for generating an interactive digital actor video, wherein generating a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database specifically includes: Generate digital actor voice based on the interactive voice, the video scene database, and the digital actor database. Generate a digital actor video stream according to the digital actor voice and the digital actor database. Generate a digital actor video based on the digital actor voice and the digital actor video stream.
[0013] The method for generating an interactive digital actor video, wherein generating digital actor voice based on the interactive voice, the video scene database, and the digital actor database specifically includes: Encode the interactive voice to obtain the voice representation of the interactive voice. Select the video description information corresponding to the interactive voice in the video scene database. Based on the video description information and the voice representation, obtain the target voice data for driving the digital actor constructed based on the target character object through a large language model. Select the target feature data of the target character object in the digital actor database. Generate the digital actor voice according to the target feature data and the target voice data.
[0014] The method for generating an interactive digital actor video, wherein generating a digital actor video stream according to the digital actor voice and the digital actor database specifically includes: Select the target feature data of the target character object in the digital actor database. Generate facial motion features for driving the digital actor constructed based on the target character object based on the target feature data and the digital actor voice. Synthesize a digital actor video stream based on the facial motion features and the target feature data.
[0015] The method for generating an interactive digital actor video, wherein synthesizing a digital actor video stream based on the facial motion features and the target feature data specifically includes: Receive the control information interacted by the user. Synthesize a digital actor video stream based on the control information, the facial movement features, and the target feature data.
[0016] The method for generating an interactive digital actor video, wherein, after generating the digital actor video based on the digital actor voice and the digital actor video stream, the method further includes: Configure a playback area, an image selection area, and an image interaction area for the digital actor video. The playback area is used to play the digital actor video, the image selection area is used to determine the digital actor constructed by the target role object, and the interaction area is used for interaction between the user and the digital actor.
[0017] The second aspect of the present application provides a system for generating an interactive digital actor video, wherein the system for generating an interactive digital actor video specifically includes: A construction module, configured to construct a video scene database and a digital actor database according to the interactive video, wherein the video scene database includes video description information of the interactive video; the digital actor database includes feature data of the role object in the interactive video; A generation module, configured to generate a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database.
[0018] The third aspect of the present application provides a computer-readable storage medium, which stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for generating an interactive digital actor video as described in any one of the above.
[0019] The fourth aspect of the present application provides a terminal device, which includes: a processor and a memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, it implements the steps in the method for generating an interactive digital actor video as described in any one of the above.
[0020] Advantageous effects: 1. The present application can process the interactive video into a video scene database and a digital actor database, and use the video scene database and the digital actor database to provide information support for generating digital actors, so as to generate a digital actor video stream for interacting with the interactive video, improving the freedom and interaction effect of video interaction.
[0021] 2. This application can receive the user's interactive voice and determine and generate a digital actor voice that is close to the role object to be interacted with by the user according to the interactive voice video scene database and the digital actor database, thereby improving the interaction effect of video interaction. At the same time, this application can directly clone the timbre of the role object based on the audio clip of the role object in the interactive video through the large language model without pre-training or fine-tuning, further reducing the interaction cost.
[0022] 3. This application can accept the user's control information and control the generated digital actor video through this control information so that the digital actor video better meets the user's needs. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions in the embodiments of this application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.
[0024] Figure 1 It is a schematic diagram of an application scenario of the method for generating an interactive digital actor video provided by an embodiment of this application.
[0025] Figure 2 It is a flowchart of the method for generating an interactive digital actor video provided by an embodiment of this application.
[0026] Figure 3 It is a schematic flowchart of an example of the method for generating an interactive digital actor video provided by an embodiment of this application.
[0027] Figure 4 It is a schematic flowchart of an example of the generation process of the video scene database and the digital actor database.
[0028] Figure 5 It is a schematic flowchart of an example of the generation process of digital actor music.
[0029] Figure 6 It is a schematic flowchart of an example of the generation process of the digital actor video stream.
[0030] Figure 7 It is an example diagram of the display interface.
[0031] Figure 8 It is a principle block diagram of the system for generating an interactive digital actor video provided by an embodiment of this application.
[0032] Figure 9 It is a principle block diagram of the terminal device provided by an embodiment of this application. Detailed implementation manners
[0033] The embodiments of the present application provide a method, a system, a device and a computer storage medium for generating an interactive digital actor video. To make the purpose, technical solutions and effects of the present application clearer and more definite, the following further describes the present application in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0034] Those skilled in the art of the present technology can understand that unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means the presence of the described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more related listed items.
[0035] Those skilled in the art of the present technology can understand that unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as the general understanding of those of ordinary skill in the art to which the present application belongs. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.
[0036] It should be understood that the sequence numbers and magnitudes of the steps in this embodiment do not mean the order of execution is prior or subsequent. The order of execution of each process is determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0037] Through research, it is found that current video platforms provide users with interaction interfaces through which users can interact with videos. Among them, the interaction interfaces are mainly divided into two ways: bullet screens or chat boxes and pre-recorded video segments with multiple branches. Bullet screens or chat boxes are used to count users' preferences in the form of pop-up windows during video playback, or to express users' attitudes by editing bullet screens. This way can only observe the communication between users, and it is often a one-sided and low-frequency information output by users, making it difficult to form large-scale effective interactions. Pre-recording video segments with multiple branches is to play corresponding segments for interaction by detecting different choices of users. Users can select different branches pre-designed in the video to view the subsequent development of the story in the corresponding branches. This way requires manual preparation of videos with different branches. In particular, for the subsequent development of the story in different branches, a large amount of time and effort are often required for design.
[0038] With the rapid development of technologies such as artificial intelligence, digital humans have been tried to be used for video interaction with users. However, when interacting with users about videos through digital humans, digital humans generally use pre-designed character tones and speaking habits to exchange with users around video content. This makes it difficult for digital humans to be flexibly applied to portrait-driven multi-modal real-time interactions, especially in video images with rich emotions and personalities.
[0039] To solve the above problems, in the embodiments of the present application, a video scene database and a digital actor database are constructed according to the interactive video; a digital actor video is generated based on the interactive voice for the interactive video, the video scene database, and the digital actor database, where the video scene database includes video description information of the interactive video; the digital actor database includes feature data of the role objects in the interactive video. The embodiments of the present application process the interactive video into a video scene database and a digital actor database, and use the video scene database and the digital actor database to provide information support for generating digital actors, so that digital actor voices can be quickly generated for the interactive voice, improving the freedom and effect of video interaction.
[0040] An application environment diagram of the method for generating an interactive digital actor video provided by the embodiments of the present application can be as Figure 1 shown. Refer to Figure 1, the generation system of the interactive digital actor video includes a user terminal 110 and a server 120. The user terminal 110 and the server 120 are connected through a network. The user terminal 110 can specifically be a desktop user terminal 110 or a mobile user terminal 110, and the mobile user terminal 110 can specifically be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented by an independent server 120 or a server cluster composed of multiple servers 120. Among them, the user terminal can play the interactive video and receive the interactive voice of the user for the interactive video. The server 120 obtains the interactive voice from the user terminal 110, processes the interactive video to form a video scene database and a digital actor database, and generates a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database. The user terminal can play the digital actor video.
[0041] The following further illustrates the application content by describing the embodiments in conjunction with the accompanying drawings.
[0042] As Figure 2 and Figure 3 shown, the generation method of the interactive digital actor video provided in this embodiment specifically includes: S10. Construct a video scene database and a digital actor database according to the interactive video.
[0043] Specifically, the interactive video can be a video uploaded by the user, a video obtained from the server according to the user's play instruction, a video locally stored in the user terminal, or a video sent by an external device, etc. Moreover, the video format of the interactive video and the language used in the video are not limited here. For example, the video format of the interactive video can be p4, avi, m4v, flv, etc., and the languages used in the video can be Chinese, English, French, etc. That is to say, the embodiments of the present application can interact with videos in any format and any language. In addition, the interactive video can include several role objects, and the role objects can be human role objects, animal role objects, etc. Each role object among the several role objects can be used to form a digital actor with a consistent role appearance, speaking tone, and style, and this digital actor is used to interact with the user for the interactive video. For example, if the interactive video is a video centered on a human role object (such as a short play, variety show, interview, talk show, TV drama, etc.), then a digital actor will be created for the human role object, and this digital actor and its corresponding human role object have the same human appearance, speaking tone, and style, and this digital actor can interact with the user for this interactive video. Among them, the interaction methods can be text interaction, voice interaction, etc., and each interaction method can use languages in any language. For example, Chinese, English, French, etc. can be used.
[0044] The video scene database includes video description information of scene video segments in the interactive video; the digital actor database includes feature data of each character object in the scene video segments. The video description information is used to describe the video scenes in the interactive video, and the feature data is the feature information of each character object, which is used to reflect the characteristics of the character object such as the appearance, speaking tone, and speaking style. Among them, the video scene database and the digital actor database are obtained by processing the interactive video, and the processing time of the interactive video can be when the interactive video is acquired, or when the interactive video is played through the video playback platform, or when the interactive voice for the interactive video is received for the first time, etc. In addition, the processing time of the interactive video can also be determined according to the source of the interactive video. For example, when the interactive video is a user-uploaded video, the time node when the interactive video is acquired can be used as the processing time; when the interactive video is a locally stored video, the time node when the interactive video is played through the video playback platform can be used as the processing time; when the interactive video is a currently playing video, the time node when the interactive voice based on the interactive video is received for the first time can be used as the processing time, etc. Of course, the processing time of the interactive video can also be preset. For example, taking the acquisition of the interactive video as the processing time, so as to pre-obtain the video scene database and the digital actor database, so that when receiving the interactive voice of the user, the video scene database and the digital actor database can be directly read, avoiding the interactive delay caused by obtaining the video scene database and the digital actor database.
[0045] In one implementation, as Figure 4 shown, the construction of the video scene database and the digital actor database according to the interactive video specifically includes: S11. Split the interactive video into several scene video segments and add segment identifiers to each scene video segment; S12. Obtain the video description information of each scene segment and the feature data of the character objects in each scene video segment; S13. Construct a video scene database according to the video description information of each scene video segment; S14. Construct a digital actor database according to the feature data of the character objects in each scene video segment.
[0046] In step S21, several scene video segments are obtained by splitting an interactive video. Among them, the splitting basis of the interactive video can be scene switching, shot switching, etc. That is to say, after obtaining the interactive video, the video content of the interactive video will be analyzed in time series to obtain the transformation time points that meet the splitting basis in the interactive video, and then each transformation time point is used as a splitting point to split several scene videos to obtain several scene video segments. At the same time, a segment identifier will also be configured for each split scene video segment, and the segment identifiers configured for each scene video segment are different, so that the scene video segment can be quickly located according to the segment identifier.
[0047] In one implementation, after obtaining the interactive video, based on scene switching and / or shot switching in the interactive video, the video content is analyzed in time series through a shot detection algorithm to identify the transformation time points of the video segments in the interactive video; and the interactive video is split into several segmented videos with the transformation time points as splitting points, and each segmented video is used as a scene video segment, and a segment identifier is configured for each scene video segment to use the segment identifier as the unique identifier of the scene video segment. Among them, the shot detection algorithm can use technologies such as pyscenedetect, and the segment identifier can be configured according to the time series order of the scene video segments in the interactive video. Of course, in other implementations, the interactive video can also be split based on the role object as the splitting basis, or after splitting several segmented videos, several segmented videos can be screened (such as retaining the segmented videos containing role objects, etc.), and the screened segmented videos are used as several scene video segments obtained by splitting, etc.
[0048] In step S12, the video description information is obtained by analyzing the content of the scene video segment, so that the video description information can reflect the segment content information of the scene video segment. The video description information can include one or more of scene text description, dialogue content, time stamp of the scene video segment, and role object identifier (the role object is included in the scene video segment), etc. For example, the scene video segment is input into a scene understanding module configured with a video understanding algorithm, and the scene understanding module identifies the scene structure and key elements in the scene video segment and outputs the video description information of the scene video segment, where the video description information includes the time stamp of the scene video segment, role object identifier, dialogue content, and scene text description, and the video understanding algorithm can use technologies such as video-captioning, HumanOmni, etc.
[0049] Further, after obtaining the video description information of each scene video segment, the feature data of the character objects in each scene video segment can be obtained, where the feature data may include voice representation data and / or face representation data, etc. The face representation data may include one or more of face three-dimensional representation, face cropped video, and face expression. In a specific implementation, the feature data of the character object includes voice representation data and face representation data, and the face representation data includes face cropped video and face three-dimensional representation. Correspondingly, as Figure 4 shown, the process of obtaining the feature data of the character objects in each scene video segment includes: S121. According to the character object, the audio data corresponding to each scene video segment is segmented into several audio segments, and serial number tags are added to each audio segment; S122. Extract the speaking style of each audio segment to obtain the voice representation data of the character object corresponding to each audio segment; S123. Detect the face in each scene video segment to obtain the face cropped video corresponding to each character object; S124. Extract the face three-dimensional representation of the face cropped video to obtain the face three-dimensional representation of each character object; S125. Use the voice representation data, the face cropped video, and the face three-dimensional representation of each character object as the feature data of each character object.
[0050] Specifically, each scene video segment may include one character object (such as a monologue scene of a character object in the scene video segment), or may include multiple character objects (such as a dialogue scene of multiple character objects in the scene video segment). In order to obtain the voice representation data of each character object in the scene video segment, the audio data corresponding to the scene video segment can be extracted first, then the voice time interval in the audio data corresponding to the scene video segment is detected according to the character object, and finally the audio data corresponding to the scene video segment is segmented into several audio segments according to the detected voice time interval. Among them, the voice time interval can be detected for the audio data through an active voice detection algorithm, and the active voice detection algorithm can use, for example, silero_vad, etc.
[0051] Further, each audio clip includes all the audio data of a character object. For example, if there are 2 character objects in a scene video clip, then the scene video clip includes the speaking video clip A1 of character object A, the speaking video clip B1 of character object B, the speaking video clip A2 of character object A, and the speaking video clip B2 of character object B. Then, 4 human voice time intervals can be detected in the scene video clip, and the corresponding audio data are respectively the human voice time intervals of the speaking video clip A1, the human voice time interval of the speaking video clip B1, the human voice time interval of the speaking video clip A2, and the human voice time interval of the speaking video clip B2. Correspondingly, the audio data of the scene video clip will be split into 2 audio clips, which are respectively the audio clip corresponding to the human voice time interval of the speaking video clip A1 and the human voice time interval of the speaking video clip A2 of character object A; the audio clip corresponding to the human voice time interval of the speaking video clip B1 and the human voice time interval of the speaking video clip B2 of character object B. Of course, in practical applications, other methods can also be used to split the scene video clip into several audio clips. Taking the above scene video clip as an example, each audio clip includes a continuous segment of audio data of a character object. Then, the audio data corresponding to the scene video clip will be split into 4 audio clips, which are respectively the audio clip corresponding to the human voice time interval of the speaking video clip A1, the audio clip corresponding to the human voice time interval of the speaking video clip B1, the audio clip corresponding to the human voice time interval of the speaking video clip A2, and the audio clip corresponding to the human voice time interval of the speaking video clip B2.
[0052] After obtaining several audio clips by splitting, the speaking style of each audio clip is extracted to obtain the speech characterization data of each character object. The speech characterization data is used to reflect the emotional changes and / or language styles of digital actors, etc., and it can include intonation, speech rate, emotion, etc. Among them, the speech characterization data can be obtained by extracting the speaking style of the character object in each audio clip through a pre-trained speech style extraction network (such as speech encoders like wav2vec, hubert, etc.).
[0053] The facial cropped video is obtained by cropping the face region of a scene video segment, and each facial cropped video corresponds to a character object. Specifically, the facial images of the character object in the scene video segment can be detected through face recognition (for example, using face recognition algorithms such as dlib, YOLO, face-alignment, etc.), and the image sequence composed of all the facial images corresponding to the character object is used as the facial cropped video to obtain the facial cropped video corresponding to each character object. In addition, after the facial cropped video corresponding to each character object, the three-dimensional facial representation of each character object can be extracted through three-dimensional reconstruction technology to obtain the three-dimensional facial representation of each character object, where the three-dimensional facial representation can adopt implicit expressions such as FLAME representation, BFM representation, and warping-vae.
[0054] After obtaining the speech representation data, facial cropped video, and three-dimensional facial representation of each character object, the speech representation data, facial cropped video, and three-dimensional facial representation can be used as the feature data of a character object.
[0055] Further, in step S13, when constructing the video scene database according to the video description information of each scene video segment, the video description information of each scene video segment can be directly used as the video scene data of the scene video segment, and the video scene data of the scene video segment is associated with each scene video segment, and then the database composed of the video scene data of all scene video segments is used as the video scene database.
[0056] In one implementation, in order to enrich the video scene database, when constructing the video scene database according to the video description information of each scene video segment, first determine the video summary of the video segment based on the video description information, and then construct the video scene database according to the video summary and video description information. Based on this, the construction of the video scene database according to the video description information of each scene video segment specifically includes: Determine the video summary of each scene video segment based on the video description information of each scene video segment; Construct the video scene database according to the video description information and video summary of each scene video segment.
[0057] Specifically, the video summary is a summary description of the scene video segment, which is determined according to the video description information. That is, the video segment can be summarized according to the video description information to obtain the video summary, and the video summary is added to the video description information, so that the video description information also includes the video summary.
[0058] Exemplarily, determining the video summary of each scene video segment based on the video description information of each scene video segment specifically includes: Perform speech recognition on each audio segment corresponding to each scene video segment to obtain the dialogue text of each audio segment; Update the video description information of each scene video segment based on all the dialogue texts corresponding to each scene video segment, and summarize the updated video description information through a large language model to obtain the video summary of each scene video segment.
[0059] Specifically, the audio segments are segmented from the audio data of the scene video segment in step S12. That is, after the audio data corresponding to the scene video segment is segmented into several audio segments, the dialogue content in each audio segment can be converted into dialogue text, and then this dialogue text is used to update the video description information. Among them, when the video description information includes dialogue text, the dialogue text in the video description information is updated with this dialogue text to update the video description information. When the video description information does not contain dialogue text, this dialogue text is added to the video description information to update the video description information.
[0060] After obtaining the updated video description information, input the video description information into a large language model (such as ChatGpt of OpenAI, deepseek series models, etc.), and summarize the video description information of each scene video segment through the large language model to generate the video summary of each scene video segment, which can remove the redundant information in the scene understanding output result and facilitate quick browsing and subsequent analysis. Here, using the dialogue content of each audio segment to update the video description information can make the video description information only include the dialogue content of the role objects, filtering out other interfering dialogue contents, thereby improving the accuracy of the video summary.
[0061] Of course, in practical applications, it is also possible not to extract the dialogue content of each audio segment, but directly summarize the video description information through a large language model to generate the video summary of the scene video segment.
[0062] Exemplarily, constructing the video scene database according to the video description information of each scene video segment specifically includes: Construct the video scene data corresponding to each scene video segment with the segment identifier of each scene video segment as the key and the video description information and summary information of each scene video segment as the value; Form a video scene database based on the video scene data corresponding to all scene video segments.
[0063] Specifically, in the video scene database, the video description information is stored in the form of key-value, where the key is the segment identifier and the value is the video description information and the video summary, so that the video scene data corresponding to each scene video segment can be quickly determined.
[0064] Further, in step S14, after obtaining the feature data of each role object in each scene video segment, as Figure 4 shown, the database composed of the feature data of each role object in all scene video segments is used as the digital actor database to provide information support for user voice interaction and digital actor generation. Among them, constructing the digital actor database according to the feature data of the role object in each scene video segment specifically includes: Constructing digital actor data with each role object as the key and the feature data of each role object as the value; Forming a digital actor database based on the digital actor data corresponding to all role objects.
[0065] Specifically, in the digital actor database, the feature data is also stored in the form of key-value, where the key is the role object identifier and the value is the feature data, and the feature data stored in the form of key-value is also associated with the segment identifier, so that the feature data of the role object in each video segment can be quickly determined.
[0066] Of course, in practical applications, the video scene database and digital actor database corresponding to the interactive video can be stored together with the video scene databases and digital actor databases of other interactive videos, and the relationship between the video scene database and digital actor database corresponding to the interactive video and the interactive video is determined.
[0067] In the embodiment of the present application, by obtaining the video scene database and digital actor database of the interactive video, the interactive video is converted into a video in which the role objects can interact. The user can interact with the role objects based on the video scene database and digital actor database, and support the interaction between the user and the user role in a real-time video manner, improving the existing video expression ability.
[0068] S20. Generate a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database.
[0069] Specifically, the interactive voice can be the voice formed by a user interacting with an interactive video. The interactive voice can be picked up by a terminal device with a sound pickup function, or transmitted by other devices, or uploaded by the user, etc. For example, a display for playing an interactive video is configured with a sound pickup function (such as a microphone, etc.), and the user voice of the user for the interactive video is picked up through the sound pickup function to obtain the interactive voice. Of course, in practical applications, in order to accurately obtain the interactive voice of the user for the interactive video, a trigger button can be configured on the playback interface of the interactive video. When the trigger button is triggered, it is determined that the user starts to input the interactive voice for the interactive video, and the user voice is picked up to obtain the interactive voice. Or, the user voice is picked up in real time, and then the relevance between the user voice and the interactive video is detected. When the detected relevance meets the requirements (for example, includes the video name of the interactive video, the actor's name, the video lines, etc.), the picked-up user voice is used as the interactive voice, etc.
[0070] For example: The user plays an interactive video through a user terminal configured with a speaker and a microphone. During the process of watching the interactive video, the user utters an evaluation voice for a character object in the interactive video. Then the user terminal will pick up the evaluation voice through the microphone and use the evaluation voice as the interactive voice for the interactive video.
[0071] The digital actor voice is the response voice corresponding to the interactive voice. The voice style of the digital actor voice is formed by cloning the speaking style of the character object, and the voice content of the digital actor voice is the response content about the interactive video formed based on the video description information. The digital actor video stream is the response video frame sequence corresponding to the interactive voice. The digital actor in the digital actor video stream is obtained by cloning the character object in the interactive video, and the digital actor adopts the facial features of its corresponding character object. The acquisition processes of the digital actor voice and the digital actor video stream are described below respectively.
[0072] Exemplarily, generating the digital actor voice based on the interactive voice, the video scene database, and the digital actor database specifically includes: H10. Encode the interactive voice to obtain the voice representation of the interactive voice; H20. Select the video description information corresponding to the interactive voice in the video scene database; H30. Based on the video description information and the voice representation, obtain the target voice data for driving the digital actor constructed based on the target character object through a large language model; H40. Select the target feature data of the target character object in the digital actor database, and generate the digital actor voice according to the target feature data and the target voice data.
[0073] Specifically, in step H10, the target role object can be determined based on user selection. For example, after receiving the interactive voice, obtain the scene video segment corresponding to the interactive voice, then read the role object identifier included in the scene video segment in the video scene database, and feedback the obtained role object identifier to the user. Determine the target role object according to the user's selection operation. Of course, in practical applications, other methods can also be used to determine the target role object. For example, directly identify the target role object from the interactive voice, or use the role object displayed in the interactive video when receiving the interactive voice as the target role object, or first let the user select the target role object. If the target role object selected by the user is not obtained, then use the role object displayed in the interactive video when receiving the interactive voice as the target role object, etc.
[0074] The voice representation can be obtained by extracting the voice representation of the interactive voice. As Figure 5 shown, after detecting the interactive voice through active voice detection (such as through a voice detection algorithm, etc.), encode the interactive voice into audio through a voice representation encoder (such as an ASR algorithm) to obtain the voice representation of the interactive voice.
[0075] In step H20, the video description information is the video description information of the scene video segment that the interactive voice needs to interact with, which is selected from the video scene database. Among them, the acquisition process of the video description information can first obtain the reception time of the interactive voice, then determine the scene video segment that the interactive voice needs to interact with based on the reception time, and then read the corresponding video description information in the video scene database based on the segment identifier of the video segment; or, it can first identify the video description information in the interactive voice, then determine the segment identifier corresponding to the interactive voice based on the video description information, and then determine the scene video segment based on the segment identifier, etc.
[0076] In step H30, as Figure 5 shown, after obtaining the voice representation and the video description information, the voice representation and the video description information can be input into a large language model first. The large language model uses the video description information as context information to generate text information for driving the digital actor constructed based on the target role object for the voice representation, and convert the text information into voice data. For example, convert each text block into voice data through a text-to-speech algorithm (such as a tts algorithm).
[0077] In step H40, after obtaining the voice data, the speaking style of the voice data is cloned into the voice color of the target character object. Specifically, after obtaining the voice data, the target feature data of the target character object is selected from the digital actor database, and then the voice color of the voice data is cloned based on the target feature data. Among them, the voice color cloning process can be cloned through a voice cloning algorithm (such as voice-craft) or a large language model. For example, the target feature data and the voice example of the target character object are used as prior knowledge, and this prior knowledge and the voice data are input into the large language model together, and the large language model clones the voice color of the target character object to generate a digital actor voice with the voice color of the target character object.
[0078] In the embodiment of the present application, the video description information is used as the context information, and the digital actor voice corresponding to the interactive voice is generated according to the feature data, and the digital actor voice can match the context of the scene video segment, ensuring that the digital actor voice is close to the character of the character object. In addition, the target feature data of the target character object is extracted in advance. When determining the digital actor voice, only the reference audio (such as 5 to 10 seconds, etc.) of the target character object is required to generate the digital actor voice corresponding to the target character object, without pre-training or fine-tuning, reducing the generation cost of the digital actor voice.
[0079] The above completes the description of the generation process of the digital actor voice. Next, the generation process of the digital actor video stream will be described. The digital actor video stream is generated based on the digital actor voice and the digital actor database. The specific process of generating the digital actor video stream according to the digital actor voice and the digital actor database includes: Select the target feature data of the target character object from the digital actor database; Generate the facial movement features for driving the digital actor constructed based on the target character object based on the target feature data and the digital actor voice; Synthesize the digital actor video stream based on the facial movement features and the target feature data.
[0080] Specifically, the digital actor video stream is a video frame sequence corresponding to the digital actor voice. Among them, each video frame in the digital actor video stream contains a digital actor constructed based on the target character object, and the facial features of the digital actor are constructed based on the facial features of the target character object. That is to say, a digital actor will be created according to the target character object, and the digital actor will perform the digital actor voice with the facial expressions and emotions of the target character object. Based on this, such as Figure 6As shown, when obtaining the target feature data, facial motion features will first be generated from the digital actor's voice and feature data (such as generating facial motion features through a seq2seq model, etc.). Among them, the facial motion features are a set of feature vectors that record the movements of the head and facial expressions of the target character object. Then, based on the facial motion features, portrait video synthesis is performed to obtain a digital actor video stream with the target character object as the digital actor. Among them, portrait video synthesis can be achieved through a pre-set portrait video synthesis network (such as a U-Net network, or the Gaussian splash method, etc.).
[0081] It should be noted that when generating the digital actor video stream, the user can also control the digital actor in the digital actor video stream. Based on this, in one implementation, as Figure 6 shown, the synthesis of the digital actor video stream based on the facial motion features and the target feature data specifically includes: Receiving control information interacted by the user; Synthesizing a digital actor video stream based on the control information, the facial motion features, and the target feature data.
[0082] Specifically, the control information is interacted by the user and is used to control the digital actor constructed based on the target character object. For example, controlling the digital actor's emotion control, facial perspective, and real-time dialogue content, etc. Among them, the control information can be one or more of text information, audio information, image information, and video information. For example, the control information is a conversation content, or the control information is a conversation content and a person image, etc. After obtaining the control information, the digital actor video stream can be synthesized based on the control information and the facial motion features, and during the synthesis process of the digital actor video stream, the facial motion features are adjusted according to the control information to generate a digital actor video stream that conforms to the control information.
[0083] After obtaining the digital actor's voice and the digital actor video stream, a digital actor video can be generated based on the digital actor's voice and the digital actor video stream. Among them, the digital actor's voice is the audio data of the digital actor video, and the digital actor video stream is the image data of the digital actor video. This digital actor video is used for interaction with the user. Among them, the interaction method can be to synchronously play the digital actor video and the interaction video for the user, or to pause the interaction video to play the digital actor video, or to play the digital actor video on the user-specified device, etc.
[0084] In one implementation, the interaction method is to configure a playback area for the digital actor video, and then play the digital actor video in this playback area. Correspondingly, after generating the digital actor video based on the digital actor's voice and the digital actor video stream, the method further includes: Configure a playback area, an image selection area, and an image interaction area for the digital actor video. The playback area is used to play the digital actor video, the image selection area is used to determine the digital actor constructed by the target character object, and the interaction area is used for interaction between the user and the digital actor.
[0085] Specifically, the playback area corresponding to the digital actor video and the playback area corresponding to the interaction video are located within the same display interface. The user can control the interaction video through the playback area corresponding to the interaction video. For example, play, pause, or adjust the video playback progress, etc. Among them, as Figure 7 shown, the display interface is deployed with an image selection area and an image interaction area. The image selection area is used to select the area to determine the character object, and the image interaction area is used to input control information. That is to say, the target character object can be determined through the image selection area, and the interaction control information can be interacted through the image interaction area.
[0086] In addition, the display interface may further include a chat history area, which displays the interaction records of each user for the interaction video, such as chat records, digital actor video playback records, etc. The user can review the past conversation history through the chat history area, view the historical digital actor videos, and quickly jump to a certain historical conversation or historical digital actor video, etc.
[0087] In summary, this embodiment provides a method for generating an interactive digital actor video. The method includes processing the interaction video to form a video scene database and a digital actor database; generating digital actor voices based on the interaction voice for the interaction video, the video scene database, and the digital actor database, and generating a digital actor video stream according to the digital actor voices and the digital actor database; generating a digital actor video based on the digital actor voices and the digital actor video stream. The embodiments of the present application process the interaction video into a video scene database and a digital actor database, and use the video scene database and the digital actor database to provide information support for generating digital actors, which can improve the freedom and interaction effect of video interaction.
[0088] Based on the above method for generating an interactive digital actor video, this embodiment provides a system for generating an interactive digital actor video, as Figure 8 shown. The system for generating an interactive digital actor video specifically includes: A receiving module 100, configured to construct a video scene database and a digital actor database according to the interaction video, where the video scene database includes video description information of the interaction video; the digital actor database includes feature data of the character object in the interaction video; A generation module 200 is configured to generate a digital actor video based on the interactive voice for the interactive video, the video scene database, and the digital actor database.
[0089] Based on the above method for generating an interactive digital actor video, this embodiment provides a computer-readable storage medium storing one or more programs, which can be executed by one or more processors to implement the steps in the method for generating an interactive digital actor video as described in the above embodiment.
[0090] Based on the above method for generating an interactive digital actor video, the present application further provides a terminal device, as Figure 9 shown, which includes at least one processor 20; a display screen 21; and a memory 22, and may further include a communication interface 23 and a bus 24. Among them, the processor 20, the display screen 21, the memory 22, and the communication interface 23 can complete mutual communication through the bus 24. The display screen 21 is set to display a user guidance interface preset in the initial setting mode. The communication interface 23 can transmit information. The processor 20 can call the logical instructions in the memory 22 to execute the method in the above embodiment.
[0091] In addition, when the logical instructions in the above memory 22 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium.
[0092] The memory 22, as a computer-readable storage medium, can be set to store software programs and computer-executable programs, such as the program instructions or modules corresponding to the method in the embodiment of the present disclosure. The processor 20 executes functional applications and data processing by running the software programs, instructions, or modules stored in the memory 22, that is, implements the method in the above embodiment.
[0093] The memory 22 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the terminal device, etc. In addition, the memory 22 may include a high-speed random access memory and may also include a non-volatile memory. For example, various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes can also be a transient storage medium.
[0094] In addition, the specific processes of loading and executing multiple instructions by the instruction processors in the above storage medium and terminal device have been described in detail in the above method, and will not be elaborated here one by one.
[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for generating an interactive digital actor video, characterized in that: The method for generating the interactive digital actor video specifically includes: Constructing a video scene database and a digital actor database according to the interactive video, wherein the video scene database includes video description information of the interactive video; the digital actor database includes feature data of role objects in the interactive video; A digital actor video is generated based on the interactive speech for the interactive video, the video scene database and the digital actor database.
2. The method for generating an interactive digital actor video according to claim 1, characterized in that: The step of constructing a video scene database and a digital actor database based on the interactive video specifically includes: Segmenting the interactive video into a plurality of scene video segments, and adding a segment identifier to each scene video segment; Obtaining video description information of each scene segment and feature data of a role object in each scene video segment; Building a video scene database according to the video description information of each scene video clip; A digital actor database is constructed according to the feature data of the role objects in each scene video clip.
3. The method for generating an interactive digital actor video according to claim 2, characterized in that: Get the feature data of the character object in each scene video clip, including: The audio data corresponding to each scene video clip is divided into several audio clips based on the character object, and a serial number label is added to each audio clip; Extracting the speech style of each audio clip to obtain the speech representation data of the character object corresponding to each audio clip; Perform face detection on each scene video clip to obtain the facial cropped video corresponding to each character object; Extracting a three-dimensional facial representation from the facial cropped video to obtain a three-dimensional facial representation of each character object; The voice representation data, the facial cropping video and the three-dimensional representation of the face of each character object are used as feature data of each character object.
4. The method for generating an interactive digital actor video according to claim 2, characterized in that: The step of constructing a video scene database according to the video description information of each scene video clip specifically includes: Determine a video summary of each scene video clip based on the video description information of each scene video clip; A video scene database is constructed according to the video description information and the video summary of each scene video clip.
5. The method for generating an interactive digital actor video according to claim 4, characterized in that: The step of determining the video summary of each scene video clip based on the video description information of each scene video clip specifically includes: Performing speech recognition on each audio segment corresponding to each scene video segment to obtain the dialogue text of each audio segment; The video description information of each scene video clip is updated based on all dialogue texts corresponding to each scene video clip, and the updated video description information is summarized through a large language model to obtain a video summary of each scene video clip.
6. The method for generating an interactive digital actor video according to claim 4, characterized in that: The step of constructing a video scene database according to the video description information and the video summary of each scene video clip specifically includes: Using the segment identifier of each scene video segment as the key and the video description information and summary information of each scene video segment as the value, construct the video scene data corresponding to each scene video segment; A video scene database is formed based on the video scene data corresponding to all scene video clips.
7. The method for generating an interactive digital actor video according to claim 2, characterized in that: The step of constructing a digital actor database according to the feature data of the role objects in each scene video clip specifically includes: Using each character object as the key and the characteristic data of each character object as the value, construct digital actor data; A digital actor database is formed based on the digital actor data corresponding to all role objects.
8. The method for generating an interactive digital actor video according to any one of claims 1 to 7, characterized in that: The generating of the digital actor video based on the interactive voice for the interactive video, the video scene database and the digital actor database specifically includes: Generate digital actor voice based on the interactive voice, the video scene database and the digital actor database; Generate a digital actor video stream according to the digital actor voice and the digital actor database; A digital actor video is generated based on the digital actor voice and the digital actor video stream.
9. The method for generating an interactive digital actor video according to claim 8, characterized in that: The generating of the digital actor voice based on the interactive voice, the video scene database and the digital actor database specifically includes: Encoding the interactive speech to obtain a speech representation of the interactive speech; Selecting video description information corresponding to the interactive voice in the video scene database; Based on the video description information and the speech representation, obtaining target speech data for driving a digital actor constructed based on a target role object through a large language model; Selecting target feature data of a target role object in the digital actor database; The digital actor voice is generated according to the target feature data and the target voice data.
10. The method for generating an interactive digital actor video according to claim 8, characterized in that: The step of generating a digital actor video stream according to the digital actor voice and the digital actor database specifically includes: Selecting target feature data of a target role object in the digital actor database; Generate facial movement features for driving a digital actor constructed based on a target character object based on the target feature data and the digital actor voice; A digital actor video stream is synthesized based on the facial motion features and the target feature data.
11. The method for generating an interactive digital actor video according to claim 10, characterized in that: The synthesizing of the digital actor video stream based on the facial motion features and the target feature data specifically includes: Receive control information from user interaction; A digital actor video stream is synthesized based on the control information, the facial motion features and the target feature data.
12. The method for generating an interactive digital actor video according to claim 8, characterized in that: After generating the digital actor video based on the digital actor voice and the digital actor video stream, the method further includes: A playback area, an image selection area and an image interaction area are configured for the digital actor video, wherein the playback area is used to play the digital actor video, the image selection area is used to determine the digital actor constructed by the target role object, and the interaction area is used for the interaction between the user and the digital actor.
13. A system for generating interactive digital actor videos, characterized in that: The interactive digital actor video generation system specifically includes: A construction module, used to construct a video scene database and a digital actor database according to the interactive video, wherein the video scene database includes video description information of the interactive video; the digital actor database includes feature data of role objects in the interactive video; A generation module is used to generate a digital actor video based on the interactive voice for the interactive video, the video scene database and the digital actor database.
14. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps in the method for generating an interactive digital actor video as described in any one of claims 1-12.
15. A terminal device, characterized in that: include: Processor and memory; The memory stores a computer-readable program executable by the processor; When the processor executes the computer-readable program, the processor implements the steps in the method for generating an interactive digital actor video as described in any one of claims 1-12.
Citation Information
Patent Citations
Visual interaction system based on holographic projection
CN113821104A
Method and system for implanting video into 3D virtual digital human
CN115002509A
Visual guiding method and system of scene video superimposed digital human, and storage medium
CN116740311A
Simulation role-based information interaction method and device and storage medium
CN117560340A