Method and device for generating plot video, computer device and storage medium
Patent Information
- Application Number
- CN202311248120.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-09-25
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2043-09-25
AI Technical Summary
[0002]在一些场景中,需要通过剧情视频让用户获知当前故事的走向,比如,故事驱动型游戏,可以通过对话的方式呈现剧情以及向玩家提供分支选项,不同的选项对应的故事走向不同,故事结局不同,用于呈现剧情或者分支选项的视频记为剧情视频,对于剧情视频的制作通常需要导演手动编辑剧情,比如,安放演员、设置拍摄机位等操作,生成剧情视频的过程较为复杂,导致视频的生成效果低,且容易出现人工操作引起的错误,导致视频的生成质量较差
[0018]本申请实施例通过获取用于生成剧情视频的视频生成资源,视频生成资源包括至少两个出场角色、目标站位模板和至少两个出场角色的对话相关内容,目标站位模板中每个角色站位配置有拍摄机位,对话相关内容中包含对白和每个对白的发言方和参与方;根据目标站位模板和至少两个出场角色,确定每个出场角色对应的角色站位;针对对话相关内容中的每个对白,根据作为发言方的出场角色的角色站位所配置的拍摄机位中,选择所述对白的拍摄机位;基于对话相关内容中的对白控制至少两个出场角色进行对话;基于每个对白对应的拍摄机位上的虚拟摄像机对至少两个出场角色进行拍摄,得到剧情视频。
Smart Images

Figure CN117354437B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of communication technology, specifically to a method, apparatus, computer device, and computer-readable storage medium for generating narrative videos, wherein the storage medium is a computer-readable storage medium. Background Technology
[0002] In some scenarios, it is necessary to use narrative videos to inform users of the current story's direction. For example, in story-driven games, the plot can be presented through dialogue, and branching options can be provided to players. Different options lead to different story directions and endings. Videos used to present the plot or branching options are called narrative videos. The production of narrative videos usually requires the director to manually edit the plot, such as placing actors and setting up camera positions. The process of generating narrative videos is relatively complex, resulting in low video quality and the possibility of errors caused by manual operation, leading to poor video quality. Summary of the Invention
[0003] This application provides a method, apparatus, computer device, and storage medium for generating narrative videos, which can improve the efficiency and quality of narrative video generation.
[0004] This application provides a method for generating a narrative video, including:
[0005] Obtain video generation resources for generating story videos. The video generation resources include at least two characters, a target position template, and dialogue-related content for the at least two characters. Each character's position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participants in each dialogue.
[0006] Based on the target position template and the at least two characters appearing in the game, determine the character position corresponding to each character appearing in the game;
[0007] For each line of dialogue in the aforementioned dialogue content, select the camera position for that line of dialogue from the camera positions configured according to the position of the character appearing as the speaker;
[0008] Based on the dialogue in the relevant content of the dialogue, control the at least two characters to engage in dialogue;
[0009] The at least two characters appearing in the scene are filmed using a virtual camera at the camera position corresponding to each dialogue, resulting in a story video.
[0010] Accordingly, this application also provides a narrative video generation apparatus, comprising:
[0011] The acquisition unit is used to acquire video generation resources for generating a story video. The video generation resources include at least two characters, a target position template, and dialogue-related content of the at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participants in each dialogue.
[0012] The positioning determination unit is used to determine the position of each character based on the target positioning template and the at least two characters appearing in the game.
[0013] The camera position determination unit is used to select the camera position for each dialogue from the camera positions configured for each dialogue in the dialogue-related content.
[0014] A control unit is used to control the at least two characters to engage in dialogue based on the dialogue-related content.
[0015] The shooting unit is used to shoot the at least two characters based on the virtual camera at the shooting position corresponding to each dialogue, so as to obtain the plot video.
[0016] Accordingly, this application also provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute any of the narrative video generation methods provided in this application.
[0017] Accordingly, embodiments of this application also provide a computer-readable storage medium for storing a computer program, which is loaded by a processor to execute any of the narrative video generation methods provided in embodiments of this application.
[0018] This application embodiment obtains video generation resources for generating a story video. The video generation resources include at least two characters, a target position template, and dialogue-related content for at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participant for each dialogue. Based on the target position template and at least two characters, the character position corresponding to each character is determined. For each dialogue in the dialogue-related content, a camera position for the dialogue is selected from the camera positions configured for the character position of the character who is the speaker. The dialogue-related content is used to control at least two characters to engage in dialogue. The virtual camera on the camera position corresponding to each dialogue is used to film the at least two characters to obtain the story video.
[0019] In this embodiment, each character's position in the positioning template is pre-set with a camera position. Based on the target positioning template, the position of each character and the corresponding camera position can be determined. Then, based on the dialogue content, the characters are controlled to engage in dialogue and be filmed. This can automatically generate a story video, improve the efficiency of generating story videos, avoid errors caused by manually editing the story, and improve the quality of the generated story videos. Attached Figure Description
[0020] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1 This is a flowchart of the narrative video generation method provided in the embodiments of this application;
[0022] Figure 2 This is a schematic diagram of the Type I station template provided in the embodiments of this application;
[0023] Figure 3 This is a schematic diagram of the L-shaped station template provided in the embodiments of this application;
[0024] Figure 4 This is a schematic diagram of the Type A station template provided in the embodiments of this application;
[0025] Figure 5 This is a schematic diagram of the camera position for the Type I station template provided in this application embodiment;
[0026] Figure 6 This is a schematic diagram of the camera position for the L-shaped station template provided in this application embodiment;
[0027] Figure 7 This is a schematic diagram of the camera position for the Type A shooting position template provided in the embodiments of this application;
[0028] Figure 8 This is a schematic diagram illustrating the relationship between audio and motion playback time provided in an embodiment of this application;
[0029] Figure 9 This is a schematic diagram of audio and motion alignment provided in an embodiment of this application;
[0030] Figure 10 This is a schematic diagram of audio playback time provided in the embodiments of this application;
[0031] Figure 11 This is a schematic diagram illustrating the audio and motion durations and playback time provided in the embodiments of this application;
[0032] Figure 12 This is another audio playback time diagram provided in an embodiment of this application;
[0033] Figure 13 This is a schematic diagram of the narrative video generation device provided in an embodiment of this application;
[0034] Figure 14 This is a schematic diagram of the structure of the computer device provided in the embodiments of this application. Detailed Implementation
[0035] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0036] This application provides a method, apparatus, computer device, and computer-readable storage medium for generating narrative videos. The narrative video generation apparatus can be integrated into a computer device, which may be a server or a terminal, etc.
[0037] The terminal may include mobile phones, wearable smart devices, tablets, laptops, personal computers (PCs), and in-vehicle computers, etc.
[0038] The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.
[0039] The following sections provide detailed descriptions of each example. It should be noted that the order in which the embodiments are described is not intended to limit the preferred order of the embodiments.
[0040] This embodiment will be described from the perspective of a narrative video generation device, which can be integrated into a computer device, such as a server or a terminal.
[0041] This application provides a method for generating narrative videos, such as... Figure 1 As shown, the specific process of this narrative video generation method can be summarized as follows:
[0042] 101. Obtain video generation resources for generating story videos. The video generation resources include at least two characters, target position templates, and dialogue-related content for at least two characters. Each character's position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participants in each dialogue.
[0043] The video generation resources can include data needed to generate story videos, such as the characters, their positions, and their dialogue.
[0044] This application embodiment can automatically generate plot videos based on the characters appearing, their positions, and dialogue content, without requiring users to manually set the characters' positions, camera angles, and subtitles, thus improving the efficiency of plot video generation, avoiding errors caused by manual plot editing, and improving the quality of the generated plot videos.
[0045] The narrative video generation method provided in this application can be applied to the fields of games and animation. In the game field, players can be given options, and different options lead to different story directions and different story endings. The plot and branching options can usually be presented through dialogue. The narrative video generation method provided in this application can generate narrative videos in games to present the plot and branching options, which can also be called cutscene videos.
[0046] The characters that appear can include those related to the plot, such as characters, heroes, mythical beasts, and animals in the game. The characters can be 3D models or 2D images. Optionally, the characters can perform corresponding actions under control.
[0047] The dialogue content can include the dialogue of each character and the corresponding participant for each dialogue. It can be assumed that the character speaking the dialogue speaks to the character participating in the dialogue.
[0048] The target position template can be a preset position template. The position template includes at least two character positions. Each character position can accommodate at least one character and each character position is equipped with a camera position.
[0049] For example, it can obtain the character and dialogue content input by the user. Optionally, the character can also be determined based on the user's selection. For example, it can provide the user with selectable characters and obtain at least two selected characters based on the user's selection.
[0050] The target positioning template can be selected based on the number of characters appearing, or alternative positioning templates can be provided to the user, and the target positioning template can be determined based on the user's selection.
[0051] In one embodiment, each character's position can be configured with at least one set of camera positions, and each set of camera positions includes at least two camera positions, a first camera position and a second camera position. There can be at least one first camera position and at least one second camera position. The first camera position is used to film the first character appearing at the character's position. The filmed characters do not include the second characters who are talking to the first character. The second camera position is used to film the character appearing at the character's position and other characters who are talking to the character.
[0052] Optionally, in each group of camera positions configured for each character's position, the second camera position may film a different character's second appearance.
[0053] For example, it can be like Figure 2 The positioning template shown includes the positioning of two characters, such as... Figure 3 and Figure 4 The positioning template shown includes three character positions, among which, Figure 3 In the shown positioning template, there is no connection between positions A and C, indicating that the characters in positions B and C appear simultaneously. This can be considered as... Figure 3 In the station template shown, stations B and C can be considered as one large station. Figure 4 The characters shown in positions B and C can appear independently of each other.
[0054] Each character's position is equipped with at least one camera position. For example, a camera position can be configured to film a solo shot of the character appearing in that position, or a camera position can be configured to film the character appearing in that position and the participants in the dialogue.
[0055] like Figure 5 The positioning template shown can be called Type I positioning. For character positioning A, the camera positions are positions 2 and 4. Position 4 can be used to film solo shots of the character appearing in position A, while position 2 can be used to film a two-person shot of the character appearing in position A and the character appearing in position B (the character interacting with the character appearing in position A). The characters appearing in positions A and B can be face-to-face. In this two-person shot, the back of the character appearing in position B is in the foreground, and the character appearing in position A is in the front. Similarly, for character positioning B, the camera positions are positions 1 and 3. Position 3 can be used to film solo shots of the character appearing in position B, while position 1 can be used to film a two-person shot of the character appearing in position A and the character appearing in position B. In this two-person shot, the back of the character appearing in position A is in the foreground.
[0056] like Figure 6 The positioning template shown can be called an L-shaped positioning. The camera positions corresponding to character position A are camera position 1 and camera position 3. Camera position 3 can be used to shoot a solo shot of the character appearing in character position A. Camera position 1 can be used to shoot a three-person shot of the character appearing in character position A and the characters appearing in character positions B and C. In the three-person shot, the back view of the character appearing in character position B is the foreground. For characters in positions B and C, camera position 2 can be used to shoot a three-person shot of the character in position A and the characters in positions B and C, with the character in position A's back in the foreground. Camera position 4 can be used to shoot a two-person shot of the characters in positions B and C. Camera position 5 shoots similar content to camera position 4, but from a greater distance than camera position 4, it shoots a two-person shot of the characters in positions B and C. It can be understood that cameras 4 and 5 can be considered as two secondary cameras for character position B, and their content does not include the character in position A, but can include other characters who are speaking or participating in dialogue along with the character in position B. Camera 6 is a symmetrical shot to camera 5, filming two characters standing in positions B and C. In the video shot by camera 5, the character standing in position B is positioned forward compared to the character standing in position C. In the video shot by camera 6, the character standing in position C is positioned forward compared to the character standing in position B. Therefore, the shot shot by camera 6 does not include the character standing in position A, but may include other characters who are speaking or participating in dialogue along with the character standing in position C.
[0057] Figure 6 The number of character positions included in the shown positioning template can be expanded. For example, character positions can be added on the positioning lines where character positions B and C are located, and then camera positions 1, 4, 5, and 6 can be adjusted so that these camera positions can capture all the characters appearing on the positioning lines where character positions B and C are located.
[0058] like Figure 7 The positioning template shown can be called Type A positioning. Type A positioning includes three positioning lines: positioning line O for characters A and B, positioning line P for characters A and C, and positioning line Q for characters B and C. Position line O corresponds to camera positions 1, 2, 3, and 4; positioning line P corresponds to camera positions 5, 6, 7, and 8; and positioning line Q corresponds to camera positions 9, 10, 11, and 12. Each of these positioning lines can be considered a positioning line in Type I positioning. The content captured by the camera position corresponding to each positioning line can be referenced. Figure 5 The camera positions shown are not described in detail here.
[0059] 102. Based on the target position template and at least two characters to appear, determine the character position corresponding to each character to appear.
[0060] For example, to determine the characters that appear in the target position template, the corresponding positions can be determined based on the independence of the characters' appearances in the dialogue content.
[0061] Assuming there are two characters appearing, the target positioning template is... Figure 2 The positioning template shown allows you to randomly select one of the two characters to appear, placing them in character position A, and the other character to appear, placing them in character position B.
[0062] If there are three characters appearing in the game, we can determine whether they are independent of each other (they can appear independently). If they are independent, then the target positioning template can be determined as follows. Figure 4 The target formation template is determined, and the three characters are randomly assigned to their respective positions; otherwise, the target formation template is determined as follows: Figure 3 The positioning template is shown. Then, the character that appears independently is identified and the corresponding position of the character is determined as position A. The positions of the other two characters are position B and position C.
[0063] Each character's position can be configured with at least one set of camera positions. Each set of camera positions can include at least two camera positions. Some camera positions film the character appearing at that position, but do not film characters appearing as participants in the dialogue. Other camera positions are used to film both parties in the dialogue. Appropriate camera positions can be selected according to the dialogue. In one embodiment, the dialogue includes dialogue text. The step "for each dialogue in the dialogue-related content, select the camera position for the dialogue text from the camera positions configured for the character's position as the speaker" can specifically include:
[0064] For each dialogue text in the aforementioned dialogue content, determine at least one set of camera positions configured for the character standing position of the speaking party;
[0065] For each dialogue text, based on the participants in the dialogue text, from the at least one set of camera positions, a second camera position is selected whose shooting content includes the participants, thus obtaining a first candidate camera position and a second candidate camera position.
[0066] For each dialogue text, if the number of texts exceeds a preset threshold, the first candidate shooting position is selected as the shooting position for the dialogue text.
[0067] For each dialogue text, if the number of texts does not exceed a preset threshold, then the second candidate camera position is selected as the camera position for the dialogue text.
[0068] For example, for each dialogue text in the dialogue content, the first character appearing as the speaker of that dialogue text is identified, along with the character's position and at least one set of corresponding camera positions. Since the second character filmed by the second camera position in each set of camera positions is different, a set of camera positions can be selected from at least one set of camera positions to film the first character and the character appearing as a participant in the dialogue text, based on the participants in the dialogue text. This set of camera positions includes at least two camera positions. If the amount of dialogue text exceeds a threshold, indicating that the first character's speaking time is relatively long, a first candidate camera position is selected from this set of camera positions to film the first character but not the second character. If the amount of dialogue text does not exceed a threshold, indicating that the first character's speaking time is relatively short, a second candidate camera position is selected from this set of camera positions to film both the first and second characters. If there are multiple first or second candidate camera positions, one can be randomly selected as the camera position for the dialogue text.
[0069] 103. For each line of dialogue in the relevant content, select the camera position for the dialogue from the camera positions configured according to the position of the character who is speaking.
[0070] In the story video, when dialogue appears, the speaker of that dialogue will usually appear in the story video. Therefore, the camera position corresponding to each dialogue can be determined based on the character's position where each speaker of the dialogue is located and the camera position configured for that character's position.
[0071] If each character's position is assigned a camera position, then that camera position will be designated as the camera position corresponding to the dialogue.
[0072] If each character's position is configured with at least two camera positions, then one of the at least two camera positions can be selected as the camera position for the dialogue. For example, it can be selected randomly, or the camera position can be determined based on the participants and content of the dialogue. That is, in one embodiment, the dialogue includes dialogue text, and each position corresponds to at least two preset candidate camera positions. The step "for each dialogue in the dialogue-related content, select the camera position for the dialogue based on the camera positions configured for the position of the character who is speaking" can specifically include:
[0073] For each line of dialogue, determine at least two candidate camera positions based on the position of the character who is speaking.
[0074] Based on the number of text units contained in each dialogue and the participants, select a camera position from at least two candidate camera positions.
[0075] For example, for each line of dialogue, determine the position of the character who speaks that line, and based on two pre-set candidate camera positions for each line of dialogue, determine at least two candidate camera positions for each line of dialogue.
[0076] The target positioning line can be determined based on the orange color of the speaker in each dialogue and the roles of the participants. From the corresponding camera positions in the target positioning line, the camera position corresponding to the role of the speaker can be obtained, and then one of them can be selected as the camera position for that dialogue. Optionally, the selection can be based on the number of text units contained in the dialogue. For example, if the number of text units exceeds a preset threshold, the camera position for shooting a single shot of the participating role can be selected; otherwise, other camera positions can be selected.
[0077] 104. Based on the dialogue content, control at least two characters to have a conversation.
[0078] Since the dialogue content includes the dialogue of each character, it is possible to control at least two characters to stand according to the target standing template and corresponding character standing positions, and to conduct dialogue. For example, based on the dialogue content, it is possible to control the character speaking to make corresponding actions during the speaking process, such as lip movements changing as they speak, or changes in body movements. In one embodiment, the dialogue includes audio, and the step "controlling at least two characters to conduct dialogue based on the dialogue in the dialogue content" can specifically include:
[0079] Feature extraction is performed on each pair of dialogue audio to obtain the audio feature information of each pair of dialogue audio;
[0080] For each dialogue audio, based on the audio feature information of the dialogue audio, the action data of the speaking character during the speaking process is generated.
[0081] For each dialogue audio, based on the motion data, the speaking character among the at least two characters is controlled to perform corresponding actions during the dialogue.
[0082] The audio information can include audio duration, speech feature information, etc., and the speech feature information can be represented by frequency domain features, time domain features, etc.
[0083] The motion data can include limb motion data and facial motion data. Facial motion data can include the motion data of the facial features, which is used to control the corresponding actions of the characters.
[0084] For example, a mapping relationship between audio feature information of dialogue audio and motion data can be preset. Features are extracted for each dialogue audio to obtain the audio information of each dialogue audio. Based on the preset mapping relationship, motion data that has a mapping relationship with the audio information of each dialogue audio can be determined.
[0085] In one embodiment, a pre-defined correspondence between body movement data and audio duration can be established so that the corresponding body movement data can be determined based on the audio duration of the dialogue.
[0086] Optionally, the body movements can be divided into multiple movement segments, where the duration of one movement segment is configured to be adjustable. By adjusting the duration of the movement segments, the body movements can be adapted to dialogues of different audio lengths. In one embodiment, the audio information includes audio duration, and the movement data includes body movement data. The step "for each dialogue, determine the movement data of the character speaking during the dialogue based on the audio information of the dialogue" can specifically include:
[0087] Acquire preset segmented action data and independent action data, as well as the audio duration of the target dialogue audio for each character. The segmented action data includes the starting action segment, the looping action segment, and the ending action segment. The duration of the looping action segment is adjustable.
[0088] For each character appearing in the story, based on the total audio duration of at least one target dialogue audio of that character and the relationship between the shortest duration of the segmented body movements, body movement data corresponding to the at least one target dialogue is selected from the segmented movement data and the independent movement data. The target dialogue audio refers to the audio of dialogue spoken by each character as the speaker.
[0089] The segmented action data can include data used to control the characters entering the scene as corresponding to the segmented actions. According to the action sequence, the segmented action can be divided into three action segments: the starting action segment, the looping action segment, and the ending action segment. The looping action segment is configured to be loopable and collapsible, which means that the action duration of the looping action segment is adjustable. When the looping segment is executed only once, the segmented action has the shortest action duration.
[0090] For example, a character's initial movement can be considered a complete preset body movement: lowering both hands, making a shrugging gesture, and then returning to lowering both hands. The initial movement segment includes the process from lowering both hands to making the shrugging gesture. The looping movement segment can include the action of maintaining the shrugging state. The ending movement segment includes the process of returning from the shrugging state to lowering both hands. The duration of the looping movement segment that maintains the shrugging state is adjustable. By adjusting the duration of the looping movement segment, the number of times the character makes the shrugging gesture can be controlled, thereby controlling the duration for which the character maintains the shrugging state. The shrugging gesture can be a single-handed shrugging gesture or a double-handed shrugging gesture, etc.
[0091] Segmented motion data can also include the process of going from having hands hanging down to having hands crossed over the chest, maintaining the crossed-over position, and returning from the crossed-over position to having hands hanging down. Segmented motion data can also represent other actions so that the character can perform the corresponding actions when speaking.
[0092] For example, dialogue can be audio. Audio features are extracted from the dialogue audio, the audio duration of each dialogue audio is extracted, segmented motion data is obtained, and the motion duration of the looping motion segments in the segmented limb motion data is adjusted according to the audio duration so that the motion duration of the adjusted segmented data is equal to or greater than the audio duration.
[0093] One, two, or more target dialogue audio clips can be used as processing units. Based on the audio duration of each target dialogue audio clip, the total audio duration of at least one target dialogue audio clip as a processing unit is obtained. If the total duration of at least one target dialogue audio clip is less than the minimum duration of a segmented action, it means that the character cannot complete the full segmented action, and an independent action can be used as the limb movement data corresponding to that at least one target dialogue audio clip. If the total duration of at least one target dialogue audio clip is greater than the minimum duration of a segmented action, it means that the character can complete the full segmented action, and the segmented action can be used as the limb movement data corresponding to that at least one target dialogue audio clip. It is understood that multiple independent actions can be set, and the action durations of multiple independent actions are different and shorter than the action duration of a segmented action.
[0094] To avoid discontinuous actions due to short speaking time, where characters fail to perform complete segmented actions or even complete action fragments, at least two target dialogue audios corresponding to a single character can be processed. Corresponding body movements can be designed for these two target dialogue audios. For two or more target dialogue audios, to ensure continuity, the first target dialogue audio needs to accommodate both the initial action fragment and the looping action fragment. In one embodiment, at least one target dialogue audio includes a first dialogue audio and a second dialogue audio. The step "For each character, based on the total audio duration of at least one target dialogue audio for that character and the relationship between the shortest action duration of the segmented action data and the independent action data, select the body movement data corresponding to the at least one target dialogue from the segmented action data and the independent action data" can include:
[0095] For each character appearing in the scene, select the first dialogue audio from the target dialogue audio corresponding to the character appearing in the scene, and determine the audio duration of the first dialogue audio.
[0096] If the audio duration is less than the shortest total duration of the starting action segment and the looping action segment, then the independent action data will be used as the limb action data corresponding to the first dialogue audio.
[0097] If the audio duration is not less than the shortest total action duration, then the action data corresponding to the starting action segment and the looping action segment are used as the limb action data corresponding to the first dialogue audio.
[0098] If the first dialogue audio corresponds to independent action data, then the second dialogue audio in the target dialogue audio that is sequentially next to the first dialogue audio is taken as the first dialogue audio, and the execution is returned. If the duration of the audio is less than the shortest total duration of the starting action segment and the loop action segment, then the independent action data is taken as the limb action data corresponding to the first dialogue audio.
[0099] If the first dialogue audio corresponds to segmented action data, then the ending action segment is used as the limb action data corresponding to the second dialogue audio, the third dialogue audio in the target dialogue audio that is sequentially next to the second dialogue audio is used as the first dialogue audio, and the process returns to execute the step where, if the audio duration is less than the shortest total duration of the starting action segment and the looping action segment, the independent action data is used as the limb action data corresponding to the first dialogue audio.
[0100] The first dialogue audio is the dialogue audio that comes first in the time sequence of the target dialogue audio corresponding to each character. The second dialogue audio is the audio that comes second in the time sequence of the first dialogue audio. The third dialogue audio is the audio that comes second in the time sequence of the second dialogue audio.
[0101] The shortest duration of the initial action segment and the loop action segment can be the duration of the initial action segment and the shortest duration of the loop action segment. The shortest duration of the loop action segment means that the corresponding action is executed only once.
[0102] For example, for each character, if the audio duration of the first dialogue audio is less than the shortest action duration of the starting action segment and the looping action segment in the segmented action data, then it is considered that the character cannot be used as a complete starting action segment and looping action segment in the first dialogue audio, and the independent action data can be used as the limb action data of the first dialogue audio.
[0103] Optionally, multiple independent actions can be set, and each action can be selected with an action duration that matches the audio duration of the first dialogue audio.
[0104] If the audio duration of the first dialogue audio is not less than the shortest action duration, then the action data corresponding to the starting action segment and the loop action segment in the segmented action data are used as the limb action data corresponding to the first dialogue audio, and the action segment duration of the loop action segment is adjusted so that the adjusted action duration matches the audio duration of the first dialogue audio.
[0105] If the body movement data corresponding to the first dialogue audio is independent movement data, it means that the character has already completed the full action when the first dialogue audio is finished. Therefore, the second dialogue audio can be used as the first dialogue audio. Based on the same processing of the first dialogue audio, the body movement data corresponding to the second dialogue audio can be determined.
[0106] If the body movement data of the first dialogue audio is the starting action segment and the looping action segment in segmented action data, it means that the character has not yet completed the complete action when finishing the first dialogue audio. In this case, the ending action segment can be used as the body movement data for the second dialogue audio, allowing the character to perform the complete action. Optionally, the actions that the second dialogue audio can include can be determined based on the action duration of the second dialogue audio, the ending action segment, and the action duration of the independent action. For example, if the audio duration of the second dialogue audio is not less than the total action duration of the ending action segment and the independent action, the ending action segment and the independent action can be used together as the body movement data of the second dialogue audio. If the audio duration of the second dialogue audio is less than the total action duration of the ending action segment and the independent action, only the ending action segment can be used as the body movement data of the second dialogue audio.
[0107] Then, the third dialogue audio is used as the first dialogue audio. Based on the same processing of the first dialogue audio, the body movement data corresponding to the third dialogue audio is determined. By iterating continuously, the body movement data corresponding to each target dialogue audio can be determined.
[0108] The following are examples.
[0109] The above-mentioned limb movement data divided into three action segments can be called a three-segment action. The two dialogue audios can be recorded as the first dialogue audio and the second dialogue audio according to the playback order, where the playback order of the second dialogue audio is next to that of the second dialogue audio.
[0110] If the duration of the first dialogue audio is not less than the total duration of the starting action segment and the shortest looping action segment, then the limb movement data corresponding to the first dialogue audio is set to the action data corresponding to the starting action segment and the looping action segment, and the limb movement data corresponding to the second dialogue audio is set to the ending action segment. The duration of the looping action segment is then adjusted so that the total duration of the starting action segment and the looping action segment is equal to the audio duration of the first dialogue audio. When the character outputs the first dialogue audio, the starting action segment and the looping action segment are performed; when the second dialogue audio is output, the ending action segment is performed. The playback time relationship between the first dialogue audio, the second dialogue audio, and the three-part action can be as follows: Figure 8 As shown.
[0111] Optionally, independent body movement data can also be set, hereinafter referred to as independent actions. To improve the richness of the characters' movements, it can be determined whether the audio duration of the second dialogue audio is not less than the sum of the duration of the ending action segment and at least one of the independent actions. If so, an independent action and the ending action segment can be selected as the body movement data corresponding to the second dialogue audio, such as... Figure 9As shown, the start time of the ending action segment can be aligned with the start time of the second dialogue audio, and the end time of the independent action can be aligned with the end time of the second dialogue audio. If not, the body movement data corresponding to the second audio segment is the ending action segment.
[0112] like Figure 10 As shown, if the audio length of the second dialogue audio is less than the ending action segment, the playback order is such that the playback time of the dialogue audio after the second dialogue audio is [time]. After the ending action segment ends, the characters corresponding to the first and second dialogue audios can be identified, and the complete three-part action can be completed.
[0113] If the duration of the first dialogue audio is less than the total duration of the starting action segment and the shortest loop action segment, then... Figure 11 The option shown allows you to select an independent action whose duration is closest to the duration of the first dialogue audio. This independent action will then be used as the limb movement data corresponding to the first dialogue audio. If the duration of the first dialogue audio is shorter than the duration of the independent action, then... Figure 12 As shown, the playback order is such that the playback time of the dialogue audio after the first dialogue audio is such that after this independent action is completed, the character corresponding to the first dialogue audio can complete the complete independent action.
[0114] Then, determine whether the audio duration of the second dialogue audio is greater than or equal to the shortest total duration of the starting action segment and the looping action segment of the three-segment action. If so, the action data corresponding to the starting action segment and the looping action segment in the three-segment action can be used as the limb action data corresponding to the second dialogue audio. If not, the independent action whose action duration is closest to the audio duration of the second dialogue audio can be selected as the limb action data corresponding to the second dialogue audio.
[0115] In one embodiment, lip-sync data can be preset. The lip-sync data is used to control the mouth movements of the characters when they speak. For example, the lip-sync data can include two actions: opening the mouth and closing the mouth. The opening and closing of the mouth can alternate. Based on the audio duration of the dialogue, the duration of the mouth opening, and the duration of the mouth closing, the arrangement of the mouth opening and closing actions in each dialogue is determined, thereby obtaining the lip-sync data of the characters corresponding to the speakers of each dialogue.
[0116] Optionally, the dialogue audio can be input into a neural network model, and the neural network model can extract audio features from the dialogue audio to obtain audio feature information for each dialogue audio. Then, the neural network model can predict the lip-sync data of the characters based on the audio feature information.
[0117] Optionally, silent segments and speech segments in the dialogue audio can be determined based on audio feature information, and corresponding lip-sync data can be generated for different audio segments. That is, in one embodiment, the audio feature information includes speech feature information, and the action data includes lip-sync data. The step "for each dialogue audio, determine the action data of the speaker's character during the speaking process based on the audio feature information of the dialogue audio" includes:
[0118] Based on the speech feature information of each pair of white audio, determine the silence segment and speech segment in each pair of white audio;
[0119] For each dialogue audio segment, first lip-sync data is generated based on the silent segment of the dialogue audio, and second lip-sync data is generated based on the speech segment.
[0120] Speech feature information may include information such as spectrum, energy, and zero-crossing rate.
[0121] Since there are usually pauses during speech, the silence segments and speech segments in each dialogue can be determined based on the speech feature information of each dialogue. A silence segment indicates that the character is not speaking, so the lip shape data of the character in the silence segment is the lip shape data corresponding to the mouth closing. A speech segment indicates that the character is speaking, so the lip shape data corresponding to the speech segment can be the lip shape data corresponding to the mouth opening and closing. Optionally, the lip shape data corresponding to the speech segment can also be predicted based on the neural network model.
[0122] 105. Using a virtual camera at the camera position corresponding to each dialogue, film at least two characters to obtain the plot video.
[0123] For example, based on the order of dialogue in the relevant content and the virtual camera at each dialogue shooting position, at least two characters can be filmed to obtain the plot video.
[0124] Suppose the dialogue contains dialogue 1, dialogue 2, and dialogue 3. The speaker of dialogue 1 is character m, and the corresponding camera position is x. The speakers of dialogue 2 and dialogue 3 are character n, and the corresponding camera position is y. Then, the dialogue 1 scene can be obtained by filming with a virtual camera on camera position x, where the subject of the filming is character m.
[0125] By using a virtual camera at camera position y, scenes of dialogue 2 and dialogue 3 are captured, with character n as the subject of the shoot. Based on the corresponding scenes of dialogue 1, dialogue 2, and dialogue 3, a narrative video can be obtained. Optionally, for dialogues with similar audio durations, the same scene can be used. For example, if the audio durations of dialogue 2 and dialogue 3 for character n are similar, the scene of dialogue 2 can be captured, and then copied to serve as the scene of dialogue 3.
[0126] Optionally, video clips of the characters performing corresponding body movements can be captured based on each preset body movement data. Then, for each character as the speaker in the target dialogue, video clips can be selected based on the body movement data corresponding to the characters and the target dialogue to obtain video clips for each character in the target dialogue.
[0127] For example, for a character m, we can film the character m making the target body movement according to each camera position corresponding to the character m's position, and obtain video clips of the character m making the target body movement under different camera positions. Similarly, we can obtain video clips of the character m for different body movements.
[0128] Then, based on the body movement data of character m for each line of dialogue, the video clips corresponding to the body movement data are obtained and spliced together to obtain the plot segment corresponding to each line of dialogue.
[0129] For dialogues that are text-based, speech synthesis can be performed on the text to obtain audio dialogues. In one embodiment, the dialogue includes text dialogues, and the step "shooting at least two characters based on a virtual camera at each corresponding camera position to obtain a story video" can specifically include:
[0130] Based on the speaker of each dialogue text, speech synthesis processing is performed on each dialogue text in the relevant content of the dialogue to obtain the corresponding dialogue audio for each dialogue text;
[0131] Based on the audio duration of the dialogue audio for each dialogue text, determine the duration for which the virtual camera will capture the speaker of each dialogue text.
[0132] Based on the camera position and shooting duration corresponding to each dialogue text, control the virtual camera to shoot at least two characters to obtain the plot video.
[0133] Among them, speech synthesis processing can convert dialogue text in text data form into dialogue in audio data form, i.e., dialogue audio, based on text to speech (TTS).
[0134] Based on the audio duration of each dialogue, the corresponding shooting duration for each dialogue can be determined. Based on the shooting duration of each dialogue text, a virtual camera is used to shoot based on the shooting position of each dialogue text to obtain the plot video.
[0135] After shooting the story video, you can add audio to the story video based on the dialogue audio, making the story video have sound.
[0136] Optionally, subtitles can be added to the story video based on the dialogue. In one embodiment, the step "shooting at least two characters based on a virtual camera at the camera position corresponding to each dialogue to obtain the story video" may specifically include:
[0137] Based on the camera position corresponding to each dialogue, control the virtual camera to film at least two characters to obtain candidate plot videos;
[0138] Generate subtitles for candidate story videos based on dialogue content;
[0139] A narrative video is generated based on candidate narrative videos and subtitles.
[0140] For example, a virtual camera at the camera position corresponding to each dialogue is used to film at least two characters to obtain candidate plot videos; then, subtitles for the candidate plot videos are generated based on the dialogue content, and a plot video is generated based on the candidate plot videos and subtitles.
[0141] As can be seen from the above, this application embodiment obtains video generation resources for generating a story video. The video generation resources include at least two characters, a target position template, and dialogue-related content for at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participant for each dialogue. Based on the target position template and at least two characters, the character position corresponding to each character is determined. For each dialogue in the dialogue-related content, the camera position for the dialogue is selected from the camera positions configured for the character position of the character who is the speaker. The dialogue is controlled to allow at least two characters to engage in dialogue based on the dialogue in the dialogue-related content. The virtual camera on the camera position corresponding to each dialogue is used to film the at least two characters to obtain the story video.
[0142] In this embodiment, each character's position in the positioning template is pre-set with a camera position. Based on the target positioning template, the position of each character and the corresponding camera position can be determined. Then, based on the dialogue content, the characters are controlled to engage in dialogue and be filmed. This can automatically generate a story video, improve the efficiency of generating story videos, avoid errors caused by manually editing the story, and improve the quality of the generated story videos.
[0143] To facilitate better implementation of the narrative video generation method provided in this application embodiment, a narrative video generation apparatus is also provided in one embodiment. The meanings of the terms used are the same as in the narrative video generation method described above, and specific implementation details can be found in the description of the method embodiment.
[0144] The narrative video generation device can be integrated into a computer device, such as... Figure 13 As shown, the narrative video generation device may include: an acquisition unit 301, a position determination unit 302, a camera position determination unit 303, a control unit 304, and a shooting unit 305, as detailed below:
[0145] (1) Acquisition unit 301 is used to acquire video generation resources for generating plot videos. The video generation resources include at least two characters, target position templates and dialogue-related content of at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participant of each dialogue.
[0146] In one embodiment, each character's position is equipped with at least one set of camera positions. Each set of camera positions includes a first camera position for filming a first character appearing at the character's position, and a second camera position for filming the first character appearing and a second character interacting with the first character appearing.
[0147] In one embodiment, the second camera position in each group of camera positions captures a different second character.
[0148] (2) Position determination unit 302 is used to determine the position of each character based on the target position template and at least two characters.
[0149] (3) Camera position determination unit 303 is used to select the camera position of the dialogue for each dialogue in the dialogue-related content, based on the camera position configured according to the position of the character who appears as the speaker.
[0150] In one embodiment, the dialogue includes dialogue text, and each character's position corresponds to at least two preset candidate camera positions. The camera position determination unit 303 includes:
[0151] The candidate camera position determination subunit is used to select, for each dialogue text, a set of camera positions from the at least one set of camera positions whose shooting content includes the participants of the dialogue text, to obtain a first candidate camera position and a second candidate camera position.
[0152] The first camera position selection subunit is used to select the first candidate shooting camera position for each dialogue text if the number of texts exceeds a preset number threshold.
[0153] The second camera selection subunit is used to select the second candidate camera position from the at least two candidate camera positions for each dialogue text if the number of texts does not exceed a preset threshold, based on the number of text units contained in each dialogue and the participants.
[0154] (4) Control unit 304, used to control at least two characters to have a dialogue based on the dialogue in the dialogue-related content.
[0155] In one embodiment, the dialogue includes dialogue audio, and the control unit 304 includes:
[0156] The extraction subunit is used to extract features from each pair of dialogue audio to obtain the audio feature information of each pair of dialogue audio.
[0157] The action determination subunit is used to generate action data of the speaker's character during the speaking process for each dialogue audio, based on the audio feature information of the dialogue audio.
[0158] The dialogue control subunit is used to control the speaking character among the at least two characters to perform corresponding actions during the dialogue, based on the action data, for each pair of audio dialogues.
[0159] In one embodiment, the action determination subunit includes:
[0160] The motion data acquisition module is used to acquire preset segmented motion data and independent motion data, as well as the audio duration of the target dialogue corresponding to each character. The segmented motion data includes motion data corresponding to the starting motion segment, the looping motion segment, and the ending motion segment. The motion duration of the looping motion segment is adjustable.
[0161] The selection module is used to select, for each character appearing, the limb movement data corresponding to the at least one target dialogue audio from the segmented limb movement data and the independent limb movement data, based on the relationship between the total audio duration of at least one target dialogue audio of the character appearing and the shortest movement duration of the segmented limb movement data.
[0162] In one embodiment, the selection module includes:
[0163] The audio selection submodule is used to select a first dialogue audio from the target dialogue audio corresponding to each character, and determine the audio duration of the first dialogue audio.
[0164] The first action determination submodule is used to use the independent action data as the limb action data corresponding to the first dialogue audio if the audio duration is less than the shortest total duration of the starting action segment and the loop action segment.
[0165] The second action determination submodule is used to take the action data corresponding to the starting action segment and the loop action segment as the limb action data corresponding to the first dialogue audio if the audio duration is not less than the shortest total action duration.
[0166] The first loop submodule is used to, if the first dialogue audio corresponds to independent action data, take the second dialogue audio in the target dialogue audio that is sequentially next to the first dialogue audio as the first dialogue audio, and return to execute the following: if the duration of the audio is less than the shortest total duration of the starting action segment and the loop action segment, take the independent action data as the limb action data corresponding to the first dialogue audio.
[0167] The second loop submodule is used to, if the first dialogue audio corresponds to segmented action data, take the ending action segment as the limb action data corresponding to the second dialogue audio, take the third dialogue audio in the target dialogue audio whose temporal order is after the second dialogue audio as the first dialogue audio, and return to execute the step of taking the independent action data as the limb action data corresponding to the first dialogue audio if the audio duration is less than the shortest total action duration of the starting action segment and the loop action segment.
[0168] In one embodiment, the audio feature information includes speech feature information, the action data includes lip-sync data, and the dialogue control subunit includes:
[0169] The segment determination module is used to determine the silent segments and speech segments in each dialogue pair based on the speech feature information of each dialogue pair audio.
[0170] The lip-sync data generation module is used to generate first lip-sync data based on the silent segments of the dialogue audio for each dialogue audio, and to generate second lip-sync data based on the speech segments.
[0171] (5) Shooting unit 305 is used to shoot at least two characters based on the virtual camera on the shooting position corresponding to each dialogue to obtain the plot video.
[0172] In one embodiment, the imaging unit 305 includes:
[0173] The story filming subunit is used to film at least two characters based on the virtual camera on the camera position corresponding to each dialogue, and to obtain candidate story videos;
[0174] The subtitle generation subunit is used to generate subtitles for candidate plot videos based on dialogue-related content;
[0175] The video generation subunit is used to generate a narrative video based on candidate narrative videos and subtitles.
[0176] In one embodiment, the dialogue includes spoken text, and the shooting unit 305 includes:
[0177] The speech synthesis subunit is used to perform speech synthesis processing on each dialogue in the relevant content of the dialogue according to the speaker of each dialogue, so as to obtain the corresponding dialogue audio for each dialogue.
[0178] The duration determination subunit is used to determine the recording duration of the virtual camera for each speaker in the dialogue based on the audio duration of the dialogue audio of each dialogue.
[0179] The video shooting subunit is used to control the virtual camera to shoot at least two characters based on the shooting position and shooting duration corresponding to each dialogue, so as to obtain the plot video.
[0180] As can be seen from the above, the narrative video generation device in this application embodiment acquires video generation resources for generating narrative videos through the acquisition unit 301. The video generation resources include at least two characters, a target position template, and dialogue-related content for at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participant for each dialogue. The position determination unit 302 determines the character position corresponding to each character based on the target position template and at least two characters. The camera position determination unit 303 determines the camera position corresponding to each dialogue for each dialogue in the dialogue-related content based on the character position of the character who is the speaker. The control unit 304 controls at least two characters to engage in dialogue based on the dialogue in the dialogue-related content. The shooting unit 305 shoots at least two characters based on the virtual camera on the camera position corresponding to each dialogue to obtain the narrative video.
[0181] In this embodiment, each character's position in the positioning template is pre-set with a camera position. Based on the target positioning template, the position of each character and the corresponding camera position can be determined. Then, based on the dialogue content, the characters are controlled to engage in dialogue and be filmed. This can automatically generate a story video, improve the efficiency of generating story videos, avoid errors caused by manually editing the story, and improve the quality of the generated story videos.
[0182] Accordingly, embodiments of this application also provide a computer device, which can be a terminal. For example... Figure 14 As shown, Figure 14This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. The computer device 500 includes a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, and a computer program stored on the memory 502 and executable on the processor. The processor 501 and the memory 502 are electrically connected. Those skilled in the art will understand that the computer device structure shown in the figure does not constitute a limitation on the computer device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0183] The processor 501 is the control center of the computer device 500. It connects various parts of the computer device 500 through various interfaces and lines. By running or loading software programs and / or modules stored in the memory 502, and calling data stored in the memory 502, it performs various functions of the computer device 500 and processes data, thereby monitoring the computer device 500 as a whole.
[0184] In this embodiment, the processor 501 in the computer device 500 loads the instructions corresponding to the processes of one or more applications into the memory 502 according to the following steps, and the processor 501 runs the applications stored in the memory 502 to achieve various functions:
[0185] Acquire video generation resources for generating story videos. The video generation resources include at least two characters, target position templates, and dialogue-related content for at least two characters. Each character's position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participants in each dialogue.
[0186] Based on the target positioning template and at least two characters to appear, determine the character positioning for each character to appear.
[0187] For each line of dialogue, select the camera position for that line of dialogue from the camera positions configured according to the position of the character who is speaking.
[0188] Control at least two characters to engage in dialogue based on the dialogue content.
[0189] The video is produced by filming at least two characters using a virtual camera positioned at the camera location corresponding to each dialogue.
[0190] In one embodiment, the dialogue includes audio data, and the step of "controlling at least two characters to converse based on dialogue in dialogue-related content" may include:
[0191] Feature extraction is performed on each pair of dialogue audio to obtain the audio feature information of each pair of dialogue audio;
[0192] For each dialogue audio, based on the audio feature information of the dialogue audio, the action data of the character who appears as the speaker during the speaking process is generated.
[0193] For each dialogue audio, based on the motion data, the speaking character among the at least two characters is controlled to perform corresponding actions during the dialogue.
[0194] In one embodiment, the audio feature information includes audio duration, and the motion data includes body motion data. The step "for each dialogue audio, generating motion data of the character appearing as the speaker during the speaking process based on the audio feature information of the dialogue audio" may include:
[0195] Acquire preset segmented action data and independent action data, and the total audio duration of the target dialogue audio for each character. The segmented action data includes the action data corresponding to the starting action segment, the looping action segment, and the ending action segment. The action duration of the looping action segment is adjustable.
[0196] For each character, based on the total audio duration of at least one target dialogue of the character and the relationship between the shortest action duration of the segmented body movement data, select at least one body movement data corresponding to the target dialogue audio from the segmented action data and the independent action data.
[0197] In one embodiment, the step "for each character appearing, based on the total audio duration of at least one target dialogue audio of the character appearing and the relationship between the shortest action duration of the segmented action data and the independent action data, select at least one body action data corresponding to the target dialogue from the segmented action data and the independent action data" may include:
[0198] For each character appearing in the game, select the first dialogue audio from the target dialogue audio corresponding to the character, and determine the audio duration of the first dialogue audio.
[0199] If the audio duration is less than the shortest total duration of the starting action segment and the looping action segment, then the independent action data will be used as the limb action data corresponding to the first dialogue audio.
[0200] If the audio duration is not less than the shortest total action duration, then the action data corresponding to the starting action segment and the looping action segment will be used as the limb action data corresponding to the first dialogue audio.
[0201] If the first dialogue audio corresponds to independent action data, then the second dialogue audio in the target dialogue audio that is sequentially next to the first dialogue audio is taken as the first dialogue audio, and the execution is returned. If the audio duration is less than the shortest total duration of the starting action segment and the loop action segment, then the independent action data is taken as the limb action data corresponding to the first dialogue audio.
[0202] If the first dialogue audio corresponds to segmented action data, then the ending action segment is used as the limb action data corresponding to the second dialogue audio. The third dialogue audio in the target dialogue audio, which is sequentially next to the second dialogue audio, is used as the first dialogue audio, and execution returns. If the audio duration is less than the shortest total duration of the starting action segment and the looping action segment, then the independent action data is used as the limb action data corresponding to the first dialogue audio. The action duration of the looping segment in the preset limb actions is adjusted to obtain the limb action data corresponding to each character and the target dialogue.
[0203] In one embodiment, the audio feature information includes speech feature information, and the action data includes lip-sync data. The step "for each dialogue audio, determine the action data of the character speaking during the dialogue based on the audio feature information of the dialogue" may include:
[0204] Based on the speech feature information of each dialogue pair, determine the silent segments and speech segments in each dialogue pair;
[0205] For each dialogue audio, first lip-sync data is generated based on the silent segments of the dialogue, and second lip-sync data is generated based on the speech segments.
[0206] In one embodiment, the step "shooting at least two characters based on a virtual camera at the camera position corresponding to each dialogue to obtain a story video" may include:
[0207] Based on the virtual camera at the shooting position corresponding to each dialogue, at least two characters appearing are filmed to obtain candidate plot videos;
[0208] Generate subtitles for candidate story videos based on dialogue content;
[0209] A narrative video is generated based on candidate narrative videos and subtitles.
[0210] In one embodiment, the dialogue includes spoken text, and the step of "filming at least two characters based on the virtual camera at the camera position corresponding to each dialogue to obtain a story video" may include:
[0211] Based on the speaker of each dialogue, each dialogue text in the relevant content of the dialogue is processed by speech synthesis to obtain the corresponding dialogue audio for each dialogue text;
[0212] Based on the audio duration of the dialogue audio for each dialogue text, determine the duration for which the virtual camera will capture the speaker of each dialogue text.
[0213] Based on the camera position and shooting duration corresponding to each dialogue, control the virtual camera to shoot at least two characters to obtain the plot video.
[0214] In one embodiment, each character's position is equipped with at least one set of camera positions. Each set of camera positions includes a first camera position for filming the first character to appear in the character's position, and a second camera position for filming the first character to appear and the second character to appear in dialogue with the first character.
[0215] In one embodiment, the second camera position in each group of camera positions captures a different second character.
[0216] In one embodiment, the dialogue includes dialogue text, and the step "for each dialogue in the dialogue-related content, select a camera position for the dialogue from the camera positions configured according to the position of the character appearing as the speaker" may include:
[0217] For each dialogue text in the dialogue content, determine at least one set of camera positions for the character standing position of the speaking party;
[0218] For each dialogue text, based on the participants in the dialogue text, select a second camera position from at least one set of camera positions. The content of the second camera position includes the camera positions of the participants, thus obtaining a first candidate camera position and a second candidate camera position.
[0219] For each dialogue text, if the number of texts exceeds a preset threshold, the first candidate shooting position is selected;
[0220] For each dialogue text, if the number of texts does not exceed a preset threshold, a second candidate camera position is selected from at least two candidate camera positions based on the number of text units contained in each dialogue and the participants.
[0221] As can be seen from the above, this application embodiment obtains video generation resources for generating a story video. The video generation resources include at least two characters, a target position template, and dialogue-related content for at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participant for each dialogue. Based on the target position template and at least two characters, the character position corresponding to each character is determined. For each dialogue in the dialogue-related content, the camera position for the dialogue is selected from the camera positions configured for the character position of the character who is the speaker. The dialogue is controlled to allow at least two characters to engage in dialogue based on the dialogue in the dialogue-related content. The virtual camera on the camera position corresponding to each dialogue is used to film the at least two characters to obtain the story video.
[0222] In this embodiment, each character's position in the positioning template is pre-set with a camera position. Based on the target positioning template, the position of each character and the corresponding camera position can be determined. Then, based on the dialogue content, the characters are controlled to engage in dialogue and be filmed. This can automatically generate a story video, improve the efficiency of generating story videos, avoid errors caused by manually editing the story, and improve the quality of the generated story videos.
[0223] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0224] Optional, such as Figure 14 As shown, the computer device 500 also includes: a touch screen display 503, a radio frequency circuit 504, an audio circuit 505, an input unit 506, and a power supply 507. The processor 501 is electrically connected to the touch screen display 503, the radio frequency circuit 504, the audio circuit 505, the input unit 506, and the power supply 507. Those skilled in the art will understand that... Figure 14 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0225] The touch display screen 503 can be used to display a graphical user interface (GUI) and receive operation commands generated by the user interacting with the GUI. The touch display screen 503 may include a display panel and a touch panel. The display panel can be used to display information input by the user or information provided to the user, as well as various graphical user interfaces of the computer device. These graphical user interfaces can be composed of graphics, text, icons, video, and any combination thereof. Optionally, the display panel can be configured using a liquid crystal display (LCD), organic light-emitting diode (OLED), or other similar technologies. The touch panel can be used to collect touch operations performed by the user on or near it (such as operations performed by the user using a finger, stylus, or any suitable object or accessory on or near the touch panel), generate corresponding operation commands, and execute the corresponding program according to the operation commands. Optionally, the touch panel may include two parts: a touch detection device and a touch controller. The touch detection device detects the user's touch location and the signal generated by the touch operation, transmitting the signal to the touch controller. The touch controller receives touch information from the touch detection device, converts it into touch point coordinates, and sends it to the processor 501. It can also receive and execute commands from the processor 501. The touch panel can cover the display panel. When the touch panel detects a touch operation on or near it, it transmits the information to the processor 501 to determine the type of touch event. Subsequently, the processor 501 provides corresponding visual output on the display panel based on the type of touch event. In this embodiment, the touch panel and the display panel can be integrated into the touch display screen 503 to achieve input and output functions. However, in some embodiments, the touch panel and the touch display screen 503 can be implemented as two independent components to achieve input and output functions. That is, the touch display screen 503 can also be used as part of the input unit 506 to achieve input functions.
[0226] The radio frequency circuit 504 can be used to transmit and receive radio frequency signals to establish wireless communication with network devices or other computer devices, and to transmit and receive signals with network devices or other computer devices.
[0227] Audio circuitry 505 can be used to provide an audio interface between a user and a computer device via a speaker and a microphone. Audio circuitry 505 converts received audio data into electrical signals, transmits them to the speaker, and the speaker converts them into sound signals for output. Conversely, the microphone converts collected sound signals into electrical signals, which are then received by audio circuitry 505, converted back into audio data, and output to processor 501 for processing. The audio data is then transmitted via radio frequency circuitry 504 to, for example, another computer device, or output to memory 502 for further processing. Audio circuitry 505 may also include an earphone jack to facilitate communication between peripheral headphones and the computer device.
[0228] The input unit 506 can be used to receive input numbers, characters, or user characteristic information (such as fingerprints, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.
[0229] Power supply 507 is used to supply power to various components of computer device 500. Optionally, power supply 507 can be logically connected to processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. Power supply 507 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0230] although Figure 14 As not shown in the diagram, the computer device 500 may also include a camera, sensor, wireless fidelity module, Bluetooth module, etc., which will not be described in detail here.
[0231] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0232] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0233] Therefore, embodiments of this application provide a computer-readable storage medium storing a plurality of computer programs that can be loaded by a processor to execute the steps in any of the virtual item marking methods provided in embodiments of this application. For example, the computer program can execute the following steps:
[0234] Acquire video generation resources for generating story videos. The video generation resources include at least two characters, target position templates, and dialogue-related content for at least two characters. Each character's position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participants in each dialogue.
[0235] Based on the target positioning template and at least two characters to appear, determine the character positioning for each character to appear.
[0236] For each line of dialogue, select the camera position for that line of dialogue from the camera positions configured according to the position of the character who is speaking.
[0237] Control at least two characters to engage in dialogue based on the dialogue content;
[0238] The video is produced by filming at least two characters using a virtual camera positioned at the camera location corresponding to each dialogue.
[0239] In one embodiment, the dialogue includes audio data, and the step of "controlling at least two characters to converse based on dialogue in dialogue-related content" may include:
[0240] Feature extraction is performed on each pair of dialogues to obtain the audio feature information of each pair of dialogues;
[0241] For each dialogue, based on the audio feature information of the dialogue audio, action data of the character who appears as the speaker during the speaking process is generated.
[0242] For each dialogue audio, based on the motion data, the speaking character among the at least two characters is controlled to perform corresponding actions during the dialogue.
[0243] In one embodiment, the audio feature information includes audio duration, and the motion data includes body motion data. The step "for each dialogue audio, generating motion data of the character appearing as the speaker during the speaking process based on the audio feature information of the dialogue audio" may include:
[0244] Acquire preset segmented action data and independent action data, and the total audio duration of the target dialogue for each character as the speaker. The segmented action data includes the action data corresponding to the starting action segment, the looping action segment, and the ending action segment. The action duration of the looping action segment is adjustable.
[0245] For each character, based on the total audio duration of at least one target dialogue of the character and the relationship between the shortest action duration of the segmented body movement data, select at least one body movement data corresponding to the target dialogue audio from the segmented action data and the independent action data.
[0246] In one embodiment, the step "for each character appearing, based on the total audio duration of at least one target dialogue audio of the character appearing and the relationship between the shortest action duration of the segmented action data and the independent action data, select at least one body action data corresponding to the target dialogue from the segmented action data and the independent action data" may include:
[0247] For each character appearing in the game, select the first dialogue audio from the target dialogue audio corresponding to the character, and determine the audio duration of the first dialogue audio.
[0248] If the audio duration is less than the shortest total duration of the starting action segment and the looping action segment, then the independent action data will be used as the limb action data corresponding to the first dialogue audio.
[0249] If the audio duration is not less than the shortest total action duration, then the action data corresponding to the starting action segment and the looping action segment will be used as the limb action data corresponding to the first dialogue audio.
[0250] If the first dialogue audio corresponds to independent action data, then the second dialogue audio in the target dialogue audio that is sequentially next to the first dialogue audio is taken as the first dialogue audio, and the execution is returned. If the audio duration is less than the shortest total duration of the starting action segment and the loop action segment, then the independent action data is taken as the limb action data corresponding to the first dialogue audio.
[0251] If the first dialogue audio corresponds to segmented action data, then the ending action segment is used as the limb action data corresponding to the second dialogue audio. The third dialogue audio in the target dialogue audio, which is sequentially next to the second dialogue audio, is used as the first dialogue audio, and execution returns. If the audio duration is less than the shortest total duration of the starting action segment and the looping action segment, then the independent action data is used as the limb action data corresponding to the first dialogue audio. The action duration of the looping segment in the preset limb actions is adjusted to obtain the limb action data corresponding to each character and the target dialogue.
[0252] In one embodiment, the audio feature information includes speech feature information, and the action data includes lip-sync data. The step "for each dialogue audio, determine the action data of the speaker's character during the speaking process based on the audio feature information of the dialogue audio" may include:
[0253] Based on the speech feature information of each dialogue audio, determine the silence segments and speech segments in each dialogue audio;
[0254] For each dialogue, first lip-sync data is generated based on the silent segments of the dialogue, and second lip-sync data is generated based on the speech segments.
[0255] In one embodiment, the step "shooting at least two characters based on a virtual camera at the camera position corresponding to each dialogue to obtain a story video" may include:
[0256] Based on the virtual camera at the shooting position corresponding to each dialogue, at least two characters appearing are filmed to obtain candidate plot videos;
[0257] Generate subtitles for candidate story videos based on dialogue content;
[0258] A narrative video is generated based on candidate narrative videos and subtitles.
[0259] In one embodiment, the dialogue includes dialogue text, and the step of "filming at least two characters based on the virtual camera at the camera position corresponding to each dialogue to obtain a story video" may include:
[0260] Based on the speaker of each dialogue, each dialogue in the relevant content of the dialogue is processed by speech synthesis to obtain the corresponding audio of each dialogue text;
[0261] Based on the audio duration of the dialogue audio for each dialogue text, determine the recording duration of the virtual camera for each speaker in the dialogue;
[0262] Based on the camera position and shooting duration corresponding to each dialogue text, control the virtual camera to shoot at least two characters to obtain the plot video.
[0263] In one embodiment, each character's position is equipped with at least one set of camera positions. Each set of camera positions includes a first camera position for filming the first character to appear in the character's position, and a second camera position for filming the first character to appear and the second character to appear in dialogue with the first character.
[0264] In one embodiment, the second camera position in each group of camera positions captures a different second character. In one embodiment, the dialogue includes dialogue text, and the step "for each dialogue in the dialogue-related content, select the camera position for the dialogue from the camera positions configured according to the character's position as the speaker" may include:
[0265] For each dialogue text in the dialogue content, determine at least one set of camera positions for the character standing position of the speaking party;
[0266] For each dialogue text, based on the participants in the dialogue text, select a second camera position from at least one set of camera positions. The content of the second camera position includes the camera positions of the participants, thus obtaining a first candidate camera position and a second candidate camera position.
[0267] For each dialogue text, if the number of texts exceeds a preset threshold, the first candidate shooting position is selected;
[0268] For each dialogue text, if the number of texts does not exceed a preset threshold, a second candidate camera position is selected from at least two candidate camera positions based on the number of text units contained in each dialogue and the participants.
[0269] As can be seen from the above, this application embodiment obtains video generation resources for generating a story video. The video generation resources include at least two characters, a target position template, and dialogue-related content for at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participant for each dialogue. Based on the target position template and at least two characters, the character position corresponding to each character is determined. For each dialogue in the dialogue-related content, the camera position corresponding to each dialogue is determined based on the character position of the character who is the speaker. The dialogue-related content controls at least two characters to engage in dialogue. The virtual camera on the camera position corresponding to each dialogue is used to film the at least two characters to obtain the story video.
[0270] In this embodiment, each character's position in the positioning template is pre-set with a camera position. Based on the target positioning template, the position of each character and the corresponding camera position can be determined. Then, based on the dialogue content, the characters are controlled to engage in dialogue and be filmed. This can automatically generate a story video, improve the efficiency of generating story videos, avoid errors caused by manually editing the story, and improve the quality of the generated story videos.
[0271] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0272] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0273] The above provides a detailed description of a method, apparatus, computer device, and computer storage medium for generating a narrative video according to embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, those skilled in the art will recognize that there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for generating a narrative video, characterized in that, include: Obtain video generation resources for generating story videos. The video generation resources include at least two characters, a target position template, and dialogue-related content for the at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participants in each dialogue. Each dialogue includes a pair of audio recordings. Based on the target position template and the at least two characters appearing in the game, determine the character position corresponding to each character appearing in the game; For each line of dialogue in the aforementioned dialogue content, the camera position for that line of dialogue is selected from the camera positions configured according to the position of the character who is speaking in the dialogue. Control the dialogue of at least two characters based on the dialogue content; The at least two characters appearing in the scene are filmed using a virtual camera at the camera position corresponding to each dialogue, to obtain the plot video; The method of controlling the dialogue of at least two characters based on the dialogue-related content includes: Audio features are extracted from each pair of white audio audio to obtain audio feature information for each pair of white audio audio, including audio duration; For each dialogue audio, based on the audio feature information of the dialogue audio, action data of the character speaking in the dialogue audio is generated during the speaking process. The action data includes body movement data. The generation steps include: acquiring preset segmented action data and independent action data, and the audio duration of the target dialogue audio corresponding to each character. The segmented action data includes action data corresponding to the starting action segment, the looping action segment, and the ending action segment, wherein the duration of the looping action segment is adjustable; for each character, based on the relationship between the total audio duration of at least one target dialogue audio of the character and the shortest action duration of the segmented action data, body movement data corresponding to the at least one target dialogue audio is selected from the segmented action data and the independent action data. For each dialogue audio, based on the motion data, the speaking character among the at least two characters is controlled to perform corresponding actions during the dialogue.
2. The method according to claim 1, characterized in that, For each character appearing in the scene, based on the relationship between the total audio duration of at least one target dialogue audio of the character and the shortest action duration of the segmented action data, the body movement data corresponding to the at least one target dialogue is selected from the segmented action data and the independent action data, including: For each character appearing in the scene, select the first dialogue audio from the target dialogue audio corresponding to the character appearing in the scene, and determine the audio duration of the first dialogue audio. Determine the relationship between the audio duration and the total shortest duration of the starting action segment and the looping action segment; If the audio duration is less than the shortest total duration of the starting action segment and the looping action segment, then the independent action data will be used as the limb action data corresponding to the first dialogue audio. If the audio duration is not less than the shortest total action duration, then the action data corresponding to the starting action segment and the looping action segment are used as the limb action data corresponding to the first dialogue audio. If the first dialogue audio corresponds to independent action data, then the second dialogue audio in the target dialogue audio that is sequentially next to the first dialogue audio is taken as the first dialogue audio, and the step of determining the relationship between the audio duration and the shortest total duration of the starting action segment and the loop action segment is returned. If the first dialogue audio corresponds to segmented action data, then the ending action segment is used as the limb action data corresponding to the second dialogue audio, the third dialogue audio in the target dialogue audio that is sequentially next to the second dialogue audio is used as the first dialogue audio, and the step of determining the relationship between the audio duration and the shortest total duration of the starting action segment and the looping action segment is returned. If neither the second nor the third dialogue audio exists, the step of determining the physical movement data of the character entering the scene is terminated.
3. The method according to claim 1, characterized in that, The audio feature information includes speech feature information, and the action data includes lip-sync data. For each dialogue audio, based on the audio feature information of the dialogue audio, the action data of the speaker during the speaking process is determined, including: Based on the speech feature information of each pair of white audio, determine the silence segment and speech segment in each pair of white audio; For each dialogue audio segment, first lip-sync data is generated based on the silent segment of the dialogue audio, and second lip-sync data is generated based on the speech segment.
4. The method according to claim 1, characterized in that, The process of filming at least two characters based on a virtual camera at a camera position corresponding to each dialogue to obtain a story video includes: Based on the virtual camera at the shooting position corresponding to each dialogue, at least two characters appearing are filmed to obtain candidate plot videos; Subtitles for the candidate story videos are generated based on the dialogue content; The story video is generated based on the candidate story videos and the subtitles.
5. The method according to claim 1, characterized in that, The dialogue includes dialogue text, and the process of filming at least two characters based on a virtual camera at a camera position corresponding to each dialogue to obtain a story video includes: Based on the speaker of each dialogue text, each dialogue text in the dialogue-related content is processed by speech synthesis to obtain the dialogue audio corresponding to each dialogue text; The duration of the virtual camera's recording of the speaker in each dialogue text is determined based on the audio duration of the dialogue audio for each dialogue text. Based on the camera position and shooting duration corresponding to each dialogue text, the virtual camera is controlled to shoot at least two characters to obtain the plot video.
6. The method according to any one of claims 1-5, characterized in that, Each character's position is equipped with at least one set of camera positions. Each set of camera positions includes a first camera position for filming the first character to appear in the position, and a second camera position for filming the first character to appear and the second character to appear in dialogue with the first character.
7. The method according to claim 6, characterized in that, The second camera in each group of cameras captures a different character making their second appearance.
8. The method according to claim 7, characterized in that, The dialogue includes dialogue text. The step of selecting a camera position for each line of dialogue within the relevant content, based on the camera positions configured according to the character's standing position as the speaker, includes: For each dialogue text in the aforementioned dialogue content, determine at least one set of camera positions configured for the character standing position of the speaking party; For each dialogue text, based on the participants in the dialogue text, from the at least one set of camera positions, a second camera position is selected whose shooting content includes the participants, thus obtaining a first candidate camera position and a second candidate camera position. For each dialogue text, if the number of texts exceeds a preset threshold, then the first candidate shooting position is selected; For each dialogue text, if the number of texts does not exceed a preset threshold, then the second candidate shooting position is selected.
9. A narrative video generation device, characterized in that, include: The acquisition unit is used to acquire video generation resources for generating a story video. The video generation resources include at least two characters, a target position template, and dialogue-related content of the at least two characters. Each character position in the target position template is configured with a camera position. The dialogue-related content includes dialogue and the speaker and participants of each dialogue. Each dialogue includes a pair of audio recordings. The positioning determination unit is used to determine the position of each character based on the target positioning template and the at least two characters appearing in the game. The camera position determination unit is used to select the camera position for each dialogue in the dialogue-related content, based on the camera positions configured according to the position of the character who is speaking. A control unit is used to control the at least two characters to engage in dialogue based on the dialogue-related content. A shooting unit is used to shoot the at least two characters based on a virtual camera at the shooting position corresponding to each dialogue, so as to obtain a plot video; An extraction subunit is used to extract audio features from each pair of white audio audios to obtain audio feature information for each pair of white audio audios, wherein the audio feature information includes audio duration. An action determination subunit is used to generate action data for each dialogue audio segment, based on the audio feature information of the dialogue audio, for the speaking character in the dialogue audio during the speaking process. The action data includes body movement data. The generation steps include: acquiring preset segmented action data and independent action data, and the audio duration of the target dialogue audio corresponding to each speaking character. The segmented action data includes action data corresponding to the starting action segment, the looping action segment, and the ending action segment, wherein the duration of the looping action segment is adjustable; for each speaking character, based on the relationship between the total audio duration of at least one target dialogue audio of the speaking character and the shortest action duration of the segmented action data, selecting the body movement data corresponding to the at least one target dialogue audio from the segmented action data and the independent action data. The dialogue control subunit is used to control the speaking character among the at least two characters to perform corresponding actions during the dialogue, based on the action data, for each pair of audio dialogues.
10. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to perform the narrative video generation method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, which is loaded by a processor to perform the narrative video generation method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Voice dialogue script generation method and device and electronic equipment
CN116312456A