A conversational video creation method and related apparatus

Through the dialogue video creation method, the plot clips of the virtual human video are visually split according to the dialogue dimension, and the video structure parameters are generated using the plot conversation area and artificial intelligence technology, which solves the problem of high complexity in the production of virtual character dialogue videos and simplifies the video editing process.

CN115115728BActive Publication Date: 2025-10-10TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210577360.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2025-10-10
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

Dialogue videos are difficult to produce, especially when controlling the content of dialogue between virtual characters. The operation is complex and the threshold for video editing is high.

Method used

The plot clips of the virtual human video are visually split according to the dialogue dimension, the plot bubbles are displayed in the plot conversation area, and the video structure parameters are generated using artificial intelligence technology to simplify the video editing process.

Benefits of technology

It reduces the operational complexity of video editing, reduces the number of interactions during the video editing process, and significantly lowers the threshold for video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115115728B_ABST
    Figure CN115115728B_ABST
Patent Text Reader

Abstract

The application discloses a dialogue video creation method and related device, which can be applied to various scenes such as audio and video, cloud technology, artificial intelligence, computer vision technology, speech technology, natural language processing, machine learning, intelligent transportation, and auxiliary driving, aiming at basic technologies such as image processing and face processing. In the method, a virtual person video including multiple plot segments is visually split in the dialogue dimension, so that the plot bubbles in the plot conversation area can intuitively display the dialogue content of different virtual objects as the social conversation interface, and the plot order of the dialogue content in the virtual person video can be clearly and intuitively sorted through the display position relationship between the plot bubbles. When the virtual person video needs to be edited, the editing can be realized in the plot bubble dimension in a manner similar to editing the social conversation, the operation complexity is reduced, the number of interactions in the video editing process is reduced, and the video editing threshold is greatly reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing, and in particular to a method and related device for creating a conversation video. Background Art

[0002] As the video industry develops, video enthusiasts are creating virtual human videos through games and official editors. Unlike traditional video production methods, virtual human videos are created using virtual characters and virtual scenes. Creating conversational videos based on virtual humans, using dialogue to advance the plot, can simplify video production.

[0003] However, the virtual characters and their conversation content still need to be controlled in the dialogue videos, and the production of dialogue videos is still quite difficult. Summary of the Invention

[0004] In order to solve the above technical problems, the present application provides a dialogue video creation method and related devices, which visually split the virtual human video including multiple plot clips according to the dialogue dimension, reduce the operation complexity, reduce the number of interactions during the video editing process, and greatly lower the threshold for video editing.

[0005] The embodiments of this application disclose the following technical solutions:

[0006] In one aspect, the present application provides a method for creating a conversation video, the method comprising:

[0007] A plot conversation area is displayed for editing a virtual human video, wherein the virtual human video to be edited includes multiple plot segments, and the plot conversation area displays a first plot bubble corresponding to a first segment of the multiple plot segments, wherein the first plot bubble is associated with a first virtual object involved in the first segment, and the first plot bubble includes a first plot text of the first virtual object in the first segment;

[0008] In response to an instruction to add a second segment from the plurality of segment segments, displaying a second segment bubble corresponding to the second segment in the segment conversation area, the second segment bubble being associated with a second virtual object related to the second segment, the second segment bubble including second segment text of the second virtual object in the second segment, the first segment bubble and the second segment bubble having a display position relationship in the segment conversation area, the display position relationship being used to identify a plot order of the first segment and the second segment in the virtual human video;

[0009] In response to a video generation instruction, video structure parameters are generated according to the first plot bubble and the second plot bubble displayed in the plot conversation area, and the video structure parameters are used to play the virtual human video when called.

[0010] In another aspect, the present application provides a device for creating a conversation video, the device comprising:

[0011] a plot conversation area display unit, configured to display a plot conversation area for editing a virtual human video, wherein the virtual human video to be edited includes multiple plot segments, and the plot conversation area displays a first plot bubble corresponding to a first segment of the multiple plot segments, wherein the first plot bubble is associated with a first virtual object involved in the first segment, and the first plot bubble includes a first plot text of the first virtual object in the first segment;

[0012] a plot bubble display unit, configured to, in response to an instruction to add a second segment from the plurality of plot segments, display a second plot bubble corresponding to the second segment in the plot conversation area, wherein the second plot bubble is associated with a second virtual object related to the second segment, the second plot bubble includes second plot text of the second virtual object in the second segment, and the first plot bubble and the second plot bubble have a display position relationship in the plot conversation area, the display position relationship being used to identify a plot order of the first segment and the second segment in the virtual human video;

[0013] The video structure parameter generating unit is used to generate video structure parameters according to the first plot bubble and the second plot bubble displayed in the plot conversation area in response to a video generation instruction, wherein the video structure parameters are used to play the virtual human video when called.

[0014] In another aspect, the present application provides a computer device, comprising a processor and a memory:

[0015] The memory is used to store a computer program and transmit the computer program to the processor;

[0016] The processor is configured to execute the conversation video creation method described in the above aspect according to instructions in the computer program.

[0017] On the other hand, an embodiment of the present application provides a computer-readable storage medium, which is used to store a computer program, and the computer program is used to execute the conversation video creation method described in the above aspects.

[0018] On the other hand, an embodiment of the present application provides a computer program product including a computer program, which, when executed on a computer device, enables the computer device to execute the method for creating a conversation video.

[0019] As can be seen from the above technical solution, a plot conversation area is displayed for editing a virtual human video. The virtual human video to be edited includes multiple plot segments, which may include a first segment and a second segment. The plot conversation area displays a first plot bubble corresponding to the first segment. In response to an instruction to add a second segment, a second plot bubble corresponding to the second segment is displayed in the plot conversation area. In other words, the plot bubbles displayed in the plot conversation area include the plot text of the corresponding plot segment, and different plot bubbles can be associated with the same or different virtual objects. The display position relationship between the plot bubbles in the plot conversation area identifies the plot order of the corresponding plot segments in the virtual human video. Based on the first and second plot bubbles displayed in the plot conversation area, video structure parameters for playing the virtual human video when called can be generated. In this way, a virtual human video including multiple plot clips can be visually split based on the dialogue dimension, so that the plot bubbles in the plot conversation area can intuitively display the dialogue content of different virtual objects like the social conversation interface. The plot sequence of the dialogue content in the virtual human video can be clearly and intuitively sorted out through the display position relationship between the plot bubbles. When the plot of the virtual human video needs to be edited, it can be achieved from the dimension of the plot bubbles in a manner similar to editing social conversations, reducing the complexity of the operation. The display of the plot bubbles in the plot conversation area can promptly and clearly reflect the increase and content of the plot clips in the virtual human video, reducing the number of interactions during the video editing process, and greatly lowering the threshold for video editing. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0021] Figure 1 A schematic diagram of an application scenario of a conversation video creation method provided in an embodiment of the present application;

[0022] Figure 2 A signaling diagram of a method for creating a conversational video provided in an embodiment of the present application;

[0023] Figure 3 A schematic diagram of a plot conversation area provided in an embodiment of the present application;

[0024] Figure 4 A schematic diagram of an operation interface provided in an embodiment of the present application;

[0025] Figure 5 A schematic diagram of another plot conversation area provided in an embodiment of the present application;

[0026] Figure 6 A schematic diagram of a display interface provided in an embodiment of the present application;

[0027] Figure 7-Figure 9 A schematic diagram of a plot conversation area provided in an embodiment of the present application;

[0028] Figure 10 A schematic diagram of video structure parameters provided in an embodiment of the present application;

[0029] Figure 11-14 A schematic diagram of another plot conversation area provided in an embodiment of the present application;

[0030] Figure 15 A schematic diagram of an action collection interface provided in an embodiment of the present application;

[0031] Figure 16 A schematic diagram of another action collection interface provided in an embodiment of the present application;

[0032] Figure 17 A schematic diagram of another action collection interface provided in an embodiment of the present application;

[0033] Figure 18 A schematic diagram of another plot conversation area provided in an embodiment of the present application;

[0034] Figure 19 A schematic diagram of another plot conversation area provided in an embodiment of the present application;

[0035] Figure 20 A functional schematic diagram of a plot conversation area provided in an embodiment of the present application;

[0036] Figure 21 A schematic diagram of information processing provided in an embodiment of the present application;

[0037] Figure 22 A schematic diagram of another information processing provided in an embodiment of the present application;

[0038] Figure 23 A schematic diagram of another video structure parameter provided in an embodiment of the present application;

[0039] Figure 24 A structural block diagram of a conversation video creation device provided in an embodiment of the present application;

[0040] Figure 25A structural diagram of a terminal device provided in an embodiment of the present application;

[0041] Figure 26 A structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0042] The embodiments of the present application are described below with reference to the accompanying drawings.

[0043] Currently, producing videos using dialogue to advance the plot can simplify video production. However, dialogue videos still require control over the virtual characters and the content of their dialogue, making the production of dialogue videos still relatively difficult.

[0044] In order to solve the above technical problems, the embodiments of the present application provide a dialogue video creation method and related devices, which visually split the virtual human video including multiple plot clips according to the dialogue dimension, reduce the operation complexity, reduce the number of interactions in the video editing process, and greatly lower the threshold for video editing.

[0045] The conversation video creation method provided in the embodiment of the present application is based on artificial intelligence (AI). Artificial intelligence is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new type of intelligent machine that can respond in a similar way to human intelligence. Artificial intelligence is to study the design principles and implementation methods of various intelligent machines, so that machines have the functions of perception, reasoning and decision-making.

[0046] Artificial intelligence (AI) technology is a comprehensive discipline encompassing a wide range of fields, encompassing both hardware and software technologies. Foundational AI technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, speech processing, natural language processing, as well as machine learning / deep learning, autonomous driving, and smart transportation.

[0047] In the embodiments of this application, the main artificial intelligence software technologies involved include the above-mentioned speech processing technology, natural language processing technology, machine learning / deep learning, etc. For example, it may involve deep learning in machine learning (ML), including various artificial neural networks (ANN).

[0048] The dialogue video creation method provided by the embodiments of the present application can be implemented by a computer device, which can be a terminal device or a server. The server can be a physical server, a server cluster composed of multiple physical servers, or a distributed system, or a cloud server providing cloud computing services. The terminal device includes but is not limited to a mobile phone, a computer, a smart voice interaction device, a smart home appliance, a vehicle-mounted terminal, and the like. The embodiments of the present application can be applied to various scenarios, including but not limited to audio and video, cloud technology, artificial intelligence, intelligent transportation, and assisted driving. The terminal device and the server can be directly or indirectly connected through wired or wireless communication, which is not limited in the present application.

[0049] The computer device with data processing has machine learning capability. Machine learning is a multi-disciplinary subject, involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory, and other disciplines. It is a subject that studies how computers simulate or implement human learning behavior to acquire new knowledge or skills, reorganize existing knowledge structure to continuously improve performance. Machine learning is the core of artificial intelligence and the fundamental approach to making computers intelligent. It is applied in various fields of artificial intelligence. Machine learning and deep learning usually include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and rule-based learning.

[0050] The computer device with data processing has voice processing technology. The key technologies of speech technology include automatic speech recognition technology, speech synthesis technology, and voiceprint recognition technology. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, and voice is one of the most promising human-computer interaction methods in the future.

[0051] The computer device with data processing has natural language processing (NLP) capability. Natural language processing is an important direction in the field of computer science and artificial intelligence. It studies various theories and methods that enable effective communication between humans and computers using natural language. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, research in this field will involve natural language, i.e., the language used in daily life, so it is closely related to linguistic research. Natural language processing technology usually includes text processing, semantic understanding, machine translation, robot question answering, knowledge graph, and other technologies.

[0052] In a conversation video creation method and related devices provided in an embodiment of the present application, the artificial intelligence model adopted mainly involves the application of natural language processing, the recognition of appearance feature points, etc., and conversation video creation is achieved through natural language processing, thereby simplifying the conversation video creation process, and the customization of actions is achieved through the recognition of appearance feature points, thereby enriching the content of the conversation video.

[0053] With the research and progress of artificial intelligence technology, artificial intelligence technology has been studied and applied in many fields, such as common smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, unmanned driving, automatic driving, drones, robots, smart medical care, smart customer service, Internet of Vehicles, automatic driving, smart transportation, etc. It is believed that with the development of technology, artificial intelligence technology will be applied in more fields and play an increasingly important role.

[0054] It is understandable that in the specific implementation of this application, related data such as user voice information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0055] In order to facilitate understanding of the technical solution provided by this application, the following will introduce a conversation video creation method provided by an embodiment of this application in combination with an actual application scenario.

[0056] See also Figure 1 , Figure 1 A schematic diagram of an application scenario of a method for creating a conversation video provided in an embodiment of the present application. Figure 1 The application scenario shown includes a terminal device 10 and a server 20. The terminal device 10 and the server 20 are connected, and the terminal device 10 is used to realize conversation video creation and conversation video display, or the server 20 is used to create a conversation video according to an operation instruction from the terminal device 10, and the terminal device 10 is used to obtain an operation instruction and display a conversation video.

[0057] The terminal device 10 can display a plot conversation area 1001 for editing a virtual human video. The virtual human video to be edited includes multiple plot segments, which can include a first segment and a second segment. The plot conversation area displays a first plot bubble 11 corresponding to the first segment. The first plot bubble 11 is associated with the first virtual object involved in the first segment. The first plot bubble 11 includes the first plot text of the first virtual object in the first segment. For example, the object identifier 12 of the first virtual object and the first plot bubble 11 are displayed in the same row, reflecting the association between the two. Among them, the first virtual object can be recorded as "Little A", and the first plot text can be "Good morning, Little B", see Figure 1 1a.

[0058] In response to the instruction to add the second segment, a second plot bubble 21 of the second segment is displayed in the plot conversation area. The second plot bubble 21 is associated with the second virtual object involved in the second segment. The second plot bubble 21 includes the second plot text of the second virtual object in the second segment. For example, the object identifier 22 of the second virtual object and the second plot bubble 21 are displayed in the same row, reflecting the association between the two. The second virtual object can be recorded as "Little B" and the second plot text is "Good morning", see Figure 1 1b. The instruction to add the second segment can be generated, for example, by triggering the segment adding control 33, the second plot text can be input, for example, through the plot text input box 32, and the second virtual object can be determined, for example, through the object selection control 31. Figure 1 1a.

[0059] That is, the plot bubbles displayed in the plot conversation area 1001 include plot text corresponding to the plot segments. Different plot bubbles may be associated with the same or different virtual objects. Different plot bubbles may be displayed at different positions in the plot conversation area. The display position relationship between different plot bubbles identifies the plot sequence of the corresponding plot segments in the virtual human video. Therefore, based on the first plot bubble and the second plot bubble displayed in the plot conversation area, video structure parameters for playing the virtual human video when called can be generated. For video structure parameters, see Figure 1 The second segment of the virtual human video can be played in the plot preview area 1002, see Figure 1 The generated video structure parameters can be stored in the terminal device 10 or in the server 20. If the video structure parameters are stored in the server 20, the video structure parameters can be called by multiple terminal devices 10 connected to the server 20, so that the virtual human video can be played on multiple terminal devices 10.

[0060] In this way, a virtual human video including multiple plot clips can be visually split based on the dialogue dimension, so that the plot bubbles in the plot conversation area can intuitively display the dialogue content of different virtual objects like the social conversation interface. The plot sequence of the dialogue content in the virtual human video can be clearly and intuitively sorted out through the display position relationship between the plot bubbles. When the plot of the virtual human video needs to be edited, it can be achieved from the dimension of the plot bubbles in a manner similar to editing social conversations, reducing the complexity of the operation. The display of the plot bubbles in the plot conversation area can promptly and clearly reflect the increase and content of the plot clips in the virtual human video, reducing the number of interactions during the video editing process, and greatly lowering the threshold for video editing.

[0061] Next, a method for creating a conversation video provided in an embodiment of the present application will be introduced in conjunction with the accompanying drawings.

[0062] See also Figure 2 , Figure 2 A signaling diagram of a method for creating a conversational video provided in an embodiment of the present application, the method comprising:

[0063] S101, a plot conversation area for editing a virtual human video is displayed.

[0064] In the embodiments of the present application, unlike traditional video shooting and production methods, virtual human videos refer to videos produced using virtual characters and virtual scenes, and virtual human videos can be produced using terminal devices. The virtual characters can be two-dimensional (2-dimension, 2D) characters or three-dimensional (3-dimension, 3D) characters, and the corresponding virtual scenes can be two-dimensional scenes or three-dimensional scenes. Producing a dialogue video based on a virtual human in the form of dialogue to advance the plot can simplify the production of the video, especially a virtual human video including a three-dimensional character and a three-dimensional scene. Since it is difficult to control details such as the framing of the virtual camera and the movement of the virtual character, producing a virtual human video in the form of a three-dimensional character's dialogue to advance the plot can greatly simplify the production of a dialogue video based on a three-dimensional virtual human.

[0065] In the embodiment of the present application, the virtual human video includes a plot segment in the form of a dialogue. The virtual human video can be edited to meet the needs of the virtual human video. The virtual human video to be edited can include multiple plot segments, and the number of multiple plot segments can be determined according to actual conditions. The following description uses the first and second segments of the multiple plot segments. In practice, the multiple plot segments of the virtual human video to be edited can include only the first and second segments, or can also include other segments in addition to the first and second segments, such as a third segment.

[0066] Specifically, a plot conversation area 1001 for editing a virtual human video may be displayed, so that the virtual human video can be edited through the plot conversation area 1001. The plot conversation area 1001 may display a first plot bubble 11 corresponding to a first segment. The first segment involves a first virtual object. The first plot bubble 11 is associated with the first virtual object and includes the first plot text of the first virtual object in the first segment. Figure 3 , which is a schematic diagram of a plot conversation area provided in an embodiment of the present application. The appearance of the first plot bubble 11 can be determined based on needs or can be set by the user. The appearance of the first plot bubble 11 can be determined based on the first virtual object. The appearance of the first plot bubble 11 includes the color, shape, and decorative elements of the first plot bubble 11.

[0067] The object identifier 12 of the first virtual object can be displayed in the plot conversation area 1001. The display position relationship between the object identifier 12 of the first virtual object and the first plot bubble 11 is used to reflect the association relationship between the first plot bubble 11 and the first virtual object. For example, the object identifier 12 of the first virtual object and the first plot bubble 11 are displayed in the same row. The first virtual object can be recorded as "Little A" and the first plot text can be "Good morning, Little B". Figure 1 1a and Figure 3 .

[0068] The plot conversation area 1001 can be displayed according to the conversation setting operation, which can be generated by triggering the conversation setting control or by a specific gesture. Figure 4 , is a schematic diagram of an operation interface provided by an embodiment of the present application, wherein the operation interface 1003 includes a dialogue setting control 35, and the display position of the dialogue setting control 35 can be set according to actual conditions, for example, it can be the lower right corner of the display screen. The dialogue setting control 35 can be displayed based on the dialogue creation instruction, see Figure 4 The dialogue creation instruction can be generated by selecting the dialogue creation control 34. When the dialogue creation control 34 is selected, the dialogue setting control 35 is displayed. When the dialogue setting control 35 is displayed, the dialogue creation control 34 can be undisplayed or displayed at the same time as the dialogue setting control 35. For example, the dialogue setting control 35 is located in the lower right corner of the display screen, and the dialogue creation control 34 is located in the lower left corner of the display screen, and both are displayed simultaneously in the same operation interface 1003.

[0069] S102 : In response to an instruction to add a second segment among the plurality of segment stories, display a second scenario bubble corresponding to the second segment in the scenario conversation area.

[0070] In an embodiment of the present application, an instruction to add a second segment among multiple segment plots may be obtained. In response to the instruction to add, a second segment plot bubble 21 corresponding to the second segment may be displayed in the segment conversation area. The second segment involves a second virtual object. The second virtual object and the first virtual object may be the same virtual object or different virtual objects. The second segment plot bubble 21 is associated with the second virtual object. The second segment plot bubble 21 includes the second segment plot text of the second virtual object in the second segment.

[0071] Among them, the instruction to add the second fragment can be generated based on the fragment adding operation for the second fragment. The fragment adding operation for the second fragment can be a triggering operation of the fragment adding control 33, or a specific gesture, or a confirmation operation of the last fragment information in the instruction to add the second fragment. The fragment information in the adding instruction includes at least the second plot text and the unique identifier of the second virtual object.

[0072] Specifically, the second plot text can be obtained through the plot text input box 32. The plot text input box 32 can be input with text. The second plot text can be determined based on the text input in the plot text input box 32. In a specific implementation, the second plot text can be the text in the plot text input box 32 when the clip adding operation is obtained. In particular, when no text is input in the plot text input box 32 and the operation is not triggered, a prompt "Type plot text here" can be displayed in the plot text input box, see Figure 3 When the plot text input box is triggered, the keyboard area 36 may be displayed to display the entered text in the plot text input box 32 according to the user's triggering operation on the keyboard area 36, ​​see Figure 5 , is a schematic diagram of another scenario conversation area provided in an embodiment of the present application, in which a keyboard area 36 is displayed below a scenario text input box 32; alternatively, when the scenario text input box 32 is triggered, a voice receiving control may be displayed to obtain a target voice from the user, and to obtain recognized text by performing voice recognition on the target voice. The recognized text may be displayed in the scenario text input box 32 as the input text in the scenario text input box 32.

[0073] Specifically, the second virtual object can be determined by selecting a target object identifier from the displayed object identifiers. In response to the selection of the target object identifier, the object corresponding to the target object identifier is used as the second virtual object, and a unique identifier for the second virtual object is obtained. The unique identifier of the second virtual object can be included in the add instruction. The unique identifier of the second virtual object is an identifier that can uniquely identify the second virtual object. It can be the number of the second virtual object in the virtual object library, or the name of the second virtual object, etc. The displayed object identifier can be a pattern corresponding to the virtual object, and can reflect the characteristics of the virtual object, such as gender characteristics, image characteristics, etc. The object identifier can be displayed on an object setting interface, and the object setting interface can be displayed in response to an object selection operation. The object selection operation can be a trigger operation on the object selection control 31, or a specific gesture.

[0074] The appearance of the second plot bubble can be determined based on demand or can be set by the user. The appearance of the second plot bubble can be the same as or different from the appearance of the first plot bubble. The appearance of the second plot bubble includes the color, shape, and decorative elements of the second plot bubble. The appearance of the second plot bubble can be determined based on the second virtual object. When the first virtual object and the second virtual object are different virtual objects, the first plot bubble and the second plot bubble can have different appearances, so as to more intuitively distinguish the plot text of the first virtual object from the plot text of the second virtual object. Figure 1 1b, the appearance of the second plot bubble 21 is the same as the appearance of the first plot bubble 11.

[0075] The object identifier of the second virtual object can be displayed in the plot conversation area. The relationship between the object identifier of the second virtual object and the display position of the second plot bubble is used to reflect the association relationship between the second plot bubble and the second virtual object. For example, see Figure 1 In 1b, the object identifier 22 of the second virtual object and the second story bubble 21 are displayed in the same row. The second virtual object can be recorded as "Little B" and the second story text is "Good morning".

[0076] In an embodiment of the present application, the first plot bubble and the second plot bubble have a display position relationship in the plot conversation area, and the display position relationship is used to identify the plot order of the first segment and the second segment in the virtual human video Zhang Hong. That is to say, the plot bubbles displayed in the plot conversation area include plot texts corresponding to the plot segments, and different plot bubbles can have an association relationship with the same or different virtual objects. In addition, different plot bubbles can be displayed in different positions in the plot conversation area, and the plot order of the corresponding plot segments in the virtual human video is identified by the display position relationship between different plot bubbles. In this way, the virtual human video including multiple plot segments is visually split according to the dialogue dimension, so that the plot bubbles in the plot conversation area can intuitively display the dialogue content of different virtual objects like the social conversation interface. The display position relationship between the plot bubbles can clearly and intuitively sort out the plot order of the dialogue content in the virtual human video, which greatly reduces the threshold for video editing.

[0077] After the second plot bubble is displayed in the plot conversation area, in response to a preview operation on the second plot bubble, the substructure parameters corresponding to the second plot bubble are determined based on the second virtual object and the second plot text. A second segment is generated based on the substructure parameters, and the second segment is displayed in the plot preview area outside the plot conversation area. In the second segment, the second virtual object speaks the second plot text, so that the user can determine whether the second virtual object and the second plot text are set reasonably based on whether the second segment displayed in the plot preview area meets the requirements. The plot preview area can be displayed simultaneously with the plot conversation area, see Figure 6 , is a schematic diagram of a display interface provided by an embodiment of the present application, wherein the plot conversation area 1001 is located on the right side of the screen, and the plot preview area 1002 is located on the left side of the screen; the plot preview area 1002 can also be displayed in the form of a new interface in the display area where the plot conversation area 1001 is located. When the plot preview area 1002 is displayed, the plot conversation area 1001 can be canceled or covered by the plot preview area 1002. Figure 1 In 1d, the plot session area 1001 is canceled.

[0078] After the second plot bubble is displayed in the plot conversation area, the display position relationship between the second plot bubble and the first plot bubble can be changed in response to a position adjustment operation on the second plot bubble. The adjustment operation on the second plot bubble can be a drag operation on the second plot bubble or a triggering operation on a position adjustment control for the second plot bubble. The position adjustment control can include an upward position control or a downward position control. For example, triggering the upward position control of the second plot bubble controls the display position of the second plot bubble to move upward, while triggering the downward position control of the second plot bubble controls the display position of the second plot bubble to move downward.

[0079] See also Figure 7 , which is a schematic diagram of another plot conversation area provided in an embodiment of the present application, in which the display position of the second plot bubble 21 can be adjusted by dragging the second plot bubble 21 upward, changing the display position relationship between the second plot bubble 21 and the first plot bubble 11, so that the second plot bubble 21 is located above the first plot bubble 11, indicating that the plot segment corresponding to the second plot bubble 21 can be played before the plot segment corresponding to the first plot bubble 11. Of course, after the display position relationship between the second plot bubble 21 and the first plot bubble 11 is changed, the display position relationship between the object identifier 22 of the second virtual object and the object identifier 12 of the first virtual object is also changed accordingly.

[0080] After the second plot bubble is displayed in the plot conversation area, a third plot bubble may be determined to be adjacent to the second plot bubble in the plot conversation area. The third plot bubble may be the first plot bubble or another plot bubble other than the first plot bubble. For example, the third plot bubble may correspond to the third segment of multiple plot segments. When the third plot bubble is the first plot bubble, the third segment is the first segment. In response to the playback order adjustment operation for the second and third plot bubbles, the playback order of the plot segments corresponding to the second and third plot bubbles may be set to be played simultaneously, or the playback order of the plot segments corresponding to the second and third plot bubbles may be set to be played sequentially.

[0081] The play order adjustment operation for the second and third plot bubbles can be a triggering operation for a play order adjustment control or a specific gesture. The play order adjustment control can be associated with the second plot bubble. For example, by selecting the second plot bubble, the play order adjustment control can be displayed, and the play order adjustment control is associated with the second plot bubble. Alternatively, the play order adjustment control can be associated with the third plot bubble. For example, by selecting the third plot bubble, the play order adjustment control can be displayed, and the play order adjustment control is associated with the third plot bubble.

[0082] Specifically, the play order adjustment control may include two options: simultaneous play and sequential play. The selection of simultaneous play or sequential play determines the play order of the second and third plot bubbles. Alternatively, if the play order of the second and third plot bubbles is sequential play by default, the play order adjustment control may include a simultaneous play option. When the simultaneous play option is selected, the play order of the plot segments corresponding to the second and third plot bubbles is determined to be simultaneous. When the simultaneous play option is not selected, the play order of the plot segments corresponding to the second and third plot bubbles is determined to be sequential.

[0083] The sequential play option may indicate the identifiers of the plot bubbles adjacent to the plot bubble corresponding to the play order adjustment control, as well as the order in which the adjacent plot bubbles are played, such as before or after the adjacent plot bubble. For example, if the plot bubble corresponding to the play order adjustment control is the second plot bubble, the second plot bubble may be indicated to be before or after the third plot bubble. Similarly, the simultaneous play option may indicate the identifiers of the plot bubbles that are played simultaneously with the plot bubble corresponding to the play order adjustment control.

[0084] See also Figure 8 , which is a schematic diagram of another plot conversation area provided by an embodiment of the present application. The first plot bubble 11 serves as the third plot bubble, and the second plot bubble is located below the third plot bubble. According to the playback order determined by the display position relationship, the second segment can be played after the plot segment corresponding to the third plot bubble (i.e., the first segment), or it can be played simultaneously with the plot segment corresponding to the third plot bubble (i.e., the first segment). By selecting the second plot bubble 21, a playback order adjustment control 37 can be displayed. The playback order adjustment control 37 can include two options: simultaneous playback and sequential playback. By selecting the simultaneous playback option, the playback order of the plot segments (i.e., the first segment) corresponding to the second plot bubble 21 and the third plot bubble can be set to simultaneous playback. The plot conversation area can also display a unique identifier 39 of the plot bubble, see Figure 8 The unique identifier of the first plot bubble is 2, the unique identifier of the second plot bubble is 3, the simultaneous play option in the play order adjustment control 37 indicates that the adjacent plot bubbles and the play order are "with 2", indicating that the play order of the plot segments corresponding to the third plot bubble and the second plot bubble 21 are played simultaneously, and the sequential play option in the play order adjustment control 37 indicates that the adjacent plot bubbles and the play order are "after 2", indicating that the plot segment corresponding to the third plot bubble is played after the second segment.

[0085] After the second plot bubble is displayed in the plot conversation area, a playback delay duration for the second segment can be determined in response to a delay adjustment operation on the second plot bubble. The playback delay duration indicates that when the playback order of the avatar video reaches the second segment, the second segment will be played after the playback delay duration. Specifically, when the playback order of the avatar video reaches the second segment, it indicates that the plot segments before the second segment have been played.

[0086] The second segment can be played alone or simultaneously with other plot segments. In the scenario where the second segment is played after the third segment corresponding to the third plot segment, when the playback order of the virtual human video reaches the second segment, it is the end playback time of the third segment, and the playback delay duration is the time interval between the start playback time of the second segment and the end playback time of the third segment; correspondingly, in the scenario where the second and third segments are played simultaneously, when the playback order of the virtual human video reaches the second segment, it is the start playback time of the third segment, and the playback delay duration is the time interval between the start playback time of the second segment and the start playback time of the third segment.

[0087] The delay adjustment operation for the second plot bubble can be a triggering operation for the delay adjustment control or a specific gesture. The delay adjustment control is associated with the second plot bubble. For example, by selecting the second plot bubble, the delay adjustment control can be displayed, and the play order adjustment control is associated with the second plot bubble. The delay adjustment control and the play order adjustment control can be displayed at the same time, or after the play order is determined by the play order adjustment control. When displayed at the same time, they can be switched by reducing the interface.

[0088] Specifically, the delay time adjustment control may include a delay time increase control and a delay time decrease control, and may also display the current delay time, so that the user can trigger the delay time increase control or the delay time decrease control according to the current delay time. For example, see Figure 8 The delay duration adjustment control 38 is displayed on the right side of the screen, where the "+" button increases the delay duration and the "-" button decreases the delay duration. If the current delay duration is 0s and the second plot bubble is after the third plot bubble, and the delay duration is 0s, playback of the second plot segment will begin at the moment the plot segment corresponding to the third plot bubble ends. Alternatively, the delay duration adjustment control may include a duration input box, or the delay duration adjustment control may include a delay duration axis, where the delay duration is determined by selecting a value on the delay duration axis.

[0089] The play order adjustment controls and delay time adjustment controls can be displayed in the sequence setting interface. The sequence setting interface displays the second plot bubble and the third plot bubble. By triggering the second plot bubble or the third plot bubble in the sequence setting interface, the play order adjustment controls and delay time adjustment controls can be displayed. Figure 8 The sequence setting interface 1004 displays the second plot bubble 21 and the first plot bubble 11 as the third plot bubble, as well as a play order adjustment control 37 and a delay time adjustment control 38. The sequence setting interface 1004 also includes a sequence setting completion control 103 for canceling the display of the sequence setting interface 1004 when triggered. In specific implementations, the sequence setting interface 1004 can be displayed by triggering the sequence setting control 102. Figure 9 , which is a schematic diagram of another plot conversation area provided in an embodiment of the present application. Sequence setting control 102 is used to display a sequence setting interface when triggered. The sequence setting interface is used to adjust the playback order, allowing the user to intuitively understand the available operations within the sequence setting interface. Of course, in other embodiments of the present application, the sequence setting interface may not be displayed. Instead, the playback order adjustment controls and delay duration adjustment controls may be displayed by triggering the second or third plot bubble in the plot conversation area.

[0090] In summary, when it is necessary to edit the plot of a virtual human video, it can be achieved from the dimension of plot bubbles in a manner similar to editing social conversations. For example, the playback order and playback timing of the plot clips can be adjusted to reduce the complexity of the operation. The display of plot bubbles in the plot conversation area can promptly and clearly reflect the addition of plot clips and the content of the plot clips in the virtual human video, reducing the number of interactions during the video editing process and greatly lowering the threshold for video editing.

[0091] S103 , in response to the video generation instruction, generating video structure parameters according to the first plot bubble and the second plot bubble displayed in the plot conversation area, wherein the video structure parameters are used to play the virtual human video when called.

[0092] In an embodiment of the present application, in response to a video generation instruction, video structure parameters can be generated based on the first plot bubble and the second plot bubble displayed in the plot conversation area. The video structure parameters are used to play the virtual human video when called. Specifically, the video structure parameters include substructure parameters corresponding to each plot bubble. The substructure parameters are the various parameters required for playing the plot segment corresponding to the plot bubble. The substructure parameters may include parameters of the virtual object and parameters of the plot text. In the playback state of the virtual human video, the real-time playback of the virtual human video can be performed by calling the various substructure parameters in the video structure parameters. In the substructure parameters corresponding to the first plot bubble, the parameters of the virtual object are the unique identifier of the first virtual object, and the parameters of the plot text are the first plot text. In the substructure parameters corresponding to the second plot bubble, the parameters of the virtual object are the unique identifier of the second virtual object, and the parameters of the plot text are the second plot text. If the plot segment corresponding to the plot bubble has been previewed, the substructure parameters corresponding to the plot bubble can be generated during the preview, and the video structure parameters can be generated using the substructure parameters. For example, see Figure 1 In 1c, the unique identifier of the first virtual object is "Xiao A", and the first plot text is "Good morning, Xiao B". The unique identifier of the second virtual object is "Xiao B", and the second plot text is "Good morning".

[0093] In response to the video generation instruction, the virtual human video identifier associated with the video structure parameters can be displayed. In response to the trigger operation on the virtual human video identifier, the video structure parameters can be called to realize the playback of the virtual human video. Specifically, the sub-structure parameters corresponding to each plot bubble can be called in turn, so as to play the plot clips corresponding to each plot bubble in turn.

[0094] The video generation instruction can be generated by triggering the video generation control, or by obtaining a specific gesture. Figure 3 、 Figure 5-Figure 7 、 Figure 8 The video generation control 101 may be located in the upper right corner of the plot conversation area. In response to a trigger operation on the video generation control 101 , the display of the plot conversation area 1001 may be canceled.

[0095] Specifically, the substructure parameters corresponding to the plot bubble may also include voice information, that is, the video structure parameters include voice information, and the voice information may be generated based on the second plot text and audio parameters. The audio parameters may include pitch and timbre. After obtaining the second plot text, the audio parameters may be determined for the second plot text. The audio parameters may be determined based on at least one of the target voice or the second virtual object. Among them, the target voice is the voice obtained when obtaining the second plot text through the plot text input box. The recognized text may be obtained by performing voice recognition on the target voice. The recognized text is displayed in the plot text input box as the input text in the plot text input box. Generating voice information based on the second plot text and audio parameters can be achieved through AI. The same virtual object may have the same timbre in different plot segments, and may have the same pitch or different pitches. By setting the pitch and timbre, users can be assisted in creating various personalized video outputs.

[0096] When the audio parameters are determined based on the target voice, pitch and timbre can be extracted based on the target voice, and voice information can be generated for the second plot text based on the extracted pitch and timbre. Of course, the target voice can also be directly used as the voice information. When the target voice is not obtained, or when the target voice is obtained, the audio parameters can be determined based on the second virtual object. For example, the audio parameters corresponding to the second plot text are determined based on the gender and personality settings of the second virtual object. When the audio parameters are determined based on the target voice and the second virtual object, the audio data corresponding to the second plot text can be first determined based on the gender and personality settings of the second virtual object, and then the audio data is combined with the audio data extracted from the target voice to obtain the audio parameters corresponding to the second plot text. The audio data includes pitch and timbre, so that the sound of the second virtual object is associated with the target voice, thereby improving the user experience.

[0097] See also Figure 10 , which is a schematic diagram of a video structure parameter provided in an embodiment of the present application, the substructure parameter corresponding to the first plot bubble also includes voice information, and the voice information has female characteristics. The substructure parameter corresponding to the second plot bubble also includes voice information, and the voice information has male characteristics.

[0098] Specifically, the substructure parameters corresponding to the plot bubble may also include subtitle information, that is, the video structure parameters include subtitle information, which may be generated based on the second plot text and subtitle parameters. The subtitle parameters may include font, font size, display time period, etc. The display time period of the subtitle information is related to the playback time period of the voice information. Figure 10 The substructure parameters corresponding to the first plot bubble and the second plot bubble respectively also include their corresponding subtitle information.

[0099] Specifically, the sub-structure parameters corresponding to the plot bubble may further include a play order parameter, which indicates whether the adjacent plot bubbles of the plot bubble and the plot segment corresponding to the plot bubble are played before or after the plot segment corresponding to the adjacent plot bubbles; the sub-structure parameters corresponding to the plot bubble may further include a play delay duration, which indicates that when the play order reaches the plot segment corresponding to the plot bubble, the plot segment will be played after the play delay duration. Figure 10 The substructure parameters corresponding to the first plot bubble also include a playback order parameter and a playback delay parameter. The playback order parameter is after the fifth segment, and the playback delay parameter is 0s. The substructure parameters corresponding to the second plot bubble also include a playback order parameter and a playback delay parameter. The playback order parameter is after the first segment, and the playback delay parameter is 0s.

[0100] In addition, plot bubbles corresponding to other plot segments besides the first and second plot bubbles can be added to enrich the content of the plot segments. For example, the other plot segments can be camera movement segments. In this case, the plot bubble corresponding to the plot segment serves as the camera movement bubble. The substructure parameters corresponding to the plot segment include virtual camera parameters, which may include camera movement time, camera movement position, camera movement curve, etc. The virtual camera parameters can be set by triggering the camera movement bubble, for example, by triggering the camera movement bubble to display a virtual camera setting control, and then triggering the virtual camera setting control to set the virtual camera parameters.

[0101] See also Figure 8 and Figure 9 The plot conversation area may further display a fourth plot bubble, which corresponds to a fourth segment among the multiple plot segments, and the fourth segment involves a fourth virtual object. An object identifier 42 of the fourth virtual object is also displayed in the plot conversation area, indicating the association between the fourth virtual object and the fourth plot bubble 41. The fourth plot bubble 41 contains the fourth plot text "I encountered a happy thing", the fourth virtual object and the first virtual object are the same virtual object, and the unique identifier of the fourth plot bubble can be recorded as 4; the plot conversation area may further display a fifth plot bubble 51, which corresponds to the fifth segment, which is a camera movement segment, and the virtual object involved in the fifth segment is the fifth virtual object. The object identifier 52 of the fifth virtual object is a virtual camera, and the unique identifier of the fifth plot bubble 51 is recorded as 1.

[0102] See also Figure 10In the substructure parameters corresponding to the fourth plot bubble, the virtual object parameter is the unique identifier of the fourth virtual object, such as "Xiao A", and the plot text parameter is the fourth plot text, such as "I encountered a happy thing." The substructure parameters corresponding to the fourth plot bubble also include voice information, which has female characteristics. The substructure parameters of the fourth plot bubble also include subtitle information. The substructure parameters corresponding to the fourth plot bubble also include playback order parameters and playback delay parameters. The playback order parameter is after the second segment, and the playback delay parameter is 0s. The substructure parameters corresponding to the fifth plot bubble include virtual camera parameters, which can include camera movement time, camera movement point, camera movement curve, etc.

[0103] In summary, a virtual human video including multiple plot clips is visually split based on the dialogue dimension, so that the plot bubbles in the plot conversation area can intuitively display the dialogue content of different virtual objects like a social conversation user interface (UI). The plot sequence of the dialogue content in the virtual human video can be clearly and intuitively sorted out through the display position relationship between the plot bubbles. When the plot of the virtual human video needs to be edited, it can be implemented from the dimension of the plot bubbles in a manner similar to editing social conversations, reducing the complexity of the operation. The display of the plot bubbles in the plot conversation area can timely and clearly reflect the increase and content of the plot clips in the virtual human video, reducing the number of interactions during the video editing process, and greatly lowering the threshold for video editing.

[0104] In an embodiment of the present application, the adding instruction may further include action parameters, which may indicate the target action of the second virtual object in the second plot segment. The action of the virtual object may be at least one of a facial expression action or a body movement. The target action of the second virtual object in the second plot segment may be at least one of a target facial expression action or a target body movement. The target facial expression action may be, for example, laughing, crying, sad, etc., and the target body movement may be, for example, waving, walking, jumping, etc., so that the plot dialogue is associated with the action and the freedom and flexibility of the plot dialogue is improved. When the target action is a target body movement, the action parameters may include the action starting point, the action end point, and the action execution location, etc. The target facial expression action and the target body movement are executed through different locations, and therefore may be executed simultaneously, such as waving while laughing, walking while crying, etc. The following introduces S102 in conjunction with the adding instruction including the action parameters.

[0105] When the add instruction further includes action parameters, as a method of reflecting the target action, the target action can be reflected in a second plot bubble, and a second plot bubble corresponding to the second segment can be displayed in the plot conversation area. Specifically, based on the action parameters in the add instruction, a target action identifier and a target position of the target action identifier in the second plot text can be determined, and based on the target action identifier and the target position, a second plot bubble including the second plot text and the target action identifier can be displayed in the plot conversation area.

[0106] The target action identifier is used to instruct the second virtual object to speak the target text in the second plot text through the target action corresponding to the target action identifier, and the target text is determined based on the target position. The target action identifier can be one or more. When the target action identifier is one, the target text is the second plot text, that is, the second virtual object speaks the second plot text through the target action, and the target action identifier can be located at the beginning, end, or middle of the second plot text. When there are multiple target action identifiers, the target text corresponding to the target action identifier can be determined based on the target position of the target action identifier.

[0107] See also Figure 11 , which is a schematic diagram of another scenario conversation area provided by an embodiment of the present application. A second scenario bubble 21 may display the second scenario text "Good morning" and a target action identifier "(smile)", where the target action identifier is located at the end of the second scenario text. A first scenario bubble 11 may display the first scenario text "Good morning, Little B", and target action identifiers "(smile)" and "(wave)", where the target action identifiers are located in the middle and at the end of the first scenario text, respectively. If the second scenario text contains multiple target action identifiers, the multiple target action identifiers set in the first scenario text may be referenced.

[0108] After determining the target action identifier based on the action parameters in the added instruction, the mapping relationship between the action feature point parameters of the target action corresponding to the target action identifier and the appearance feature points of the second virtual object can be determined based on the target action identifier. In the process of calling the video structure parameters to play the virtual human video, the mapping relationship is used to instruct the second virtual human to perform the target action, thereby realizing action reuse. There is no need to draw actions for different virtual objects one by one, reducing implementation costs. The action feature point parameters of the target action can be determined based on the appearance feature point position corresponding to the target action and the temporal change of the appearance feature point position. The appearance feature point can be a skeletal feature point. When the target action is a target facial expression action, the appearance feature point is a facial skeletal feature point. When the target action is a target limb action, the appearance feature point is a skeletal feature point of the action execution site.

[0109] Specifically, when there are multiple target action identifiers, the target action identifiers may include a first action identifier and a second action identifier. The first action identifier is located at the first target position of the second plot text, and the second action identifier is located at the second target position of the second plot text. In the second plot text, the first target position is before the second target position. The first target text corresponding to the first action identifier is determined based on the text before the first target position in the second plot text, and the second target text corresponding to the second action identifier is determined based on the text between the first target position and the second target position in the second plot text. In other words, when the second plot bubble includes the first action identifier and the second action identifier, the second virtual object can be instructed to speak the first target text using the first action corresponding to the first action identifier and to speak the second target text using the second action corresponding to the second action identifier. A single plot bubble can be used to instruct the second virtual object to switch actions while speaking the second plot text, thus enriching the configuration of the second virtual object and better meeting actual needs.

[0110] It should be noted that if the first action identifier and the second action identifier correspond to the target facial expression and the target body movement respectively, then the first target position and the second target position may be adjacent positions, there may be no text between the two, and the first action identifier and the second action identifier may correspond to the same target text. That is to say, when it is determined that the text between the second target position and the second target position is empty, it can be determined that the first action identifier and the second action identifier may correspond to the same target text, and the target text is determined based on the text before the first target position in the second plot text. For example, if the second plot bubble displays "Good morning (smile) (wave)", it means that the first action identifier is "(smile)" and the second action identifier is "(wave)". The text between the first action identifier and the second action identifier is empty, and the two correspond to the same target text. The action of the second virtual object in the virtual human video is reflected as the second virtual object smiling, waving, and saying "Good morning".

[0111] If the target action includes a target body movement, the execution parameters of the target body movement can also be determined according to the length of the second plot text. The action parameters also include execution parameters of the target body movement. The execution parameters may include at least one of the number of executions or the execution speed, so that the target body movement matches the second plot text, and the execution duration of the target body movement is close to the duration of the voice information corresponding to the second plot text, avoiding the problem of poor video quality caused by the mismatch between the action duration and the voice duration, such as the action in a silent state, rhythm jamming or animation interruption caused by short voice duration, or the problem of voice broadcast without action caused by long voice duration.

[0112] Among them, when the target body movement is an end-to-end action, the number of executions can be positively correlated with the length of the second plot text. When the second plot text is longer, the number of executions of the target body movement is more, and when the second plot text is shorter, the number of executions of the target body movement is less, which can better match the second plot text. For example, when the second plot text is long, the number of waving times can be set to multiple times to match the second plot text; the execution speed can be inversely correlated with the length of the second plot text. When the second plot text is long, the execution speed of the target body movement is slower, and when the second plot text is shorter, the execution speed of the target body movement is faster, which can also better match the second plot text. For example, when the second plot text is long, the waving speed can be set to be slower to match the second plot text.

[0113] If the target action includes a target facial expression action, the execution parameters of the target facial expression action can also be determined according to the length of the second plot text. The action parameters also include the execution parameters. The execution parameters may include at least one of the number of executions or the maintenance duration, so as to better match the second plot text, so that the execution duration of the target body action is close to the duration of the voice information corresponding to the second plot text, avoiding the problem of poor video quality caused by the mismatch between the action duration and the voice duration, such as the action in a silent state, rhythm jamming or animation interruption caused by short voice time, or the problem of voice broadcast without action caused by long voice time.

[0114] Among them, when the target facial expression action is an action connected from beginning to end (such as speaking, etc.), the number of executions can be positively correlated with the length of the second plot text. When the second plot text is longer, the number of executions of the target facial expression action is more, and when the second plot text is shorter, the number of executions of the target facial expression action is less, which can better match the second plot text; the maintenance time can be positively correlated with the length of the second plot text. When the second plot text is longer, the maintenance time of the target body movement is longer, and when the second plot text is shorter, the maintenance time of the target body movement is shorter, which can also better match the second plot text. For example, when the second plot text is long, the duration of the smile can be set to be consistent with the duration of the voice information corresponding to the second plot text, so as to match the second plot text.

[0115] Of course, in other embodiments of the present application, the execution parameters of the target action may be independent of the length of the second plot text. In this case, regardless of the length of the second plot text, the execution parameters of the target action remain constant. For example, if the second plot text is long, only a waving action or a smiling action may be performed at the beginning of the second plot text being spoken, and no action is performed thereafter.

[0116] Specifically, the target position can be determined based on the selection of the action adding position in the second plot text. The selection of the action adding position can be performed during the input process of the second plot text, or after the second plot text is input. After determining the action adding position, the target action identifier to be added to the action adding position in the second plot text can be determined based on the action selection operation from the action identifiers displayed in the action selection area, and then the action parameters are generated based on the target action identifier and the target position. In response to the segment adding operation for the second segment, the adding instruction for the second segment is generated based on the action parameters. For the relevant description of the segment adding operation, please refer to the description in S102 and will not be repeated here. In this way, the dialogue plot is divided into multiple segments. By individually regulating the content corresponding to each segment, the dialogue plot can be controlled, giving the dialogue plot a higher degree of freedom and controllability.

[0117] As a possible implementation, when the second plot text is determined through the plot text input box, the selection of the action addition location can be determined through an action addition operation. The action addition operation is for the action addition location of the text already entered in the plot text input box. The action addition operation can include a triggering operation for location selection and a triggering operation for action addition. The location selection triggering operation can be, for example, a single-clicking operation at the action addition location where the text has been entered. The action addition triggering operation can be, for example, a triggering operation for action addition control. The action addition operation can also include a specific gesture or triggering operation at the action addition location, such as a double-clicking operation at the action addition location. In response to the action addition operation, an action selection area can be displayed so that an action identifier in the action selection area can be selected.

[0118] See also Figure 12 , is a schematic diagram of another plot conversation area provided by an embodiment of the present application. After the action adding position is determined to be "good morning" by the trigger operation of the position selection, the action adding control 61 can be triggered to display the action selection area. Figure 13 , which is a schematic diagram of another plot conversation area provided in an embodiment of the present application, wherein the action selection area 62 can display multiple action identifiers, and the target action identifier can be determined by triggering the action identifier. The selected action identifier in the action selection area can be highlighted, and when a segment adding operation for a second segment is obtained, the selected action identifier is used as the target action identifier. For example, the action identifier corresponding to a smile is highlighted and used as the target action identifier when a segment adding operation for the second segment is obtained.

[0119] From the user's side, the addition of the second segment can be achieved through the following operations: calling out the object setting interface through the object selection operation, determining the second virtual object by selecting the target object identifier in the object setting interface, entering text in the plot text input box as the input text, and selecting the action addition position in the input text to determine the action addition position. Then, a trigger operation for adding an action can be performed, thereby calling out the action selection area, and selecting the target action identifier in the action identifier in the action selection area to achieve the selection of the target action, and then performing a segment addition operation to generate an addition instruction for the second segment. The addition instruction includes the unique identifier of the second virtual object, the second plot text, and the action parameters. The input text is used as the second plot text, and the action parameters can be determined based on the action addition position and the target action identifier. This operation is similar to chatting or playing games, which greatly reduces the threshold for video editing.

[0120] As another possible implementation, when determining the second plot text via the plot text input box, the action addition position can be determined as follows: semantic analysis is performed on the target text segment within the text already entered in the plot text input box to obtain the target semantics, a recommended action corresponding to the target semantics is determined, and the action addition position corresponding to the action identifier of the recommended action is determined to be the position after the target text segment. After the recommended action is determined, the action identifier of the recommended action can be displayed in the action selection area so that the action identifier in the action selection area can be selected. The recommended action can be one or more, and the action selection area can be located above or below the plot text input box.

[0121] See also Figure 14 , a schematic diagram of another scenario conversation area provided in an embodiment of the present application, wherein semantic analysis is performed on the target text segment "good morning" in the input text to obtain the target semantics, and the recommended action corresponding to the target semantics is determined to be smiling. After the action is added to the location "good morning", the action identifier of the recommended action is displayed in the action selection area. When a selection operation is obtained for the action identifier, the action identifier can be used as the target action identifier. The action selection area 62 is located above the scenario text input box 32.

[0122] From the user side, the addition of the second segment can be achieved through the following operations: calling out the object setting interface through the object selection operation, determining the second virtual object by selecting the target object identifier in the object setting interface, entering text in the plot text input box as the input text, entering text in the plot text input box as the input text, performing semantic analysis on the target text segment in the input text and displaying the action identifier of the recommended action, then determining the target action by selecting the action identifier of the recommended action, and then performing the segment addition operation to generate the addition instruction of the second segment, the addition instruction including the unique identifier of the second virtual object, the second plot text and the action parameters, the input text as the second plot text, and the action parameters can be determined according to the action addition position and the target action identifier. This operation is similar to chatting or playing games, which greatly reduces the threshold for video editing.

[0123] The action identifiers displayed in the action selection area are optional action identifiers. The action identifiers displayed in the action selection area are generated based on the action library. In addition, the action library can be enriched in a customized form to meet the needs of more users. Specifically, the action video can be obtained by collecting the physical object that performs the custom action. According to the position of the shape feature points of the physical object in the action video and the temporal change of the shape feature point position, the action feature point parameters of the custom action are extracted, and the action identifier of the custom action is displayed in the action selection area according to the action feature point parameters. Among them, when the target action is a target facial expression action, the shape feature points are facial bone feature points, and when the target action is a target limb action, the shape feature points are bone feature points of the action execution part. The shape feature points can be identified by AI technology, and the offset of the shape feature points can be reflected as the bone offset or deformation of the physical object, which can be used to reflect the bone offset or deformation of the virtual object.

[0124] The collection of entity objects can be triggered by creating an action. The creating action can be a triggering operation of creating an action control or a specific gesture. The creating action control can be displayed in the action selection area. In this way, by triggering the creating action control in the action selection area, the newly created action identifier will be displayed in the action selection area, which is convenient for selecting the action identifier. The collection of entity objects can be controlled through the action collection interface. Figure 13 , the new action control 63 is located in the action selection area 62.

[0125] Specifically, before the collection begins, the action collection interface may include a collection start control and an action display area. The collection start control is used to start the collection of physical objects when triggered, and the action display area is used to display the custom actions of the collected physical objects. In addition, the action collection interface may also display an action preview area. The action preview area is used to display the custom actions of virtual objects generated in real time based on the custom actions of the actual objects, so that the reuse effect of the custom actions can be known through the action preview area. The virtual object can be a second virtual object or a general virtual object in the action library. The action collection interface also includes a completion control and a cancel control. The completion control is used to generate an action identifier based on the custom action that has been collected when triggered, and the cancel control is used to cancel the action collection and cancel the display of the action collection interface when triggered.

[0126] See also Figure 15 , is a schematic diagram of an action collection interface provided in an embodiment of the present application. The action collection interface 1005 includes a collection start control 73 and an action display area 74. The action collection interface may also include an action preview area 79. The action display area 74 displays the custom action of the physical object, such as a smiling action. The action preview area 79 displays the custom action of the virtual object generated in real time according to the action of the physical object, such as a smiling action of the first virtual object. The completion control 72 and the cancel control 71 are located above the action collection interface 1005.

[0127] During the acquisition process of the physical object, the action acquisition interface can display the acquisition stop control and the action display area. The acquisition stop control is used to stop the acquisition of the physical object and obtain the action video when it is triggered. In addition, the action acquisition interface can also display the action preview area. Figure 16 , which is a schematic diagram of another action collection interface provided in an embodiment of the present application. The action collection interface 1005 includes a collection stop control 75 and an action display area 74. The action collection interface may also include an action preview area 79. The action display area 74 displays custom actions of the physical object, such as a smile action and a scissor hand action. The action preview area 79 displays custom actions of the virtual object generated in real time according to the action of the physical object, such as a smile action and a scissor hand action of the first virtual object.

[0128] After the capture stops, the action capture interface can display the capture preview control, which is used to preview the action video when triggered. The action capture interface can also display the action preview area, see Figure 17As another embodiment of the action collection interface provided by the present application, the action collection interface includes a preview control 76 and an action preview area 79, and the action preview area 79 displays the action video. Of course, the action collection interface can also include a video clipping control for obtaining the action video by clipping the recorded video, and specifically, the time axis of the recorded video can be displayed, and one of the video clips can be selected as the action video through the time axis, thereby ensuring the accuracy of the action video. In specific implementation, the start and end time of the action video in the recorded video can be determined through the time axis, so as to determine the action video from the recorded video. Referring to Figure 17 The video between the two time points on the time axis 77 can be selected as the action video.

[0129] The action collection interface also includes a complete control and a cancel control. The complete control is used to generate the action identifier according to the custom action that has been collected when triggered, and the cancel control is used to cancel the action collection and cancel the display of the action collection interface when triggered. Referring to Figure 17 The complete control 72 and the cancel control 71 are located above the action collection interface 1005.

[0130] If the determined target action identifier is the action identifier of the custom action, the mapping relationship between the action feature point parameters and the contour feature points of the second virtual object can be determined, wherein the action feature point parameters of the custom action are extracted according to the contour feature point positions of the entity object in the action video and the time sequence changes of the contour feature point positions. In the process of calling the video structure parameters to play the virtual human video, the second virtual object is instructed to make the custom action through the mapping relationship. In this way, the reuse of the custom action can be realized, the reuse rate of the action video can be improved, the action does not need to be drawn one by one for different virtual objects, the action library can be effectively expanded, the implementation cost can be reduced, and the degree of freedom of action selection can be improved.

[0131] When the adding instruction can also include the action parameter, as another embodiment of the target action, the object identifier of the second virtual object can be displayed, and the target action can be embodied through the object identifier. Specifically, in response to the adding instruction of the second clip, the target action identifier can be determined according to the action parameter, the target action identifier is used to instruct the second virtual object to speak the second plot text through the target action corresponding to the target action identifier, and after the object identifier of the second virtual object is configured with the target action indicated by the target action identifier, the object identifier of the second virtual object is displayed in the plot conversation area. The object identifier of the second virtual object and the second plot bubble are associated through the display position relationship to identify the association relationship. The object identifier of the second virtual object can be displayed in the same row as the second plot bubble to identify the association relationship between them.

[0132] The target action identifier can be one or two. When the target action identifier is one, the second virtual object speaks the second plot text through the target action; when the target action identifier is two, the two target action identifiers correspond to a target facial expression action and a target body action respectively, and the second virtual object speaks the second plot text through the target facial expression action and the target body action. For example, see Figure 18 Another schematic diagram of a plot conversation area provided by an embodiment of the present application is shown, the first plot bubble 11 can correspond to a smiling action, and the object identifier 12 of the first virtual object embodies the smiling action; the second plot bubble 21 can correspond to the smiling action, and the object identifier 22 of the second virtual object embodies the smiling action; and the fourth plot bubble 41 can correspond to the smiling action and a scissors hand action, and the object identifier 42 of the fourth virtual object embodies the smiling action and the scissors hand action.

[0133] After the target action identifier is determined according to the action parameter in the adding instruction, the mapping relationship between the action feature point parameter of the target action corresponding to the target action identifier and the contour feature point of the second virtual object can be determined, and in the process of calling the video structure parameter to play the virtual human video, the second virtual human is instructed to make the custom action through the mapping relationship. The action feature point parameter of the target action can be determined according to the contour feature point position corresponding to the target action and the time sequence change of the contour feature point position. The contour feature point can be a skeletal feature point. When the target action is a target facial expression action, the contour feature point is a facial skeletal feature point. When the target action is a target body action, the contour feature point is a skeletal feature point of an action execution part.

[0134] Specifically, the target action identifier can be determined from the action identifiers displayed in the action selection area according to the action selection operation, and then the adding instruction for the second segment can be generated according to the action parameter of the target action identifier in response to the segment adding operation for the second segment. The related description of the segment adding operation is referred to the description in S102, which is not repeated here. The action selection area can be in the object setting interface, so that the object identifier and the action identifier can be displayed in the same interface, reducing the refreshing frequency of the interface and improving the user experience. The object setting interface can be displayed through the triggering operation of the object setting control. In addition, the action library can be enriched in a custom form. The foregoing description can be referred to, which is not repeated here.

[0135] Referring to Figure 19As shown in FIG. 6, the scenario conversation area provided by the embodiment of the present application is illustrated again. The object setting interface is displayed by triggering the object setting control 31. The object setting interface displays the object identifier 64. The object setting interface can include the action selection area 62. The object identifier 64 is used as the target object identifier corresponding to the second virtual object when selected. The action identifier in the action selection area 62 is used as the target action identifier when selected. The second scenario text can be obtained by using the scenario text input box 32. After the target object identifier corresponding to the second virtual object and the target action identifier are determined, the adding instruction of the second clip can be generated by triggering the clip adding control 33. The adding instruction includes the unique identifier of the second virtual object, the second text and the action parameter.

[0136] From the user side, the adding of the second clip can be achieved by the following operations: the object setting interface is called by the object selection operation. The object setting interface displays the object identifier and the action identifier in the action selection area. The second virtual object can be selected by the selection operation of the object identifier. The target action identifier can be selected by the selection operation of the action identifier. The input text is input by the scenario text input box. Then, the clip adding operation is performed to generate the adding instruction of the second clip. The adding instruction includes the unique identifier of the second virtual object, the second scenario text and the action parameter. The input text is used as the second scenario text. The action parameter is determined according to the target action identifier. Such operation is similar to chatting or playing games, which greatly reduces the video editing threshold.

[0137] If the target action includes the target limb action, the execution parameter of the target limb action can also be determined according to the length of the second scenario text. The action parameter also includes the execution parameter of the target limb action. The execution parameter can include at least one of the execution times or the execution speed, so that the target limb action and the second scenario text are matched, and the execution duration of the target limb action is close to the duration of the voice information corresponding to the second scenario text, avoiding the problem that the video quality is poor due to the mismatch between the action duration and the voice duration. If the target action includes the target facial expression action, the execution parameter of the target facial expression action can also be determined according to the length of the second scenario text. The action parameter also includes the execution parameter. The execution parameter can include at least one of the execution times or the maintenance duration, so as to better match the second scenario text. The execution duration of the target limb action is close to the duration of the voice information corresponding to the second scenario text, avoiding the problem that the video quality is poor due to the mismatch between the action duration and the voice duration. Of course, in other embodiments of the present application, the execution parameter of the target action can be irrelevant to the length of the second scenario text. Regardless of the length of the second scenario text, the execution parameter of the target action is a fixed value.

[0138] In summary, the substructure parameters corresponding to the second plot bubble may include the unique identifier of the second virtual object, the second plot text, and may also include at least one of voice information, subtitle information, action parameters, playback order, and playback delay duration. By editing each substructure parameter in the editing state, the entire substructure parameter can be edited, and the video structure parameters can also be edited.

[0139] See also Figure 20 , which is a functional schematic diagram of a plot conversation area provided in an embodiment of the present application. Through the plot conversation area, virtual objects can be set to achieve the setting of plot segments corresponding to dialogue bubbles, such as determining the unique identifier of the virtual object, the plot text corresponding to the virtual object, the target action corresponding to the virtual object, etc. When the target action is a target physical action, the starting point, end point and execution location of the target physical action can also be determined; through the plot conversation area, virtual camera parameters can also be set to achieve the setting of plot segments corresponding to camera movement bubbles, such as determining the camera movement time, camera movement point, camera movement curve, etc.; through the plot conversation area, the playback order and playback delay duration of the plot segments corresponding to each plot bubble can also be set; through the plot conversation area, a preview of the plot segments can also be triggered, and the preview of the plot segments can be displayed through the plot preview area.

[0140] See also Figure 21 , which is a schematic diagram of information processing provided by an embodiment of the present application. Regarding text processing, based on the input plot text, voice information can be obtained by combining audio parameters, and subtitle information can be obtained by combining subtitle parameters. The display duration of the subtitle information is determined based on the length of the plot text. Regarding virtual objects, virtual objects and their actions can be determined. An action can include at least one of a facial expression or a body movement, causing the virtual object to perform the action. The execution parameters of the action can be determined based on the length of the plot text.

[0141] See also Figure 22, which is a schematic diagram of another information processing provided by an embodiment of the present application. According to the acquired text and virtual object information, the substructure parameters of each plot bubble can be obtained. For example, for the first plot bubble corresponding to the first segment, the substructure parameters include the unique identifier of the first virtual object, the first plot text, the target action, voice information, audio parameters, etc., wherein the unique identifier of the first virtual object is "Little A", the first plot text is "Good morning, Little B", the target action is a smile, the content of the voice information is the same as the first plot text, the voice information is generated according to the audio parameters, and the audio parameters have female timbre characteristics; for the second plot bubble corresponding to the second segment, the substructure parameters include the unique identifier of the second virtual object, the second plot text, the target action, the voice information, and the audio parameters. The substructure parameters include the unique identifier of the fourth virtual object, the fourth plot text, the target action, voice information, audio parameters, etc., wherein the unique identifier of the second virtual object is "Xiao B", the second plot text is "Good morning", the target action is smiling, the content of the voice information is the same as the second plot text, the voice information is generated according to the audio parameters, and the audio parameters have male timbre characteristics; for the fourth plot bubble corresponding to the fourth segment, the substructure parameters include the unique identifier of the fourth virtual object, the fourth plot text, the target action, voice information, audio parameters, etc., wherein the unique identifier of the fourth virtual object is "Xiao A", the fourth plot text is "I encountered a happy thing", the target action is smiling and scissors hand gestures, the content of the voice information is the same as the fourth plot text, the voice information is generated according to the audio parameters, and the audio parameters have female timbre characteristics.

[0142] See also Figure 23 , which is a schematic diagram of another video structure parameter provided in an embodiment of the present application. The substructure parameters corresponding to the fifth plot bubble include virtual camera parameters, which may include camera movement time, camera movement point, camera movement curve, etc.; the substructure parameters corresponding to the first plot bubble include the unique identifier of the first virtual object "Xiao A", the first plot text "Good morning, Xiao B", the target action, voice information, subtitle information, playback order, and playback delay duration, etc.; the substructure parameters corresponding to the second plot bubble include the unique identifier of the second virtual object "Xiao B", the second plot text "Good morning", the target action, voice information, subtitle information, playback order, and playback delay duration, etc.; the substructure parameters corresponding to the fourth plot bubble include the unique identifier of the fourth virtual object "Xiao A", the fourth plot text "I encountered a happy thing", the target action, voice information, subtitle information, playback order, and playback delay duration, etc.

[0143] Based on the method for creating a conversation video provided in the above embodiment, the present application also provides a device for creating a conversation video. Figure 24 , Figure 24 This is a structural block diagram of a conversation video creation device provided in an embodiment of the present application. The conversation video creation device 1300 includes:

[0144] The plot conversation area display unit 1301 is configured to display a plot conversation area for editing a virtual human video, wherein the virtual human video to be edited includes multiple plot segments, and the plot conversation area displays a first plot bubble corresponding to a first segment of the multiple plot segments, wherein the first plot bubble is associated with a first virtual object involved in the first segment, and the first plot bubble includes a first plot text of the first virtual object in the first segment;

[0145] A plot bubble display unit 1302 is configured to, in response to an instruction to add a second segment from the plurality of plot segments, display a second plot bubble corresponding to the second segment in the plot conversation area, wherein the second plot bubble is associated with a second virtual object related to the second segment, the second plot bubble includes second plot text of the second virtual object in the second segment, and the first plot bubble and the second plot bubble have a display position relationship in the plot conversation area, the display position relationship being used to identify a plot order of the first segment and the second segment in the virtual human video;

[0146] The video structure parameter generating unit 1303 is configured to generate video structure parameters according to the first plot bubble and the second plot bubble displayed in the plot conversation area in response to the video generating instruction, wherein the video structure parameters are used to play the virtual human video when called.

[0147] Optionally, the adding instruction includes an action parameter, and the plot bubble display unit 1302 includes:

[0148] a display parameter determining unit configured to, in response to an instruction to add a second segment from the plurality of scenario segments, determine a target action identifier and a target position of the target action identifier in the second scenario text based on the action parameters, wherein the target action identifier is used to instruct the second virtual object to speak a target text in the second scenario text through a target action corresponding to the target action identifier, and the target text is determined based on the target position;

[0149] The plot bubble display subunit is configured to display a second plot bubble including the second plot text and the target action identifier in the plot conversation area according to the target action identifier and the target position.

[0150] Optionally, the target action identifier includes a first action identifier and a second action identifier, the first action identifier is located at a first target position of the second plot text, the second action identifier is located at a second target position of the second plot text, and in the second plot text, the first target position is located before the second target position;

[0151] The first action identifier corresponds to a first target word determined according to a word before the first target position in the second plot word, and the second action identifier corresponds to a second target word determined according to a word between the first target position and the second target position in the second plot word.

[0152] Optionally, the device further comprises:

[0153] A position determination unit, configured to determine the target position according to selection of an action addition position in the second plot word;

[0154] An action identifier determination unit, configured to determine, from action identifiers displayed in the action selection area, a target action identifier added at the action addition position in the second plot word according to an action selection operation;

[0155] An action parameter generation unit, configured to generate the action parameter according to the target action identifier and the target position;

[0156] An addition instruction generation unit, configured to generate an addition instruction for the second segment according to the action parameter in response to a segment addition operation for the second segment.

[0157] Optionally, the position determination unit is specifically configured to:

[0158] Determine the target position according to an action addition operation for a plot word input box, wherein the plot word input box is used to obtain the second plot word, the second plot word is a word in the plot word input box when the segment addition operation is obtained, and the action addition operation is for an action addition position in an input word in the plot word input box;

[0159] The device further comprises:

[0160] An action selection area display unit, configured to display the action selection area in response to the action addition operation.

[0161] Optionally, the position determination unit comprises:

[0162] A semantic analysis unit, configured to perform semantic analysis on a target word segment in an input word in a plot word input box to obtain a target semantic, wherein the plot word input box is used to obtain the second plot word, and the second plot word is a word in the plot word input box when the segment addition operation is obtained;

[0163] A position determination sub-unit, configured to determine a recommended action corresponding to the target semantic, and determine an action addition position corresponding to an action identifier of the recommended action as a position after the target word segment.

[0164] Determining the target location according to the action adding location corresponding to the action identifier of the recommended action;

[0165] The device further comprises:

[0166] The action selection area display unit is used to display the action identifier of the recommended action in the action selection area.

[0167] Optionally, the device further includes:

[0168] A video acquisition unit, configured to acquire an action video by capturing an entity object performing a custom action;

[0169] an action feature point parameter extraction unit, configured to extract action feature point parameters of the custom action based on the positions of the shape feature points of the physical object in the action video and the temporal changes of the positions of the shape feature points;

[0170] An action identifier display unit is used to display the action identifier of the custom action in the action selection area according to the action feature point parameters.

[0171] Optionally, the device further includes:

[0172] a mapping relationship determining unit, configured to determine a mapping relationship between the action feature point parameters and the appearance feature points of the second virtual object if the target action identifier is the action identifier of the custom action;

[0173] An instructing unit is configured to instruct the second virtual object to perform the custom action through the mapping relationship during the process of calling the video structure parameters to play the virtual human video.

[0174] Optionally, the adding instruction includes an action parameter, and the device further includes:

[0175] an action identifier determining unit, configured to determine, in response to an instruction to add the second segment, a target action identifier according to the action parameter, wherein the target action identifier is used to instruct the second virtual object to speak the second plot text through a target action corresponding to the target action identifier;

[0176] The object identification display unit is configured to configure the target action indicated by the target action identifier for the object identification of the second virtual object, and then display the object identification of the second virtual object in the plot conversation area, wherein the object identification of the second virtual object and the second plot bubble are associated by displaying a position relationship identifier.

[0177] Optionally, the adding instruction includes a unique identifier of the second virtual object, and the apparatus further includes:

[0178] An object setting interface display unit, configured to display an object setting interface in response to an object selection operation, wherein the object setting interface displays an object identifier and includes an action selection area;

[0179] a unique identifier determining unit, configured to, in response to a selection operation on a target object identifier in the object identifier, use the object corresponding to the target object identifier as a second virtual object and obtain a unique identifier for the second virtual object;

[0180] a target action identifier determining unit, configured to determine the target action identifier from the action identifiers displayed in the action selection area according to an action selection operation;

[0181] An action parameter generating unit is configured to generate the action parameter according to the target action identifier.

[0182] Optionally, the target action includes at least one of a target facial expression action or a target limb action. When the target action is the target limb action, the action parameters include an action starting point, an action end point, and an action execution location.

[0183] Optionally, if the target action includes a target limb action, the device further includes:

[0184] An execution parameter determination unit is used to determine the execution parameters of the target limb movement according to the length of the second plot text, wherein the action parameters also include execution parameters of the target limb movement, and the execution parameters include at least one of the number of executions or the execution speed.

[0185] Optionally, the device further includes:

[0186] a recognition text input unit, configured to input recognition text obtained by performing speech recognition on a target speech in the plot text input box, the recognition text being used as input text in the plot text input box;

[0187] an audio parameter determination unit, configured to determine audio parameters for the second plot text after acquiring the second plot text, wherein the audio parameters are determined based on at least one of the target voice or the second virtual object, and the audio parameters include pitch and timbre;

[0188] A voice information generating unit is configured to generate voice information according to the second plot text and the audio parameters, wherein the video structure parameters also include the voice information.

[0189] Optionally, the device further includes:

[0190] a second segment generating unit configured to generate a second segment based on the second virtual object and the second segment text in response to a preview operation on the second segment bubble after a second segment bubble corresponding to the second segment is displayed in the segment conversation area;

[0191] The second segment display unit is configured to display the second segment in a scenario preview area outside the scenario conversation area.

[0192] Optionally, the device further includes:

[0193] A position adjustment unit is configured to change a display position relationship between the second plot bubble and the first plot bubble in response to a position adjustment operation on the second plot bubble.

[0194] Optionally, in the plot conversation area, a third plot bubble corresponding to a third segment among the plurality of plot segments and the second plot bubble are displayed in an adjacent position relationship, and the device further includes:

[0195] The play order adjustment unit is configured to set the play order of the plot segments respectively corresponding to the second plot bubble and the third plot bubble to be played simultaneously in response to the play order adjustment operation on the second plot bubble and the third plot bubble.

[0196] Optionally, the device further includes:

[0197] A delay duration adjustment unit is used to determine the playback delay duration of the second segment in response to a delay duration adjustment operation on the second plot bubble, wherein the playback delay duration is used to indicate that when the playback order of the virtual human video reaches the second segment, the second segment is played after the playback delay duration.

[0198] The present application also provides a computer device, which is the aforementioned computer device and may include a terminal device or a server. The aforementioned conversation video creation apparatus may be configured within the computer device. The computer device includes a processor and a memory. The memory is configured to store a computer program and transmit the computer program to the processor. The processor is configured to execute the aforementioned conversation video creation method according to the instructions in the computer program. The computer device is described below with reference to the accompanying drawings.

[0199] If the computer device is a terminal device, see Figure 25 As shown, the embodiment of the present application provides a terminal device, taking a mobile phone as an example:

[0200] Figure 25 The block diagram shows a partial structure of a mobile phone related to the terminal device provided in the embodiment of the present application. Figure 25The mobile phone includes components such as a radio frequency (RF) circuit 1410, a memory 1420, an input unit 1430, a display unit 1440, a sensor 1450, an audio circuit 1460, a wireless fidelity (WiFi) module 1470, a processor 1480, and a power supply 1490. Those skilled in the art will understand that Figure 25 The mobile phone structure shown in the figure does not constitute a limitation to the mobile phone, and may include more or fewer components than shown in the figure, or combine certain components, or arrange the components differently.

[0201] The following combination Figure 25 A detailed introduction to the various components of a mobile phone:

[0202] The RF circuit 1410 may be used for receiving and sending signals during information transmission or calls. In particular, after receiving downlink information from the base station, it is sent to the processor 1480 for processing. In addition, the designed uplink data is sent to the base station.

[0203] Memory 1420 can be used to store software programs and modules. Processor 1480 executes the various functional applications and data processing of the mobile phone by running the software programs and modules stored in memory 1420. Memory 1420 may mainly include a program storage area and a data storage area. The program storage area may store an operating system and at least one application required for a function (such as a sound playback function, an image playback function, etc.); the data storage area may store data created based on the use of the mobile phone (such as audio data, a phone book, etc.). In addition, memory 1420 may include high-speed random access memory and non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device.

[0204] The input unit 1430 may be configured to receive input digital or character information and generate key signal input related to user settings and function control of the mobile phone. Specifically, the input unit 1430 may include a touch panel 1431 and other input devices 1432 .

[0205] The display unit 1440 may be configured to display information input by the user or information provided to the user, as well as various menus of the mobile phone. The display unit 1440 may include a display panel 1441 .

[0206] The mobile phone may also include at least one sensor 1450, such as a light sensor, a motion sensor, and other sensors.

[0207] The audio circuit 1460 , the speaker 1461 , and the microphone 1462 can provide an audio interface between the user and the mobile phone.

[0208] WiFi belongs to short-range wireless transmission technology, and the mobile phone can help users send and receive emails, browse web pages and access streaming media through the WiFi module 1470, which provides users with wireless broadband Internet access.

[0209] The processor 1480 is the control center of the mobile phone, which connects all parts of the mobile phone through various interfaces and lines, and performs various functions and processes data of the mobile phone by running or executing software programs and / or modules stored in the memory 1420 and calling data stored in the memory 1420.

[0210] The mobile phone also includes a power supply 1490 (such as a battery) for powering various components.

[0211] In this embodiment, the processor 1480 included in the terminal device also has the following functions:

[0212] The scenario conversation area for editing the virtual person video is displayed, the virtual person video to be edited includes a plurality of scenario clips, the scenario conversation area displays a first scenario bubble corresponding to a first clip in the plurality of scenario clips, the first scenario bubble has an association relationship with a first virtual object involved in the first clip, and the first scenario bubble includes first scenario text of the first virtual object in the first clip;

[0213] In response to an adding instruction of a second clip in the plurality of scenario clips, a second scenario bubble corresponding to the second clip is displayed in the scenario conversation area, the second scenario bubble has an association relationship with a second virtual object involved in the second clip, the second scenario bubble includes second scenario text of the second virtual object in the second clip, the first scenario bubble and the second scenario bubble have a display position relationship in the scenario conversation area, and the display position relationship is used to identify a scenario order of the first clip and the second clip in the virtual person video;

[0214] In response to a video generation instruction, a video structure parameter is generated according to the first scenario bubble and the second scenario bubble displayed in the scenario conversation area, and the video structure parameter is used to play the virtual person video when called.

[0215] If the computer device is a server, the embodiment of the application further provides a server, please refer to Figure 26 , Figure 26The structural diagram of the server 1500 provided in the embodiment of the present application, the server 1500 may have relatively large differences due to different configurations or performances, and may include one or more processors 1522, such as central processing units (CPUs), memories 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) storing application programs 1542 or data 1544. Among them, the memories 1532 and the storage media 1530 can be temporary storage or permanent storage. The program stored in the storage medium 1530 may include one or more modules (not shown in the figure), each module may include a series of instruction operations on the server. Furthermore, the processor 1522 can be configured to communicate with the storage medium 1530 to execute a series of instruction operations in the storage medium 1530 on the server 1500.

[0216] The server 1500 may also include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input and output interfaces 1558, and / or one or more operating systems 1541, such as Windows Server 2003. TM , Mac OS X TM , Unix TM ,Linux TM , FreeBSD TM etc.

[0217] The steps performed by the server in the above embodiment can be based on Figure 26 The server structure shown.

[0218] In addition, an embodiment of the present application further provides a storage medium, which is used to store a computer program, and the computer program is used to execute the method provided by the above embodiment.

[0219] An embodiment of the present application further provides a computer program product including a computer program, which, when executed on a computer device, enables the computer device to execute the method provided in the above embodiment.

[0220] Those skilled in the art will understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to computer program instructions, and the aforementioned computer program can be stored in a computer-readable storage medium. When the computer program is executed, it executes the steps of the above-mentioned method embodiments; and the aforementioned storage medium can be at least one of the following media: read-only memory (English: Read-only Memory, abbreviated: ROM), RAM, magnetic disk or optical disk, and other media that can store program codes.

[0221] It should be noted that the various embodiments in this specification are described in a progressive manner, and the same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments. The device and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0222] The above is only one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the scope of protection of the present application. Moreover, based on the implementation methods provided in the above aspects, the present application can also be further combined to provide more implementation methods. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

Claims

1. A method for creating a conversation video, characterized in that: The method comprises: A plot conversation area is displayed for editing a virtual human video, wherein the virtual human video to be edited includes multiple plot segments. The plot conversation area displays a first plot bubble corresponding to a first segment of the multiple plot segments. The first plot bubble is associated with a first virtual object related to the first segment. The first plot bubble includes a first plot text of the first virtual object in the first segment. The plot conversation area includes a segment addition control, a plot file input box, and an object selection control. In response to an instruction to add a second segment from the plurality of segment segments, a second scenario bubble corresponding to the second segment is displayed in the scenario conversation area. The second scenario bubble is associated with a second virtual object related to the second segment. The second scenario bubble includes second scenario text of the second virtual object in the second segment. The first scenario bubble and the second scenario bubble have a display position relationship in the scenario conversation area. The display position relationship is used to identify a plot order of the first segment and the second segment in the virtual human video. The scenario bubbles in the scenario conversation area display conversation content of different virtual objects. The plot order of the conversation content in the virtual human video is sorted out through the display position relationship between the scenario bubbles. In response to a video generation instruction, generating video structure parameters according to the first plot bubble and the second plot bubble displayed in the plot conversation area, wherein the video structure parameters are used to play the virtual human video when called; The adding instruction includes an action parameter, and the displaying of the second plot bubble corresponding to the second segment in the plot conversation area includes: determining a target action identifier and a target position of the target action identifier in the second plot text according to the action parameter, wherein the target action identifier is used to instruct the second virtual object to speak target text in the second plot text through a target action corresponding to the target action identifier, and the target text is determined according to the target position; A second scenario bubble including the second scenario text and the target action identifier is displayed in the scenario conversation area according to the target action identifier and the target position.

2. The method according to claim 1, characterized in that The target action identifier includes a first action identifier and a second action identifier, the first action identifier is located at a first target position of the second plot text, and the second action identifier is located at a second target position of the second plot text, and in the second plot text, the first target position is located before the second target position; The first target text corresponding to the first action identifier is determined based on the text before the first target position in the second plot text, and the second target text corresponding to the second action identifier is determined based on the text between the first target position and the second target position in the second plot text.

3. The method according to claim 1, characterized in that The method further comprises: Determining the target position according to the selection of the action adding position in the second plot text; Determining, from the action identifiers displayed in the action selection area, a target action identifier to be added at the action adding position in the second plot text according to the action selection operation; generating the action parameters according to the target action identifier and the target position; In response to the segment adding operation for the second segment, an adding instruction for the second segment is generated according to the action parameter.

4. The method according to claim 3, characterized in that The step of determining the target position according to the selection of the action adding position in the second plot text includes: Determining the target position according to an action adding operation on a plot text input box, wherein the plot text input box is used to obtain the second plot text, the second plot text being the text in the plot text input box when the segment adding operation is obtained, and the action adding operation is performed on the action adding position within the text already inputted in the plot text input box; The method further comprises: In response to the action adding operation, the action selection area is displayed.

5. The method according to claim 3, characterized in that The step of determining the target position according to the selection of the action adding position in the second plot text includes: Performing semantic analysis on a target text segment in text input into a plot text input box to obtain target semantics, wherein the plot text input box is used to obtain the second plot text, the second plot text being the text in the plot text input box when the segment adding operation is performed; Determining a recommended action corresponding to the target semantics, and determining that an action adding position corresponding to an action identifier of the recommended action is a position after the target text segment; Determining the target location according to the action adding location corresponding to the action identifier of the recommended action; The method further comprises: An action identifier of the recommended action is displayed in the action selection area.

6. The method according to claim 3, characterized in that The method further comprises: By collecting the entity objects that perform custom actions, action videos are obtained; Extracting action feature point parameters of the custom action based on the positions of the shape feature points of the physical object in the action video and the temporal changes of the positions of the shape feature points; The action identifier of the custom action is displayed in the action selection area according to the action feature point parameters.

7. The method according to claim 6, characterized in that If the target action identifier is the action identifier of the custom action, the method further includes: Determining a mapping relationship between the motion feature point parameters and the shape feature points of the second virtual object; In the process of calling the video structure parameters to play the virtual human video, the second virtual object is instructed to perform the custom action through the mapping relationship.

8. The method according to claim 1, characterized in that The adding instruction includes an action parameter, and the method further includes: In response to the instruction to add the second segment, determining a target action identifier according to the action parameter, wherein the target action identifier is used to instruct the second virtual object to speak the second plot text through a target action corresponding to the target action identifier; After configuring the target action indicated by the target action identifier for the object identifier of the second virtual object, the object identifier of the second virtual object is displayed in the plot conversation area, and the object identifier of the second virtual object and the second plot bubble are associated by displaying a position relationship identifier.

9. The method according to claim 8, characterized in that The adding instruction includes a unique identifier of the second virtual object, and the method further includes: In response to the object selection operation, displaying an object setting interface, wherein the object identification is displayed in the object setting interface, and the object setting interface includes an action selection area; In response to a selection operation on a target object identifier in the object identifiers, taking the object corresponding to the target object identifier as a second virtual object, and obtaining a unique identifier of the second virtual object; Determining the target action identifier from the action identifiers displayed in the action selection area according to an action selection operation; The action parameter is generated according to the target action identifier.

10. The method according to any one of claims 1 to 9, characterized in that The target action includes at least one of a target facial expression action or a target limb action. When the target action is the target limb action, the action parameters include an action starting point, an action end point, and an action execution location.

11. The method according to claim 10, characterized in that If the target action includes a target body movement, the method further includes: The execution parameters of the target limb action are determined according to the length of the second plot text, and the action parameters also include execution parameters of the target limb action, and the execution parameters include at least one of the number of executions or the execution speed.

12. The method according to claim 4 or 5, characterized in that The method further comprises: Inputting the recognized text obtained by performing speech recognition on the target speech in the plot text input box, and the recognized text serves as the input text in the plot text input box; After acquiring the second plot text, determining audio parameters for the second plot text, the audio parameters being determined based on at least one of the target voice or the second virtual object, the audio parameters including pitch and timbre; Voice information is generated according to the second plot text and the audio parameters, and the video structure parameters also include the voice information.

13. The method according to any one of claims 1 to 9, characterized in that After displaying the second plot bubble corresponding to the second segment in the plot conversation area, the method further includes: In response to a preview operation on the second plot bubble, determining a substructure parameter corresponding to the second plot bubble according to the second virtual object and the second plot text, wherein the video structure parameter includes the substructure parameter; generating a second segment according to the substructure parameters; The second segment is displayed in a plot preview area outside the plot conversation area.

14. The method according to claim 1, wherein The method further comprises: In response to a position adjustment operation on the second storyline bubble, a display position relationship between the second storyline bubble and the first storyline bubble is changed.

15. The method according to claim 1, wherein In the plot conversation area, a third plot bubble corresponding to a third segment among the plurality of plot segments and the second plot bubble are displayed in an adjacent position relationship, and the method further includes: In response to the play order adjustment operation on the second plot bubble and the third plot bubble, the play order of the plot segments respectively corresponding to the second plot bubble and the third plot bubble is set to be played simultaneously.

16. The method according to claim 14 or 15, characterized in that The method further comprises: In response to the delay duration adjustment operation on the second plot bubble, the playback delay duration of the second segment is determined, and the playback delay duration is used to indicate that when the playback order of the virtual human video reaches the second segment, the second segment is played after the playback delay duration.

17. A conversation video creation device, characterized in that: The device comprises: A plot conversation area display unit is configured to display a plot conversation area for editing a virtual human video, wherein the virtual human video to be edited includes multiple plot segments, and the plot conversation area displays a first plot bubble corresponding to a first segment of the multiple plot segments, wherein the first plot bubble is associated with a first virtual object involved in the first segment, and the first plot bubble includes a first plot text of the first virtual object in the first segment. The plot conversation area includes a segment addition control, a plot file input box, and an object selection control; a plot bubble display unit, configured to, in response to an instruction to add a second segment from the plurality of plot segments, display a second plot bubble corresponding to the second segment in the plot conversation area, wherein the second plot bubble is associated with a second virtual object related to the second segment, the second plot bubble including second plot text of the second virtual object in the second segment, the first plot bubble and the second plot bubble having a display position relationship in the plot conversation area, the display position relationship being used to identify a plot sequence of the first segment and the second segment in the virtual human video, the plot bubbles in the plot conversation area displaying conversation content of different virtual objects, and the plot sequence of the conversation content in the virtual human video being sorted out by the display position relationship between the plot bubbles; a video structure parameter generating unit, configured to generate video structure parameters according to the first plot bubble and the second plot bubble displayed in the plot conversation area in response to a video generation instruction, wherein the video structure parameters are used to play the virtual human video when called; The adding instruction includes action parameters, and the plot bubble display unit includes: a display parameter determining unit configured to determine a target action identifier and a target position of the target action identifier in the second plot text based on the action parameter, wherein the target action identifier is configured to instruct the second virtual object to speak a target text in the second plot text by performing a target action corresponding to the target action identifier, and the target text is determined based on the target position; The plot bubble display subunit is configured to display a second plot bubble including the second plot text and the target action identifier in the plot conversation area according to the target action identifier and the target position.

18. The device according to claim 17, characterized in that The target action identifier includes a first action identifier and a second action identifier, the first action identifier is located at a first target position of the second plot text, and the second action identifier is located at a second target position of the second plot text, and in the second plot text, the first target position is located before the second target position; The first target text corresponding to the first action identifier is determined based on the text before the first target position in the second plot text, and the second target text corresponding to the second action identifier is determined based on the text between the first target position and the second target position in the second plot text.

19. The device according to claim 17, characterized in that The device further comprises: a position determination unit, configured to determine the target position according to the selection of the action adding position in the second plot text; an action identifier determining unit, configured to determine, from the action identifiers displayed in the action selection area, a target action identifier to be added at the action adding position in the second plot text according to an action selection operation; an action parameter generating unit, configured to generate the action parameter according to the target action identifier and the target position; An adding instruction generating unit is configured to generate an adding instruction for the second segment according to the action parameter in response to the segment adding operation for the second segment.

20. The device according to claim 19, characterized in that The position determination unit is specifically configured to: Determining the target position according to an action adding operation on a plot text input box, wherein the plot text input box is used to obtain the second plot text, the second plot text being the text in the plot text input box when the segment adding operation is obtained, and the action adding operation is performed on the action adding position within the text already inputted in the plot text input box; The device further comprises: The action selection area display unit is used to display the action selection area in response to the action adding operation.

21. The device according to claim 19, characterized in that The position determination unit includes: a semantic analysis unit, configured to perform semantic analysis on a target text segment in text input into a plot text input box to obtain target semantics, wherein the plot text input box is used to obtain the second plot text, the second plot text being the text in the plot text input box when the segment adding operation is performed; a position determination subunit, configured to determine a recommended action corresponding to the target semantics, and determine that the action adding position corresponding to the action identifier of the recommended action is the position after the target text segment; and determine the target position according to the action adding position corresponding to the action identifier of the recommended action; The device further comprises: The action selection area display unit is used to display the action identifier of the recommended action in the action selection area.

22. The device according to claim 19, characterized in that The device further comprises: A video acquisition unit, configured to acquire an action video by capturing an entity object performing a custom action; an action feature point parameter extraction unit, configured to extract action feature point parameters of the custom action based on the positions of the shape feature points of the physical object in the action video and the temporal changes of the positions of the shape feature points; An action identifier display unit is used to display the action identifier of the custom action in the action selection area according to the action feature point parameters.

23. The device according to claim 22, characterized in that If the target action identifier is the action identifier of the custom action, the device further includes: a mapping relationship determining unit, configured to determine a mapping relationship between the action feature point parameters and the shape feature points of the second virtual object; An instructing unit is configured to instruct the second virtual object to perform the custom action through the mapping relationship during the process of calling the video structure parameters to play the virtual human video.

24. The device according to claim 17, wherein The adding instruction includes an action parameter, and the device further includes: an action identifier determining unit, configured to determine, in response to an instruction to add the second segment, a target action identifier according to the action parameter, wherein the target action identifier is used to instruct the second virtual object to speak the second plot text through a target action corresponding to the target action identifier; The action identifier display unit is configured to configure the target action indicated by the target action identifier for the object identifier of the second virtual object, and then display the object identifier of the second virtual object in the plot conversation area, wherein the object identifier of the second virtual object and the second plot bubble are associated by displaying a position relationship identifier.

25. The device according to claim 24, characterized in that The adding instruction includes a unique identifier of the second virtual object, and the device further includes: An object setting interface display unit, configured to display an object setting interface in response to an object selection operation, wherein the object setting interface displays an object identifier and includes an action selection area; a unique identifier determining unit, configured to, in response to a selection operation on a target object identifier in the object identifier, use the object corresponding to the target object identifier as a second virtual object and obtain a unique identifier for the second virtual object; a target action identifier determining unit, configured to determine the target action identifier from the action identifiers displayed in the action selection area according to an action selection operation; An action parameter generating unit is configured to generate the action parameter according to the target action identifier.

26. The device according to any one of claims 17 to 25, characterized in that The target action includes at least one of a target facial expression action or a target limb action. When the target action is the target limb action, the action parameters include an action starting point, an action end point, and an action execution location.

27. The device according to claim 26, characterized in that If the target action includes a target limb action, the device further includes: An execution parameter determination unit is used to determine the execution parameters of the target limb movement according to the length of the second plot text, wherein the action parameters also include execution parameters of the target limb movement, and the execution parameters include at least one of the number of executions or the execution speed.

28. The device according to claim 20 or 21, characterized in that The device further comprises: a recognition text input unit, configured to input recognition text obtained by performing speech recognition on a target speech in the plot text input box, the recognition text being used as input text in the plot text input box; an audio parameter determination unit, configured to determine audio parameters for the second plot text after acquiring the second plot text, wherein the audio parameters are determined based on at least one of the target voice or the second virtual object, and the audio parameters include pitch and timbre; A voice information generating unit is configured to generate voice information according to the second plot text and the audio parameters, wherein the video structure parameters also include the voice information.

29. The device according to any one of claims 17 to 25, characterized in that The device further comprises: a second segment generating unit configured to, after displaying a second scenario bubble corresponding to the second segment in the scenario conversation area, determine, in response to a preview operation on the second scenario bubble, substructure parameters corresponding to the second scenario bubble based on the second virtual object and the second scenario text, wherein the video structure parameters include the substructure parameters; and generate a second segment based on the substructure parameters; The second segment display unit is configured to display the second segment in a scenario preview area outside the scenario conversation area.

30. The device according to claim 17, wherein The device further comprises: A position adjustment unit is configured to change a display position relationship between the second plot bubble and the first plot bubble in response to a position adjustment operation on the second plot bubble.

31. The device according to claim 17, wherein In the plot conversation area, a third plot bubble corresponding to a third segment among the plurality of plot segments and the second plot bubble are displayed in an adjacent position relationship, and the device further includes: The play order adjustment unit is configured to set the play order of the plot segments respectively corresponding to the second plot bubble and the third plot bubble to be played simultaneously in response to the play order adjustment operation on the second plot bubble and the third plot bubble.

32. The device according to claim 30 or 31, characterized in that The device further comprises: A delay duration adjustment unit is used to determine the playback delay duration of the second segment in response to a delay duration adjustment operation on the second plot bubble, wherein the playback delay duration is used to indicate that when the playback order of the virtual human video reaches the second segment, the second segment is played after the playback delay duration.

33. A computer device, characterized in that: The computer device includes a processor and a memory: The memory is used to store a computer program and transmit the computer program to the processor; The processor is configured to execute the conversation video creation method according to any one of claims 1 to 16 according to instructions in the computer program.

34. A computer-readable storage medium, characterized in that The computer-readable storage medium is used to store a computer program, and the computer program is used to execute the conversation video creation method according to any one of claims 1 to 16.

35. A computer program product comprising a computer program, characterized in that When the method is run on a computer device, the computer device is enabled to execute the conversation video creation method described in any one of claims 1 to 16.