Video generation methods, apparatus, electronic devices and storage media
By displaying multiple templates and replacing object information during the video generation process, the target video is generated, solving the problem of low video generation efficiency and achieving more efficient and flexible video interaction effects.
Patent Information
- Application Number
- CN202411147076.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-20
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2044-08-20
AI Technical Summary
Existing technologies have low video generation efficiency, resulting in poor interactive effects.
By displaying multiple video templates, after selecting the target video template, the object information collection page is displayed, and a collection confirmation command is responded to. The facial image and audio of the template object are replaced with the facial image and audio of the collected object to generate the target video.
It improves the efficiency and flexibility of video generation and enhances the interactive effect of videos.
Smart Images

Figure CN119255037B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of multimedia technology, and in particular to a video generation method, apparatus, electronic device, and storage medium. Background Technology
[0002] With the development of the Internet and multimedia technology, video-based interactive methods have become more diverse. In general, users need to create their own videos for interaction. Video production involves shooting and editing, which leads to low efficiency and poor interactive effects. Summary of the Invention
[0003] This disclosure provides a video generation method, apparatus, electronic device, and storage medium to at least address the problem of how to improve video generation efficiency in related technologies. The technical solution of this disclosure is as follows:
[0004] According to a first aspect of the present disclosure, a video generation method is provided, comprising:
[0005] In response to video generation commands, display multiple video templates;
[0006] When the target video template is selected from the plurality of video templates, the object information collection page is displayed;
[0007] In response to a collection confirmation command triggered by the object information collection page, the target video corresponding to the target video template is displayed; the target video is obtained by replacing the template object's face image with the collected object's face image and replacing the template's audio with the collected object's audio.
[0008] The facial image and audio of the target object are obtained by collecting the target object information from the target information collection page under the collection confirmation command. The facial image of the template object is the facial image of the template object in the target video template, and the audio of the template is the audio in the target video template.
[0009] In one possible implementation, the object information acquisition page includes an image acquisition page and an audio acquisition page; the step of displaying the target video corresponding to the target video template in response to an acquisition confirmation command triggered based on the object information acquisition page includes:
[0010] In response to an image acquisition completion command triggered by the image acquisition page, the facial image of the acquired object is displayed;
[0011] If the facial image of the subject being captured is confirmed, the audio capture page is displayed.
[0012] In response to an audio acquisition completion command triggered by the audio acquisition page, the audio of the acquired object is displayed;
[0013] If the audio of the acquired object is confirmed, the target video is displayed.
[0014] In one possible implementation, the image acquisition page displays a face description input area, and the acquired face image includes a generated face image; the method further includes:
[0015] In response to the input completion operation of inputting facial description information in the facial description input area, the image acquisition completion command is generated;
[0016] The step of displaying the facial image of the captured object in response to an image capture completion command triggered by the image capture page includes:
[0017] In response to the image acquisition completion command, the generated facial image is displayed.
[0018] In one possible implementation, the image acquisition page displays a shooting control, and the captured facial image includes a captured facial image; the step of displaying the captured facial image in response to an image acquisition completion command triggered based on the image acquisition page includes:
[0019] In response to the image acquisition completion command triggered by the shooting control, the captured facial image is displayed.
[0020] In one possible implementation, the audio acquisition page displays multiple preset keywords; the acquired audio is a first acquired audio; the method further includes:
[0021] When at least one preset keyword is selected from the plurality of preset keywords and a selection confirmation operation is triggered, a first audio text and a first audio acquisition trigger control corresponding to the first audio text are displayed; the first audio text is generated based on the at least one preset keyword;
[0022] When the first audio acquisition trigger control is triggered, the audio corresponding to the first audio text is acquired, and when the audio acquisition corresponding to the first audio text is acquired, the audio acquisition completion instruction is generated.
[0023] The step of displaying the audio of the captured object in response to an audio capture completion command triggered by the audio capture page includes:
[0024] In response to the audio acquisition completion command, the first acquired audio is displayed. The first acquired audio is the audio corresponding to the first audio text acquired within a first time period. The start time of the first time period is the time when the first audio acquisition trigger control is triggered, and the end time of the first time period is the time corresponding to the audio acquisition completion command.
[0025] In one possible implementation, the audio acquisition page displays a custom text input area; the acquired audio is a second acquired audio; the method further includes:
[0026] When custom text is detected in the input area and the text input is completed, the second audio text corresponding to the custom text and the second audio acquisition trigger control corresponding to the second audio text are displayed.
[0027] When the second audio acquisition trigger control is triggered, the audio corresponding to the second audio text is acquired, and when the audio acquisition corresponding to the second audio text is acquired, the audio acquisition completion instruction is generated.
[0028] The step of displaying the audio of the captured object in response to an audio capture completion command triggered by the audio capture page includes:
[0029] In response to the audio acquisition completion command, the second acquired audio is displayed. The second acquired audio is the audio corresponding to the second audio text acquired during the second time period. The start time of the second time period is the time when the second audio acquisition trigger control is triggered, and the end time of the second time period is the time corresponding to the audio acquisition completion command.
[0030] In one possible implementation, the audio acquisition page displays preset audio text and an audio acquisition trigger control; the acquired audio is a third-party acquired audio; the method further includes:
[0031] When the audio acquisition trigger control is triggered, the audio corresponding to the preset audio text is acquired, and when the audio acquisition corresponding to the preset audio text is completed, the audio acquisition completion command is generated.
[0032] The step of displaying the audio of the captured object in response to an audio capture completion command triggered by the audio capture page includes:
[0033] In response to the audio acquisition completion command, the third acquired audio is displayed. The third acquired audio is obtained by replacing the timbre in the template audio with the timbre in the audio corresponding to the preset audio text when the audio acquisition trigger control is triggered.
[0034] In one possible implementation, the method further includes:
[0035] Display the session page;
[0036] In the session page, the video generation instruction is generated in response to a message interaction request based on a preset event;
[0037] The process of displaying the target video corresponding to the target video template includes:
[0038] The target video, which is in the target display state, and the object identifier of the first session object that sent the target video are displayed on the session page; the target display state is used to indicate whether the video content of the target video is displayed.
[0039] In one possible implementation, the conversation page displays a conversation video sent by a second conversation object, in a state where the message content is not displayed, and a reply control matching the conversation video; the conversation video is a conversation message based on the preset event; the message interaction request is a message reply interaction request; the target display state is a state where the message content is displayed; the method further includes:
[0040] When the reply control is triggered, the message reply interaction request is generated;
[0041] The object identifier of the first session object that sent the target video and displays the target video in the target display state on the session page includes:
[0042] The target video, which is in the state of displaying message content, and the object identifier of the first session object that sent the target video are displayed on the session page, and the session video is switched to the state of displaying message content.
[0043] In one possible implementation, the session page displays a message interaction control that triggers the preset event; the message interaction request is a message-initiated interaction request; the target display state is a state where no message content is displayed; the method further includes:
[0044] When the message interaction control is triggered, the message is generated to initiate an interaction request;
[0045] The object identifier of the first session object that sent the target video and displays the target video in the target display state on the session page includes:
[0046] The target video, which is in the state of not displaying message content, the object identifier of the first session object, and the reply control matching the target video are displayed on the session page.
[0047] According to a second aspect of the present disclosure, a video generation apparatus is provided, the apparatus comprising:
[0048] The video template display module is configured to respond to video generation instructions and display multiple video templates;
[0049] The information collection module is configured to display an object information collection page when a target video template is selected from the plurality of video templates.
[0050] The target video display module is configured to execute a collection confirmation command triggered based on the object information collection page, and display the target video corresponding to the target video template; the target video is obtained by replacing the template object's face image with the collected object's face image and replacing the template's audio with the collected object's audio;
[0051] The facial image and audio of the target object are obtained by collecting the target object information from the target information collection page under the collection confirmation command. The facial image of the template object is the facial image of the template object in the target video template, and the audio of the template is the audio in the target video template.
[0052] In one possible implementation, the object information acquisition page includes an image acquisition page and an audio acquisition page; the target video display module includes:
[0053] The image acquisition and display unit is configured to execute an image acquisition completion command triggered based on the image acquisition page and display the facial image of the acquired object;
[0054] The audio acquisition page display unit is configured to display the audio acquisition page when a confirmation operation is performed on the facial image of the acquisition object.
[0055] The audio display unit for the acquired object is configured to display the audio of the acquired object in response to an audio acquisition completion command triggered based on the audio acquisition page.
[0056] The target video display unit is configured to display the target video when the audio of the acquired object is confirmed.
[0057] In one possible implementation, the image acquisition page displays a face description input area, and the acquired face image includes a generated face image; the device further includes:
[0058] The first acquisition completion instruction module is configured to execute an input completion operation in response to the input of facial description information in the facial description input area, and generate the image acquisition completion instruction;
[0059] Accordingly, the image display unit includes:
[0060] The face image display subunit is configured to display the generated face image in response to the image acquisition completion command.
[0061] In one possible implementation, the image acquisition page displays shooting controls, and the captured facial image includes a captured facial image; the acquired image display unit includes:
[0062] The face image display subunit is configured to execute and display the captured face image in response to the image acquisition completion command triggered by the shooting control.
[0063] In one possible implementation, the audio acquisition page displays multiple preset keywords; the acquired audio is a first acquired audio; the device may further include:
[0064] The first audio text display module is configured to display a first audio text and a first audio acquisition trigger control corresponding to the first audio text when at least one preset keyword is selected from the plurality of preset keywords and a selection confirmation operation is triggered; the first audio text is generated based on the at least one preset keyword;
[0065] The second acquisition completion instruction module is configured to start acquiring the audio corresponding to the first audio text when the first audio acquisition trigger control is triggered, and generate the audio acquisition completion instruction when the acquisition of the audio corresponding to the first audio text ends.
[0066] Accordingly, the audio display unit for the collected object includes:
[0067] The first audio acquisition display subunit is configured to display the first acquired audio in response to the audio acquisition completion instruction. The first acquired audio is the audio corresponding to the first audio text acquired within a first time period. The start time of the first time period is the time when the first audio acquisition trigger control is triggered, and the end time of the first time period is the time corresponding to the audio acquisition completion instruction.
[0068] In one possible implementation, the audio acquisition page displays a custom text input area; the acquired audio is a second acquired audio; the device may further include:
[0069] The second audio text display module is configured to display the second audio text corresponding to the custom text and the second audio acquisition trigger control corresponding to the second audio text when the input of custom text is detected in the input area and the text input is completed.
[0070] The third acquisition completion instruction module is configured to start acquiring the audio corresponding to the second audio text when the second audio acquisition trigger control is triggered, and generate the audio acquisition completion instruction when the acquisition of the audio corresponding to the second audio text ends.
[0071] Accordingly, the audio display unit for the collected object includes:
[0072] The second audio acquisition display subunit is configured to display the second acquired audio in response to the audio acquisition completion instruction. The second acquired audio is the audio corresponding to the second audio text acquired during the second time period. The start time of the second time period is the time when the second audio acquisition trigger control is triggered, and the end time of the second time period is the time corresponding to the audio acquisition completion instruction.
[0073] In one possible implementation, the audio acquisition page displays preset audio text and an audio acquisition trigger control; the acquired audio is a third-party acquired audio; the device may further include:
[0074] The fourth acquisition completion instruction module is configured to start acquiring the audio corresponding to the preset audio text when the audio acquisition trigger control is triggered, and generate the audio acquisition completion instruction when the acquisition of the audio corresponding to the preset audio text ends.
[0075] Accordingly, the audio display unit for the collected object includes:
[0076] The third audio acquisition display subunit is configured to execute a response to the audio acquisition completion instruction and display the third acquired audio. The third acquired audio is obtained by acquiring the timbre from the audio corresponding to the preset audio text and replacing the timbre in the template audio when the audio acquisition trigger control is triggered.
[0077] In one possible implementation, the device may further include:
[0078] The session page display module is configured to display the session page;
[0079] The video instruction generation module is configured to execute on the session page, in response to a message interaction request based on a preset event, to generate the video generation instruction;
[0080] The target video display module may include:
[0081] The display unit is configured to display the target video in the target display state and the object identifier of the first session object that sent the target video on the session page; the target display state is used to indicate whether the video content of the target video is displayed.
[0082] In one possible implementation, the conversation page displays a conversation video sent by a second conversation object, in a state where the message content is not displayed, and a reply control matching the conversation video; the conversation video is a conversation message based on the preset event; the message interaction request is a message reply interaction request; the target display state is a state where the message content is displayed; the device further includes:
[0083] The message reply interaction module is configured to generate the message reply interaction request when the reply control is triggered;
[0084] Accordingly, the display unit includes:
[0085] The first display subunit is configured to display the target video in the display message content state and the object identifier of the first session object that sent the target video in the session page, and switch the session video to the display message content state.
[0086] In one possible implementation, the session page displays a message interaction control that triggers the preset event; the message interaction request is a message-initiated interaction request; the target display state is a state where no message content is displayed; the device may further include:
[0087] The message initiation interaction module is configured to generate the message initiation interaction request when the message interaction control is triggered;
[0088] Accordingly, the display unit includes:
[0089] The second display subunit is configured to display the target video in the non-display message content state, the object identifier of the first session object, and the reply control matching the target video on the session page.
[0090] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the method as described in any one of the first aspects above.
[0091] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided such that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, the electronic device is enabled to perform any of the methods described in the first aspect of the present disclosure.
[0092] According to a fifth aspect of the present disclosure, a computer program product is provided, including computer instructions that, when executed by a processor, cause a computer to perform the method described in any one of the first aspects of the present disclosure.
[0093] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:
[0094] In response to a video generation command, multiple video templates are displayed. When a target video template is selected, an object information collection page is displayed. In response to a collection confirmation command triggered by the object information collection page, a target video corresponding to the target video template is displayed. The target video is obtained by replacing the template object's face image with the collected object's face image and replacing the template's audio with the collected object's audio. This replacement of the face image and audio in the video template makes video generation more efficient and flexible, thereby improving the effect of video-based interaction.
[0095] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0096] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0097] Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment.
[0098] Figure 2 This is a flowchart illustrating a video generation method according to an exemplary embodiment.
[0099] Figure 3 This is a schematic diagram illustrating a plurality of video templates according to an exemplary embodiment.
[0100] Figure 4 This is a schematic diagram illustrating an image acquisition page according to an exemplary embodiment.
[0101] Figure 5 This is a schematic diagram illustrating an image acquisition and audio acquisition process according to an exemplary embodiment.
[0102] Figure 6 This is a schematic diagram illustrating a session page according to an exemplary embodiment.
[0103] Figure 7 This is a schematic diagram illustrating the display of a target video in a conversation page when a message reply interaction request is triggered based on a reply control, according to an exemplary embodiment.
[0104] Figure 8 This is a schematic diagram illustrating the display of a target video in a session page when an interaction request is initiated based on a message interaction control, according to an exemplary embodiment.
[0105] Figure 9 This is a block diagram of a video generation apparatus according to an exemplary embodiment.
[0106] Figure 10 This is a block diagram illustrating an electronic device for video generation according to an exemplary embodiment.
[0107] Figure 11 This is a block diagram illustrating an electronic device for video generation based on an exemplary embodiment. Detailed Implementation
[0108] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0109] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0110] Please see Figure 1 , Figure 1 This is a schematic diagram illustrating an application environment according to an exemplary embodiment, such as... Figure 1 As shown, the application environment may include server 01 and terminal 02.
[0111] In an optional embodiment, terminal 02 can be used for video generation and display processing. Specifically, terminal 02 can be, but is not limited to, electronic devices such as smartphones, desktop computers, tablets, laptops, smart speakers, digital assistants, augmented reality (AR) / virtual reality (VR) devices, and smart wearable devices. Optionally, the operating system running on the electronic device can be, but is not limited to, Android, iOS, Linux, and Windows.
[0112] In an optional embodiment, server 01 can be used for backend support of video generation. Specifically, server 01 can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.
[0113] In addition, it should be noted that, Figure 1 The example shown is merely one application environment of the video generation method provided in this disclosure.
[0114] In the embodiments described in this specification, the server 01 and the terminal 02 can be directly or indirectly connected through wired or wireless communication, and this application does not impose any restrictions on this.
[0115] It should be noted that the following diagram illustrates one possible sequence of steps, and it is not strictly required to follow this order. Some steps can be performed in parallel without interdependence. The user information (including but not limited to user device information, user personal information, user behavior information, etc.) and data (including but not limited to data used for display, training data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties.
[0116] Figure 2 This is a flowchart illustrating a video generation method according to an exemplary embodiment. Figure 2 As shown, the steps may include the following.
[0117] In step S201, in response to the video generation instruction, multiple video templates are displayed.
[0118] In the embodiments of this specification, a video template can refer to a template used to generate a video. As an example, it can be a template used to generate a blessing video, which allows for the efficient generation of blessing videos based on the video template for blessing interactions. This disclosure does not limit the video template.
[0119] In one possible implementation, a video generation command can be triggered within a multimedia application. Accordingly, in response to the video generation command, multiple video templates can be displayed, such as... Figure 3 The templates shown are for generating blessing videos, i.e., video blessing templates. The content of each template (1 through 3) is as follows: Figure 3 It is not shown in the text.
[0120] In step S203, if the target video template is selected from multiple video templates, the object information collection page is displayed.
[0121] In practical applications, any video template can be selected for the video generation process; the selected video template can be called the target video template. Correspondingly, if the target video template is selected from multiple video templates, an object information collection page, such as a shooting page, will be displayed.
[0122] For example, the object information collection page can be used to collect facial images for subsequent replacement of the facial images of template objects in the target video template, and can be used to collect audio of the generated video for subsequent replacement of the template audio in the target video template. This disclosure does not limit the presentation method of the object information collection page, as long as the above collection functions can be achieved.
[0123] In step S205, in response to the collection confirmation command triggered by the object information collection page, the target video corresponding to the target video template is displayed; the target video is obtained by replacing the template object's face image with the collected object's face image and replacing the template's audio with the collected object's audio.
[0124] Among them, the facial image and audio of the subject being collected are obtained from the subject information collection page under the collection confirmation command, the facial image of the template object is the facial image of the template object in the target video template, and the audio of the template is the audio in the target video template.
[0125] In practical applications, after capturing facial images and audio, a capture confirmation command can be triggered. Correspondingly, in response to the capture confirmation command triggered by the object information capture page, the target video corresponding to the target video template can be displayed.
[0126] In response to a video generation command, multiple video templates are displayed. When a target video template is selected, an object information collection page is displayed. In response to a collection confirmation command triggered by the object information collection page, a target video corresponding to the target video template is displayed. The target video is obtained by replacing the template object's face image with the collected object's face image and replacing the template's audio with the collected object's audio. This replacement of the face image and audio in the video template makes video generation more efficient and flexible, thereby improving the effect of video-based interaction.
[0127] In one possible implementation, the object information acquisition page may include an image acquisition page and an audio acquisition page. Accordingly, the above-mentioned response to an acquisition confirmation command triggered based on the object information acquisition page, displaying the target video corresponding to the target video template, may include:
[0128] In response to the image acquisition completion command triggered by the image acquisition page, display the facial image of the acquired object;
[0129] If the facial image of the subject is confirmed, the audio capture page will be displayed.
[0130] In response to an audio capture completion command triggered by the audio capture page, display the audio of the captured object;
[0131] After confirming the audio of the collected object, the target video is displayed.
[0132] In the embodiments of this specification, the image acquisition page can refer to a page used for acquiring image sequences, such as a shooting page, and this disclosure does not limit this. The audio acquisition page can refer to a page used for acquiring audio, such as a recording page, a shooting page, etc., and this disclosure does not limit this either. By setting up image acquisition pages and audio acquisition pages, facial image acquisition and audio acquisition can be performed separately, making the acquisition of object information more flexible and the operation more convenient.
[0133] In one possible implementation, the image acquisition page described above can display a facial description input area, for example... Figure 5 As shown. Accordingly, the aforementioned acquisition of the subject's facial image may include generating a facial image, that is, generating a facial image using AI technology based on the facial description information input in the facial description input area. This makes facial image acquisition more intelligent and flexible, and applicable to scenarios involving virtual human interaction. Based on this, the above method may further include: generating an image acquisition completion instruction in response to an input completion operation of inputting facial description information in the facial description input area. For example, clicking... Figure 4 The "Complete" control in the text can be seen as a way to complete the input operation.
[0134] Accordingly, the above-mentioned response to the image acquisition completion command triggered by the image acquisition page, displaying the facial image of the acquired object, may include: displaying a generated facial image in response to the image acquisition completion command. The generated facial image may utilize AI-generated image processing technology to generate a facial image matching facial description information.
[0135] In another possible implementation, the image capture page could refer to a page capable of capturing images, allowing for the capture of facial images used to replace the video template, making image capture more convenient and efficient. Based on this, the aforementioned image capture page could display shooting controls, such as... Figure 5 As shown in the second image above, 51. The acquisition of the subject's facial image can include capturing a facial image. Accordingly, the above-mentioned response to an image acquisition completion command triggered by the image acquisition page, displaying the subject's facial image, can include: displaying the captured facial image in response to an image acquisition completion command triggered by the shooting control. For example, a preset operation can be performed on the shooting control to trigger the image acquisition completion command. The preset operation can include, but is not limited to, clicking, long-pressing, etc.
[0136] The above provides an exemplary description of the image acquisition page. The following provides an exemplary description of the audio acquisition page, which can be used to acquire at least one of audio content and audio timbre.
[0137] In one possible implementation, the aforementioned audio capture page can display multiple preset keywords, such as... Figure 5 As shown in the third image above. Accordingly, the audio of the acquired object is the first acquired audio. Based on this, the above method may also include:
[0138] When at least one preset keyword is selected from multiple preset keywords and a selection confirmation operation is triggered, a first audio text and a corresponding first audio acquisition trigger control are displayed; the first audio text is generated based on at least one preset keyword. For example, refer to... Figure 5 As shown in the upper right image, the confirmation option may include clicking "Generate Blessing Message," but this disclosure does not limit this option. (See reference...) Figure 5 As shown in the upper right image, selecting at least one preset keyword can include "family," "health," "happiness," and "holiday." (Refer to...) Figure 5 As shown in the lower right image, the first audio text can be the text under the generated blessing, and the first audio acquisition trigger control can be the "Confirm" control.
[0139] When the first audio capture trigger control is triggered, audio capture corresponding to the first audio text begins, and when audio capture corresponding to the first audio text ends, an audio capture completion command is generated. For example, when the first audio capture trigger control is triggered, a recording page can be displayed, such as... Figure 5 The middle figure below illustrates this. Thus, the audio corresponding to the first audio text can be captured on this recording page. Upon completion of the audio capture corresponding to the first audio text, an audio capture completion command can be generated. For example, clicking "Complete" confirms the end of audio capture; this disclosure does not limit this approach.
[0140] Accordingly, the above-mentioned response to an audio acquisition completion command triggered by the audio acquisition page, displaying the acquired audio, may include: responding to the audio acquisition completion command by displaying a first acquired audio, which may be the audio corresponding to the first audio text acquired within a first time period; the start time of the first time period may be the time when the first audio acquisition trigger control is triggered, and the end time of the first time period may be the time corresponding to the audio acquisition completion command. The first acquired audio may be as follows: Figure 5 As shown in Figure 52 in the lower left corner.
[0141] By combining keyword-based audio content with the collected timbre, the audio of the target video is generated. The audio content does not require manual input by the user, making the audio content generation more efficient. Furthermore, the timbre can be flexibly set, making the audio of the target video more flexible and efficient.
[0142] In another possible implementation, the audio capture page can display an input area with custom text; correspondingly, the audio to be captured can be a second audio file. Based on this, the method may further include:
[0143] When custom text is detected in the input area and the text input is completed, the second audio text corresponding to the custom text and the second audio acquisition trigger control corresponding to the second audio text are displayed.
[0144] When the second audio capture trigger control is triggered, the audio corresponding to the second audio text begins to be captured, and when the audio capture corresponding to the second audio text ends, an audio capture completion command is generated. The capture of the audio corresponding to the second audio text can also be done in... Figure 5 The recording page shown in the middle image below only displays the second audio text. The specific acquisition process can be found in the corresponding content above, and will not be repeated here.
[0145] Accordingly, the above-mentioned response to the audio acquisition completion instruction triggered by the audio acquisition page, displaying the acquired audio, may include: in response to the audio acquisition completion instruction, displaying a second acquired audio, which may be the audio corresponding to the second audio text acquired during a second time period; the start time of the second time period may be the time when the second audio acquisition trigger control is triggered, and the end time of the second time period may be the time corresponding to the audio acquisition completion instruction.
[0146] This disclosure does not limit the way the second audio text is displayed or the second audio acquisition trigger control.
[0147] By combining custom audio content with the captured timbre, the audio of the target video is generated. The audio content can be customized by the user without restriction, and the timbre can be flexibly set, so that the audio content of the target video can be differentiated based on different users, making the target video more personalized.
[0148] In another possible implementation, the audio capture page can display preset audio text and an audio capture trigger control. The preset audio text can refer to a pre-defined text used for capturing audio timbre, such as a piece of text. Correspondingly, the audio to be captured can be a third-party audio capture. Based on this, the method can further include: starting to capture the audio corresponding to the preset audio text when the audio capture trigger control is triggered, and generating an audio capture completion command when the audio capture corresponding to the preset audio text ends. Here, capturing the audio corresponding to the preset audio text can also be done in... Figure 5 The recording page shown in the middle image below only displays preset audio text. The specific acquisition process can be found in the corresponding content above, and will not be repeated here.
[0149] Accordingly, the above-mentioned response to the audio capture completion instruction triggered by the audio capture page, displaying the captured audio, may include: in response to the audio capture completion instruction, displaying a third captured audio, which can be obtained by replacing the timbre in the template audio with the timbre from the audio corresponding to the preset audio text when the audio capture trigger control is triggered. That is, the audio content of the target video is the same as the audio content in the target video template, but the timbre is different.
[0150] By replacing the timbre in the template audio with the collected timbre, the generation of the target video becomes more efficient, and the timbre in the target video can be flexibly set by the user.
[0151] As an example, the video generation method in the embodiments of this specification can be performed in a conversational scenario. Therefore, the above method may further include: displaying a conversation page, for example... Figure 6As shown. Furthermore, in the chat page, in response to message interaction requests based on preset events, video generation instructions can be generated.
[0152] In the embodiments of this specification, a preset event can refer to a pre-set event used to trigger message interaction, such as a time-sensitive event, for example, a holiday greeting event or a sports event. This disclosure does not limit this. It should be noted that the embodiments of this specification use a holiday greeting event to describe the conversation display method, and do not limit this disclosure.
[0153] In practical applications, the two parties in a conversation on a conversation page can interact based on preset events, such as exchanging greeting messages based on a holiday greeting event. The two parties in a conversation on the conversation page can include a first conversation partner and a second conversation partner; the first conversation partner can refer to oneself, and the second conversation partner can refer to the other party. In conversation interactions based on preset events, the target application provides multiple video templates corresponding to the preset events to make the conversation interactions more convenient and flexible. Based on this, a message interaction request based on a preset event can be triggered on the conversation page. Correspondingly, in response to this message interaction request, multiple video templates corresponding to the preset event can be displayed for the use of the conversation participants. The target application can refer to the application to which the conversation page belongs, such as a short video application, etc., and this disclosure does not limit this.
[0154] In one example, a message interaction request can be actively triggered by the first session object. For example... Figure 6 As shown, it can be triggered Figure 6 61 of them, "AI face-swapping to send blessings," generates message interaction requests.
[0155] In another example, a response can be made based on a pre-set event-triggered video message sent by a second session object, triggering a message interaction request. For example... Figure 6 As shown, for example, it can trigger Figure 6 The second session object sends a 62-second "rewind" of the session video to the right to generate a message interaction request.
[0156] Accordingly, the target video corresponding to the aforementioned target video template may include: the target video displayed in the target display state on the session page and the object identifier of the first session object that sent the target video.
[0157] The target display status can be used to indicate whether or not the video content of the target video is displayed. For example, the target display status can be the status for displaying message content, as shown in [reference needed]. Figure 3 As shown in figures 34 and 35 in the lower left corner; the target display state is a state where no message content is displayed, as can be seen in the following figures. Figure 663 indicates a non-display message content state; alternatively, the target display state can be a display message content state.
[0158] By responding to message interaction requests based on preset events on the conversation page and triggering video generation instructions, a convenient conversation interaction based on preset events and templates is realized. In the conversation interaction, the face image of the target video template can be replaced with the face image of the captured object, and the audio of the captured object can be replaced with the audio of the target, which increases the fun and richness of the conversation interaction, enriches the functions of the conversation interaction, and improves the utilization rate of conversation resources.
[0159] In one possible implementation, the conversation page may display a conversation video sent by a second conversation object in a state where the message content is not displayed, and a reply control matching the conversation video. The conversation video is a conversation message based on the preset event; the message interaction request is a message reply interaction request; and the target display state is a state where the message content is displayed. For example, the conversation video sent by the second conversation object in a state where the message content is not displayed can be as follows: Figure 6 As shown in Figure 63. The reply control can be as follows: Figure 6 As shown in Figure 62. Accordingly, the method may further include: generating the aforementioned message reply interaction request when the reply control is triggered, that is, the aforementioned message reply interaction request can be triggered based on the reply control.
[0160] Based on the aforementioned triggering of the message interaction request, the object identifier of the first session object that sent the target video and is in the target display state can be displayed on the session page, and may include: displaying the target video and the object identifier of the first session object that sent the target video and is in the message content display state on the session page, and switching the session video to the message content display state, for example... Figure 7 As shown in the lower left image.
[0161] In one example, the object identifier of the first session object may include, but is not limited to, the object avatar, object account, object name, etc. of the first session object, and this disclosure does not limit it.
[0162] Using the above message interaction methods, after replying on your side, both parties' conversation information will display the message content. For example... Figure 7 As shown, the other party's blessing message is initially in the waiting exchange phase. After the other party responds, both parties' blessing messages are displayed. This interactive display of conversation messages makes the conversation process richer and more flexible, and improves the utilization rate of conversation resources.
[0163] In one possible implementation, the session page may display a message interaction control that triggers a preset event; the message interaction request is a message-initiated interaction request; the target display state is a state where no message content is displayed; the method may further include: generating a message-initiated interaction request when the message interaction control is triggered.
[0164] Accordingly, the object identifier of the first session object that sends the target video and displays the target video in the target display state on the session page may include: displaying the target video in a non-display message content state, the object identifier of the first session object, and a reply control matching the target video on the session page.
[0165] For example, a message interaction control can be as follows: Figure 6 As shown in section 61, by triggering this message interaction control, the first conversation object can actively trigger a message interaction request instead of replying to the second conversation object's video. In this case, after the first conversation object sends the target video, it can also, similar to the conversation video sent by the second conversation object, not display any content and remain in a non-display message content state. The reply control matching the target video can be as follows: Figure 8 81 shown.
[0166] Figure 9 This is a block diagram illustrating a video generation apparatus according to an exemplary embodiment. (Refer to...) Figure 9 The device may include:
[0167] The video template display module 901 is configured to display multiple video templates in response to video generation instructions.
[0168] The information acquisition module 903 is configured to display an object information acquisition page when a target video template is selected among the plurality of video templates;
[0169] The target video display module 905 is configured to execute a collection confirmation command triggered based on the object information collection page, and display the target video corresponding to the target video template; the target video is obtained by replacing the template object's face image with the collected object's face image and replacing the template's audio with the collected object's audio;
[0170] The facial image and audio of the target object are obtained by collecting the target object information from the target information collection page under the collection confirmation command. The facial image of the template object is the facial image of the template object in the target video template, and the audio of the template is the audio in the target video template.
[0171] In response to a video generation command, multiple video templates are displayed. When a target video template is selected, an object information collection page is displayed. In response to a collection confirmation command triggered by the object information collection page, a target video corresponding to the target video template is displayed. The target video is obtained by replacing the template object's face image with the collected object's face image and replacing the template's audio with the collected object's audio. This replacement of the face image and audio in the video template makes video generation more efficient and flexible, thereby improving the effect of video-based interaction.
[0172] In one possible implementation, the object information acquisition page includes an image acquisition page and an audio acquisition page; the target video display module 905 may include:
[0173] The image acquisition and display unit is configured to execute an image acquisition completion command triggered based on the image acquisition page and display the facial image of the acquired object;
[0174] The audio acquisition page display unit is configured to display the audio acquisition page when a confirmation operation is performed on the facial image of the acquisition object.
[0175] The audio display unit for the acquired object is configured to display the audio of the acquired object in response to an audio acquisition completion command triggered based on the audio acquisition page.
[0176] The target video display unit is configured to display the target video when the audio of the acquired object is confirmed.
[0177] In one possible implementation, the image acquisition page displays a face description input area, and the acquired face image includes a generated face image; the device may further include:
[0178] The first acquisition completion instruction module is configured to execute an input completion operation in response to the input of facial description information in the facial description input area, and generate the image acquisition completion instruction;
[0179] Accordingly, the aforementioned image display unit may include:
[0180] The face image display subunit is configured to display the generated face image in response to the image acquisition completion command.
[0181] In one possible implementation, the image acquisition page displays shooting controls, and the captured facial image includes a captured facial image; the aforementioned image acquisition display unit may include:
[0182] The face image display subunit is configured to execute and display the captured face image in response to the image acquisition completion command triggered by the shooting control.
[0183] In one possible implementation, the audio acquisition page displays multiple preset keywords; the acquired audio is a first acquired audio; the device may further include:
[0184] The first audio text display module is configured to display a first audio text and a first audio acquisition trigger control corresponding to the first audio text when at least one preset keyword is selected from the plurality of preset keywords and a selection confirmation operation is triggered; the first audio text is generated based on the at least one preset keyword;
[0185] The second acquisition completion instruction module is configured to start acquiring the audio corresponding to the first audio text when the first audio acquisition trigger control is triggered, and generate the audio acquisition completion instruction when the acquisition of the audio corresponding to the first audio text ends.
[0186] Accordingly, the aforementioned audio display unit for the collected object may include:
[0187] The first audio acquisition display subunit is configured to display the first acquired audio in response to the audio acquisition completion instruction. The first acquired audio is the audio corresponding to the first audio text acquired within a first time period. The start time of the first time period is the time when the first audio acquisition trigger control is triggered, and the end time of the first time period is the time corresponding to the audio acquisition completion instruction.
[0188] In one possible implementation, the audio acquisition page displays a custom text input area; the acquired audio is a second acquired audio; the device may further include:
[0189] The second audio text display module is configured to display the second audio text corresponding to the custom text and the second audio acquisition trigger control corresponding to the second audio text when the input of custom text is detected in the input area and the text input is completed.
[0190] The third acquisition completion instruction module is configured to start acquiring the audio corresponding to the second audio text when the second audio acquisition trigger control is triggered, and generate the audio acquisition completion instruction when the acquisition of the audio corresponding to the second audio text ends.
[0191] Accordingly, the aforementioned audio display unit for the collected object may include:
[0192] The second audio acquisition display subunit is configured to display the second acquired audio in response to the audio acquisition completion instruction. The second acquired audio is the audio corresponding to the second audio text acquired during the second time period. The start time of the second time period is the time when the second audio acquisition trigger control is triggered, and the end time of the second time period is the time corresponding to the audio acquisition completion instruction.
[0193] In one possible implementation, the audio acquisition page displays preset audio text and an audio acquisition trigger control; the acquired audio is a third-party acquired audio; the device may further include:
[0194] The fourth acquisition completion instruction module is configured to start acquiring the audio corresponding to the preset audio text when the audio acquisition trigger control is triggered, and generate the audio acquisition completion instruction when the acquisition of the audio corresponding to the preset audio text ends.
[0195] Accordingly, the aforementioned audio display unit for the collected object may include:
[0196] The third audio acquisition display subunit is configured to execute a response to the audio acquisition completion instruction and display the third acquired audio. The third acquired audio is obtained by acquiring the timbre from the audio corresponding to the preset audio text and replacing the timbre in the template audio when the audio acquisition trigger control is triggered.
[0197] In one possible implementation, the device may further include:
[0198] The session page display module is configured to display the session page;
[0199] The video instruction generation module is configured to execute on the session page, in response to a message interaction request based on a preset event, to generate the video generation instruction;
[0200] The target video display module 905 may include:
[0201] The display unit is configured to display the target video in the target display state and the object identifier of the first session object that sent the target video on the session page; the target display state is used to indicate whether the video content of the target video is displayed.
[0202] In one possible implementation, the conversation page displays a conversation video sent by a second conversation object, in a state where the message content is not displayed, and a reply control matching the conversation video. The conversation video is a conversation message based on the preset event; the message interaction request is a message reply interaction request; the target display state is a state where the message content is displayed; the device may further include:
[0203] The message reply interaction module is configured to generate the message reply interaction request when the reply control is triggered;
[0204] Accordingly, the above-mentioned display unit may include:
[0205] The first display subunit is configured to display the target video in the display message content state and the object identifier of the first session object that sent the target video in the session page, and switch the session video to the display message content state.
[0206] In one possible implementation, the session page displays a message interaction control that triggers the preset event; the message interaction request is a message-initiated interaction request; the target display state is a state where no message content is displayed; the device may further include:
[0207] The message initiation interaction module is configured to generate the message initiation interaction request when the message interaction control is triggered;
[0208] Accordingly, the above-mentioned display unit may include:
[0209] The second display subunit is configured to display the target video in the non-display message content state, the object identifier of the first session object, and the reply control matching the target video on the session page.
[0210] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0211] Figure 10 This is a block diagram illustrating an electronic device for video generation according to an exemplary embodiment. The electronic device may be a terminal, and its internal structure diagram may be as follows: Figure 10As shown, the electronic device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a video generation method. The display screen can be a liquid crystal display (LCD) or an e-ink display. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad mounted on the device's casing, or an external keyboard, touchpad, or mouse.
[0212] Those skilled in the art will understand that Figure 10 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0213] Figure 11 This is a block diagram illustrating an electronic device for video generation based on an exemplary embodiment. The electronic device may be a server, and its internal structure diagram may be as follows: Figure 11 As shown, the electronic device includes a processor, memory, and a network interface connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with external terminals via a network connection. When the computer program is executed by the processor, it implements a video generation method.
[0214] Those skilled in the art will understand that Figure 11 The structure shown is merely a block diagram of a portion of the structure related to the present disclosure and does not constitute a limitation on the electronic device to which the present disclosure is applied. A specific electronic device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0215] In an exemplary embodiment, an electronic device is also provided, including: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement the video generation method as described in the embodiments of this disclosure.
[0216] In an exemplary embodiment, a computer-readable storage medium is also provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the video generation method of the present disclosure. The computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, or optical data storage device, etc.
[0217] In an exemplary embodiment, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the video generation method of the present disclosure embodiments.
[0218] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and RAMbus dynamic RAM (RDRAM), etc.
[0219] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0220] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A video generation method, characterized in that, include: In the chat page, in response to a message interaction request based on a preset event, a video generation instruction is generated, and in response to the video generation instruction, multiple video templates are displayed; When the target video template is selected from the plurality of video templates, the object information collection page is displayed; In response to the collection confirmation command triggered by the object information collection page, the target video corresponding to the target video template is displayed; The display of the target video corresponding to the target video template includes: displaying the target video in the target display state and the object identifier of the first session object that sent the target video in the session page; the target display state is used to indicate whether to display the video content of the target video; the target video is obtained by replacing the template object's face image with the collected object's face image and replacing the template's audio with the collected object's audio; The facial image and audio of the target object are obtained by collecting data based on the object information collection page under the collection confirmation command. The facial image of the template object is the facial image of the template object in the target video template, and the audio of the template is the audio in the target video template.
2. The method according to claim 1, characterized in that, The object information acquisition page includes an image acquisition page and an audio acquisition page; the step of displaying the target video corresponding to the target video template in response to the acquisition confirmation command triggered based on the object information acquisition page includes: In response to an image acquisition completion command triggered by the image acquisition page, the facial image of the acquired object is displayed; If the facial image of the subject being captured is confirmed, the audio capture page is displayed. In response to an audio acquisition completion command triggered by the audio acquisition page, the audio of the acquired object is displayed; If the audio of the acquired object is confirmed, the target video is displayed.
3. The method according to claim 2, characterized in that, The image acquisition page displays a facial description input area, and the acquired facial image includes a generated facial image; the method further includes: In response to the input completion operation of inputting facial description information in the facial description input area, the image acquisition completion command is generated; The step of displaying the facial image of the captured object in response to an image capture completion command triggered by the image capture page includes: In response to the image acquisition completion command, the generated facial image is displayed.
4. The method according to claim 2, characterized in that, The image acquisition page displays a shooting control, and the captured facial image includes a captured facial image; the step of displaying the captured facial image in response to an image acquisition completion command triggered based on the image acquisition page includes: In response to the image acquisition completion command triggered by the shooting control, the captured facial image is displayed.
5. The method according to claim 2, characterized in that, The audio acquisition page displays multiple preset keywords; the audio to be acquired is the first acquired audio; the method further includes: When at least one preset keyword is selected from the plurality of preset keywords and a selection confirmation operation is triggered, a first audio text and a first audio acquisition trigger control corresponding to the first audio text are displayed; the first audio text is generated based on the at least one preset keyword; When the first audio acquisition trigger control is triggered, the audio corresponding to the first audio text is acquired, and when the audio acquisition corresponding to the first audio text is completed, the audio acquisition completion instruction is generated. The step of displaying the audio of the captured object in response to an audio capture completion command triggered by the audio capture page includes: In response to the audio acquisition completion command, the first acquired audio is displayed. The first acquired audio is the audio corresponding to the first audio text acquired within a first time period. The start time of the first time period is the time when the first audio acquisition trigger control is triggered, and the end time of the first time period is the time corresponding to the audio acquisition completion command.
6. The method according to claim 2 or 5, characterized in that, The audio capture page displays a custom text input area; The audio to be acquired is the second acquired audio; the method further includes: When custom text is detected in the input area and the text input is completed, the second audio text corresponding to the custom text and the second audio acquisition trigger control corresponding to the second audio text are displayed. When the second audio acquisition trigger control is triggered, the audio corresponding to the second audio text is acquired, and when the audio acquisition corresponding to the second audio text is acquired, the audio acquisition completion instruction is generated. The step of displaying the audio of the captured object in response to an audio capture completion command triggered by the audio capture page includes: In response to the audio acquisition completion command, the second acquired audio is displayed. The second acquired audio is the audio corresponding to the second audio text acquired during the second time period. The start time of the second time period is the time when the second audio acquisition trigger control is triggered, and the end time of the second time period is the time corresponding to the audio acquisition completion command.
7. The method according to claim 2, characterized in that, The audio acquisition page displays preset audio text and an audio acquisition trigger control; the acquired audio is a third-party acquired audio; the method further includes: When the audio acquisition trigger control is triggered, the audio corresponding to the preset audio text is acquired, and when the audio acquisition corresponding to the preset audio text is completed, the audio acquisition completion command is generated. The step of displaying the audio of the captured object in response to an audio capture completion command triggered by the audio capture page includes: In response to the audio acquisition completion command, the third acquired audio is displayed. The third acquired audio is obtained by replacing the timbre in the template audio with the timbre in the audio corresponding to the preset audio text when the audio acquisition trigger control is triggered.
8. The method according to claim 7, characterized in that, The conversation page displays a conversation video sent by the second conversation object, which is in a state where no message content is displayed, and a reply control that matches the conversation video. The conversation video is a conversation message based on the preset event. The message interaction request is a message reply interaction request; The target display state is the state of displaying message content; The method further includes: When the reply control is triggered, the message reply interaction request is generated; The object identifier of the first session object that sent the target video and displays the target video in the target display state on the session page includes: The target video, which is in the state of displaying message content, and the object identifier of the first session object that sent the target video are displayed on the session page, and the session video is switched to the state of displaying message content.
9. The method according to claim 7, characterized in that, The session page displays message interaction controls that trigger the preset event; the message interaction request is a message-initiated interaction request. The target display state is a state where no message content is displayed; The method further includes: When the message interaction control is triggered, the message is generated to initiate an interaction request; The object identifier of the first session object that sent the target video and displays the target video in the target display state on the session page includes: The target video, which is in the state of not displaying message content, the object identifier of the first session object, and the reply control matching the target video are displayed on the session page.
10. A video generation apparatus, characterized in that, include: The session page display module is configured to display the session page; The video instruction generation module is configured to execute on the session page, in response to a message interaction request based on a preset event, to generate a video generation instruction; The video template display module is configured to display multiple video templates in response to the video generation command; The information collection module is configured to display an object information collection page when a target video template is selected from the plurality of video templates. The target video display module is configured to execute a collection confirmation command triggered based on the object information collection page and display the target video corresponding to the target video template. The target video is obtained by replacing the template object's face image with the face image of the captured object and replacing the template's audio with the audio of the captured object; Wherein, the facial image and audio of the target object are obtained by collecting the target object information based on the target object information collection page under the collection confirmation command, the facial image of the template object is the facial image of the template object in the target video template, and the audio of the template is the audio in the target video template; The target video display module includes: The display unit is configured to display the target video in the target display state and the object identifier of the first session object that sent the target video on the session page; the target display state is used to indicate whether the video content of the target video is displayed.
11. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the video generation method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the video generation method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that, It includes computer instructions, which, when executed by a processor, cause the computer to perform the video generation method as described in any one of claims 1 to 9.
Citation Information
Patent Citations
Video processing method and device, storage medium and electronic equipment
CN114222077A
Session processing method and apparatus, and device and storage medium
WO2024169975A1