Video generation method, apparatus, device, and storage medium

By recording voice, automatically identifying text and searching multimedia materials to generate video clips, the complex operation of video production software is solved, and the user-friendly video generation process is realized, which reduces learning costs and improves efficiency.

CN113518160BActive Publication Date: 2025-07-25TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202110035480.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-01-12
Publication Date
2025-07-25
Estimated Expiration
2041-01-12

AI Technical Summary

Technical Problem

The existing video production software has complex human-computer interaction, which leads to high learning costs and low efficiency, making it difficult to meet the needs of lightweight video production scenarios.

Method used

By recording voice content, automatically identify text materials and searching for multimedia materials, generate video clips, simplify the video production process, and reduce user operation complexity.

Benefits of technology

Users can quickly generate videos without professional knowledge, reduce learning costs and material search time, and improve video production efficiency and human-computer interaction efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113518160B_ABST
    Figure CN113518160B_ABST
Patent Text Reader

Abstract

The present application discloses a video generation method, apparatus, device and medium, which are applied in the field of artificial intelligence. The method includes: responding to a first recording operation, recording first voice content; displaying first text material and first multimedia material corresponding to the first voice content; responding to a video generation operation, displaying a video having a first video segment, where the first video segment is generated based on the first text material and the first multimedia material. This method enables users to generate videos with relatively simple operations.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of artificial intelligence, and particularly to a video generation method, apparatus, device, and storage medium. Background Art

[0002] Video is a relatively high-quality information dissemination carrier in Internet information. Many online information websites, video websites, and short video APPs (Applications) use video as the main presentation method.

[0003] In the related art, professional video production software is provided. Users prepare material files such as text materials, picture materials, audio materials, video materials, and transition animations in advance. Then, various materials are edited in the video production software to generate a video file.

[0004] Since the human-computer interaction operation of video production software is relatively complex, for many lightweight video production scenarios, the learning cost for users to use video production software is too high and the efficiency is relatively low. Summary of the Invention

[0005] The present application provides a video generation method, apparatus, device, and storage medium, which can generate tutorial and training videos with a relatively low learning cost and a relatively simple recording operation. The technical solutions are as follows:

[0006] According to one aspect of the present application, a video generation method is provided. The method includes:

[0007] In response to a first recording operation, record first voice content;

[0008] Display a first text material and a first multimedia material corresponding to the first voice content;

[0009] In response to a video generation operation, display a video having a first video segment, where the first video segment is generated based on the first text material and the first multimedia material.

[0010] According to another aspect of the present application, a video generation method is provided. The method is applied to a server and includes:

[0011] In response to a first recording operation in a terminal, obtain a first text material and a first multimedia material, where the first text material is obtained by performing speech recognition on the first voice content, and the first multimedia material is searched based on at least one of the first voice content and the first text material;

[0012] In response to a video generation operation in the terminal, send a video with a first video segment to the terminal, where the first video segment is generated based on the first text material and the first multimedia material.

[0013] According to another aspect of the present application, there is provided a video generation device, the device includes:

[0014] A recording module, configured to record first voice content in response to a first recording operation;

[0015] A display module, configured to display the first text material and the first multimedia material corresponding to the first voice content;

[0016] The display module is further configured to display a video with a first video segment in response to a video generation operation, where the first video segment is generated based on the first text material and the first multimedia material.

[0017] In an alternative design of the present application, the display module is further configured to display the first text material obtained by performing speech recognition on the first voice content during the recording process.

[0018] In an alternative design of the present application, the display module is further configured to display the first multimedia material corresponding to the first voice content after the recording ends.

[0019] In an alternative design of the present application, the display module is further configured to display the first multimedia material corresponding to the first speech recognition content after recognizing the first speech recognition content during the recording process; after recognizing the second speech recognition content, display the updated first multimedia material; where the updated first multimedia material corresponds to the second speech recognition content, or the updated first multimedia material corresponds to the first speech recognition content and the second speech recognition content.

[0020] In an alternative design of the present application, the display module is further configured to display the first multimedia material corresponding to the first speech recognition content after recognizing the first speech recognition content during the recording process; after recognizing the second speech recognition content, display the first multimedia material corresponding to the second speech recognition content.

[0021] In an alternative design of the present application, the display module is further configured to display a video production interface, and the video production interface includes a recording control.

[0022] The recording module is further configured to record the first voice content in response to a first recording operation on the recording control.

[0023] In an alternative design of the present application, the first multimedia material includes: picture material; the video frames of the first video clip are generated based on the picture material, and the audio frames of the first video clip are generated based on the speech content; or, the first multimedia material includes: video material; the video frames of the first video clip are generated based on the video material, and the audio frames of the first video clip are generated based on the speech content; or, the first multimedia material includes: audio material; the audio frames of the first video clip are generated based on the speech content and the audio material.

[0024] In an alternative design of the present application, the display module is further configured to display the edited first text material in response to an editing operation on the first text material; wherein, the editing operation includes at least one of: an operation of adding text, an operation of deleting text, an operation of searching for text, an operation of modifying text, an operation of replacing text, an operation of moving text, and an operation of changing the format.

[0025] In an alternative design of the present application, the display module is further configured to display a replacement control corresponding to the first multimedia material.

[0026] The display module is further configured to display the alternative multimedia material as the first multimedia material in response to a replacement operation on the replacement control, and the alternative multimedia material is searched based on at least one of the speech content and the first text material.

[0027] In an alternative design of the present application, the display module is further configured to display a deletion control corresponding to the first multimedia material; in response to a deletion operation on the deletion control, delete the first multimedia material and display an import control; in response to an import operation on the import control, display the imported multimedia material as the first multimedia material.

[0028] In an alternative design of the present application, the recording module is further configured to record a second speech content in response to a second recording operation.

[0029] The display module is further configured to display a second text material and a second multimedia material corresponding to the second speech content, the second text material is obtained by performing speech recognition on the second speech content, and the second multimedia material is searched based on at least one of the speech content and the second text material.

[0030] The display module is further configured to display a video having a first video clip and a second video clip in response to a video generation operation, and the second video clip is generated based on the second text material and the second multimedia material.

[0031] In an alternative design of the present application, there is also a transition animation between the first video segment and the second video segment.

[0032] In an alternative design of the present application, the first text material is displayed in the first video segment in the form of subtitles.

[0033] In an alternative design of the present application, the display module is further configured to display a sharing control corresponding to the video having the first video segment.

[0034] The device further includes:

[0035] A communication module, configured to, in response to a sharing operation on the sharing control, send the video having the first video segment to other terminals; or, in response to the sharing operation on the sharing control, send the video having the first video segment to the network space.

[0036] In an alternative design of the present application, the communication module is further configured to send the first voice content to the server.

[0037] The communication module is further configured to receive the first text material and the first multimedia material replied by the server.

[0038] In an alternative design of the present application, the device further includes:

[0039] An identification module, configured to perform speech recognition on the first voice content to obtain the first text material.

[0040] The communication module is further configured to send the first text material to the server.

[0041] The communication module is further configured to receive the first multimedia material replied by the server based on the first text material.

[0042] According to another aspect of the present application, there is provided a video generation device, the device including:

[0043] An acquisition module, configured to acquire a first text material and a first multimedia material, where the first text material is obtained by performing speech recognition on a first voice content, and the first multimedia material is searched based on at least one of the first voice content and the first text material; the first voice content is recorded by a first recording operation on the terminal;

[0044] A synthesis module, configured to generate a video having a first video segment according to the first text material and the first multimedia material, where the first video segment is generated based on the first text material and the first multimedia material;

[0045] A sending module, configured to send the video with the first video segment to the terminal.

[0046] In an alternative design of the present application, the obtaining module is further configured to receive the first voice content sent by the terminal.

[0047] The synthesizing module is further configured to perform speech recognition on the first voice content to obtain the first text material; extract keywords from the first text material; and search for and obtain the first multimedia material based on the keywords.

[0048] In an alternative design of the present application, the obtaining module is further configured to receive the first text material sent by the terminal, where the first text material is obtained by the terminal performing speech recognition on the first voice content.

[0049] The synthesizing module is further configured to extract the keywords from the first text material; and search for and obtain the first multimedia material based on the keywords.

[0050] According to one aspect of the present application, there is provided a computer device, including: a processor and a memory, where the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the video generation method as described above.

[0051] According to another aspect of the present application, there is provided a computer-readable storage medium, where the storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the video generation method as described above.

[0052] According to another aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A controller reads the computer instructions from the computer-readable storage medium, and the controller executes the computer instructions, so that the display device executes the video generation method as described in the above aspect.

[0053] The beneficial effects brought by the technical solution provided by the embodiments of the present application at least include:

[0054] The user only needs to record the first voice content and perform a simple selection operation to quickly and simply generate a video, without the user spending a lot of time searching for multimedia materials, nor does the user need to have professional video editing knowledge. This can not only reduce the learning cost of the user using video production software, but also eliminate the process of the user searching for materials, improving the efficiency of video production and the human-computer interaction efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0056] Figure 1 It is an exemplary interface schematic diagram of a video generation method provided by an exemplary embodiment of the present application;

[0057] Figure 2 It is a schematic structural diagram of a computer system provided by an exemplary embodiment of the present application;

[0058] Figure 3 It is a schematic flowchart of a video generation method provided by an exemplary embodiment of the present application;

[0059] Figure 4 It is an exemplary interface schematic diagram of a video generation method provided by an exemplary embodiment of the present application;

[0060] Figure 5 It is a schematic flowchart of a video generation method provided by an exemplary embodiment of the present application;

[0061] Figure 6 It is an exemplary interface schematic diagram of a video generation method provided by an exemplary embodiment of the present application;

[0062] Figure 7 It is an exemplary interface schematic diagram of a video generation method provided by an exemplary embodiment of the present application;

[0063] Figure 8 It is an exemplary interface schematic diagram of a video generation method provided by an exemplary embodiment of the present application;

[0064] Figure 9 It is a schematic flowchart of a video generation method provided by an exemplary embodiment of the present application;

[0065] Figure 10 It is a schematic flowchart of a video generation method provided by an exemplary embodiment of the present application;

[0066] Figure 11 It is a schematic flowchart of a video generation method provided by an exemplary embodiment of the present application;

[0067] Figure 12 It is an exemplary service architecture provided by an exemplary embodiment of the present application;

[0068] Figure 13It is an exemplary architecture for converting exemplary speech content into text materials provided by an exemplary embodiment of the present application;

[0069] Figure 14 It is an exemplary structural diagram of the background system structure provided by an exemplary embodiment of the present application;

[0070] Figure 15 It is an exemplary structural schematic diagram of a Windows server provided by an exemplary embodiment of the present application;

[0071] Figure 16 It is a schematic flowchart of a video synthesis method provided by an exemplary embodiment of the present application;

[0072] Figure 17 It is a block diagram of a video synthesis device provided by an exemplary embodiment of the present application;

[0073] Figure 18 It is a block diagram of a video synthesis device provided by an exemplary embodiment of the present application;

[0074] Figure 19 It is a block diagram of a server provided by an exemplary embodiment of the present application. Detailed implementation manners

[0075] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all the implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0076] It should be understood that the "several" mentioned herein refers to one or more, and "multiple" refers to two or more. "And / or" describes the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0077] For the convenience of understanding, the nouns involved in the embodiments of the present application are described below.

[0078] Artificial Intelligence (AI): Artificial intelligence uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, including theories, methods, technologies, and application systems for perceiving the environment, acquiring knowledge, and using knowledge to achieve the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines to enable machines to have the functions of perception, reasoning, and decision-making.

[0079] Artificial intelligence technology is an interdisciplinary subject with a wide range of fields, including both hardware-level and software-level technologies. The basic technologies of artificial intelligence generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. The software technologies of artificial intelligence mainly include several major directions such as computer vision technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0080] Machine Learning (ML): Machine learning is an interdisciplinary subject that involves multiple disciplines such as probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers simulate or implement human learning behaviors to acquire new knowledge or skills and reorganize the existing knowledge structure to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent, and its applications cover all fields of artificial intelligence. Machine learning and deep learning usually include technologies such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and teaching learning.

[0081] Natural Language Processing (NLP): It is an important direction in the fields of computer science and artificial intelligence. It studies various theories and methods that can achieve effective communication between humans and computers using natural languages. Natural language processing is a science that integrates linguistics, computer science, and mathematics. Therefore, the research in this field will involve natural languages, that is, the languages used by people in daily life, so it has a close connection with the research of linguistics. Natural language processing technologies usually include technologies such as text processing, semantic understanding, machine translation, robot question answering, and knowledge graphs.

[0082] The solution provided in the embodiments of this application involves technologies such as natural language processing, image processing, and machine learning in artificial intelligence, which will be specifically described through the following embodiments:

[0083] Exemplarily, such as Figure 1As shown below, taking the video production by a doctor as an example, the video generation solution provided by this application will be described. First, the user interface 11 is a video production interface. At the top of the user interface 11, the interface content "Creation Content" is displayed. Below this text, the video title "Symptoms and Treatments of Eczema" is displayed. In the middle of the user interface 11, the text "Press the button below to start recording" is displayed. The doctor can press the recording control 101 on the user interface 11 to prepare to record the voice content.

[0084] After clicking the recording control 101, the user interface 12 is displayed. On the user interface 12, the aforementioned interface title "Creation Content" and video title "Symptoms and Treatments of Eczema" are still displayed. And below the video title, a text reminder and a graphical reminder of "Recording" are displayed to remind the doctor that the voice content is being recorded. Below the reminder of "Recording", the working state of "Voice to Text" is displayed, and below this working state, the voice content that has been input is displayed in text form for the doctor to view at any time. The doctor can click the pause control 102 located at the center of the bottom of the user interface 12 to temporarily stop the recording of the voice. Similarly, the doctor can also click the completion control 103 located on the right side of the pause control to complete the input of this section of voice content.

[0085] After the doctor clicks the completion control 103, the user interface 13 is displayed. Similarly, on the user interface 13, the interface title "Creation Content" and video title "Symptoms and Treatments of Eczema" are retained. At the middle position of the user interface 13, a play control 104, a first text material 105, and a first multimedia material 106 are displayed from top to bottom. Among them, the play control 104 is used to play the input voice content. The first text material 105 is the text converted from the input voice content. The first multimedia material 106 is a picture related to the voice content and / or the first text material 105. For example, a picture of a patient with eczema. A deletion control 115 is also displayed in the upper right corner of the user interface 13. If the doctor is not satisfied with the input voice content, the doctor can click the deletion control 115 to delete all or part of the play control 104, the first text material 105, and the first multimedia material 106 on the user interface 13. A replacement control 107 is also displayed on the first multimedia material 106. If the doctor is not satisfied with the generated first multimedia material 106, the doctor can click the replacement control 107 to replace the first multimedia material 106 with other multimedia materials. Similarly, a recording control 101 and a completion control 103 are also displayed at the bottom of the user interface 13. By clicking the recording control 101, the doctor can continue to record the next section of voice content; by clicking the completion control 103, the recording of the voice content is completed and the next step is entered.

[0086] Assume that the doctor has recorded two pieces of voice content, and the user interface 14 is displayed, on which the second text material 108 and the second multimedia material 109 are shown. The doctor can also replace the second multimedia material 109 or play the voice content, which will not be elaborated here.

[0087] After the doctor has completed the input of all voice content, click the completion control 103 on the user interface 14 to display the user interface 15. The user can perform editing operations on the first text material 105, the first multimedia material 106, the second text material 108, and the second multimedia material 109 shown on the user interface 15. For example, modify the text content of the first text material 105, or click the deletion control 110 on the upper right of the first multimedia material 106 to delete the first multimedia material 106.

[0088] After the doctor has completed the editing operation, click the video generation control 111 on the user interface 15 to display the user interface 16. The user interface 16 shows a video 112, and below the video 112, a text material 113 is shown. The text material 113 is obtained based on the first text material 105 and the second text material 108. The doctor can also click the sharing control 114 to share the generated video 112 and the text material 113 with other users or upload them to a preset network space.

[0089] Figure 2 The structural schematic diagram of a computer system provided by an exemplary embodiment of the present application is shown. The computer system 200 includes: a terminal 220 and a server 240.

[0090] A client related to video generation is installed on the terminal 220. The client can be a small program in an APP, or a dedicated application program, or a web client. The user performs operations related to video generation on the terminal 220. The terminal 220 is at least one of a smart phone, a tablet computer, an e-book reader, an MP3 player, an MP4 player, a laptop portable computer, and a desktop computer.

[0091] The terminal 220 is connected to the server 240 through a wireless network or a wired network.

[0092] The server 240 may be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The server 240 is used to provide background services for the application program that supports video generation. Optionally, the server 240 undertakes the main computing work, and the terminal 220 undertakes the secondary computing work; or, the server 240 undertakes the secondary computing work, and the terminal 220 undertakes the main computing work; or, the server 240 and the terminal 220 both adopt a distributed computing architecture for collaborative computing.

[0093] Figure 3 FIG. shows a schematic flowchart of a video generation method provided by an exemplary embodiment of the present application.

[0094] This method may be executed by Figure 2 the terminal 220 shown in FIG., and this method includes the following steps:

[0095] Step 302: In response to the first recording operation, record the first voice content.

[0096] The first recording operation refers to the operation of the user to record voice. The first recording operation is to execute the recording of voice by pressing one or more preset physical buttons. The user can also execute the first recording operation through the signals generated by releasing, long-pressing, clicking, double-clicking, and / or swiping on the touch screen.

[0097] The first voice content refers to the voice recorded by the user in real time. Optionally, the first voice content is obtained by downloading through the network, or the first voice content is obtained by querying the audio data stored locally, or the first voice content is sent by other terminals. This embodiment is described by taking the first voice content as the voice recorded by the user in real time as an example.

[0098] Exemplarily, such as Figure 4As shown, the user interface 41 is a video production interface. On the user interface 41, a recording control 401 is displayed. The user can click on the recording control 401 to start recording the first voice content. In addition, a video title and / or an interface title can also be displayed on the user interface 41. The video title is used to represent the main content of the video to be generated, and the interface title represents the main content of this user interface. After clicking on the recording control 401, the user interface 42 is displayed. The user interface 42 is a recording interface. The video title and / or the interface title displayed on the user interface 41 are retained and displayed on the user interface, and a reminder icon is displayed. The reminder icon is used to remind the user that the terminal is receiving the user's first voice content. The reminder icon is composed of text and / or graphics. A text material 403 can also be displayed on the interface 42. By converting a part of the input first voice content into text, the text material 403 is obtained, that is, the text material 403 is the part that has completed the conversion of voice to text. A pause control 404 is also displayed on the interface 42. The user can click on the pause control 404 to temporarily stop the input of the first voice content. A completion control 405 is also displayed on the interface 42. The user can click on the completion control 405 to complete the input of the first voice content.

[0099] Step 304: Display the first text material and the first multimedia material corresponding to the first voice content.

[0100] The first text material refers to the text obtained by recognizing the first voice content. Exemplarily, if the first voice content is "How to treat a cold", then the corresponding first text material is "How to treat a cold". The first text material is the text representation of the first voice content, and the semantics expressed by the first text material and the first voice content are the same.

[0101] Speech recognition can convert the voice content input by the user into text. There are various ways to achieve speech recognition. For example, a database corresponding to speech and text is established. When a piece of speech is input, the corresponding text is searched in the database. Or, a trained speech recognition neural network is used. This speech recognition neural network can output the input speech as text.

[0102] The first text material can be obtained by segmenting and translating the first voice content, or by translating the first voice content as a whole. Exemplarily, if the first voice content is "It's going to rain. Remember to bring an umbrella", when using segmental translation, after inputting the voice content "It's going to rain", this voice content is translated. After inputting the language content "Remember to bring an umbrella", this voice content is translated. The two translation results are spliced to obtain the first text material "It's going to rain. Remember to bring an umbrella". When using overall translation, after inputting the complete voice content "It's going to rain. Remember to bring an umbrella", this complete voice content is translated to obtain the first text material "It's going to rain. Remember to bring an umbrella".

[0103] The first multimedia material is obtained by searching based on at least one of the first voice content and the first text material, and there is a direct or indirect relationship between the first multimedia material, the first voice content and the first text material. The first multimedia material includes at least one of pictures, videos, and audios. Exemplarily, if the first text material is "cold", then the first multimedia material is a video on how to treat a cold, or a picture of a man sneezing.

[0104] Exemplarily, as Figure 4 shown, the user interface 42 is an intermediate user interface for voice recording. When the user finishes voice recording and clicks the completion control 405 on the user interface 42, the user interface 43 is displayed. The video title and / or interface title are still displayed on the user interface 43. A play control 406 is also displayed on the user interface 43. Clicking the play control 406 will play the first voice content recorded by the user. A progress bar is also displayed at the peripheral position of the play control 406, and the progress bar is used to indicate the playing progress of the first voice content. The user can also slide the aforementioned progress bar to adjust the played content. A first text material 407 is also displayed on the user interface 43, which is obtained by performing voice recognition on the first voice content. A first multimedia material 408 is also displayed on the user interface 43. A recording control 401 is also displayed on the user interface for the user to record the next voice content. A completion control 414 is also displayed on the user interface 43. The function of the completion control 414 is different from that of the completion control 405 in the user interface 42. Here, the function of the completion control 414 is to complete the input of all voice content and prepare to generate a video. Optionally, a delete control 413 can also be displayed on the user interface 43. Clicking the delete control 413 can delete display elements such as the play control 406, the first text material 407, and the first multimedia material 408 displayed on the user interface, and the user interface 41 is redisplayed.

[0105] Step 306: In response to the video generation operation, display a video with a first video segment, where the first video segment is generated based on the first text material and the first multimedia material.

[0106] The video generation operation is used to generate a video with a first video segment based on the first text material and the first multimedia material. The video generation operation is to generate the video by pressing one or more preset physical buttons. The user can also perform the video generation operation through signals generated by releasing, long-pressing, clicking, double-clicking, and / or sliding on the touch screen.

[0107] Optionally, the first text material is displayed in the form of subtitles in the first video clip. The subtitles can be arranged horizontally at the lower or upper side of the first video clip, or vertically at the left or right side of the first video clip. The present application does not limit the specific display position of the subtitles.

[0108] A video is composed of multiple video frames and multiple audio frames. According to different first multimedia materials, the situations of generating the first video clip include but are not limited to the following:

[0109] 1. The first multimedia material includes: picture materials.

[0110] All or part of the video frames of the first video clip are generated based on the above picture materials, and all or part of the audio frames of the first video clip are generated based on the voice content.

[0111] 2. The first multimedia material includes: video materials.

[0112] All or part of the video frames of the first video clip are generated based on the above video materials, and all or part of the audio frames of the first video clip are generated based on the voice content.

[0113] 3. The first multimedia material includes: audio materials.

[0114] All or part of the audio frames of the first video clip are generated based on the above audio materials and voice content.

[0115] The generation methods of the video frames and audio frames of the first video clip in the present application are at least one or a combination of the above.

[0116] The first video clip refers to a video generated based on the first text material and the first multimedia material.

[0117] Optionally, in response to the video generation operation, a video with the first video clip and the first text material is displayed.

[0118] Optionally, the video is transcoded into videos with different resolutions and bitrates to adapt to different types of terminals.

[0119] Exemplarily, such as Figure 4As shown, the user clicks the completion control 414 on the user interface 43, and the user interface 44 is displayed. A video 410 is displayed on the user interface 44. The video 410 is generated according to the first multimedia material 408 and the first text material 407. A text material 411 is also displayed on the user interface 44. The text material 411 includes all or part of the content of the first text material 407. The first text material 407 may also include text other than the first text material. Optionally, a sharing control 412 is also displayed on the user interface 44. When the user clicks the sharing control 412, the generated video 410 and the text material 411 can be sent to other users or sent to a specified network space.

[0120] In summary, in this embodiment, the user only needs to record voice content and perform simple operations to quickly obtain a video, without the user spending a lot of time searching for materials, nor does the user need to have professional editing knowledge. This can not only reduce the learning cost of the user using video production software, but also eliminate the process of the user searching for materials, improving the efficiency of video production and the efficiency of human-computer interaction.

[0121] Figure 5 The flowchart of the video generation method provided by an exemplary embodiment of the present application is shown.

[0122] This method can be executed by Figure 2 the terminal 220 shown, and this method includes the following steps:

[0123] Step 501: Display a video production interface, and the video production interface includes a recording control.

[0124] The video production interface is the initial interface for video production, and the user starts to produce a video through this interface.

[0125] The recording control is used to start recording the user's voice.

[0126] Optionally, a video title and / or a video overview are also displayed on the video production interface.

[0127] Exemplarily, as Figure 1 shown, the user interface 11 is the video production interface, and a recording control 101 is displayed on the interface 11. An interface title or a video title can also be displayed.

[0128] Step 502: In response to a first recording operation on the recording control, record the first voice content.

[0129] The first recording operation refers to the operation of the user recording voice. The first recording operation is performed by pressing one or more preset physical buttons to record voice. The user can also perform the first recording operation through signals generated by releasing, long pressing, clicking, double clicking, and / or swiping on the touch screen.

[0130] In this embodiment, the first voice content refers to the voice recorded by the user in real time.

[0131] Exemplarily, as Figure 1 shown, the user clicks on the recording control 101 on the interface 11 to start recording the voice content. During the recording of the voice content, the interface of the terminal is displayed as the user interface 12, on which there are text reminders and graphic reminders to remind the user that the voice content is being recorded. There is also a pause control 102 displayed on the user interface 12 to temporarily stop the recording of the voice. There is also a completion control 103 displayed on the user interface 12 for completing the input of the first voice content.

[0132] Step 503: During the recording process, display the first text material obtained by performing speech recognition on the first voice content.

[0133] The first text material is obtained by performing speech recognition on the first voice content.

[0134] Speech recognition can convert the voice content input by the user into text. There are various ways to implement speech recognition. For example, a database corresponding to speech and text is established. When a piece of speech is input, the corresponding text is searched for in the database. Or, a trained speech recognition neural network is used, which can output the input speech as text.

[0135] Optionally, after the recording is completed, display the first text material obtained by performing speech recognition on the first voice content.

[0136] Exemplarily, as Figure 1 shown, on the user interface 12, the user is inputting the first voice content, and there is a first text material 105 displayed on the user interface 12. Here, the first text material is obtained by performing speech recognition on the first voice content that has been input.

[0137] Step 504: After the recording ends, display the first multimedia material corresponding to the first voice content.

[0138] The first multimedia material is searched for based on at least one of the first voice content and the first text material.

[0139] Optionally, after the recording ends, display at least one of the first multimedia material and the first text material.

[0140] Exemplarily, as Figure 1 shown, after the recording ends, the user interface 13 is displayed. On the user interface 13, there can also be a play control 104, a first text material 105, and a first multimedia material 106 displayed.

[0141] Optionally, after the recording ends, at least one of picture material, video material, and audio material corresponding to the first voice content is displayed.

[0142] Exemplarily, as Figure 7 shown, after the recording is completed, the user interface 71 is displayed, and in the user interface 71, multiple picture materials 701 are displayed.

[0143] Exemplarily, as Figure 8 shown, after the recording is completed, the user interface 81 is displayed, and in the user interface 81, multiple picture materials 801 and audio material 802 are simultaneously displayed.

[0144] Optionally, the first text material includes first speech recognition content and second speech recognition content, and the first speech recognition content and the second speech recognition content are obtained by performing speech recognition on different parts of the first voice content.

[0145] Optionally, after the first speech recognition content is recognized during the recording process, the first multimedia material corresponding to the first speech recognition content is displayed; after the second speech recognition content is recognized, the updated first multimedia material is displayed; wherein, the updated first multimedia material corresponds to the second speech recognition content, or the updated first multimedia material corresponds to the first speech recognition content and the second speech recognition content. Exemplarily, the first voice content input by the user is "On a sunny day, I go to the playground to run". After the user inputs the first voice content, the terminal first obtains the first speech recognition content "On a sunny day", and the terminal displays a picture related to "sunny day" on the user interface according to the first speech recognition content, such as a picture of the sun. When the terminal obtains the second speech recognition content "I go to the playground to run", the first multimedia material is updated, the original first multimedia material is cancelled from display, and the updated first multimedia material is displayed, for example, a picture of a person running.

[0146] Optionally, after the first speech recognition content is recognized during the recording process, the first multimedia material corresponding to the first speech recognition content is displayed; after the second speech recognition content is recognized, the first multimedia material corresponding to the second speech recognition content is displayed. Exemplarily, the user inputs the first voice content "On a sunny day, I go to the playground to run". When the terminal obtains the first speech recognition content "On a sunny day", the first multimedia material corresponding to the first speech recognition content is displayed on the user interface, such as a picture related to "sunny day". When the user inputs the second speech recognition content, the first multimedia material is kept displayed, and the first multimedia material corresponding to the second speech recognition content is displayed in other areas of the user interface, for example, a picture of a person running.

[0147] Step 505: Display a replacement control corresponding to the first multimedia material.

[0148] Exemplarily, as Figure 1 shown, a replacement control 107 is displayed at the bottom of the first multimedia 106 in the user interface 13, and the replacement control 107 is displayed with a text identifier of "Change to Another".

[0149] Step 506: In response to a replacement operation on the replacement control, display the alternative multimedia material as the first multimedia material.

[0150] The replacement operation is used for the user to replace the first multimedia material with the alternative multimedia material. The replacement operation is to execute the replacement of the first multimedia material by pressing one or more preset physical buttons, and the user can also execute the replacement operation through signals generated by releasing, long-pressing, clicking, double-clicking, and / or swiping on the touch screen.

[0151] The alternative multimedia material is obtained by searching based on at least one of the first voice content and the first text material.

[0152] Optionally, arrange at least one multimedia material according to the search results, set the first multimedia material in the arranged order as the first multimedia material, and set the remaining multimedia materials as alternative multimedia materials.

[0153] Optionally, the first multimedia material and the alternative multimedia materials are arranged according to the degree of association with the first text material, or according to the degree of association with the first voice content.

[0154] Optionally, the user can repeat step 506.

[0155] Exemplarily, the user can click the replacement control 107 in the user interface 13 to replace the first multimedia material 106 with the alternative multimedia material.

[0156] Step 507: In response to a second recording operation, record the second voice content.

[0157] The second recording operation refers to the operation of the user recording voice. The second recording operation refers to the operation of the user recording voice. The second recording operation is to execute the recording of voice by pressing one or more preset physical buttons, and the user can also execute the second recording operation through signals generated by releasing, long-pressing, clicking, double-clicking, and / or swiping on the touch screen. Optionally, the second recording operation is the same as or different from the first recording operation.

[0158] The second voice content refers to the voice recorded by the user in real time. Optionally, the second voice content is obtained by downloading through the network, or the second voice content is obtained by querying the audio data stored locally, or the second voice content is sent by other terminals. In this embodiment, the second voice content is taken as an example of being recorded by the user in real time for illustration.

[0159] Exemplarily, such as Figure 1 As shown, click the recording control 101 on the user interface 13 to record the second voice content. For the user interface during the recording process, please refer to the user interface 12, which will not be elaborated here.

[0160] Step 508: Display the second text material and the second multimedia material corresponding to the second voice content.

[0161] The second text material is obtained by performing speech recognition on the second voice content.

[0162] The second multimedia material is searched based on at least one of the second voice content and the second text material. Optionally, the second multimedia material includes at least one of pictures, videos, and audios.

[0163] Exemplarily, such as Figure 1 As shown, after the second voice content is recorded, the user interface 14 is displayed. On the interface 14, a play control 104, a second text material 108, and a second multimedia material 109 are displayed. A recording control 101 and a completion control 103 are also displayed on the user interface 14. The user can click the recording control 101 to record more voice content. When the user completes the input of the above voice content, click the completion control 103 to end the input of the voice content.

[0164] Step 509: In response to an editing operation on the first text material, display the edited first text material.

[0165] The editing operation is used to modify the first text material. Among them, the editing operation includes at least one of an operation of adding text, an operation of deleting text, an operation of searching for text, an operation of modifying text, an operation of replacing text, an operation of moving text, and an operation of changing the format.

[0166] The editing operation on the first text material and the first multimedia material can be performed before step 505. The present application does not limit the specific timing of the operation.

[0167] Optionally, in response to an editing operation on the second text material, display the edited second text material.

[0168] The user can repeat step 509.

[0169] Exemplarily, such as Figure 1As shown, after all the voice content is input, click the completion control 103 on the user interface 14 to display the user interface 15. On the user interface 15, a first text material 105, a first multimedia material 106, a second text material 108, and a second multimedia material 109 are displayed. The user can directly click on the first text material 105 or the second text material 108 on the user interface 15 to edit the text content therein. Optionally, more text materials or multimedia materials can also be displayed on the user interface 15, or fewer text materials or multimedia materials can be displayed. This application does not make any limitations in this regard.

[0170] Step 510: Display a deletion control corresponding to the first multimedia material.

[0171] The deletion control is used to cancel the display of the first multimedia material.

[0172] Optionally, display a deletion control corresponding to the second multimedia material.

[0173] Exemplarily, as Figure 6 shown, a deletion control 110 is displayed on the user interface 15. The deletion control 110 is superimposed and displayed above the first multimedia material 106.

[0174] Step 511: In response to a deletion operation on the deletion control, delete the first multimedia material and display an import control.

[0175] The deletion operation is to delete the first multimedia material by pressing one or more preset physical buttons. The user can also perform the deletion operation through signals generated by releasing, long-pressing, clicking, double-clicking, and / or swiping on the touch screen.

[0176] Optionally, in response to a deletion operation on the deletion control corresponding to the second multimedia material, delete the second multimedia material and display an import control.

[0177] Optionally, display an import control corresponding to the second multimedia material. In response to an import operation on the import control corresponding to the second multimedia material, import the second multimedia material. At this time, the first multimedia material and the second multimedia material will be displayed on the user interface at the same time. The user can also choose to display more multimedia materials.

[0178] Exemplarily, as Figure 6As shown, when the deletion control 110 on the user interface 15 is clicked, the user interface 61 is displayed, the first multimedia material 106 is deleted and the display is cancelled, and an import control 601 is displayed at the position of the original first multimedia material 106. The import control 601 can be displayed as a piece of text. For example, the text "Import other pictures" is displayed on the user interface 61, or it can be displayed in the form of a button. This application does not make any restrictions on this.

[0179] Step 512: In response to the import operation on the import control, display the imported multimedia material as the first multimedia material.

[0180] The import operation is used for the user to import the required multimedia material. The import operation is to execute the import of the multimedia material by pressing one or more preset physical buttons. The user can also execute the import operation through the signals generated by releasing, long-pressing, clicking, double-clicking, and / or swiping on the touch screen.

[0181] Optionally, in response to the deletion operation on the deletion control, delete the first multimedia material and display the imported multimedia material. At this time, there is no need for the user to perform an import operation, and the terminal can directly import and display the multimedia material on the default path.

[0182] Exemplarily, as Figure 6 shown, when the import control 601 is clicked, the user interface 62 is displayed, where in the user interface 62, the multimedia material is replaced from the first multimedia material 106 to the first multimedia material 602.

[0183] Optionally, steps 510 to 512 can be replaced with: display the deletion control corresponding to the second multimedia material; in response to the deletion operation on the deletion control, delete the second multimedia material and display the import control; in response to the import operation on the import control, display the imported multimedia material as the second multimedia material. After the replacement, steps 510 to 512 can be carried out together with the original steps 510 to 512.

[0184] The user can repeat steps 510 to 512.

[0185] Step 509 and steps 510 to 512 are not in a sequential order in terms of timing.

[0186] Step 513: In response to the video generation operation, display a video with a first video segment and a second video segment.

[0187] The video generation operation is used for the operation of generating a video. Optionally, the video generation operation is to click, double-click, or press a video generation control or a video completion control, or click, double-click, or press the keys of the physical keyboard to generate a video.

[0188] There is also a transition animation between the first video clip and the second video clip. The transition animation is used to connect the first video clip and the second video clip, so that the video screen is played more smoothly.

[0189] The second video segment is generated based on the second text material and the second multimedia material.

[0190] Optionally, a sharing control corresponding to the video with the first video clip is also displayed on the user interface, and in response to a sharing operation on the sharing control, the video with the first video clip is sent to other terminals; or, in response to a sharing operation on the sharing control, the video with the first video clip is sent to cyberspace.

[0191] For example, Figure 1 As shown, click the generate video control 111 on the user interface 15 to display the user interface 16, wherein the user interface 16 displays a video 112, the video 112 has a first video segment and a second video segment, and a text material 113 is displayed below the video 112, the text material 113 is obtained according to the first text material 105 and the second text material 108. The user interface 16 also displays a sharing control 114, and the user can share the video 112 and the text material 113 with other users or upload them to a preset network space through the sharing control 114.

[0192] In summary, this embodiment can reduce the learning cost of users using video production software, and also eliminate the process of users searching for materials, thereby improving the efficiency of video production and human-computer interaction.

[0193] Moreover, users can splice multiple videos into one video to extend the length of the video, enrich the content of the video, and further improve the efficiency of video production and human-computer interaction.

[0194] Furthermore, the user can edit the first multimedia material and the first text material to improve the quality of the video and make the content of the video closer to reality.

[0195] It should be noted in advance that the present application involves: a speech recognition process, a keyword extraction process, a multimedia material search process and a video production process.

[0196] Each of the above four processes can be implemented by the client or the server, and can be divided into at least the following possible implementation methods:

[0197] 1. The video production process is implemented by the client, and the speech recognition process, keyword extraction process and multimedia material search process are implemented by the server.

[0198] 2. The speech recognition process and the video production process are implemented by the client, and the keyword extraction process and the multimedia material search process are implemented by the server.

[0199] 3. The speech recognition process, the keyword extraction process, the multimedia material search process, and the video production process are all implemented by the client.

[0200] 4. The speech recognition process is implemented by the client, and the keyword extraction process, the multimedia material search process, and the video production process are implemented by the server.

[0201] 5. The speech recognition process and the keyword extraction process are implemented by the client, and the multimedia material search process and the video production process are implemented by the server.

[0202] 6. The speech recognition process, the keyword extraction process, and the video production process are implemented by the client, and the multimedia material search process is implemented by the server.

[0203] The following uses Figure 9 the following embodiments to introduce the above implementation method 1. The method includes the following steps:

[0204] Step 901: In response to the first recording operation, the terminal records the first speech content.

[0205] Step 902: The terminal sends the first speech content to the server.

[0206] Step 903: The server receives the first speech content.

[0207] The first text material is obtained by performing speech recognition on the first speech content.

[0208] The first multimedia material is searched based on at least one of the first speech content and the first text material.

[0209] The terminal receives the first text material and the first multimedia material replied by the server.

[0210] Step 904: The server performs speech recognition on the first speech content to obtain the first text material.

[0211] Step 905: The server extracts keywords from the first text material.

[0212] Keyword extraction refers to extracting words from text materials that can express the core idea. There is at least one keyword in a text material. Exemplarily, for the text material "How to treat a cold?", the keywords in this text material are "cold" and "treatment". There are various methods for keyword extraction. For example, input the text material into a keyword extraction neural network, whose function is to extract the keywords in the text material and output them. Or, based on the relationship between a large number of text materials and keywords, establish a corresponding database and retrieve the keywords from this database.

[0213] Step 906: The server searches for and obtains the first multimedia material based on the keyword.

[0214] The terminal sends the first text material to the server.

[0215] Optionally, the server uses a search engine to search for the first voice content and the first text material to obtain the first multimedia material. Or, the server queries the local memory to obtain the first multimedia material. In the above local memory, there is a corresponding relationship between the first voice material and the first multimedia material or a corresponding relationship between the first text material and the first multimedia material. This application does not make any limitations in this regard.

[0216] Step 907: The server sends the first text material and the first multimedia material to the terminal.

[0217] The terminal receives the first multimedia material replied by the server based on the first text material.

[0218] Step 908: The terminal receives the first text material and the first multimedia material sent by the server.

[0219] Speech recognition refers to converting the first voice content into the first text material, and the two express the same meaning. Exemplarily, the first voice content "How to treat a cold" is recognized as the first text material "How to treat a cold".

[0220] Step 909: Based on the first text material and the first multimedia material, the terminal generates a video.

[0221] This video has a first video segment, and the first video segment is generated according to the first text material and the first multimedia material.

[0222] Optionally, the recording operation is carried out in multiple times and generates multiple video segments. The transition animation between adjacent video segments is automatically generated or set by the user.

[0223] Step 910: The terminal displays the video.

[0224] The terminal sends the keyword to the server.

[0225] Exemplarily, as Figure 1 shown, a video 112 is displayed on the user interface 16.

[0226] In summary, this embodiment provides a method for generating a video. Since the processing capabilities of various terminals are different, in this embodiment, the video production process is processed at the terminal. When the performance of the terminal is poor, the method provided in this embodiment can still be implemented, reducing the processing pressure on the terminal.

[0227] Next, the following Figure 10 illustrated embodiment is used to introduce the above implementation manner 2. The method includes the following steps:

[0228] For the specific implementation processes in the following steps, reference can be made to steps 901 to 910. Although there will be differences in the implementation entities of the specific steps, it does not affect the specific implementation process.

[0229] Step 1001: In response to a first recording operation, the terminal records first voice content.

[0230] Step 1002: The terminal performs speech recognition on the first voice content to obtain first text material.

[0231] Step 1003: The terminal sends the first text material to the server.

[0232] Step 1004: The server receives the first text material.

[0233] Step 1005: The server extracts keywords from the first text material.

[0234] Step 1006: Based on the keywords, the server searches for and obtains first multimedia material.

[0235] Step 1007: The server sends the first multimedia material to the terminal.

[0236] Step 1008: The terminal receives the first multimedia material sent by the server.

[0237] Step 1009: Based on the first text material and the first multimedia material, the terminal generates a video.

[0238] Step 1010: The terminal displays the video.

[0239] In summary, in this embodiment, the speech recognition process and the video production process are implemented by the client, and the keyword extraction process and the multimedia material search process are implemented by the server. By having the terminal undertake part of the computing tasks, the pressure on the server can be effectively reduced.

[0240] The above two implementation manners can be easily thought of based on the above 2 embodiments, and will not be elaborated here one by one.

[0241] Figure 11 The flowchart shows the video generation method provided by an exemplary embodiment of the present application. This method is applied to a server and can be executed by the Figure 2 server 240 as shown. The method includes the following steps:

[0242] Step 1101: Obtain a first text material and a first multimedia material. The first text material is obtained by performing speech recognition on a first speech content, and the first multimedia material is searched based on at least one of the first speech content and the first text material. The first speech content is obtained by a first recording operation on a terminal.

[0243] The server obtains the first text material and the first multimedia material.

[0244] Optionally, the first text material and the first multimedia material are sent to the server by the terminal, or the first text material and the first multimedia material are pre-stored in the server.

[0245] Step 1102: Generate a video with a first video segment based on the first text material and the first multimedia material. The first video segment is generated based on the first text material and the first multimedia material.

[0246] Step 1103: Send the video with the first video segment to the terminal.

[0247] In summary, in this embodiment, the user only needs to record the speech content and perform simple operations to quickly obtain a video, without the need for the user to spend a lot of time searching for materials, nor does the user need to have professional editing knowledge. It can not only reduce the learning cost of the user using video production software, but also eliminate the process of the user searching for materials, improving the efficiency of video production and the human-computer interaction efficiency. Implementing the processing process in the server can reduce the computing pressure on the terminal. At the same time, since the processing power of the server is generally stronger than that of the terminal, the efficiency can be improved.

[0248] Exemplarily, list the service architecture of an exemplary server of the present application, as Figure 12 shown:

[0249] This service architecture can be divided into four layers, namely the data access layer 1201, the business logic layer 1202, the data access layer 1203, and the persistence layer 1204. The application accesses the server through the restful (a design style and development method of network application programs) interface in the data access layer 1201.

[0250] The data access layer 1001 includes a restful interface and an access server. Among them, the restful interface is used to access the application program, and after the application program is accessed, it receives and responds to requests; the role of the access server is to receive http (a simple request-response protocol) requests, convert the http requests into Grpc (Google remote procedure call, a remote procedure call method developed by Google), and call the business logic layer.

[0251] The business logic layer 1202 is implemented using Golang (a compiled language), exposes a Grpc service to the upper layer, and at the same time provides rpc (remote procedure call) calls internally. The business logic layer 1002 includes at least one of speech synthesis, speech conversion, image search, graphic synthesis, graphic abstract, and video generation. Among them, speech synthesis refers to generating corresponding speech content from text materials, which is the reverse process of speech conversion; speech conversion is to obtain corresponding text materials from speech content; image search refers to searching for corresponding images based on speech content and text materials; graphic synthesis refers to generating a corresponding article according to images and text materials, and the article includes the aforementioned images and text materials; graphic abstract refers to extracting keywords from text materials and forming corresponding abstracts; video generation refers to generating videos according to images and speech content. The business logic layer 1202 also includes at least one of service warning, content warning, user management, article management, and information configuration. Among them, service warning is used to detect whether there are risks in the services provided by the server; content warning is used to determine whether the content provided by the server meets the preset conditions; user management provides users with the permission to manage some or all of the content of the server; article management is used to manage the text materials stored in the server; information configuration is used to configure various types of information of the server content. The business logic layer also includes log records. Among them, the log records record the historical records of the server, which are convenient for users and technicians to consult. The business logic layer may also include other content, which is not specifically limited in this application.

[0252] The data access layer 1203 is used to provide the server with methods for accessing data. The data access layer 1203 includes at least one of an Aggregation Pipeline, an Application Programming Interface Caller (API Caller), a Mysql connection pool (Mysql is an open-source relational database management system, and Pool represents a connection pool for storing various connections), and a Redis connection pool (Redis is an open-source database, and Pool is the same as above, a connection pool). Among them, the Aggregation Pipeline is a data aggregation framework modeled based on the concept of a data processing pipeline, which can convert the input documents into aggregated results; the Application Programming Interface Caller is used to call the application corresponding to the interface; the Mysql connection pool is used to provide connections to Mysql; the Redis connection pool is used to provide connections to Redis. Optionally, the data access layer 1003 also includes at least one of database interaction, cache interaction, and application programming interface interaction.

[0253] The persistence layer 1204 is mainly used to store various types of data. The persistence layer includes at least one of a cluster and the Tencent Cloud Distributed File System (Tencent Cloud China Operating System, Tencent Cloud COS). Among them, a cluster is a group of independent computers interconnected by a high-speed network. They form a group and are managed in the mode of a single system. When a client interacts with a cluster, the cluster acts like an independent server, and a cluster includes a master node and slave nodes; the Tencent Cloud Distributed File System can provide distributed storage services.

[0254] In the message queue of this service architecture, there is at least one of a penetration testing tool (such as Sparta Nginx), a remote procedure call (such as Golong grpc, Golong is a programming language based on grpc), a read instruction (such as a Crontab instruction), and an open-source logging component (such as Logback). Among them, the Nginx penetration testing tool is used for port scanning; the read instruction can read the standard instructions of the input device and store them in a specified file for subsequent reading and execution; the open-source logging component can store logs.

[0255] Exemplarily, an exemplary architecture for converting voice content into text materials is given, as Figure 13 shown:

[0256] This architecture includes at least one of a service access layer 1301, a capability combination 1302, a speech recognition foundation 1303, a corpus part 1304, and an external capability 1305.

[0257] The service access layer 1301 is used to access other external applications. The service access layer 1301 includes at least one of Web Services (an independent application), the restful interface of the Hypertext Transfer Protocol, and the Software Development Kit (SDK). Among them, Web Services can provide a platform for data interaction or integration; the restful interface of the Hypertext Transfer Protocol can support both the Hypertext Transfer Protocol interface and the restful interface, facilitating data interaction; the Software Development Kit is used to provide application programming interfaces for applications.

[0258] The capability combination 1302 combines the services provided by the server, packages specific services, and provides a standard service interface externally. The capability combination 1302 includes at least one of different capability combination call application programming interfaces, speech recognition development application programming interfaces, and semantic understanding development application programming interfaces. Among them, the different capability combination call application programming interface is used to provide an interface that can implement multiple different services externally; the speech recognition development application programming interface is used to provide a speech recognition interface externally; the semantic understanding development application programming interface is used to provide a semantic understanding interface externally. The capability combination 1302 can also provide other standard service interfaces, which will not be elaborated here, and the present application does not make specific limitations on this.

[0259] The speech recognition foundation 1303 refers to various basic services that the server can provide. The speech recognition foundation 1303 includes at least one of text transcription, voiceprint recognition, recording segmentation, pinyin annotation, and silence detection. Among them, text transcription can transcribe text materials into text in other formats; voiceprint recognition is used to recognize the voiceprint corresponding to the speech content; recording segmentation is used to segment the speech content to facilitate speech recognition; pinyin annotation is used to annotate the corresponding pinyin on the text materials; silence detection is used to determine whether the speech content is in a silent state and whether speech recognition needs to be performed on it. The speech recognition foundation 1303 also includes services such as word segmentation, synonyms, annotation, grammar analysis, stop words, and pinyin retrieval. Among them, the word segmentation service is used to determine whether there is word segmentation in the speech content; the synonyms service is used to replace some or all of the text in the text materials with synonyms; the annotation service is used to annotate pinyin or annotations on the text materials; the grammar analysis service is used to analyze whether the grammar of the text materials is correct and correct it; the pinyin retrieval service is used to retrieve the pinyin of the text in the text materials.

[0260] The corpus part 1304 is used to store various types of corpora and services. The corpus part 1304 includes corpus resources and service resources. The corpus part includes at least one of a general corpus database, a professional corpus database, and a special corpus database. Among them, the general corpus database stores language materials for daily use; the professional corpus database stores language materials used in professional fields; the special corpus database stores language materials corresponding to some special vocabulary. The service resources include at least one of a general service database, an account service database, and a call service database. Among them, the general service database stores data required for common services, for example, data corresponding to display elements on a page; the account service database stores user account information; the call service database stores call interfaces corresponding to various services.

[0261] The external capability 1305 is used to provide services to other external application programs. The external capability 1305 includes at least one of business call interface management, user management, access service management, and statistical analysis management. Among them, the business call interface management is used to provide various services that can be realized by this architecture to the outside world, for example, individuals or combinations of services such as text transcription and pinyin annotation; the user management is used to provide a method for managing users externally; the access service management is used for other external application programs to access this architecture and call all or part of its functions; the statistical analysis management is used to count various data of this architecture, for example, the number of accesses, the number of uses, etc., and analyze them to obtain corresponding analysis results.

[0262] In this embodiment, the core of the architecture is to modularly distinguish various basic capabilities that need to be processed by voice big data, and define various modular external service interfaces, so that the processing of voice big data is more oriented to the business requirements of application software systems and analysis systems, and the value contained in the big data can be fully mined. It should be noted that semantic understanding technology is also a core technology in big data mining. In fact, if pure speech recognition technology is not fully integrated with semantic understanding technology, the effect of voice big data mining and application will be greatly reduced.

[0263] Exemplarily, Figure 14 FIG. shows an exemplary structural diagram of a background system provided by an exemplary embodiment of the present application.

[0264] The background system includes a front end 1401, a server end 1402, and a data end 1403. The overall background system adopts the LNMP architecture (L refers to Linux, a commonly used operating system; N refers to Nginx, a high-performance HTTP and reverse proxy web server; M refers to the Mysql database; P refers to PHP, Personal Home Page, a powerful server-side scripting language for creating dynamic interactive sites).

[0265] The front - end 1401 consists of HyperText Markup Language (HTML, a standard markup language for creating web pages), Cascading Style Sheets (CSS, a computer language used to present the style of HTML files), and JQUERY (a fast, small, and feature - rich JavaScript library that can traverse and manipulate HTML documents).

[0266] The server - end 1402 consists of PHP, NGINX, FFMPEG (an open - source computer program that can be used to record, convert digital audio and video, and convert them into streams), LINUX, graphic video processing software (such as Adobe After Effects, a graphic video processing software launched by Adobe), and Windows Server.

[0267] The data - end 1403 consists of a database (such as the MYSQL database).

[0268] Exemplarily, Figure 15 shows the Windows Server architecture provided by an exemplary embodiment of the present application.

[0269] In this architecture, the web client (front-end) 1501 is connected to the gateway (Gateway) 1505 through an API interface and is connected to the web client (back-end) 1504 through programming language instructions (page). Both the application platform 1502 and the mobile application 1503 are connected to the web client (back-end) 1504 through an API interface. The gateway 1505 is connected to the conditional random field algorithm-manager scalable nodes 1506 (conditional random field algorithm-managerscalable-nodes, crf-manager scalable-nodes) through an HTTP API. The conditional random field algorithm-manager scalable nodes 1506 include at least one of a resource-based network router (crf-manager scalable-nodes), a message push space (socketpush rooms, emit), an authentication center / security center (auth / security middleware), a storage service, a database, a file cache (storage service db, file, cache), the termination and establishment of a message queue (broker mqpublish), and a message queue list (task-queue scheduler). The conditional random field algorithm-manager scalable nodes 1506 are connected to the conditional random field algorithm-renderer scalable nodes 1507 (conditional random field algorithm-renderer scalable-nodes, crf-renderer scalable-nodes) through an AMQC remote procedure call (where AMQP refers to the Advanced Message Queuing Protocol, an application layer standard advanced message queue protocol that provides a unified message service). The conditional random field algorithm-renderer scalable nodes 1507 include at least one of a local task management template, jobs / subjobs (local task management templates, jobs / subjobs), a local configuration topic, channel (local configuration topic, channel), an event reservation, notification (event emergingsubscriber, notification), and a renderer adapter layer multiple engine support (renderer adapter layermultiple engine support).The Windows server architecture also includes a KV cache 1508 (Key-Value Cache, where KV refers to a design of computer cache) and a message queue 1509 (Message Queue). The Windows server architecture also includes a deployment infrastructure 1510, where the deployment infrastructure 1510 includes at least one of continuous integration lint, test, build (CIlint, test, build), version control staging production, tools for defining and running applications (such as docker compose), Tencent Cloud servers (qcloud services CloudVirtual Machine, qcloud services cvm), and load balance monitors (load balance monitor).

[0270] Exemplarily, Figure 16 The flowchart of a video synthesis method provided by an exemplary embodiment of the present application is shown. The method includes the following steps:

[0271] Step 1601: Standardize the first multimedia material, the first voice content, and the first text material.

[0272] Use FFMPEG to standardize the first multimedia material, the first voice content, and the first text material. The processing content includes at least one of size and encoding format.

[0273] Step 1602: Send the first multimedia material, the first voice content, and the first text material to the Windows server.

[0274] Send the first multimedia material, the first voice content, and the first text material to the video synthesis service on the Windows server in JSON format.

[0275] Step 1603: The Windows server returns a video with the first video clip.

[0276] Among them, the video synthesis service on the Windows server parses the JSON content, extracts the first multimedia material, the first voice content, the first text material, and the video configuration content, and synthesizes a video of the first video clip from the first multimedia material, the first voice content, and the first text material through the After Effect interface.

[0277] Optionally, execute the command line program aerender (a command line execution program of the professional video generation software After Effects) of After Effects, and call the template xml (Extensible Markup Language) to add corresponding special effects to the video, generating a video with the first video segment.

[0278] Exemplarily, the startup command is "aerender -project test.aepx -comp “test” -RStemplate “test_1” -Omtemplate “test_2” -output test.mov".

[0279] The specific parameter explanations are as follows:

[0280] The parameter project indicates that the current project template file is test.aepx;

[0281] The parameter comp indicates that the name of the compositor used in this composition is test;

[0282] The parameter RStemplate indicates that the name of the rendering template is test_1;

[0283] The parameter Omtemplate indicates that the name of the video output template is test_2;

[0284] The parameter output indicates that the name of the output video is test.mov.

[0285] Optionally, multiple segments of effects can be added to the After Effects template, and aerender is called through chained loops to repeatedly superimpose effects on the generated video.

[0286] Step 1604: Transcode the video with the first video segment.

[0287] Transcode the video with the first video segment through FFMPEG to generate a format adapted to each terminal.

[0288] Optionally, perform encoding, audio integration, and subtitle integration processing on the video with the first video segment through FFMPEG.

[0289] Optionally, after the aerender processing, perform further processing through FFMPEG to supplement the effect content, such as adding a title, adding an end credit, adding sound, video encoding, etc.

[0290] Step 1605: Send the transcoded video to the terminal.

[0291] In summary, this embodiment provides an exemplary way to implement video synthesis, offering a technical possibility. All the user needs to do is record the voice content and perform simple operations to quickly obtain a video, without the need for the user to spend a lot of time searching for materials or have professional editing knowledge. This can not only reduce the learning cost of the user using video production software, but also eliminate the process of the user searching for materials, improving the efficiency of video production and the human-computer interaction efficiency.

[0292] Figure 17 The block diagram of a video synthesis device provided by an exemplary embodiment of the present application is shown. The device 1700 includes:

[0293] A recording module 1701, configured to record first voice content in response to a first recording operation;

[0294] A display module 1702, configured to display a first text material and a first multimedia material corresponding to the first voice content, where the first text material is obtained by performing speech recognition on the first voice content, and the first multimedia material is searched based on at least one of the first voice content and the first text material;

[0295] In an alternative design of the present application, the display module 1702 is further configured to display the first multimedia material corresponding to the first speech recognition content after the first speech recognition content is recognized during the recording process; and display the updated first multimedia material after the second speech recognition content is recognized, where the updated first multimedia material corresponds to the second speech recognition content, or the updated first multimedia material corresponds to the first speech recognition content and the second speech recognition content.

[0296] In an alternative design of the present application, the display module 1702 is further configured to display the first multimedia material corresponding to the first speech recognition content after the first speech recognition content is recognized during the recording process; and display the first multimedia material corresponding to the second speech recognition content after the second speech recognition content is recognized.

[0297] The display module 1702 is further configured to display a video with a first video segment in response to a video generation operation, where the first video segment is generated based on the first text material and the first multimedia material.

[0298] In an alternative design of the present application, the display module 1702 is further configured to display the first text material obtained by performing speech recognition on the first voice content during the recording process.

[0299] In an alternative design of the present application, the display module 1702 is further configured to display the first multimedia material corresponding to the first voice content after the recording ends.

[0300] In an alternative design of the present application, the display module 1702 is further configured to display a video production interface, and the video production interface includes a recording control.

[0301] The recording module 1701 is further configured to record the first voice content in response to a first recording operation on the recording control.

[0302] In an alternative design of the present application, the first multimedia material includes: picture material; the video frames of the first video segment are generated based on the picture material, and the audio frames of the first video segment are generated based on the voice content; or, the first multimedia material includes: video material; the video frames of the first video segment are generated based on the video material, and the audio frames of the first video segment are generated based on the voice content; or, the first multimedia material includes: audio material; the audio frames of the first video segment are generated based on the voice content and the audio material.

[0303] In an alternative design of the present application, the display module 1702 is further configured to display the edited first text material in response to an editing operation on the first text material; wherein, the editing operation includes at least one of: an operation of adding text, an operation of deleting text, an operation of searching for text, an operation of modifying text, an operation of replacing text, an operation of moving text, and an operation of changing the format.

[0304] In an alternative design of the present application, the display module 1702 is further configured to display a replacement control corresponding to the first multimedia material.

[0305] The display module 1702 is further configured to display the alternative multimedia material as the first multimedia material in response to a replacement operation on the replacement control, and the alternative multimedia material is searched based on at least one of the voice content and the first text material.

[0306] In an alternative design of the present application, the display module 1702 is further configured to display a deletion control corresponding to the first multimedia material; in response to a deletion operation on the deletion control, delete the first multimedia material and display an import control; in response to an import operation on the import control, display the imported multimedia material as the first multimedia material.

[0307] In an alternative design of the present application, the recording module 1701 is further configured to record a second voice content in response to a second recording operation.

[0308] The display module 1702 is further configured to display the second text material and the second multimedia material corresponding to the second voice content, where the second text material is obtained by performing speech recognition on the second voice content, and the second multimedia material is searched based on at least one of the voice content and the second text material.

[0309] The display module 1702 is further configured to, in response to a video generation operation, display a video having a first video segment and a second video segment, where the second video segment is generated based on the second text material and the second multimedia material.

[0310] In an alternative design of the present application, there is also a transition animation between the first video segment and the second video segment.

[0311] In an alternative design of the present application, the first text material is displayed in the first video segment in the form of subtitles.

[0312] In an alternative design of the present application, the display module 1702 is further configured to display a sharing control corresponding to the video having the first video segment.

[0313] The device 1700 further includes:

[0314] A communication module 1703, configured to, in response to a sharing operation on the sharing control, send the video having the first video segment to other terminals; or, in response to a sharing operation on the sharing control, send the video having the first video segment to a network space.

[0315] In an alternative design of the present application, the communication module 1703 is further configured to send the first voice content to a server.

[0316] The communication module 1703 is further configured to receive the first text material and the first multimedia material replied by the server.

[0317] In an alternative design of the present application, the device 1700 further includes:

[0318] An identification module 1704, configured to perform speech recognition on the first voice content to obtain the first text material.

[0319] The communication module 1703 is further configured to send the first text material to a server.

[0320] The communication module 1703 is further configured to receive the first multimedia material replied by the server based on the first text material.

[0321] In an alternative design of the present application, the recognition module 1704 is further configured to perform speech recognition on the first speech content to obtain the first text material; and extract keywords from the first text material.

[0322] The communication module 1703 is further configured to send the keywords to the server; and receive the first multimedia material replied by the server based on the keywords.

[0323] In summary, in this embodiment, the user only needs to record the speech content and perform simple operations to quickly obtain a video, without the user spending a lot of time searching for materials, nor does the user need to have professional editing knowledge. This can not only reduce the learning cost of the user using the video production software, but also eliminate the process of the user searching for materials, improving the efficiency of video production and the human-computer interaction efficiency.

[0324] Figure 18 The block diagram of a video synthesis device provided by an exemplary embodiment of the present application is shown. The device 1800 includes:

[0325] An acquisition module 1801, configured to acquire a first text material and a first multimedia material, where the first text material is obtained by performing speech recognition on a first speech content, and the first multimedia material is searched based on at least one of the first speech content and the first text material; the first speech content is recorded by a first recording operation on the terminal;

[0326] A synthesis module 1802, configured to generate a video with a first video segment according to the first text material and the first multimedia material, where the first video segment is generated based on the first text material and the first multimedia material;

[0327] A sending module 1803, configured to send the video with the first video segment to the terminal.

[0328] In an alternative design of the present application, the acquisition module 1801 is further configured to receive the first speech content sent by the terminal.

[0329] The synthesis module 1802 is further configured to perform speech recognition on the first speech content to obtain the first text material; extract keywords from the first text material; and search for and obtain the first multimedia material based on the keywords.

[0330] In an alternative design of the present application, the acquisition module 1801 is further configured to receive the first text material sent by the terminal, where the first text material is obtained by the terminal performing speech recognition on the first speech content.

[0331] The synthesis module 1802 is further configured to extract the keyword from the first text material; and search for and obtain the first multimedia material based on the keyword.

[0332] In summary, in this embodiment, the user only needs to record voice content and perform simple operations to quickly obtain a video, without the user spending a lot of time searching for materials or having professional editing knowledge. This can not only reduce the learning cost of the user using video production software, but also eliminate the process of the user searching for materials, improving the efficiency of video production and the man-machine interaction efficiency.

[0333] Figure 19 FIG. is a schematic structural diagram of a server shown according to an exemplary embodiment. The server 1900 includes a central processing unit (CPU) 1901, a system memory 1904 including a random access memory (RAM) 1902 and a read-only memory (ROM) 1903, and a system bus 1905 connecting the system memory 1904 and the central processing unit 1901. The server 1900 further includes a basic input / output system (I / O system) 1906 for facilitating information transmission between various components within the computer device, and a mass storage device 1907 for storing an operating system 1913, application programs 1914, and other program modules 1915.

[0334] The basic input / output system 1906 includes a display 1908 for displaying information and an input device 1909 such as a mouse, keyboard, etc. for user input of information. Both the display 1908 and the input device 1909 are connected to the central processing unit 1901 through an input / output controller 1910 connected to the system bus 1905. The basic input / output system 1906 may further include an input / output controller 1910 for receiving and processing inputs from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1910 also provides outputs to a display screen, printer, or other types of output devices.

[0335] The large-capacity storage device 1907 is connected to the central processing unit 1901 through a large-capacity storage controller (not shown) connected to the system bus 1905. The large-capacity storage device 1907 and its associated computer-readable medium provide non-volatile storage for the server 1900. That is, the large-capacity storage device 1907 may include computer-readable media (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.

[0336] Without loss of generality, the computer-readable media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), CD-ROM, digital video disc (DVD), or other optical storage, magnetic tape cartridges, tapes, disk storage, or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media is not limited to the above several types. The above system memory 1904 and large-capacity storage device 1907 may be collectively referred to as memory.

[0337] According to various embodiments of the present disclosure, the server 1900 may also run by connecting to a remote computer device on the network through a network such as the Internet. That is, the server 1900 may be connected to the network 1911 through a network interface unit 1912 connected to the system bus 1905, or in other words, the network interface unit 1912 may also be used to connect to other types of networks or remote computer device systems (not shown).

[0338] The memory further includes one or more programs, and the one or more programs are stored in the memory. The central processing unit 1901 implements all or part of the steps of the above video synthesis method by executing the one or more programs.

[0339] In an exemplary embodiment, a computer-readable storage medium is further provided. At least one instruction, at least one segment of program, a code set, or an instruction set is stored in the computer-readable storage medium. The at least one instruction, the at least one segment of program, the code set, or the instruction set is loaded and executed by a processor to implement the video synthesis method provided in each of the above method embodiments.

[0340] This application further provides a computer-readable storage medium. At least one instruction, at least one segment of program, a code set, or an instruction set is stored in the storage medium. The at least one instruction, the at least one segment of program, the code set, or the instruction set is loaded and executed by a processor to implement the video synthesis method provided in the above method embodiments.

[0341] Optionally, this application further provides a computer program product or a computer program. The computer program product or the computer program includes computer instructions. The computer instructions are stored in a computer-readable storage medium. A controller reads the computer instructions from the computer-readable storage medium, and the controller executes the computer instructions to enable the display device to execute the video generation method described in the above aspects.

[0342] The serial numbers of the embodiments of the present application above are only for description and do not represent the advantages and disadvantages of the embodiments.

[0343] Those of ordinary skill in the art can understand that all or part of the steps for implementing the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium. The storage medium mentioned above can be a read-only memory, a magnetic disk, or an optical disc, etc.

[0344] The above are only optional embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A video generation method, characterized in that, The method includes: Recording first voice content in response to a first recording operation; Displaying first text material and first multimedia material corresponding to the first voice content, where a replacement control corresponds to the first multimedia material; In response to a replacement operation on the replacement control, replacing the first multimedia material with alternative multimedia material, where the first multimedia material and the alternative multimedia material are respectively the first multimedia material and the remaining multimedia material among at least two arranged multimedia materials, and the at least two multimedia materials are searched based on at least one of the first voice content and the first text material; In response to a video generation operation, displaying a video with a first video segment, where the first video segment is generated based on the first text material and the replaced first multimedia material, and the video generation operation is an operation for generating the video.

2. The method according to claim 1, wherein The displaying the first text material corresponding to the first voice content includes: During the recording process, displaying the first text material obtained by performing speech recognition on the first voice content.

3. The method according to claim 1, characterized in that, The displaying the first multimedia material corresponding to the first voice content includes: After the recording ends, displaying the first multimedia material corresponding to the first voice content.

4. The method according to claim 1, characterized in that, The first text material includes first speech recognition content and second speech recognition content, and the first speech recognition content is recognized before the second speech recognition content; The displaying the first multimedia material corresponding to the first voice content includes: After recognizing the first speech recognition content during the recording process, displaying the first multimedia material corresponding to the first speech recognition content; After recognizing the second speech recognition content, displaying updated first multimedia material; wherein the updated first multimedia material corresponds to the second speech recognition content, or the updated first multimedia material corresponds to the first speech recognition content and the second speech recognition content.

5. The method according to claim 1, characterized in that The first text material includes first speech recognition content and second speech recognition content, and the first speech recognition content is recognized before the second speech recognition content; The displaying the first multimedia material corresponding to the first voice content includes: After recognizing the first speech recognition content during the recording process, displaying the first multimedia material corresponding to the first speech recognition content; After recognizing the second speech recognition content, displaying the first multimedia material corresponding to the second speech recognition content.

6. The method according to any one of claims 1 to 5, characterized in that The replaced first multimedia material includes: picture material; the video frames of the first video segment are generated based on the picture material, and the audio frames of the first video segment are generated based on the first voice content; or, The replaced first multimedia material includes: video material; the video frames of the first video segment are generated based on the video material, and the audio frames of the first video segment are generated based on the first voice content; or, The replaced first multimedia material includes: audio material; the audio frames of the first video segment are generated based on the first speech content and the audio material.

7. According to the method as claimed in any one of claims 1 to 5, characterized in that The method further includes: displaying the edited first text material in response to an editing operation on the first text material; wherein the editing operation includes at least one of an operation of adding text, an operation of deleting text, an operation of searching for text, an operation of modifying text, an operation of replacing text, an operation of moving text, and an operation of changing format.

8. The method according to any one of claims 1 to 5, characterized in that, The method further includes: displaying a deletion control corresponding to the first multimedia material; deleting the first multimedia material and displaying an import control in response to a deletion operation on the deletion control; displaying the imported multimedia material as the first multimedia material in response to an import operation on the import control.

9. The method according to any one of claims 1 to 5, characterized in that, The method further includes: recording a second speech content in response to a second recording operation; displaying a second text material and a second multimedia material corresponding to the second speech content, where the second text material is obtained by performing speech recognition on the second speech content, and the second multimedia material is searched based on at least one of the second speech content and the second text material; The displaying a video with a first video segment in response to a video generation operation includes: displaying a video with a first video segment and a second video segment in response to a video generation operation, where the second video segment is generated based on the second text material and the second multimedia material.

10. A video generation method, characterized in that, Applied to a server, the method includes: obtaining a first text material and a first multimedia material, where the first text material is obtained by performing speech recognition on a first speech content; the first speech content is recorded by a first recording operation on a terminal; sending the first text material and the first multimedia material to the terminal, where the first text material and the first multimedia material are used for display on the terminal; the first multimedia material corresponds to a replacement control on the terminal, and the terminal is configured to replace the first multimedia material with an alternative multimedia material in response to a replacement operation on the replacement control, where the first multimedia material and the alternative multimedia material are respectively the first multimedia material and the remaining multimedia material among at least two arranged multimedia materials, and the at least two multimedia materials are searched based on at least one of the first speech content and the first text material; generating a video with a first video segment based on the first text material and the replaced first multimedia material in response to a video generation operation on the terminal, where the first video segment is generated based on the first text material and the replaced first multimedia material; sending the video with the first video segment to the terminal, where the video is used for display on the terminal.

11. A video generation device, characterized in that, The device includes: a recording module, configured to record a first speech content in response to a first recording operation; a display module, configured to display a first text material and a first multimedia material corresponding to the first speech content, where the first multimedia material corresponds to a replacement control; The display module is further configured to, in response to a replacement operation on the replacement control, replace the first multimedia material with an alternative multimedia material, where the first multimedia material and the alternative multimedia material are respectively the first multimedia material and the remaining multimedia materials among at least two arranged multimedia materials, and the at least two multimedia materials are obtained by searching based on at least one of the first speech content and the first text material; The display module is further configured to, in response to a video generation operation, display a video having a first video segment, where the first video segment is generated based on the first text material and the replaced first multimedia material, and the video generation operation is an operation for generating the video.

12. The device according to claim 11, characterized in that The display module is configured to display the first text material obtained by performing speech recognition on the first speech content during the recording process.

13. The device according to claim 11, characterized in that, The display module is configured to display the first multimedia material corresponding to the first speech content after the recording ends.

14. The device according to claim 11, characterized in that, The first text material includes first speech recognition content and second speech recognition content, and the first speech recognition content is recognized before the second speech recognition content; the display module is configured to, after recognizing the first speech recognition content during the recording process, display the first multimedia material corresponding to the first speech recognition content; and after recognizing the second speech recognition content, display the updated first multimedia material; where the updated first multimedia material corresponds to the second speech recognition content, or the updated first multimedia material corresponds to the first speech recognition content and the second speech recognition content.

15. The device according to claim 11, wherein The first text material includes first speech recognition content and second speech recognition content, and the first speech recognition content is recognized before the second speech recognition content; the display module is configured to, after recognizing the first speech recognition content during the recording process, display the first multimedia material corresponding to the first speech recognition content; After recognizing the second speech recognition content, display the first multimedia material corresponding to the second speech recognition content.

16. The device according to any one of claims 11 to 15, wherein The replaced first multimedia material includes: a picture material; the video frames of the first video segment are generated based on the picture material, and the audio frames of the first video segment are generated based on the first speech content; or, the replaced first multimedia material includes: a video material; the video frames of the first video segment are generated based on the video material, and the audio frames of the first video segment are generated based on the first speech content; or, the replaced first multimedia material includes: an audio material; the audio frames of the first video segment are generated based on the first speech content and the audio material.

17. The device according to any one of claims 11 to 15, characterized in that, The display module is further configured to display the edited first text material in response to an editing operation on the first text material; wherein, the editing operation includes at least one of an operation of adding text, an operation of deleting text, an operation of searching for text, an operation of modifying text, an operation of replacing text, an operation of moving text, and an operation of changing format.

18. The device according to any one of claims 11 to 15, characterized in that The display module is further configured to display a deletion control corresponding to the first multimedia material; and in response to a deletion operation on the deletion control, delete the first multimedia material and display an import control. In response to an import operation on the import control, display the imported multimedia material as the first multimedia material.

19. The device according to any one of claims 11 to 15, characterized in that The recording module is further configured to record second voice content in response to a second recording operation. The display module is further configured to display a second text material and a second multimedia material corresponding to the second voice content, where the second text material is obtained by performing speech recognition on the second voice content, and the second multimedia material is searched based on at least one of the second voice content and the second text material. The display module is further configured to display a video having a first video segment and a second video segment in response to a video generation operation, where the second video segment is generated based on the second text material and the second multimedia material.

20. A video generation device, characterized in that, The device includes: An acquisition module, configured to acquire a first text material and a first multimedia material, where the first text material is obtained by performing speech recognition on a first voice content; and the first voice content is recorded by a first recording operation on a terminal. A sending module, configured to send the first text material and the first multimedia material to the terminal, where the first text material and the first multimedia material are used for display on the terminal; the first multimedia material corresponds to a replacement control on the terminal, and the terminal is configured to replace the first multimedia material with an alternative multimedia material in response to a replacement operation on the replacement control, where the first multimedia material and the alternative multimedia material are respectively the first multimedia material and the remaining multimedia material among at least two arranged multimedia materials, and the at least two multimedia materials are searched based on at least one of the first voice content and the first text material. A synthesis module, configured to generate a video having a first video segment according to the first text material and the replaced first multimedia material in response to a video generation operation on the terminal, where the first video segment is generated based on the first text material and the replaced first multimedia material. The sending module is further configured to send the video having the first video segment to the terminal, where the video is used for display on the terminal.

21. A computer device, characterized in that, The computer device includes a processor and a memory, and the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the video generation method according to any one of claims 1 to 10.

22. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and the computer program is loaded and executed by a processor to implement the video generation method according to any one of claims 1 to 10.

23. A computer program product, characterized in that, The computer program product includes computer instructions, the computer instructions are stored in a computer-readable storage medium, a controller reads the computer instructions from the computer-readable storage medium, and the controller executes the computer instructions to implement the video generation method according to any one of claims 1 to 10.

Citation Information

Patent Citations

  • Media data display method, server and display device

    CN111885400A

  • Music short video generation method and device, electronic equipment and storage medium

    CN111935537A