Video generation method and electronic equipment

By analyzing the content input in the session, determining the material characteristics and matching candidate materials, the problem of high complexity in existing video generation technology is solved and a convenient video generation experience is achieved.

CN120281961APending Publication Date: 2025-07-08HONOR DEVICE CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202311867656.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing video generation technology has complex operation and poor user experience, making it difficult to generate videos quickly and easily.

Method used

By analyzing the content input in session, determining the material characteristics, matching candidate materials, generating target videos, simplifying the operation process, and optimizing user portrait data using natural language processing and learning memory library.

Benefits of technology

It reduces the operational complexity of video generation, improves convenience and user experience, and can quickly generate videos that meet users' intentions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120281961A_ABST
    Figure CN120281961A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video generation method and electronic equipment. The method comprises the following steps: acquiring session input content; analyzing the session input content to obtain a session analysis result; under the condition that the session analysis result indicates a video generation intention, determining a first material feature indicated by the session input content; based on the first material feature, displaying a first candidate material matched with the first material feature; in response to a received material slicing request, generating a target video based on the first candidate material; and displaying the target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the technical field of terminals, and in particular, to a video generation method and an electronic device. Background Art

[0002] Video is one of the main carriers for people to obtain information and enjoy entertainment in daily life. With the improvement of the hardware performance of the client and the continuous progress of artificial intelligence technology, the demand for quickly generating videos through electronic devices is increasing day by day.

[0003] In the related technologies of video generation, there are defects such as complex operations and poor convenience. Summary of the Invention

[0004] To solve the above technical problems, the present application provides a video generation method, an electronic device, and a computer-readable storage medium. In this method, in response to the acquired session input content, the session input content is parsed to obtain a session parsing result. When the session parsing result indicates a video generation intention, the first material feature indicated by the session input content is determined, and the first candidate material matching the first material feature is presented. Then, in response to receiving a material compilation request, a target video is generated based on the first candidate material and displayed. It can effectively reduce the operational complexity of video generation, effectively improve the convenience of video generation, and is beneficial to improving the user experience.

[0005] In a first aspect, an embodiment of the present application provides a video generation method applied to an electronic device. The method includes: acquiring session input content; parsing the session input content to obtain a session parsing result; when the session parsing result indicates a video generation intention, determining a first material feature indicated by the session input content; based on the first material feature, displaying a first candidate material matching the first material feature; in response to receiving a material compilation request, generating a target video based on the first candidate material; and displaying the target video.

[0006] When the session parsing result indicates a video generation intention, determining the first material feature, presenting the first candidate material matching the first material feature, and generating a target video based on the first candidate material. It is possible to directly and quickly generate a target video without complex operations, which can effectively improve the user experience.

[0007] According to the first aspect, when the session parsing result indicates a video generation intention, determining the first material feature indicated by the session input content includes: when the session parsing result indicates the video generation intention, extracting at least one of the following keyword in the session parsing result as the first material feature: time keyword, location keyword, person keyword, and event keyword.

[0008] Extracting the keywords in the session parsing result as the first material feature is beneficial to improving the matching degree between the presented first candidate material and the session input content, beneficial to generating a target video that better conforms to the user's intention, and can effectively improve the user experience.

[0009] According to the first aspect, or any implementation manner of the above first aspect, presenting the first candidate material that matches the first material feature based on the first material feature includes: searching for candidate image materials that match the first material feature according to each of the keywords in the first material feature; screening the candidate image materials to obtain the first candidate material; and displaying the thumbnail of the first candidate material. The candidate image materials include static image materials and dynamic image materials.

[0010] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: in the case where the candidate image materials cannot be matched based on the target keyword in the first material feature, determining a second material feature according to the target keyword and a preset learning memory library; and matching the candidate image materials according to the first material feature and the second material feature. The target keyword includes any keyword in the first material feature, and the learning memory library includes pre-stored user portrait data.

[0011] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: in the case where the second material feature cannot be determined according to the target keyword and the learning memory library, initiating an inquiry session for the target keyword; determining a third material feature in response to the received answer result for the inquiry session; and matching the candidate image materials according to the first material feature and the third material feature.

[0012] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: generating new user portrait data according to the inquiry session and the answer result; and writing the new user portrait data into the learning memory library.

[0013] According to the first aspect, or any implementation manner of the above first aspect, screening the candidate image materials to obtain the first candidate material includes: screening the candidate image materials based on at least one of the following indicators to obtain the first candidate material: user semantics, material size, material duration, image clarity, composition overlap degree, and time overlap degree.

[0014] Exemplarily, the candidate image materials can be roughly selected based on at least one of user semantics, material size, material duration, and image clarity to obtain an intermediate image material set; and the intermediate image material set can be finely selected based on the composition overlap degree and the time overlap degree.

[0015] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: in response to detecting a material editing operation, updating the first candidate material according to an editing instruction generated based on the operation. Wherein, the editing instruction includes at least one of a sequential update instruction, a deletion instruction, and a duration adjustment instruction.

[0016] After determining the first candidate material, at least one of the operations of sequential adjustment, deletion, and duration adjustment can be performed on it.

[0017] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: in response to detecting a material addition instruction, presenting a material database that marks each material in the first candidate material; in response to detecting a selection operation for an unmarked material in the material database, adding the corresponding material to the first candidate material to obtain an updated first candidate material.

[0018] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: in response to detecting an unmarking operation for a marked material in the material database, deleting the corresponding material from the first candidate material to obtain an updated first candidate material.

[0019] According to the first aspect, or any implementation manner of the above first aspect, the method further includes: presenting the updated first candidate material when there is an update to the first candidate material; in response to receiving the material compilation request, generating the target video based on the updated first candidate material.

[0020] According to the first aspect, or any implementation manner of the above first aspect, the generating the target video based on the first candidate material in response to receiving the material compilation request includes: in response to receiving the material compilation request, determining the target video theme according to the content features of each material in the first candidate material; generating the target video based on the first candidate material according to a video template matching the target video theme.

[0021] Determining the target video theme according to the content features of each material in the first candidate material and generating the target video according to a video template matching the target video theme. Generating the target video based on the video template can significantly reduce the complexity of video generation and can effectively improve the efficiency of video generation.

[0022] According to the first aspect, or any implementation manner of the above first aspect, the video template includes a plurality of placeholders, and the placeholders are used to indicate at least one of text, pictures, and videos; generating the target video based on the first candidate material according to the video template matching the target video theme includes: importing each material in the first candidate material to the position of the corresponding placeholder in the video template; rendering the image frames of the video template after importing each material to obtain the target video.

[0023] According to the first aspect, or any implementation manner of the above first aspect, rendering the image frames of the video template after importing each material to obtain the target video includes: identifying a renderer corresponding to a target placeholder in the video template; rendering the image frames of the video template after importing each material according to the rendering effect of the renderer to obtain the target video. Wherein, the target placeholder includes any placeholder in the video template.

[0024] According to the first aspect, or any implementation manner of the above first aspect, parsing the session input content in response to the obtained session input content to obtain a session parsing result includes: performing natural language processing on the session input content in response to the obtained session input content to obtain the session parsing result. The natural language processing includes at least one of the following processes: word segmentation processing, stemming extraction, part-of-speech tagging processing, part-of-speech restoration processing, named entity recognition processing, and chunk normalization processing.

[0025] According to the first aspect, or any implementation manner of the above first aspect, the session input is performed by triggering at least one of the following entrances of the electronic device: system global entrance, desktop entrance, function recommendation page, application embedded entrance, and application floating entrance.

[0026] According to the first aspect, or any implementation manner of the above first aspect, the system global entrance includes at least one of the following entrances: power button, voice input entrance, gesture recognition entrance, face recognition entrance, and fingerprint recognition entrance; the desktop entrance includes at least one of the following entrances: desktop quick options, preset function cards, and global search entrance.

[0027] In a second aspect, an embodiment of the present application provides an electronic device, including: one or more processors, a memory, and one or more computer programs, where the one or more computer programs are stored on the memory, and when the computer programs are executed by the one or more processors, the electronic device is caused to perform the following steps: obtaining session input content; parsing the session input content to obtain a session parsing result; in a case where the session parsing result indicates a video generation intention, determining a first material feature indicated by the session input content; based on the first material feature, displaying a first candidate material that matches the first material feature; in response to receiving a material compilation request, generating a target video based on the first candidate material; and displaying the target video.

[0028] The second aspect and any implementation manner of the second aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the second aspect and any implementation manner of the second aspect, reference may be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, and details are not described herein again.

[0029] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, including a computer program, which, when running on an electronic device, causes the electronic device to execute the instructions of the method in the first aspect or any possible implementation manner of the first aspect.

[0030] In a fourth aspect, an embodiment of the present application provides a computer program, which includes instructions for executing the method in the first aspect or any possible implementation manner of the first aspect.

[0031] In a fifth aspect, an embodiment of the present application provides a chip, which includes a processing circuit and transceiver pins. Wherein, the transceiver pins and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the method in the first aspect or any possible implementation manner of the first aspect to control the receiving pin to receive a signal and control the sending pin to send a signal. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 It is a schematic structural diagram of an exemplary electronic device;

[0033] Figure 2 It is a schematic software structure block diagram of an exemplary electronic device;

[0034] Figure 3 It is a schematic diagram of a triggering video generation method provided by an embodiment of the present application;

[0035] Figure 4 It is a schematic diagram of a process for determining a material feature provided by an embodiment of the present application;

[0036] Figure 5 Schematic diagram of the process for determining a person image provided by an embodiment of the present application;

[0037] Figure 6 Schematic diagram of the process for determining candidate materials provided by an embodiment of the present application;

[0038] Figure 7 Schematic diagram of the process for forming a material into a finished product provided by an embodiment of the present application;

[0039] Figure 8 Schematic diagram of the process for adding captions provided by an embodiment of the present application;

[0040] Figure 9 Schematic diagram of the process for replacing the video template of a finished video provided by an embodiment of the present application;

[0041] Figure 10 Schematic diagram of the process for replacing the music of a finished video provided by an embodiment of the present application;

[0042] Figure 11 Schematic diagram of the process for simultaneously replacing the music and video template of a finished video provided by an embodiment of the present application;

[0043] Figure 12 Schematic diagram of the process for deleting materials provided by an embodiment of the present application;

[0044] Figure 13 Schematic diagram of the process for updating the order of materials provided by an embodiment of the present application;

[0045] Figure 14 Schematic diagram of the process for adding materials provided by an embodiment of the present application;

[0046] Figure 15 Schematic diagram of the process for video generation provided by an embodiment of the present application. Detailed implementation manners

[0047] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without making creative efforts shall fall within the protection scope of the present application.

[0048] The term "and / or" in this article is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone.

[0049] In the description of the embodiments of the present application, the terms "first", "second", etc. in the specification and claims are used to distinguish different objects, rather than to describe a specific order of the objects. For example, the first target object and the second target object, etc. are used to distinguish different target objects, rather than to describe a specific order of the target objects.

[0050] In the embodiments of the present application, words such as "exemplarily" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design solution described as "exemplarily" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplarily" or "for example" is intended to present relevant concepts in a specific manner.

[0051] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more. For example, a plurality of processing units refers to two or more processing units; a plurality of systems refers to two or more systems.

[0052] Please refer to Figure 1 , Figure 1 which is a schematic structural diagram of the electronic device 100 provided by the embodiments of the present application. Optionally, the electronic device 100 may be referred to as a terminal or a terminal device. The specific product form of the electronic device 100 may be a smart terminal, such as a mobile phone, a tablet computer, a wearable device, an augmented reality / virtual reality device, a laptop computer, a vehicle-mounted device, a personal digital assistant (PDA), etc., which are electronic devices with a video generation function. Specifically, the functional modules involved in the present application may be deployed on the DSP chip of the relevant device, specifically an application program or software therein. A video generation function can be realized through software installation or upgrade, and through the cooperation of hardware calls.

[0053] It should be understood that Figure 1 the illustrated electronic device 100 is only an example of an electronic device, and the electronic device 100 may have more or fewer components than those shown in the figure, may combine two or more components, or may have different component configurations. Figure 1 The various components shown in

[0054] The electronic device 100 may include: a processor 110, an external memory interface 120, an internal memory 121, a universal serial bus (USB) interface 130, a charging management module 140, a power management module 141, a battery 142, an antenna 1, an antenna 2, a mobile communication module 150, a wireless communication module 160, an audio module 170, a sensor module 180, a key 190, a motor 191, an indicator 192, a camera 193, a display screen 194, and a subscriber identification module (SIM) card interface 195, etc. Among them, the sensor module 180 may include a pressure sensor, a gyroscope sensor, an acceleration sensor, a temperature sensor, a motion sensor, a barometric pressure sensor, a magnetic sensor, a distance sensor, a proximity light sensor, a fingerprint sensor, a touch sensor, an ambient light sensor, a bone conduction sensor, etc.

[0055] The processor 110 may include one or more processing units. For example, the processor 110 may include an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a controller, a memory, a video codec, a digital signal processor (DSP), a baseband processor, and / or a neural-network processing unit (NPU), etc. Among them, different processing units may be independent devices or integrated in one or more processors.

[0056] Among them, the controller may be the nerve center and command center of the electronic device 100. The controller may generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching and executing instructions.

[0057] A memory may also be provided in the processor 110 for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory.

[0058] The USB interface 130 is an interface that conforms to the USB standard specification, and may specifically be a Mini USB interface, a Micro USB interface, a USB Type C interface, etc.

[0059] The charging management module 140 is used to receive a charging input from a charger. The charger can be a wireless charger or a wired charger. While the charging management module 140 charges the battery 142, it can also supply power to the electronic device through the power management module 141. The power management module 141 is used to connect the battery 142, the charging management module 140, and the processor 110. The power management module 141 receives inputs from the battery 142 and / or the charging management module 140 and supplies power to the processor 110, the internal memory 121, the external memory, the display screen 194, the camera 193, and the wireless communication module 160, etc.

[0060] The wireless communication function of the electronic device 100 can be implemented by the antenna 1, the antenna 2, the mobile communication module 150, the wireless communication module 160, the modulation and demodulation processor, and the baseband processor, etc.

[0061] The antenna 1 and the antenna 2 are used to transmit and receive electromagnetic wave signals. Each antenna in the electronic device 100 can be used to cover a single or multiple communication frequency bands. Different antennas can also be multiplexed to improve the utilization rate of the antennas.

[0062] The mobile communication module 150 can provide solutions for wireless communications including 2G / 3G / 4G / 5G, etc. applied to the electronic device 100. The mobile communication module 150 can include at least one filter, switch, power amplifier, low noise amplifier (LNA), etc.

[0063] The wireless communication module 160 can provide solutions for wireless communications including wireless local area networks (WLAN) (such as wireless fidelity (Wi-Fi) networks), Bluetooth (BT), global navigation satellite system (GNSS), frequency modulation (FM), near field communication (NFC), infrared technology (IR), etc. applied to the electronic device 100.

[0064] In some embodiments, the antenna 1 of the electronic device 100 is coupled to the mobile communication module 150, and the antenna 2 is coupled to the wireless communication module 160, so that the electronic device 100 can communicate with the network and other devices through wireless communication technologies.

[0065] The electronic device 100 implements the display function through a GPU, a display screen 194, an application processor, etc. The processor 110 may include one or more GPUs, which execute program instructions to generate or change display information.

[0066] The display screen 194 is used to display images, videos, etc. The display screen 194 includes a display panel. In some embodiments, the electronic device 100 may include one or N display screens 194, where N is a positive integer greater than 1.

[0067] The electronic device 100 can implement the shooting function through an ISP, a camera 193, a video codec, a GPU, a display screen 194, an application processor, etc.

[0068] The ISP is used to process the data fed back by the camera 193. For example, when taking a photo, the shutter is opened, and light passes through the lens and is transmitted to the camera photosensitive element. The optical signal is converted into an electrical signal, and the camera photosensitive element transmits the electrical signal to the ISP for processing and converts it into an image visible to the naked eye.

[0069] The camera 193 is used to capture static images or videos. An object generates an optical image through the lens and projects it onto the photosensitive element. The photosensitive element can be a charge coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The photosensitive element converts the optical signal into an electrical signal, and then transmits the electrical signal to the ISP to convert it into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into an image signal in a standard RGB, YUV, etc. format. In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0070] Among them, the camera 193 can be located in the edge area of the electronic device, and can be an under-screen camera or a liftable camera. The camera 193 can include a rear camera and can also include a rear camera. The specific position and form of the camera 193 in the embodiments of the present application are not limited. The electronic device 100 may include cameras of one or more focal lengths. For example, cameras of different focal lengths may include a telephoto camera, a wide-angle camera, an ultra-wide-angle camera, or a panoramic camera, etc.

[0071] The external memory interface 120 can be used to connect an external memory card, such as a Micro SD card, to expand the storage capacity of the electronic device 100. The external memory card communicates with the processor 110 through the external memory interface 120 to implement the data storage function.

[0072] The internal memory 121 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 110 executes various functional applications and data processing of the electronic device 100 by running the instructions stored in the internal memory 121. For example, the electronic device 100 implements the task processing method in the embodiments of the present application. The internal memory 121 can include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.). The data storage area can store data created during the use of the electronic device 100 (such as audio data, a phone book, etc.). In addition, the internal memory 121 can include a high-speed random access memory, and can also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, a universal flash storage (UFS), etc.

[0073] The electronic device 100 can implement an audio function through the audio module 170 and an application processor, etc. For example, music playback, recording, etc.

[0074] The audio module 170 is used to convert digital audio information into an analog audio signal for output, and is also used to convert an analog audio input into a digital audio signal. The audio module 170 can also be used to encode and decode audio signals. In some embodiments, the audio module 170 can be disposed in the processor 110, or some functional modules of the audio module 170 can be disposed in the processor 110.

[0075] The touch sensor, also known as the "touch panel". The touch sensor can be disposed on the display screen 194, and the touch sensor and the display screen 194 form a touch screen, also known as the "touch screen". The touch sensor is used to detect a touch operation acting thereon or nearby. The touch sensor can transmit the detected touch operation to the application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194.

[0076] The pressure sensor is used to sense a pressure signal and can convert the pressure signal into an electrical signal. In some embodiments, the pressure sensor can be disposed on the display screen 194. The electronic device 100 can also calculate the position of the touch according to the detection signal of the pressure sensor.

[0077] The gyroscope sensor can be used to determine the motion posture of the electronic device 100. In some embodiments, the angular velocity of the electronic device 100 around three axes (i.e., the x, y, and z axes) can be determined through the gyroscope sensor.

[0078] The acceleration sensor can detect the magnitude of the acceleration of the electronic device 100 in various directions (generally three axes). When the electronic device 100 is stationary, the acceleration sensor can detect the magnitude and direction of gravity. The acceleration sensor can also be used to identify the attitude of the electronic device and is applied to applications such as horizontal and vertical screen switching and pedometers.

[0079] The button 190 includes a power-on button (or power button), volume buttons, etc. The button 190 can be a mechanical button or a touch button. The electronic device 100 can receive button inputs and generate key signal inputs related to the user settings and function control of the electronic device 100.

[0080] The software system of the electronic device 100 can adopt a layered architecture, event-driven architecture, microkernel architecture, microservices architecture or cloud architecture. In the embodiments of the present invention, taking the Android system with a layered architecture as an example, the software structure of the electronic device 100 is exemplarily described.

[0081] Such as Figure 2 As an exemplary software structure block diagram of the electronic device 100 shown, the layered architecture of the electronic device 100 divides the software into several layers, and each layer has a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into four layers, from top to bottom, namely the application layer, application framework layer, Android runtime, system layer and kernel layer.

[0082] The application layer can include a series of application packages. Such as Figure 2 shown, the application packages can include applications such as a gallery, creation assistant, quick application engine, light editing service, etc. The creation assistant can, for example, parse the content of the user's session input and call other applications or middleware to generate videos according to the session parsing results. The quick application engine can, for example, provide a JS (JavaScript) card download service. The light editing service can, for example, generate videos based on the searched static image materials and / or video materials.

[0083] The application framework layer provides application programming interfaces (Application Programming Interface, API) and programming frameworks for the applications in the application layer, and includes various components and services to support developers' Android development. The application framework layer includes some predefined functions. Such as Figure 2 shown, the application framework layer can include a window manager, content provider, notification manager, resource manager, search engine, learning memory library and visual image module, etc.

[0084] A search engine can be used to search for image materials that match the session input content. The user portrait data of the device owner is stored in the learning memory bank. The visual image module can determine a video theme for video generation based on the content features of the image materials, for example.

[0085] The image materials can include static image materials and dynamic image materials. The static image materials can be images that remain unchanged within a preset time period. Such image materials are usually captured or created instantaneously and do not exhibit dynamic characteristics that change over time. The static image materials can be, for example, photos, paintings, charts, icons, etc. The dynamic image materials (such as videos or animations) contain change information in the time dimension and can present motion or changes through consecutive image frames.

[0086] It should be noted that the image materials involved in the embodiments of the present application can be obtained by an electronic device through shooting, or downloaded from a server by the electronic device, or received from other electronic devices by the electronic device. The embodiments of the present application do not make any limitations in this regard.

[0087] The window manager is used to manage window programs. The window manager can obtain the display screen size, determine whether there is a status bar, lock the screen, capture the screen, etc.

[0088] The content provider is used to store and obtain data and enable these data to be accessible by application programs. The data can include videos, images, audio, incoming and outgoing calls, browsing history and bookmarks, phone books, etc.

[0089] The resource manager can provide various resources for application programs, such as localized strings, icons, pictures, layout files, video files, and so on.

[0090] The notification manager enables application programs to display notification information in the status bar. It can be used to convey notification-type messages, which can automatically disappear after a short stay without user interaction. For example, the notification information is used to inform that the download is completed, message reminders, etc. The notification information can also be a notification that appears in the system top status bar in the form of a chart or scroll bar text, such as the notification of a background-running application program, or a notification that appears in the form of a dialogue window on the screen. For example, prompt text information is displayed in the status bar, a prompt sound is emitted, the electronic device vibrates, the indicator light flashes, etc.

[0091] Libraries and the Android Runtime. The system libraries can include multiple functional modules, such as an image rendering library, an image composition library, function libraries, and a media library, etc. The Android Runtime includes core libraries and a virtual machine, and the Android Runtime is responsible for the scheduling and management of the Android system. The core libraries contain two parts: one part is the functional functions that need to be called by the Java language, and the other part is the core libraries of Android. The application layer and the application framework layer run in the virtual machine, and the virtual machine executes the Java files of the application layer and the application framework layer as binary files. The virtual machine is used to perform functions such as object lifecycle management, stack management, thread management, security and exception management, and garbage collection.

[0092] It can be understood that Figure 2 The components included in the illustrated system framework layer, system libraries, and runtime layer do not constitute a specific limitation on the electronic device 100. In other embodiments of the present application, the electronic device 100 may include more or fewer components than shown in the figure, or combine certain components, or split certain components, or have different component arrangements.

[0093] The kernel layer is the layer between the hardware and the above software layer. The kernel layer at least includes a display driver, a camera driver, and a sensor driver. The hardware may include devices such as a camera, a display screen, a microphone, a processor, and a memory, etc.

[0094] The embodiments of the present application provide a video generation method, which can be applied to an electronic device. The electronic device obtains session input content, parses the session input content, and obtains a session parsing result. When the session parsing result indicates a video generation intention, the electronic device determines a first material feature indicated by the session input content. The electronic device presents a first candidate material that matches the first material feature, and in response to receiving a material composition request, the electronic device generates a target video based on the first candidate material and displays it. It can effectively reduce the operation complexity of video generation, effectively improve the convenience of video generation, and is beneficial to improving the user experience.

[0095] The electronic device may include a mobile phone, a tablet computer, a smart watch, a laptop computer, a virtual-reality fusion device, a super mobile personal computer, a smart TV, a smart screen, a high-definition TV, a smart speaker, a smart projector, etc. The embodiments of the present application can be applied to various scenarios that require video generation.

[0096] Before explaining the technical solutions of the embodiments of the present application, first, the triggering methods of video generation in the embodiments of the present application will be explained with reference to the accompanying drawings. Exemplarily, video generation can be triggered through system global entrances, desktop entrances, function recommendation pages, application embedded entrances, application floating entrances, etc. The desktop entrances may include, for example, creation assistant cards, global search, desktop cards, and other entrances.

[0097] Figure 3 A schematically shows triggering video generation through the system global entry. For example, the power button 101 can be physically touched in a preset manner to trigger video generation, or video generation can be triggered through the voice input entry 102, or video generation can also be triggered through methods such as breath wake-up, fingerprint recognition, gesture recognition, and face recognition.

[0098] Exemplarily, in response to detecting a triggering operation on the power button based on a preset manner, the creation assistant takes the default session as the user's session input content. The default session can be, for example, "Please perform video generation". In response to detecting the user's voice input operation, the creation assistant takes the detected voice input content as the user's session input content. The voice input content can be, for example, "Generate a video of the child in the past five years". Or, in response to detecting the user's text input operation, the creation assistant takes the detected text input content as the user's session input content. The text input content can be, for example, "Generate a video of the second daughter's birthday".

[0099] Figure 3 B - 3E schematically show triggering video generation through the desktop entry. The desktop entry can include, for example, desktop quick options, preset function cards, and the global search entry. As Figure 3 shown in B, video generation can be triggered through the desktop quick option 103 (such as the YOYO assistant). As Figure 3 shown in C, video generation can be triggered through the creation assistant card 104. The creation assistant can be, for example, the YOYO assistant. As Figure 3 shown in D, video generation can be triggered through the global search entry 105. For example, enter "Intelligent Video Composition" in the global search entry 105 to perform video generation. As Figure 3 shown in E, video generation can be triggered through the intelligent video composition card 106.

[0100] Figure 3 F schematically shows triggering video generation through the function recommendation page. For example, video generation is triggered through the intelligent video composition page 107. Figure 3 G schematically shows triggering video generation through the application embedded entry. For example, video generation is triggered through the intelligent video composition embedded entry 108. Figure 3 H schematically shows triggering video generation through the application floating entry. For example, video generation is triggered through the floating entry 109 in the gallery application.

[0101] The following combines Figures 4 to 15 , and schematically illustrates the video generation process of the embodiments of the present application.

[0102] Figure 4The figure shows a schematic diagram of the process for determining material features in an embodiment of the present application. JS (JavaScript) cards are widely used in mobile applications and web applications. A JS card can be an independent module containing specific information, functions, or operations. JS cards can make the user interface easy to understand and navigate. In addition, JS cards can also contain interactive elements such as buttons, links, or forms, etc., so that users can perform corresponding operations based on the interactive elements. The use of JS cards requires the introduction of corresponding JavaScript files and ensures that they can be correctly loaded. As Figure 4 shown, the service management framework initiates a card registration service to the media processing middleware, and the media processing middleware returns a card registration result to the service management framework. In the case where the card registration result indicates successful registration, in the user interface layout, information, images, or other content can be displayed through JS cards, or interactive elements can be provided through JS cards so that users can perform corresponding operations based on the interactive elements.

[0103] In response to the obtained session input content, the creation assistant parses the session input content to obtain a session parsing result. The session input content can include, for example, content in text form and voice form. Exemplarily, in response to the obtained session input content, the creation assistant performs natural language processing on the session input content to obtain a session parsing result. Natural language processing includes at least one of the following processes: word segmentation processing, stemming, part-of-speech tagging processing, lemmatization processing, named entity recognition processing, and chunking normalization processing. In natural language processing, chunking normalization processing includes chunking and normalization processing for text. Chunking is the process of splitting text into smaller chunks or phrases, and chunking is usually used to mark and extract specific information in the text, such as named entities (person names, place names, etc.) or phrase chunks. After chunking, further normalization processing can be performed, such as case conversion, spelling correction, synonym replacement, etc., to achieve a unified representation of the text.

[0104] In the case where the session parsing result indicates a video generation intention, the creation assistant initiates a process initialization request to the media processing middleware, and the media processing middleware creates and starts a finished video service. The creation assistant can extract time keywords, location keywords, person keywords, and event keywords in the session parsing result as material features for video generation.

[0105] In the case where the creation assistant cannot determine the time keyword based on the session input content, the creation assistant can query the corresponding time information from the learning memory bank, and determine the time keyword for searching image materials based on the queried time information. In the case of a query failure, the creation assistant can initiate an inquiry session for the time keyword, and the content of the inquiry session can include, for example, content in text form and voice form. The creation assistant can determine the time keyword based on the user's response result, as a material feature for video generation. In addition, the creation assistant can also write the corresponding time information into the learning memory bank for the learning memory bank to improve the user portrait data based on the time information.

[0106] Similarly, in the case where the creation assistant cannot determine the location keyword based on the session input content, the creation assistant can query the corresponding location information from the learning memory bank, and determine the location keyword for searching image materials based on the queried location information. In the case of a query failure, the creation assistant can initiate an inquiry session for the location keyword, and determine the location keyword for searching image materials based on the user's response result. The creation assistant can also write the corresponding location information into the learning memory bank for the learning memory bank to improve the user portrait data based on the location information. In the case where the creation assistant cannot determine the event keyword based on the session input content, the creation assistant queries or inquires about the corresponding event information based on a similar process, which will not be elaborated here.

[0107] The time keyword, location keyword, person keyword, and event keyword for searching image materials can be determined in any order, and this application does not limit this.

[0108] For example, the creation assistant parses the session input content "Generate a video of my daughter when she was three years old in her hometown", and obtains the session parsing results "daughter", "three years old", "in her hometown", "video". The creation assistant cannot determine the time keyword for searching image materials based on the session input content. The creation assistant can query the age information of the machine owner's daughter from the learning memory bank, and determine which year the machine owner's daughter was three years old based on the queried age information. In the case of a failure to query the age information, the creation assistant initiates an inquiry session "May I ask which year your daughter was three years old". The creation assistant determines the time keyword "2022" for searching image materials based on the user's response result "2022". In addition, the creation assistant can also write the time information "My daughter was three years old in 2022" into the learning memory bank so that the learning memory bank can improve the user portrait data based on the time information.

[0109] Figure 5 The process schematic diagram of determining the person image in the embodiment of the present application is shown, as Figure 5As shown, the creation assistant sends a person confirmation request to the service management framework, and the service management framework forwards the person confirmation request to the media processing middle platform. The media processing middle platform queries the person relationship from the learning memory bank. The learning memory bank determines the person relationship based on the user portrait data and returns the person relationship to the media processing middle platform. The media processing middle platform matches the person relationship label according to the confirmed person and queries the person image data from the image library according to the label. The image library returns the person image set to the media processing middle platform.

[0110] The media processing middle platform sends the person confirmation card URL (Uniform Resource Locator), the person image data, and the person label to the creation assistant. The creation assistant forwards the person confirmation card URL to the quick application engine, and the quick application engine returns the person confirmation card for the user to view the selected person images based on the person confirmation card and select the corresponding operation options through the person confirmation card.

[0111] Exemplarily, in response to the user clicking the "View All" option, the creation assistant displays all the thumbnails of the selected person images. In response to the user clicking the "Add" option, the creation assistant pulls up a more person selection page from the image library through the Internet. The image library displays the person selection page for the user to perform the operation of selecting a person. After the user completes the operation of selecting a person, the image library returns the person label and the person image data of the selected person to the creation assistant, and the creation assistant updates the person confirmation card based on the received person image data.

[0112] In response to the user clicking the "Label Modify" option, the creation assistant pulls up a person label change page from the image library through the Internet. The image library displays the person label modification page for the user to perform the operation of modifying the person label. After the user completes the operation of modifying the person label, the image library returns the modified person label to the creation assistant and updates the modified person label to the database. In response to the user clicking the "Confirm" option, the creation assistant saves the changed person label.

[0113] Figure 6 The process schematic diagram for determining candidate materials according to the embodiments of the present application is shown. As Figure 6 shown, the creation assistant sends a material search request to the service management framework, and the service management framework forwards the material search request to the media processing middle platform. The material search request includes the user semantics obtained based on the session input content. The user semantics includes semantic slots and semantic content. The important information units mentioned by the user when expressing intentions or needs constitute semantic slots. The semantic slots can be, for example, key parameters in a specific field or task, such as date, time, location, quantity, etc.

[0114] The media processing middleware sends a material search request to the search engine, and the search engine determines a material set that meets the semantics based on the material search request. The media processing middleware sequentially performs rough material selection and fine material selection on the material set that meets the semantics to obtain a candidate material set to be recommended. Exemplarily, the media processing middleware performs rough material selection on the material set that meets the semantics based on at least one of user semantics, material size, material duration, and image clarity to obtain an intermediate material set. The media processing middleware performs fine material selection on the intermediate material set based on the composition overlap degree and the time overlap degree to obtain a candidate material set.

[0115] The media processing middleware sends the candidate material set card URL, card ID, candidate material set, and session ID to the creation assistant. The creation assistant downloads the candidate material set card from the quick application engine based on the card URL. In response to the user clicking the "View All" option, the creation assistant displays the candidate material set and the session ID through the candidate material set card. In addition, the creation assistant pulls up the candidate material details page from the picture library through the Internet. The creation assistant simultaneously sends the session ID and the hash ID of the candidate material set to the picture library, and the picture library displays the candidate material details page. In response to the user clicking the "Add" option, the picture library displays a material browsing and selection page, and the material browsing and selection page marks each material in the candidate material set. In response to the user's operation of selecting a material, the picture library updates the selected material to the material details page.

[0116] Exemplarily, in response to the user's selection operation on the unmarked material in the material browsing and selection page, the corresponding material is added to the material details page to obtain an updated candidate material set. In addition, in response to the user's operation of canceling the mark on the marked material in the material browsing and selection page, the corresponding material is deleted from the material details page to obtain an updated candidate material set.

[0117] In response to the user clicking the "OK" option, the picture library sends the selected material of the user to the creation assistant. The creation assistant sends a material update request to the media processing middleware, and the material update request carries the session ID and the hash ID of the updated candidate material set. The media processing middleware performs material update based on the received material update request and returns the updated candidate material set card URL, card ID, candidate material set, and session ID to the creation assistant. The creation assistant downloads the updated candidate material set card from the quick application engine based on the card URL for the user to view the selected materials based on the updated candidate material set card and select corresponding operation options through the updated candidate material set card.

[0118] Figure 7 Shows a schematic diagram of the process of forming a material into a finished product in an embodiment of the present application, such as Figure 7As shown in the figure, the creation assistant sends a material compilation request to the service management framework, and the hash ID of the material to be compiled is carried in the material compilation request. The service management framework forwards the material compilation request to the media processing middle platform. The media processing middle platform queries the material content features from the picture library according to the hash ID of the material to be compiled, and the picture library returns the material content features to the media processing middle platform. The media processing middle platform obtains the video theme from the search engine according to the material content features, and the search engine returns the video theme to the media processing middle platform.

[0119] The media processing middle platform sends a compilation request to the light editing service. The light editing service compiles the video according to the video template that matches the video theme, and returns the compilation status to the media processing middle platform at a preset time interval. Among them, the compilation status includes the compilation percentage. In response to receiving the compilation status, the media processing middle platform sends the compilation waiting card URL, card ID, and compilation percentage to the creation assistant, and the creation assistant downloads the compilation waiting card from the quick application engine based on the card URL.

[0120] Exemplarily, the video template may include multiple placeholders, and the placeholders are used to indicate at least one of text, pictures, and videos. The light editing service imports each material in the candidate material set into the position of the corresponding placeholder in the video template, and renders the image frames of the video template after importing each material to compile the video and obtain the compiled video. Specifically, the light editing service identifies the renderer corresponding to the target placeholder in the video template, and renders the image frames of the video template after importing each material according to the rendering effect of the renderer to obtain the compiled video. The target placeholder includes any placeholder in the video template.

[0121] The light editing service returns that the compilation is completed to the media processing middle platform, and the media processing middle platform sends the compilation display card URL, card ID, video theme, and video cover to the creation assistant. The creation assistant downloads the compilation display card from the quick application engine based on the card URL for the user to view the compiled video.

[0122] Figure 8 The process schematic diagram of adding captions according to the embodiments of the present application is shown, as Figure 8 As shown in the figure, the creation assistant sends a caption generation request to the service management framework, and the service management framework forwards the caption generation request to the media processing middle platform. The media processing middle platform generates captions based on the caption generation request, and sends the caption card URL, session ID, card ID, and video theme to the creation assistant. The creation assistant downloads the caption card from the quick application engine based on the card URL for the user to view the captions based on the caption card and select corresponding operation options through the caption card.

[0123] In response to the user clicking the "Change" option, the creation assistant generates a blessing message based on the video theme and updates the caption based on the newly generated blessing message. In response to the user clicking the "Confirm" option, the creation assistant saves the updated caption.

[0124] The creation assistant sends a caption addition request to the service management framework. The caption addition request carries the session ID, caption, and finished video JSON. The service management framework forwards the caption addition request to the media processing middleware. The media processing middleware sends the caption addition request to the light editing service. The light editing service modifies the finished video based on the caption addition request and returns the modified finished video to the media processing middleware.

[0125] The media processing middleware sends the finished video display card URL, finished video JSON, and card ID of the modified finished video to the creation assistant. The creation assistant downloads the finished video display card from the quick application engine based on the card URL for the user to view the finished video based on the finished video display card and select corresponding operation options through the finished video display card.

[0126] Figure 9 The process schematic diagram of replacing the video template of the finished video in the embodiment of the present application is shown. As Figure 9 shown, the creation assistant sends a template replacement request to the service management framework. The template replacement request carries the session ID. The service management framework forwards the template replacement request to the media processing middleware. The media processing middleware sends a query theme template request to the light editing service. The light editing service returns a theme template set to the media processing middleware. The media processing middleware determines a video template with a high matching degree according to the video theme to obtain the replaced video template. The media processing middleware sends the template card URL, template list, and card ID to the creation assistant. The creation assistant downloads the template card from the quick application engine based on the card URL for the user to view the replaced video template based on the template card and select corresponding operation options through the template card.

[0127] In response to the user clicking the "OK" option, the creation assistant sends a template change request to the service management framework. The template change request carries the session ID and template type. The service management framework forwards the template change request to the media processing middleware. The media processing middleware sends the template change request to the light editing service. The light editing service changes the template of the finished video to obtain the finished video after template change and returns the finished video after template change to the media processing middleware. The media processing middleware sends the finished video display card URL, finished video JSON, theme, and card ID to the creation assistant. The creation assistant downloads the finished video display card from the quick application engine based on the card URL for the user to view the finished video after template change based on the finished video display card and select corresponding operation options through the finished video display card.

[0128] Figure 10The process schematic diagram of replacing the music of a finished video in an embodiment of the present application is shown. As Figure 10 shown, the creation assistant sends a music replacement request to the service management framework. The video theme and session ID are carried in the music replacement request. The service management framework forwards the music replacement request to the media processing middleware. The media processing middleware sends a music style query request to the light editing service, and the light editing service returns a music style set to the media processing middleware. The media processing middleware determines the music with high matching degree according to the video theme to obtain the replaced music. The media processing middleware sends the music card URL, music list and card ID to the creation assistant. The creation assistant downloads the music card from the quick application engine based on the card URL for the user to view the replaced music based on the music card and select corresponding operation options through the music card.

[0129] In response to the user clicking the "OK" option, the creation assistant sends a music change request to the service management framework. The session ID and music type are carried in the music change request. The service management framework forwards the music change request to the media processing middleware. The media processing middleware sends a music change request to the light editing service. The light editing service changes the music of the finished video to obtain the finished video after music change and returns the finished video after music change to the media processing middleware. The media processing middleware sends the finished video display card URL, finished video JSON, theme and card ID to the creation assistant. The creation assistant downloads the finished video display card from the quick application engine based on the card URL for the user to view the finished video after music change based on the finished video display card and select corresponding operation options through the finished video display card.

[0130] Figure 11 The process schematic diagram of simultaneously replacing the music and video template of a finished video in an embodiment of the present application is shown. As Figure 11 shown, the creation assistant sends a music and template replacement request to the service management framework. The video theme and session ID are carried in the music and template replacement request. The service management framework forwards the music and template replacement request to the media processing middleware. The media processing middleware sends a music style and theme template query request to the light editing service, and the light editing service returns a music style set and a theme template set to the media processing middleware. The media processing middleware determines the music with high matching degree and the video template with high matching degree according to the video theme to obtain the replaced music and the replaced video template. The media processing middleware sends the secondary editing function card URL, card ID, music set, template set and session ID to the creation assistant. The creation assistant downloads the secondary editing function card from the quick application engine based on the card URL.

[0131] In response to the user clicking the "Select" option and the "OK" option, the creation assistant sends a music and template change request to the service management framework. The music and template change request carries the session ID, template type, music type, and duration. The service management framework forwards the music and template change request to the media processing middleware. The media processing middleware sends the music and template change request to the light editing service. The light editing service changes the music and video template of the finished video to obtain the finished video with the changed music and video template. The light editing template returns the finished video with the changed music and video template to the media processing middleware. The media processing middleware sends the URL of the finished video display card, the finished video JSON, the theme, and the card ID to the creation assistant. The creation assistant downloads the finished video display card from the quick application engine based on the card URL.

[0132] Figure 12 The process schematic diagram of material deletion according to the embodiment of the present application is shown, as Figure 12 shown, in response to the obtained material deletion instruction, the creation assistant sends a material deletion request to the service management framework. The material deletion instruction may include, for example, a deletion instruction in voice form and a deletion instruction in text form. Correspondingly, the material deletion request carries the semantic session ID and the serial number session ID. The service management framework forwards the material deletion request to the media processing middleware.

[0133] In the case where the material deletion instruction is a deletion instruction in voice form, the media processing middleware sends a material deletion request to the search engine. The search engine returns the set of materials to be deleted to the media processing middleware. The media processing middleware deletes the materials based on the set of materials to be deleted. In the case where the material deletion instruction is a deletion instruction in text form, the media processing middleware deletes the materials based on the serial number to be deleted.

[0134] The media processing middleware sends the URL of the selected material set card, the session ID, the selected material set, the candidate material set hash-id, and the card ID to the creation assistant. The creation assistant downloads the selected material set card from the quick application engine based on the card URL. The creation assistant displays the selected material set card.

[0135] Figure 13 The process schematic diagram of material order update according to the embodiment of the present application is shown, as Figure 13 shown, in response to the obtained material order update instruction, the creation assistant sends a material order update request to the service management framework. The material order update instruction may include, for example, an instruction in voice form and an instruction in text form. The material order update request carries the materials to be updated quickly and the session ID. The service management framework forwards the material order update request to the media processing middleware.

[0136] The media processing center updates the material order according to the material order update request, and sends the selected material set card URL, session ID, selected material set, candidate material set hash-id and card ID to the creative assistant. The creative assistant downloads the selected material set card from the quick application engine based on the card URL. The creative assistant displays the selected material set card.

[0137] Figure 14 A schematic diagram of the process of adding materials in an embodiment of the present application is shown. Figure 14 As shown, in response to the acquired material adding instruction, the creation assistant sends a material adding request to the service management framework. The material adding instruction may include, for example, an instruction in voice form and an instruction in text form, and the material adding request carries a session ID, semantics, and slot. The service management framework forwards the material adding request to the media processing center.

[0138] The media processing center sends a material search request to the search engine, which carries semantics and slots. The search engine determines the material set that meets the semantics based on the material search request, and returns the material set that meets the semantics to the media processing center. The media processing center performs rough material selection and fine material selection on the material set that meets the semantics in turn to obtain the candidate material set to be added. The media processing center deduplicates and merges the candidate material set to be added with the original material set.

[0139] The media processing center sends the selected material set card URL, session ID, selected material set, candidate material set hash-id and card ID to the creative assistant. The creative assistant downloads the selected material set card from the quick application engine based on the card URL.

[0140] Figure 15 A schematic diagram of the video generation process of an embodiment of the present application is shown. Figure 15 As shown, the creation assistant parses the session input content in response to the acquired session input content to obtain the session parsing result. When the session parsing result indicates the video generation intention, the creation assistant initiates a process initialization request to the media processing center, and the media processing center creates and starts the filming service. The creation assistant determines the material features indicated by the session input content based on the session input content.

[0141] The creation assistant sends a person confirmation request to the service management framework, and the service management framework forwards the person confirmation request to the media processing middle platform. The media processing middle platform queries the person relationship from the learning memory library. The learning memory library determines the person relationship based on the user portrait data and returns the person relationship to the media processing middle platform. The media processing middle platform matches the person relationship tags according to the confirmed person and queries the person image data from the picture library. The picture library returns the person image set to the media processing middle platform. The media processing middle platform sends the person confirmation card URL, person image data, and person tags to the creation assistant. The creation assistant forwards the person confirmation card URL to the quick application engine, and the quick application engine returns the person confirmation card for the user to view the selected person images based on the person confirmation card and select the corresponding operation options through the person confirmation card.

[0142] In response to the user clicking the "Select" option and the "Confirm" option, the creation assistant sends a material search request to the service management framework, and the service management framework forwards the material search request to the media processing middle platform. The media processing middle platform sends a material search request to the search engine, and the search engine determines the material set that conforms to the semantics based on the material search request. The media processing middle platform sequentially performs rough material selection and fine material selection on the material set that conforms to the semantics to obtain the candidate material set to be recommended. The media processing middle platform sends the candidate material set card URL, card ID, candidate material set, and session ID to the creation assistant, and the creation assistant downloads the candidate material set card from the quick application engine based on the card URL.

[0143] In response to the user clicking the "Generate Video" option, the creation assistant sends a material into-video request to the service management framework. The service management framework forwards the material into-video request to the media processing middle platform. The media processing middle platform queries the material content features from the picture library according to the hash ID of the material to be made into a video, and the picture library returns the material content features to the media processing middle platform. The media processing middle platform obtains the video theme from the search engine according to the material content features, and the search engine returns the video theme set to the media processing middle platform.

[0144] The media processing middle platform sends an into-video request to the light editing service. The light editing service makes a video according to the video template that matches the video theme and returns the video-making status to the media processing middle platform at preset time intervals. In response to receiving the video-making status, the media processing middle platform sends the into-video waiting card URL, card ID, and into-video percentage to the creation assistant, and the creation assistant downloads the into-video waiting card from the quick application engine based on the card URL. The light editing service returns that the video-making is completed to the media processing middle platform, and the media processing middle platform sends the into-video display card URL, card ID, into-video JSON, video theme, video cover, and video duration to the creation assistant. The creation assistant downloads the into-video display card from the quick application engine based on the card URL for the user to view the into-video.

[0145] In response to the user clicking the "Play" option, the creation assistant launches the light editing service through an intent.

[0146] It can be understood that, in order for the electronic device to implement the above functions, it includes the corresponding hardware and / or software modules for executing each function. Combining the algorithm steps of each example described in the embodiments disclosed in this article, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving the hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application in combination with the embodiments, but such implementation should not be considered to exceed the scope of the present application.

[0147] All relevant contents of the steps involved in the above method embodiments can be cited in the function descriptions of the corresponding function modules and will not be repeated here.

[0148] This embodiment also provides an electronic device, including: one or more processors, a memory, and one or more computer programs, where one or more computer programs are stored on the memory, and when the computer programs are executed by one or more processors, the electronic device is caused to perform the following steps: obtaining session input content; parsing the session input content to obtain a session parsing result; in the case where the session parsing result indicates a video generation intent, determining a first material feature indicated by the session input content; presenting a first candidate material that matches the first material feature based on the first material feature; and generating and displaying a target video based on the first candidate material in response to receiving a material completion request.

[0149] This embodiment also provides a computer storage medium, including a computer program, and when the computer program runs on an electronic device, the electronic device is caused to execute the above related method steps to implement the video generation method in the above embodiments.

[0150] This embodiment also provides a computer program product, and when the computer program product runs on a computer, the computer is caused to execute the above related steps to implement the video generation method in the above embodiments.

[0151] In addition, an embodiment of the present application also provides a device, which may specifically be a chip, a component, or a module. The device may include a processing circuit and transceiver pins. Among them, the transceiver pins and the processing circuit communicate with each other through an internal connection path, and the processing circuit executes the video generation method in each of the above method embodiments to control the receiving pin to receive a signal and control the sending pin to send a signal.

[0152] Among them, the electronic device, computer storage medium, computer program product or chip provided in this embodiment are all used to execute the corresponding methods provided above. Therefore, the beneficial effects they can achieve can refer to the beneficial effects in the corresponding methods provided above, and will not be elaborated here.

[0153] From the description of the above embodiments, those skilled in the art can understand that for the convenience and simplicity of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0154] In the several embodiments provided in this application, it should be understood that the disclosed device and method can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0155] The unit described as a separated component may or may not be physically separated. The component displayed as a unit may be a physical unit or multiple physical units, that is, it can be located in one place, or it can be distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0156] In addition, each functional unit in the various embodiments of this application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0157] Any content in the various embodiments of this application, as well as any content in the same embodiment, can be freely combined. Any combination of the above content is within the scope of this application.

[0158] When an integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0159] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them belong to the protection scope of the present application.

[0160] The steps of the methods or algorithms described in connection with the disclosed content of the embodiments of the present application can be implemented in a hardware manner or by a processor executing software instructions. The software instructions can be composed of corresponding software modules. The software modules can be stored in a random access memory (RAM), flash memory, read only memory (ROM), erasable programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), registers, hard disks, mobile hard disks, CD-ROMs, or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC.

[0161] Those skilled in the art should be able to realize that in one or more of the above examples, the functions described in the embodiments of the present application can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium. The computer-readable medium includes computer storage media and communication media, where the communication media includes any medium that facilitates the transmission of a computer program from one place to another. The storage media can be any available medium accessible by a general-purpose or special-purpose computer.

[0162] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims, and all of them fall within the protection scope of the present application.

Claims

1. A video generation method, applied to an electronic device, characterized in that The method includes: Obtaining session input content; Parsing the session input content to obtain a session parsing result; When the session parsing result indicates a video generation intention, determining a first material feature indicated by the session input content; Based on the first material feature, displaying a first candidate material that matches the first material feature; In response to receiving a material editing request, generating a target video based on the first candidate material; and Displaying the target video.

2. The method according to claim 1, characterized in that, The step of, when the session parsing result indicates a video generation intention, determining a first material feature indicated by the session input content includes: When the session parsing result indicates the video generation intention, extracting at least one of the following keywords in the session parsing result as the first material feature: Time keyword, location keyword, person keyword, and event keyword.

3. The method according to claim 2, wherein The step of, based on the first material feature, presenting a first candidate material that matches the first material feature includes: Searching for candidate image materials that match the first material feature according to each keyword in the first material feature; Screening the candidate image materials to obtain the first candidate material; and Displaying a thumbnail of the first candidate material; Wherein, the candidate image materials include static image materials and dynamic image materials.

4. The method according to claim 2, characterized in that The method further includes: For a target keyword in the first material feature, when no candidate image material can be matched based on the target keyword, determining a second material feature according to the target keyword and a preset learning and memory library; and Matching the candidate image materials according to the first material feature and the second material feature, Wherein, the target keyword includes any keyword in the first material feature, and the learning and memory library includes pre-saved user portrait data.

5. The method according to claim 4, characterized in that The method further includes: When the second material feature cannot be determined according to the target keyword and the learning and memory library, initiating an inquiry session for the target keyword; In response to the received answer result for the inquiry session, determining a third material feature; and Matching the candidate image materials according to the first material feature and the third material feature.

6. The method according to claim 5, wherein The method further includes: Generating new user portrait data according to the inquiry session and the answer result; and Writing the new user portrait data into the learning and memory library.

7. The method according to claim 3, characterized in that, The step of screening the candidate image materials to obtain the first candidate material includes: Screening the candidate image materials based on at least one of the following indicators to obtain the first candidate material: User semantics, material size, material duration, image clarity, composition overlap, and time overlap.

8. The method according to claim 3, wherein The method further includes: In response to detecting a material editing operation, updating the first candidate material according to an editing instruction generated based on the operation, Wherein, the editing instruction includes at least one of a sequential update instruction, a deletion instruction, and a duration adjustment instruction.

9. The method according to claim 3, characterized in that, The method further includes: In response to detecting a material addition instruction, presenting a material database, where each material in the first candidate materials is marked; and In response to detecting a selection operation on an unmarked material in the material database, adding the corresponding material to the first candidate materials to obtain updated first candidate materials.

10. The method according to claim 9, wherein The method further includes: In response to detecting an unmarking operation on a marked material in the material database, deleting the corresponding material from the first candidate materials to obtain updated first candidate materials.

11. The method according to any one of claims 8 to 10, characterized in that The method further includes: When there is an update to the first candidate materials, presenting the updated first candidate materials; and In response to receiving the material video generation request, generating the target video based on the updated first candidate materials.

12. The method according to claim 1, characterized in that, The generating the target video based on the first candidate materials in response to receiving the material video generation request includes: In response to receiving the material video generation request, determining a target video theme according to the content characteristics of each material in the first candidate materials; and Generating the target video based on the first candidate materials according to a video template matching the target video theme.

13. The method according to claim 12, wherein The video template includes a plurality of placeholders for indicating at least one of text, pictures, and videos; The generating the target video based on the first candidate materials according to a video template matching the target video theme includes: Importing each material in the first candidate materials to the position of the corresponding placeholder in the video template; Rendering the image frames of the video template after importing each material to obtain the target video.

14. The method according to claim 13, wherein The rendering the image frames of the video template after importing each material to obtain the target video includes: Identifying a renderer corresponding to a target placeholder in the video template; Rendering the image frames of the video template after importing each material according to the rendering effects of the renderer to obtain the target video, where the target placeholder includes any placeholder in the video template.

15. The method according to claim 1, characterized in that The parsing the session input content in response to the obtained session input content to obtain a session parsing result includes: In response to the obtained session input content, performing natural language processing on the session input content to obtain the session parsing result, where the natural language processing includes at least one of the following processes: word segmentation processing, stemming extraction, part-of-speech tagging processing, part-of-speech restoration processing, named entity recognition processing, and chunk standardization processing.

16. The method according to claim 1, wherein The session input is performed by triggering at least one of the following entrances of the electronic device: System global entrance, desktop entrance, function recommendation page, application embedded entrance, and application floating entrance.

17. According to the method of claim 16, wherein The system global entrance includes at least one of the following entrances: power button, voice input entrance, gesture recognition entrance, face recognition entrance, and fingerprint recognition entrance; The desktop entrance includes at least one of the following entrances: desktop quick options, preset function cards, and global search entrance.

18. An electronic device, characterized in that, including: One or more processors, a memory, and one or more computer programs, wherein the one or more computer programs are stored in the memory and, when executed by the one or more processors, cause the electronic device to perform the following steps: Obtain session input content; Parse the session input content to obtain a session parsing result; When the session parsing result indicates a video generation intention, determine a first material feature indicated by the session input content; Based on the first material feature, display a first candidate material that matches the first material feature; And In response to receiving a material compilation request, generate a target video based on the first candidate material; And Display the target video.

19. A computer-readable storage medium, characterized in that, A computer program which, when running on an electronic device, causes the electronic device to perform the video generation method according to any one of claims 1 to 17.

Citation Information

Cited By

  • Method, device, equipment and product for generating video

    CN120935430A

  • Methods, apparatuses, devices, and products for generating video

    CN120935430B