Video generation method and device, electronic equipment, storage medium and program product

By introducing multiple input identifiers and data types to the video generation system, the problem of insufficient video expression caused by a single material input is solved, and the quality of video generation and creative expression are improved.

CN120455799APending Publication Date: 2025-08-08BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510725671.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

In the existing video generation technology, the single-modal material input method leads to insufficient video expression and lack of attractiveness, and the richness of video content and creative space are limited, making it difficult to meet users' needs for high-quality creative videos.

Method used

Displays multiple types of input identifiers, including links, multimedia and files, in the input area, through which multiple types of input data are indicated and parsed to generate target video.

Benefits of technology

Through the flexible combination of multimodal materials, the quality, flexibility and creative expression of video generation are improved, the needs of complex and changeable application scenarios are met, and the efficiency and effect of video generation are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120455799A_ABST
    Figure CN120455799A_ABST
Patent Text Reader

Abstract

The invention provides a video generation method and device, electronic equipment, a storage medium and a program product. The method comprises the following steps: in response to an input operation for an input area in a display interface, displaying input content in the input area; wherein the input content comprises at least one prompt text and a plurality of input identifiers, and the plurality of input identifiers are used for indicating various types of input data; the input data of the multiple types are analyzed, and corresponding description information is obtained; according to the sequence of the prompt text and the input identifier in the input content, generating a target text from the corresponding description information; and generating a target video based on the target text and the multiple types of input data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of video generation, and in particular to a video generation method, device, electronic device, storage medium, and program product. Background Art

[0002] In video generation technologies, users are often limited to generating videos by uploading footage from a single modality. However, this single input method significantly limits the expressiveness of the resulting videos, resulting in a lack of overall appeal and poor results. Furthermore, this limits the richness of video content and the creative potential, making it difficult to meet user demands for high-quality, creative videos. Summary of the Invention

[0003] The present disclosure proposes a video generation method, device, electronic device, storage medium and program product, which at least to a certain extent solve technical problems in related technologies such as poor creative effects of video generation.

[0004] In a first aspect, the present disclosure provides a video generation method, comprising:

[0005] In response to an input operation on an input area in a display interface, displaying input content in the input area; wherein the input content includes at least one prompt text and multiple input identifiers, and the multiple input identifiers are used to indicate multiple types of input data;

[0006] Parsing the multiple types of input data to obtain corresponding description information;

[0007] Generating a target text from the corresponding description information according to the order of the prompt text and the input identifier in the input content;

[0008] A target video is generated based on the target text and the multiple types of input data.

[0009] In a second aspect of the present disclosure, a video generation device is provided, comprising:

[0010] A display module, configured to display input content in an input area in a display interface in response to an input operation on the input area; wherein the input content includes at least one prompt text and multiple input identifiers, wherein the multiple input identifiers are used to indicate multiple types of input data;

[0011] A data parsing module, configured to parse the multiple types of input data to obtain corresponding description information;

[0012] A target text module, configured to generate a target text from the corresponding description information according to the order of the prompt text and the input identifier in the input content;

[0013] A video generation module is used to generate a target video based on the target text and the multiple types of input data.

[0014] According to a third aspect of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the first aspect when executing the computer program.

[0015] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the method described in the first aspect.

[0016] According to a fifth aspect of the present disclosure, a computer program product is provided, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method according to the first aspect.

[0017] As can be seen from the foregoing, the video generation method, apparatus, electronic device, storage medium, and program product provided herein generate a target video based on input content in various forms, such as links, multimedia, or files, entered into an input area. This is no longer limited to a single form of input material, but rather allows for the flexible combination of multimodal materials in various forms, thereby improving the quality, flexibility, and creative expression of video generation. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0019] Figure 1 A schematic diagram of a video generation architecture according to an embodiment of the present disclosure.

[0020] Figure 2 Schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.

[0021] Figure 3 Schematic diagram of the video generation method according to an embodiment of the present disclosure.

[0022] Figure 4 Schematic diagram of the input area of an embodiment of the present disclosure.

[0023] Figure 5 A schematic diagram of plug-in installation prompt information and plug-in installation identification according to an embodiment of the present disclosure.

[0024] Figure 6 A schematic diagram of a plug-in usage prompt information according to an embodiment of the present disclosure.

[0025] Figure 7 Schematic diagram of the link conversion principle of the embodiment of the present disclosure.

[0026] Figure 8 A schematic diagram of the principle of multimedia material input according to an embodiment of the present disclosure.

[0027] Figure 9 A schematic diagram of the principle of file material input according to an embodiment of the present disclosure.

[0028] Figure 10 A schematic diagram illustrating the principle of attribute input according to an embodiment of the present disclosure.

[0029] Figure 11 Schematic diagram of a form page according to an embodiment of the present disclosure.

[0030] Figure 12 A schematic diagram of a history record and asset library according to an embodiment of the present disclosure.

[0031] Figure 13 A schematic diagram of an asset form page according to an embodiment of the present disclosure.

[0032] Figure 14 Schematic diagram of a video generating device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0033] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.

[0034] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.

[0035] It is understood that before using the technical solutions disclosed in each embodiment of the present disclosure, the user should be informed of the type, scope of use, and usage scenarios of the personal information involved in the present disclosure and obtain the user's authorization in an appropriate manner in accordance with relevant laws and regulations. For example, in response to receiving a user's active request, a prompt message is sent to the user to clearly remind the user that the operation requested will require the acquisition and use of the user's personal information. In this way, the user can independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operation of the technical solution of the present disclosure based on the prompt message.

[0036] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0037] Figure 1 FIG2 is a schematic diagram showing a video generation architecture of an embodiment of the present disclosure. Figure 1 The video generation architecture 100 may include a server 110, a terminal 120, and a network 130 that provides a communication link. The server 110 and the terminal 120 may be connected via a wired or wireless network 130. The server 110 may be an independent physical server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and CDN.

[0038] Terminal 120 can be implemented in hardware or software. For example, when implemented in hardware, terminal 120 can be any electronic device with a display screen that supports page display, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, and desktop computers. When terminal 120 is implemented in software, it can be installed in the electronic devices listed above; it can be implemented as multiple software or software modules (such as software or software modules used to provide distributed services), or it can be implemented as a single software or software module, and no specific limitations are given here.

[0039] It should be noted that the video generation method provided in the embodiment of the present disclosure can be executed by the terminal 120, can be executed by the server 110, or can be executed by the terminal 120 and the server 110 together. Figure 1 The number of terminals, networks, and servers in the embodiment is for illustration only and is not intended to limit the number of terminals, networks, and servers.

[0040] Figure 2FIG. 2 shows a schematic diagram of the hardware structure of an exemplary electronic device 200 provided in an embodiment of the present disclosure. Figure 2 As shown, electronic device 200 may include: processor 202, memory 204, network module 206, peripheral interface 208 and bus 210. Processor 202, memory 204, network module 206 and peripheral interface 208 are connected to each other through bus 210 in communication with each other within electronic device 200.

[0041] The processor 202 may be a central processing unit (CPU), a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or one or more integrated circuits. The processor 202 may be used to perform functions related to the technology described in this disclosure. In some embodiments, the processor 202 may also include multiple processors integrated into a single logical component. For example, Figure 2 As shown, the processor 202 may include a plurality of processors 202a, 202b, and 202c.

[0042] The memory 204 may be configured to store data (eg, instructions, computer code, etc.). Figure 2 As shown, the data stored in the memory 204 may include program instructions (e.g., program instructions for implementing the video generation method of the embodiment of the present disclosure) and data to be processed (e.g., the memory may store configuration documents of other modules, etc.). The processor 202 may also access the program instructions and data stored in the memory 204 and execute the program instructions to operate on the data to be processed. The memory 204 may include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 204 may include a random access memory (RAM), a read-only memory (ROM), an optical disk, a magnetic disk, a hard disk, a solid-state drive (SSD), a flash memory, a memory stick, etc.

[0043] The network module 206 can be configured to provide the electronic device 200 with communication with other external devices via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC)), a cellular network, the Internet, or a combination thereof. It will be appreciated that the type of network is not limited to the specific examples above. In some embodiments, the network module 206 can include any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, and the like.

[0044] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to implement information input and output. For example, the peripheral devices can include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, and various sensors, and output devices such as a display, a speaker, a vibrator, and an indicator light.

[0045] The bus 210 can be configured to transmit information between the various components of the electronic device 200 (e.g., the processor 202, the memory 204, the network module 206, and the peripheral interface 208), such as an internal bus (e.g., a processor-memory bus), an external bus (USB port, PCI-E bus), etc.

[0046] It should be noted that although the architecture of the electronic device 200 shown above only shows the processor 202, the memory 204, the network module 206, the peripheral interface 208, and the bus 210, in a specific implementation, the architecture of the electronic device 200 may also include other components necessary for normal execution. In addition, it will be understood by those skilled in the art that the architecture of the electronic device 200 may also include only the components necessary to implement the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0047] Related technologies for video generation often rely on a limited input format, such as supporting only a single URL or direct input of video or image elements. However, this approach fails to meet the complex demands of video generation in diverse scenarios, where users may desire diverse content from a variety of sources. Existing input limitations significantly reduce the system's flexibility and practicality. Therefore, improving the effectiveness and flexibility of video generation, as well as expanding the creative possibilities, have become pressing technical challenges.

[0048] In view of this, embodiments of the present disclosure provide a video generation method, apparatus, electronic device, storage medium, and program product. By inputting content in various forms, such as links, multimedia, or files, into an input area, a target video is generated based on this input content. This is no longer limited to a single form of material input, but rather allows for the flexible combination of multimodal materials in various forms, thereby improving the quality, flexibility, and creative expression of video generation.

[0049] See also Figure 3 , Figure 3 The schematic flow chart of the video generation method according to the embodiment of the present disclosure is shown. The video generation method according to the embodiment of the present disclosure can be deployed on a server or a terminal. Figure 3 In the embodiment, the video generation method 300 may further include the following steps.

[0050] In step S310, in response to an input operation on an input area in a display interface, input content is displayed in the input area; wherein the input content includes at least one prompt text and multiple input identifiers, and the multiple input identifiers are used to indicate multiple types of input data.

[0051] Among them, the display interface may refer to a visual interface for interacting with an application or system. The display interface may include input areas, buttons, icons, text and other visual elements, through which you can interact with the system. The input area may refer to the part of the display interface used to receive input, such as an input box or other input-enabled components, in which you can enter text, select options or perform other forms of input. An input operation may refer to any input activity performed in the input area, including keyboard input, mouse clicks, paste operations, drag and drop operations, etc. These operations or combinations of operations can be captured by the system and used to trigger corresponding functions or responses. An input identifier may refer to a visual element displayed in the input area to indicate the corresponding input data. Multiple types of input data may refer to different types of data or information that can be obtained through the input area. These types may include, but are not limited to: links, multimedia or files.

[0052] It can be seen that compared with the traditional single input form, the present disclosure integrates multiple types of data in one input area, which greatly expands the boundaries of information input and meets the needs of complex and changing application scenarios. When performing input operations, the input content is indicated by the input identifier, which can more intuitively identify and manage the input content and improve the efficiency of video generation. In addition, by inputting data of multiple modalities at the same time, through the organic combination of multiple forms, it is possible to express information more comprehensively and understand the meaning and intention of the input content more accurately, thereby improving the effect of generating videos.

[0053] Specifically, see Figure 4 , Figure 4 A schematic diagram illustrating an input area according to an embodiment of the present disclosure is shown. Figure 4In the display interface 400, an input area 410 is displayed. The input area 410 can display input guide information, such as "Enter creation instructions" to prompt that the input area 410 can be used to input relevant instructions for creating videos. Various input operations can be performed in the input area 410, and the input guide information can be hidden at this time. Through input operations, prompt texts 411-413 and various types of input data can be entered in the input area 410. Corresponding input identifiers 421 (link type), input identifiers 422-423 (multimedia type), and input identifier 424 (file type) can be displayed in the input area 410. The input identifier 421 can indicate the data of the link link1, the input identifiers 422-423 can indicate the multimedia data media1 and media2, and the input identifier 424 can indicate file data.

[0054] In some embodiments, the type includes a link, and a link identifier is displayed in the display interface;

[0055] In response to an input operation on an input area in a display interface, displaying input content in the input area includes:

[0056] In response to a triggering operation on the link identifier, displaying a preset link prompt text and a corresponding link content identifier in the input area;

[0057] Entering a target link in the link content identifier;

[0058] Alternatively, in response to inputting a target link in the input area, a preset link prompt text and a corresponding link content identifier are displayed in the input area, and at least a portion of the target link is displayed in the link content identifier.

[0059] Among them, the link can be a URL (Uniform Resource Locator) used to point to a web page or resource. In the input area, the link can be displayed in text form in the corresponding input identifier, and the target resource can be accessed by clicking it. The link identifier can refer to an identifier in the display interface used to input or trigger operations related to the link. When the user interacts with the link identifier (such as clicking), the system will execute the function related to the link. The preset link prompt text can refer to a pre-set text displayed in the input area after the link identifier is triggered, which is used to prompt the use of the resource of the entered link address. The link content identifier can refer to an identifier in the input area for receiving and displaying the link address entered by the user. The link content identifier can be a text box or input area with a fixed length, in which the address of the target link can be entered. When the address length of the target link is greater than the above-mentioned fixed length, the excess part can be hidden, and the target link in the link content identifier can be edited. The target link is the network link address that is expected to be entered or used, pointing to a specific web page or resource. After triggering the link identifier, the corresponding input identifier - the link content identifier - is displayed in the input area, and the target link address can be entered in the link content identifier. Alternatively, the target link address can be directly entered in the input area, and the input area can automatically recognize and trigger the display of the link content identifier, while displaying part or all of the link address entered by the user in the link content identifier. One or more link addresses can be entered in a variety of ways, improving input flexibility and user experience.

[0060] Specifically, if Figure 4 As shown, a link identifier 431 (for example, an icon or button) can be pre-set and displayed in the display interface 400. When interacting with the link identifier 431 (such as clicking or other triggering operations), a preset link prompt text 411 (i.e., prompt text 411, for example, "Use the information in the link") and a corresponding link content identifier 421 (i.e., input identifier 421) can be displayed in the input area. The address of the target link "https: / / www.****.com" can be directly entered in the link content identifier 421. Alternatively, the target link address "https: / / www.****.com" can be directly entered in the input area 410. The input area will automatically recognize and trigger the display of the link content identifier 421, and at the same time, part or all of the entered link address will be displayed in the link content identifier.

[0061] In some embodiments, method 300 further includes:

[0062] In response to detecting that the target link indicated by the link content identifier cannot be resolved, performing at least one of the following operations:

[0063] Displaying abnormal prompt information in an area associated with the link content identifier, wherein the abnormal prompt information is used to indicate that the target link cannot be resolved;

[0064] Alternatively, in response to detecting that the link resolution plug-in is not installed, displaying plug-in installation prompt information and / or a plug-in installation identifier in the display interface, wherein the plug-in installation prompt information is used to prompt the installation of the link resolution plug-in, and the plug-in installation identifier is used to trigger the installation of the link resolution plug-in;

[0065] Alternatively, in response to detecting that the link resolution plug-in has been installed, plug-in usage prompt information is displayed in the display interface to prompt the user to use the link resolution plug-in.

[0066] Among them, when it is detected that the target link indicated by the link content identifier cannot be resolved, an abnormal prompt message is displayed in the associated area of the link content identifier to prompt that there is a problem with the target link and it cannot be resolved normally. The link resolution plug-in may refer to a component used to resolve the link address entered by the user. When it is unable to directly resolve certain links by itself, it may be necessary to use the link resolution plug-in to complete the resolution work. When it is detected that the link resolution plug-in is not installed, a plug-in installation prompt message can be displayed in the display interface to prompt that the plug-in can be installed so that the link can be resolved normally. When it is detected that the link resolution plug-in has been installed, a plug-in usage prompt message can be displayed in the display interface to prompt how to use the plug-in to resolve the link, helping users to better utilize the plug-in function.

[0067] By detecting link resolution status and providing corresponding prompts, the system can better handle various link input scenarios, improving compatibility and processing capabilities. Prompting users to install or use a link resolution plug-in expands the functional boundaries of the input area, enabling the system to handle more types of links and meet diverse user needs. This also prevents users from being lost or confused when encountering link resolution issues, reduces the learning curve and operational difficulty, and makes it easier for users to input and resolve links using the system.

[0068] Specifically, if Figure 4 As shown, when it is detected that the target link address https: / / www.****.com cannot be resolved, an abnormal prompt message 410 can be displayed in the associated area, such as "Sorry, the link cannot be resolved." Figure 5 , Figure 5A schematic diagram of a plug-in installation prompt message and a plug-in installation identifier according to an embodiment of the present disclosure is shown. If the link resolution plug-in is not installed, the plug-in installation prompt message 510 and / or the plug-in installation identifier 520 may be displayed in a pop-up window or other manner in the display interface. For example, the plug-in installation prompt message 510 may be "Install the link resolution plug-in", and triggering the plug-in installation identifier 520 may cause the link resolution plug-in to be installed. If the link resolution plug-in is already installed, the plug-in installation prompt message 510, such as "Install the link resolution plug-in", may be displayed in a pop-up window or other manner in the display interface. See Figure 6 , Figure 6 A schematic diagram of a plug-in usage prompt according to an embodiment of the present disclosure is shown. If a link resolution plug-in has been installed, plug-in usage prompts 610-620 may be displayed in a pop-up window or other manner in the display interface to prompt how to use the link resolution plug-in. For example, plug-in usage prompts 610-620 may refer to steps 1 and 2, respectively.

[0069] In some embodiments, method 300 further includes:

[0070] In response to a triggering operation on the plug-in identifier corresponding to the link resolution plug-in, displaying a plug-in panel in an area associated with the plug-in identifier, wherein a link conversion identifier is displayed in the plug-in panel;

[0071] In response to a triggering operation on the link conversion identifier, converting the target link to obtain a link conversion result, and displaying a conversion completion identifier;

[0072] In response to a triggering operation on the conversion completion mark, at least a portion of the link conversion result is displayed in a link content mark used to indicate the link conversion result.

[0073] Among them, the plug-in identifier may refer to an identifier used to represent the link resolution plug-in in the display interface, and the plug-in identifier can be interacted with by triggering operations such as clicking and touching to start functions related to the link resolution plug-in. The plug-in panel may refer to an operation panel displayed in the associated area of the plug-in identifier (such as a nearby blank space, a pop-up window, etc.) after the plug-in identifier is triggered. The panel contains various function identifiers and operation options related to the link resolution plug-in, which is used to provide a more detailed plug-in operation interface. The link conversion identifier may refer to an identifier located in the plug-in panel for indicating a link conversion operation. When the user triggers the link conversion identifier, the target link can be parsed and converted. For example, format conversion, encoding conversion, and regeneration of the link conversion result after parsing. The conversion completion identifier may refer to an identifier displayed in the plug-in panel after the link conversion operation is completed to indicate that the link conversion has been completed. Subsequent operations can be performed by triggering the conversion completion identifier, which is used to display the link conversion result in the link content identifier.

[0074] The plugin tag and plugin panel centralize link conversion functions in a single, easy-to-use interface. Simply triggering the plugin tag allows quick access to the plugin panel for link conversion operations, reducing the need to switch between different interfaces or functional modules and improving operational efficiency. The addition of link conversion expands the functional boundaries of input, better meeting users' diverse link processing needs and improving the efficiency and effectiveness of video generation.

[0075] See also Figure 7 , Figure 7 A schematic diagram showing the principle of link conversion according to an embodiment of the present disclosure is shown. Figure 7 In the example, when the plug-in identifier (not shown) is triggered, the plug-in panel 710 can be displayed. The plug-in panel 710 displays a link conversion identifier 711, which can be triggered to parse and convert the target link that cannot be resolved to obtain a link conversion result. The plug-in panel 710 can display a resolution completion prompt message and a conversion completion identifier 721. The conversion completion identifier 721 can be triggered, and then the link conversion result can be obtained. Figure 4 The link conversion result is displayed in the link content identifier 421 in the input area 410 shown in . At this time, it can be displayed in the input area of the newly created input page, or it can be replaced by the unresolved target link in the input area of the original input page, which is not limited here.

[0076] In some embodiments, method 300 further includes:

[0077] In response to detecting that a preset text is displayed in the link content identifier, displaying a permission confirmation identifier for the link access authority on the display interface;

[0078] In response to a confirmation operation on the permission confirmation identifier, allowing generation of the target video;

[0079] In response to detecting that the permission confirmation identifier is not confirmed, refusing to generate the target video, and / or displaying operation prompt information about the permission confirmation identifier on the display interface.

[0080] The permission confirmation indicator for link access rights may refer to an indicator that appears in the display interface for confirming link access rights, and interaction with the permission confirmation indicator (e.g., clicking a confirmation button) indicates whether the user has the relevant link access rights. Operation prompt information may refer to prompt text or information displayed in the display interface when the permission confirmation indicator is not confirmed, used to indicate the current operation status (e.g., unconfirmed permission) and subsequent operation instructions. When a preset text is detected in the link content indicator, the permission confirmation indicator for link access rights is displayed in an appropriate location on the display interface (e.g., a nearby blank area, a pop-up window, etc.). When a confirmation operation (e.g., a selection operation) is performed on the permission confirmation indicator, it indicates that the user has the link access rights, and the target video is allowed to be generated. The video generation operation is performed according to the preset process and rules. If it is detected that the permission confirmation indicator is not confirmed (e.g., the user does not click the confirmation button, or the operation times out, etc.), the target video is denied. At the same time, operation prompt information about the permission confirmation indicator can be displayed in the display interface to inform the user of the unconfirmed permission and that the confirmation operation can be performed.

[0081] By requiring users to confirm link access permissions under specific conditions (link content identification displays preset text), unauthorized link access and potential security risks can be effectively prevented, protecting data security.

[0082] Figure 4 In the embodiment of the present invention, when a preset text is displayed in the detected link content identifier 421, the preset text may indicate the presence of link information, such as the presence of http: / / , then information related to the link access permission 442 and a permission confirmation identifier 443 for the link access permission may be displayed on the display interface. The information related to the link access permission 442 may include instructions for confirming the link access permission. When the permission confirmation identifier 443 is in a first state (e.g., selected), it may indicate that the link access permission has been confirmed. When the permission confirmation identifier 443 is in a second state (e.g., unselected), it may indicate that the link access permission has not been confirmed.

[0083] In some embodiments, the type includes multimedia, and a multimedia logo is displayed in the display interface;

[0084] In response to an input operation on an input area in a display interface, displaying input content in the input area includes:

[0085] In response to a triggering operation on the multimedia identifier, displaying a multimedia source identifier in an area associated with the multimedia identifier;

[0086] In response to a triggering operation on the multimedia source identifier, displaying multimedia material corresponding to the multimedia source identifier;

[0087] In response to a selection operation on the multimedia material, displaying a multimedia content identifier corresponding to the selected multimedia material in the input area;

[0088] or,

[0089] Display multimedia materials;

[0090] In response to a drag operation on the multimedia material, displaying a multimedia content identifier corresponding to the multimedia material in the input area;

[0091] or,

[0092] In response to detecting the first input for the multimedia material, a multimedia material prompt text and a corresponding multimedia content identifier are displayed in the input area.

[0093] Among them, multimedia can refer to media content in the form of images, audio, video, etc. The multimedia identifier can refer to an identifier used to represent multimedia-related operations in the display interface. Clicking or triggering the multimedia identifier can perform multimedia-related operation processes. The multimedia source identifier can refer to an identifier that indicates the source of the multimedia material. The multimedia content identifier can refer to an identifier displayed in the input area and used to represent the multimedia material selected by the user. The dragging operation of the multimedia material can refer to the operation of dragging the multimedia material from the original position to the input area using a device such as a mouse or a touch screen.

[0094] Specifically, multimedia materials can be input into the input area through a multimedia identifier, and the multimedia identifier is displayed in the display interface. When the user performs a trigger operation (such as clicking) on the multimedia identifier, the multimedia source identifier is displayed in the associated area of the multimedia identifier. When a trigger operation (such as clicking) is performed on the multimedia source identifier, the multimedia materials corresponding to the multimedia source identifier can be displayed. When a multimedia material is selected (such as clicking to select), the multimedia content identifier corresponding to the selected multimedia material can be displayed in the input area. For example, Figure 8 In the display interface, multimedia identifier 432 is displayed. Triggering multimedia identifier 432 displays multimedia source identifiers 4321-4324 in the associated area of multimedia identifier 432 to indicate different sources, such as device, asset, mobile phone, and cloud. Triggering a multimedia source identifier allows you to select a path for uploading multimedia materials and display multimedia materials from the corresponding source. You can select a target material from the multimedia materials and enter it into the input area, which displays the corresponding multimedia content identifiers 422-423.

[0095] Alternatively, a target multimedia material can be dragged from the displayed multimedia material (which can be displayed based on a path identified by a multimedia source or by other paths) to the input area, whereupon the multimedia content identifier corresponding to the multimedia material is displayed in the input area. When the first input (e.g., a click, a drag attempt, etc.) for the multimedia material is detected, multimedia material prompt text 412 (i.e., prompt text 412) and corresponding multimedia content identifiers 422-423 can be displayed in the input area.

[0096] As can be seen, the present disclosure provides multiple ways to input multimedia materials, including through triggering identifiers and dragging operations, further expanding the input function to meet different user habits and quickly and conveniently select the required multimedia materials. It also facilitates the subsequent expansion of more multimedia sources and material types by simply adding corresponding material display and selection functions in the area corresponding to the multimedia source identifier.

[0097] In some embodiments, the type includes a file, and a file identifier is displayed in the display interface;

[0098] In response to an input operation on an input area in a display interface, displaying input content in the input area includes:

[0099] In response to a trigger operation on the file identifier, displaying a file source identifier in an area associated with the file identifier;

[0100] In response to a triggering operation on the file source identifier, displaying the file material corresponding to the selected file source identifier;

[0101] In response to a selection operation on the file material, displaying a file content identifier corresponding to the selected file material in the input area;

[0102] or,

[0103] Display file material;

[0104] In response to a drag operation on the file material, displaying a file content identifier corresponding to the file material in the input area;

[0105] or,

[0106] In response to detecting that the file material cannot be identified, file abnormality prompt information is displayed on the display interface.

[0107] Among them, files may refer to electronic files such as uploaded or input documents, spreadsheets, PDFs, presentation documents, etc. The file identifier may refer to an identifier used to represent file-related operations on the display interface, and the file input operation process can be carried out by triggering the file identifier. The file source identifier may refer to an identifier used to indicate the source of the file material, and the file material of the corresponding source can be viewed by triggering the file source identifier. The file content identifier may refer to an identifier displayed in the input area and used to indicate the selected file material. The dragging operation of the file material may refer to the operation of dragging the file material from the original location to the input area using a device such as a mouse or touch screen. The file abnormality prompt information may refer to the prompt information displayed when the file material cannot be identified.

[0108] For example, Figure 9 In the example, a file identifier 433 may be displayed on the display interface 400. When the file identifier is triggered (e.g., clicked), file source identifiers 4331-4332 may be displayed in the associated area of the file identifier to indicate the file source. The file source identifier may be triggered (e.g., clicked) to display the file material corresponding to the file source identifier. The file material may be selected (e.g., clicked to select), and the file content identifier 424 corresponding to the selected file material may be displayed in the input area 410.

[0109] Alternatively, a target file material may be dragged from the displayed file material (which may be displayed based on a path identified by a multimedia source or by another path) to the input area 410, whereupon a file content identifier 424 corresponding to the file material is displayed. If an unrecognizable file material is detected, a file abnormality prompt message 4333, such as "File abnormality", may be displayed in the input area to indicate that the file is abnormal.

[0110] It can be seen that the present disclosure provides multiple intuitive ways to select file materials, which improves the flexibility and convenience of file material input, as well as the friendliness and ease of use of the interface.

[0111] In some embodiments, the display interface includes a setting indicator;

[0112] In response to an input operation on an input area in a display interface, displaying input content in the input area includes:

[0113] In response to a trigger operation on the setting identifier, displaying at least one attribute setting identifier in an area associated with the setting identifier;

[0114] In response to a setting operation on the attribute setting identifier, the corresponding attribute identifier is displayed in the input area; the attribute identifier displays the attribute value set by the setting operation.

[0115] Among them, the setting identifier can refer to an identifier used to trigger the setting function in the display interface. Triggering the setting identifier can enter the setting-related operation process. The attribute setting identifier can be used to represent different settable attributes. Triggering the identifier can perform setting operations on the corresponding attributes. The setting operation can refer to the operation performed on the attribute setting identifier, such as selecting an option in a drop-down menu, entering a value, switching the switch state, etc., to determine the specific attribute value of the attribute. The attribute identifier can refer to the attribute information determined by the setting operation, in which the attribute value set by the setting operation will be displayed.

[0116] like Figure 10 As shown, a setting indicator 434 can be displayed in the display interface 400. When a trigger operation (e.g., a click) is performed on the setting indicator 434, at least one property setting indicator 4341-4343 is displayed in the associated area of the setting indicator 434. A setting operation (e.g., selecting an option, entering a value, etc.) can be performed on the property setting indicator, and the corresponding property indicator is displayed in the input area, and the property value set by the setting operation is displayed in the property indicator. By using the setting indicator and the property setting indicator, the relevant properties can be personalized according to needs, thereby improving the customizability and flexibility of the input.

[0117] In some embodiments, the attribute setting identifier includes a virtual object setting identifier, which is used to set whether a virtual object is used in the target video.

[0118] In some embodiments, the attribute setting identifier includes a duration identifier for setting the duration of the target video.

[0119] In some embodiments, the attribute setting identifier includes a language identifier for setting the language of the target video.

[0120] Among them, the virtual object setting identifier may refer to an identifier used to control the use of virtual objects in the target video. When performing a setting operation on the virtual object setting identifier (such as clicking a toggle switch, checking an option, etc.), if it is set to use a virtual object, then in the subsequent process of generating the target video, the virtual object will be integrated into the target video according to the preset rules or algorithms; if it is set not to use a virtual object, then the generated target video will not have any virtual object-related elements. The duration identifier can be used to set the duration of the target video. The setting operation is performed on the duration identifier, such as entering a specific duration value (such as 30 seconds, 2 minutes, etc.), or selecting a preset duration option through a slider, drop-down menu, etc. When generating the target video, the playback duration of the video will be controlled according to the duration set by the user. The language identifier can be used to set the language used by the target video. The setting operation is performed on the language identifier, which can be to select the desired language from the drop-down menu. When generating the target video, the language-related elements such as voice and subtitles in the video are configured according to the language selected by the user. Such as Figure 10 As shown, the attribute setting identifiers 4341-4343 are used to set the use of the virtual object, the duration of the target video, and the language of the target video respectively.

[0121] In some embodiments, the method 300 further includes: in response to a move operation on the input identifier, moving the input identifier to a cursor position of the move operation.

[0122] Among them, the position update can be triggered by detecting the movement operation of the input identifier (such as mouse dragging or touch sliding). When the start of the movement operation is detected, the position change of the cursor (or finger touch point) during the movement process can be tracked in real time, and the current coordinates of the cursor can be continuously calculated during the operation. Subsequently, the display position of the input identifier is dynamically adjusted to the real-time position of the cursor until the movement operation ends (such as the mouse is released or the finger leaves the screen), and finally the input identifier stays at the position where the cursor was last located. In this way, by freely and arbitrarily adjusting the position of the material, compared with the limitations of the traditional fixed layout, the flexibility of the input is greatly increased, and the material can be placed in any suitable position in the input area according to the creative ideas or actual needs, thereby completing the video creation more efficiently and accurately, and significantly improving the creative efficiency and video effect.

[0123] In some embodiments, method 300 further includes:

[0124] In response to a triggering operation on the input identifier, previewing the input content in an area associated with the input identifier;

[0125] Alternatively, in response to a triggering operation on the input identifier, displaying a preview identifier in an area associated with the input identifier;

[0126] In response to a triggering operation on the preview indicator, the corresponding input content is previewed.

[0127] Among them, the input identifier can include a link content identifier, a multimedia content identifier, a file content identifier, and an attribute identifier. The input identifier can be triggered (for example, hovering, clicking, etc.), and the preview information of the input content can be directly displayed in the associated area. Specifically, for a link content identifier, the preview information can be a complete link; for a multimedia content identifier, the preview information can be the corresponding multimedia material; for a file content identifier, the preview information can be the complete file name and file type; for an attribute identifier, the preview information can be the attribute name and the corresponding attribute value. It is also possible to display only a preview identifier (such as a thumbnail, icon, etc.) when the input identifier is triggered, and further trigger the preview identifier to expand the preview of the complete input content.

[0128] In step S320, the multiple types of input data are parsed to obtain corresponding description information.

[0129] Among them, it is possible to analyze various types of multimodal input data (such as text, images, audio, video, etc.). For text data, its core themes and key points can be obtained through semantic recognition and keyword extraction, or the text data can be left unprocessed. For image data, image recognition algorithms can be used to identify objects, scenes, color features, and other information in the picture. For audio data, descriptions such as its text content and rhythm frequency can be obtained through the use of speech recognition and audio analysis technology. For video data, comprehensive description information such as the plot, shot changes, and sound coordination can be extracted by combining image and audio analysis results, thus converting various types of input data into structured and understandable description information.

[0130] In some embodiments, parsing the multiple types of input data to obtain corresponding description information includes:

[0131] Acquiring the corresponding input data based on the address identifier or identity identifier of the input identifier;

[0132] Parsing the image in the input data to obtain a first text description;

[0133] For a video in the input data, determining corresponding key frames based on scene changes of the video and parsing the key frames to obtain a second text description;

[0134] The files in the input data are parsed to obtain a third text description.

[0135] Among them, various types of input data corresponding to the input identifier can be located and obtained based on the address identifier (such as network path, local storage path, etc.) or identity identifier (such as data unique ID) carried by the input identifier. For images in the input data, image analysis can be performed and natural language processing technology can be used to convert the visual information of the image data into an accurate and vivid first text description. For videos in the input data, key frames can be accurately determined by analyzing the scene changes of the video. Key frames can be important frames in the video that are representative and can summarize the content of the scene. Subsequently, each key frame is deeply analyzed, and the elements and features in the key frame are extracted by combining image recognition technology, and the corresponding second text description is generated. For files in the input data (such as documents, tables, etc.), a file parsing algorithm can be used to deeply analyze the content structure, logical relationship and key information of the file and convert it into a third text description.

[0136] Based on the textual information originally present in the input data and at least one of the first, second, and third textual descriptions obtained through the aforementioned parsing, this textual information is integrated and edited with the corresponding image and video assets in accordance with the creative requirements and logical structure of the target video. By adding appropriate subtitles, dubbing, special effects, and other elements, the textual information and visual assets are organically combined to ultimately generate the desired target video.

[0137] In some embodiments, the display interface displays a first generation indicator;

[0138] The method further comprises:

[0139] In response to a triggering operation on the first generation identifier, a form page corresponding to the input content is displayed; the form page includes preset form items and corresponding item content, and the item content corresponds to the input content.

[0140] Among them, the first generation identifier can be used to generate a form page. When a trigger operation (such as clicking, long pressing, etc.) for the first generation identifier is executed, a form page corresponding to the input content can be displayed. The form page includes multiple preset form items, and the item content corresponding to the preset items can be filled based on the descriptive information obtained from the analysis of the input data, and displayed in the form of a form page. In this way, there is no need to manually search and organize in a large amount of information. You only need to view, edit or confirm the corresponding item content in the form page, which greatly improves the efficiency and accuracy of input information processing. See Figure 11 , Figure 11 A schematic diagram of a form page according to an embodiment of the present disclosure is shown. Figure 11 In the example, the form page 1100 includes a plurality of preset form items 1110 - 1140 and corresponding item contents.

[0141] In some embodiments, the project content corresponds to the input content, including:

[0142] Determining corresponding item content from the description information based on the preset form item;

[0143] The corresponding area of the preset form item is automatically filled with the item content.

[0144] Among them, based on the descriptive information obtained by parsing the input data, matching can be performed according to the preset form items. Each preset form item corresponds to a specific information category, and information fragments related to the item can be filtered out from the descriptive information. For example, assuming that the preset form items include "product functions", "performance indicators" and "applicable scenarios", the corresponding parts of the information can be extracted from the descriptive information respectively as candidate content for each form item. After the matching is completed, the filtered item content can be automatically filled into the corresponding area of the preset form item, so that the pre-filled information that is highly relevant to the input content can be seen efficiently and quickly on the form page, which improves the operational efficiency and experience.

[0145] like Figure 11 As shown, the preset form item 1110 may be an item about multimedia information, in which case the relevant content of the multimedia information is matched from the description information and filled in. The preset form item 1120 may be an item about multimedia material, in which case the corresponding multimedia material is matched based on the description information and filled in, or all multimedia materials involved in the input area may be filled in. The preset form item 1130 may be an item about a script, in which case the new custom script content is matched from the description information, and the item content 1131 of the preset form item 1130 may be generated based on the custom script content. The preset form item 1140 may be an item about a promotional price, in which case the promotional price content is matched from the description information and the item content "original price AAA", "promotional price BBB", and "promotional time period T1-T2" of the preset form item 1140 are automatically filled in.

[0146] In some embodiments, method 300 further includes:

[0147] In response to an editing operation on the project content, updating the project content;

[0148] or,

[0149] Displaying an update indicator in an area associated with the project content;

[0150] In response to a triggering operation on the update identifier, the project content is updated.

[0151] You can directly edit the item content to update it, or display an update indicator in the associated area of the item content. When the user triggers the update indicator, the item content is updated. For example, the item content of the preset form item 1140, "Promotional Time Period T1-T2," can be directly modified to "Promotional Time Period T1-T3." This allows for more intuitive updating of video creation input information through the form page, improving the efficiency and accuracy of video creation.

[0152] In some embodiments, the form page includes a view indicator and / or an edit indicator for the input content;

[0153] The method further comprises:

[0154] In response to a triggering operation on the viewing indicator, displaying all the input content in an area associated with the viewing indicator;

[0155] or,

[0156] In response to a trigger operation on the edit mark, the input area is returned to display; the input content of the input area corresponds to the current content of the form page.

[0157] The view mark can be used to view the input content corresponding to the form page, and the edit mark can be used to return to the input area to edit the input content. Figure 11 As shown, when the view indicator 1150 is triggered (e.g., by sliding over or hovering), the entire input content can be displayed in the area associated with the view indicator, making it easy to view the complete information. When the edit indicator 1160 is triggered, the input area can be returned to display, and the content presented in the input area is consistent with the content currently displayed on the form page, so that editing and modification can be performed based on this.

[0158] In some embodiments, a second generation indicator is displayed on the form page;

[0159] The method further comprises:

[0160] In response to a triggering operation for the second generation identifier, the target text is updated based on current item content of the form page.

[0161] Among them, generating the target video based on the current project content of the form page can ensure that the content of the target video is highly relevant to the information on the form page, avoiding errors caused by inconsistent information. At the same time, the input information is integrated into the form page to reduce redundancy and dispersion, promote data sharing and interaction, and improve the data processing efficiency of video creation. Figure 11 As shown, a second generation mark 1170 is also displayed on the form page 1100. Triggering the second generation mark 1170 can update the target text based on the current project content of the form page. Furthermore, the current project content can be combined with the input content, input data, etc. to generate a video.

[0162] In some embodiments, in response to a triggering operation on the first generation identifier, displaying a form page corresponding to the input content includes:

[0163] In response to a trigger operation for the first generation identifier and detecting that no material can be extracted based on the input area,

[0164] Determining the item content corresponding to the preset form item in the form page in a designated material library based on the prompt text; and displaying the form page and the corresponding item content;

[0165] and / or,

[0166] Displays input prompt information for uploading input data.

[0167] Among them, when the first generation identifier is triggered, if it is detected that no material can be extracted based on the input area, for example, no material has been uploaded, and no material has been extracted from the parsing result of the link / file, then the project content corresponding to the preset form item in the form page can be accurately located and determined in the specified material library based on the text information in the input area, and then the form page and its corresponding project content can be displayed; at the same time, input prompt information for uploading input data can also be displayed to guide the user to further operate, such as emphasizing the display Figure 11 The upload identifier is 1123. This not only ensures the automation and accuracy of the form page content generation, but also enhances the interactive guidance between the user and the system through prompt information.

[0168] In some embodiments, further comprising:

[0169] In response to detecting that no input information exists in the input area, displaying the historical records and / or the asset library on the display interface;

[0170] In response to a triggering operation on the historical record, displaying historical input content corresponding to the historical record in the input area;

[0171] Or, in response to a trigger operation for an asset identifier in the asset library, an asset form page of the target asset corresponding to the asset identifier is displayed, and the content of the asset form page is obtained based on the description information corresponding to the asset identifier and the corresponding input content in the input area.

[0172] When it is detected that there is no input information in the input area, the historical records and / or asset library can be presented on the display interface. Figure 12 As shown, Figure 12 A schematic diagram of a history record and an asset library according to an embodiment of the present disclosure is shown. Figure 12In the example, when it is detected that there is no input information in the input area, the history record 450 and / or the asset library 460 can be presented on the display interface 400. The history record 450 displays historical input contents 1-3, and the asset library includes asset identifiers 461-465, ... The historical input contents 1-3 can be triggered (for example, clicked), and the historical input contents are displayed in the input area. For example, if the historical input content 2 is clicked, the historical input content 2 is displayed in the input area. The asset identifier can be triggered to display the asset form page of the corresponding asset. For example, the asset identifier 461 can be triggered, and the asset form page corresponding to asset 1 can be displayed, such as Figure 13 As shown in the asset form 1300, Figure 13 A schematic diagram of an asset form page according to an embodiment of the present disclosure is shown. When input information is detected in the input area, the history record and / or the asset library can be hidden.

[0173] It can be seen that if the user triggers the history record, the corresponding historical input content can be displayed in the input area, which greatly facilitates the reuse of past inputs, avoids repeatedly entering the same content, saves time and energy, and is especially suitable for situations where users often need to repeatedly fill in similar information, thereby improving operational efficiency and user experience.

[0174] When a user triggers an asset identifier in the asset library, the asset form page for the target asset corresponding to that identifier is displayed. This form page is generated based on the asset identifier's descriptive information and the corresponding input in the input area. This not only enables rapid access to asset information but also integrates it with the current input for display, allowing users to more comprehensively and conveniently access and process asset-related information. This enhances the system's flexibility and practicality in asset information management, helping users perform asset-related operations and information management more efficiently.

[0175] In some embodiments, the asset form page displays a third generation indicator and a synchronization indicator;

[0176] The method further comprises:

[0177] In response to a trigger operation for the third generation identifier and the synchronization identifier being in a preset state, all descriptive information about the target asset currently corresponding to the asset form page is synchronized to the asset library.

[0178] The third generation identifier may refer to the generation identifier in the asset form page, such as Figure 13As shown in the third generation identifier 1310. When the third generation identifier 1310 is triggered and the synchronization identifier 1320 is in a preset state (for example, a selected state), all descriptive information of the target asset (i.e., asset 1) currently corresponding to the asset form page 1300 can be automatically synchronized to the asset library. On the one hand, after the target asset description information is modified or improved on the asset form page, it can be quickly synchronized to the asset library through a simple trigger operation, avoiding manual data updates between different pages or systems, greatly improving the efficiency and accuracy of information maintenance; on the other hand, it ensures the real-time and consistency of the target asset information in the asset library, so that the asset data in the entire system is always kept up to date, providing a reliable data foundation for subsequent operations based on asset information (such as query, analysis, decision-making, etc.), enhancing the stability and reliability of the system, and also improving the convenience and experience of asset information management for users.

[0179] In step S330, according to the order of the prompt text and the input identifier in the input content, the corresponding description information is generated into a target text.

[0180] The prompt text can guide the input content, for example, it can be a prompt or explanation of the input content, and can contain specific keywords, questions or format requirements (such as Figure 4 The description information may include explanations, examples, and descriptions of the input content. The target text may refer to the final text generated based on the order of the prompt text and input identifiers, combined with the corresponding description information. This text integrates various guidance and user input information during the input process to form a complete text that meets specific formats or requirements.

[0181] Specifically, the order of the prompt text and input identifier in the input content can be clarified according to the layout in the input area. Each prompt text and the corresponding input identifier are associated with the corresponding descriptive information. For example, based on a pre-set rule, for example, the prompt text and the input identifier have the same identifier or belong to the same input group, the descriptive information is associated with the group. According to the order of the prompt text and the input identifier, the corresponding descriptive information is integrated into the target text in sequence. Specifically, for each group of prompt text, input identifier and descriptive information, the descriptive information is inserted into the corresponding position of the target text according to a certain format. For example, the corresponding descriptive information can be inserted immediately after the prompt text, and then the next group of data is processed to generate the target text. In this way, the target text can be automatically generated according to the order of the prompt text and the input identifier, which improves the efficiency of information integration and is particularly suitable for scenarios where a large amount of input content is processed. At the same time, the consistency of the generated target text in format and content can be ensured, thereby improving the quality of video generation.

[0182] In step S340 , a target video is generated based on the target text and the multiple types of input data.

[0183] Among them, the processed materials can be synthesized in the order and manner planned by the target text, and subtitles, animations and other elements can also be added to finally generate a complete target video.

[0184] In some embodiments, generating a target video based on the target text and the multiple types of input data includes: generating the target video based on at least one of the text in the input data, the first text description, the second text description, or the third text description.

[0185] Among them, the target text and the input data can be fused to obtain the target video. For example, with the target text as the reference framework, at the same time, the original text in the input data (such as the explanatory text directly entered by the user), the first text description obtained by image analysis (such as the text summary of the screen elements), the second text description of the video key frame analysis (such as the text extraction of the scene content) and the third text description of the file analysis (such as the paraphrase of the core ideas of the document) are integrated. These text information can act alone or complement each other to provide content guidance and semantic basis for video generation. It is also possible to match relevant materials from the corresponding input data such as images, videos, files, etc. based on the selected text information, and combine the text content with the materials in the form of subtitles, dubbing, animation, etc., and finally generate a target video that is consistent with the text semantics and expressive.

[0186] It should be noted that the method of the embodiment of the present disclosure can be performed by a single device, such as a computer or server. The method of this embodiment can also be applied in a distributed scenario, where multiple devices cooperate to complete the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiment of the present disclosure, and the multiple devices will generate video with each other to complete the method.

[0187] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0188] Based on the same technical concept, corresponding to any of the above embodiments and methods, the present disclosure also provides a video generation device, see Figure 14 , the video generating device includes:

[0189] A display module, configured to display input content in an input area in a display interface in response to an input operation on the input area; wherein the input content includes at least one prompt text and multiple input identifiers, wherein the multiple input identifiers are used to indicate multiple types of input data;

[0190] A data parsing module, configured to parse the multiple types of input data to obtain corresponding description information;

[0191] A target text module, configured to generate a target text from the corresponding description information according to the order of the prompt text and the input identifier in the input content;

[0192] A video generation module is used to generate a target video based on the target text and the multiple types of input data.

[0193] In some embodiments, the type includes a link, and a link identifier is displayed in the display interface;

[0194] The display module is also used for:

[0195] In response to a triggering operation on the link identifier, a preset link prompt text and a corresponding link content identifier are displayed in the input area; and a target link is input in the link content identifier;

[0196] Alternatively, in response to inputting a target link in the input area, a preset link prompt text and a corresponding link content identifier are displayed in the input area, and at least a portion of the target link is displayed in the link content identifier.

[0197] In some embodiments, the display module is further configured to:

[0198] In response to detecting that the target link indicated by the link content identifier cannot be resolved, performing at least one of the following operations:

[0199] Displaying abnormal prompt information in an area associated with the link content identifier, wherein the abnormal prompt information is used to indicate that the target link cannot be resolved;

[0200] Alternatively, in response to detecting that the link resolution plug-in is not installed, displaying plug-in installation prompt information and / or a plug-in installation identifier in the display interface, wherein the plug-in installation prompt information is used to prompt the installation of the link resolution plug-in, and the plug-in installation identifier is used to trigger the installation of the link resolution plug-in;

[0201] Alternatively, in response to detecting that the link resolution plug-in has been installed, plug-in usage prompt information is displayed in the display interface to prompt the user to use the link resolution plug-in.

[0202] In some embodiments, the display module is further configured to: in response to a triggering operation on a plug-in identifier corresponding to the link resolution plug-in, display a plug-in panel in an area associated with the plug-in identifier, wherein the plug-in panel displays a link conversion identifier;

[0203] The video generating device further includes:

[0204] a link conversion module, configured to convert the target link to obtain a link conversion result and display a conversion completion mark in response to a triggering operation on the link conversion mark;

[0205] The display module is further configured to: in response to a triggering operation on the conversion completion identifier, display at least a portion of the link conversion result in a link content identifier indicating the link conversion result.

[0206] In some embodiments, the display module is further configured to:

[0207] In response to detecting that a preset text is displayed in the link content identifier, displaying a permission confirmation identifier for the link access authority on the display interface;

[0208] The video generating device further includes:

[0209] The permission confirmation module is used to allow the generation of the target video in response to a confirmation operation on the permission confirmation identifier; and in response to detecting that the permission confirmation identifier has not been confirmed, refuse to generate the target video, and / or display operation prompt information about the permission confirmation identifier on the display interface.

[0210] In some embodiments, the type includes multimedia, and a multimedia logo is displayed in the display interface;

[0211] The display module is also used for:

[0212] In response to an input operation on an input area in a display interface, displaying input content in the input area includes:

[0213] In response to a triggering operation on the multimedia identifier, displaying a multimedia source identifier in an area associated with the multimedia identifier;

[0214] In response to a triggering operation on the multimedia source identifier, displaying multimedia material corresponding to the multimedia source identifier;

[0215] In response to a selection operation on the multimedia material, displaying a multimedia content identifier corresponding to the selected multimedia material in the input area;

[0216] or,

[0217] Display multimedia materials;

[0218] In response to a drag operation on the multimedia material, displaying a multimedia content identifier corresponding to the multimedia material in the input area;

[0219] or,

[0220] In response to detecting the first input for the multimedia material, a multimedia material prompt text and a corresponding multimedia content identifier are displayed in the input area.

[0221] In some embodiments, the type includes a file, and a file identifier is displayed in the display interface;

[0222] The display module is also used for:

[0223] In response to an input operation on an input area in a display interface, displaying input content in the input area includes:

[0224] In response to a trigger operation on the file identifier, displaying a file source identifier in an area associated with the file identifier;

[0225] In response to a triggering operation on the file source identifier, displaying the file material corresponding to the selected file source identifier;

[0226] In response to a selection operation on the file material, displaying a file content identifier corresponding to the selected file material in the input area;

[0227] or,

[0228] Display file material;

[0229] In response to a drag operation on the file material, displaying a file content identifier corresponding to the file material in the input area;

[0230] or,

[0231] In response to detecting that the file material cannot be identified, file abnormality prompt information is displayed on the display interface.

[0232] In some embodiments, the display interface includes a setting indicator;

[0233] The display module is also used for:

[0234] In response to an input operation on an input area in a display interface, displaying input content in the input area includes:

[0235] In response to a trigger operation on the setting identifier, displaying at least one attribute setting identifier in an area associated with the setting identifier;

[0236] In response to a setting operation on the attribute setting identifier, the corresponding attribute identifier is displayed in the input area; the attribute identifier displays the attribute value set by the setting operation.

[0237] In some embodiments, the attribute setting identifier includes a virtual object setting identifier, which is used to set whether a virtual object is used in the target video;

[0238] Alternatively, the attribute setting identifier includes a duration identifier, which is used to set the duration of the target video;

[0239] Alternatively, the attribute setting identifier includes a language identifier, which is used to set the language of the target video.

[0240] In some embodiments, the display interface displays a first generation indicator;

[0241] The display module is also used for:

[0242] In response to a triggering operation on the first generation identifier, a form page corresponding to the input content is displayed; the form page includes preset form items and corresponding item content, and the item content corresponds to the input content.

[0243] In some embodiments, the video generation module further includes:

[0244] The item content module is configured to determine corresponding item content from the description information based on the preset form item; and automatically fill in the item content in a corresponding area of the preset form item.

[0245] In some embodiments, the project content module is further configured to:

[0246] In response to an editing operation on the project content, updating the project content;

[0247] or,

[0248] Displaying an update indicator in an area associated with the project content;

[0249] In response to a triggering operation on the update identifier, the project content is updated.

[0250] In some embodiments, the form page includes a view indicator and / or an edit indicator for the input content;

[0251] The display module is also used for:

[0252] In response to a triggering operation on the viewing indicator, displaying all the input content in an area associated with the viewing indicator;

[0253] or,

[0254] In response to a trigger operation on the edit mark, the input area is returned to display; the input content of the input area corresponds to the current content of the form page.

[0255] In some embodiments, a second generation indicator is displayed on the form page;

[0256] The target text module is also used to:

[0257] In response to a triggering operation for the second generation identifier, the target text is updated based on current item content of the form page.

[0258] In some embodiments, the display module is further configured to:

[0259] In response to a triggering operation on the first generation identifier, displaying a form page corresponding to the input content includes:

[0260] In response to a trigger operation for the first generation identifier and detecting that no material can be extracted based on the input area,

[0261] Determining the item content corresponding to the preset form item in the form page in a designated material library based on the prompt text; and displaying the form page and the corresponding item content;

[0262] and / or,

[0263] Displays input prompt information for uploading input data.

[0264] In some embodiments, the display module is further configured to:

[0265] In response to detecting that no input information exists in the input area, displaying the historical records and / or the asset library on the display interface;

[0266] In response to a triggering operation on the historical record, displaying historical input content corresponding to the historical record in the input area;

[0267] Or, in response to a trigger operation for an asset identifier in the asset library, an asset form page of the target asset corresponding to the asset identifier is displayed, and the content of the asset form page is obtained based on the description information corresponding to the asset identifier and the corresponding input content in the input area.

[0268] In some embodiments, the asset form page displays a third generation indicator and a synchronization indicator;

[0269] The video generating device further includes:

[0270] A synchronization module is configured to synchronize all descriptive information about the target asset currently corresponding to the asset form page to the asset library in response to a trigger operation on the third generation identifier and when the synchronization identifier is in a preset state.

[0271] In some embodiments, the video generating apparatus further comprises:

[0272] A preview module, configured to preview the input content in an area associated with the input identifier in response to a triggering operation on the input identifier;

[0273] Alternatively, in response to a triggering operation on the input identifier, displaying a preview identifier in an area associated with the input identifier;

[0274] In response to a triggering operation on the preview indicator, the corresponding input content is previewed.

[0275] In some embodiments, the video generation module further includes:

[0276] The moving module is configured to move the input identifier to a cursor position of the moving operation in response to a moving operation on the input identifier.

[0277] In some embodiments, the data parsing module is further configured to:

[0278] Acquiring the corresponding input data based on the address identifier or identity identifier of the input identifier;

[0279] Parsing the image in the input data to obtain a first text description;

[0280] For a video in the input data, determining corresponding key frames based on scene changes of the video and parsing the key frames to obtain a second text description;

[0281] Parsing the file in the input data to obtain a third text description;

[0282] In addition, the video generation module is also used to: generate a target video based on the target text and the multiple types of input data, including: generating the target video based on at least one of the text in the input data, the first text description, the second text description or the third text description.

[0283] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.

[0284] The apparatus of the above embodiment is used to implement the corresponding video generation method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0285] Based on the same technical concept, corresponding to any of the above-mentioned embodiments and methods, the present disclosure also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the video generation method described in any of the above embodiments.

[0286] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable multimedia, and information storage can be achieved by any method or technology. Information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.

[0287] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the video generation method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.

[0288] Based on the same inventive concept, corresponding to the video generation method in any of the above embodiments, the present disclosure further provides a computer program product comprising computer program instructions. In some embodiments, when the computer program instructions are executed on a computer, the computer executes each step in each embodiment of the video generation method. For each step in each embodiment of the video generation method, the processor executing the corresponding step may belong to the corresponding execution entity.

[0289] The computer program product of the above embodiment is used to enable a processor to execute the video generation method described in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.

[0290] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.

[0291] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.

[0292] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.

[0293] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.

Claims

1. A video generation method, comprising: In response to an input operation on an input area in a display interface, displaying input content in the input area; wherein the input content includes at least one prompt text and multiple input identifiers, and the multiple input identifiers are used to indicate multiple types of input data; Parsing the multiple types of input data to obtain corresponding description information; Generating a target text from the corresponding description information according to the order of the prompt text and the input identifier in the input content; A target video is generated based on the target text and the multiple types of input data.

2. The method according to claim 1, wherein The type includes a link, and a link logo is displayed in the display interface; In response to an input operation on an input area in a display interface, displaying an input content input identifier in the input area includes: In response to a triggering operation on the link identifier, displaying a preset link prompt text and a corresponding link content identifier in the input area; Entering a target link in the link content identifier; Alternatively, in response to inputting a target link in the input area, a preset link prompt text and a corresponding link content identifier are displayed in the input area, and at least a portion of the target link is displayed in the link content identifier.

3. The method according to claim 2, further comprising: In response to detecting that the target link indicated by the link content identifier cannot be resolved, performing at least one of the following operations: Displaying abnormal prompt information in the associated area of the link content identifier, wherein the abnormal prompt information is used to indicate that the target link cannot be resolved; Alternatively, in response to detecting that the link resolution plug-in is not installed, displaying plug-in installation prompt information and / or a plug-in installation identifier in the display interface, wherein the plug-in installation prompt information is used to prompt the installation of the link resolution plug-in, and the plug-in installation identifier is used to trigger the installation of the link resolution plug-in; Alternatively, in response to detecting that the link resolution plug-in has been installed, plug-in usage prompt information is displayed in the display interface to prompt the user to use the link resolution plug-in.

4. The method according to claim 3, further comprising: In response to a triggering operation on the plug-in identifier corresponding to the link resolution plug-in, displaying a plug-in panel in an area associated with the plug-in identifier, wherein a link conversion identifier is displayed in the plug-in panel; In response to a triggering operation on the link conversion identifier, converting the target link to obtain a link conversion result, and displaying a conversion completion identifier; In response to a triggering operation on the conversion completion mark, at least a portion of the link conversion result is displayed in a link content mark used to indicate the link conversion result.

5. The method according to claim 2, further comprising: In response to detecting that a preset text is displayed in the link content identifier, displaying a permission confirmation identifier for the link access authority on the display interface; In response to a confirmation operation on the permission confirmation identifier, allowing generation of the target video; In response to detecting that the permission confirmation identifier is not confirmed, refusing to generate the target video, and / or displaying operation prompt information about the permission confirmation identifier on the display interface.

6. The method according to claim 1, wherein The type includes multimedia, and a multimedia logo is displayed in the display interface; In response to an input operation on an input area in a display interface, displaying input content in the input area includes: In response to a triggering operation on the multimedia identifier, displaying a multimedia source identifier in an area associated with the multimedia identifier; In response to a triggering operation on the multimedia source identifier, displaying multimedia material corresponding to the multimedia source identifier; In response to a selection operation on the multimedia material, displaying a multimedia content identifier corresponding to the selected multimedia material in the input area; or, Display multimedia materials; In response to a drag operation on the multimedia material, displaying a multimedia content identifier corresponding to the multimedia material in the input area; or, In response to detecting the first input for the multimedia material, a multimedia material prompt text and a corresponding multimedia content identifier are displayed in the input area.

7. The method according to claim 1, wherein The type includes file, and a file identifier is displayed in the display interface; In response to an input operation on an input area in a display interface, displaying input content in the input area includes: In response to a trigger operation on the file identifier, displaying a file source identifier in an area associated with the file identifier; In response to a triggering operation on the file source identifier, displaying the file material corresponding to the selected file source identifier; In response to a selection operation on the file material, displaying a file content identifier corresponding to the selected file material in the input area; or, Display file material; In response to a drag operation on the file material, displaying a file content identifier corresponding to the file material in the input area; or, In response to detecting that the file material cannot be identified, file abnormality prompt information is displayed on the display interface.

8. The method according to claim 1, wherein The display interface includes a setting indicator; In response to an input operation on an input area in a display interface, displaying input content in the input area includes: In response to a trigger operation on the setting identifier, displaying at least one attribute setting identifier in an area associated with the setting identifier; In response to a setting operation on the attribute setting identifier, the corresponding attribute identifier is displayed in the input area; the attribute identifier displays the attribute value set by the setting operation.

9. The method according to claim 8, wherein The attribute setting identifier includes a virtual object setting identifier, which is used to set whether a virtual object is used in the target video; Alternatively, the attribute setting identifier includes a duration identifier, which is used to set the duration of the target video; Alternatively, the attribute setting identifier includes a language identifier, which is used to set the language of the target video.

10. The method according to claim 1, wherein A first generation identifier is displayed in the display interface; The method further comprises: In response to a triggering operation on the first generation identifier, a form page corresponding to the input content is displayed; the form page includes preset form items and corresponding item content, and the item content corresponds to the input content.

11. The method according to claim 10, wherein the project content corresponds to the input content, comprising: Determining corresponding item content from the description information based on the preset form item; The corresponding area of the preset form item is automatically filled with the item content.

12. The method according to claim 10, further comprising: In response to an editing operation on the project content, updating the project content; or, Displaying an update indicator in an area associated with the project content; In response to a triggering operation on the update identifier, the project content is updated.

13. The method according to claim 10, wherein: The form page includes a viewing mark and / or an editing mark for the input content; The method further comprises: In response to a triggering operation on the viewing indicator, displaying all the input content in an area associated with the viewing indicator; or, In response to a trigger operation on the edit mark, the input area is returned to display; the input content of the input area corresponds to the current content of the form page.

14. The method according to claim 10, wherein: A second generation indicator is displayed on the form page; The method further comprises: In response to a triggering operation for the second generation identifier, the target text is updated based on current item content of the form page.

15. The method according to claim 10, wherein In response to a triggering operation on the first generation identifier, displaying a form page corresponding to the input content includes: In response to a trigger operation for the first generation identifier and detecting that no material can be extracted based on the input area, Determining the item content corresponding to the preset form item in the form page in a designated material library based on the prompt text; and displaying the form page and the corresponding item content; and / or, Displays input prompt information for uploading input data.

16. The method according to claim 1, further comprising: In response to detecting that no input information exists in the input area, displaying the historical records and / or the asset library on the display interface; In response to a triggering operation on the historical record, displaying historical input content corresponding to the historical record in the input area; Or, in response to a trigger operation for an asset identifier in the asset library, an asset form page of the target asset corresponding to the asset identifier is displayed, and the content of the asset form page is obtained based on the description information corresponding to the asset identifier and the corresponding input content in the input area.

17. The method according to claim 16, wherein The asset form page displays a third generation indicator and a synchronization indicator; The method further comprises: In response to a trigger operation for the third generation identifier and the synchronization identifier being in a preset state, all descriptive information about the target asset currently corresponding to the asset form page is synchronized to the asset library.

18. The method of claim 1, further comprising: In response to a triggering operation on the input identifier, previewing the input content in an area associated with the input identifier; Alternatively, in response to a triggering operation on the input identifier, displaying a preview identifier in an area associated with the input identifier; In response to a triggering operation on the preview indicator, the corresponding input content is previewed.

19. The method of claim 1, further comprising: In response to a move operation on the input mark, the input mark is moved to a cursor position of the move operation.

20. The method according to claim 1, wherein Parsing the multiple types of input data to obtain corresponding description information includes: Acquiring the corresponding input data based on the address identifier or identity identifier of the input identifier; Parsing the image in the input data to obtain a first text description; For a video in the input data, determining corresponding key frames based on scene changes of the video and parsing the key frames to obtain a second text description; Parsing the file in the input data to obtain a third text description; And, generating a target video based on the target text and the multiple types of input data includes: generating the target video based on at least one of the text in the input data, the first text description, the second text description, or the third text description.

21. A video generating device, comprising: A display module, configured to display input content in an input area in a display interface in response to an input operation on the input area; wherein the input content includes at least one prompt text and multiple input identifiers, wherein the multiple input identifiers are used to indicate multiple types of input data; A data parsing module, configured to parse the multiple types of input data to obtain corresponding description information; A target text module, configured to generate a target text from the corresponding description information according to the order of the prompt text and the input identifier in the input content; A video generation module is used to generate a target video based on the target text and the multiple types of input data.

22. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method according to any one of claims 1 to 20 when executing the computer program.

23. A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method according to any one of claims 1 to 20.

24. A computer program product comprising computer program instructions, which, when executed on a computer, cause the computer to perform the method according to any one of claims 1 to 20.