Video processing method and related device
Through multimodal encoding, the problem of relying on manual design for video effects is solved, and efficient and automated video processing is achieved, reducing labor costs and improving presentation effects.
Patent Information
- Application Number
- PCT/CN2024/137517
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-26
- Filing Date
- 2024-12-06
- Publication Date
- 2025-07-03
AI Technical Summary
In the prior art, video effects are added relying on manual design strategies, resulting in high labor costs, low efficiency and poor results.
Through multimodal encoding technology, the pending video and input text are encoded, effect configuration information is generated, and the video is rendered based on this information to achieve automatic effect addition.
It reduces labor costs, improves video processing efficiency and presentation effects, can cope with complex scenarios and meet personalized needs.
Smart Images

Figure CN2024137517_03072025_PF_FP_ABST
Abstract
Description
Video processing method and related equipment
[0001] This application claims priority to the Chinese invention patent application with application number 202311815841.3 filed on December 26, 2023 and titled “Video Processing Method and Related Equipment”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present disclosure relates to the field of computer technology, and in particular to a video processing method and related equipment. Background Art
[0003] Currently, adding effects to videos often relies on manually designed logic, breaking the effect-adding task into multiple modules and then concatenating the results from each module. This makes the video effect highly dependent on design strategies, resulting in increased labor costs and low video processing efficiency. Summary of the Invention
[0004] The present disclosure provides a video processing method, apparatus, device, storage medium and program product to solve the technical problem of low efficiency of video processing to a certain extent.
[0005] In a first aspect of the present disclosure, a video processing method is provided, comprising: obtaining a video to be processed and an input text; performing multimodal encoding based on the video to be processed and the input text to obtain a multimodal encoding result; generating effect configuration information based on the multimodal encoding result; and rendering the video to be processed based on the effect configuration information to obtain a target video.
[0006] In a second aspect of the present disclosure, a video processing device is provided, including: an acquisition module for acquiring a video to be processed and an input text; an encoding module for performing multimodal encoding based on the video to be processed and the input text to obtain a multimodal encoding result; a model module for generating effect configuration information based on the multimodal encoding result; and a rendering module for rendering the video to be processed based on the effect configuration information to obtain a target video.
[0007] In a third aspect of the present disclosure, an electronic device is provided, comprising one or more processors, a memory; and one or more programs, wherein the one or more programs are stored in the memory and executed by the one or more processors, and the programs include instructions for executing the method described in the first aspect.
[0008] According to a fourth aspect of the present disclosure, a non-volatile computer-readable storage medium containing a computer program is provided. When the computer program is executed by one or more processors, the processors are caused to execute the method described in the first aspect.
[0009] In a fifth aspect of the present disclosure, a computer program product is provided, comprising computer program instructions, which, when executed on a computer, cause the computer to execute the method described in the first aspect.
[0010] As can be seen from the above, the video processing method and related equipment provided by this disclosure perform multimodal encoding of the video to be processed and input text, generate corresponding effect configuration information, and then render the video to be processed based on this effect configuration information to add effects. This method can automatically add appropriate effects to videos, reducing labor costs and improving video processing efficiency and presentation quality. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following briefly introduces the drawings required for use in the embodiments or related technical descriptions. Obviously, the drawings described below are only embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0012] FIG1 is a schematic diagram of a video processing architecture according to an embodiment of the present disclosure.
[0013] FIG2 is a schematic diagram of the hardware structure of an exemplary electronic device according to an embodiment of the present disclosure.
[0014] FIG3 is a schematic flowchart of a video processing method according to an embodiment of the present disclosure.
[0015] FIG4 is a schematic diagram of a video processing method according to an embodiment of the present disclosure.
[0016] FIG5 is a schematic diagram of a language model according to an embodiment of the present disclosure.
[0017] FIG6 is a schematic diagram of a video processing device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0018] In order to make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the present disclosure is further described in detail below in conjunction with specific embodiments and with reference to the accompanying drawings.
[0019] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the usual meanings understood by people with ordinary skills in the field to which the present disclosure belongs. The "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. "Include" or "comprise" and similar words mean that the elements or objects appearing before the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects. "Connect" or "connected" and similar words are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Up", "down", "left", "right" and the like are only used to indicate relative position relationships. When the absolute position of the described object changes, the relative position relationship may also change accordingly.
[0020] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0021] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.
[0022] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0023] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0024] FIG1 shows a schematic diagram of a video processing architecture according to an embodiment of the present disclosure. Referring to FIG1 , the video processing architecture 100 may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 may be connected via a wired or wireless network 130. In some embodiments, the server 110 may be an independent physical server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and CDN.
[0025] Terminal 120 can be implemented in hardware or software. For example, when implemented in hardware, terminal 120 can be any electronic device with a display screen that supports page display, including but not limited to smartphones, tablet computers, e-book readers, laptop computers, and desktop computers. When terminal 120 is implemented in software, it can be installed in the electronic devices listed above; it can be implemented as multiple software or software modules (such as software or software modules used to provide distributed services), or it can be implemented as a single software or software module, and no specific limitations are given here.
[0026] It should be noted that the video processing method provided in the embodiments of the present application can be executed by the terminal 120 or by the server 110. It should be understood that the number of terminals, networks, and servers in FIG1 is for illustration only and is not intended to limit the number thereof. Any number of terminals, networks, and servers can be used as required.
[0027] Figure 2 shows a schematic diagram of the hardware structure of an exemplary electronic device 200 provided by an embodiment of the present disclosure. As shown in Figure 2, electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. In some embodiments, processor 202, memory 204, network module 206, and peripheral interface 208 are communicatively connected to each other within electronic device 200 via bus 210.
[0028] The processor 202 may be a central processing unit (CPU), a video processor, a neural network processor (NPU), a microcontroller (MCU), a programmable logic device (PLD), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), or one or more integrated circuits. The processor 202 may be configured to perform functions related to the technology described in this disclosure. In some embodiments, the processor 202 may also include multiple processors integrated into a single logical component. For example, as shown in FIG2 , the processor 202 may include multiple processors 202 a, 202 b, and 202 c.
[0029] The memory 204 can be configured to store data (e.g., instructions, computer codes, etc.). As shown in Figure 2, the data stored in the memory 204 can include program instructions (e.g., program instructions for implementing the video processing method of the embodiment of the present disclosure) and data to be processed (e.g., the memory can store configuration files of other modules, etc.). The processor 202 can also access the program instructions and data stored in the memory 204, and execute the program instructions to operate on the data to be processed. The memory 204 can include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 204 can include a random access memory (RAM), a read-only memory (ROM), an optical disk, a magnetic disk, a hard disk, a solid-state drive (SSD), a flash memory, a memory stick, etc.
[0030] The network module 206 can be configured to provide the electronic device 200 with communication with other external devices via a network. The network can be any wired or wireless network capable of transmitting and receiving data. For example, the network can be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC)), a cellular network, the Internet, or a combination thereof. It will be appreciated that the type of network is not limited to the specific examples above. In some embodiments, the network module 306 can include any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, and the like.
[0031] The peripheral interface 208 can be configured to connect the electronic device 200 to one or more peripheral devices to implement information input and output. For example, the peripheral devices can include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, and various sensors, and output devices such as a display, a speaker, a vibrator, and an indicator light.
[0032] The bus 210 can be configured to transmit information between the various components of the electronic device 200 (e.g., the processor 202, the memory 204, the network module 206, and the peripheral interface 208), such as an internal bus (e.g., a processor-memory bus), an external bus (USB port, PCI-E bus), etc.
[0033] It should be noted that although the architecture of the electronic device 200 shown above only includes the processor 202, the memory 204, the network module 206, the peripheral interface 208, and the bus 210, in a specific implementation, the architecture of the electronic device 200 may also include other components necessary for normal operation. In addition, those skilled in the art will understand that the architecture of the electronic device 200 may also include only the components necessary to implement the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.
[0034] Currently, adding effects to videos often involves breaking down the effect-addition task into multiple task modules based on manually designed logic, and then concatenating the results of each task module. For example, various tasks are broken down through manually designed logic, and algorithms are then developed separately for modules such as keyword extraction, sound effects and special effects addition rules, and text recommendation, which are then assembled together. This makes video effects highly dependent on design strategies and consumes excessive manpower during development iterations. Not only does each module require manual maintenance and updates, resulting in significant communication costs, increased labor costs, and low video processing efficiency, but because each task module is developed and maintained by different personnel, the effects may be inconsistent across the modules. For example, differences in style between modules can lead to poor presentation after the effects are added. Therefore, improving the presentation and processing efficiency of added video effects and reducing labor costs have become urgent technical challenges.
[0035] In light of this, embodiments of the present disclosure provide a video processing method and related devices. This method performs multimodal encoding of the video to be processed and input text, generates corresponding effect configuration information, and then renders the video to be processed based on this effect configuration information to achieve the addition of effects. This method automatically adds appropriate effects to the video, reducing labor costs and improving video processing efficiency and presentation quality.
[0036] 3 , which shows a schematic flow chart of a video processing method according to an embodiment of the present disclosure. In FIG3 , the video processing method 300 may further include the following steps S310 , S320 , S330 , and S340 .
[0037] In step S310, a video to be processed and input text are obtained.
[0038] In some embodiments, the video to be processed can be obtained based on local upload or via the network. The input text can refer to text input by the user based on the input interface, or obtained by other modal data input by the user. For example, the user can input audio data, and voice recognition is performed based on the audio data to obtain the input text. For example, an input interface can be provided to the user through a display interface, and the user can perform input operations in the input interface to input text text. Referring to Figure 4, Figure 4 shows a schematic diagram of a video processing method according to an embodiment of the present disclosure. In Figure 4, the user can input at least one video to be processed video and input text text based on the input interface.
[0039] In step S320, multimodal encoding is performed based on the video to be processed and the input text to obtain a multimodal encoding result.
[0040] In some embodiments, multimodal encoding may refer to encoding data of multiple modalities, where the multiple modalities may refer to text modules, video modalities, audio modalities, and the like.
[0041] In some embodiments, multimodal encoding is performed based on the video to be processed and the input text to obtain a multimodal encoding result, including: encoding based on the video to be processed to obtain a first encoding result; generating corresponding effect condition information based on the input text; text encoding the input text and the effect condition information to obtain a second encoding result; and fusing the second encoding result and the first encoding result based on a preset format to obtain the multimodal encoding result.
[0042] In some embodiments, effect condition information may refer to information about the conditions that must be met when adding effects to a video. Adding effects to a video may include adding stickers, music, sound effects, or other special effects based on the original video material input by the user. Effect conditions can be set by the user or default. The effect condition information may include at least one of the following: effect object, which may refer to the object to which the effect is added, for example, if an effect is added to input text, picture or title, the input text, picture or title is the effect object; effect type, which may refer to the category of the added effect, and each category may include one or more effect types, for example, the effect type is text template, sticker, flower font, the effect type of the text template may include several different text effect types, and the effect type of the sticker may include several sticker patterns with different effects; priority of effect type, for example, using stickers as much as possible may mean that stickers have a higher priority than other effect types, using flower fonts first may mean that flower fonts have a higher priority than other effect types, etc.; constraint conditions for effect type, for example, the number of effect types of effect type A should not exceed i, and the number of effect type B should not exceed j; applicable scenarios, such as non-marketing scenarios; effect frequency, which may refer to the number of objects with added effects in the total number of objects, for example, the effect frequency is controlled to be less than or equal to k.
[0043] Specifically, the user can input the video to be processed and the input text, and the multimodal encoding result may include a fusion of the encoding results of the two. The preset format may refer to a format for connecting the first encoding result and the second encoding result, for example, splicing the second encoding result after the first encoding result. Referring to Figure 4, the position can be reserved before the text is input based on the video frame encoding results frame_m1, frame_m2, ..., frame_m (m is a positive integer) of the first video frame in the preset format.<frame_m1> 、<frame_m2> 、……、<frame_m> The video frame encoding results frame_m1, frame_m2, ..., frame_m can be added to the corresponding reserved positions respectively.<frame_m1> 、<frame_m2> 、……、<frame_m> , thereby obtaining the first encoding result.
[0044] In some embodiments, encoding is performed based on the video to be processed to obtain a first encoding result, including: determining a first video frame in the video to be processed; visually encoding the first video frame to obtain a visual encoding result; mapping the visual encoding result to a preset feature space based on the semantics of the visual encoding result to obtain a video frame encoding result; and obtaining a first encoding result based on the video frame encoding result.
[0045] In some embodiments, the preset feature space may refer to the feature space corresponding to the trained language model. Referring to Figure 4, at least part of the video frames in the processed video may be visually encoded to obtain a visual encoding result, and then the visual encoding result may be semantically mapped to the preset feature space of the language model, and the visual encoding result may be converted into an encoding result that the language model can understand, thereby achieving semantic alignment between the visual encoding result and the language model. Specifically, a data set with a high semantic relevance (for example, greater than or equal to a certain threshold) may be used as training data to train a visual mapper, and under the premise of keeping the weight parameters of the language model unchanged, the weight parameters of the visual mapper may be adjusted until the training requirements are met.
[0046] In some embodiments, determining the first video frame in the video to be processed includes: in response to determining that the duration of the video to be processed is greater than or equal to a preset duration, extracting a video frame in the video to be processed based on a preset time interval to obtain the first video frame; in response to determining that the duration of the video to be processed is less than a preset duration, extracting a video frame in the video to be processed to obtain the first video frame.
[0047] In some embodiments, when the user does not input text associated with the video to be processed, and the text cannot be obtained based on the video to be processed, the video frame can be directly extracted from the video to be processed. For example, for a video to be processed that is greater than or equal to a preset duration (e.g., N seconds), one frame is extracted based on a preset time interval, and for a video to be processed that is less than the preset duration, one frame is extracted as the first video frame. In some embodiments, the preset duration is the same as the preset time interval.
[0048] In some embodiments, determining the first video frame in the video to be processed includes: segmenting the input text to obtain at least one text segment; and determining the video frame in the video to be processed that has the greatest semantic relevance to the text segment as the first video frame.
[0049] In some embodiments, input text may refer to text entered by a user related to the video to be processed, such as text describing the content presented in the video to be processed or dialogue text in the video to be processed. Input audio may refer to audio related to the video to be processed, such as dialogue audio in the video to be processed; it may also refer to audio independent of the video to be processed, such as music. The first text may be text derived based on the video to be processed, the input text, or the input audio. Specifically, the priority of the input text is higher than the audio related to the video to be processed. When there is input text, the input text can be used as the first text; when there is no input text but there is audio related to the video to be processed, the audio related to the video to be processed can be speech recognized to obtain the first text; when there is no input text and no audio related to the video to be processed, the first text can be obtained based on the video to be processed itself. For example, when text is displayed in the video to be processed, text recognition can be performed on the video to be processed, and the text extracted from it can be used as the first text; when there is no text but audio is present in the video to be processed, audio data can be extracted from the video to be processed, and speech recognition can be performed on the audio data to obtain the first text; when there is neither text nor audio in the video to be processed, the corresponding first text can be determined based on the image content displayed by the video to be processed. Referring to Figure 4, the input text text is segmented to obtain at least one text segment context_1, context_2, context_3, ..., context_n, where n is a positive integer.
[0050] In some embodiments, a first encoding result is obtained based on the video frame encoding result, including: performing text encoding on the text segment to obtain a text encoding result including a reserved position in the preset feature space; adding the video frame encoding result to the reserved position and fusing it with the text encoding result to obtain the first encoding result.
[0051] In some embodiments, the text segment can be encoded, and a reserved position can be reserved in the text encoding result for the video frame encoding result of the first video frame corresponding to the text segment so that the two can be fused. As shown in FIG4 , the reserved position can be set after the text segments context_1, context_2, context_3, ..., context_n.<frame_n1> 、<frame_n2> 、……、<frame_n> , perform text encoding to obtain the text encoding result. After visual encoding and semantic mapping of the first video frame, the video frame encoding results frame_n1, frame_n2, ..., frame_n corresponding to the text segment are obtained. The video frame encoding results frame_n1, frame_n2, ..., frame_n can be added to the corresponding reserved positions respectively.<frame_n1> 、<frame_n2> 、……、<frame_n> , thereby obtaining the first encoding result of fusion multi-modality.
[0052] In step S330, effect configuration information is generated based on the multimodal encoding result.
[0053] In some embodiments, the effect configuration information may refer to configuration information of an effect added to a video to be processed, and the effect configuration information may have a fixed format.
[0054] In some embodiments, generating effect configuration information based on the multimodal encoding result includes: inputting the multimodal encoding result into a trained language model to generate the effect configuration information; the effect configuration information includes at least one of the following: effect object, presentation position of the effect object, presentation time of the effect object, size of the effect object, effect parameters of the effect object and corresponding parameter values, and path of the effect object.
[0055] In some embodiments, the presentation position of the effect object may refer to the position where it is displayed in the video; the presentation time may refer to the start and end time of its appearance in the video; and the path may refer to the storage location or network resource address of the effect object. Specifically, the multimodal encoding result may be input into a trained language model LLM, which is used to generate corresponding effect configuration information based on the video to be processed and the input text input by the user. The format of the effect configuration information may include effect object-effect type-effect parameter value, or effect type-effect parameter, or music: <music type><music style>-music name. For example, as shown in Figure 4, the effect configuration information may include: the text segment context-n1 does not have any effects added, part of the text text1 in the text segment context-n2 uses the first effect in the text template with the effect type, and also uses the first sticker in the sticker with the effect type to add effects, the text segment context-n3 uses the second flower character effect in the flower character with the effect type, and part of the text text2 in the text segment context-n4 uses the first effect in the text template with the effect type and the first flower character effect in the flower character with the effect type to add, and the entire video to be processed can add music music: <music type><music style>-music name.
[0056] In some embodiments, the language model includes a pre-trained language model and an adaptation model, the adaptation model is connected in parallel with at least part of the linear layers in the pre-trained language model, and the adaptation model includes a dimensionality reduction network and a dimensionality increase network.
[0057] In some embodiments, the pre-trained language model may be a pre-trained model that generates text data based on input multimodal data. The adaptation model may be provided in parallel with at least some of the linear layers in the pre-trained language model.
[0058] In some embodiments, method 300 also includes: inputting the training data into the pre-trained language model to obtain a first processing result; inputting the training data into the adaptation model for dimensionality reduction processing and then dimensionality increase processing to obtain a second processing result; determining a loss function based on the sum of the first processing result and the second processing result; adjusting the weight parameters of the adaptation model to minimize the loss function to obtain the trained language model.
[0059] In some embodiments, the training data may include a multimodal training coding result obtained by multimodal coding based on the training sample, and corresponding effect training information. The training sample may include a training video, a corresponding description training text, and an effect prompt training text. The video frame is extracted from the training video for feature extraction and semantic mapping to obtain a video frame training coding result, and the description training text and the effect prompt training text are respectively text-coded to obtain a description text training coding result and an effect prompt training coding result. The corresponding video frame training coding result, the description text training coding result, and the effect prompt training coding result can be spliced to obtain a multimodal training coding result. The corresponding effect training information can be manually constructed based on the training sample. The training data is input into the pre-trained language model and the adaptation model, the weight parameters of the pre-trained language model are kept unchanged, and the weight parameters in the adaptation model are adjusted to obtain a trained language model. Specifically, referring to Figure 5, Figure 5 shows a schematic diagram of a language model according to an embodiment of the present disclosure. As shown in Figure 5, the modal training encoding result x is input into the pre-trained language model to obtain the first processing result. The modal training encoding result x of length d is input into the adaptation model. First, dimensionality reduction processing is performed to obtain an intermediate processing result of length r, and then dimensionality increase processing is performed to obtain the second processing result. The first processing result and the second processing result are added together to obtain the output result h. The loss function h can be obtained based on the difference between the output result h and the corresponding effect training information. The weight parameters of the dimensionality reduction network and the dimensionality increase network in the adaptation model are adjusted to minimize the loss function h, and the trained language model can be obtained.
[0060] In step S340, the video to be processed is rendered based on the effect configuration information to obtain a target video.
[0061] In some embodiments, the target video may refer to a video after the effect is added to the video to be processed. Specifically, the effect information may be formatted, rendering parameters may be configured, and portions that are not renderable or do not need to be rendered may be removed. Rendering may then be performed to obtain the target video with the effect added.
[0062] According to the video processing method of the embodiment of the present disclosure, the task of adding effects to the video is learned using a trained language model, and a data set is constructed based on the user's input and output. The input may include the user's audio, video and text, and the output is formatted effect configuration information, including packaging points, added effects, added text and text arrangement. The trained language model can also generate effect information based on a large amount of high-quality online data, with lower labor costs, and is data-driven rather than relying on manual design strategies, and can better cope with complex scenarios. Through simple training, tens of thousands of special effects, stickers, etc. can be achieved. At the same time, based on real online data, the generated target video results are more suitable and can better meet the personalized needs of users.
[0063] It should be noted that the method of the embodiments of the present disclosure can be performed by a single device, such as a computer or server. The method of the embodiments of the present disclosure can also be applied in a distributed scenario, where multiple devices cooperate to perform the method. In such a distributed scenario, one of the multiple devices may only perform one or more steps of the method of the embodiments of the present disclosure, and the multiple devices will interact with each other to complete the method.
[0064] It should be noted that the above description is limited to some embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in an order different from that described in the above embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0065] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a video processing device, see Figure 6, the video processing device includes: an acquisition module, used to acquire the video to be processed and the input text; an encoding module, used to perform multimodal encoding based on the video to be processed and the input text to obtain a multimodal encoding result; a model module, used to configure effect information based on the multimodal encoding result; a rendering module, used to render the video to be processed based on the effect configuration information to obtain a target video.
[0066] For the convenience of description, the above devices are described as being functionally divided into various modules. Of course, when implementing the present disclosure, the functions of each module can be implemented in the same or multiple software and / or hardware.
[0067] The apparatus of the above embodiment is used to implement the corresponding video processing method in any of the above embodiments, and has the beneficial effects of the corresponding method embodiment, which will not be described in detail here.
[0068] Based on the same technical concept, corresponding to any of the above-mentioned embodiment methods, the present disclosure also provides a non-transitory computer-readable storage medium, which stores computer instructions, and the computer instructions are used to enable the computer to execute the video processing method described in any of the above embodiments.
[0069] The computer-readable media of this embodiment include permanent and non-permanent, removable and non-removable media that can be used to store information by any method or technology. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, read-only compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device.
[0070] The computer instructions stored in the storage medium of the above embodiment are used to enable the computer to execute the video processing method described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0071] Those skilled in the art should understand that the discussion of any of the above embodiments is merely illustrative and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples. Within the scope of the present disclosure, the technical features in the above embodiments or different embodiments may be combined, the steps may be implemented in any order, and there are many other variations of the different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of simplicity.
[0072] In addition, to simplify the description and discussion, and so as not to obscure the embodiments of the present disclosure, known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided figures. In addition, devices may be shown in the form of block diagrams to avoid obscuring the embodiments of the present disclosure, and this also takes into account the fact that the details of the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the purview of those skilled in the art). Where specific details (e.g., circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be implemented without these specific details or with variations in these specific details. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0073] Although the present disclosure has been described in conjunction with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those skilled in the art based on the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed.
[0074] The embodiments of the present disclosure are intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the embodiments of the present disclosure should be included in the scope of protection of the present disclosure.
Claims
1. A video processing method, the method comprising: Obtaining a video to be processed and input text; Performing multi-modal encoding based on the video to be processed and the input text to obtain a multi-modal encoding result; Generating effect configuration information based on the multi-modal encoding result; Rendering the video to be processed based on the effect configuration information to obtain a target video.
2. The method according to claim 1, wherein, Performing multi-modal encoding based on the video to be processed and the input text to obtain a multi-modal encoding result, including: Encoding the video to be processed to obtain a first encoding result; Generating corresponding effect condition information based on the input text; Performing text encoding on the input text and the effect condition information to obtain a second encoding result; Fusing the second encoding result and the first encoding result based on a preset format to obtain the multi-modal encoding result.
3. The method according to claim 2, wherein Encoding the video to be processed to obtain a first encoding result, including: Determining a first video frame in the video to be processed; Performing visual encoding on the first video frame to obtain a visual encoding result; Mapping the visual encoding result to a preset feature space based on the semantics of the visual encoding result to obtain a video frame encoding result; Obtaining a first encoding result based on the video frame encoding result.
4. The method according to claim 3, wherein, Determining a first video frame in the video to be processed, including: In response to determining that the duration of the video to be processed is greater than or equal to a preset duration, extracting video frames in the video to be processed at a preset time interval to obtain the first video frame; In response to determining that the duration of the video to be processed is less than the preset duration, extracting one video frame in the video to be processed to obtain the first video frame.
5. The method according to claim 3, wherein, Determining a first video frame in the video to be processed, including: Performing sentence segmentation on the input text to obtain at least one text segment; Determining the video frame in the video to be processed with the highest semantic relevance to the text segment as the first video frame; And obtaining a first encoding result based on the video frame encoding result, including: Performing text encoding on the text segment to obtain a text encoding result including reserved positions in the preset feature space; Adding the video frame encoding result to the reserved positions and fusing it with the text encoding result to obtain the first encoding result.
6. The method according to claim 1, wherein Generating full effect configuration information based on the multi-modal encoding result, including: Inputting the multi-modal encoding result into a trained language model to generate the effect configuration information; the effect configuration information includes at least one of the following: an effect object, the presentation position of the effect object, the presentation time of the effect object, the size of the effect object, the effect parameters of the effect object and the corresponding parameter values, the path of the effect object.
7. The method according to claim 6, wherein The language model includes a pre-trained language model and an adaptation model, the adaptation model is connected in parallel with at least some linear layers in the pre-trained language model, and the adaptation model includes a dimensionality reduction network and a dimensionality increase network; The method further includes: Inputting training data into the pre-trained language model to obtain a first processing result; Inputting the training data into the adaptation model for dimensionality reduction processing and then performing dimensionality increase processing to obtain a second processing result; Determine a loss function based on the sum of the first processing result and the second processing result; Adjust the weight parameters of the adaptation model to minimize the loss function, and obtain the trained language model.
8. A video processing device, comprising: An acquisition module, configured to acquire a video to be processed and input text; An encoding module, configured to perform multi-modal encoding based on the video to be processed and the input text to obtain a multi-modal encoding result; A model module, configured to generate effect configuration information based on the multi-modal encoding result; A rendering module, configured to render the video to be processed based on the effect configuration information to obtain a target video.
9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the method according to any one of claims 1 to 7 when executing the program.
10. A non-transitory computer-readable storage medium, the non-transitory computer-readable storage medium stores computer instructions, and the computer instructions are used to cause a computer to execute the method according to any one of claims 1 to 7.
11. A computer program product, the computer program product is tangibly stored in a computer-readable storage medium and includes computer instructions, and the computer instructions cause a device to execute the method according to any one of claims 1 to 7 when executed by the device.
Citation Information
Patent Citations
Video generation method and device, equipment and medium
CN115022732A
Multi-modal generative abstract acquisition method based on cross fusion and reconstruction
CN115544244A
Image-text video processing method, device and equipment based on multi-mode neural network model
CN116485949A
Flocculation video generation method based on multi-mode and dynamic view angle adjustment
CN117177005A
Personalizing videos with nonlinear playback
US20220138474A1
Cited By
AI automatic editing method and related equipment
CN121309911A