Video processing method and related equipment

By multimodal encoding of video and text to generate effect configuration information and video rendering, the problem of low video processing efficiency in the prior art is solved, and automated and efficient video effects are achieved.

CN120223974APending Publication Date: 2025-06-27BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311815841.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-26
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing video processing technology is not efficient and relies on manual design strategies, resulting in high labor costs, low processing efficiency and poor performance.

Method used

By multi-modal encoding of the video to be processed and the input text, the effect configuration information is generated, and the video is rendered based on this information to achieve automatic effect addition.

Benefits of technology

It reduces labor costs, improves video processing efficiency and presentation effects, and realizes the automation and efficiency of adding video automatic effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120223974A_ABST
    Figure CN120223974A_ABST
Patent Text Reader

Abstract

The invention provides a video processing method and related equipment. The method comprises the following steps: acquiring a to-be-processed video and an input text; performing multi-modal coding based on the to-be-processed video and the input text to obtain a multi-modal coding result; generating effect configuration information based on the multi-modal coding result; and rendering the to-be-processed video based on the effect configuration information to obtain a target video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of computer technology, and in particular, to a video processing method and related devices. Background Art

[0002] Currently, adding effects to videos often decomposes the effect addition task into multiple task modules based on manually designed logic, and then stitches together the results of each task module. This makes the video effects highly dependent on the design strategy, resulting in increased labor costs and low efficiency of video processing. Summary of the Invention

[0003] The present disclosure provides a video processing method, apparatus, device, storage medium, and program product to solve the technical problem of low efficiency of video processing to a certain extent.

[0004] In a first aspect of the present disclosure, there is provided a video processing method, including:

[0005] Obtaining a video to be processed and input text;

[0006] Performing multimodal encoding based on the video to be processed and the input text to obtain a multimodal encoding result;

[0007] Generating effect configuration information based on the multimodal encoding result;

[0008] Rendering the video to be processed based on the effect configuration information to obtain a target video.

[0009] In a second aspect of the present disclosure, there is provided a video processing apparatus, including:

[0010] An obtaining module, configured to obtain a video to be processed and input text;

[0011] An encoding module, configured to perform multimodal encoding based on the video to be processed and the input text to obtain a multimodal encoding result;

[0012] A model module, configured to generate effect configuration information based on the multimodal encoding result;

[0013] A rendering module, configured to render the video to be processed based on the effect configuration information to obtain a target video.

[0014] In a third aspect of the present disclosure, there is provided an electronic device, including one or more processors, a memory; and one or more programs, where the one or more programs are stored in the memory and executed by the one or more processors, and the programs include instructions for executing the method according to the first aspect.

[0015] In a fourth aspect of the present disclosure, there is provided a non-volatile computer-readable storage medium containing a computer program, which, when executed by one or more processors, causes the processors to execute the method described in the first aspect.

[0016] In a fifth aspect of the present disclosure, there is provided a computer program product including computer program instructions, which, when running on a computer, cause the computer to execute the method described in the first aspect.

[0017] As can be seen from the above, a video processing method and related devices provided by the present disclosure perform multi-modal encoding on a video to be processed and input text, and generate corresponding effect configuration information; then, based on the effect configuration information, the video to be processed is rendered to add effects. It can automatically add appropriate effects to the video, reduce labor costs, and improve the processing efficiency and presentation effect of the video. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the present disclosure or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are only embodiments of the present disclosure. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0019] Figure 1 It is a schematic diagram of the video processing architecture of an embodiment of the present disclosure.

[0020] Figure 2 It is a schematic diagram of the hardware structure of an exemplary electronic device of an embodiment of the present disclosure.

[0021] Figure 3 It is a schematic flowchart of the video processing method of an embodiment of the present disclosure.

[0022] Figure 4 It is a schematic diagram of the video processing method of an embodiment of the present disclosure.

[0023] Figure 5 It is a schematic diagram of the language model of an embodiment of the present disclosure.

[0024] Figure 6 It is a schematic diagram of the video processing device of an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0025] To make the objectives, technical solutions, and advantages of the present disclosure clearer and more understandable, the following further elaborates on the present disclosure in detail with reference to specific embodiments and the accompanying drawings.

[0026] It should be noted that, unless otherwise defined, the technical terms or scientific terms used in the embodiments of the present disclosure should have the ordinary meanings understood by those of ordinary skill in the art to which the present disclosure belongs. The "first", "second" and similar terms used in the embodiments of the present disclosure do not denote any order, quantity or importance, but are only used to distinguish different components. Words such as "including" or "comprising" mean that the elements or items appearing before the word cover the elements or items listed after the word and their equivalents, without excluding other elements or items. Words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. "Upper", "lower", "left", "right", etc. are only used to represent relative positional relationships, and when the absolute position of the object being described changes, the relative positional relationship may also change accordingly.

[0027] It can be understood that, before using the technical solutions disclosed in the embodiments of the present disclosure, the types, usage scopes, usage scenarios, etc. of the personal information involved in the present disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0028] For example, when a user's active request is received, a prompt message is sent to the user to clearly prompt the user that the operation requested by the user will require obtaining and using the user's personal information. Thus, the user can autonomously choose whether to provide personal information to software or hardware such as an electronic device, an application program, a server or a storage medium that performs the operations of the technical solutions of the present disclosure according to the prompt message.

[0029] As an optional but non-limiting implementation manner, the manner of sending a prompt message to the user in response to receiving the user's active request may be, for example, in the form of a pop-up window, and the prompt message may be presented in text in the pop-up window. In addition, the pop-up window may also carry a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0030] It can be understood that the above process of notifying and obtaining the user's authorization is only illustrative and does not limit the implementation manners of the present disclosure, and other manners that meet relevant laws and regulations can also be applied to the implementation manners of the present disclosure.

[0031] Figure 1 shows a schematic diagram of the video processing architecture of the embodiments of the present disclosure. Refer to Figure 1, the video processing architecture 100 may include a server 110, a terminal 120, and a network 130 providing a communication link. The server 110 and the terminal 120 may be connected through the wired or wireless network 130. Among them, the server 110 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, security services, and CDN.

[0032] The terminal 120 may be implemented by hardware or software. For example, when the terminal 120 is implemented by hardware, it may be various electronic devices with a display screen and supporting page display, including but not limited to smartphones, tablets, e-book readers, laptop computers, desktop computers, and so on. When the terminal 120 device is implemented by software, it may be installed in the above-listed electronic devices; it may be implemented as multiple software or software modules (such as software or software modules for providing distributed services), or it may be implemented as a single software or software module, which is not specifically limited herein.

[0033] It should be noted that the video processing method provided by the embodiments of the present application may be executed by the terminal 120 or by the server 110. It should be understood that Figure 1 the numbers of the terminal, network, and server in

[0034] Figure 2 shows a schematic diagram of the hardware structure of an exemplary electronic device 200 provided by the embodiments of the present disclosure. As Figure 2 shown, the electronic device 200 may include: a processor 202, a memory 204, a network module 206, a peripheral interface 208, and a bus 210. Among them, the processor 202, the memory 204, the network module 206, and the peripheral interface 208 are communicatively connected to each other inside the electronic device 200 through the bus 210.

[0035] The processor 202 may be a central processing unit (CPU), a video processor, a neural network processor (NPU), a microcontroller (MCU), a programmable logic device, a digital signal processor (DSP), an application specific integrated circuit (ASIC), or one or more integrated circuits. The processor 202 may be used to execute functions related to the technology described in the present disclosure. In some embodiments, the processor 202 may further include multiple processors integrated as a single logic component. For example, asFigure 2 As shown, the processor 202 may include multiple processors 202a, 202b, and 202c.

[0036] The memory 204 may be configured to store data (e.g., instructions, computer code, etc.). As Figure 2 shown, the data stored in the memory 204 may include program instructions (e.g., program instructions for implementing the video processing method of the embodiments of the present disclosure) and data to be processed (e.g., the memory may store configuration files of other modules, etc.). The processor 202 may also access the program instructions and data stored in the memory 204 and execute the program instructions to operate on the data to be processed. The memory 204 may include a volatile storage device or a non-volatile storage device. In some embodiments, the memory 204 may include a random access memory (RAM), a read-only memory (ROM), an optical disc, a magnetic disk, a hard disk, a solid state drive (SSD), a flash memory, a memory stick, etc.

[0037] The network module 206 may be configured to provide communication with other external devices to the electronic device 200 via a network. The network may be any wired or wireless network capable of transmitting and receiving data. For example, the network may be a wired network, a local wireless network (e.g., Bluetooth, WiFi, near field communication (NFC), etc.), a cellular network, the Internet, or a combination of the above. It can be understood that the type of the network is not limited to the above specific examples. In some embodiments, the network module 306 may include any combination of any number of network interface controllers (NICs), radio frequency modules, transceivers, modems, routers, gateways, adapters, cellular network chips, etc.

[0038] The peripheral interface 208 may be configured to connect the electronic device 200 to one or more peripheral devices to implement information input and output. For example, the peripheral devices may include input devices such as a keyboard, a mouse, a touchpad, a touch screen, a microphone, various sensors, etc. and output devices such as a display, a speaker, a vibrator, an indicator light, etc.

[0039] The bus 210 may be configured to transmit information between various components of the electronic device 200 (e.g., the processor 202, the memory 204, the network module 206, and the peripheral interface 208), such as an internal bus (e.g., a processor-memory bus), an external bus (a USB port, a PCI-E bus), etc.

[0040] It should be noted that although the architecture of the above-mentioned electronic device 200 only shows the processor 202, the memory 204, the network module 206, the peripheral interface 208, and the bus 210, in the specific implementation process, the architecture of the electronic device 200 may further include other components necessary for normal operation. In addition, those skilled in the art can understand that the architecture of the above-mentioned electronic device 200 may also only include the components necessary to implement the solution of the embodiments of the present disclosure, and does not necessarily include all the components shown in the figure.

[0041] Currently, adding effects to videos often disassembles the effect addition task into multiple task modules based on manually designed logic, and then stitches together the results of each task module. For example, various tasks are disassembled through manually designed logic, and then algorithms are developed separately, such as keyword extraction, sound effect addition rules, copywriting recommendation and other modules, and assembled together. This makes the video effects strongly dependent on the design strategy, and consumes too much manpower during development and iteration. Not only does each module require manual maintenance and update, there is a large communication cost, resulting in an increase in labor costs and low video processing efficiency; moreover, since each task module is developed and maintained by different personnel, the effect consistency of each task module may be poor, such as style differences between each task module, resulting in poor presentation effects after the video is added with effects. Therefore, how to improve the presentation effect and processing efficiency after adding video effects and reduce labor costs has become an urgent technical problem to be solved.

[0042] In view of this, the embodiments of the present disclosure provide a video processing method and related devices. By performing multi-modal encoding on the video to be processed and the input text, and generating corresponding effect configuration information; then rendering the video to be processed based on the effect configuration information to achieve effect addition. It can automatically add appropriate effects to the video, reduce labor costs, and improve the processing efficiency and presentation effect of the video.

[0043] See Figure 3 , Figure 3 shows a schematic flowchart of the video processing method according to the embodiments of the present disclosure. Figure 3 In [the figure], the video processing method 300 may further include the following steps.

[0044] In step S310, obtain the video to be processed and the input text.

[0045] Among them, the video to be processed can be obtained based on local upload or via the network. The input text may refer to the text input by the user based on the input interface, or obtained from other modal data input by the user. For example, the user can input audio data, and the input text is obtained by performing speech recognition on the audio data. For example, an input interface can be provided to the user through the display interface, and the user can perform an input operation in the input interface to input the text text. SeeFigure 4 , Figure 4 shows a schematic diagram of a video processing method according to an embodiment of the present disclosure. Figure 4 In this, the user can input at least one video to be processed and input text based on the input interface.

[0046] In step S320, multimodal encoding is performed based on the video to be processed and the input text to obtain a multimodal encoding result.

[0047] Among them, multimodal encoding can refer to encoding data of multiple modalities, and multiple modalities can refer to a text module, a video modality, an audio modality, etc.

[0048] In some embodiments, performing multimodal encoding based on the video to be processed and the input text to obtain a multimodal encoding result includes:

[0049] Encoding the video to be processed to obtain a first encoding result;

[0050] Generating corresponding effect condition information based on the input text;

[0051] Performing text encoding on the input text and the effect condition information to obtain a second encoding result;

[0052] Fusing the second encoding result and the first encoding result based on a preset format to obtain the multimodal encoding result.

[0053] Among them, the effect condition information can refer to the information on the conditions that need to be met when adding effects to the video to be processed. Among them, adding effects to the video can refer to operations such as adding stickers, music sound effects, and various special effects based on the original video material input by the user. The effect conditions can be set by the user or can be default. Among them, the effect condition information can include at least one of the following: the effect object, which can refer to the object to which the effect is added. For example, when adding effects to the input text, the picture, or the title, the input text, the picture, or the title is the effect object; the effect type, which can refer to the category of the added effect, and each category can include one or more effect types. For example, the effect type is a text template, a sticker, or a fancy word. The effect type of the text template can include several different text effect types, and the effect type of the sticker can include several different sticker patterns with effects; the priority of the effect type. For example, try to use stickers, which means that the priority of the stickers is higher than other effect types, and give priority to using fancy words, which means that the priority of the fancy words is higher than other effect types, etc.; the constraint conditions of the effect type. For example, the number of effect types of using effect type A does not exceed i, and the number of effect type B does not exceed j; the applicable scenario, such as a non-marketing scenario; the effect frequency, which can refer to the number of objects to which the effect is added in the total number of objects. For example, the effect frequency is controlled to be less than or equal to k.

[0054] Specifically, the user can input the video to be processed and the input text, and the multi-modal encoding result can include the fusion of the encoding results of both. The preset format can refer to the format of concatenating the first encoding result and the second encoding result. For example, the second encoding result is spliced after the first encoding result. Refer to Figure 4 , for the video frame encoding results frame_m1, frame_m2, ……, frame_m (m is a positive integer) of the first video frame based on the preset format, positions <frame_m1>, <frame_m2>, ……, <frame_m> are reserved before the input text. The video frame encoding results frame_m1, frame_m2, ……, frame_m can be added to the corresponding reserved positions <frame_m1>, <frame_m2>, ……, <frame_m> respectively, so as to obtain the first encoding result.

[0055] In some embodiments, encoding based on the video to be processed to obtain a first encoding result includes:

[0056] Determine the first video frame in the video to be processed;

[0057] Perform visual encoding on the first video frame to obtain a visual encoding result;

[0058] Map the visual encoding result to a preset feature space based on the semantics of the visual encoding result to obtain a video frame encoding result;

[0059] Obtain the first encoding result based on the video frame encoding result.

[0060] Among them, the preset feature space can refer to the feature space corresponding to the trained language model. Refer to Figure 4 , visual encoding can be performed on at least some video frames in the video to be processed to obtain a visual encoding result, and then the visual encoding result is semantically mapped to the preset feature space of the language model, and the visual encoding result is converted into an encoding result that the language model can understand, so as to achieve the semantic alignment between the visual encoding result and the language model. Specifically, a dataset with relatively high semantic relevance (for example, greater than or equal to a certain threshold) can be used as training data to train a visual mapper. On the premise of keeping the weight parameters of the language model unchanged, the weight parameters of the visual mapper are adjusted until the training requirements are met.

[0061] In some embodiments, determining the first video frame in the video to be processed includes:

[0062] In response to determining that the duration of the video to be processed is greater than or equal to the preset duration, extract video frames from the video to be processed based on the preset time interval to obtain the first video frame;

[0063] In response to determining that the duration of the video to be processed is less than a preset duration, extract a video frame from the video to be processed to obtain the first video frame.

[0064] Wherein, when the user does not input text associated with the video to be processed and the text cannot be obtained based on the video to be processed, the video frame extraction of the video to be processed can be directly performed. For example, for a video to be processed with a duration greater than or equal to the preset duration (e.g., N seconds), extract one frame at a preset time interval, and for a video to be processed with a duration less than the preset duration, extract one frame as the first video frame. In some embodiments, the preset duration is the same as the preset time interval.

[0065] In some embodiments, determining the first video frame in the video to be processed includes:

[0066] Perform sentence segmentation on the input text to obtain at least one text segment;

[0067] Determine the video frame in the video to be processed with the highest semantic relevance to the text segment as the first video frame.

[0068] Wherein, the input text may refer to the text input by the user related to the video to be processed, such as the description text of the content presented in the video to be processed or the dialogue text in the video to be processed, etc. The input audio may refer to the audio related to the video to be processed, such as the dialogue audio of the video to be processed; it may also refer to the audio independent of the video to be processed, such as music. The first text may be the text obtained based on the video to be processed or the input text or the input audio. Specifically, the priority of the input text is higher than that of the audio related to the video to be processed. When there is input text, the input text can be used as the first text; when there is no input text but there is audio related to the video to be processed, speech recognition can be performed on the audio related to the video to be processed to obtain the first text; when there is no input text and audio related to the video to be processed, the first text can be obtained based on the video to be processed itself. For example, when there is text displayed in the video to be processed, text recognition can be performed on the video to be processed, and the text extracted therefrom is used as the first text; when there is no text displayed in the video to be processed but there is audio, audio data can be extracted from the video to be processed, and speech recognition is performed on the audio data to obtain the first text; when there is neither text displayed nor audio in the video to be processed, the corresponding first text can be determined based on the image content displayed in the video to be processed. See Figure 4 Perform sentence segmentation on the input text text to obtain at least one text segment context_1, context_2, context_3,..., context_n, where n is a positive integer.

[0069] In some embodiments, obtaining a first encoding result based on the video frame encoding result includes:

[0070] Performing text encoding on the text segment to obtain a text encoding result including reserved positions in the preset feature space;

[0071] Based on the video frame encoding result, adding it to the reserved positions and fusing it with the text encoding result to obtain the first encoding result.

[0072] Among them, text encoding can be performed on the text segment, and reserved positions can be left in the text encoding result for the video frame encoding result of the first video frame corresponding to the text segment, so that the two can be fused. As Figure 4 shown, reserved positions <frame_n1>, <frame_n2>,..., <frame_n> can be set behind the text segments context_1, context_2, context_3,..., context_n, and text encoding is performed to obtain a text encoding result. After performing visual encoding and semantic mapping on the first video frame, video frame encoding results frame_n1, frame_n2,..., frame_n corresponding to the text segments are obtained. The video frame encoding results frame_n1, frame_n2,..., frame_n can be added to the corresponding reserved positions <frame_n1>, <frame_n2>,..., <frame_n> respectively, so as to obtain the first encoding result that fuses multiple modalities.

[0073] In step S330, generating effect configuration information based on the multimodal encoding result.

[0074] Among them, the effect configuration information can refer to the configuration information of the effects added to the video to be processed, and the effect configuration information can have a fixed format.

[0075] In some embodiments, generating effect configuration information based on the multimodal encoding result includes:

[0076] Inputting the multimodal encoding result into a trained language model to generate the effect configuration information; the effect configuration information includes at least one of the following: the effect object, the presentation position of the effect object, the presentation time of the effect object, the size of the effect object, the effect parameters of the effect object and the corresponding parameter values, and the path of the effect object.

[0077] Among them, the presentation position of the effect object can refer to its position displayed in the video; the presentation time can refer to the start time and end time when it appears in the video; the path can refer to the storage location of the effect object or the network resource address. Specifically, the multi-modal encoding result can be input into the trained language model LLM, which is used to generate corresponding effect configuration information based on the video to be processed and the input text input by the user. The format of the effect configuration information can include effect object - effect type - effect parameter value, or effect type - effect parameter, or music: <music type><music style> - music name. For example, as Figure 4 shown, the effect configuration information can include: no effect is added to the text segment context-n1, part of the text text1 in the text segment context-n2 adopts the first effect in the text template as the effect type, and also adopts the first sticker in the sticker as the effect type for effect addition, the text segment context-n3 adopts the second fancy word effect in the fancy word as the effect type, part of the text text2 in the text segment context-n4 adopts the first effect in the text template and the first fancy word effect in the fancy word as the effect type for addition, and music music: <music type><music style> - music name can be added to the entire video to be processed.

[0078] In some embodiments, the language model includes a pre-trained language model and an adaptation model, the adaptation model is connected in parallel with at least some of the linear layers in the pre-trained language model, and the adaptation model includes a dimensionality reduction network and a dimensionality increase network.

[0079] Among them, the pre-trained language model can be a model that generates text data based on the input multi-modal data after pre-training. For at least some of the linear layers in the pre-trained language model, an adaptation model can be connected in parallel.

[0080] In some embodiments, method 300 further includes:

[0081] Inputting the training data into the pre-trained language model to obtain a first processing result;

[0082] Inputting the training data into the adaptation model for dimensionality reduction processing and then dimensionality increase processing to obtain a second processing result;

[0083] Determining a loss function based on the sum of the first processing result and the second processing result;

[0084] Adjusting the weight parameters of the adaptation model to minimize the loss function to obtain the trained language model.

[0085] Among them, the training data may include the multi-modal training coding results obtained by multi-modal coding based on training samples, and the corresponding effect training information. Among them, the training samples may include training videos, corresponding description training texts, and effect prompt training texts. Video frames are extracted from the training videos for feature extraction and then semantic mapping to obtain video frame training coding results, and the description training texts and effect prompt training texts are respectively subjected to text coding to obtain description text training coding results and effect prompt training coding results. The corresponding video frame training coding results, description text training coding results, and effect prompt training coding results can be concatenated to obtain the multi-modal training coding results. The corresponding effect training information can be constructed manually based on the training samples. Inputting the training data into the pre-trained language model and the adaptation model, keeping the weight parameters of the pre-trained language model unchanged, and adjusting the weight parameters in the adaptation model, a trained language model can be obtained. Specifically, see Figure 5 , Figure 5 which shows a schematic diagram of the language model according to an embodiment of the present disclosure. As Figure 5 shown, the modal training coding result x is respectively input into the pre-trained language model to obtain the first processing result. The modal training coding result x with a length of d is input into the adaptation model, first subjected to dimensionality reduction processing to obtain an intermediate processing result with a length of r, and then subjected to dimensionality increase processing to obtain the second processing result. The first processing result and the second processing result are added to obtain the output result h. Based on the difference between the output result h and the corresponding effect training information, a loss function h can be obtained. Adjusting the weight parameters of the dimensionality reduction network and the dimensionality increase network in the adaptation model to minimize the loss function h, a trained language model can be obtained.

[0086] In step S340, the video to be processed is rendered based on the effect configuration information to obtain a target video.

[0087] Among them, the target video may refer to the video after adding effects to the video to be processed. Specifically, the effect information can be formatted, rendering parameters can be configured, and the parts that cannot be rendered or do not need to be rendered can be removed. On this basis, rendering can be performed to obtain the target video with added effects.

[0088] According to the video processing method of the present disclosure embodiments, the effect addition task of the video is learned using a trained language model, and a dataset is constructed based on the user's input and output. The input can include the user's audio, video, and text, and the output is formatted effect configuration information, including packaging points, added effects, added text, and the arrangement of the text. The trained language model can also generate effect information based on a large amount of high-quality online data, with lower labor costs, and is data-driven without relying on manual design strategies, making it more capable of handling complex scenarios. Through simple training, tens of thousands of special effects, stickers, etc. can be achieved, and at the same time, based on the real online data, the generated target video results are more in line with and can better meet the personalized needs of users.

[0089] It should be noted that the method of the present disclosure embodiments can be executed by a single device, such as a computer or a server, etc. The method of this embodiment can also be applied to a distributed scenario and completed by multiple devices cooperating with each other. In this case of a distributed scenario, one of the multiple devices can only execute one or more steps of the method of the present disclosure embodiments, and these multiple devices will interact with each other to complete the described method.

[0090] It should be noted that some embodiments of the present disclosure have been described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the above embodiments and still achieve the desired results. Additionally, the processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0091] Based on the same technical concept, corresponding to the method of any of the above embodiments, the present disclosure also provides a video processing device. See Figure 6 , the video processing device includes:

[0092] An acquisition module, configured to acquire the video to be processed and the input text;

[0093] An encoding module, configured to perform multimodal encoding based on the video to be processed and the input text to obtain a multimodal encoding result;

[0094] A model module, configured to configure information based on the multimodal encoding result;

[0095] A rendering module, configured to render the video to be processed based on the effect configuration information to obtain a target video.

[0096] For the convenience of description, when describing the above device, it is divided into various modules according to functions for separate description. Of course, when implementing the present disclosure, the functions of each module can be implemented in one or more software and / or hardware.

[0097] The device in the above embodiment is used to implement the corresponding video processing method in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0098] Based on the same technical concept, corresponding to any of the above method embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the video processing method described in any of the foregoing embodiments.

[0099] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic tape magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0100] The computer instructions stored in the storage medium of the above embodiment are used to cause the computer to execute the video processing method described in any of the foregoing embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be elaborated here.

[0101] Those of ordinary skill in the art should understand that: the discussion of any of the above embodiments is only exemplary and is not intended to imply that the scope of the present disclosure (including the claims) is limited to these examples; under the concept of the present disclosure, the technical features in the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations in different aspects of the embodiments of the present disclosure as described above, which are not provided in detail for the sake of brevity.

[0102] Additionally, for simplicity of explanation and discussion, and so as not to render the embodiments of the present disclosure difficult to understand, well-known power / ground connections to integrated circuit (IC) chips and other components may or may not be shown in the provided drawings. Further, the devices may be shown in block diagram form in order to avoid rendering the embodiments of the present disclosure difficult to understand, and this also takes into account the fact that details regarding the implementation of these block diagram devices are highly dependent on the platform on which the embodiments of the present disclosure are to be implemented (i.e., these details should be fully within the understanding of those skilled in the art). In cases where specific details (such as circuits) are set forth to describe exemplary embodiments of the present disclosure, it will be apparent to those skilled in the art that the embodiments of the present disclosure may be practiced without these specific details or with variations of these specific details. Accordingly, these descriptions are to be regarded as illustrative rather than restrictive.

[0103] Although the present disclosure has been described in connection with specific embodiments thereof, many alternatives, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other memory architectures (such as dynamic RAM (DRAM)) may be used with the embodiments discussed.

[0104] Embodiments of the present disclosure are intended to cover all such alternatives, modifications, and variations that fall within the broad scope of the appended claims. Accordingly, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the embodiments of the present disclosure shall be included within the scope of protection of the present disclosure.

Claims

1. A video processing method, the method comprising: Obtaining a video to be processed and input text; Performing multi-modal encoding based on the video to be processed and the input text to obtain a multi-modal encoding result; Generating effect configuration information based on the multi-modal encoding result; Rendering the video to be processed based on the effect configuration information to obtain a target video.

2. The method according to claim 1, wherein Performing multi-modal encoding based on the video to be processed and the input text to obtain a multi-modal encoding result, including: Encoding the video to be processed to obtain a first encoding result; Generating corresponding effect condition information based on the input text; Performing text encoding on the input text and the effect condition information to obtain a second encoding result; Fusing the second encoding result and the first encoding result based on a preset format to obtain the multi-modal encoding result.

3. The method according to claim 2, wherein Encoding the video to be processed to obtain a first encoding result, including: Determining a first video frame in the video to be processed; Performing visual encoding on the first video frame to obtain a visual encoding result; Mapping the visual encoding result to a preset feature space based on the semantics of the visual encoding result to obtain a video frame encoding result; Obtaining a first encoding result based on the video frame encoding result.

4. The method according to claim 3, wherein Determining a first video frame in the video to be processed, including: In response to determining that the duration of the video to be processed is greater than or equal to a preset duration, extracting video frames in the video to be processed at a preset time interval to obtain the first video frame; In response to determining that the duration of the video to be processed is less than the preset duration, extracting one video frame in the video to be processed to obtain the first video frame.

5. The method according to claim 3, wherein Determining a first video frame in the video to be processed, including: Performing sentence segmentation on the input text to obtain at least one text segment; Determining the video frame in the video to be processed with the highest semantic relevance to the text segment as the first video frame; And obtaining a first encoding result based on the video frame encoding result, including: Performing text encoding on the text segment to obtain a text encoding result including reserved positions in the preset feature space; Adding the video frame encoding result to the reserved positions and fusing it with the text encoding result to obtain the first encoding result.

6. The method according to claim 1, wherein Generating full effect configuration information based on the multi-modal encoding result, including: Inputting the multi-modal encoding result into a trained language model to generate the effect configuration information; the effect configuration information includes at least one of the following: an effect object, the presentation position of the effect object, the presentation time of the effect object, the size of the effect object, the effect parameters of the effect object and the corresponding parameter values, the path of the effect object.

7. The method according to claim 6, wherein The language model includes a pre-trained language model and an adaptation model, the adaptation model is connected in parallel with at least some of the linear layers in the pre-trained language model, and the adaptation model includes a dimensionality reduction network and a dimensionality increase network; The method further includes: Inputting training data into the pre-trained language model to obtain a first processing result; Inputting the training data into the adaptation model for dimensionality reduction processing and then dimensionality increase processing to obtain a second processing result; Determine a loss function based on the sum of the first processing result and the second processing result; Adjust the weight parameters of the adaptation model to minimize the loss function, and obtain the trained language model.

8. A video processing device, comprising: An acquisition module, configured to acquire a video to be processed and input text; An encoding module, configured to perform multimodal encoding based on the video to be processed and the input text to obtain a multimodal encoding result; A model module, configured to generate effect configuration information based on the multimodal encoding result; A rendering module, configured to render the video to be processed based on the effect configuration information to obtain a target video.

9. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the program, the method according to any one of claims 1 to 7 is implemented.

10. A non-transitory computer-readable storage medium, which stores computer instructions for causing a computer to execute the method according to any one of claims 1 to 7.