Video processing method and device, equipment and medium
By obtaining video copy and using the target big model to generate text of packaging elements, the rendering operation is solved, and the video packaging editing in the existing technology is low efficiency and poor effect is achieved, and efficient and effective video packaging editing is achieved.
Patent Information
- Application Number
- CN202510134478.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-06
- Publication Date
- 2025-05-09
AI Technical Summary
The prior art is less efficient and has poor results when packaging and editing videos, making it difficult to meet users' demand for video enhancement.
By obtaining the original video, extracting the video copy in text form, and using the target model to generate target text containing multiple text-form packaging elements, performing rendering operations to generate target videos with packaging effects.
It realizes efficient packaging and editing of videos, improves the fun and user experience of videos, and improves the efficiency of packaging and editing.
Smart Images

Figure CN119967253A_ABST
Abstract
Description
Technical Field
[0001] Embodiments of the present disclosure relate to the field of video processing technology, and more particularly to a video processing method, apparatus, device, and medium. Background Art
[0002] With the continuous development of internet and digital media technologies, digital media is increasingly being used in people's daily lives and work, providing numerous conveniences. Currently, numerous platforms offering video services have emerged, which not only promotes the dissemination of videos online but also provides more ways for people to modify and process videos. To make videos more interesting, existing videos can be packaged and edited to create a lively and engaging effect. Currently, a video processing solution is needed. Summary of the Invention
[0003] The embodiments of the present disclosure describe a video processing method, apparatus, device, and medium.
[0004] According to a first aspect, a method is provided for obtaining an original video; obtaining a video copy in text form based on the original video; generating a target text for adding a packaging effect to the original video using a target macro model based on the original video and the video copy; the target text including a plurality of packaging elements in text form; and rendering the original video based on the target text to obtain a target video with a packaging effect.
[0005] According to a second aspect, a video processing device is provided, which includes: a first acquisition unit, configured to acquire an original video; a second acquisition unit, configured to acquire a video copy in text form based on the original video; a generation unit, configured to generate a target text for adding a packaging effect to the original video using a target macro model based on the original video and the video copy; the target text includes multiple packaging elements in text form; and a rendering unit, configured to perform a rendering operation on the original video based on the target text to obtain a target video with a packaging effect.
[0006] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute any one of the methods in the first aspect.
[0007] According to a fourth aspect, an electronic device is provided, comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, any one of the above methods in the first aspect is implemented.
[0008] According to the video processing solution provided by the embodiments of the present disclosure, an original video is obtained, and based on the original video, a text-based video copy is obtained. Based on the original video and the video copy, a target text is generated using a target macro model to add a packaging effect to the original video. The target text includes multiple text-based packaging elements. The original video is then rendered based on the target text to obtain a target video with the packaging effect. This enables packaging editing of existing videos, resulting in better effects after the packaging editing, improving the efficiency of packaging editing, and enhancing the user experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] Figure 1 1 is a schematic diagram of a video processing scenario according to an exemplary embodiment of the present disclosure;
[0010] Figure 2 is a schematic diagram of an exemplary system architecture to which the embodiments of the present disclosure are applied;
[0011] Figure 3 is a flowchart of a video processing method according to an exemplary embodiment of the present disclosure;
[0012] Figure 4 is a block diagram of a video processing device according to an exemplary embodiment of the present disclosure;
[0013] Figure 5 This is a schematic block diagram of an electronic device provided in some embodiments of the present disclosure. DETAILED DESCRIPTION
[0014] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.
[0015] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the technical solution of this disclosure based on the prompt message.
[0016] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.
[0017] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.
[0018] The technical solutions provided by the present disclosure are further described in detail below in conjunction with the accompanying drawings and embodiments. It will be understood that the specific embodiments described herein are merely for explaining the relevant inventions and are not intended to limit the inventions. It should also be noted that, for ease of description, only the portions relevant to the relevant inventions are shown in the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments of the present disclosure and the features therein may be combined with each other.
[0019] With the continuous development of Internet technology and digital media technology, digital media is increasingly used in people's work and life, providing many conveniences for people's work and life. At present, many platforms that provide video services for users have emerged, which not only promotes the dissemination of videos on the Internet, but also provides people with more ways to modify and process videos.
[0020] To make videos more interesting, existing videos can be packaged and edited to create a more engaging and engaging video. However, currently, in related technologies, this can only be accomplished manually by the user or through template editing. Consequently, the efficiency of video packaging and editing is low, resulting in poor video quality.
[0021] This disclosure provides a video processing solution that obtains an original video, obtains a text-based video copy based on the original video, and then uses a target macro model to generate a target text for adding a packaging effect to the original video. The target text includes multiple text-based packaging elements. The original video is then rendered based on the target text to obtain a target video with the packaging effect. This enables packaging editing of existing videos, resulting in better effects after the packaging editing, improving packaging editing efficiency, and enhancing the user experience.
[0022] See also Figure 1 , is a schematic diagram of a video processing scenario according to an exemplary embodiment.
[0023] like Figure 1As shown, first, a user can input video A to be packaged and edited into a video processing client through a terminal device. The video processing client can use ASR (Automatic Speech Recognition) technology to perform speech recognition on video A and obtain text-based copy B. Copy B can be the text content of the speech in video A converted into text. Copy B can include multiple copy sentences and the timestamp corresponding to each copy sentence.
[0024] Next, video A and document B can be preprocessed to obtain feature sequence C. Specifically, a text segmenter can be used to segment document B, obtaining feature sequence C1 in the text space. Simultaneously, video A is processed using visual encoding processing to obtain feature sequence C2 in the visual space. A pre-trained adapter is then used to convert feature sequence C2 into feature sequence C3 in the text space. Feature sequences C1 and C3 are then fused to obtain feature sequence C in the text space.
[0025] Next, feature sequence C is input into the large model M, which generates text D based on feature sequence C. Text D is a text with text-based packaging elements added to copy B. Finally, a rendering operation is performed based on text D to produce video E. Video E is a video that adds packaging effects to video A. These packaging effects may include, but are not limited to, displaying subtitles, dynamic subtitle effects, subtitle playback format, sound effects, animation effects, background music, and filter effects.
[0026] It should be noted that the visual encoder and text segmenter can be models trained in any reasonable way, and the large model M and the adapter can be trained at the same time.
[0027] It should be noted that Figure 1 The embodiment is described by taking the media processing client directly processing video A as an example. In other embodiments, the media processing client can also transmit video A to the media processing server deployed on the service platform through the network. The media processing server processes video A to obtain video E, and transmits video E to the media processing client through the network to provide video E to the user. For details, see Figure 2 .
[0028] Figure 2 A schematic diagram of an exemplary system architecture for applying an embodiment of the present disclosure.
[0029] like Figure 2 As shown, the system architecture 200 may include a terminal device 202, a network 203 and a server 204. It should be understood that Figure 2The number or type of terminal devices, networks, and servers in the embodiment are merely illustrative. Any number or type of terminal devices, networks, and servers may be used as required.
[0030] The network 203 is used to provide a medium for communication links between terminal devices and servers. The network 203 can include various connection types, such as wired or wireless communication links or optical fiber cables.
[0031] The terminal device 202 has a media processing client installed therein, and the terminal device 202 can interact with the server via the network 203 to receive or send requests or information. The terminal device 202 can be any electronic device, including but not limited to a smartphone, a tablet computer, a laptop computer, a desktop computer, and a smart wearable device.
[0032] Server 204 includes a media processing server. Server 204 can store and analyze received data, and can also send control commands or requests to terminal devices or other servers. The server can provide media processing services in response to user service requests. It is understood that a server can provide one or more services, and the same service can be provided by multiple servers.
[0033] based on Figure 2 In the system architecture shown, in the embodiment of the present disclosure, user 201 can input the original video to be processed through terminal device 202, and terminal device 202 can transmit the original video to server 204 via network 203. After receiving the original video, server 204 can obtain the video copy in text form based on the original media. Based on the original video and the video copy, it uses the target macro model to generate a target text for adding a packaging effect to the original video. The original video is then rendered based on the target text to obtain a target video with a packaging effect. Finally, server 204 returns the target video to terminal device 202 via network 203, allowing user 201 to view and save the target video through terminal device 202.
[0034] The present disclosure will be described in detail below with reference to specific embodiments.
[0035] Figure 3 The flowchart of a video processing method according to an exemplary embodiment is shown. This method can be applied to a terminal device. In this embodiment, for ease of understanding, an example is provided using a terminal device capable of installing a media processing client. Those skilled in the art will appreciate that the terminal device may include, but is not limited to, a mobile terminal device such as a smartphone, a smart wearable device, a tablet computer, and the like. The method may include the following steps:
[0036] like Figure 3 As shown, in step 301, an original video is obtained, and in step 302, a video copy in text form is obtained based on the original video.
[0037] In this embodiment, the original video is a video that needs to increase the packaging effect, and the original video may include at least an audio part and an image part. ASR technology can be used to perform speech recognition on the audio part of the original video to obtain a video copy in text form. The video copy may include at least a plurality of copy sentences, each copy sentence may correspond to multiple consecutive video frames in the original video, and therefore, the video copy can be used to produce subtitles for the original video. Optionally, the video copy may also include a timestamp corresponding to each copy sentence, or a video frame identifier corresponding to each copy sentence, and each copy sentence in the video copy may be marked in sequence with a serial number. The timestamp corresponding to the copy sentence can be used to indicate the starting moment corresponding to the copy sentence in the original video, and the video frame identifier corresponding to the copy sentence can be used to indicate the multiple video frames corresponding to the copy sentence in the original video.
[0038] For example, the video copy may include the following content (taking the video copy including the video frame identifier corresponding to the copy sentence as an example):
[0039] [<Video Frame A> 1. You choose the brand, style and price
[0040] <Video Frame B> 2. Ultimately, it all comes down to the quality and durability of the product itself.
[0041] <Video Frame C> 3. This range hood I recommend to you today
[0042] <Video frame D>4. Help you solve the above series of problems
[0043] <Video Frame E> 5. First of all, its entire appearance is made of black gold tempered glass
[0044] <Video Frame F> 6. Showcasing Understated Luxury
[0045] <Video Frame G> 7. Stylish and atmospheric
[0046] <Video Frame H> 8. Nano waterproof coating
[0047] 9. Better hygiene
[0048] <Video Frame Rate> 10, more energy-saving and power-saving
[0049] …………………………】
[0050] Among them, taking the first line as an example, the text "You choose the brand, style and price" is a copy sentence. The "1" before the copy sentence represents the serial number for marking the copy sentence in sequence, and <video frame A> represents the identifier of the multiple video frames corresponding to the copy sentence.
[0051] In step 303, based on the original video and the video text, a target text for increasing the packaging effect of the original video is generated using a target macro model.
[0052] In this embodiment, the packaging effect added to the original video may include but is not limited to displaying subtitles, dynamic effects of subtitles (such as subtitles jumping and rotating, etc.), the playback form of subtitles, sound effects, animation effects, background music, filter effects, etc. Based on the original video and video copy, the target large model can be used to generate a target text for adding a packaging effect to the original video. Specifically, first, the video copy can be segmented using a text segmenter to obtain a first feature sequence in the text space, and the original video can be visually encoded using a visual encoder to obtain a second feature sequence in the visual space. The pre-trained target adapter is then used to convert the second feature sequence into a third feature sequence in the text space. Then, the first feature sequence and the third feature sequence in the text space are feature fused to obtain a target feature sequence in the text space. Finally, the target feature sequence is input into the target large model to obtain the target text generated by the target large model.
[0053] Among them, the target text can be used to add a packaging effect to the original video. The target text can include multiple copy sentences, a timestamp corresponding to each copy sentence, and multiple packaging elements in text form. Each copy sentence can correspond to one or more packaging elements. According to the target object, the packaging elements can at least be divided into global elements for the entire video and sentence elements for the copy sentences. Among them, the global element can be located before or after the first copy sentence in the target text, and any sentence element can be located after the copy sentence corresponding to the sentence element in the target text. For the content of any packaging element, the content can at least include packaging effect description information, and the packaging effect description information can include the category corresponding to the packaging effect and the effect description for the packaging effect.
[0054] Furthermore, at least part of the content of the packaging element may also include packaging effect parameter information, which may include one or more of the following: a time offset corresponding to the packaging effect, a duration corresponding to the packaging effect, and an effect intensity corresponding to the packaging effect. The time offset corresponding to the packaging effect may be the time difference between the start time of the text sentence corresponding to the packaging effect in the original video and the time when the packaging effect appears. For example, if text sentence a starts at time t1 in the original video, and packaging effect b corresponding to text sentence a starts at time t2 in the original video, then the time offset corresponding to packaging effect b may be t2-t1.
[0055] Taking the video copy shown in step 302 as an example, the content of the target text that can be generated based on the video copy is as follows:
[0056]
【Default text color (#FFFFFF), font (Simplified Kaiti), font size (8)
[0057] 1. You choose the brand, style and price
[0058] Full sentence, animation (gold dust falling, duration 1.0)
[0059] Music (Fashion Rhythm Charm Bloom-Only Time Will Tell), offset (-0.1)
[0060] Filter (transparent), intensity (0.8), offset (-0.1), duration (16.5)
[0061] Keywords (brand), special effects (cool blue 3D floral lettering)
[0062] Keywords (style), special effects (cool blue 3D floral lettering)
[0063] Keywords (price), special effects (cool blue 3D floral text)
[0064] 2. Ultimately, it all comes down to the quality and durability of the product itself
[0065] Full sentence, animation (gold dust falling, duration 1.0)
[0066] Keywords (product), special effects (cool blue 3D floral lettering)
[0067] Keywords (quality), special effects (cool blue 3D floral text)
[0068] Keywords (durability), special effects (cool blue 3D floral lettering)
[0069] 3. The range hood I recommend to you today
[0070] Animation (gold dust falling, duration 1.0)
[0071] Keywords (range hood), special effects (cool blue 3D floral lettering)
[0072] 4. Help you solve the above series of problems
[0073] Animation (gold dust falling, duration 1.0)
[0074] Keywords (solve), special effects (cool blue 3D floral text)
[0075] Keywords (a series of questions), special effects (cool blue 3D floral letters)
[0076] …………………………
[0077] …………………………】
[0078] Among them, the text "You choose the brand, style, and price" is the first copy sentence (hereinafter referred to as copy sentence 1). The effect element "default text color (#FFFFFF), font (simplified regular script), font size (8)" before copy sentence 1 is a global element for the entire video. The effect elements "music (fashion rhythm charm bloom-Only Time Will Tell), offset (-0.1)" and "filter (transparent), intensity (0.8), offset (-0.1), duration (16.5)" after copy sentence 1 are also global elements for the entire video. And the effect elements "whole sentence, animation (gold powder falling, duration 1.0)" after copy sentence 1, "keyword (brand), special effect (cool blue three-dimensional flower characters)", "keyword (style), special effect (cool blue three-dimensional flower characters)" and "keyword (price), special effect (cool blue three-dimensional flower characters)" are all sentence elements for copy sentence 1.
[0079] For example, take the effect element "Music (Fashion Rhythm Charm Blooms - Only Time Will Tell), offset (-0.1)" as an example. The "Music (Fashion Rhythm Charm Blooms - Only Time Will Tell)" included in the content of the effect element is the packaging effect description information, "Music" is the category corresponding to the packaging effect (i.e., adding a music effect), and "Fashion Rhythm Charm Blooms - Only Time Will Tell" is the effect description for the packaging effect (i.e., the style of the added music, etc.). The "offset (-0.1)" included in the content of the effect element is the packaging effect parameter information, indicating that the time offset corresponding to the packaging effect is -0.1 (for example, indicating that the music packaging effect begins to be added 0.1 seconds after the start of copy sentence 1).
[0080] For another example, take the effect element "Filter (Transparent), Intensity (0.8), Offset (-0.1), Duration (16.5)" as an example, where the "Filter (Transparent)" included in the content of the effect element is the packaging effect description information, "Filter" is the category corresponding to the packaging effect (i.e., adding a filter effect), and "Transparent" is the effect description for the packaging effect (i.e., the added filter effect is a transparent effect). The "Intensity (0.8), Offset (-0.1), Duration (16.5)" included in the content of the effect element is the packaging effect parameter information. Among them, Intensity (0.8) indicates that the effect intensity corresponding to the packaging effect is 0.8 (i.e., the intensity of the added filter is 0.8), Offset (-0.1) indicates that the time offset corresponding to the packaging effect is -0.1, and Duration (16.5) indicates that the duration corresponding to the packaging effect is 16.5s.
[0081] For example, take the effect element "Animation (Gold Dust Falling, Duration 1.0)" as an example. The "Animation (Gold Dust Falling)" included in the content of this effect element is the packaging effect description information, "Animation" is the corresponding category of the packaging effect (i.e., adding an animation effect), and "Gold Dust Falling" is the effect description for this packaging effect (i.e., the added animation effect is the gold dust falling effect). The "Duration 1.0" included in the content of this effect element is the packaging effect parameter information, indicating that the duration of the packaging effect is 1.0s.
[0082] For another example, consider the effect element "Keywords (brand), special effects (cool blue 3D floral text)." The content of this effect element includes "Keywords (brand), special effects (cool blue 3D floral text)" as a description of the packaging effect, and does not include effect parameter information. "Keywords" and "Special Effects" represent the categories of the packaging effect (i.e., using special effects to highlight keywords), and "Brand" and "Cool Blue 3D Floral Text" represent the effect description for this packaging effect (i.e., using the font effect of "Cool Blue 3D Floral Text" to highlight the keyword "Brand").
[0083] In this embodiment, the target large model can be a pre-trained large language model. The target large model can be trained as follows: First, a pre-trained initial large model can be obtained. The initial large model has the ability to generate packaging effect text, which is used to add packaging effect to the video. A sample video and sample text corresponding to the sample video are obtained. The sample video can be a video containing speech, including audio and image components. ASR technology can be used to perform speech recognition on the audio portion of the sample video to obtain the sample text corresponding to the sample video.
[0084] Then, a packaging element in the form of text can be added to the sample copy to obtain a sample packaging effect text. Specifically, an initial text can be generated based on the sample copy, and the initial text can include multiple sample copy sentences and timestamps corresponding to the sample copy sentences. After at least part of the sample copy sentences in the initial text, a packaging element for the sample copy sentence is added to obtain a sample packaging effect text. For example, a packaging element for the sample copy sentence can be manually added directly after part or all of the sample copy sentences to obtain a sample packaging effect text. Alternatively, a packaging effect operation can be added to the sample video through a preset editing program, and the copy after the packaging effect operation is extracted as the sample packaging effect text.
[0085] The initial large model is then used to determine the predicted packaging effect text corresponding to the sample video. The prediction loss is calculated based on the sample packaging effect text and the predicted packaging effect text. To minimize the prediction loss, the model parameters of the initial large model are adjusted to iteratively update the initial large model. After multiple rounds of iterative updates, the target large model is obtained.
[0086] In step 304, a rendering operation is performed on the original video based on the target text to obtain a target video with a packaging effect.
[0087] In this embodiment, the target text can be used to perform a rendering operation on the original video, and the packaging effect can be added to the video through the rendering operation to obtain a target video with the packaging effect.
[0088] This disclosure provides a video processing method that obtains an original video, obtains a text-based video copy based on the original video, and then uses a target macro model to generate a target text for adding a packaging effect to the original video. The target text includes multiple text-based packaging elements. The original video is then rendered based on the target text to obtain a target video with the packaging effect. This method enables packaging editing of existing videos, resulting in better effects after the packaging editing, improving packaging editing efficiency, and enhancing the user experience.
[0089] It should be noted that although the operations of the methods of the embodiments of the present disclosure are described in a specific order in the above embodiments, this does not require or imply that these operations must be performed in this specific order, or that all of the operations shown must be performed to achieve the desired results. On the contrary, the steps depicted in the flowcharts can be performed in a different order. Additionally or alternatively, certain steps can be omitted, multiple steps can be combined into one step, and / or one step can be decomposed into multiple steps.
[0090] Corresponding to the aforementioned video processing method embodiment, the present disclosure also provides an embodiment of a video processing device.
[0091] like Figure 4 As shown, Figure 4 4 is a block diagram of a video processing device according to an exemplary embodiment of the present disclosure. The device may include: a first acquiring unit 401, a second acquiring unit 402, a generating unit 403 and a rendering unit 404.
[0092] The first acquisition unit 401 is configured to acquire the original video.
[0093] The second acquisition unit 402 is configured to acquire a video copy in text form based on the original video.
[0094] The generating unit 403 is configured to generate a target text for adding a packaging effect to the original video based on the original video and the video copy using the target macro model, wherein the target text includes a plurality of packaging elements in text form.
[0095] The rendering unit 404 is configured to perform a rendering operation on the original video based on the target text to obtain a target video with a packaging effect.
[0096] In some embodiments, the target text also includes multiple text sentences and corresponding timestamps for the text sentences. The packaging element includes a global element for the entire video and a sentence element for the text sentences. The global element is located before or after the first text sentence in the target text, and any sentence element is located after the text sentence corresponding to the sentence element in the target text.
[0097] In other implementations, the content of the packaging element includes packaging effect description information, and the packaging effect description information includes a category corresponding to the packaging effect and an effect description for the packaging effect.
[0098] In other embodiments, at least part of the content of the packaging element also includes packaging effect parameter information, which may include one or more of the following: the time offset corresponding to the packaging effect; the duration corresponding to the packaging effect; and the effect intensity corresponding to the packaging effect.
[0099] In other embodiments, the generation unit 403 is configured to: perform word segmentation processing on the video text to obtain a first feature sequence in the text space, perform visual encoding processing on the original video to obtain a second feature sequence in the visual space, use a target adapter to convert the second feature sequence into a third feature sequence in the text space, fuse the first feature sequence and the third feature sequence to obtain a target feature sequence in the text space, input the target feature sequence into the target macro model, and obtain the target text generated by the target macro model.
[0100] In other embodiments, the target large model can be trained by obtaining a pre-trained initial large model capable of generating packaging effect text, which is used to add packaging effect to a video. A sample video and sample text corresponding to the sample video are obtained, and packaging elements in the form of text are added to the sample text to obtain sample packaging effect text. The initial large model is used to determine predicted packaging effect text corresponding to the sample video. A prediction loss is calculated based on the sample packaging effect text and the predicted packaging effect text. The initial large model is updated based on the prediction loss to obtain the target large model.
[0101] In other embodiments, a sample packaging effect text is obtained by adding a text packaging element to the sample copy as follows: an initial text is generated based on the sample copy, the initial text including a plurality of sample copy sentences and timestamps corresponding to the sample copy sentences. After at least some of the sample copy sentences in the initial text, a packaging element specific to the sample copy sentence is added to obtain the sample packaging effect text.
[0102] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the embodiment scheme of the present disclosure. A person of ordinary skill in the art can understand and implement it without paying any creative work.
[0103] Reference below Figure 5 , Figure 5 A schematic block diagram of an electronic device provided for some embodiments of the present disclosure. The electronic device 920 is suitable for implementing the video processing method provided in the embodiments of the present disclosure, for example. The electronic device 920 may be a terminal device, etc., and may be used to implement a client or a server. The electronic device 920 may include but is not limited to mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable electronic devices, etc., and fixed terminals such as digital TVs, desktop computers, smart home devices, etc. It should be noted that, Figure 5 The electronic device 920 shown is only an example and does not limit the functions and scope of use of the embodiments of the present disclosure.
[0104] like Figure 5 As shown, the electronic device 920 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 921, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 922 or a program loaded from a storage device 928 into a random access memory (RAM) 923. Various programs and data required for the operation of the electronic device 920 are also stored in the RAM 923. The processing device 921, the ROM 922, and the RAM 923 are connected to each other via a bus 924. An input / output (I / O) interface 925 is also connected to the bus 924.
[0105] Typically, the following devices may be connected to the I / O interface 925: an input device 926 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 927 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 928 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 929. The communication device 929 may allow the electronic device 920 to communicate with other electronic devices wirelessly or by wire to exchange data. Although Figure 5 The electronic device 920 is shown with various devices, but it should be understood that it is not required to implement or possess all of the devices shown, and the electronic device 920 may instead implement or possess more or fewer devices. Figure 5 Each block shown in the figure may represent one device, or may represent multiple devices as needed.
[0106] According to an embodiment of the present disclosure, the above-mentioned video processing method can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes a program code for executing the above-mentioned video processing method. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 929, or installed from the storage device 928, or installed from the ROM 922. When the computer program is executed by the processing device 921, the functions defined in the video processing method provided by the embodiment of the present disclosure can be implemented.
[0107] An embodiment of the present disclosure further provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed in a computer, the computer is caused to execute the method provided in the present disclosure.
[0108] It should be noted that the computer-readable medium described in the embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, device, or device. In the embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or convey a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code embodied on the computer-readable medium may be conveyed using any suitable medium, including but not limited to wires, optical cables, RF (Radio Frequency), or any suitable combination thereof.
[0109] Computer program code for performing the operations of the disclosed embodiments may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages such as Java, Smalltalk, C++, and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0110] The various embodiments of this disclosure are described in a progressive manner. Similar portions between the various embodiments can be referenced to each other, and each embodiment focuses on the differences between the other embodiments. In particular, the storage medium and computing device embodiments are described briefly because they are generally similar to the method embodiments. For relevant portions, reference can be made to the description of the method embodiments.
[0111] Those skilled in the art will appreciate that, in one or more of the above examples, the functions described in the embodiments of the present disclosure may be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions may be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.
[0112] The specific implementation methods described above further illustrate the purpose, technical solutions, and beneficial effects of the embodiments of the present disclosure. It should be understood that the above description is only a specific implementation method of the embodiments of the present disclosure and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present disclosure shall be included in the scope of protection of the present invention.
Claims
1. A video processing method, the method comprising: Get the original video; Based on the original video, obtaining a video copy in text form; Based on the original video and the video copy, a target text for adding a packaging effect to the original video is generated using a target macro model; the target text includes a plurality of packaging elements in text form; The original video is rendered based on the target text to obtain a target video with a packaging effect.
2. The method according to claim 1, wherein: The target text also includes multiple copy sentences and timestamps corresponding to the copy sentences; the packaging element includes a global element for the entire video and a sentence element for the copy sentences; wherein the global element is located before or after the first copy sentence in the target text, and any sentence element is located after the copy sentence corresponding to the sentence element in the target text.
3. The method according to claim 1, wherein: The content of the packaging element includes packaging effect description information; the packaging effect description information includes a category corresponding to the packaging effect and an effect description for the packaging effect.
4. The method according to claim 3, wherein: At least part of the contents of the packaging elements also include packaging effect parameter information; the packaging effect parameter information includes one or more of the following: The time offset corresponding to the packaging effect; The duration of the packaging effect; The effect strength corresponding to the packaging effect.
5. The method according to claim 1, wherein: The method of generating a target text for increasing the packaging effect of the original video by using a target macro model based on the original video and the video text includes: Performing word segmentation processing on the video text to obtain a first feature sequence in the text space; Performing visual coding processing on the original video to obtain a second feature sequence in a visual space; converting the second feature sequence into a third feature sequence in text space using a target adapter; The first feature sequence and the third feature sequence are merged to obtain a target feature sequence in the text space; The target feature sequence is input into the target macro model to obtain the target text generated by the target macro model.
6. The method according to claim 1, wherein: The target large model is trained in the following way: Acquire a pre-trained initial large model, wherein the initial large model has the ability to generate packaging effect text; the packaging effect text is used to add packaging effect to the video; Obtaining a sample video and a sample copy corresponding to the sample video; Adding packaging elements in the form of text to the sample copy to obtain a sample packaging effect text; Determine the predicted packaging effect text corresponding to the sample video by using the initial large model; Calculating a predicted loss based on the sample packaging effect text and the predicted packaging effect text; The initial large model is updated based on the predicted loss to obtain a target large model.
7. The method according to claim 6, wherein: The step of adding packaging elements in the form of text to the sample copy to obtain a sample packaging effect text includes: Based on the sample copy, an initial text is generated; the initial text includes a plurality of sample copy sentences and timestamps corresponding to the sample copy sentences; After at least part of the sample copy sentences in the initial text, a packaging element for the sample copy sentences is added to obtain a sample packaging effect text.
8. A video processing device, the device comprising: A first acquisition unit, configured to acquire an original video; A second acquisition unit is configured to acquire a video copy in text form based on the original video; A generating unit configured to generate a target text for increasing a packaging effect on the original video using a target macro model based on the original video and the video copy; The target text includes a plurality of packaging elements in text form; The rendering unit is configured to perform a rendering operation on the original video based on the target text to obtain a target video with a packaging effect.
9. A computer-readable storage medium having a computer program stored thereon, which, when executed in a computer, causes the computer to execute the method according to any one of claims 1 to 7.
10. An electronic device comprising a memory and a processor, wherein the memory stores executable code, and when the processor executes the executable code, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Video editing method and device, storage medium and program product
CN120455805A
Image generation method, computing device, computer readable storage medium and computer program product
CN120640064A
Dynamic special effect graph generation method, computing device, readable storage medium and computer program product
CN120956940A