Audio and video editing method and device based on large model, medium, equipment and product

By segmenting and processing audio and video data using large models, the target content is determined, edited, and optimized, solving the problem of inaccurate editing information in existing technologies and achieving high-quality audio and video editing effects.

CN121462831AActive Publication Date: 2026-02-03BEIJING VOLCANO ENGINE TECH CO LTD
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202511973131.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-02-03
Estimated Expiration
2045-12-24

AI Technical Summary

Technical Problem

In existing technologies, when processing original videos directly using trained models, the editing information is not accurate enough, resulting in poor video editing quality.

Method used

By segmenting the raw audio and video data, determining the target content, using a large model to obtain target editing information, and then editing and optimizing according to optimization strategies, including processing video screen text, audio text, and video screen descriptions, high-quality target audio and video are generated.

Benefits of technology

It improves the accuracy and quality of audio and video editing, ensuring the effectiveness of video editing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121462831A_ABST
    Figure CN121462831A_ABST
Patent Text Reader

Abstract

The invention provides an audio and video editing method and device based on a large model, a medium, equipment and a product, and relates to the technical field of audio and video processing, the method comprises the following steps: segmenting original audio and video data to obtain an audio and video data set, the audio and video data set comprising a plurality of audio and video segments; determining a target content corresponding to the audio and video data set, wherein the target content comprises a video picture text, an audio text and a video picture description; according to the target content, obtaining target edited information through a target large model; identifying the audio and video data set to obtain a target optimization strategy; editing the audio and video data set according to the target editing information to obtain a plurality of target editing segments; and performing video optimization and audio optimization on the plurality of target edited fragments according to the target optimization strategy to obtain a target audio and video. Therefore, the accuracy of the obtained target editing information can be improved, and the quality of the obtained target audio and video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to the technical field of audio and video processing, in particular, to an audio and video clipping method and device based on a large model, a medium, equipment and products. BACKGROUND

[0002] Nowadays, there are many clipping videos of TV series, movies, short dramas and live broadcasts on short video platforms. For video clipping, in related technologies, in order to improve the clipping efficiency, a completed model is directly used to process the original video to obtain clipping information, and the clipping video is directly obtained by clipping the clipping information. However, due to the complexity of the original video, directly processing the original video by using the completed model may lead to inaccurate clipping information, and thus the quality of the clipping video obtained by clipping the clipping information is not good enough. SUMMARY

[0003] This summary is provided to introduce a selection of concepts that are further described below in the detailed description. This summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0004] In a first aspect, the present disclosure provides an audio and video clipping method based on a large model, comprising: splitting original audio and video data to obtain an audio and video data set, the audio and video data set comprising a plurality of audio and video segments; determining target content corresponding to the audio and video data set, the target content comprising video picture text, audio text and video picture description; obtaining target clipping information by a target large model according to the target content; identifying the audio and video data set to obtain a target optimization strategy; clipping the audio and video data set according to the target clipping information to obtain a plurality of target clipping segments; performing video optimization and audio optimization on the plurality of target clipping segments according to the target optimization strategy to obtain a target audio and video.

[0005] In a second aspect, the present disclosure provides an audio and video clipping device based on a large model, comprising: a splitting module configured to split original audio and video data to obtain an audio and video data set, the audio and video data set comprising a plurality of audio and video segments; a determining module configured to determine target content corresponding to the audio and video data set, the target content comprising video picture text, audio text and video picture description; An obtaining module is configured to obtain target clipping information according to the target content through a target large model; An identifying module is configured to identify the audio-video data set to obtain a target optimization strategy; A clipping module is configured to clip the audio-video data set according to the target clipping information to obtain a plurality of target clip segments; An optimization module is configured to perform video optimization and audio optimization on the plurality of target clip segments according to the target optimization strategy to obtain a target audio-video.

[0006] In a third aspect, the present disclosure provides a computer readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method of the first aspect.

[0007] In a fourth aspect, the present disclosure provides an electronic device, comprising: a storage device having a computer program stored thereon; a processing device configured to execute the computer program in the storage device to implement the steps of the method of the first aspect.

[0008] In a fifth aspect, the present disclosure provides a computer program product comprising a computer program, which, when executed by a processor, implements the steps of the method of the first aspect.

[0009] Based on the above technical solution, the original audio-video data is first divided to obtain an audio-video data set, so that a more accurate target content corresponding to the audio-video data set can be obtained, the target content including video picture text, audio text and video picture description, and then the target content is processed through a target large model to obtain more accurate target clipping information, and then the audio-video data set is clipped through the target clipping information to obtain more accurate target clip segments, and then the target clip segments are video-optimized and audio-optimized through the target optimization strategy corresponding to the audio-video data set, so that accurate and high-quality target audio-video can be obtained.

[0010] Other features and advantages of the present disclosure will be described in detail in the following detailed description. BRIEF DESCRIPTION OF DRAWINGS

[0011] The above and other features, advantages, and aspects of embodiments of the present disclosure will become more apparent by describing in detail some embodiments with reference to the attached drawings. The same or similar reference numerals in the drawings denote the same or similar elements. It should be understood that the drawings are schematic, and the original and elements are not necessarily drawn according to the scale. In the drawings: Figure 1 is a schematic diagram of an application scenario of a large model-based audio-video clipping method according to some embodiments.

[0012] Figure 2 is a flowchart of a large model-based audio-video clip method according to some embodiments.

[0013] Figure 3 is a system architecture diagram of a large model-based audio-video clip method according to some embodiments.

[0014] Figure 4 is a generation illustration diagram of a large model-based audio-video clip method according to some embodiments.

[0015] Figure 5 is a structure diagram of a large model-based audio-video clip device according to some embodiments.

[0016] Figure 6 is a diagram of an electronic device according to some embodiments. DETAILED DESCRIPTION

[0017] Embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. While certain embodiments of the present disclosure will be illustrated and described, it is to be understood that the present disclosure is not limited to the embodiments thereof. Rather, the present disclosure is capable of many embodiments, and various omissions, substitutions and changes in the form of the methods and systems disclosed herein can be made without departing from the spirit of the present disclosure. It is understood that the drawings and embodiments are only for illustrative purposes and not intended to limit the scope of protection of the present disclosure.

[0018] It should be understood that each of the steps in the method embodiments of the present disclosure can be performed in a different order and / or in parallel. In addition, the method embodiments can include additional steps and / or omit the steps shown. The scope of the present disclosure is not limited in this respect.

[0019] The term “comprising” and variations thereof as used herein are used inclusively, i.e., “comprising but not limited to.” The term “based on” means “based at least in part on.” The term “one embodiment” means “at least one embodiment”; the term “another embodiment” means “at least one additional embodiment”; the term “some embodiments” means “at least some embodiments.” Related definitions are given below.

[0020] It should be noted that the terms “first”, “second”, and the like used in the present disclosure are merely used to distinguish different devices, modules or units, and do not imply the order or interdependence of the functions performed by these devices, modules or units.

[0021] It should be noted that the modification of "one", "multiple" mentioned in the present disclosure is illustrative but not restrictive, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".

[0022] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not used to limit the scope of the messages or information.

[0023] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type, use range, use scenario, etc. of the personal information involved in the present disclosure should be informed to the user and the consent of the user should be obtained in a proper manner according to relevant laws and regulations.

[0024] For example, in response to receiving the active request of the user, the prompt information is sent to the user to explicitly prompt the user that the operation requested to be performed will require obtaining and using the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the software or hardware such as electronic device, application program, server or storage medium, etc. that performs the operation of the technical solutions of the present disclosure according to the prompt information.

[0025] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the prompt information can be sent to the user in the form of a pop-up window, and the prompt information can be presented in the form of text in the pop-up window. In addition, the pop-up window can also carry selection controls for the user to select "agree" or "disagree" to provide the personal information to the electronic device.

[0026] It can be understood that the above notification and obtaining of user consent process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners meeting the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.

[0027] At the same time, it can be understood that the data (including but not limited to the data itself, the acquisition or use of the data) involved in the present technical solution should comply with the requirements of the relevant laws and regulations and relevant provisions.

[0028] Figure 1Fig. 1 is a schematic diagram of an application scenario of a large model-based audio and video clip method according to some embodiments, which can be applied to a video clip system, and can be used for live audio and video clip and on-demand audio and video clip. The video clip system can include a user terminal, a short video platform, a clip gateway, an order service, a clip hardware cluster, a callback service, a message queue, a result query, and a result storage. The clip gateway can be used to respond to clip requests, such as creation, destruction, task query traversal, and task detail query; can also be used for clip on-demand task queue maintenance and live task reservation start; can also be used for clip task life cycle management and task start decision cluster and resources; and can also be used for clip task message interaction in the clip hardware cluster, including heartbeat, task callback, and task end message. The order service is a basic component service, which is used for basic scheduling of clip tasks, including cluster, machine, and resource dimension scheduling. The message queue is used to cache processing result information and processing failure information for user terminal consumption. The callback service is used to receive processing result information and processing failure information, and uniformly call the information to the user terminal. The result query is a strongly consistent storage component, which is used to store processing result information and processing failure information. When the user terminal queries the task details through the gateway query interface, the processing result information and processing failure information of the task are queried from the result query. The result storage is used to store the final generated target audio and video.

[0029] Figure 2 Fig. 2 is a flowchart of a large model-based audio and video clip method according to some embodiments, Figure 3 Fig. 3 is a schematic diagram of a system architecture of a large model-based audio and video clip method according to some embodiments, as shown in Figure 2 and Figure 3 The present disclosure provides a large model-based audio and video clip method, which can be specifically performed by a large model-based audio and video clip device. The device can be implemented by software and / or hardware. The method can include the following steps.

[0030] In step S210, the original audio and video data is divided to obtain an audio and video data set, which includes a plurality of audio and video segments.

[0031] In this embodiment, the original audio and video data can be live data or on-demand data, that is, the live stream can be pulled to obtain the original audio and video data, or the on-demand file can be downloaded to obtain the original audio and video data.

[0032] The splitting of the original audio-video data can include at least one of coarse splitting and fine splitting. For example, the original audio-video data can be coarsely split to obtain an audio-video data set. The original audio-video data can also be directly fine split to obtain an audio-video data set. The original audio-video data can also be coarsely split to obtain coarse split segments, and then the coarse split data set is fine split to obtain an audio-video data set. The audio-video data set can include a plurality of audio-video segments obtained by splitting, and each audio-video segment can correspond to a video segment and an audio segment.

[0033] For coarse splitting, the original audio-video data can be sequentially split into coarse split segments with a preset time length. For example, the preset time length can be 3-5 seconds. For fine splitting, the original audio-video data or each coarse split segment after coarse splitting can be split. The original audio-video data or each coarse split segment after coarse splitting can be fine split by a shot algorithm to obtain an audio-video data set.

[0034] In step S220, the target content corresponding to the audio-video data set is determined. The target content includes video picture text, audio text, and video picture description.

[0035] In this embodiment, the video picture of the audio-video data set can be subjected to text recognition to obtain video picture text. The audio of the audio-video data set can be subjected to text conversion to obtain audio text. The audio-video data set can be subjected to frame extraction analysis to obtain video picture description.

[0036] For example, the audio-video data set can be processed by various models that have been pre-trained to obtain the target content corresponding to the audio-video data set. The audio-video data set can be processed by an optical character recognition model to obtain video picture text. The optical character recognition model can be an OCR (Optical Character Recognition) model. The target audio in the audio-video data set can be subjected to sound-to-text processing by an ASR (Automatic Speech Recognition) model to obtain audio text. The video in the audio-video data set can be subjected to frame extraction analysis by a VLM (Vision-Language Model) to obtain video picture description.

[0037] In step S230, target clip information is obtained by a target large model according to the target content.

[0038] In the embodiment, the target large model can be a large language model (LLM), and the target content can be input into the target large model to obtain target clip information, which can include a timestamp of a target segment. For commentary type audio and video data, the target clip information can further include commentary content. Optionally, the target clip information can be used for clipping highlight audio and video, and the target clip information can further include a highlight score, a confidence, a highlight description, corresponding source material, highlight highlight information, and subtitle position information, etc. Optionally, the target large model can output target clip information corresponding to a target business scenario according to an audio and video data set corresponding to the target business scenario. The target business scenario can include scenarios such as short plays, e-commerce, and football.

[0039] In step S240, the audio and video data set is identified to obtain a target optimization strategy.

[0040] In the embodiment, the audio and video data set can be identified, which can include content identification and original sound loudness identification, to obtain a corresponding target optimization strategy. The target optimization strategy can include capability arrangement decisions and audio and video parameter decisions. The capability arrangement decisions can be used to represent whether video subtitle synthesis is needed, whether subtitle erasure is needed, whether commentary subtitle synthesis is needed, whether commentary audio synthesis is needed, whether background music generation is needed, etc. The audio and video parameter decisions can include video parameter decisions and audio parameter decisions. The video parameter decisions can include at least one of a resolution, a frame rate, a time reference, a pixel format, a transition effect, a picture rotation, and watermark addition corresponding to a video frame of a video. The audio parameter decisions can include at least one of a sampling rate, a sampling number, a channel layout, a volume, and an audio format corresponding to an audio frame of an audio.

[0041] In step S250, the audio and video data set is clipped according to the target clip information to obtain a plurality of target clip segments.

[0042] In the embodiment, after the target clip information is obtained, each audio and video segment in the audio and video data set can be clipped according to the target clip information to obtain a plurality of target clip segments. The target clip information can include a timestamp of a target segment, and the audio and video data set can be clipped according to the timestamp of the target segment to obtain a plurality of target clip segments.

[0043] In step S260, the plurality of target clip segments are video optimized and audio optimized according to the target optimization strategy to obtain a target audio and video.

[0044] In the embodiment, the plurality of target clip segments obtained by clipping can be video optimized and audio optimized according to the determined target optimization strategy corresponding to the current audio and video data set, so as to obtain a target audio and video with higher quality.

[0045] In the embodiment, the original audio-video data is first divided to obtain an audio-video data set, so that more accurate target content corresponding to the audio-video data set can be obtained, the target content including video picture text, audio text and video picture description, and then the target content is processed by a target large model to obtain more accurate target clip information, and the audio-video data set is clipped by the target clip information to obtain more accurate target clip segment, and the target clip segment is subjected to video optimization and audio optimization by a target optimization strategy corresponding to the audio-video data set, so that accurate and high-quality target audio-video can be obtained.

[0046] In some possible embodiments, determining the target content corresponding to the audio-video data set can include: obtaining historical cache content, the historical cache content being obtained by processing the audio-video data set in a historical case, the historical cache content including at least one of video picture text, audio text and video picture description; and determining the target content corresponding to the audio-video data set according to the historical cache content.

[0047] In the embodiment, for processing of the audio-video data set, an intermediate result obtained in processing the audio-video data set to obtain the target content can be stored, for example, the intermediate result can be cached locally and in the cloud. In a case where the processing of the audio-video data set to obtain the target content fails, that is, complete target content is not obtained, the audio-video data set can be processed again to obtain the target content, and in a case where the processing of the audio-video data set to obtain the target content fails, a preset number of retries can be performed. Then, the intermediate result corresponding to the audio-video data set, that is, the historical cache content, can be obtained, and the target content corresponding to the audio-video data set can be obtained according to the historical cache content, so as to avoid repeated processing of the processed data in the audio-video data set, thereby reducing the data processing amount and improving the data processing efficiency. In addition, in a case where a historical task uses data partially or completely same as the audio-video data set and obtains partial or complete target content and is cached, the historical cache content can be directly obtained, so as to avoid repeated processing of the processed data in the audio-video data set, thereby reducing the data processing amount and improving the data processing efficiency.

[0048] The historical cache content includes at least one of video picture text, audio text and video picture description. The method for determining the historical cache content can refer to the method for determining the target content corresponding to the audio-video data set, which will not be described herein again.

[0049] In a possible embodiment, the target content corresponding to the audio-video data set is determined according to the historical cache content, including: According to the historical cache content, determine the unprocessed data in the audio-video data set; process the unprocessed data to obtain the current content; and determine the historical cache content and the current content as the target content corresponding to the audio-video data set.

[0050] In the embodiment, the processed content in the audio-video data set can be determined according to the historical cache content, and then the unprocessed data in the audio-video data set can be determined, so that only the unprocessed data needs to be processed to obtain the corresponding current content, and then the current content and the historical cache content can be determined as the target content corresponding to the audio-video data set, so as to avoid repeated processing of the processed data in the audio-video data set, thereby reducing the data processing amount and improving the data processing efficiency.

[0051] The current content can include at least one of video picture text, audio text and video picture description, and the determination method of the current content can refer to the method of determining the target content corresponding to the audio-video data set, which will not be described here.

[0052] In a possible implementation, the target clip information is obtained from the target large model according to the target content, including: The original clip information is obtained from the target large model according to the target content, the original clip segment is determined according to the original clip information, and the target clip information is obtained by post-processing the original clip segment, the post-processing including at least one of segment merging and output verification.

[0053] In the embodiment, the original clip information output by the target large model can also be post-processed. The original clip segment can be determined based on the original clip information, and the original clip segment can be merged and output verified, so that more accurate and high-quality target clip information can be obtained, and higher-quality target audio-video can be obtained.

[0054] In a possible implementation, the post-processing includes segment merging.

[0055] The original clip segment is post-processed to obtain the target clip information, including: The adjacent original clip segments with a time stamp interval less than a first time length are merged to obtain a first clip segment, and the target clip information is obtained according to the first clip segment.

[0056] In the embodiment, the original clip segments can be merged, the start timestamp and the end timestamp of each original clip segment can be determined, the timestamp interval of adjacent original clip segments can be further determined, the adjacent original clip segments with a timestamp interval less than the first time length can be merged to obtain the first clip segments, and more accurate target clip information can be obtained, so as to reduce the amount of clip, improve the clip efficiency, and improve the quality of the target audio and video obtained by clip.

[0057] In a possible implementation, the post-processing includes output verification.

[0058] The original clip segments are post-processed to obtain target clip information, including: The original clip segments with a time length less than the second time length are deleted to obtain second clip segments, and the target clip information is obtained according to the second clip segments.

[0059] In the embodiment, the original clip segments can be output verified, the start timestamp and the end timestamp of each original clip segment can be determined, the time length of each original clip segment can be further determined, the original clip segments with a time length less than the second time length can be deleted to obtain second clip segments, and more accurate target clip information can be obtained according to the second clip segments, so as to reduce the amount of clip, improve the clip efficiency, and improve the quality of the target audio and video obtained by clip.

[0060] In some possible implementations, the original clip segments can be merged and output verified at the same time, the adjacent original clip segments with a timestamp interval less than the first time length can be merged to obtain first clip segments, the first clip segments with a time length less than the second time length can be deleted to obtain second clip segments, and the target clip information is obtained according to the second clip segments.

[0061] In some possible implementations, in a case where the target optimization strategy represents that video caption synthesis needs to be performed, the target audio and video is obtained by performing video optimization and audio optimization on the multiple target clip segments according to the target optimization strategy, including: The multiple target clip segments are subjected to video format conversion, video caption synthesis, and audio format conversion to obtain multiple first optimized clip segments, and the target audio and video is obtained by splicing the multiple first optimized clip segments.

[0062] In the embodiment, for the live broadcast scene such as e-commerce live broadcast, the target clip segments obtained by clip can be subjected to video caption synthesis, so that the target audio and video obtained by clip can enable the user to more clearly know the content spoken by the host. At this time, the target optimization strategy can represent that video caption synthesis needs to be performed.

[0063] Each target clip segment can include a video segment and an audio segment, the video segment can be subjected to video format conversion and video subtitle synthesis, and the audio segment can be subjected to audio format conversion. The video format conversion of the video segment can be performed according to a video parameter decision, the video parameter decision can include at least one of resolution, frame rate, time reference, pixel format, transition effect, picture rotation, and watermark adding, the audio format conversion of the audio segment can be performed through an audio parameter decision, the audio parameter decision can include at least one of sampling rate, sampling number, channel layout, volume, and audio format of an audio frame of the audio, the target audio in the audio segment can be converted into text, and the text can be rendered into subtitles, so as to complete the video subtitle synthesis of the video segment. The target audio can be the voice of a target user. Through the above optimization method, each target clip segment can be optimized to obtain a plurality of first optimized clip segments, and the plurality of first optimized clip segments can be spliced in time sequence, so as to obtain a target audio video with high quality and including video subtitles.

[0064] In a possible implementation, in a case where the target optimization strategy indicates that background music generation is required, the target audio video is obtained by performing video optimization and audio optimization on the plurality of target clip segments according to the target optimization strategy, and includes: The video format conversion and the audio format conversion are performed on the plurality of target clip segments to obtain a plurality of second optimized clip segments, the plurality of second optimized clip segments are spliced to obtain a candidate clip audio video, the background music corresponding to the candidate clip audio video is generated according to the feature information of the candidate clip audio video, and the target audio video is obtained by mixing the background music and the original sound of the candidate clip audio video.

[0065] In the embodiment, in order to improve the quality of the target audio-video, background music can also be generated. Each target clip segment can include a video segment and an audio segment, the video segment can be converted in video format, and the audio segment can be converted in audio format. The video segment can be converted in video format according to a video parameter decision, the video parameter decision can include at least one of resolution, frame rate, time reference, pixel format, transition effect, picture rotation, and watermark adding, the audio segment can be converted in audio format according to an audio parameter decision, the audio parameter decision can include at least one of sampling rate, sampling number, channel layout, volume, and audio format of an audio frame of the audio. Through the above method, a plurality of second optimized clip segments can be obtained, and the plurality of second optimized clip segments can be spliced in the arrangement order of the original audio-video frames to obtain a candidate clip audio-video, wherein the audio frame of the audio gap segment can be filled. According to the characteristic information of the candidate clip audio-video, the background music corresponding to the candidate clip audio-video can be determined from the background music library. The characteristic information can be video style, time length, and video belonging scene, so that the corresponding background music can be matched from the background music library based on the characteristic information. The original sound of the candidate clip audio-video and the background music can be mixed to obtain the target audio-video. The background music file can be pulled to obtain an audio packet, and the background audio frame can be decoded to obtain the background audio frame. The original sound and the background audio frame can be set in loudness, resampled, set in sampling rate and sampling format, and mixed. Through the generation of the background music, the quality of the obtained target audio-video can be improved.

[0066] In a possible implementation, video caption generation and background music generation can be performed on a plurality of target clip segments at the same time. The generation method can refer to the above embodiments, which will not be described herein.

[0067] In a possible implementation, the target clip information includes a timestamp of a target segment.

[0068] According to the target clip information, the audio-video data set is clipped to obtain a plurality of target clip segments, including: According to the timestamp of the target segment, the audio-video data set is clipped to obtain a plurality of target clip segments.

[0069] In the embodiment, the audio-video data set can be decoded, and the audio-video data set can be finely filtered according to the timestamp of the target segment to obtain a plurality of target clip segments. Each audio-video segment in the audio-video data set can be video decoded to obtain a decoded video frame, and each audio-video segment in the audio-video data set can be audio decoded to obtain a decoded audio frame. The decoded video frame and the decoded audio frame can be finely filtered according to the timestamp of the target segment to obtain more accurate target clip segments.

[0070] Figure 4 is a generated diagram illustrating a large model-based audio-video clip method according to some embodiments, as Figure 4 As shown in a possible implementation, the target clip information further includes the explanation content.

[0071] In the case where the target optimization strategy representation needs to be generated, the target audio-video is obtained by performing video optimization and audio optimization on the plurality of target clip segments according to the target optimization strategy, including: The video format conversion and the audio format conversion are performed on the plurality of target clip segments to obtain a plurality of second optimization clip segments; the plurality of second optimization clip segments are spliced to obtain a candidate clip audio-video; and the explanation audio and the explanation subtitle are added to the candidate clip audio-video according to the explanation content to obtain the target audio-video.

[0072] In the present implementation, for the steps of performing the video format conversion and the audio format conversion on the plurality of target clip segments to obtain the plurality of second optimization clip segments, and splicing the plurality of second optimization clip segments to obtain the candidate clip audio-video, reference can be made to the above embodiments, which will not be described here again. In the case where the target optimization strategy representation needs to be generated, after the candidate clip video is obtained, the explanation audio and the explanation subtitle can be added to the candidate clip audio-video according to the explanation content in the target clip information, so as to obtain the target audio-video including the explanation audio and the explanation subtitle, so as to improve the quality of the target audio-video and reduce the manual clip workload of the user.

[0073] Continuing to refer to Figure 4 In a possible implementation, the explanation audio and the explanation subtitle are added to the candidate clip audio-video according to the explanation content to obtain the target audio-video, including: determining background music corresponding to the candidate clip audio-video; generating the explanation audio and the explanation subtitle according to the explanation content; obtaining target audio according to the original sound of the candidate clip audio-video, the explanation audio, and the background music; obtaining target video according to the original subtitle of the candidate clip audio-video and the explanation subtitle; and obtaining the target audio-video according to the target video and the target audio.

[0074] In the present implementation, the background music corresponding to the candidate clip audio-video can be the original background music of the candidate clip audio-video, or the background music generated automatically through the above embodiments. The explanation audio and the explanation subtitle can be generated according to the explanation content, so as to mix the original sound of the candidate clip audio-video, the explanation audio, and the background music to obtain the target audio, and mix the original subtitle of the candidate clip audio-video and the explanation subtitle to obtain the target video. Then, the target video and the target audio are aligned to determine the target audio-video.

[0075] Continuing to refer to Figure 4In a possible implementation, the target audio is obtained according to the original sound, the commentary audio and the background music of the candidate cut audio-video, and includes the following steps: The commentary audio and the background music are merged to obtain merged audio, and the original sound of the candidate cut audio-video is inserted into time periods in which the merged audio does not exist in the candidate cut audio-video to obtain the target audio.

[0076] In this embodiment, the priority of the commentary audio and the background music is higher than that of the original sound, and the commentary audio and the background music can be merged first. Alternatively, the commentary audio packet can be obtained by pulling the audio packet from the commentary audio file, and the commentary audio frame is obtained by decoding, and the background music packet is obtained by pulling the audio packet from the background music file, and the background audio frame is obtained by decoding; the commentary audio frame and the background audio frame are subjected to loudness setting, sampling rate and sampling format setting of resampling, and are merged to obtain merged audio; and the merged audio and the original sound are alternately inserted, that is, the original sound of the candidate cut audio-video is inserted into time periods in which the merged audio does not exist in the candidate cut audio-video, and the original sound is deleted in time periods in which the merged audio exists in the candidate cut audio-video, so that the target audio can be obtained. Thus, the user can preferentially hear the commentary audio and the background music, so as to improve the quality of the target audio-video.

[0077] With reference to the foregoing Figure 4 In a possible implementation, the target video is obtained according to the original subtitle and the commentary subtitle of the candidate cut audio-video, and includes the following steps: The original subtitle corresponding to a time period in which the original subtitle of the candidate cut audio-video coincides with the commentary subtitle is replaced by the commentary subtitle to obtain the target video.

[0078] In this embodiment, the priority of the commentary subtitle is higher than that of the original subtitle, and the original subtitle corresponding to a time period in which the original subtitle of the candidate cut audio-video coincides with the commentary subtitle is replaced by the commentary subtitle to obtain the target video. The original subtitle of the time period in which the original subtitle coincides with the commentary subtitle is erased based on the position of the original subtitle, and the corresponding commentary subtitle is rendered to the corresponding position, so that the mixed subtitle is obtained, and the target video including the commentary subtitle is obtained.

[0079] In a possible implementation, after the target audio-video is obtained, the target audio-video can be uploaded to a storage system, and a storage index of the target audio-video is obtained. For the case that the target audio-video is a highlight video, a processing result or a processing failure information can be obtained based on the target audio-video. The processing result can include the storage index, a highlight time point, a highlight score, a highlight classification, description information and the like, and the failure information can be reported.

[0080] Figure 5 FIG. 1 is a structural schematic diagram of a large model-based audio-video clipping device according to some embodiments. As shown in FIG. 1, the large model-based audio-video clipping device includes a large model-based audio-video clipping unit 100, a storage system 200 and a server 300. Figure 5As shown, the embodiment of the present disclosure provides a large model-based audio and video clipping device 500, which comprises: A segmentation module 501 configured to segment original audio and video data to obtain an audio and video data set, wherein the audio and video data set comprises a plurality of audio and video clips; A determination module 502 configured to determine target content corresponding to the audio and video data set, wherein the target content comprises video picture text, audio text and video picture description; An obtaining module 503 configured to obtain target clipping information through a target large model according to the target content; An identification module 504 configured to identify the audio and video data set to obtain a target optimization strategy; A clipping module 505 configured to clip the audio and video data set according to the target clipping information to obtain a plurality of target clipping clips; An optimization module 506 configured to perform video optimization and audio optimization on the plurality of target clipping clips according to the target optimization strategy to obtain target audio and video.

[0081] In some possible implementation manners, the determination module 502 is configured to: perform text recognition on a video picture of the audio and video data set to obtain the video picture text; perform text conversion on audio of the audio and video data set to obtain the audio text; perform frame extraction analysis on the audio and video data set to obtain the video picture description.

[0082] In some possible implementation manners, the determination module 502 is configured to: obtain historical cache content, wherein the historical cache content is obtained by processing the audio and video data set in a historical case, and the historical cache content comprises at least one of video picture text, audio text and video picture description; determine the target content corresponding to the audio and video data set according to the historical cache content.

[0083] In some possible implementation manners, the determination module 502 is configured to: determine unprocessed data in the audio and video data set according to the historical cache content; perform processing on the unprocessed data to obtain current content; determine the historical cache content and the current content as the target content corresponding to the audio and video data set.

[0084] In some possible implementation manners, the obtaining module 503 is configured to: obtain original clip information according to the target content and the target large model; determine original clip segments according to the original clip information; perform post-processing on the original clip segments to obtain the target clip information, the post-processing including at least one of segment merging and output checking.

[0085] In some possible implementation manners, the post-processing includes segment merging. The obtaining module 503 is configured to: merge adjacent original clip segments with a time stamp interval less than a first time length to obtain first clip segments; obtain the target clip information according to the first clip segments.

[0086] In some possible implementation manners, the post-processing includes output checking. The obtaining module 503 is configured to: delete original clip segments with a time length less than a second time length to obtain second clip segments; obtain the target clip information according to the second clip segments.

[0087] In some possible implementation manners, when the target optimization strategy indicates that video subtitle synthesis needs to be performed, the optimization module 506 is configured to: perform video format conversion, video subtitle synthesis, and audio format conversion on the plurality of target clip segments to obtain a plurality of first optimized clip segments; splice the plurality of first optimized clip segments to obtain the target audio video.

[0088] In some possible implementation manners, when the target optimization strategy indicates that background music generation needs to be performed, the optimization module 506 is configured to: perform video format conversion and audio format conversion on the plurality of target clip segments to obtain a plurality of second optimized clip segments; splice the plurality of second optimized clip segments to obtain a candidate clip audio video; generate background music corresponding to the candidate clip audio video according to feature information of the candidate clip audio video; mix the background music and original sound of the candidate clip audio video to obtain the target audio video.

[0089] In some possible implementation manners, the target clip information includes a time stamp of a target segment. The clipping module 505 is configured to: clip the audio-video data set according to the timestamps of the target segments to obtain a plurality of target clip segments.

[0090] In some possible implementation manners, the target clip information further includes commentary content. In a case where the target optimization strategy indicates that commentary generation needs to be performed, the optimization module 506 is configured to: perform video format conversion and audio format conversion on the plurality of target clip segments to obtain a plurality of second optimization clip segments; splice the plurality of second optimization clip segments to obtain a candidate clip audio-video; add commentary audio and commentary subtitles to the candidate clip audio-video according to the commentary content to obtain the target audio-video.

[0091] In some possible implementation manners, the optimization module 506 is configured to: determine background music corresponding to the candidate clip audio-video; generate commentary audio and commentary subtitles according to the commentary content; obtain target audio according to original sound of the candidate clip audio-video, the commentary audio, and the background music; obtain target video according to original subtitles of the candidate clip audio-video and the commentary subtitles; obtain the target audio-video according to the target video and the target audio.

[0092] In some possible implementation manners, the optimization module 506 is configured to: merge the commentary audio and the background music to obtain merged audio; insert original sound of the candidate clip audio-video into time periods in which the merged audio does not exist in the candidate clip audio-video to obtain the target audio.

[0093] In some possible implementation manners, the optimization module 506 is configured to: replace original subtitles corresponding to time periods in which original subtitles of the candidate clip audio-video coincide with the commentary subtitles with the commentary subtitles to obtain the target video.

[0094] The functions performed by each functional module in the above-described audio-video clipping apparatus 500 based on a large model have been described in detail in the part about the method, and will not be repeated here.

[0095] Reference will be made to the accompanying drawings below. Figure 6The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as head-mounted devices, mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0096] like Figure 6 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0097] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0098] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.

[0099] It is noted that the aforementioned computer-readable medium of the present disclosure can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example and without limitation, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium can include a computer-readable program code transmitted by a computer-readable medium or a carrier wave in a baseband or as part of a carrier wave. Such a propagated computer-readable signal medium can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the foregoing. The computer-readable signal medium can also be any computer-readable medium that is not a computer-readable storage medium and that can be used to carry or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), or the like, or any suitable combination of the foregoing.

[0100] In some embodiments, the quality assurance system and the business system can communicate using any currently known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communications (e.g., a communications network) of any form or medium (e.g., wireline, wireless, etc.). Examples of communications networks include local area networks ("LANs"), wide area networks ("WANs"), internetworks (e.g., the Internet), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future developed networks.

[0101] The aforementioned computer-readable medium can be included within the aforementioned electronic device; or can exist separately from the electronic device, without being incorporated in the electronic device.

[0102] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: split original audio and video data to obtain an audio and video data set, the audio and video data set comprising a plurality of audio and video segments; determine a target content corresponding to the audio and video data set, the target content comprising video picture text, audio text, and video picture description; obtain target clip information through a target large model according to the target content; identify the audio and video data set to obtain a target optimization strategy; clip the audio and video data set according to the target clip information to obtain a plurality of target clip segments; and perform video optimization and audio optimization on the plurality of target clip segments according to the target optimization strategy to obtain a target audio and video.

[0103] Computer program code for carrying out operations of the present disclosure can be written in any one or more of a variety of programming languages or combinations of languages, including an object-oriented programming language such as Java, Smalltalk, C++, or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer, or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).

[0104] The flow diagrams and the block diagrams in the drawings are illustrations of architectures, functionalities, and operations of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams or block diagrams can represent a module, a segment, or a portion of code, which comprises one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently or the blocks may

[0105] The modules involved in the embodiments of the present disclosure can be implemented in the form of software, or can be implemented in the form of hardware. In some cases, the name of the module does not constitute a limitation on the module itself.

[0106] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, and without limitation, example types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Application-specific Integrated Circuits (ASICs), Application-specific Standard Products (ASSPs), System-on-a-chip (SOCs), Complex Programmable Logic Devices (CPLDs), etc.

[0107] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of the machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fibers, portable compact disc read-only memories (CD-ROMs), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0108] The above description is merely exemplary of the present disclosure and the application of the principles thereof. It is not intended to limit the disclosed concepts to the precise forms disclosed. Rather, it is intended to cover such departures from the present disclosure as come within the scope of the concepts disclosed herein and the patentable scope of the present disclosure. For example, the features described above and other features disclosed herein (but not limited to) can be interchanged or eliminated in any manner not specifically described above, but which will be apparent within the scope of the disclosure.

[0109] Moreover, while operations can be depicted in a particular, serial order, this should not be understood as requiring or implying that such operations be performed in the order illustrated, or in sequential order, or that all illustrated operations be performed, to achieve desirable results. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, while specific implementation details are contained in the above discussion, these should not be construed as limiting the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination. Accordingly, the particular implementation described herein is illustrative only and not limiting.

[0110] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended definitions is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the definitions appended to this specification. With respect to the devices in the above-described embodiments, in which various modules perform operations, the specific manner in which the various modules perform the operations has been described in detail in the embodiments relating to the method. Here, no detailed explanation will be given.

Claims

1. A method for audio and video editing based on a large model, characterized in that, include: The original audio and video data is segmented to obtain an audio and video dataset, which includes multiple audio and video segments. Determine the target content corresponding to the audio and video dataset, wherein the target content includes video screen text, audio text, and video screen description; Based on the target content, the target clip information is obtained through the target large model; The audio and video dataset is identified to obtain the target optimization strategy; Based on the target clip information, the audio and video dataset is edited to obtain multiple target clip segments; Based on the target optimization strategy, video and audio optimizations are performed on the multiple target clips to obtain target audio and video.

2. The audio and video editing method based on a large model according to claim 1, characterized in that, Determining the target content corresponding to the audio and video dataset includes: Text recognition is performed on the video frames of the audio and video dataset to obtain the text of the video frames; The audio in the audio and video dataset is converted into text to obtain the audio text; Frame-by-frame analysis is performed on the audio and video dataset to obtain a description of the video frame.

3. The audio and video editing method based on a large model according to claim 1, characterized in that, Determining the target content corresponding to the audio and video dataset includes: Obtain historical cached content, which is obtained by processing the audio and video dataset under historical conditions. The historical cached content includes at least one of video screen text, audio text, and video screen description. Based on the historical cached content, determine the target content corresponding to the audio and video dataset.

4. The audio and video editing method based on a large model according to claim 3, characterized in that, The step of determining the target content corresponding to the audio and video dataset based on the historical cached content includes: Based on the historical cached content, determine the unprocessed data in the audio and video dataset; The unprocessed data is processed to obtain the current content; The historical cached content and the current content are determined as the target content corresponding to the audio and video dataset.

5. The audio and video editing method based on a large model according to claim 1, characterized in that, The step of obtaining target clip information based on the target content through the target large model includes: Based on the target content, the original editing information is obtained through the target large model; Based on the original editing information, determine the original editing segment; The original clip is post-processed to obtain the target clip information. The post-processing includes at least one of clip merging and output verification.

6. The audio and video editing method based on a large model according to claim 5, characterized in that, The post-processing includes fragment merging; The post-processing of the original clip to obtain the target clip information includes: Merge adjacent original clips with timestamp intervals less than the first duration to obtain the first clip; The target clip information is obtained based on the first clip.

7. The audio and video editing method based on a large model according to claim 5, characterized in that, The post-processing includes output verification; The post-processing of the original clip to obtain the target clip information includes: Delete the original clip that is shorter than the second clip to obtain the second clip; The target clip information is obtained based on the second clip.

8. The audio and video editing method based on a large model according to claim 1, characterized in that, When the target optimization strategy indicates that video subtitle synthesis is required, the step of performing video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video includes: The multiple target clips are subjected to video format conversion, video subtitle synthesis, and audio format conversion to obtain multiple first optimized clips; The target audio and video are obtained by splicing together the multiple first optimized clips.

9. The audio and video editing method based on a large model according to claim 1, characterized in that, When the target optimization strategy indicates that background music generation is required, the step of performing video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video includes: The multiple target clips are subjected to video and audio format conversion to obtain multiple second optimized clips; The multiple second optimized clips are spliced ​​together to obtain candidate clip audio and video; Based on the feature information of the candidate audio and video clips, background music corresponding to the candidate audio and video clips is generated; The background music and the original sound of the candidate clip audio / video are mixed to obtain the target audio / video.

10. The audio and video editing method based on a large model according to claim 1, characterized in that, The target clip information includes the timestamp of the target segment; The step involves editing the audio and video dataset based on the target clip information to obtain multiple target clip segments, including: Based on the timestamp of the target segment, the audio and video dataset is edited to obtain multiple target clip segments.

11. The audio and video editing method based on a large model according to claim 10, characterized in that, The target clip information also includes narration content; When the target optimization strategy representation requires narration generation, the step of performing video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video includes: The multiple target clips are subjected to video and audio format conversion to obtain multiple second optimized clips; The multiple second optimized clips are spliced ​​together to obtain candidate clip audio and video; Based on the narration content, add narration audio and narration subtitles to the candidate edited audio and video to obtain the target audio and video.

12. The audio and video editing method based on a large model according to claim 11, characterized in that, The step of adding narration audio and narration subtitles to the candidate edited audio and video based on the narration content to obtain the target audio and video includes: Determine the background music corresponding to the candidate audio / video clips; Based on the narration content, generate narration audio and narration subtitles; The target audio is obtained based on the original sound of the candidate edited audio and video, the narration audio, and the background music; The target video is obtained based on the original subtitles and the narration subtitles of the candidate edited audio and video. The target audio and video are obtained based on the target video and the target audio.

13. The audio and video editing method based on a large model according to claim 12, characterized in that, The step of obtaining the target audio based on the original sound of the candidate edited audio / video, the narration audio, and the background music includes: The narration audio and the background music are merged to obtain a merged audio; The original audio of the candidate clip is inserted into the time period during which the merged audio is not present in the candidate clip audio / video to obtain the target audio.

14. The audio and video editing method based on a large model according to claim 12, characterized in that, The step of obtaining the target video based on the original subtitles and the narration subtitles of the candidate edited audio and video includes: The original subtitles corresponding to the time periods when the original subtitles of the candidate edited audio and video overlap with the narration subtitles are replaced with the narration subtitles to obtain the target video.

15. An audio and video editing device based on a large model, characterized in that, include: The segmentation module is configured to segment the original audio and video data to obtain an audio and video dataset, which includes multiple audio and video segments; The determination module is configured to determine the target content corresponding to the audio and video dataset, wherein the target content includes video screen text, audio text, and video screen description; The acquisition module is configured to obtain target clip information based on the target content through the target large model; The recognition module is configured to recognize the audio and video dataset to obtain a target optimization strategy; The editing module is configured to edit the audio and video dataset according to the target editing information to obtain multiple target editing segments; The optimization module is configured to perform video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video.

16. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processing device, it implements the steps of the method according to any one of claims 1-14.

17. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-14.

18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-14.

Citation Information

Patent Citations

  • Video processing method and device, computing equipment and computer storage medium

    CN117319765A

  • Intelligent video editing method and system based on large model

    CN117812386A

  • Video editing method and device based on large model, equipment, medium and product

    CN118338072A

  • Video editing method and device based on large model, equipment and medium

    CN119583897A

  • Video editing method and device

    CN119835500A