Audio and video clip method, device, medium, equipment and product based on large model
By segmenting audio and video data and processing it with a large model, the target content is determined and optimization strategies are implemented, solving the problem of inaccurate editing information in existing technologies and achieving high-quality audio and video editing effects.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-24
- Publication Date
- 2026-04-14
AI Technical Summary
In existing technologies, directly processing the original video using a trained model can easily lead to inaccurate editing information, resulting in poor video editing quality.
By segmenting the raw audio and video data, determining the target content, using a large model to obtain target editing information, and then editing and optimizing according to optimization strategies, including processing video screen text, audio text, and video screen descriptions, high-quality target audio and video are generated.
It improves the accuracy and quality of audio and video editing, ensures the precision of edited segments and the effectiveness of optimization strategies, and generates high-quality target audio and video.
Smart Images

Figure CN121462831B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of audio and video processing technology, and more specifically, to an audio and video editing method, apparatus, medium, device, and product based on a large model. Background Technology
[0002] Currently, short video platforms contain numerous edited videos of TV dramas, movies, short dramas, and live streams. Regarding video editing, to improve efficiency, some technologies directly process the original video using a trained model to obtain editing information, and then edit the video using this information. However, due to the complexity of the original video, directly processing it using a trained model can easily lead to inaccurate editing information, resulting in a lower quality edited video. Summary of the Invention
[0003] This summary section is provided to briefly introduce the concepts, which will be described in detail in the detailed description section below. This summary section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0004] Firstly, this disclosure provides an audio / video editing method based on a large model, including:
[0005] The original audio and video data is segmented to obtain an audio and video dataset, which includes multiple audio and video segments.
[0006] Determine the target content corresponding to the audio and video dataset, wherein the target content includes video screen text, audio text, and video screen description;
[0007] Based on the target content, the target clip information is obtained through the target large model;
[0008] The audio and video dataset is identified to obtain the target optimization strategy;
[0009] Based on the target clip information, the audio and video dataset is edited to obtain multiple target clip segments;
[0010] Based on the target optimization strategy, video and audio optimizations are performed on the multiple target clips to obtain target audio and video.
[0011] Secondly, this disclosure provides an audio and video editing device based on a large model, comprising:
[0012] The segmentation module is configured to segment the original audio and video data to obtain an audio and video dataset, which includes multiple audio and video segments;
[0013] The determination module is configured to determine the target content corresponding to the audio and video dataset, wherein the target content includes video screen text, audio text, and video screen description;
[0014] The acquisition module is configured to obtain target clip information based on the target content through the target large model;
[0015] The recognition module is configured to recognize the audio and video dataset to obtain a target optimization strategy;
[0016] The editing module is configured to edit the audio and video dataset according to the target editing information to obtain multiple target editing segments;
[0017] The optimization module is configured to perform video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video.
[0018] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0019] Fourthly, this disclosure provides an electronic device, comprising:
[0020] A storage device on which computer programs are stored;
[0021] A processing device for executing the computer program in the storage device to implement the steps of the method described in the first aspect.
[0022] Fifthly, this disclosure provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described in the first aspect.
[0023] Based on the above technical solution, the original audio and video data is first segmented to obtain an audio and video dataset, so as to obtain more accurate target content corresponding to the audio and video dataset. The target content includes video screen text, audio text, and video screen description. Then, the target content is processed by the target big model to obtain more accurate target editing information. Then, the audio and video dataset is edited with the target editing information to obtain more accurate target editing segments. Finally, the target optimization strategy corresponding to the audio and video dataset is used to optimize the video and audio of the target editing segments, so as to obtain accurate and high-quality target audio and video.
[0024] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0025] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0026] Figure 1 This is a schematic diagram illustrating an application scenario of a large-model-based audio and video editing method, based on some embodiments.
[0027] Figure 2 This is a flowchart illustrating a large-model-based audio and video editing method according to some embodiments.
[0028] Figure 3 This is a schematic diagram of the system architecture of a large-model-based audio and video editing method, illustrated according to some embodiments.
[0029] Figure 4 This is a diagram illustrating the generation of a large-model-based audio and video editing method, based on some embodiments.
[0030] Figure 5 This is a schematic diagram of the structure of a large-model-based audio and video editing device, as shown in some embodiments.
[0031] Figure 6 This is a schematic diagram of an electronic device according to some embodiments. Detailed Implementation
[0032] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0033] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0034] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0035] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0036] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0037] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0038] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and their consent should be obtained.
[0039] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0040] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0041] It is understood that the above notification and user consent process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0042] Meanwhile, it is understood that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations and related provisions.
[0043] Figure 1This diagram illustrates an application scenario of a large-model-based audio and video editing method, illustrated with several embodiments. It can be applied to video editing systems, including live and on-demand audio and video editing. The video editing system may include a user client, a short video platform, an editing gateway, an order service, an editing hardware cluster, a callback service, a message queue, result querying, and result storage. The editing gateway can respond to editing requests, such as creating, destroying, querying, and detailing tasks; it can also maintain the on-demand editing task queue and schedule live task initiation; it can manage the lifecycle of editing tasks and make cluster and resource decisions for task initiation; and it can interact with editing tasks in the editing hardware cluster, including heartbeats, task callbacks, and task completion messages. The order service is a fundamental component service used for basic scheduling of editing tasks, including scheduling at the cluster, machine, and resource levels. The message queue caches processing results and failure information for user client consumption. The callback service receives processing results and failure information and sends this information back to the user client. The result query is a strongly consistent storage component used to store processing result information and processing failure information. When a user queries task details through the gateway query interface, it will retrieve the processing result information and processing failure information for that task from the result query. The result storage is used to store the final generated target audio and video.
[0044] Figure 2 This is a flowchart illustrating a large-model-based audio and video editing method according to some embodiments. Figure 3 This is a system architecture diagram of a large-model-based audio and video editing method, illustrated according to some embodiments, such as... Figure 2 and Figure 3 As shown, this disclosure provides a large-model-based audio and video editing method, specifically executed by a large-model-based audio and video editing device, which can be implemented in software and / or hardware. The method may include the following steps.
[0045] In step S210, the original audio and video data is segmented to obtain an audio and video dataset, which includes multiple audio and video segments.
[0046] In this embodiment, the original audio and video data can be live streaming data or video-on-demand data. That is, the original audio and video data can be obtained by pulling the live stream or by downloading the video-on-demand file.
[0047] The segmentation of raw audio and video data can include at least one of coarse segmentation and fine segmentation. For example, the raw audio and video data can be coarsely segmented to obtain an audio and video dataset; the raw audio and video data can also be finely segmented directly to obtain an audio and video dataset; or the raw audio and video data can be coarsely segmented first to obtain coarse segments, and then the coarse segments can be finely segmented to obtain an audio and video dataset. This audio and video dataset can include multiple segmented audio and video segments, each of which can correspond to both video and audio segments.
[0048] For coarse cutting, the original audio and video data can be sequentially divided into coarse segments of a preset duration, for example, 3 to 5 seconds. For fine cutting, the original audio and video data or each coarse segment after coarse cutting can be further segmented. By using a storyboard algorithm, the original audio and video data or each coarse segment after coarse cutting can be finely cut to obtain an audio and video dataset.
[0049] In step S220, the target content corresponding to the audio and video dataset is determined. The target content includes video screen text, audio text, and video screen description.
[0050] In this embodiment, text recognition can be performed on the video frames of the audio and video dataset to obtain video frame text; text conversion can be performed on the audio of the audio and video dataset to obtain audio text; and frame extraction analysis can be performed on the audio and video dataset to obtain video frame descriptions.
[0051] For example, various pre-trained models can be used to process audio and video datasets to obtain the target content corresponding to the dataset. Specifically, an Optical Character Recognition (OCR) model can be used to process the audio and video dataset to obtain text from the video images. An Automatic Speech Recognition (ASR) model can be used to convert the target audio in the audio and video dataset into text, obtaining audio text. A Vision-Language Model (VLM) can be used to perform frame extraction analysis on the video in the audio and video dataset to obtain video image descriptions.
[0052] In step S230, target clipping information is obtained by using the target large model based on the target content.
[0053] In this embodiment, the target large model can be an LLM (Large Language Model). The target content can be input into the target large model to obtain target clip information. This target clip information may include the timestamp of the target segment. For narration-type audio and video data, the target clip information may also include the narration content. Optionally, this target clip information can be used to edit highlight audio and video. The target clip information may also include highlight score, confidence level, highlight description, corresponding source material, highlight compilation information, and subtitle position information. Optionally, the target large model can output target clip information corresponding to the target business scenario based on the target business scenario corresponding to the audio and video dataset. This target business scenario may include scenarios such as short dramas, e-commerce, and football.
[0054] In step S240, the audio and video dataset is identified to obtain the target optimization strategy.
[0055] In this embodiment, the audio and video dataset can be identified, including content recognition and loudness recognition of the original sound, to obtain a corresponding target optimization strategy. This target optimization strategy can include capability orchestration decisions and audio / video parameter decisions. Capability orchestration decisions can be used to characterize whether video subtitle synthesis, subtitle erasure, narration subtitle synthesis, narration audio synthesis, and background music generation are needed. Audio and video parameter decisions can include video parameter decisions and audio parameter decisions. Video parameter decisions can include at least one of the following: resolution, frame rate, time base, pixel format, transition effects, screen rotation, and watermarking for the video frame. Audio parameter decisions can include at least one of the following: sampling rate, number of samples, channel layout, volume, and audio format for the audio frame.
[0056] In step S250, the audio and video dataset is edited according to the target clip information to obtain multiple target clip segments.
[0057] In this embodiment, after obtaining the target clip information, the audio and video segments in the audio and video dataset can be edited according to the target clip information to obtain multiple target clip segments. The target clip information may include the timestamps of the target segments; the audio and video dataset can be edited according to the timestamps of the target segments to obtain multiple target clip segments.
[0058] In step S260, according to the target optimization strategy, video and audio optimization are performed on multiple target clips to obtain target audio and video.
[0059] In this embodiment, video and audio optimization can be performed on multiple target clips obtained from editing, based on the determined target optimization strategy corresponding to the current audio and video dataset, in order to obtain higher quality target audio and video.
[0060] In this embodiment, the original audio and video data is first segmented to obtain an audio and video dataset, which enables the acquisition of more accurate target content corresponding to the audio and video dataset. This target content includes video screen text, audio text, and video screen description. Then, the target content is processed using a target big model to obtain more accurate target editing information. The audio and video dataset is then edited using the target editing information to obtain more accurate target editing segments. Finally, the target optimization strategy corresponding to the audio and video dataset is used to optimize the video and audio of the target editing segments, resulting in accurate and high-quality target audio and video.
[0061] In some possible implementations, determining the target content corresponding to the audio and video dataset may include:
[0062] Retrieve historical cached content, which is obtained by processing the audio and video dataset under historical conditions. Historical cached content includes at least one of video frame text, audio text, and video frame description. Based on the historical cached content, determine the target content corresponding to the audio and video dataset.
[0063] In this embodiment, during the processing of the audio and video dataset to obtain the target content, intermediate results can be stored, for example, cached locally or in the cloud. If processing the audio and video dataset to obtain the target content fails (i.e., the complete target content is not obtained), the audio and video dataset can be reprocessed to obtain the target content. A preset number of retries can be performed if processing fails to obtain the target content. Furthermore, the intermediate results corresponding to the audio and video dataset, i.e., the historical cached content, can be obtained, and the target content corresponding to the audio and video dataset can be obtained based on the historical cached content. This avoids duplicate processing of already processed data in the audio and video dataset, thereby reducing the amount of data processing and improving data processing efficiency. Additionally, if a historical task uses data that is partially or completely identical to that in the audio and video dataset and obtains part of the target content and caches it, the historical cached content can be directly obtained to avoid duplicate processing of already processed data in the audio and video dataset, thereby reducing the amount of data processing and improving data processing efficiency.
[0064] The historical cached content includes at least one of the following: video image text, audio text, and video image description. The method for determining the historical cached content can refer to the method described above for determining the target content corresponding to the audio and video dataset, and will not be repeated here.
[0065] In one possible implementation, determining the target content corresponding to the audio / video dataset based on historical cached content includes:
[0066] Based on the historical cached content, identify the unprocessed data in the audio and video dataset; process the unprocessed data to obtain the current content; and determine the historical cached content and the current content as the target content corresponding to the audio and video dataset.
[0067] In this embodiment, the processed content in the audio and video dataset can be determined based on the historical cached content, and then the unprocessed data in the audio and video dataset can be determined. This allows the processing of only the unprocessed data to obtain the corresponding current content. The current content and the historical cached content can then be determined as the target content corresponding to the audio and video dataset, thereby avoiding the repeated processing of the processed data in the audio and video dataset, thus reducing the amount of data processing and improving data processing efficiency.
[0068] The current content may include at least one of the following: video screen text, audio text, and video screen description. The method for determining the current content can refer to the method for determining the target content corresponding to the audio and video dataset described above, and will not be repeated here.
[0069] In one possible implementation, target clip information is obtained based on the target content through a target large model, including:
[0070] Based on the target content, the original editing information is obtained through the target large model; based on the original editing information, the original editing segments are determined; the original editing segments are post-processed to obtain the target editing information, and the post-processing includes at least one of segment merging and output verification.
[0071] In this embodiment, the original edited information output by the target large model can also be post-processed. Specifically, the original edited segments can be determined based on the original edited information, and these segments can be merged and output verified to obtain more accurate and higher-quality target edited information, thereby achieving higher-quality target audio and video.
[0072] In one possible implementation, post-processing includes fragment merging.
[0073] Post-processing of the original clips yields the target clip information, including:
[0074] Merge adjacent original clips with timestamp intervals less than the first duration to obtain the first clip; obtain the target clip information based on the first clip.
[0075] In this embodiment, the original clips can be merged. The start and end timestamps of each original clip can be determined so that the timestamp interval between adjacent original clips can be further determined. Adjacent original clips with timestamp intervals less than a first duration are merged to obtain a first clip, thereby obtaining more accurate target clip information, so as to reduce the amount of editing, improve editing efficiency and improve the quality of the target audio and video obtained by editing.
[0076] In one possible implementation, post-processing includes output verification.
[0077] Post-processing of the original clips yields the target clip information, including:
[0078] Delete the original clip with a duration shorter than the second clip to obtain the second clip; obtain the target clip information based on the second clip.
[0079] In this embodiment, the original clips can be output verified to determine the start and end timestamps of each original clip, so as to further determine the duration of each original clip. This allows the deletion of original clips with a duration shorter than the second duration, resulting in the second clip. More accurate target clip information can then be obtained based on the second clip, thereby reducing the amount of editing, improving editing efficiency, and enhancing the quality of the resulting target audio and video.
[0080] In some possible implementations, the original clip segments can be merged and output verified simultaneously. Adjacent original clip segments with timestamp intervals less than a first duration can be merged to obtain a first clip segment. The first clip segment with a duration less than a second duration can be deleted to obtain a second clip segment. The target clip information can be obtained based on the second clip segment.
[0081] In some possible implementations, when the target optimization strategy characterization requires video subtitle synthesis, multiple target clips are optimized for video and audio according to the target optimization strategy to obtain target audio and video, including:
[0082] Multiple target clips are converted into video formats, video subtitles are synthesized, and audio formats are converted to obtain multiple first-optimized clips; these multiple first-optimized clips are then spliced together to obtain the target audio and video.
[0083] In this embodiment, for live streaming scenarios, such as e-commerce live streaming, video subtitles can be synthesized on the edited target clips to enable users to more clearly understand the content being presented by the streamer. In this case, the target optimization strategy represents the need for video subtitle synthesis.
[0084] Each target clip can include both video and audio segments. Video segments can undergo video format conversion and subtitle synthesis, while audio segments can undergo audio format conversion. Specifically, video format conversion can be performed based on video parameters, including at least one of resolution, frame rate, time base, pixel format, transition effects, screen rotation, and watermarking. Audio format conversion can be performed based on audio parameters, including at least one of the sampling rate, number of samples, channel layout, volume, and audio format of the audio frames. The target audio in the audio segments is then converted into text, and the text is rendered as subtitles, thus completing the video subtitle synthesis for the video segments. The target audio can be the voice of the target user. Using this optimization method, each target clip can be optimized separately to obtain multiple first-optimized clips. These multiple first-optimized clips are then spliced together in time sequence to obtain a high-quality target audio-visual product including video subtitles.
[0085] In one possible implementation, when the target optimization strategy characterization requires background music generation, video and audio optimization are performed on multiple target clips according to the target optimization strategy to obtain target audio and video, including:
[0086] Multiple target clips are converted into video and audio formats to obtain multiple second-optimized clips; these second-optimized clips are then spliced together to obtain candidate audio-visual clips; background music is generated based on the feature information of the candidate audio-visual clips; and the background music and the original sound of the candidate audio-visual clips are mixed to obtain the target audio-visual clips.
[0087] In this embodiment, background music can be generated to improve the quality of the target audio and video. Each target clip can include a video clip and an audio clip. Video clips can be converted to video formats, and audio clips can be converted to audio formats. Video format conversion can be performed on video clips based on video parameter decisions, which may include at least one of resolution, frame rate, time base, pixel format, transition effects, screen rotation, and watermarking. Audio format conversion can be performed on audio clips based on audio parameter decisions, which may include at least one of the sampling rate, number of samples, channel layout, volume, and audio format corresponding to the audio frames. Using this method, multiple second optimized clips can be obtained. These second optimized clips are then concatenated according to the original audio and video frame order to obtain candidate audio and video clips. Audio frames can be padded to fill in any missing audio segments. Finally, based on the feature information of the candidate audio and video clips, the corresponding background music is determined from a background music library. The feature information may include video style, duration, and the scene to which the video belongs, allowing the matching of corresponding background music from the background music library based on this feature information. The background music and the original audio from the candidate clips are then mixed to obtain the target audio and video. This can be done by extracting the audio package from the background music file, decoding it to obtain background audio frames, and then adjusting the loudness, resampling, sampling rate, and sampling format of the original audio and background audio frames before mixing them. Generating background music improves the quality of the resulting target audio and video.
[0088] In one possible implementation, video subtitles and background music can be generated simultaneously for multiple target clips. The generation method can be referred to the above embodiments and will not be repeated here.
[0089] In one possible implementation, the target clip information includes the timestamp of the target segment.
[0090] Based on the target clip information, the audio and video dataset is edited to obtain multiple target clip segments, including:
[0091] Based on the timestamps of the target segments, the audio and video dataset is edited to obtain multiple target clip segments.
[0092] In this embodiment, the audio and video dataset can be decoded, and the dataset can be further refined using the timestamps of the target segments to obtain multiple target clips. Specifically, each audio and video segment in the dataset can be individually decoded to obtain decoded video frames, and each segment can also be individually decoded to obtain decoded audio frames. The timestamps of the target segments can then be used to refine the decoded video and audio frames, resulting in more accurate multiple target clips.
[0093] Figure 4 These are explanatory diagrams illustrating a large-model-based audio and video editing method, as shown in some embodiments. Figure 4 As shown, in one possible implementation, the target clip information also includes narration content.
[0094] When the target optimization strategy representation requires narration generation, video and audio optimization are performed on multiple target clips according to the target optimization strategy to obtain target audio and video, including:
[0095] Multiple target clips are converted into video and audio formats to obtain multiple second-optimized clips; these second-optimized clips are then spliced together to obtain candidate clip audio and video; based on the narration content, narration audio and subtitles are added to the candidate clip audio and video to obtain the target audio and video.
[0096] In this embodiment, the steps of converting the video and audio formats of multiple target clips to obtain multiple second optimized clips, and splicing the multiple second optimized clips to obtain candidate clip audio-visual content, can be referred to the above embodiments and will not be repeated here. When the target optimization strategy characterization requires narration generation, after obtaining the candidate clip video, narration audio and subtitles can be added to the candidate clip audio-visual content according to the narration content in the target clip information, so as to obtain target audio-visual content including narration audio and subtitles, thereby improving the quality of the target audio-visual content and reducing the user's manual editing workload.
[0097] Continue to refer to Figure 4 In one possible implementation, based on the narration content, narration audio and narration subtitles are added to the candidate audio-visual clips to obtain the target audio-visual file, including:
[0098] Determine the background music corresponding to the candidate audio / video clips; generate narration audio and narration subtitles based on the narration content; obtain the target audio based on the original sound, narration audio, and background music of the candidate audio / video clips; obtain the target video based on the original subtitles and narration subtitles of the candidate audio / video clips; obtain the target audio / video based on the target video and target audio.
[0099] In this embodiment, the background music corresponding to the candidate audio / video clip can be the original background music of the candidate audio / video clip, or it can be background music automatically generated through the above embodiments. Furthermore, narration audio and subtitles can be generated based on the narration content, so that the original sound, narration audio, and background music of the candidate audio / video clip can be mixed to obtain the target audio, and the original subtitles and narration subtitles of the candidate audio / video clip can be mixed to obtain the target video. Finally, the target video and target audio are aligned and determined as the target audio / video clip.
[0100] Continue to refer to Figure 4 In one possible implementation, the target audio is obtained based on the original soundtrack, narration audio, and background music of the candidate clip audio / video, including:
[0101] The narration audio and background music are merged to obtain the merged audio; the original audio of the candidate clip is inserted into the time period where the merged audio does not exist in the candidate clip audio and video to obtain the target audio.
[0102] In this embodiment, the narration audio and background music have higher priority than the original audio, and can be merged first. Optionally, an audio packet can be retrieved from the narration audio file to obtain the narration audio packet, and decoded to obtain the narration audio frame; similarly, an audio packet can be retrieved from the background music file to obtain the background music packet, and decoded to obtain the background audio frame. The loudness, resampling rate, and sampling format of the narration audio frame and the background audio frame can be set and merged to obtain the merged audio. Then, the merged audio is interspersed with the original audio, that is, the original audio of the candidate clip audio / video is inserted into the time period when the merged audio / video is not present, and the original audio is deleted into the time period when the merged audio / video is present, thus obtaining the target audio. This allows the user to hear the narration audio and background music first, thereby improving the quality of the target audio / video.
[0103] Continue to refer to Figure 4 In one possible implementation, the target video is obtained based on the original subtitles and the narration subtitles of the candidate clip audio / video, including:
[0104] Replace the original subtitles with the narration subtitles during the time periods when the original subtitles and narration subtitles of the candidate edited audio and video with the narration subtitles to obtain the target video.
[0105] In this embodiment, the narration subtitles have a higher priority than the original subtitles. The original subtitles corresponding to the time periods when the original subtitles and narration subtitles overlap in the candidate edited audio / video can be replaced with the narration subtitles to obtain the target video. Specifically, based on the position of the original subtitles, the original subtitles during the time periods when the original subtitles and narration subtitles overlap can be erased, and the corresponding narration subtitles can be rendered to their corresponding positions, thereby obtaining a mixed subtitle and thus a target video including the narration subtitles.
[0106] In one possible implementation, after obtaining the target audio and video, the target audio and video can be uploaded to a storage system, and the storage index of the target audio and video can be obtained. If the target audio and video is a highlight video, processing results or processing failure information can be obtained based on the target audio and video. The processing results may include the storage index, highlight time points, highlight scores, highlight classifications, descriptive information, etc., and failure information can be reported.
[0107] Figure 5 This is a structural schematic diagram of a large-model-based audio and video editing device, illustrated according to some embodiments. For example... Figure 5 As shown, this disclosure provides a large-model-based audio and video editing device 500, which includes:
[0108] The segmentation module 501 is configured to segment the original audio and video data to obtain an audio and video dataset, wherein the audio and video dataset includes multiple audio and video segments;
[0109] The determination module 502 is configured to determine the target content corresponding to the audio and video dataset, wherein the target content includes video screen text, audio text, and video screen description;
[0110] The module 503 is configured to obtain target clip information based on the target content through the target large model;
[0111] The recognition module 504 is configured to recognize the audio and video dataset to obtain a target optimization strategy;
[0112] The editing module 505 is configured to edit the audio and video dataset according to the target editing information to obtain multiple target editing segments;
[0113] The optimization module 506 is configured to perform video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video.
[0114] In some possible implementations, the determining module 502 is configured to:
[0115] Text recognition is performed on the video frames of the audio and video dataset to obtain the text of the video frames;
[0116] The audio in the audio and video dataset is converted into text to obtain the audio text;
[0117] Frame-by-frame analysis is performed on the audio and video dataset to obtain a description of the video frame.
[0118] In some possible implementations, the determining module 502 is configured to:
[0119] Obtain historical cached content, which is obtained by processing the audio and video dataset under historical conditions. The historical cached content includes at least one of video screen text, audio text, and video screen description.
[0120] Based on the historical cached content, determine the target content corresponding to the audio and video dataset.
[0121] In some possible implementations, the determining module 502 is configured to:
[0122] Based on the historical cached content, determine the unprocessed data in the audio and video dataset;
[0123] The unprocessed data is processed to obtain the current content;
[0124] The historical cached content and the current content are determined as the target content corresponding to the audio and video dataset.
[0125] In some possible implementations, the obtaining module 503 is configured to:
[0126] Based on the target content, the original editing information is obtained through the target large model;
[0127] Based on the original editing information, determine the original editing segment;
[0128] The original clip is post-processed to obtain the target clip information. The post-processing includes at least one of clip merging and output verification.
[0129] In some possible implementations, the post-processing includes fragment merging;
[0130] The obtaining module 503 is configured as follows:
[0131] Merge adjacent original clips with timestamp intervals less than the first duration to obtain the first clip;
[0132] The target clip information is obtained based on the first clip.
[0133] In some possible implementations, the post-processing includes output verification;
[0134] The obtaining module 503 is configured as follows:
[0135] Delete the original clip that is shorter than the second clip to obtain the second clip;
[0136] The target clip information is obtained based on the second clip.
[0137] In some possible implementations, when the target optimization strategy characterization requires video subtitle synthesis, the optimization module 506 is configured to:
[0138] The multiple target clips are subjected to video format conversion, video subtitle synthesis, and audio format conversion to obtain multiple first optimized clips;
[0139] The target audio and video are obtained by splicing together the multiple first optimized clips.
[0140] In some possible implementations, when the target optimization strategy characterization requires background music generation, the optimization module 506 is configured to:
[0141] The multiple target clips are subjected to video and audio format conversion to obtain multiple second optimized clips;
[0142] The multiple second optimized clips are spliced together to obtain candidate clip audio and video;
[0143] Based on the feature information of the candidate audio and video clips, background music corresponding to the candidate audio and video clips is generated;
[0144] The background music and the original sound of the candidate clip audio / video are mixed to obtain the target audio / video.
[0145] In some possible implementations, the target clip information includes a timestamp of the target segment;
[0146] The editing module 505 is configured as follows:
[0147] Based on the timestamp of the target segment, the audio and video dataset is edited to obtain multiple target clip segments.
[0148] In some possible implementations, the target clip information may also include narration content;
[0149] When the target optimization strategy representation requires interpretation generation, the optimization module 506 is configured as follows:
[0150] The multiple target clips are subjected to video and audio format conversion to obtain multiple second optimized clips;
[0151] The multiple second optimized clips are spliced together to obtain candidate clip audio and video;
[0152] Based on the narration content, add narration audio and narration subtitles to the candidate edited audio and video to obtain the target audio and video.
[0153] In some possible implementations, the optimization module 506 is configured to:
[0154] Determine the background music corresponding to the candidate audio / video clips;
[0155] Based on the narration content, generate narration audio and narration subtitles;
[0156] The target audio is obtained based on the original sound of the candidate edited audio and video, the narration audio, and the background music;
[0157] The target video is obtained based on the original subtitles and the narration subtitles of the candidate edited audio and video.
[0158] The target audio and video are obtained based on the target video and the target audio.
[0159] In some possible implementations, the optimization module 506 is configured to:
[0160] The narration audio and the background music are merged to obtain a merged audio;
[0161] The original audio of the candidate clip is inserted into the time period during which the merged audio is not present in the candidate clip audio / video to obtain the target audio.
[0162] In some possible implementations, the optimization module 506 is configured to:
[0163] The original subtitles corresponding to the time periods when the original subtitles of the candidate edited audio and video overlap with the narration subtitles are replaced with the narration subtitles to obtain the target video.
[0164] The functional logic executed by each functional module in the aforementioned large-model-based audio and video editing device 500 has been explained in detail in the section on methods, and will not be repeated here.
[0165] The following is for reference. Figure 6 The diagram illustrates a structural schematic of an electronic device (e.g., a terminal device or a server) 600 suitable for implementing embodiments of the present disclosure. The terminal device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as head-mounted devices, mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 6 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0166] like Figure 6As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0167] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 6 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0168] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0169] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0170] In some implementations, quality assurance systems and business systems can communicate using any currently known or future-developed network protocol, such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communications of any form or medium (e.g., communication networks). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0171] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0172] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: segment the original audio and video data to obtain an audio and video dataset, the audio and video dataset including multiple audio and video segments; determine the target content corresponding to the audio and video dataset, the target content including video screen text, audio text, and video screen description; obtain target editing information based on the target content through a target big model; identify the audio and video dataset to obtain a target optimization strategy; edit the audio and video dataset according to the target editing information to obtain multiple target editing segments; and perform video optimization and audio optimization on the multiple target editing segments according to the target optimization strategy to obtain target audio and video.
[0173] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0174] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0175] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules are not, in some cases, intended to limit the functionality of the module itself.
[0176] The functions described above in this document can be performed at least in part by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), system-on-a-chip (SoCs), complex programmable logic devices (CPLDs), and so on.
[0177] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0178] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0179] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0180] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended content is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the appended content. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A method for audio and video editing based on a large model, characterized in that, include: The original audio and video data is segmented to obtain an audio and video dataset, which includes multiple audio and video segments. The segmentation of the original audio and video data includes at least one of coarse segmentation and fine segmentation. The coarse segmentation is used to sequentially segment the original audio and video data into coarse segments of a preset duration. The fine segmentation is used to segment the original audio and video data or each coarse segment after coarse segmentation using a scene segmentation algorithm. The original audio and video data is live data or on-demand data. Determine the target content corresponding to the audio and video dataset, wherein the target content includes video screen text, audio text, and video screen description; Based on the target content, the target big model is used to obtain the target clip information corresponding to the target business scenario based on the target business scenario corresponding to the audio and video dataset. The target clip information includes the timestamp of the target segment. The audio and video dataset is identified to obtain a target optimization strategy. The identification of the audio and video dataset includes content identification and loudness identification of the original sound. The target optimization strategy includes capability arrangement decision and audio and video parameter decision. Based on the target clip information, the audio and video dataset is edited to obtain multiple target clip segments; Based on the target optimization strategy, video and audio optimizations are performed on the multiple target clips to obtain target audio and video.
2. The audio and video editing method based on a large model according to claim 1, characterized in that, Determining the target content corresponding to the audio and video dataset includes: Text recognition is performed on the video frames of the audio and video dataset to obtain the text of the video frames; The audio in the audio and video dataset is converted into text to obtain the audio text; Frame-by-frame analysis is performed on the audio and video dataset to obtain a description of the video frame.
3. The audio and video editing method based on a large model according to claim 1, characterized in that, Determining the target content corresponding to the audio and video dataset includes: Obtain historical cached content, which is obtained by processing the audio and video dataset under historical conditions. The historical cached content includes at least one of video screen text, audio text, and video screen description. Based on the historical cached content, determine the target content corresponding to the audio and video dataset.
4. The audio and video editing method based on a large model according to claim 3, characterized in that, The step of determining the target content corresponding to the audio and video dataset based on the historical cached content includes: Based on the historical cached content, determine the unprocessed data in the audio and video dataset; The unprocessed data is processed to obtain the current content; The historical cached content and the current content are determined as the target content corresponding to the audio and video dataset.
5. The audio and video editing method based on a large model according to claim 1, characterized in that, The step of obtaining target clip information based on the target content through the target large model includes: Based on the target content, the original editing information is obtained through the target large model; Based on the original editing information, determine the original editing segment; The original clip is post-processed to obtain the target clip information. The post-processing includes at least one of clip merging and output verification.
6. The audio and video editing method based on a large model according to claim 5, characterized in that, The post-processing includes fragment merging; The post-processing of the original clip to obtain the target clip information includes: Merge adjacent original clips with timestamp intervals less than the first duration to obtain the first clip; The target clip information is obtained based on the first clip.
7. The audio and video editing method based on a large model according to claim 5, characterized in that, The post-processing includes output verification; The post-processing of the original clip to obtain the target clip information includes: Delete the original clip that is shorter than the second clip to obtain the second clip; The target clip information is obtained based on the second clip.
8. The audio and video editing method based on a large model according to claim 1, characterized in that, When the target optimization strategy indicates that video subtitle synthesis is required, the step of performing video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video includes: The multiple target clips are subjected to video format conversion, video subtitle synthesis, and audio format conversion to obtain multiple first optimized clips; The target audio and video are obtained by splicing together the multiple first optimized clips.
9. The audio and video editing method based on a large model according to claim 1, characterized in that, When the target optimization strategy indicates that background music generation is required, the step of performing video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video includes: The multiple target clips are subjected to video and audio format conversion to obtain multiple second optimized clips; The multiple second optimized clips are spliced together to obtain candidate clip audio and video; Based on the feature information of the candidate audio and video clips, background music corresponding to the candidate audio and video clips is generated; The background music and the original sound of the candidate clip audio / video are mixed to obtain the target audio / video.
10. The audio and video editing method based on a large model according to claim 1, characterized in that, The target clip information includes the timestamp of the target segment; The step involves editing the audio and video dataset based on the target clip information to obtain multiple target clip segments, including: Based on the timestamp of the target segment, the audio and video dataset is edited to obtain multiple target clip segments.
11. The audio and video editing method based on a large model according to claim 10, characterized in that, The target clip information also includes narration content; When the target optimization strategy representation requires narration generation, the step of performing video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video includes: The multiple target clips are subjected to video and audio format conversion to obtain multiple second optimized clips; The multiple second optimized clips are spliced together to obtain candidate clip audio and video; Based on the narration content, add narration audio and narration subtitles to the candidate edited audio and video to obtain the target audio and video.
12. The audio and video editing method based on a large model according to claim 11, characterized in that, The step of adding narration audio and narration subtitles to the candidate edited audio and video based on the narration content to obtain the target audio and video includes: Determine the background music corresponding to the candidate audio / video clips; Based on the narration content, generate narration audio and narration subtitles; The target audio is obtained based on the original sound of the candidate edited audio and video, the narration audio, and the background music; The target video is obtained based on the original subtitles and the narration subtitles of the candidate edited audio and video. The target audio and video are obtained based on the target video and the target audio.
13. The audio and video editing method based on a large model according to claim 12, characterized in that, The step of obtaining the target audio based on the original sound of the candidate edited audio / video, the narration audio, and the background music includes: The narration audio and the background music are merged to obtain a merged audio; The original audio of the candidate clip is inserted into the time period during which the merged audio is not present in the candidate clip audio / video to obtain the target audio.
14. The audio and video editing method based on a large model according to claim 12, characterized in that, The step of obtaining the target video based on the original subtitles and the narration subtitles of the candidate edited audio and video includes: The original subtitles corresponding to the time periods when the original subtitles of the candidate edited audio and video overlap with the narration subtitles are replaced with the narration subtitles to obtain the target video.
15. An audio and video editing device based on a large model, characterized in that, include: The segmentation module is configured to segment the original audio and video data to obtain an audio and video dataset, which includes multiple audio and video segments. The segmentation of the original audio and video data includes at least one of coarse segmentation and fine segmentation. The coarse segmentation is used to sequentially segment the original audio and video data into coarse segments of a preset duration. The fine segmentation is used to segment the original audio and video data or each coarse segment after coarse segmentation using a scene segmentation algorithm. The original audio and video data is live data or on-demand data. The determination module is configured to determine the target content corresponding to the audio and video dataset, wherein the target content includes video screen text, audio text, and video screen description; The acquisition module is configured to obtain target clip information corresponding to the target business scenario based on the target content through a target big model. The target big model obtains target clip information corresponding to the target business scenario based on the audio and video dataset. The target clip information includes the timestamp of the target segment. The recognition module is configured to recognize the audio and video dataset to obtain a target optimization strategy, wherein the recognition of the audio and video dataset includes content recognition and loudness recognition of the original sound, and the target optimization strategy includes capability arrangement decision and audio and video parameter decision. The editing module is configured to edit the audio and video dataset according to the target editing information to obtain multiple target editing segments; The optimization module is configured to perform video and audio optimization on the multiple target clips according to the target optimization strategy to obtain target audio and video.
16. A computer-readable medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processing device, it implements the steps of the method according to any one of claims 1-14.
17. An electronic device, characterized in that, include: A storage device on which computer programs are stored; A processing device for executing the computer program in the storage device to implement the steps of the method according to any one of claims 1-14.
18. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-14.
Citation Information
Patent Citations
Method for automatically understanding, editing and explaining movie and television play based on large model
CN119967234A