Video and audio integrated packaging and editing method, device and equipment

Through the disassembly and semantic analysis of audio and video data, the clip positioning axis is generated for automatic editing and format packaging, which solves the problem of low video and video clip efficiency, and achieves efficient and accurate multi-terminal adaptation and content dissemination.

CN120547404APending Publication Date: 2025-08-26BEIJING FILM & VIDEO ORIGIN FILM & TELEVISION CULTURE MEDIA CO LTD

Patent Information

Application Number
CN202510895161.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-08-26

AI Technical Summary

Technical Problem

In the prior art, audio and video editing is low, and the deep semantic relationship between video and audio cannot be fully explored, and it is difficult to adapt to the format conversion and cross-platform compatibility problems of multi-terminal devices and platforms.

Method used

By disassembling and separating the original audio-visual data, the pre-trained audio-visual information analysis model is used for semantic analysis, the semantic sequence of audio and video is generated, key events are identified and the clip positioning axis is generated, and the audio-visual editing and format encapsulation is finally carried out to meet the needs of multi-term media platforms.

Benefits of technology

It improves the editing efficiency of audio and video clips, reduces manual intervention, can accurately adapt to the needs of different platforms, and improves the usability and dissemination effect of audio and video content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120547404A_ABST
    Figure CN120547404A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of multimedia processing, and discloses a video and audio integrated packaging and editing method, device and equipment, and the method comprises the steps: disassembling original video and audio data, respectively extracting audio and video information streams, carrying out semantic analysis through a pre-trained video and audio information analysis model, and generating a semantic sequence of the audio and the video; based on the sequences, key events are identified, an editing positioning shaft is generated, finally, video and audio editing and format packaging are carried out according to the positioning shaft, the requirements of a multi-end platform are met, and corresponding editing files are generated, according to the method, through intelligent analysis and automatic editing, the editing efficiency is greatly improved, manual intervention is reduced, and the editing efficiency is improved. Moreover, the method can accurately adapt to the demands of different platforms, improves the availability and propagation effect of video and audio contents, and solves a problem that the video and audio editing efficiency is low in the prior art.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of multimedia processing, and in particular to a method, device and equipment for integrated video and audio packaging and editing. Background Art

[0002] With the rapid growth of multimedia content consumption, especially on video platforms and social media, users' demand for the production, editing and sharing of audio and video content has gradually increased. This trend has prompted the rapid development of video editing and packaging technologies to meet the requirements of multiple devices and platforms. Current traditional video editing methods mainly rely on manual operations, which is not only time-consuming and labor-intensive, but also unable to fully explore the deep semantic relationship between video and audio. In particular, when audio and video information needs to be integrated and adapted to multiple platforms (such as mobile phones, TVs, web pages, etc.), problems such as format conversion, content distortion and cross-platform compatibility often arise. Summary of the Invention

[0003] The purpose of the present invention is to provide a method, device and equipment for integrated audio and video packaging editing, aiming to solve the problem of low efficiency in editing audio and video in the prior art.

[0004] The present invention is implemented as follows: In a first aspect, the present invention provides a method for integrated video and audio packaging and editing, comprising: Separate the original audio and video data into audio data and video data to obtain audio information stream and video information stream; Performing semantic parsing on the audio information stream and the video information stream using a pre-trained audio-visual information parsing model to generate an audio semantic sequence and a video semantic sequence; Identifying key events of the original audiovisual data based on the audio semantic sequence and the video semantic sequence to generate an editing positioning axis; The audio information stream and the video information stream are subjected to audio and video editing and format packaging adapted to a multi-terminal media platform according to the editing positioning axis to generate an editing file corresponding to the multi-terminal media platform.

[0005] In a second aspect, the present invention provides a video and audio integrated packaging and editing system, which is used to implement the video and audio integrated packaging and editing method described in any one of the first aspects, comprising: A data separation module is used to separate the original audio and video data into audio data and video data to obtain an audio information stream and a video information stream; A semantic parsing module, configured to perform semantic parsing on the audio information stream and the video information stream using a pre-trained audio-visual information parsing model to generate an audio semantic sequence and a video semantic sequence; An event recognition module, configured to recognize key events in the original audio-visual data based on the audio semantic sequence and the video semantic sequence, so as to generate an editing positioning axis; The file editing module is used to perform audio and video editing and format packaging on the audio information stream and the video information stream according to the editing positioning axis to adapt to the multi-terminal media platform, so as to generate an editing file corresponding to the multi-terminal media platform.

[0006] In a third aspect, the present invention provides an integrated audio and video packaging and editing device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements an integrated audio and video packaging and editing method as described in any one of the first aspects.

[0007] The present invention provides a method for integrated video and audio packaging and editing, which has the following beneficial effects: The present invention disassembles the original audio and video data, extracts the audio and video information streams respectively, performs semantic analysis using a pre-trained audio and video information analysis model, generates semantic sequences of audio and video, identifies key events based on these sequences and generates editing positioning axes, and finally performs audio and video editing and format packaging according to the positioning axes to adapt to the requirements of multiple platforms and generate corresponding editing files. This method greatly improves editing efficiency and reduces manual intervention through intelligent analysis and automated editing, and can accurately adapt to the needs of different platforms, thereby improving the availability and dissemination effect of audio and video content and solving the problem of low efficiency in editing audio and video in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0008] Figure 1 This is a schematic diagram of the steps of a method for integrated video and audio packaging and editing provided by an embodiment of the present invention; Figure 2 It is a structural diagram of an audio-visual integrated packaging and editing system provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0009] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0010] The implementation of the present invention is described in detail below with reference to specific embodiments.

[0011] Reference Figure 1 、 Figure 2 As shown, a preferred embodiment of the present invention is provided.

[0012] In a first aspect, the present invention provides a method for integrated video and audio packaging and editing, comprising: S1: Decomposing and separating the original audio and video data into audio data and video data to obtain audio information stream and video information stream; S2: performing semantic parsing on the audio information stream and the video information stream using a pre-trained audio-visual information parsing model to generate an audio semantic sequence and a video semantic sequence; S3: identifying key events in the original audio-visual data based on the audio semantic sequence and the video semantic sequence to generate an editing positioning axis; S4: The audio information stream and the video information stream are subjected to audio and video editing and format packaging adapted to a multi-terminal media platform according to the editing positioning axis to generate an editing file corresponding to the multi-terminal media platform.

[0013] Specifically, in step S1 of the embodiment provided by the present invention, the original audio and video data is extracted from the video file (such as MP4, AVI, MOV and other formats), and decoded into original audio and video streams using an audio and video decoder (such as FFmpeg or other dedicated decoding library). During the decoding process, the audio and video data are respectively converted into corresponding original data formats: audio is usually converted into PCM format or other lossless audio formats (such as WAV), and video is converted into frame data or each frame image of the video stream. Decoding converts the audio and video data from a compressed format into uncompressed original data, which is convenient for subsequent further analysis and processing. After the audio and video are separated, they can be processed independently, and audio signal analysis and video image analysis can be performed respectively. This operation ensures that the subsequent steps can perform independent semantic parsing and event recognition on the audio and video content respectively, thereby improving the flexibility and operability of the entire process.

[0014] More specifically, after decoding, both audio and video streams contain timestamp information. The audio data stream and the video data stream will be synchronized and aligned according to the output of the decoder to ensure that the timestamps of each frame of video and each audio segment correspond to each other. The timestamps of each frame of video and audio will be recorded to ensure that the audio and video data streams are synchronized and consistent on the timeline. Time synchronization and timestamp marking provide the basis for subsequent semantic parsing and event recognition, so that audio and video data can be accurately matched in different analysis modules, which is crucial for dual event positioning and causal analysis of audio and video content. Timestamp information ensures that the logical order of audio and video is not disrupted, ensuring that the edited content can be played smoothly according to the original rhythm and plot.

[0015] More specifically, the decoded audio stream and video stream are stored separately. The audio stream can be saved in WAV, MP3 or FLAC format, while the video stream is saved as image frames (for example, saved in JPEG, PNG or H.264 encoding format). For the audio stream, features such as spectrum, rhythm, and pitch can be extracted; for the video stream, features such as inter-frame differences and image content can be extracted to prepare for subsequent semantic analysis. The audio and video data streams are stored and processed independently so that they can be independently analyzed and optimized in subsequent stages to avoid mutual interference. This independent processing method ensures the integrity of the audio and video content and facilitates the further use of deep learning or traditional machine learning methods to perform refined analysis of the semantic information of audio and video.

[0016] More specifically, metadata extracted from audio streams includes: audio duration, frequency, volume, spectral information, etc.; metadata from video streams includes: frame rate, resolution, content features of each frame (such as facial recognition, scene changes, etc.), and semantic labels of image content. Generated through automated algorithms or manual annotation, this metadata will be used for subsequent semantic parsing and editing decisions. Generating audio and video metadata provides rich information for higher-level analysis of audio and video content, such as sentiment analysis, content analysis, and scene recognition. This metadata is crucial for subsequent editing decisions. Audio and video metadata provides a detailed description of the raw data, helping the system identify key events in audio and video content and providing a precise basis for editing.

[0017] More specifically, the independently processed audio information stream and video information stream are integrated through the timeline to ensure that the audio and video can be matched according to the timestamp when performing semantic analysis and event positioning. Even if they are physically stored as independent files, their synchronization can be logically guaranteed. The audio and video information streams will be passed to the semantic parsing model, event recognition model, etc. for analysis in subsequent steps. The integration of audio and video streams provides a complete audio-visual information foundation, ensuring that subsequent semantic analysis and editing can be based on complete and accurate data streams. Through synchronized audio and video information streams, the system can perform more precise analysis and positioning during the editing process, thereby improving the quality and accuracy of the editing effect.

[0018] More specifically, the quality of the separated audio and video streams is evaluated and verified to ensure that no key information is lost after decoding. Compensation and reconstruction strategies are used to repair possible errors or missing data. The quality of audio and video is analyzed in multiple dimensions to ensure that they can adapt to subsequent editing needs. Quality assessment ensures that the separated audio and video streams can meet the requirements of subsequent processing and avoid data problems that affect subsequent analysis and editing. Data verification and repair technology can ensure that the final edited video and audio are distortion-free, providing the best viewing experience.

[0019] It is understandable that by disassembling and independently processing the audio and video streams, special optimization can be performed for each type of data stream, thereby improving the accuracy of overall analysis and editing. After the audio and video data are disassembled, they can be processed separately and optimized and analyzed using different algorithms, which improves the flexibility and intelligence of the system. By marking timestamps, the audio and video can be kept synchronized during the editing process to avoid misalignment or incoordination. Through precise metadata and event positioning, key events can be accurately identified, and corresponding editing optimization can be performed according to the needs of different platforms, thereby improving the quality and adaptability of the final work. The technical implementation of this step will provide strong data support and technical guarantee for subsequent audio and video editing, semantic analysis and multi-platform adaptation.

[0020] Specifically, in step S2 of the embodiment provided by the present invention, feature extraction is performed on the audio information stream, including the audio spectrum, pitch, volume, duration, etc., and audio features are extracted using MFCC (Mel-Frequency Cepstral Coefficients), Chroma, Zero-Crossing Rate and other technologies. Feature extraction is performed on the video information stream, mainly extracting edges, textures, color distribution, motion information, etc. in the image frames, and image feature extraction is performed using convolutional neural networks (CNN) or optical flow methods and other technologies. Audio feature extraction can help the model understand the basic characteristics of sound, such as pitch changes, pauses, and tone in language. Video feature extraction can capture changes in details in the image, helping the model recognize objects, actions and other visual information in the scene. These features provide basic data for subsequent semantic analysis, enabling the model to better understand the content of audio and video.

[0021] More specifically, pre-trained audio and video information parsing models (such as deep neural networks, convolutional neural networks, recurrent neural networks (RNNs), etc.) are used to perform semantic parsing of audio and video streams. The semantic parsing of audio streams mainly focuses on language content, sentiment analysis, speech recognition, etc., while the semantic parsing of video streams mainly focuses on object detection, scene recognition, action recognition, etc. Multimodal learning technology is used to combine audio and video data to improve the accuracy of semantic understanding. Pre-trained models usually have strong feature extraction and pattern recognition capabilities and can automatically learn deep semantic information from audio and video data. Speech recognition models can recognize speech content, sentiment analysis models can understand the speaker's emotional changes, and video analysis models can identify key objects and actions in images. Combined with multimodal learning (such as joint training of joint features of audio and video), the model can perform comprehensive analysis of the two information streams to increase the accuracy and diversity of semantic understanding.

[0022] More specifically, based on the audio features and model analysis results, an audio semantic sequence is generated. This sequence contains key information in the audio, such as keywords, sentence sentiment, and audio content classification (for example, dialogue, background music, ambient sound, etc.). The audio time segments are associated with their corresponding semantic information to form a time-ordered audio semantic sequence. For example, LSTM (Long Short-Term Memory) networks can be used to process audio time series data and maintain temporal consistency. The generation of audio semantic sequences facilitates high-level understanding of the audio content, enabling the system to understand contextual changes, emotional fluctuations, and other non-linguistic features in the audio (such as changes in ambient noise). The semantic sequences generated by time segments can accurately locate key events in the audio (such as the start and end of a conversation, emotional transitions), providing an important basis for subsequent editing and analysis.

[0023] More specifically, convolutional neural networks (CNNs) or more complex visual recognition models (such as YOLO and Faster R-CNN) analyze video frames frame by frame, extracting visual features such as scenes, characters, and actions. Key elements in the video, such as people, objects, actions, and scene transitions, are identified. A semantic sequence of the video is generated in chronological order, recording the visual information and corresponding semantic labels at each time point (e.g., character appearance, scene transition, and important action). This video semantic sequence provides structured information for understanding video content, enabling the system to clearly identify each key node in the video. By processing the video frame by frame, the system can accurately capture dynamic changes and analyze the visual content and events in the video. This provides a reliable basis for subsequent event analysis and editing.

[0024] More specifically, audio and video semantic sequences are fused, and fusion models (such as multimodal neural networks and cross-modal learning) are used to combine the semantic information in audio and video. This fusion can be achieved through weighted averaging, time-series alignment, or joint representation learning in deep learning. This allows the semantic information of audio and video to be analyzed within the same framework. Multimodal semantic fusion can enhance the system's overall semantic understanding capabilities. For example, changes in a scene in a video can be combined with information such as the conversation content and emotion in the audio to achieve a more complete semantic understanding. Through cross-modal semantic fusion, the model's performance in complex scenarios can be improved, making it more advantageous in applications such as video editing and event recognition.

[0025] More specifically, the generated audio and video semantic sequences are post-processed, such as removing redundancy, filling missing semantic information, and optimizing the quality of the sequences. Sequence labeling techniques (such as conditional random fields (CRF)) can be used to further optimize the semantic sequences to make the semantic information more accurate and coherent. Post-processing and optimization can improve the accuracy and coherence of semantic sequences, and avoid noise or inaccurate information in the sequences affecting subsequent tasks (such as editing, event analysis, etc.). The optimized semantic sequences can provide more accurate support for subsequent multi-platform adaptive editing, content extraction, etc.

[0026] It is understandable that the pre-trained model can extract rich semantic information from audio and video, and improve the depth and accuracy of understanding through multimodal fusion. Through parallel processing of audio and video and semantic sequence generation, it can efficiently extract key information and provide support for subsequent editing, analysis, recommendation and other tasks. The fusion of multimodal information enables the system to perform accurate event recognition and semantic reasoning in more complex scenarios, thereby providing strong technical support for practical applications (such as automatic editing, intelligent recommendation, etc.).

[0027] Specifically, in step S3 of the embodiment provided by the present invention, the audio semantic sequence is analyzed to identify key events in the audio, such as keywords in speech recognition, emotional fluctuations, the beginning and end of dialogue paragraphs, sudden sound effects, etc. Common events include the start of dialogue, emotional turning points, important sound effects, etc. The video semantic sequence is analyzed to identify key events in the video, such as the appearance of characters, scene switching, action occurrence, behavior of important objects or characters, etc. The key events in the video are determined by object detection, scene recognition, action detection and other technologies. The key events in audio and video are the most informative part of the entire content, which can help the system focus on important information such as dialogue, action, emotional changes, etc. Analyzing audio and video semantic sequences can systematically extract valuable events from different angles, providing a basis for subsequent editing positioning and event analysis.

[0028] More specifically, the key events of audio and video are aligned according to timestamps to ensure that the key events of audio and video are synchronized in time. A time alignment algorithm (such as DTW - Dynamic Time Warping) or a direct alignment method based on timestamps can be used. Through alignment, relevant events in audio and video are ensured to be synchronized in the same time period for subsequent editing positioning. The key events of audio and video need to be consistent in time to ensure that the relationship between them is not disrupted. Time alignment ensures the smoothness of audio and video content. This process can ensure that there will be no audio and video mismatch problems during the editing process, and ensure that the time of events is accurate and the content is coherent.

[0029] More specifically, a weight is assigned to each event based on the importance of key events in the audio and video. For example, core information in a conversation, drastic emotional changes, or key action scenes are assigned higher weights, while background noise, irrelevant actions, etc. are assigned lower weights. Machine learning models or expert rules can be used to classify and weight events. For example, natural language processing (NLP) can be used to perform sentiment analysis on conversations, or computer vision models can be used to identify the appearance of key objects. Assigning weights to key events can help the system determine which events are more important and which events have a greater impact on the editing effect. Weight assignment and sorting make the editing process more intelligent, and the system can automatically select the most valuable clips for the storyline, emotional expression, or visual presentation.

[0030] More specifically, based on the timing, weight and priority of key events in audio and video, an editing positioning axis is generated. The editing positioning axis is a sequence of time periods that marks the time points when key events occur and the corresponding editing segments. On the editing positioning axis, the system marks each key event and arranges it in chronological order. The positioning axis provides clear time guidance for subsequent automatic editing. The editing positioning axis identifies key events in audio and video in time and provides precise time nodes for editing, which makes the editing process more efficient and accurate, reduces manual intervention, and the generated positioning axis can provide clear editing guidance to ensure that the content presentation meets the expected structure and fluency.

[0031] More specifically, based on the editing positioning axis, automatic editing suggestions are generated to determine which key events should be retained and which can be deleted or modified. The editing suggestions can be optimized based on the weight, importance and editing goals of the events. For example, for the emotional climax, more time segments are retained, while repetitive or insignificant scenes can be deleted or skipped. By optimizing the editing suggestions, the system can edit according to the focus of key events and improve the quality of the editing effects. For example, the climax part of the plot development should be retained, while some irrelevant or redundant segments can be omitted. Through automated editing suggestions, editing efficiency can be greatly improved, while reducing the deviation of human decision-making and ensuring the quality of content.

[0032] More specifically, the editing results generated according to the editing positioning axis will be verified to check whether key events are accurately retained, whether the audio and video synchronization is good, and whether the plot is smooth. In the actual editing process, if the system fails to correctly identify certain key events or there are editing problems, adjustments can be made through manual intervention or automated correction mechanisms. Verification of editing results is an important part of ensuring the quality of the final video. Verification can ensure that the audio and video content is not distorted due to editing errors. The automated correction mechanism can improve the intelligence of the system, allowing the editing process to self-optimize and ensure that the output content meets the expected quality standards.

[0033] More specifically, the final edited video is generated based on the editing positioning axis and optimized editing suggestions. By synchronizing, aligning, and editing the audio and video, the system stitches together key events in the correct order and rhythm into a complete, plot-aligned edited video. The resulting video, based on key audio and video events, ensures a coherent and rhythmic plot, conveying the correct emotions and information. Automated editing significantly improves efficiency and ensures the consistency and quality of editing.

[0034] It is understandable that through automated key event identification and editing positioning axis generation, the system can quickly and accurately provide a basis for the editing process, reduce manual intervention, and weight assignment and event priority sorting make the editing results more in line with story requirements, with more prominent emotional and visual effects. The intelligent editing process combined with audio and video semantic analysis can greatly improve the accuracy and effect of editing, reduce human errors, and ensure the quality of the final work. By continuously optimizing editing suggestions and verification and correction mechanisms, the system can adapt to different types of audio and video content and provide flexible and accurate editing solutions.

[0035] Specifically, in step S4 of the embodiment provided by the present invention, the corresponding time period content in the audio information stream and the video information stream is extracted according to the generated editing positioning axis. Specifically, the editing time period of audio and video should be accurately intercepted according to the timestamp range in the positioning axis to ensure the accuracy of the editing. During the extraction process, it is necessary to ensure the synchronization of audio and video. The audio editing time period should correspond to the video time period to avoid the problem of audio and video being out of sync. By accurately intercepting the audio and video content, the editing result is guaranteed to be consistent with the positioning axis to avoid unnecessary content loss or dislocation. This process ensures the integrity and continuity of the audio and video content, and lays a solid foundation for subsequent format encapsulation.

[0036] More specifically, audio and video content is adapted according to the characteristics and needs of the target platform. For example, different platforms have different screen resolutions and device types (mobile phones, tablets, PCs, etc.) with different requirements for video resolution, frame rate, audio encoding, etc. Therefore, it is necessary to adjust parameters such as video resolution, bit rate, and audio sampling rate according to the specific requirements of the target platform. For video, its encoding format (such as H.264, HEVC) and resolution (such as 1080p, 720p, etc.) need to be adjusted to adapt to the playback environment of different platforms. For audio, it is necessary to select and adjust the audio encoding based on the audio encoding format supported by the platform (such as AAC, MP3, Opus, etc.), and select an appropriate bit rate based on the performance of the device. Multi-end adaptation ensures that the generated clip files can be played smoothly on different devices and platforms, and the user experience will not be affected by format incompatibility or low resolution. The adaptation of video and audio can maximize the viewing and listening experience of multi-end users, ensuring that the content can achieve the best effect regardless of the device it is played on.

[0037] More specifically, the extracted audio and video data is encoded and compressed using efficient video and audio coding formats (such as H.264, HEVC, AAC, etc.) to reduce file size while maintaining good visual and auditory quality. Appropriate compression parameters are selected to balance the relationship between file size and quality. For video, reasonable bit rate, resolution and frame rate can be set; for audio, appropriate sampling rate and bit rate can be set. Encoding and compression processing can effectively reduce the size of the generated file, facilitate transmission and storage, and increase loading speed. Efficient encoding can ensure that video and audio files are minimized while ensuring quality, so that they can adapt to limited network bandwidth and take into account the performance requirements of playback devices.

[0038] More specifically, audio and video content is encapsulated according to the format requirements of multiple platforms. Common encapsulation formats include MP4, MKV, WebM, AVI, etc. These formats have different support and applicability on different platforms. According to the support of the target platform, the appropriate encapsulation format is selected. For example, MP4 format can be selected for mobile devices, WebM format can be selected for network playback platforms, and MKV format can be selected for high-quality playback devices. The audio and video streams are encapsulated, and the audio stream and video stream are merged into the final media file. At the same time, metadata (such as title, cover, chapter information, etc.) and appropriate subtitles, language tracks and other additional information are added during the encapsulation process. Encapsulation ensures that the generated file can be played smoothly on multiple platforms and provides appropriate metadata and additional information. By selecting the appropriate encapsulation format, the file can be compatible with players and devices on different platforms, ensuring cross-platform playback performance.

[0039] More specifically, after the generated multi-end clip files are encapsulated, they are optimized and verified. Optimization includes adjusting the file's fault tolerance, optimizing file header information, cleaning up redundant data, and performing integrity and consistency checks on the generated files to ensure there are no issues such as file corruption, audio and video asynchrony, and resolution incompatibility. File integrity checks can be performed using methods such as hash checks and MD5 checks. File optimization and verification can improve file stability and compatibility, ensuring error-free playback on different devices. This process ensures that users do not encounter playback interruptions or substandard quality issues during playback, thereby improving the overall viewing experience.

[0040] More specifically, the multi-terminal adapted editing files that have been formatted, optimized and verified are uploaded to the target media platform. Different platforms (such as YouTube, Bilibili, TikTok, etc.) have different upload specifications, so corresponding adjustments need to be made according to the requirements of each platform. During the upload process, data compression or resolution change is required according to the platform's requirements, and the upload time and data size must be in line with the platform's upload restrictions to ensure that the file can be successfully uploaded to the target platform and played normally on the platform to achieve the expected multi-platform dissemination effect. The distribution process ensures that the file reaches users quickly and safely, improving the visibility and dissemination efficiency of the content.

[0041] It is understandable that through multi-terminal adaptation processing, encoding and compression optimization, the generated clip files can be played without obstacles on various devices and platforms. The compression and encoding optimization of video and audio improves the storage and transmission efficiency of files, ensuring a better user experience. At the same time, it also ensures that there will be no delay or freeze during file transmission. Through reasonable packaging and format selection, high-quality audio and video content can be provided on different platforms, and meet the technical requirements of each platform to ensure that the content can be displayed smoothly. Through the synchronization, encoding and compression processing of audio and video content, the playback effect is improved, and technical problems that may be encountered during playback are avoided, providing a smooth viewing experience.

[0042] The present invention provides a method for integrated video and audio packaging and editing, which has the following beneficial effects: The present invention disassembles the original audio and video data, extracts the audio and video information streams respectively, performs semantic analysis using a pre-trained audio and video information analysis model, generates semantic sequences of audio and video, identifies key events based on these sequences and generates editing positioning axes, and finally performs audio and video editing and format packaging according to the positioning axes to adapt to the requirements of multiple platforms and generate corresponding editing files. This method greatly improves editing efficiency and reduces manual intervention through intelligent analysis and automated editing, and can accurately adapt to the needs of different platforms, thereby improving the availability and dissemination effect of audio and video content and solving the problem of low efficiency in editing audio and video in the prior art.

[0043] Preferably, the step of separating the original audio and video data into audio data and video data to obtain the audio information stream and the video information stream includes: S11: acquiring original audio and video data, and performing audio waveform decoding and video frame decoding on the original audio and video data, respectively, to obtain audio waveform decoding data and video frame decoding data separated from the original audio and video data; S12: During the decoding process of the original audio and video data, synchronously generating a timestamp, and assigning the timestamp at each moment to the audio waveform decoded data and the video frame decoded data; S13: Constructing a unified timeline based on the timestamp marks of each moment, and performing information mapping on the audio waveform decoded data and the video frame decoded data relative to the unified timeline, so as to disassemble the audio waveform decoded data into a plurality of audio metadata and assign them to the unified timeline to obtain an audio information stream, and disassemble the video information stream into a plurality of video metadata and assign them to the unified timeline to obtain a video information stream.

[0044] Specifically, the original data stream containing the audio and video content is obtained, usually a composite media file (such as MP4, MKV, MOV, etc.). This file contains the combined data of video frames and audio waveforms. This step is the starting point of the disassembly process, ensuring that all necessary audio and video data are extracted from the unified media container in preparation for subsequent decoding and disassembly.

[0045] More specifically, the original audio and video data is decoded to extract audio data and video data respectively. Audio waveform decoding refers to encoding and decoding audio data (for example, decoding from AAC and MP3 formats to PCM waveform data); video frame decoding is decoding video data into image data of each frame (for example, decoding from H.264 format to RGB or YUV image data). During the decoding process, the corresponding codec (such as FFmpeg, GStreamer, etc.) is used to decode the data. The decoding step extracts the audio and video streams from the original data and converts them into a format that the computer can process (such as audio PCM waveform and video image frame). This is the basis for disassembling audio and video data. The use of specialized decoders ensures the accuracy and efficiency of data format conversion, making subsequent audio and video processing smoother.

[0046] More specifically, during the audio and video decoding process, a timestamp is generated for each frame of video and each audio segment according to the decoded frame rate (such as the frame rate of the video and the sampling rate of the audio). The timestamp represents the playback moment of the audio and video data, that is, the playback time corresponding to each frame of video and each audio segment. The timestamp mark corresponds one-to-one with the audio waveform data and the video frame data during the decoding process to ensure the synchronization of audio and video. The synchronous generation of timestamps is the key to ensuring the synchronization of audio and video streams. In multimedia processing, audio and video synchronization is very important. Through accurate timestamp marking, the playback time of audio and video can be ensured to be consistent, avoiding the phenomenon of audio and video being out of sync. The timestamp provides a time basis for subsequent mapping and rearrangement, ensuring the consistency and accuracy of the processing flow.

[0047] More specifically, the timestamps of the audio and video are arranged in chronological order, and combined with the frame rate or sampling rate of the audio and video data, a unified timeline is constructed. This timeline integrates the timestamps of the video frames and audio clips to form a unified timeline. Each moment on the timeline corresponds to an audio metadata and video metadata. Each time point of the audio and video data is mapped to the unified timeline, ensuring that the playback order and duration of each audio and video data segment are precisely controlled. The unified timeline is a key technical link to ensure the synchronization of audio and video data. By unifying the timeline, it can be ensured that the audio and video data are played in the correct order and duration on multiple devices. The construction of the timeline enables the audio and video data to be accurately aligned on the timeline, avoiding synchronization problems caused by inconsistent timelines.

[0048] More specifically, according to the unified timeline, the audio data will be decomposed into multiple audio metadata blocks (for example, each audio metadata block can be an audio clip or audio data within a time period). Each audio metadata block will be assigned to the unified timeline according to the timestamp mark to generate an audio information stream. The audio information stream contains the audio clip data corresponding to each moment on the timeline. This step ensures that the playback time of each audio clip is accurate by decomposing the audio data into smaller audio metadata blocks and mapping them according to the timeline, avoiding audio loss or misalignment. The decomposed audio data is convenient for subsequent processing and playback, and is also convenient for more efficient storage and transmission.

[0049] More specifically, similarly, based on the unified timeline, the video data will be disassembled into multiple video metadata blocks (for example, each video metadata block is a video frame or a time period of a video frame), and each video metadata block will be assigned to the unified timeline according to the timestamp mark to generate a video information stream. The video information stream contains the video frame data corresponding to each moment on the timeline. Disassembling the video data and mapping it to the unified timeline can ensure that each frame of the video is played in chronological order. The disassembly and mapping of the video frames enable the video information stream to be accurately presented to the user with minimal delay. This step ensures that the video content can be displayed smoothly on multiple devices and maintains its picture smoothness.

[0050] More specifically, the audio metadata stream and the video metadata stream are based on a unified timeline, and the audio information stream and the video information stream are merged or saved separately in formats that adapt to different needs. Through further processing of the audio and video streams, these information streams can be encapsulated into target media formats, such as MP4, WebM, AVI, etc. The core of this step is to convert the disassembled audio and video information streams into file formats suitable for playback and storage, preparing for subsequent multi-terminal distribution and playback. The encapsulation of the information stream provides adaptability for playback on different platforms, ensuring that it can meet the playback needs of various devices while ensuring the stability and smoothness of the playback.

[0051] It can be understood that the marking of timestamps and the construction of a unified timeline ensure the precise synchronization of audio and video data, avoiding the problem of audio and video asynchrony. The disassembly and mapping of audio and video can make each data unit independent and can accurately control the playback time and duration. The final generated audio and video information stream can be further processed and packaged according to the needs of different platforms and devices to adapt to the playback needs of multiple platforms and multiple devices. Through precise timeline mapping and data disassembly, the smooth playback of audio and video data is ensured, and the user experience is improved.

[0052] Preferably, the step of performing semantic parsing on the audio information stream and the video information stream by using a pre-trained audio-visual information parsing model to generate an audio semantic sequence and a video semantic sequence includes: S21: Inputting the audio information stream and the video information stream into a pre-trained audio-visual information parsing model, so as to perform information stream semantic parsing using multiple parsing methods on the audio information stream and the video information stream respectively through the audio-visual information parsing model, thereby obtaining semantic parsing information corresponding to various parsing methods for the audio information stream and the video information stream; wherein the audio-visual information parsing model is a pre-trained convolutional neural network model; S22: performing information fusion of the semantic parsing information of each audio metadata on the unified timeline based on the semantic parsing information corresponding to each parsing method of the audio information stream to obtain audio semantic units corresponding to each audio metadata and assigning them to the unified timeline to obtain an audio semantic sequence; S23: Based on the semantic parsing information of the video information stream corresponding to various parsing methods, the semantic parsing information of each video metadata on the unified timeline is fused to obtain the video semantic unit corresponding to each video metadata and assign it to the unified timeline to obtain a video semantic sequence.

[0053] Specifically, the audio information stream and video information stream mapped through the timeline are input into a pre-trained audio-visual information parsing model, which is usually composed of multiple parallel convolutional neural networks (CNNs) or fusion networks (such as CNN+LSTM, CNN+Transformer). The audio information stream (such as PCM, Mel spectrogram) and video frame sequence (such as RGB frame, optical flow frame) will be sent to different sub-models respectively. The pre-trained model has learned the expression features of audio events (such as speaking, clapping, explosion) and video semantics (such as human actions, object recognition, scene type, etc.) on large-scale audio-visual corpus. The use of CNN architecture can effectively extract local spatiotemporal features, which is the mainstream method for processing audio spectrograms and image frames. It significantly improves the accuracy and robustness of semantic recognition, reduces the workload of manually designed features, and realizes end-to-end feature extraction and recognition.

[0054] More specifically, the model uses multiple semantic parsing methods for audio and video, such as: audio: speech recognition (ASR), sound event detection (SED), and emotion recognition (SER); video: object recognition, action recognition, scene recognition, face recognition, etc. Each parsing method outputs a set of time-aligned semantic labels or vector representations. Different semantic parsing methods focus on different information dimensions. Multi-dimensional semantic parsing can build a more complete semantic understanding. The multi-task learning (MTL) framework can parse multiple dimensions in parallel, improve computational efficiency and model generalization capabilities, and achieve more comprehensive and fine-grained semantic extraction of audio and video information. The synchronized parsing results ensure the cross-modal correspondence of audio and video semantics.

[0055] More specifically, the audio semantic information from different parsing methods (such as text transcription, sound tags, and emotion tags) is fused for each audio metadata on a unified timeline to obtain structured audio semantic units. These semantic units are assigned to each audio segment on the unified timeline in chronological order to form an audio semantic sequence. The features or labels output by each parsing method are complementary, and the fusion processing can improve semantic consistency. The timeline assignment ensures that the semantic information has a time context, which is convenient for subsequent reasoning or retrieval, and improves the accuracy and expressiveness of the audio semantic recognition results. A clearly structured audio semantic sequence is output, which is convenient for visualization, subtitle generation, or knowledge graph construction.

[0056] More specifically, similarly, the video semantic information under various parsing methods (such as target detection, action classification, and scene labeling) is fused for each video metadata, and the fused semantic units are mapped to a unified time axis to generate a video semantic sequence. There is a temporal correlation between video frames. The fusion processing can enhance semantic continuity and context understanding. The unified time axis processing enables the subsequent cross-modal alignment and joint analysis of audio and video semantic sequences, and constructs a structured video semantic description stream for subsequent event detection, behavior understanding, cross-modal retrieval and other applications, thereby improving the integrity and diversity of semantic extraction.

[0057] Preferably, the parsing method of the audio-visual information parsing model includes a metadata timing parsing method, and the step of performing information flow semantic parsing of the audio information stream and the video information stream in the metadata timing parsing method using the audio-visual information parsing model includes: S2101: Assigning parsing sequence numbers to each audio metadata of the audio information stream and each video metadata of the video information stream according to the information correspondence between the audio information stream and the video information stream relative to the unified time axis; S2102: performing semantic parsing and vectorization expression on each audio metadata and each video metadata in sequence according to the parsing sequence number of each audio metadata and each video metadata, so as to generate a semantic parsing feature matrix of each audio metadata and each video metadata; S2103: Arranging the semantic parsing feature matrices corresponding to the audio metadata and the video metadata according to the parsing sequence numbers, and interactively verifying the semantic parsing feature vectors of the sequentially arranged semantic parsing feature matrices to assign confidence information to the semantic parsing feature matrices corresponding to the audio metadata and the video metadata; S2104: Collaboratively analyze each semantic parsing feature matrix based on its confidence information to extract a unique semantic parsing feature vector from each semantic parsing feature matrix as semantic parsing information of the metadata timing parsing method corresponding to the audio information stream and the video information stream.

[0058] Specifically, according to the correspondence between the audio information stream and the video information stream relative to the unified time axis, a parsing sequence number is assigned to each audio metadata and video metadata. This sequence number indicates the sequential position of the metadata in the time axis. In the process of processing audio and video data, the audio and video metadata on the unified time axis need to be parsed and fused in sequence. Assigning a parsing sequence number helps to establish a temporal alignment relationship between audio and video. By clearly identifying the timestamp and sequence of each metadata, it ensures that the audio and video data can be processed in the correct time sequence in subsequent steps, ensuring that the audio and video data are accurately aligned on the time axis, avoiding time sequence confusion, enhancing the consistency of cross-modal data processing, and providing a clear structure for subsequent semantic parsing and feature vector generation.

[0059] More specifically, the audio metadata and video metadata assigned with parsing numbers are semantically parsed in turn, and the parsing results are converted into vectorized expressions (for example, speech transcription and event detection results for audio, object recognition and action recognition for video). A semantic parsing feature matrix will be generated for each audio and video metadata, which contains all the semantic features of the metadata. Semantic parsing and vectorization are key steps in abstracting and compressing the original audio and video data. The content of audio and video can be converted into vectors in a high-dimensional feature space, so that the model can understand the audio and video content by calculating the similarity between features. The use of vectorized expression can convert complex spatiotemporal information into digital features, which is convenient for subsequent processing and analysis, and improves the semantic abstraction ability of audio and video information, so that the original data can be efficiently represented and processed. The vectorized feature matrix can capture the details in the audio and video content and is suitable for subsequent deep learning model training and reasoning.

[0060] More specifically, the semantic parsing feature matrices corresponding to the audio and video metadata are arranged according to the parsing sequence number to ensure that the order of each feature matrix is ​​consistent with the time axis. The arranged semantic parsing feature matrices are interactively verified for semantic parsing feature vectors to check the consistency and reliability of these feature vectors in cross-modal data. In the process of audio and video information processing, the feature matrices of audio and video must be arranged in chronological order to maintain temporal consistency. Through interactive verification, the semantic consistency of feature vectors between different modalities (audio, video) can be ensured, thereby improving the accuracy of semantic parsing, ensuring the correct arrangement of semantic parsing feature matrices, avoiding timing deviations, and ensuring the spatiotemporal synchronization of audio and video data. The reliability of semantic parsing is improved through feature verification, especially in the process of cross-modal semantic fusion.

[0061] More specifically, based on the confidence information of each semantic parsing feature matrix (such as the model's prediction probability, similarity score, etc.), each feature matrix is ​​collaboratively analyzed to extract a unique semantic parsing feature vector. This feature vector will serve as the final semantic parsing information of the metadata temporal parsing method corresponding to the audio and video information streams. In multimodal information parsing, assigning confidence to each feature matrix helps to indicate the model's trust in the parsing results. A higher confidence level means a more reliable result. Collaborative analysis can eliminate noise from different parsing methods or data modalities, ensure that the output semantic feature vector is representative, and can capture the core semantic information of the audio and video content to the greatest extent, improve the robustness and accuracy of the model in the multimodal fusion process, and ensure that the output semantic feature vector can truly reflect the semantics of the audio and video content. Confidence information and collaborative analysis make the audio and video semantic parsing results more credible and suitable for subsequent intelligent reasoning and analysis applications.

[0062] Preferably, the parsing method of the audio-visual information parsing model includes an object-based parsing method, and the steps of performing information flow semantic parsing of the audio information stream and the video information stream in the object-based parsing method using the audio-visual information parsing model include: S21001: Assigning parsing sequence numbers to each audio metadata of the audio information stream and each video metadata of the video information stream according to the information correspondence between the audio information stream and the video information stream relative to the unified time axis; S21002: performing content location of data feedback objects for each audio metadata and each video metadata to obtain feedback object information of each audio metadata and each video metadata; S21003: performing object summarization processing on the audio information stream and the video information stream based on the feedback object information of each audio metadata and each video metadata, so as to connect each audio metadata having feedback object information that meets a preset standard with each video metadata, thereby generating a plurality of audio sub-information streams and a plurality of video sub-information streams; S21004: Analyze the object feedback frequency and object feedback time distribution of each audio sub-information stream and each video sub-information stream according to the parsing sequence number corresponding to each audio sub-information stream and each video sub-information stream, and assign an object weight curve to each audio sub-information stream and each video sub-information stream based on the analysis result; wherein the object weight curve is used to describe the object feedback weight of the sub-information stream in each time interval; S21005: Performing a multi-dimensional comparative analysis on the object weight curves of each of the audio sub-information streams and each of the video sub-information streams to obtain semantic parsing direction features of each time interval of each of the audio sub-information streams and each of the video sub-information streams; S21006: Based on the semantic parsing pointing features of each time interval of each audio sub-information stream and each video sub-information stream, semantic parsing of each audio metadata and each video metadata is performed on each audio sub-information stream and each video sub-information stream respectively, to serve as semantic parsing information of the summary parsing method of the corresponding objects of the audio information stream and the video information stream.

[0063] Specifically, according to the correspondence between the audio information stream and the video information stream on the unified timeline, a parsing sequence number is assigned to each audio and video metadata. The audio and video data must be precisely aligned with the timeline to ensure that the time series data can be accurately matched and fused, which is crucial for subsequent object induction and semantic analysis. By assigning parsing sequences, it can be ensured that the audio and video data can be processed in chronological order during parsing, ensuring the consistency of the audio and video data time series, and ensuring that multimodal data is accurately processed within the same time period, laying the foundation for subsequent semantic analysis of audio and video metadata and object feedback analysis.

[0064] More specifically, the content of the data feedback object is located for each audio and video metadata to obtain feedback object information. This information indicates the objects involved in the audio and video content (such as objects, people, events, etc.). Identifying the key objects involved in audio and video data is a key step in object summarization and semantic parsing. By using feedback object information, the core elements of the audio and video can be accurately located, facilitating subsequent summarization and processing. By obtaining feedback object information, a basis is provided for the effective integration of different audio and video metadata, improving the accuracy and reliability of object summarization processing, providing key support for subsequent content connection and data analysis, accurately locating key information in audio and video data, and improving the semantic parsing capabilities of the model.

[0065] More specifically, object summarization is performed based on the feedback object information of the audio and video metadata, and the audio metadata that meets the preset standards is connected with the video metadata to generate several audio sub-information streams and video sub-information streams. Through object summarization, the audio and video data are grouped and integrated according to the objects they involve. This is to reasonably divide unrelated information streams so that audio and video information can be effectively connected to generate more meaningful information streams. By dividing the sub-information streams in audio and video, subsequent analysis and processing are facilitated, especially in the interactive analysis of multimodal data. It helps to obtain more fine-grained semantic information, optimize the structure of the audio and video information stream, make the connection of data content more reasonable, facilitate subsequent in-depth analysis, and improve the fusion of audio and video data. Especially in terms of semantic information integration, it can improve the efficiency and accuracy of data processing.

[0066] More specifically, the frequency and time distribution of object feedback are analyzed for each audio sub-information stream and video sub-information stream, and an object weight curve is assigned to each sub-information stream based on the analysis results. By analyzing the object feedback frequency in the audio and video sub-information streams and their distribution on the timeline, it is possible to identify which objects dominate in a specific time period, or which objects are related to the main plot or events of the audio and video content. By assigning a weight curve to each sub-information stream, the degree of influence of the sub-information stream on the overall semantic parsing in each time interval can be further characterized, thereby improving the temporal sensitivity of analyzing object feedback in the audio and video information streams, and helping to reveal the evolution process of the audio and video content and the dynamic relationship between objects. The weight curve can provide support for subsequent semantic parsing, so that audio and video data in different time intervals can be reasonably evaluated and analyzed.

[0067] More specifically, a multi-dimensional comparative analysis is performed on the object weight curves of each audio sub-information stream and the video sub-information stream to extract the semantic analysis pointing features of each time interval. By performing a multi-dimensional comparison of the object weight curves, the semantic features of the audio and video data can be analyzed from different angles (such as time, frequency, feedback intensity, etc.) to identify the core patterns and laws behind them. The purpose of this step is to extract the semantic pointing features in each time period through comparative analysis. These features help to determine the theme or plot of the audio and video content. Through multi-dimensional analysis, richer and deeper semantic information is provided, making the understanding of the audio and video content more comprehensive. Comparative analysis helps to identify the similarities and differences between audio and video data, and provides an important basis for subsequent semantic fusion and reasoning.

[0068] More specifically, based on the semantic parsing pointing features of each time interval, the audio sub-information stream and the video sub-information stream are semantically parsed to obtain the semantic parsing information of the object summary parsing method corresponding to the final audio information stream and video information stream. By parsing the semantic features of the audio and video sub-information streams, a unified semantic understanding can ultimately be provided for multimodal data. The purpose of this step is to extract the overall semantics of audio and video from the feedback object information and weight curves, providing a clear semantic framework for the overall understanding of audio and video content, facilitating subsequent applications (such as intelligent recommendation, automatic summarization, etc.). Through semantic parsing, a unified and in-depth semantic understanding can be provided for audio and video data, promoting the further development of audio and video data analysis technology. This process greatly improves the accuracy of semantic understanding and can provide strong support for various practical applications, such as intelligent monitoring and automatic video editing.

[0069] Preferably, the step of identifying key events of the original audiovisual data based on the audio semantic sequence and the video semantic sequence to generate an editing positioning axis includes: S31: performing dual event positioning on the original audio-visual data based on the audio semantic sequence and the video semantic sequence to obtain deduced events of the original audio-visual data at various time nodes; S32: combining the deduced events of the original video and audio data at each time node, performing a causal correlation analysis on the deduced events at each time node to generate causal correlation parameters between the deduced events at each time node; S33: performing a connection simulation on the deduced events of the original video and audio data at each time node according to the causal correlation parameters between the deduced events at each time node, so as to obtain a plurality of deduced event information chains; S34: performing overall semantic analysis on each of the deduced event information chains, and performing event criticality assessment on the original audio-visual data based on the overall semantic analysis result, so as to obtain an event criticality curve of the original audio-visual data; S35: Based on the causal correlation parameters between the event criticality curve and the deduced events at each time node, the editing value of the deduced events of the original audio and video data at each time node is evaluated relative to the multi-terminal media platform to obtain the editing value positioning information of the original audio and video data at each time node, so as to jointly constitute the editing positioning axis of the original audio and video data.

[0070] Specifically, based on the audio semantic sequence and the video semantic sequence, dual event positioning is performed on the original audio and video data to obtain the deduced events of the original audio and video data at each time node. The audio and video information each carry different semantic features. Only by combining the two can events be identified and located more comprehensively. By analyzing the semantic sequences of audio and video, the occurrence of events can be determined at more precise time nodes. Since events usually have specific temporal sequences, the combination of audio and video helps to improve the accuracy of event recognition, reduce the error of a single mode, improve the accuracy of event positioning, ensure that the synchronous events of audio and video data are accurately extracted, and help analyze complex interactions in audio and video data. It is suitable for the field of multimodal data processing.

[0071] More specifically, by combining the deduced events of the original audio and video data at various time nodes, a causal correlation analysis is performed on the deduced events, and causal correlation parameters between the deduced events at each time node are generated. In multimodal data, the relationship between events is not only parallel, but often causal. By analyzing the causal relationship between the deduced events, the internal logic of the development of events can be revealed. By analyzing the causal relationship, the influence of events and their propagation paths can be better understood, so that reasonable content integration can be carried out in the subsequent processing process, which enhances the depth and accuracy of event understanding, especially can reveal the potential causal relationship in audio and video data, make the audio and video data more orderly and logical in time, and can provide clearer guidance for subsequent editing and editing.

[0072] More specifically, based on the causal correlation parameters, the deduced events of the original audio and video data are connected and simulated at each time node to generate several deduced event information chains. The connection simulation of the deduced events is a key step to ensure that the events at each time node can be naturally connected semantically and logically. By simulating the connection between events, the development context of the events can be better reconstructed. Audio and video data usually contains multiple interrelated events. By simulating the connection of events, the coherence and integrity between events can be better captured, so that the logical relationship between events is further strengthened, ensuring that the event transition in the editing process is smoother and more natural, which helps to generate more plot-coherent editing content and enhance the viewing experience of multi-terminal media platforms.

[0073] More specifically, all deduced event information chains are semantically analyzed as a whole, and the event criticality in the original audio-visual data is evaluated based on the analysis results to generate an event criticality curve. By semantically analyzing the deduced event information chains, the meaning of each information chain can be fully understood, and its importance in the entire audio-visual data can be accurately evaluated. The evaluation of event criticality helps to provide guidance for editing, ensuring that the most critical events can be given priority, improving the relevance of the editing content, and ensuring that the final edited content can accurately reflect the core and key of the event, providing clear quantitative indicators for video editing, and avoiding editing methods with strong subjectivity.

[0074] More specifically, based on the event criticality curve and the causal correlation parameters of the deduced events, the editing value of the deduced events of the audio and video data at each time node is evaluated to generate editing value positioning information. Different time nodes and events have different editing values. Based on the event criticality and causal correlation, the editing value can be evaluated more objectively. This evaluation helps to decide which parts should be retained or reconstructed during the editing process. When evaluating the editing value, the needs of different platforms also need to be considered. For example, social media prefers short and concise content, while long video platforms focus on plot development and details. The editing of video content is optimized to ensure that the key parts of the editing can attract the target audience group to the greatest extent. When distributing content on multiple platforms, more accurate content customization is provided, which enhances the dissemination effect of audio and video data.

[0075] More specifically, based on the editing value positioning information and the causal correlation parameters between the deduced events at each time node, an editing positioning axis is generated. By comprehensively considering the causal relationship and editing value of the events, an editing positioning axis can be generated that can accurately reflect the criticality and value of the events to guide the final editing of the audio-visual data. The generated editing positioning axis can adapt to the needs of different platforms, optimize the presentation form of the final editing, accurately reflect the core events and key parts of the audio-visual data, improve the accuracy of the editing, and provide a quantitative basis for content production on multiple platforms, ensuring that the editing content not only conforms to the characteristics of the target platform, but also meets the viewing needs of users.

[0076] Preferably, the step of performing overall semantic analysis on each of the deduced event information chains, and performing event criticality assessment on the original audio-visual data based on the overall semantic analysis result to obtain an event criticality curve of the original audio-visual data includes: S341: performing overall semantic analysis on each of the deduction event information chains to obtain overall semantic information of each of the deduction event information chains; S342: performing interactive collaborative supervision based on the overall semantic information of each deduction event information chain to obtain the information missing degree and information redundancy degree of each deduction event information chain; S343: Using the information missing degree and the information redundancy as supervision conditions, analyze the expression value of each event in each deduction event information chain to obtain an event value curve for each deduction event information chain; wherein the event value curve is used to describe the expression value of the deduction event at each time node in the deduction event information chain; S344: Based on the event value curves of each of the deduced event information chains, a comprehensive analysis of the event criticality of the deduced events of the original audio and video data at each time node is performed to obtain an event criticality curve of the original audio and video data.

[0077] Specifically, a holistic semantic analysis is performed on each deduction event information chain to obtain the overall semantic information of each deduction event information chain. Holistic semantic analysis is the core of the entire process. Through in-depth analysis of each deduction event information chain, we can understand the meaning of each event in the context and its relevance to other events. When the semantic content of audio and video data is very complex, holistic semantic analysis can help extract the core elements of the information chain and its role in the overall event. By comprehensively analyzing the information chain, we can accurately extract the semantic hierarchy of events and their associations, and improve the system's ability to understand multimodal information. Comprehensive semantic information provides accurate basic data for subsequent missingness, redundancy analysis, and event value analysis.

[0078] More specifically, based on the overall semantic information, interactive collaborative supervision is performed on each deduced event information chain to obtain the information missingness and information redundancy of each deduced event information chain. In multimodal event recognition, information missingness may lead to misunderstanding or mispositioning of events. The evaluation of missingness helps reveal which events are described too briefly or insufficiently. Redundancy refers to repeated and unnecessary parts of information. Through redundancy evaluation, redundant event descriptions can be removed, making the final event information more concise and clear. By combining audio and video information, two-way information supplementation and redundancy elimination are performed, which can more accurately locate information missing and redundancy problems. By reducing redundant information and making up for missing information, the information chain can be ensured to be more concise and accurate, which is conducive to the efficient execution of subsequent steps. Missing and redundancy evaluation can significantly optimize event data, making subsequent event value analysis more reliable.

[0079] More specifically, the degree of information loss and redundancy is used as supervision conditions, and based on these conditions, the expression value of each event in each deduced event information chain is analyzed to obtain the event value curve of each deduced event information chain. In multimodal data processing, not all events are equally important. By evaluating the expression value of events, it is possible to judge from a holistic perspective which events have higher value for the advancement of the story line, emotional expression or other key goals. The expression value analysis can effectively screen out the most representative and meaningful events from the information chain and exclude those parts with low information content or irrelevant parts. The event expression value curve can help editors better identify and extract event parts with higher expression value, thereby optimizing the editing content. By giving priority to retaining events with high expression value, the viewing experience of the editing content can be improved, and the audience's appeal and participation can be enhanced.

[0080] More specifically, based on the event value curve of each deduced event information chain, a comprehensive analysis of the event criticality of the deduced events of the original audio and video data at each time node is performed, thereby generating an event criticality curve for the original audio and video data. The criticality of an event represents the relative importance of an event in the entire audio and video content. A comprehensive analysis of event values ​​in multiple dimensions (including expression value, information missing and redundancy, etc.) can assign a reasonable weight to each event, thereby revealing its criticality in the overall content. Comprehensive consideration of multiple factors (such as information value, temporal importance, plot relevance, etc.) helps to arrive at a more objective and accurate assessment of event criticality. The event criticality curve obtained after comprehensive analysis provides a clear quantitative reference for editing and post-processing, ensuring that key information and events are presented first. Through comprehensive analysis of event criticality, the most important events can be quickly identified and processed, reducing manual screening and repetitive work, and improving editing efficiency.

[0081] More specifically, the generated event criticality curve is based on the above analysis results, and ultimately forms a comprehensive criticality assessment of the original audio and video data. The event criticality curve is a comprehensive indicator that reflects the importance of each event in the video on the timeline, and can provide a clear direction for video editing. Through the event criticality curve, the editing strategy can be dynamically adjusted to prioritize the retention of high-criticality events and delete low-criticality events to optimize the video content. The generation of the criticality curve helps editors make decisions based on data, thereby improving the accuracy of editing. The event criticality curve provides a basis for multi-platform content adaptation, ensuring that the core part of the content can be optimized and displayed on each platform.

[0082] Preferably, the step of performing audio and video clipping and formatting adapted to a multi-terminal media platform on the audio information stream and the video information stream according to the clip positioning axis to generate a clipping file corresponding to the multi-terminal media platform includes: S41: selecting content to be edited from the audio information stream and the video information stream according to the editing positioning axis, thereby cutting and connecting the content to be edited, and obtaining a plurality of initial editing information suitable for the multi-terminal media platforms; wherein the multi-terminal media platforms include streaming media, broadcast television, and mobile terminals; S42: Encapsulating various initial clipping information adapted to the multi-terminal media platforms in corresponding formats according to the respective media platform format requirements of the multi-terminal media platforms to obtain various encapsulated clipping information corresponding to the multi-terminal media platforms; S43: Perform a quality assessment of the image quality loss of each packaged editing information, and send a data compensation request to the original audio and video data based on the quality assessment result, so as to optimize the packaged editing information and generate various editing files corresponding to the multi-terminal media platform.

[0083] S44: Selecting content to be edited that is suitable for the multi-terminal media platform for the audio information stream and the video information stream according to the editing positioning axis, cutting and connecting the content to obtain a plurality of initial editing information that is suitable for the multi-terminal media platform.

[0084] Specifically, the editing positioning axis refers to a reference framework based on time or content, which is used to determine the key events and time nodes in the video and audio. Through this axis, the content to be edited can be accurately selected. Different platforms have different requirements for the length, rhythm, transitions, etc. of the content. According to the editing positioning axis, the video and audio parts can be more effectively connected to ensure the smooth conversion of the overall content between different platforms. By accurately selecting and connecting audio and video content, the presentation effect on different platforms is ensured to be consistent. The editing positioning axis is used to quickly select the content to be edited, reduce manual operations, and improve editing efficiency and accuracy.

[0085] More specifically, according to the format requirements of different multi-terminal media platforms, the initial editing information adapted to the multi-terminal media platforms is packaged in corresponding formats to obtain the packaged editing information corresponding to the multi-terminal media platforms. Streaming media, broadcast television and mobile terminals have their own specifications in terms of video resolution, frame rate, encoding format, audio quality, etc. For example, streaming media platforms usually require higher resolution and compression ratio, while mobile terminals pay more attention to optimizing data traffic. Format packaging is to convert the original video and audio data into a format suitable for playback on a specific platform. According to the needs of different platforms, reasonable format packaging can ensure playback quality and smoothness. Format packaging ensures the compatibility of video and audio content on different platforms, avoiding problems such as playback failure or screen freeze. Adapting to the packaging format of different platforms can enable the content to achieve optimal performance on each platform and ensure user experience.

[0086] More specifically, a quality assessment of image quality loss is performed on each item of packaged editing information, and a data compensation request is sent to the original audio and video data based on the quality assessment result. The packaged editing information is optimized, and editing files corresponding to multi-terminal media platforms are generated. When performing format packaging, especially compression format, video quality loss may occur. Different platforms (such as mobile terminals and large-screen TVs) have different tolerances for image quality loss. Therefore, a quality assessment of image quality loss is required. The quality assessment result provides a basis for data compensation request. The system can automatically optimize based on the quality assessment to ensure that the quality of the final generated video content meets the requirements of the platform. Through quality assessment and data compensation, image quality loss is reduced to ensure that the output video file meets high standards in visual effects. Automatic compensation through quality assessment can adjust video and audio parameters during the format packaging process to optimize the final output effect and reduce human intervention.

[0087] More specifically, all optimized packaged editing information is integrated to eventually generate various editing files suitable for streaming media, broadcast television, and mobile platforms. The final generated editing files need to meet the playback requirements of different platforms to ensure that users on each platform have the best experience when watching content. This step integrates all previous content to ensure that the final file format, quality, timeline, etc. meet the specific needs of the platform. By generating editing files that meet the needs of different platforms, it can ensure smooth playback of content on various terminals. Through multi-terminal adapted editing files, it can provide the best viewing experience and improve user satisfaction.

[0088] Reference Figure 2 As shown, in a second aspect, the present invention provides a video and audio integrated packaging and editing system, which is used to implement the video and audio integrated packaging and editing method described in any one of the first aspects, including: A data separation module is used to separate the original audio and video data into audio data and video data to obtain an audio information stream and a video information stream; A semantic parsing module, configured to perform semantic parsing on the audio information stream and the video information stream using a pre-trained audio-visual information parsing model to generate an audio semantic sequence and a video semantic sequence; An event recognition module, configured to recognize key events in the original audio-visual data based on the audio semantic sequence and the video semantic sequence, so as to generate an editing positioning axis; The file editing module is used to perform audio and video editing and format packaging on the audio information stream and the video information stream according to the editing positioning axis to adapt to the multi-terminal media platform, so as to generate an editing file corresponding to the multi-terminal media platform.

[0089] In this embodiment, for the specific implementation of each module in the above system embodiment, please refer to the above method embodiment, which will not be repeated here.

[0090] In a third aspect, the present invention provides an integrated audio and video packaging and editing device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, and when the processor executes the computer program, it implements an integrated audio and video packaging and editing method as described in any one of the first aspects.

[0091] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A video and audio integrated packaging editing method, characterized in that: include: Separate the original audio and video data into audio data and video data to obtain audio information stream and video information stream; Performing semantic parsing on the audio information stream and the video information stream using a pre-trained audio-visual information parsing model to generate an audio semantic sequence and a video semantic sequence; Identifying key events in the original audiovisual data based on the audio semantic sequence and the video semantic sequence to generate an editing positioning axis; The audio information stream and the video information stream are subjected to audio and video editing and format packaging adapted to a multi-terminal media platform according to the editing positioning axis to generate an editing file corresponding to the multi-terminal media platform.

2. The video and audio integrated packaging and editing method according to claim 1, characterized in that: The steps of separating the audio data and the video data from the original audio and video data to obtain the audio information stream and the video information stream include: Acquire original audio and video data, and perform audio waveform decoding and video frame decoding on the original audio and video data respectively, to obtain audio waveform decoding data and video frame decoding data disassembled and separated from the original audio and video data; In the process of decoding the original audio and video data, synchronously generating a timestamp mark, and assigning the timestamp mark of each moment to the audio waveform decoded data and the video frame decoded data; A unified timeline is constructed based on the timestamp mark of each moment, and information mapping is performed on the audio waveform decoded data and the video frame decoded data relative to the unified timeline, so as to disassemble the audio waveform decoded data into a plurality of audio metadata and assign them to the unified timeline to obtain an audio information stream, and disassemble the video information stream into a plurality of video metadata and assign them to the unified timeline to obtain a video information stream.

3. The video and audio integrated packaging editing method according to claim 2, characterized in that: The steps of performing semantic parsing on the audio information stream and the video information stream by using a pre-trained audio-visual information parsing model to generate an audio semantic sequence and a video semantic sequence include: Inputting the audio information stream and the video information stream into a pre-trained audio-visual information parsing model, so as to perform information stream semantic parsing of the audio information stream and the video information stream using multiple parsing methods by the audio-visual information parsing model, and obtaining semantic parsing information of the audio information stream and the video information stream corresponding to various parsing methods; According to the semantic parsing information corresponding to each parsing method of the audio information stream, the semantic parsing information of each audio metadata on the unified time axis is fused to obtain audio semantic units corresponding to each audio metadata and assign them to the unified time axis to obtain an audio semantic sequence; According to the semantic parsing information of the video information stream corresponding to various parsing methods, the semantic parsing information of each video metadata on the unified timeline is fused to obtain the video semantic unit corresponding to each video metadata and assign it to the unified timeline to obtain a video semantic sequence.

4. The video and audio integrated packaging and editing method according to claim 3, characterized in that: The parsing method of the audio-visual information parsing model includes a metadata timing parsing method, and the steps of performing information flow semantic parsing of the audio information stream and the video information stream in the metadata timing parsing method using the audio-visual information parsing model include: Assigning parsing sequence numbers to each audio metadata of the audio information stream and each video metadata of the video information stream according to the information correspondence between the audio information stream and the video information stream relative to the unified time axis; According to the parsing sequence number of each audio metadata and each video metadata, semantic parsing and vectorization expression are performed on each audio metadata and each video metadata in sequence to generate a semantic parsing feature matrix of each audio metadata and each video metadata; Arranging the semantic parsing feature matrices corresponding to the audio metadata and the video metadata according to the parsing sequence numbers, and interactively verifying the semantic parsing feature vectors of the sequentially arranged semantic parsing feature matrices to assign confidence information to the semantic parsing feature matrices corresponding to the audio metadata and the video metadata; The confidence information of each semantic parsing feature matrix is ​​used to collaboratively analyze each semantic parsing feature matrix to extract a unique semantic parsing feature vector from each semantic parsing feature matrix as the semantic parsing information of the metadata timing parsing method corresponding to the audio information stream and the video information stream.

5. The video and audio integrated packaging and editing method according to claim 3, characterized in that: The parsing method of the audio-visual information parsing model includes an object-based parsing method. The steps of performing information flow semantic parsing of the audio information flow and the video information flow in the object-based parsing method using the audio-visual information parsing model include: Assigning parsing sequence numbers to each audio metadata of the audio information stream and each video metadata of the video information stream according to the information correspondence between the audio information stream and the video information stream relative to the unified time axis; Performing content positioning of data feedback objects on each audio metadata and each video metadata to obtain feedback object information of each audio metadata and each video metadata; performing object summarization processing on the audio information stream and the video information stream according to the feedback object information of each audio metadata and each video metadata, so as to connect each audio metadata having feedback object information that meets a preset standard with each video metadata, thereby generating a plurality of audio sub-information streams and a plurality of video sub-information streams; Analyzing the object feedback frequency and object feedback time distribution of each audio sub-information stream and each video sub-information stream according to the parsed sequence numbers corresponding to each audio sub-information stream and each video sub-information stream, and assigning an object weight curve to each audio sub-information stream and each video sub-information stream based on the analysis results; wherein the object weight curve is used to describe the object feedback weight of the sub-information stream in each time interval; Performing a multi-dimensional comparative analysis on the object weight curves of each of the audio sub-information streams and each of the video sub-information streams to obtain semantic parsing direction features of each time interval of each of the audio sub-information streams and each of the video sub-information streams; According to the semantic parsing pointing features of each time interval of each audio sub-information stream and each video sub-information stream, semantic parsing of each audio metadata and each video metadata is performed on each audio sub-information stream and each video sub-information stream respectively, to serve as semantic parsing information of the summary parsing method of the corresponding objects of the audio information stream and the video information stream.

6. The video and audio integrated packaging and editing method according to claim 1, characterized in that: The step of identifying key events of the original audiovisual data based on the audio semantic sequence and the video semantic sequence to generate an editing positioning axis includes: Performing dual event positioning on the original audio-visual data based on the audio semantic sequence and the video semantic sequence to obtain deduced events of the original audio-visual data at various time nodes; Combining the deduced events of the original video and audio data at each time node, performing a causal correlation analysis on the deduced events at each time node to generate causal correlation parameters between the deduced events at each time node; According to the causal correlation parameters between the deduced events at each time node, the deduced events of the original audio and video data at each time node are connected and simulated to obtain a plurality of deduced event information chains; Performing overall semantic analysis on each of the deduced event information chains, and performing event criticality assessment on the original audio-visual data based on the overall semantic analysis result, so as to obtain an event criticality curve of the original audio-visual data; Based on the causal correlation parameters between the event criticality curve and the deduced events at each time node, the editing value of the deduced events of the original audio and video data at each time node is evaluated relative to the multi-terminal media platform to obtain the editing value positioning information of the original audio and video data at each time node, which together constitute the editing positioning axis of the original audio and video data.

7. The video and audio integrated packaging and editing method according to claim 6, characterized in that: The steps of performing overall semantic analysis on each of the deduced event information chains, and evaluating the event criticality of the original audio-visual data based on the overall semantic analysis results to obtain an event criticality curve for the original audio-visual data include: Performing overall semantic analysis on each of the deduction event information chains to obtain overall semantic information of each of the deduction event information chains; Based on the overall semantic information of each deduction event information chain, interactive collaborative supervision is performed to obtain the information missing degree and information redundancy degree of each deduction event information chain; Using the information missing degree and the information redundancy as supervision conditions, the expression value of each event on each deduction event information chain is analyzed to obtain an event value curve for each deduction event information chain; wherein the event value curve is used to describe the expression value of the deduction event at each time node on the deduction event information chain; According to the event value curve of each deduced event information chain, a comprehensive analysis of the event criticality of the deduced events of the original audio and video data at each time node is performed to obtain the event criticality curve of the original audio and video data.

8. The video and audio integrated packaging and editing method according to claim 1, wherein: The steps of performing audio and video clipping and formatting of the audio information stream and the video information stream adapted to the multi-terminal media platform according to the clip positioning axis to generate a clipping file corresponding to the multi-terminal media platform include: Selecting content to be edited from the audio information stream and the video information stream according to the editing positioning axis, so as to cut and connect the content to be edited, thereby obtaining a plurality of initial editing information suitable for the multi-terminal media platform; wherein the multi-terminal media platform includes streaming media, broadcast television, and mobile terminals; According to the respective media platform format requirements of the multiple media platforms, the various initial clipping information adapted to the multiple media platforms is encapsulated in corresponding formats to obtain the various encapsulated clipping information corresponding to the multiple media platforms; Perform a quality assessment of image quality loss on each packaged editing information, and send a data compensation request to the original audio and video data based on the quality assessment result, so as to optimize the packaged editing information and generate various editing files corresponding to the multi-terminal media platform.

9. An integrated video and audio packaging and editing system, characterized in that: A method for implementing the video and audio integrated packaging and editing method according to any one of claims 1 to 8, comprising: A data separation module is used to separate the original audio and video data into audio data and video data to obtain an audio information stream and a video information stream; A semantic parsing module, configured to perform semantic parsing on the audio information stream and the video information stream using a pre-trained audio-visual information parsing model to generate an audio semantic sequence and a video semantic sequence; An event recognition module, configured to recognize key events in the original audio-visual data based on the audio semantic sequence and the video semantic sequence, so as to generate an editing positioning axis; The file editing module is used to perform audio and video editing and format packaging on the audio information stream and the video information stream according to the editing positioning axis to adapt to the multi-terminal media platform, so as to generate an editing file corresponding to the multi-terminal media platform.

10. An integrated video and audio packaging and editing device, comprising a memory and a processor, wherein the memory stores a computer program that can be run on the processor, characterized in that: When the processor executes the computer program, the video and audio integrated packaging editing method described in any one of claims 1-8 is implemented.

Citation Information

Patent Citations

  • Audio-visual video analysis device and method based on multi-scale semantic network

    CN114519809A

  • Video generation method, video generation device, electronic equipment and storage medium

    CN114786059A

  • Video clip positioning system based on space-time semantic decomposition

    CN115309939A

  • Video post-editing and video synthesis optimization method

    CN116847123A

  • Video data processing method and device, computer equipment and storage medium

    CN117579858A

Cited By

  • Micro-wave audio data intelligent processing and storage optimization system

    CN121256083A