Fast exporting method and device for video stream cutting and merging and related medium

By only decoding and encoding key frames during the video cutting and merging process, the low efficiency problem caused by the need to decode and encode all frames in the existing technology of video cutting and merging is solved, efficient video cutting and merging processing is achieved, and the export speed and processing efficiency are improved.

CN120751219APending Publication Date: 2025-10-03AFIRSTSOFT CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511007084.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-22
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

In the existing technology, all frames need to be decoded and encoded during the video cutting and merging process, resulting in long processing time and low efficiency, especially when processing high-definition or long-duration videos.

Method used

By obtaining a video processing request, determining the start and end time points of the target video segment, locating key frames, building a mapping relationship, standardizing the target video segment and correcting its timestamp, extracting the frames to be encoded and copied in segments, decoding and encoding only the key frames, retaining the original encoding structure of other frames, and building a continuously increasing timestamp sequence, the video export result is finally merged.

Benefits of technology

It significantly reduces computing resource consumption and processing time, and improves the efficiency of video cutting and merging. Especially when a single cut segment is long or multiple segments are processed, the export speed can be increased by one to dozens of times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120751219A_ABST
    Figure CN120751219A_ABST
Patent Text Reader

Abstract

The invention discloses a video stream cutting and merging rapid exporting method and device and a related medium, and the method comprises the steps: obtaining a video processing request, and determining the starting time point and the ending time point of a plurality of target video clips; constructing a mapping relation based on key frame positioning; performing standardization processing and timestamp correction on the video clip, and extracting a to-be-coded frame and a to-be-copied frame; decoding and recoding the to-be-coded frame, and copying the to-be-copied frame by keeping the original coding structure; uniformly adjusting timestamps of all frames and constructing a continuous timestamp sequence; and finally, combining the target coding frame with the target copy frame to obtain a video export result. Compared with the prior art, the method has the advantages that only the key position frames are decoded and coded, so that the resource consumption and the processing time are reduced, the efficient export of the video in the cut and merged scene is realized, and the export efficiency of the video is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of multimedia processing technology, and in particular to a method and device for quickly exporting video stream cutting and merging, and related media. Background Art

[0002] In existing video editing applications, users often need to cut and merge raw videos, for example, extracting specific segments from a video or merging multiple similar segments into a complete video file. To achieve these operations, existing technologies typically use a "decode-edit-re-encode" process, decoding each frame within the target segment and then re-encoding it into a new video stream after the editing process is complete.

[0003] However, this approach has significant technical drawbacks: the comprehensive decoding and re-encoding process consumes significant computing resources and takes a long time to process, which can lead to inefficient exporting, especially for high-definition or long-duration videos. In actual cutting and merging scenarios, most video processing involves only frame-level cropping and splicing, and does not involve complex image processing such as image scaling and filter transformations. Therefore, re-encoding all video frames is unnecessary. If traditional processing methods are still used, export efficiency will be significantly reduced. Summary of the Invention

[0004] The embodiments of the present invention provide a method, device and related media for quickly exporting video stream cutting and merging, aiming to solve the problem in the prior art that video cutting and merging requires decoding and encoding all frames before exporting, resulting in low export efficiency.

[0005] In a first aspect, an embodiment of the present invention provides a method for quickly exporting video streams by cutting and merging, comprising:

[0006] Obtain a video processing request and determine the start and end time points of multiple target video segments;

[0007] Based on the starting time point and the ending time point, performing key frame positioning processing on the target video segment to establish a mapping relationship;

[0008] Performing standardization processing on the target video clip and performing timestamp correction to obtain video frame data;

[0009] Segmenting the video frame data based on the mapping relationship to extract frames to be encoded and frames to be copied respectively;

[0010] Decoding the frame to be encoded to obtain a valid image frame, and encoding the valid image frame to generate a target encoded frame;

[0011] Performing a copy process on the frame to be copied, retaining the original coding structure, and generating a target copy frame;

[0012] Standardizing the timestamp information of the target coded frame and the target copied frame respectively to construct a continuously increasing timestamp sequence;

[0013] After the processing of the plurality of target video segments is completed, the target coded frame and the target copied frame are merged based on the timestamp sequence to obtain a video export result.

[0014] In a second aspect, an embodiment of the present invention provides a fast exporting device for cutting and merging video streams, comprising:

[0015] A node acquisition unit, configured to acquire a video processing request and determine the start time points and end time points of multiple target video segments;

[0016] a node positioning unit, configured to perform key frame positioning processing on the target video segment based on the starting time point and the ending time point to establish a mapping relationship;

[0017] A video processing unit, configured to perform standardization processing on the target video segment and perform timestamp correction to obtain video frame data;

[0018] A segmentation processing unit, configured to perform segmentation processing on the video frame data based on the mapping relationship, and extract frames to be encoded and frames to be copied respectively;

[0019] a video encoding unit, configured to decode the frame to be encoded to obtain a valid image frame, and perform encoding processing based on the valid image frame to generate a target encoded frame;

[0020] A video copying unit, configured to copy the frame to be copied, retain the original coding structure, and generate a target copy frame;

[0021] A sequence construction unit, configured to perform standardization processing on the timestamp information of the target coded frame and the target copied frame respectively, and to construct a continuously increasing timestamp sequence;

[0022] The video merging unit is configured to merge the target coded frame and the target copied frame based on the timestamp sequence after the processing of the plurality of target video segments is completed, so as to obtain a video export result.

[0023] In a third aspect, an embodiment of the present invention provides a computer device comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method for rapidly exporting video stream cutting and merging according to the first aspect is implemented.

[0024] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium, wherein a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the method for quickly exporting video stream cutting and merging according to the first aspect is implemented.

[0025] An embodiment of the present invention provides a fast export method for video stream cutting and merging, comprising obtaining a video processing request and determining the start and end time points of multiple target video segments; performing key frame positioning processing on the target video segments based on the start and end time points to establish a mapping relationship; performing standardization processing on the target video segments and performing timestamp correction to obtain video frame data; performing segmentation processing on the video frame data based on the mapping relationship to extract frames to be encoded and frames to be copied respectively; decoding the frames to be encoded to obtain valid image frames, and encoding processing based on the valid image frames to generate target encoded frames; copying the frames to be copied, retaining the original encoding structure, to generate target copied frames; performing standardization processing on the timestamp information of the target encoded frames and target copied frames respectively to establish a continuously increasing timestamp sequence; after the processing of the multiple target video segments is completed, merging the target encoded frames with the target copied frames based on the timestamp sequence to obtain a video export result. The present invention reduces resource consumption and processing time by only decoding and encoding frames at key positions, thereby achieving efficient video export in cutting and merging scenarios and improving export speed and processing efficiency.

[0026] The embodiment of the present invention also provides a fast exporting device, computer equipment and storage medium for cutting and merging video streams, which also have the above-mentioned beneficial effects. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0028] Figure 1 A schematic diagram of a process flow of a fast export method for cutting and merging video streams provided by an embodiment of the present invention;

[0029] Figure 2 A schematic block diagram of a fast exporting device for cutting and merging video streams provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0031] It will be understood that when used in this specification and the appended claims, the terms “comprises” and “comprising” indicate the presence of described features, integers, steps, operations, elements and / or components, but do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.

[0032] It should also be understood that the terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the present invention. As used in the specification and appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms unless the context clearly indicates otherwise.

[0033] It should be further understood that the term "and / or" used in the present description and the appended claims refers to and includes any and all possible combinations of one or more of the associated listed items.

[0034] See below Figure 1 , Figure 1 A flowchart of a fast export method for cutting and merging video streams provided by an embodiment of the present invention specifically includes steps S101 to S108.

[0035] S101, obtaining a video processing request, and determining the start time points and end time points of multiple target video segments;

[0036] S102: performing key frame positioning processing on the target video segment based on the starting time point and the ending time point to establish a mapping relationship;

[0037] S103, performing standardization processing on the target video segment and performing timestamp correction to obtain video frame data;

[0038] S104, segmenting the video frame data based on the mapping relationship to extract frames to be encoded and frames to be copied respectively;

[0039] S105, decoding the frame to be encoded to obtain a valid image frame, and encoding the valid image frame to generate a target encoded frame;

[0040] S106, performing a copy process on the frame to be copied, retaining the original coding structure, and generating a target copy frame;

[0041] S107, respectively standardizing the timestamp information of the target coded frame and the target copied frame to construct a continuously increasing timestamp sequence;

[0042] S108: After the processing of the plurality of target video segments is completed, the target coded frame and the target copied frame are merged based on the timestamp sequence to obtain a video export result.

[0043] In step S101, a video processing request is obtained, and the start time points and end time points of multiple target video segments are determined. The video processing request can be initiated by a user in video editing software, and specifies several target video segments to be cut or merged.

[0044] In step S102, the target video segment is decapsulated by calling a preset video processing interface (such as FFmpegAPI), and key frame positioning is performed at each time point to identify the position of the key frame (I frame) of the video group (GOP) in which it is located, thereby establishing a correspondence between the cutting time point and the key frame.

[0045] In one embodiment, step S102 includes:

[0046] Decapsulate the target video clip using a preset video processing interface to extract corresponding video stream information;

[0047] Taking the starting time point or the ending time point as an input parameter, and calling a jump positioning function to perform jump positioning on the video stream information;

[0048] After the jump position, read the timestamp of the video frame;

[0049] Determine whether the timestamp of the video frame is greater than the start time point or the end time point; if not, use the current frame as a key frame of the video group corresponding to the time point;

[0050] If yes, the time point is offset adjusted and the adjusted time point is re-jumped until the preset maximum jump limit is met or the offset time is zero, and the key frame is output;

[0051] A mapping relationship between the start time point and the end time point and the corresponding key frame position is established.

[0052] In this embodiment, a preset video processing interface (such as FFmpegAPI) is first used to decapsulate the target video segment and extract the corresponding video stream information; then, the start time point or the end time point is used as an input parameter, and a jump positioning function (such as av_seek_frame) is called to perform a jump operation (seek) on the video stream information, so that the decoder is positioned at a frame position close to that time point.

[0053] Furthermore, the video frames are read frame by frame starting from the current position, and it is determined whether the read frame is a key frame (I frame). If it is a key frame, its display timestamp (PTS) is further compared to see whether it is less than or equal to the input time point. If so, the I frame is the starting key frame position of the video group (GOP) at the current time point. If the PTS is greater than the current time point, the time point needs to be offset, that is, the set time offset (offset = number of jumps × 1 second) is subtracted, and then the jump operation is performed again. This process can be repeated until the maximum number of jumps (such as 5 times) is met or the offset time is zero. After identifying the key frame position corresponding to the starting time point or the ending time point, a mapping relationship can be established between the time point and the corresponding key frame. For scenes where the first and last time points of the target video clip cover the complete video content, there is no need to construct a key frame mapping relationship.

[0054] In step S103, the encoding format of the video frame is unified (for example, the H264 / H265 frame is converted from the AVCC format to the Annex B format), and the timestamp value of the audio and video stream is subtracted from the minimum start time of all streams and converted into a unified time unit (such as milliseconds) to ensure the consistency and comparability of subsequent timestamps.

[0055] In one embodiment, step S103 includes:

[0056] Decapsulating the target video segment and converting the video frames in the target video segment into a standard format;

[0057] After completing the format conversion, calculating a minimum start timestamp value based on the target video segment;

[0058] The minimum starting timestamp value is used to perform timestamp correction on the target video segment to obtain video frame data.

[0059] In this embodiment, the target video clip is decapsulated and the video frame is extracted. During the decapsulation process, for the extracted video frames, if their encoding format is H264 or H265 and their encapsulation structure is AVCC format, they need to be uniformly converted to Annex B format. This conversion operation can be completed by calling a preset video stream bitstream filter (such as FFmpeg bitstream filter) to ensure that the video frame format is consistent with the output format required by the encoder. At the same time, some encapsulation formats only support the writing of Annex B format frames, and the conversion operation can improve subsequent write compatibility.

[0060] After format conversion is complete, the start timestamp information of the audio and video streams of the target video clip is uniformly analyzed to calculate the minimum start timestamp value. This timestamp usually refers to the presentation timestamp (PTS) or decoding timestamp (DTS) of the frame and serves as a unified calibration benchmark. This minimum start timestamp value is used to calibrate the timestamp information of all video frames. This means that the minimum value is subtracted from the original timestamp of each frame, ensuring that the timestamps of all frames increment from a unified starting point. The time unit is then converted to a unified unit (such as milliseconds), thereby obtaining standardized video frame data.

[0061] In step S104, the video frames are read in chronological order, their timestamps and frame types are compared, and combined with the key frame position, it is determined whether the current frame is in the cut start GOP, the cut end GOP, or in between, and thus marked as a frame to be encoded or a frame to be copied.

[0062] In one embodiment, the step S104 includes:

[0063] Read the video frame data sequentially from front to back, and compare the key frame position recorded in the mapping relationship based on the timestamp and frame type information of each frame;

[0064] When the read frame is located before the key frame of the video group where the starting time point is located, discarding the read frame;

[0065] When a key frame of the video group at the starting time point is read, all frames of the video group are marked as frames to be encoded; or, when a key frame of the video group at the ending time point is read, the video group and its subsequent frames are marked as frames to be encoded;

[0066] The frames of the video group that are between two key frames and do not need to be encoded are marked as frames to be copied.

[0067] In this embodiment, based on the mapping relationship between the key frame positions of the video group (GOP) corresponding to the start time point and the end time point constructed in the analysis phase, it is determined whether a mapping relationship currently exists. If a mapping relationship exists, only the mapped key frame area is decoded and encoded, and the remaining frames are copied in the original frame manner; if a mapping relationship does not exist, the entire cut segment (i.e., the target video segment, the same below) is considered to be equivalent to the complete video content, and all frames are copied in the original manner. In the case of a mapping relationship, the processing process for each cut segment is as follows:

[0068] Read the video frame data from front to back and extract the presentation timestamp (PTS) and frame type (I-frame, P-frame, B-frame) of each frame. Then compare the PTS of the current frame with the key frame position in the mapping relationship to determine whether the frame belongs to the starting GOP, ending GOP or middle area of ​​the cut segment. The specific determination includes the following five sub-steps:

[0069] The first step is frame discarding: frames that have not yet reached the key frame position of the GOP where the cutting start time point is located are directly discarded without any processing;

[0070] The second step is decoding the cut start point: When the I frame of the GOP where the cut start time point is located is detected, all frames in the GOP are sent to the decoder for decoding. For the decoded image frame, its PTS is further determined to see if it is within the cut segment range. If it is within the range, it is retained as a frame to be encoded; otherwise, it is discarded. The retained image frame is then input to the encoder for encoding.

[0071] Step 3: Intermediate segment copying: When the next I-frame is read (i.e., the start point of the next GOP after the cut-start GOP), the decoder is first refreshed. All image frames that have not yet been output by the decoder in the previous stage but whose timestamps are still in the cut segment area are sent to the encoder for encoding. Subsequently, all subsequent frames starting from this I-frame are no longer decoded and are instead treated as frames to be copied, that is, the original encoding structure is retained and directly written to the output.

[0072] Step 4: Decoding the end point of the cut: When reading the I frame of the GOP where the cut ends, the I frame and subsequent frames will be sent to the decoder for decoding again to ensure that the end boundary of the cut segment is completely restored.

[0073] Step 5: End segment refresh processing: When the PTS of the frame output by the decoder is greater than the end time point of the cut, the decoder needs to be refreshed, and the image frames still in the cut segment area will continue to be sent to the encoder to complete the encoding, and the encoder status will be refreshed to mark the completion of the processing flow of the cut segment.

[0074] Through the above steps, the frames in the cut segment can be effectively divided into frames to be re-encoded and frames to be copied that can be directly reused, ensuring the editing accuracy of the start and end boundaries and greatly improving the overall export efficiency.

[0075] In step S105, the decoder is first called to decode the frames to be encoded, and valid image frames whose timestamps fall within the range of the cut segment are screened out. Then, the frames are input to the encoder and re-encoded based on the encoding parameter configuration consistent with the original video.

[0076] In one embodiment, step S105 includes:

[0077] Decoding the frame to be encoded using a decoder to obtain a valid image frame;

[0078] Determining whether the timestamp of the valid image frame is within the time range of the target video segment, and if not, discarding the valid image frame;

[0079] If so, the valid image frame is input to the encoder for encoding processing to generate a target encoding frame.

[0080] In this embodiment, a preset decoder is used to decode the frames marked as to be encoded to obtain valid image frames (such as YUV frames, etc.). For each valid image frame obtained by decoding, it is determined whether its display timestamp (PTS) is within the time range of the target video clip. If it is not within the range, it means that the frame is invalid, does not participate in the clipping output, and is discarded; if it is within the cutting time period, the image frame is considered valid and is sent to the encoder for encoding. The encoding operation is set based on a configuration consistent with the source video encoding parameters to ensure that the generated target encoded frame is consistent with the source video in terms of format, resolution, encoding type, etc., so as to achieve seamless merging of the final video stream. The target encoded frame finally output will participate in the timestamp unified processing and export operation together with the target copy frame.

[0081] In one embodiment, inputting the valid image frame into an encoder for encoding to generate a target coded frame includes:

[0082] Obtaining encoding parameter information of the target video segment;

[0083] Constructing an encoder configuration file according to the encoding parameter information, and setting the parameters of the encoder using the encoder configuration file;

[0084] After the parameters of the encoder are set, the global header information flag item of the encoder is turned off, and the valid image frame is input into the encoder to generate a target coded frame.

[0085] In this embodiment, encoding parameter information corresponding to the target video segment is obtained. The parameter information generally includes encoding ID, resolution, frame format, color space information, encoding profile, and encoding level. The encoding parameters can be obtained by analyzing the encoding structure of the source video or reading the video header information, aiming to ensure that the target encoded frame generated by re-encoding is basically consistent with the source video in encoding characteristics. Furthermore, based on the above encoding parameter information, an encoder configuration file is constructed, and the encoder is initialized using the configuration file. During the setting process, special attention should be paid to setting the time base (timebase) to a globally universal unit time base. The time base should be consistent with the time unit used in the clip processing process, so as to ensure that the display timestamp (PTS) of the encoded output frame is aligned with the timestamp of the original input frame, thereby ensuring the accuracy and consistency of the exported video time sequence.

[0086] After completing the encoder parameter configuration, disable the encoder's global header information flag (GLOBAL HEADER option) to ensure that the output video stream contains the necessary encoding description information, such as the SPS (Sequence Parameter Set), in each I-frame. This setting is particularly critical for video frames using video coding standards such as H264 / H265, ensuring decoder compatibility and independent decoding capabilities of the exported video. Finally, input the valid image frame into the configured encoder, perform the encoding operation, and generate the target coded frame. This target coded frame will be used in the subsequent timestamp processing and segment merging steps, ultimately forming the key part of the exported video.

[0087] In step S106 , for video frames that do not need to be re-encoded, their original binary data is directly retained to avoid unnecessary decoding and encoding processes, thereby improving processing efficiency.

[0088] In step S107, the frame timestamps in each segment are uniformly started from zero, and the decoding timestamp (dts) is ensured to meet the monotonically increasing requirement. If not, it is adjusted to increase.

[0089] In one embodiment, step S107 includes:

[0090] subtracting the display timestamp of the first frame of the target video segment from the display timestamp of the current frame of the target video segment, so that the display timestamp of the target video segment increases from zero, to obtain a timestamp sequence;

[0091] Alternatively, subtracting the decoding timestamp of the first frame of the target video segment from the decoding timestamp of the current frame of the target video segment, so that the decoding timestamp of the target video segment increments from zero;

[0092] Determine whether the adjusted decoding timestamp meets the monotonically increasing condition; if not, set the decoding timestamp of the current frame of the target video to the decoding timestamp of the previous output frame plus one; if so, output a timestamp sequence.

[0093] In this embodiment, the minimum value of the start time of all audio and video streams in the file is subtracted from the original presentation timestamp (PTS) and decoding timestamp (DTS) as the time reference point, thereby normalizing the timestamps of the audio and video streams. At the same time, the original time unit is converted into a standard unit (such as milliseconds) to ensure the time base consistency of multi-stream data and lay the foundation for subsequent timestamp calculations. At each target video segment processing stage, the timestamp values ​​of all frames processed and output in the segment are uniformly incremented from zero. Specifically, the PTS value of the current frame is subtracted from the PTS value of the first frame in the cut segment, or the DTS value of the current frame is subtracted from the DTS value of the first frame in the cut segment, thereby constructing a standardized timestamp sequence within the cut segment.

[0094] In order to further ensure the increasing continuity of the output frame timestamp, the DTS value needs to be judged for monotonicity. If the DTS value of a frame is less than or equal to the DTS value of the previous output frame and does not meet the increasing condition, the DTS value of the frame is set to the DTS value of the previous frame plus one, ensuring that the decoding timestamps of all frames are strictly monotonically increasing, thereby avoiding disorder or abnormal decoder behavior during video playback. Finally, after multiple cut segments have independently completed timestamp standardization, the timestamps need to be adjusted for overall continuity during the segment merging stage. Specifically, for the frame in the currently processed segment, its PTS / DTS value is added to the total duration of all segments before the current segment, thereby forming a complete, continuous, and non-overlapping global timestamp sequence after all segments are spliced ​​together.

[0095] In step S108, if multiple clips come from different files, it is necessary to first ensure the consistency of the video stream parameters (such as resolution, encoding format, etc.), and then complete the export and merging in the order of the clips.

[0096] The embodiment of the present invention accurately identifies key frames in the video cutting and merging process, and only decodes and encodes the video groups at the beginning and end of the cut segment, while directly copying the original encoded frames for the video frames in the middle part, avoiding the high-overhead operation of decoding and re-encoding all video frames, and significantly reducing computing resource consumption and processing time.

[0097] Furthermore, when processing multiple clips, the present invention prioritizes consistency checks for parameters (including resolution, encoding format, color space, rotation information, interlacing information, pixel aspect ratio, etc.) across the video streams to which each clip belongs. Only when these parameters are consistent will a unified export be performed, effectively ensuring the playback compatibility and structural stability of the output file. By continuously adjusting the processed frame timestamps (PTS / DTS) of each clip, multiple clips can be seamlessly merged into a single output file, ensuring timeline coherence.

[0098] Based on the above technical solutions, the present invention achieves efficient processing of video cutting and merging tasks on the basis of ensuring video editing accuracy and playback quality. Especially when a single cut segment is long or multiple segments are processed, the export speed can be increased by one to dozens of times compared with the traditional full transcoding solution, which has significant performance advantages and engineering application value.

[0099] Combine Figure 2 As shown, Figure 2 A schematic block diagram of a fast exporting apparatus for cutting and merging video streams provided by an embodiment of the present invention is provided. The fast exporting apparatus 200 for cutting and merging video streams includes:

[0100] The node acquisition unit 201 is used to obtain a video processing request and determine the start time point and the end time point of multiple target video segments;

[0101] A node positioning unit 202 is configured to perform key frame positioning processing on the target video segment based on the start time point and the end time point to establish a mapping relationship;

[0102] The video processing unit 203 is configured to perform standardization processing on the target video segment and perform timestamp correction to obtain video frame data;

[0103] A segmentation processing unit 204 is configured to segment the video frame data based on the mapping relationship to extract frames to be encoded and frames to be copied respectively;

[0104] The video encoding unit 205 is configured to decode the to-be-encoded frame to obtain a valid image frame, and perform encoding processing based on the valid image frame to generate a target coded frame;

[0105] The video copying unit 206 is configured to copy the frame to be copied, retain the original coding structure, and generate a target copy frame;

[0106] A sequence construction unit 207 is configured to perform standardization processing on the time stamp information of the target coded frame and the target copied frame respectively, and construct a continuously increasing time stamp sequence;

[0107] The video merging unit 208 is configured to merge the target coded frame and the target copied frame based on the timestamp sequence after the processing of the plurality of target video segments is completed, to obtain a video export result.

[0108] In this embodiment, the node acquisition unit 201 obtains a video processing request and determines the starting time point and the ending time point of multiple target video segments; the node positioning unit 202 performs key frame positioning processing on the target video segments based on the starting time point and the ending time point to establish a mapping relationship; the video processing unit 203 performs standardization processing on the target video segments and performs timestamp correction to obtain video frame data; the segmentation processing unit 204 performs segmentation processing on the video frame data based on the mapping relationship, and extracts the frames to be encoded and the frames to be copied respectively; the video encoding unit 205 performs segmentation processing on the target video segments and extracts the frames to be encoded and the frames to be copied respectively; the video encoding unit 205 performs segmentation processing on the target video segments and extracts the frames to be encoded and the frames to be copied respectively. The frame to be encoded is decoded to obtain a valid image frame, and encoding is performed based on the valid image frame to generate a target encoded frame; the video copying unit 206 copies the frame to be copied, retains the original encoding structure, and generates a target copied frame; the sequence construction unit 207 standardizes the timestamp information of the target encoded frame and the target copied frame respectively, and constructs a continuously increasing timestamp sequence; after the processing of multiple target video clips is completed, the video merging unit 208 merges the target encoded frame with the target copied frame based on the timestamp sequence to obtain a video export result.

[0109] In one embodiment, the node location unit 202 is specifically configured to:

[0110] Decapsulate the target video clip using a preset video processing interface to extract corresponding video stream information;

[0111] Taking the starting time point or the ending time point as an input parameter, and calling a jump positioning function to perform jump positioning on the video stream information;

[0112] After the jump position, read the timestamp of the video frame;

[0113] Determine whether the timestamp of the video frame is greater than the start time point or the end time point; if not, use the current frame as a key frame of the video group corresponding to the time point;

[0114] If yes, the time point is offset adjusted and the adjusted time point is re-jumped until the preset maximum jump limit is met or the offset time is zero, and the key frame is output;

[0115] A mapping relationship between the start time point and the end time point and the corresponding key frame position is established.

[0116] In one embodiment, the video processing unit 203 is specifically configured to:

[0117] Decapsulating the target video segment and converting the video frames in the target video segment into a standard format;

[0118] After completing the format conversion, calculating a minimum start timestamp value based on the target video segment;

[0119] The minimum starting timestamp value is used to perform timestamp correction on the target video segment to obtain video frame data.

[0120] In one embodiment, the segment processing unit 204 is specifically configured to:

[0121] Read the video frame data sequentially from front to back, and compare the key frame position recorded in the mapping relationship based on the timestamp and frame type information of each frame;

[0122] When the read frame is located before the key frame of the video group where the starting time point is located, discarding the read frame;

[0123] When a key frame of the video group at the starting time point is read, all frames of the video group are marked as frames to be encoded; or, when a key frame of the video group at the ending time point is read, the video group and its subsequent frames are marked as frames to be encoded;

[0124] The frames of the video group that are between two key frames and do not need to be encoded are marked as frames to be copied.

[0125] In one embodiment, the video encoding unit 205 is specifically configured to:

[0126] Decoding the frame to be encoded using a decoder to obtain a valid image frame;

[0127] Determining whether the timestamp of the valid image frame is within the time range of the target video segment, and if not, discarding the valid image frame;

[0128] If so, the valid image frame is input to the encoder for encoding processing to generate a target encoding frame.

[0129] In one embodiment, the video encoding unit 205 is further configured to:

[0130] Obtaining encoding parameter information of the target video segment;

[0131] Constructing an encoder configuration file according to the encoding parameter information, and setting the parameters of the encoder using the encoder configuration file;

[0132] After the parameters of the encoder are set, the global header information flag item of the encoder is turned off, and the valid image frame is input into the encoder to generate a target coded frame.

[0133] In one embodiment, the sequence construction unit 207 is specifically configured to:

[0134] subtracting the display timestamp of the first frame of the target video segment from the display timestamp of the current frame of the target video segment, so that the display timestamp of the target video segment increases from zero, to obtain a timestamp sequence;

[0135] Alternatively, subtracting the decoding timestamp of the first frame of the target video segment from the decoding timestamp of the current frame of the target video segment, so that the decoding timestamp of the target video segment increments from zero;

[0136] Determine whether the adjusted decoding timestamp meets the monotonically increasing condition; if not, set the decoding timestamp of the current frame of the target video to the decoding timestamp of the previous output frame plus one; if so, output a timestamp sequence.

[0137] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.

[0138] The present invention also provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed, the steps provided in the above embodiments can be implemented. The storage medium may include a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, among other media capable of storing program code.

[0139] The present invention also provides a computer device that may include a memory and a processor. The memory stores a computer program, and when the processor calls the computer program in the memory, the steps provided in the above embodiment can be implemented. Of course, the computer device may also include various network interfaces, a power supply, and other components.

[0140] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the system disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

[0141] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusions, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

Claims

1. A fast export method for cutting and merging video streams, characterized in that: include: Obtain a video processing request and determine the start and end time points of multiple target video segments; Based on the starting time point and the ending time point, performing key frame positioning processing on the target video segment to establish a mapping relationship; Performing standardization processing on the target video clip and performing timestamp correction to obtain video frame data; Segmenting the video frame data based on the mapping relationship to extract frames to be encoded and frames to be copied respectively; Decoding the frame to be encoded to obtain a valid image frame, and encoding the valid image frame to generate a target encoded frame; Performing a copy process on the frame to be copied, retaining the original coding structure, and generating a target copy frame; Standardizing the timestamp information of the target coded frame and the target copied frame respectively to construct a continuously increasing timestamp sequence; After the processing of the plurality of target video segments is completed, the target coded frame and the target copied frame are merged based on the timestamp sequence to obtain a video export result.

2. The method for rapidly exporting video streams by cutting and merging according to claim 1, characterized in that: The performing key frame positioning processing on the target video segment based on the starting time point and the ending time point to construct a mapping relationship includes: Decapsulate the target video clip using a preset video processing interface to extract corresponding video stream information; Taking the starting time point or the ending time point as an input parameter, and calling a jump positioning function to perform jump positioning on the video stream information; After the jump position, read the timestamp of the video frame; Determine whether the timestamp of the video frame is greater than the start time point or the end time point; if not, use the current frame as a key frame of the video group corresponding to the time point; If yes, the time point is offset adjusted and the adjusted time point is re-jumped until the preset maximum jump limit is met or the offset time is zero, and the key frame is output; A mapping relationship between the start time point and the end time point and the corresponding key frame position is established.

3. The method for rapidly exporting video streams by cutting and merging according to claim 1, characterized in that: The step of normalizing the target video segment and performing timestamp correction on the target video segment to obtain video frame data includes: Decapsulating the target video segment and converting the video frames in the target video segment into a standard format; After completing the format conversion, calculating a minimum start timestamp value based on the target video segment; The minimum starting timestamp value is used to perform timestamp correction on the target video segment to obtain video frame data.

4. The method for rapidly exporting video streams by cutting and merging according to claim 1, wherein: The segmenting of the video frame data based on the mapping relationship to extract frames to be encoded and frames to be copied respectively includes: Read the video frame data sequentially from front to back, and compare the key frame position recorded in the mapping relationship based on the timestamp and frame type information of each frame; When the read frame is located before the key frame of the video group where the starting time point is located, discarding the read frame; When a key frame of the video group at the starting time point is read, all frames of the video group are marked as frames to be encoded; or, when a key frame of the video group at the ending time point is read, the video group and its subsequent frames are marked as frames to be encoded; The frames of the video group that are between two key frames and do not need to be encoded are marked as frames to be copied.

5. The method for rapidly exporting video streams by cutting and merging according to claim 1, characterized in that: The decoding process is performed on the frame to be encoded to obtain a valid image frame, and encoding process is performed based on the valid image frame to generate a target encoding frame, including: Decoding the frame to be encoded using a decoder to obtain a valid image frame; Determining whether the timestamp of the valid image frame is within the time range of the target video segment, and if not, discarding the valid image frame; If so, the valid image frame is input to the encoder for encoding processing to generate a target encoding frame.

6. The method for rapidly exporting video streams by cutting and merging according to claim 5, characterized in that: The step of inputting the valid image frame into an encoder for encoding to generate a target encoding frame includes: Obtaining encoding parameter information of the target video segment; Constructing an encoder configuration file according to the encoding parameter information, and setting the parameters of the encoder using the encoder configuration file; After the parameters of the encoder are set, the global header information flag item of the encoder is turned off, and the valid image frame is input into the encoder to generate a target coded frame.

7. The method for rapidly exporting video streams by cutting and merging according to claim 1, characterized in that: The standardizing of the timestamp information of the target coded frame and the target copied frame to construct a continuously increasing timestamp sequence includes: subtracting the display timestamp of the first frame of the target video segment from the display timestamp of the current frame of the target video segment, so that the display timestamp of the target video segment increases from zero, to obtain a timestamp sequence; Alternatively, subtracting the decoding timestamp of the first frame of the target video segment from the decoding timestamp of the current frame of the target video segment, so that the decoding timestamp of the target video segment increments from zero; Determine whether the adjusted decoding timestamp meets the monotonically increasing condition; if not, set the decoding timestamp of the current frame of the target video to the decoding timestamp of the previous output frame plus one; if so, output a timestamp sequence.

8. A fast export device for cutting and merging video streams, characterized in that: include: A node acquisition unit, configured to acquire a video processing request and determine the start time points and end time points of multiple target video segments; a node positioning unit, configured to perform key frame positioning processing on the target video segment based on the starting time point and the ending time point to establish a mapping relationship; A video processing unit, configured to perform standardization processing on the target video segment and perform timestamp correction to obtain video frame data; A segmentation processing unit, configured to perform segmentation processing on the video frame data based on the mapping relationship, and extract frames to be encoded and frames to be copied respectively; a video encoding unit, configured to decode the frame to be encoded to obtain a valid image frame, and perform encoding processing based on the valid image frame to generate a target encoded frame; A video copying unit, configured to copy the frame to be copied, retain the original coding structure, and generate a target copy frame; A sequence construction unit, configured to perform standardization processing on the timestamp information of the target coded frame and the target copied frame respectively, and to construct a continuously increasing timestamp sequence; The video merging unit is configured to merge the target coded frame and the target copied frame based on the timestamp sequence after the processing of the plurality of target video segments is completed, so as to obtain a video export result.

9. A computer device, characterized in that: The method comprises a memory, a processor and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the method for rapidly exporting video streams for cutting and merging according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the fast exporting method for cutting and merging video streams as described in any one of claims 1 to 7.