Billions of pixel fusion video review method, device and system
By using unified timestamp and overlapping frame group technology in the billion-pixel computational imaging system, the synchronization and continuity problems in the fusion video playback are solved, and the accurate picture synchronization and time continuity of the fusion video are achieved.
Patent Information
- Application Number
- CN202311653804.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-10-10
- Estimated Expiration
- 2043-12-05
AI Technical Summary
In the billion-pixel computational imaging system, independent exposure and encoding of each local camera causes video frames with the same frame number to be out of sync when the fused video is played back, resulting in time differences, causing frame loss and poor continuity in the fused video.
By starting the video recording of the next recording period with an interval of at least one frame group length before the end time of the unit recording period, it is ensured that there is overlapping video frame group data in the video files of adjacent recording periods, and splicing and fusion are performed based on a unified timestamp to obtain the splicing start and end time reference of the target video file, thereby achieving synchronization and continuity of video frames.
It ensures the accurate synchronization of the fusion video images, improves the continuity of the fusion video playback, avoids the problem of video frame loss, and ensures the continuity of the images and time.
Smart Images

Figure CN119094681B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of video processing technology, and in particular to a method and device for reviewing billion-pixel fusion videos, a billion-pixel computational imaging system, computer equipment, storage medium, and computer program product. Background Art
[0002] Video surveillance has been widely used in the security field due to its advantages such as high reliability, timelyness, and ease of viewing. When the monitoring area is large, to achieve ultra-high-definition video surveillance of large scenes, billion-pixel computational imaging systems based on array cameras have emerged. These systems use multiple local cameras with long focal lengths and narrow fields of view to capture multiple HD videos of the monitored scene. These local videos are then stitched and fused together to produce ultra-high-definition fused videos with a large or wide field of view, reaching up to billion-pixel levels. These systems are suitable for video capture and security monitoring of large scenes such as airports, highways, parks, sports stadiums, border crossings, and sea surfaces.
[0003] In related technologies, billion-pixel computational imaging systems can also provide a fused video playback function. To achieve this, each partial video is usually recorded in segments, resulting in multiple sets of video files. When the fused video is to be played back, the corresponding frames in each set of video files are spliced and fused.
[0004] However, since each local camera performs exposure and encoding independently, it is difficult to synchronize the shooting time of each local video. In the video files of each local camera recorded in the same period (a group of video files), the video frames with the same frame number (such as the first frame, the second frame, etc.) are not shot synchronously. Therefore, the start time and end time of each local video in a group of video files are inconsistent, and there is a certain time difference. This time difference will result in some video frames in each group of video files that cannot be used for synchronous stitching (stitching videos shot synchronously), which will lead to frame loss in the fused video and poor continuity in video playback. Summary of the Invention
[0005] Based on this, it is necessary to provide a method, device, billion-pixel computational imaging system, computer equipment, storage medium and computer program product for reviewing billion-pixel fused videos that can improve the continuity of fused video review in response to the above technical problems.
[0006] In a first aspect, the present application provides a method for reviewing a billion-pixel fusion video. The method comprises:
[0007] In response to a review instruction for the fused video, determining a video review time, and obtaining a plurality of target video files corresponding to the video review time from each group of video files corresponding to each unit recording period stored;
[0008] The set of video files corresponding to a unit recording period is obtained by recording, for each local video stream corresponding to the fused video, the video frames after the first key frame after the start time of the unit recording period and before the last reference frame of the video frame group corresponding to the end time of the unit recording period; of two adjacent unit recording periods, the start time of the next unit recording period is earlier than the end time of the previous unit recording period, and the time difference between the start time of the next unit recording period and the end time of the previous unit recording period is not less than the frame group length of the local video stream;
[0009] Splicing and fusion are performed based on each group of the target video files to obtain the fused video for review.
[0010] In one embodiment, the time difference is n times the frame group length of the local video stream, where n is a positive integer not less than 1; and the splicing and fusion based on each group of the target video files to obtain the fused video for playback includes:
[0011] For a first group of target video files in each group of target video files, determining a first timestamp of a first video frame and a second timestamp of a last video frame of each target video file in the group, using the largest first timestamp among the first timestamps as a start time reference for splicing the group of target video files, and using the smallest second timestamp among the second timestamps as an end time reference for splicing the group of target video files;
[0012] For the target video files in the other groups except the first group of target video files, determine the third timestamp of the first video frame and the fourth timestamp of the last video frame of each target video file, use the timestamp of the (n+1)th key frame in the target video file with the largest third timestamp in the group as the start time reference for splicing the target video files in the group, and use the smallest fourth timestamp among the fourth timestamps as the end time reference for splicing the target video files in the group;
[0013] Based on the splicing start time reference and the splicing end time reference of each group of target video files, each group of target video files is spliced and fused to obtain the fused video for playback.
[0014] In one embodiment, the time difference is the frame group duration of the partial video stream; before responding to the review instruction for the fused video, the method further includes:
[0015] In response to a recording instruction of the fused video, determining the start time and end time of each unit recording period according to the recording start time, the unit recording duration, and the frame group duration of the partial video stream, wherein, of two adjacent unit recording periods, the start time of the next unit recording period is earlier than the end time of the previous unit recording period, and the time difference between the start time of the next unit recording period and the end time of the previous unit recording period is the frame group duration;
[0016] When the current time reaches the start time of the current unit recording period, for each of the local video streams, the video frames after the first key frame after the start time of the current unit recording period and before the last reference frame of the video frame group corresponding to the end time of the current unit recording period are recorded as the recording files corresponding to the current unit recording period, and the recording files of each of the local video streams in the same unit recording period are stored as a group of recording files.
[0017] In one embodiment, the method further comprises:
[0018] The camera timestamp of the first key frame obtained after the start time of the video recording is used as the reference timestamp;
[0019] The difference between the camera timestamp of each video frame in each video file and the reference timestamp is used as the video timestamp of each video frame, and the video timestamp is written into the video file.
[0020] In one embodiment, the method is applied to a server, the server including a collaboration module and two recording modules, the time difference being the duration of a frame group of the local video stream; and before responding to a review instruction for the fused video, the method further includes:
[0021] The collaborative module sends a first start recording instruction to the first recording module in response to the recording instruction of the fused video. The first recording module obtains each local video stream corresponding to the fused video in response to the first start recording instruction. For each local video stream, starting from the first key frame of the local video stream obtained after the current moment, the collaborative module records the first key frame and subsequent video frames as a video file for the current unit recording period.
[0022] When the current time reaches the start time of the next unit recording period, the collaboration module sends a second start recording instruction to the second recording module. In response to the second start recording instruction, the second recording module obtains each local video stream corresponding to the fused video, and for each local video stream, starting from the first key frame of the local video stream obtained after the current time, records the first key frame and subsequent video frames as a video file for the next unit recording period.
[0023] When the current time reaches the end time of the current unit recording period, the collaboration module sends a first stop recording instruction to the first recording module. In response to the first stop recording instruction, the first recording module writes all video frames of the video frame group in which the video frame acquired at the current moment belongs to the video file of the current unit recording period and then stops recording.
[0024] When the current time reaches the end time of the next unit recording period, the collaboration module sends a second stop recording instruction to the second recording module. In response to the second stop recording instruction, the second recording module writes all video frames of the video frame group in which the video frame acquired at the current moment belongs to the video file of the next unit recording period and then stops recording.
[0025] The start time of the next unit recording period is earlier than the end time of the current unit recording period, and the time difference between the start time of the next unit recording period and the end time of the current unit recording period is the frame group length of the local video stream.
[0026] In a second aspect, the present application also provides a device for reviewing billion-pixel fusion videos. The device comprises:
[0027] An acquisition module is configured to determine a video review time in response to a review instruction for a fused video, and to obtain a plurality of target video files corresponding to the video review time from the respective groups of video files corresponding to the respective stored unit recording time periods; wherein a group of video files corresponding to a unit recording time period is obtained by recording, for each local video stream corresponding to the fused video, the video frames after the first key frame after the start time of the unit recording time period and before the last reference frame of the video frame group corresponding to the end time of the unit recording time period; of two adjacent unit recording time periods, the start time of the next unit recording time period is earlier than the end time of the previous unit recording time period, and the time difference between the start time of the next unit recording time period and the end time of the previous unit recording time period is not less than the frame group length of the local video stream;
[0028] The fusion module is used to perform splicing and fusion based on each group of target video files to obtain the fused video for playback.
[0029] In a third aspect, the present application also provides a billion-pixel computational imaging system. The billion-pixel computational imaging system comprises an array camera and a server, the array camera comprising a plurality of local cameras, wherein:
[0030] The array camera is configured to capture local videos by each of the local cameras, and send each of the local video streams to the server;
[0031] The server is configured to, in response to a recording instruction of a fusion video, store each of the acquired local video streams as a plurality of groups of recording files in a plurality of unit recording time periods; wherein a group of recording files corresponding to one unit recording time period is obtained by recording each video frame after a first key frame after a start time of the unit recording time period and before a last reference frame of a group of video frames corresponding to an end time of the unit recording time period for each local video stream corresponding to the fusion video; in adjacent two unit recording time periods, a start time of a next unit recording time period is earlier than an end time of a previous unit recording time period, and a time difference between the start time of the next unit recording time period and the end time of the previous unit recording time period is not less than a group time length of the local video stream;
[0032] The server is further configured to, in response to a review instruction for the fusion video, determine a video review time, acquire a plurality of groups of target recording files corresponding to the video review time from each of the groups of recording files corresponding to each of the unit recording time periods, and perform splicing and fusion based on each of the groups of target recording files to obtain the fusion video for review.
[0033] In a fourth aspect, the present application also provides a computer device. The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the steps of the method of the first aspect when executing the computer program.
[0034] In a fifth aspect, the present application also provides a computer readable storage medium. The computer readable storage medium stores a computer program, and the computer program implements the steps of the method of the first aspect when executed by a processor.
[0035] In a sixth aspect, the present application also provides a computer program product. The computer program product comprises a computer program, and the computer program implements the steps of the method of the first aspect when executed by a processor.
[0036] The above-mentioned billion-pixel fused video review method, apparatus, billion-pixel computational imaging system, computer equipment, storage medium, and computer program product start video recording for the next recording period at a time before the end of the current recording period and at least one frame group duration, so that at least one overlapping video frame group data exists in the video files of adjacent recording periods. Because the shooting time of each local camera is not absolutely synchronized, when reviewing, the fused video obtained based on the video file of the previous period will end with the last video frame with the smallest timestamp. The last video frame used for splicing in other video files in the group may be an intermediate reference frame (not necessarily the last video frame) in the last video frame group of the video file. Because there is at least one overlapping video frame group data between the next period and the previous period, a video frame continuous with the end time of the previous period can be found in the video file of the next period, and the video file also includes the key frame of the video frame group to which the video frame belongs. Therefore, the continuous video frames can be decoded to obtain a complete image, thereby splicing a fused video with accurate picture synchronization and temporal continuity. Therefore, based on this solution, fusion video playback can ensure the accuracy of fusion video picture synchronization while improving the continuity of fusion video playback, and can solve the problems of frame loss and poor continuity of fusion video in related technologies. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1a A schematic diagram of an example of a billion-pixel computational imaging system;
[0038] Figure 1b to Figure 1c The following is a schematic diagram of recording a local video stream in an example; Figure 1b The middle part shows the situation where all local videos are shot simultaneously. Figure 1c This is the case where each local video is shot asynchronously;
[0039] Figure 2 1. A flowchart of a method for reviewing a billion-pixel fusion video according to an embodiment;
[0040] Figure 3 This is a schematic diagram of recording a local video stream based on this solution in an example;
[0041] Figure 4 This is a structural block diagram of a device for reviewing billion-pixel fusion video in one embodiment;
[0042] Figure 5 FIG. 1 is a diagram showing the internal structure of a computer device in one embodiment. DETAILED DESCRIPTION
[0043] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.
[0044] First, before specifically introducing the technical solutions of the embodiments of the present application, the technical background or technical evolution of the embodiments of the present application is introduced. In order to achieve ultra-high-definition video surveillance of large scenes, billion-pixel computational imaging systems have emerged, such as Figure 1a In the illustrated exapixel computational imaging system, the array camera 102 and the display terminal 106 communicate with a server 104 via a network. The array camera 102 includes multiple local cameras with different shooting angles, each capable of capturing localized high-definition video of a portion of the monitored scene. The server then stitches and fuses these local videos to produce a large or wide field of view, exapixel-level fused video. This video is then played or displayed on a display terminal. This system is suitable for video capture and security monitoring of large scenes such as airports, highways, parks, sports stadiums, border crossings, and ocean surfaces.
[0045] To meet users' needs for playback and viewing of fused videos from historical time periods, the billion-pixel computational imaging system can also provide a fused video playback function, which involves recording and storing ultra-high-definition fused videos. In practical applications, to conserve computing resources, the local videos captured by each local camera of the array camera are typically recorded separately, resulting in multiple video files. When a user needs to review the fused video of a certain time period, the local video data for that time period can be spliced and fused. This avoids splicing and fusion of the local videos that do not need to be reviewed, saving computing resources. In addition, if a long local video stream (high-definition video, with a large data volume) is recorded as a single file, the video file will be large, which is not conducive to storage and use. Therefore, to improve the convenience and flexibility of storage and use, each local video is typically recorded in segments, resulting in multiple sets of time-continuous video files. For example, the duration of each video file or each segment can be set to L (e.g., 120,000 milliseconds, i.e., 2 minutes).
[0046] The current coding standards used for local video streams (such as H264, etc.) are usually encoded using a video frame group (GOP, Group of picture) method. The first frame of each video frame group is a key frame (I frame) containing complete image information, and the other frames after the key frame are reference frames (P frames or B frames) containing partial image information (usually difference information from the video frame before the frame). The reference frame must be based on the data of other video frames after (including the key frame) and before the frame of the group (possibly including the reference frame after it) in order to decode and obtain a complete image.
[0047] like Figure 1b As shown, this example uses the local video recording of three local cameras (A11, A12, and A13) as an example to illustrate. The continuous line segment in the middle represents the local video stream shot by the local camera. Each small line segment represents a video frame group (including the front end point but not the back end point). The front end point of the line segment is the key frame (first frame) of the video frame group, and the back end point is the key frame of the next video frame group. The time to start recording is S0. When the server receives the recording instruction, it starts to pull the local video stream of each local camera, where I 1_A11 The first key frame of the local camera A11 obtained after the video starts (since after S0, I 1_A11 Previously, it was only a reference frame of the previous video frame group and could not be fully decoded, so recording started from the key frame). 1_A11 Start recording and generate video file F 1_A11 When the recording time reaches the set L, the current complete video frame group is written into the video file F 1_A11 (This file contains keyframes I 1_A11 , corresponding to the solid dots in the figure, excluding keyframe I m1_A11 , corresponding to the hollow dots in the figure), the next video frame group (I m1_A11 The frame group where it is located) is recorded as a new video file F 2_A11 (Including keyframe I m1_A11 ), so as to realize segmented recording. Similarly, the video files of other local cameras can be obtained (F 1_A12 、F 1_A13 The video files of each local camera recorded in the same segment are taken as a group. When the fused video needs to be reviewed, the video files of each group are obtained and stitched and fused.
[0048] If the exposure time of each local camera is exactly the same, then each frame with the same sequence number (such as the first frame I) in a group of video files in the same segment will be recorded by each local camera. 1_A11 , I 1_A12 , I 1_A13 ) are shot synchronously (with the same shooting time), so starting from the first frame, each frame is stitched and fused sequentially (every first frame is stitched, every second frame is stitched, etc.), and an accurate fused video can be obtained. However, the local cameras of current array cameras usually perform exposure and encoding independently. In fact, the shooting time of each local video cannot be absolutely consistent, and there is a certain time difference. Figure 1cAs shown in the figure, after the server receives the recording instruction, the shooting time of the first key frame of each local video stream pulled out is different. Therefore, each frame with the same sequence number in each video file is not shot synchronously. If the frames are stitched and fused sequentially starting from the first frame, a fused video image with severe picture distortion and poor imaging quality will be obtained. In related technologies, the timestamp reference of each frame in each video file can be unified, and the video images shot synchronously can be stitched and fused based on the timestamp of the unified reference. For example, the first key frame I shot by the local camera A11 after starting the recording is 1_A11 The time is T0, and the first key frame (I 1_A12 and I 1_A13 ) are all earlier than T0, and the video frame shot at T0 is a reference frame in the video frame group where the first key frame is located. For the video file group F shot during this period (S0 to S0+L), 1_A11 、F 1_A12 、F 1_A13 , you can use the time after T0 (including T0) and T 12 Before time (excluding T 12 ) The corresponding video frames shot are spliced and fused to obtain a fused video with synchronized images and accurate splicing. For the video file group F shot in the next period (S0+L~S0+2L), 2_A11 、F 2_A12 、F 2_A13 , you can use T 12 After time (including T 12 ) are spliced and fused accordingly.
[0049] The above method can achieve synchronization and accuracy of the playback screen of the fusion video. However, this method has the problem of frame loss in the fusion video, and the continuity of the video playback is poor. Figure 1c In the example, based on a set of video files (F 1_A11 、F 1_A12 、F 1_A13 ) and a set of video files for the next period (F 2_A11 、F 2_A12 、F 2_A13 ), all cannot be spliced and fused to obtain time T 12 To T 11 The fused video between the two, that is, the frame with the earliest shooting time and the smallest timestamp in the last frame of each local video stream of the previous period (in this example, I m1_A12 ), and the frame with the latest shooting time and the largest timestamp (in this example, m1_A11 If there is a time gap between the previous frame and the previous frame, the fused video frame within the time gap will be lost. If something important happens during this time gap, it will affect the review and evidence of the video.
[0050] Against this backdrop, the applicant, through long-term research and development and experimental verification, has proposed the present invention's method for replaying billion-pixel fused videos. This method ensures accurate synchronization of the fused video while improving the continuity of the fused video playback, resolving the issues of frame loss and poor continuity in the fused video found in related technologies. Furthermore, it should be noted that the applicant has expended considerable creative effort in discovering the technical issues of this application and in developing the technical solutions described in the following embodiments.
[0051] In one embodiment, Figure 2 As shown in the figure, a method for reviewing billion-pixel fusion video is provided, which can be applied to Figure 1a The server in the billion-pixel computational imaging system shown in FIG. The server can be implemented as an independent server or a server cluster consisting of multiple servers. In this embodiment, the method includes the following steps:
[0052] Step 201 : In response to a review instruction for a fused video, determine a video review time, and obtain a plurality of target video files corresponding to the video review time from the stored video files corresponding to each unit recording period.
[0053] In implementation, the user can select a video review time and trigger a review instruction for the fused video at that time, so that the server can respond to the review instruction and obtain the video review time. The video review time can be a time period, such as selecting to review the video of a certain day, or reviewing the video of a certain period of time on a certain day, or it can be a video review start time. It can be understood that if the video review time is a time period, the server can obtain each group of target video files corresponding to the time period for splicing and fusion to obtain a fused video; if the video review time is the review start time, the server can determine the unit recording time period where the start time is located, and obtain the unit recording time period and several groups of target video files thereafter, so as to splice and fuse to obtain a continuous fused video, until the user triggers the stop review instruction, or the recorded video ends.
[0054] The video files are pre-recorded and stored. The array camera can send the captured local video streams to the server, so that the server can obtain the local video streams. If the user currently needs to view the real-time fused video, the server can splice and fuse the local video streams in real time, and send the fused video to the display terminal for playback. The server can also record the acquired local video streams in segments based on the user's instructions. The server can store the video files recorded by each local video stream in the same unit time period (such as a duration of L) as a group of video files, such as storing them in the same folder, and each unit recording time period corresponds to a folder. The start time or end time of the recording time period can be used as the name of the folder, which is convenient for indexing the video files of each unit recording time period that is continuous in time based on the folder name.
[0055] A set of video files corresponding to a recording period is obtained by recording, for each local video stream corresponding to the fused video, all video frames from the first key frame after the start time of the recording period (including the key frame) to the last reference frame of the video frame group corresponding to the end time of the recording period (including the last reference frame). To ensure that each video frame in the recording file can be decoded into a complete image, each recording file is typically recorded starting from the key frame and ending with the last frame of the key frame group. In two adjacent unit recording periods, the start time of the next unit recording period is earlier than the end time of the previous unit recording period, and the time difference between the start time of the next unit recording period and the end time of the previous unit recording period is no less than the frame group duration of the local video stream (denoted as G). The frame group duration G is the duration of a video frame group from the first frame (key frame) to the last frame (the duration of a GOP group). It can be pre-set or calculated by the server based on the timestamps of two adjacent key frames in the local video stream (timestamp difference). Typically, the frame group duration of the local video streams of each local camera is the same.
[0056] For example, Figure 3 As shown, S0 is the start time of the video recording, which is also the start time of the first unit recording period. The end time of the first unit recording period is S0+L. The start time of the second unit recording period is S0+LG and the end time is S0+2L. A group of video files (F 1_A11 、F 1_A12 、F 1_A13 ), the first video frame is I 1_A11 , I 1_A12 , I 1_A13 , the last video frame is I m1_A11 , I m1_A12 , I m1_A13The previous frame of (the hollow circle does not include the endpoint). A set of video files of the second unit recording period (F 2_A11 、F 2_A12 、F 2_A13 ), the first video frame is I m1-1_A11 , I m1-1_A12 , I m1-1_A13 In this example, the start time of the second unit recording period differs from the end time of the first unit recording period by a frame group length G. Therefore, the video file of the second unit recording period and the video file of the first unit recording period have an overlapping video frame group. In this example, the overlapping frame group is key frame I. m1-1 The video frame group.
[0057] In other examples, the time difference between the start time of the next unit recording period and the end time of the previous unit recording period can be greater than the length of a frame group, such as n times the length of the frame group (n is a positive integer not less than 1). Therefore, the video files of two adjacent unit recording periods will have n overlapping video frame groups.
[0058] Step 202: performing splicing and fusion based on each set of target video files to obtain a fused video for playback.
[0059] In practice, after the server obtains each set of target video files, it can stitch and merge each set of video files in sequence according to the time sequence. Since the partial videos are not shot synchronously, it is not possible to stitch together the frames with the same frame number in sequence. Instead, it is necessary to stitch and merge the frames with the same or similar timestamps according to the timestamps of each video frame (the timestamps after the reference is unified). Figure 3 As shown, for a set of video files (F 1_A11 、F 1_A12 、F 1_A13 ), start stitching from the video frame shot at time T0 (including the video frame at that time), until T 12 The video frame ends at time (excluding T 12 For a set of video files (F 2_A11 、F 2_A12 、F 2_A13 ), from T 12 The video frames captured at the same moment in time are stitched together. This ensures that a complete image can be decoded from the set of video files at each stitching moment, and that the stitched images are shot synchronously, resulting in accurate synchronization of the resulting fused video. Furthermore, the last fused frame of the previous set of video files is temporally continuous with the first fused frame of the next set of video files, thus preventing frame dropouts in the fused video.
[0060] In one embodiment, the time difference between the start time of the next unit recording period and the end time of the previous unit recording period may be n times the frame group length of the local video stream, where n is a positive integer not less than 1. The process of splicing and fusing based on each group of target video files in step 202 specifically includes the following steps:
[0061] In step 2021, for the first group of target video files in each group of target video files, determine the first timestamp of the first video frame and the second timestamp of the last video frame in each target video file, use the largest first timestamp among the first timestamps as the start time reference for splicing the group of target video files, and use the smallest second timestamp among the second timestamps as the end time reference for splicing the group of target video files.
[0062] The timestamps of each video frame are based on a unified reference and can reflect the order in which each video frame was captured. For example, the camera times of each local camera can be synchronized and then the timestamps of each video frame can be determined based on the camera timestamps.
[0063] In one implementation, based on existing video coding standards (such as H.264), the timestamps of each video frame in a video file are typically calculated using a time difference method. The camera timestamp of the first frame in the video file is generally used as a benchmark, and the timestamp of each video frame is subtracted from this benchmark to obtain the video timestamp of each video frame, such as 0 for the first frame and 40 for the second frame. This timestamp only reflects the chronological order of the video frames in the video file and cannot reflect the absolute chronological order of the video frames in different video files. For two video files shot at asynchronous times, two video frames with the same video timestamp were not shot synchronously. Splicing and fusion based on this timestamp will not produce a fused video with accurate image synchronization. Therefore, it is necessary to unify the benchmarks of the timestamps of each video frame in the video file.
[0064] Specifically, the server can use the camera timestamp of the first key frame obtained after the recording start time S0 as the reference timestamp, and then use the difference between the camera timestamp of each video frame in each recording file and the reference timestamp as the video timestamp of each video frame, and write the video timestamp into the recording file. Figure 3 As shown, the first key frame obtained after the video start time S0 is the I shot by the local camera A12. 1_A12 , so the local camera A12 can shoot I 1_A12The camera timestamp at the time of recording is used as the reference timestamp. The difference between each video frame's camera timestamp and this reference timestamp is then written into the recording file as the video timestamp. This ensures that the timestamp reference for each video frame is consistent, ensuring that video frames with the same or similar timestamps in each recording file were captured synchronously. By splicing and fusing based on these video timestamps, a fused video with accurate image synchronization can be obtained.
[0065] It should be noted that the first set of target video files refers to the set of target video files with the earliest recording time among the several sets of target video files corresponding to the acquired video playback time, and is not the same as the first set of video files after the recording start time. For example, if the first set of target video files is a set of video files recorded during the second unit recording period (F 2_A11 、F 2_A12 、F 2_A13 ), the first video frame of each target video file in the group is I m1-1_A11 , I m1-1_A12 , I m1-1_A13 , the last video frame is I m2_A11 , I m2_A12 , I m2_A13 The previous frame, or I m2-1_A11 , I m2-1_A12 , I m2-1_A13 The last video frame of the video frame group. m1-1_A11 , I m1-1_A12 , I m1-1_A13 middle, I m1-1_A11 The timestamp of I is the largest (the shooting time is the latest), m2_A11 , I m2_A12 , I m2_A13 middle, I m2_A12 The timestamp of the local camera A11 is the smallest (the shooting time is the earliest), so the video file F corresponding to the local camera A11 is 2_A11 The timestamp of the first video frame in the target video file is used as the starting time reference for the splicing of the target video file. The video file F corresponding to the local camera A12 is 2_A12 The last video frame in m2_A11 The timestamp of the previous frame is used as the time reference for the end of splicing the target video files.
[0066] In step 2022, for the target video files in the other groups except the first group, the third timestamp of the first video frame and the fourth timestamp of the last video frame in each target video file are determined. The timestamp of the (n+1)th key frame in the target video file with the largest third timestamp in the group is used as the start time reference for splicing the target video files in the group, and the smallest fourth timestamp among the fourth timestamps is used as the end time reference for splicing the target video files in the group.
[0067] For example, for a set of target video files F 3_A11 、F 3_A12 、F 3_A13 , target video file F 3_A11 The first video frame I m2-1_A11 The timestamp of the target video file is the largest. 3_A12 The timestamp of the last video frame is the smallest, so the target video file F 3_A11 The n+1th key frame, in this example n=1, that is, the second key frame I m2_A11 The timestamp of is used as the starting time reference for the splicing of the target video files (the next frame of the ending time reference for the splicing of the previous target video files). The target video file F 3_A12 The timestamp of the last video frame is used as the splicing end time reference for the group of target video files.
[0068] Step 2023: Based on the splicing start time reference and the splicing end time reference of each group of target video files, each group of target video files is spliced and fused to obtain a fused video for playback.
[0069] During implementation, after the server determines the splicing start time benchmark and the splicing end time benchmark of each group of target video files, it can sequentially splice and fuse the video frames after the splicing start time benchmark (including this benchmark) in each group of video files until the video frame corresponding to the splicing end time benchmark ends (including this benchmark). The timestamps of the sequentially spliced video frames are the same or similar, which is synchronous video splicing. In adjacent recording periods, the time of the last frame of the fused video in the previous recording period is continuous with the time of the first frame of the fused video in the next recording period, avoiding the frame loss problem.
[0070] In one embodiment, the server can control the two recording modules to record video files through a collaborative module. Specifically, a user can trigger a recording instruction (e.g., the server can be set to automatically record when it is turned on). The collaborative module can respond to the recording instruction of the fused video by sending a first start recording instruction to the first recording module, which instruction includes the recording start time S0. In response to the first start recording instruction, the first recording module can obtain each local video stream corresponding to the fused video and, for each local video stream, record the first key frame and subsequent video frames, starting from the first key frame of the local video stream obtained after the current time, as a video file for the current recording period.
[0071] After the collaboration module sends the first start recording instruction, it can start a timer. When the current time reaches the start time of the next recording period, the collaboration module can send a second start recording instruction to the second recording module. The start time of the next recording period is earlier than the end time of the current recording period, and the time difference between the start time of the next recording period and the end time of the current recording period is the frame group length of the local video stream.
[0072] The second recording module can respond to the second start recording instruction to obtain the local video streams corresponding to the fused video. For each local video stream, starting from the first key frame of the local video stream obtained after the current moment, the first key frame and subsequent video frames are recorded as the recording file of the next recording period.
[0073] When the current time reaches the end time of the current recording period, the collaborative module sends a first stop recording instruction to the first recording module. When the current time reaches the end time of the next recording period, the collaborative module sends a second stop recording instruction to the second recording module. In response to the first stop recording instruction, the first recording module writes all video frames of the video frame group to which the video frame currently acquired belongs to the recording file for the current recording period and then stops recording. In response to the second stop recording instruction, the second recording module writes all video frames of the video frame group to which the video frame currently acquired belongs to the recording file for the next recording period and then stops recording.
[0074] When the current time reaches the start time of the next recording period of the above-mentioned next recording period (which can be used as the new current recording period), a first start recording instruction can be sent to the first recording module. In this way, the recording files are recorded in a loop by the two recording modules, so that an overlapping video frame group can be present in the recording files of two adjacent recording periods to ensure the continuity of the fused video.
[0075] The above-mentioned billion-pixel fusion video playback method starts the video recording of the next recording period at a time before the end time of the current recording period and at a time interval of at least one frame group length, so that there is at least one overlapping video frame group data in the video files of adjacent recording periods. Because the shooting time of each local camera is not absolutely synchronized, when reviewing, the fusion video obtained based on the video file of the previous period will end with the last video frame with the smallest timestamp. The last video frame used for splicing in other video files in the group will be an intermediate reference frame (not the last video frame) in the last video frame group of the video file. Because there is at least one overlapping video frame group data between the next period and the previous period, a video frame that is continuous with the end time of the previous period can be found in the video file of the next period, and the video file also includes the key frame of the video frame group in which the video frame is located. Therefore, the continuous video frames can be decoded to obtain a complete image, and thus a fusion video with accurate picture synchronization and temporal continuity can be spliced. Therefore, based on this solution, fusion video playback can improve the continuity of fusion video playback while ensuring accurate picture synchronization of the fusion video, and can solve the problems of frame loss and poor continuity of fusion video in related technologies.
[0076] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.
[0077] Based on the same inventive concept, the embodiments of the present application also provide a device for replaying billion-pixel fusion videos for implementing the aforementioned method for replaying billion-pixel fusion videos. The implementation solution provided by this device is similar to the implementation solution described in the aforementioned method. Therefore, the specific limitations in the embodiments of one or more devices for replaying billion-pixel fusion videos provided below can be found in the above-mentioned limitations on the method for replaying billion-pixel fusion videos, and will not be repeated here.
[0078] In one embodiment, Figure 4 As shown, a device 400 for reviewing a billion-pixel fusion video is provided, comprising: an acquisition module 401 and a fusion module 402, wherein:
[0079] The acquisition module 401 is used to determine the video review time in response to the review instruction for the fused video, and obtain several groups of target video files corresponding to the video review time in each group of video files corresponding to each stored unit recording time period; wherein, a group of video files corresponding to a unit recording time period is obtained by recording each video frame after the first key frame after the start time of the unit recording time period and before the last reference frame of the video frame group corresponding to the end time of the unit recording time period for each local video stream corresponding to the fused video; in two adjacent unit recording time periods, the start time of the next unit recording time period is earlier than the end time of the previous unit recording time period, and the time difference between the start time of the next unit recording time period and the end time of the previous unit recording time period is not less than the frame group length of the local video stream.
[0080] The fusion module 402 is used to perform splicing and fusion based on each group of the target video files to obtain the fused video for playback.
[0081] In one embodiment, the time difference is n times the duration of the frame group of the local video stream, where n is a positive integer not less than 1. The fusion module 402 is specifically configured to: for a first group of target video files in each group of target video files, determine a first timestamp of a first video frame and a second timestamp of a last video frame of each target video file in the group, use the largest first timestamp among the first timestamps as a start time reference for splicing the target video files in the group, and use the smallest second timestamp among the second timestamps as an end time reference for splicing the target video files in the group; for other groups of target video files except the first group of target video files, determine a third timestamp of a first video frame and a fourth timestamp of a last video frame of each target video file, use the timestamp of the (n+1)th key frame in the target video file with the largest third timestamp in the group as a start time reference for splicing the target video files in the group, and use the smallest fourth timestamp among the fourth timestamps as an end time reference for splicing the target video files in the group; and splice and fuse the target video files in each group based on the splicing start time reference and the splicing end time reference of each group of target video files to obtain the fused video for playback.
[0082] In one of the embodiments, the time difference is the frame group duration of the local video stream. The device further comprises a recording module configured to, in response to a recording instruction of the fusion video, determine the start time and the end time of each unit recording period according to the recording start time, the unit recording duration and the frame group duration of the local video stream, wherein the start time of a next unit recording period is earlier than the end time of a previous unit recording period, and the time difference between the start time of the next unit recording period and the end time of the previous unit recording period is the frame group duration; when the current time reaches the start time of a current unit recording period, for each of the local video streams, record each video frame after the first key frame after the start time of the current unit recording period and before the last reference frame of the video frame group corresponding to the end time of the current unit recording period as a recording file corresponding to the current unit recording period, and store the recording files of each of the local video streams in the same unit recording period as a group of recording files.
[0083] In one of the embodiments, the device further comprises a writing module configured to write the camera timestamp of the first key frame acquired after the recording start time as a reference timestamp, write the difference between the camera timestamp of each video frame in each recording file and the reference timestamp as a video timestamp of each video frame, and write the video timestamp in the recording file.
[0084] In one of the embodiments, the device comprises a coordination module and two recording modules, wherein:
[0085] The coordination module is configured to send a first start recording instruction to a first recording module in response to a recording instruction of the fusion video.
[0086] The first recording module is configured to, in response to the first start recording instruction, acquire each local video stream corresponding to the fusion video, and for each of the local video streams, record each video frame after the first key frame of the local video stream acquired after the current time as a recording file of the current unit recording period.
[0087] The coordination module is further configured to send a second start recording instruction to a second recording module when the current time reaches the start time of a next unit recording period. The start time of the next unit recording period is earlier than the end time of the current unit recording period, and the time difference between the start time of the next unit recording period and the end time of the current unit recording period is the frame group duration of the local video stream.
[0088] The second recording module is used to respond to the second start recording instruction to obtain the local video streams corresponding to the fused video, and for each of the local video streams, starting from the first key frame of the local video stream obtained after the current moment, record the first key frame and subsequent video frames as the recording file of the next unit recording period.
[0089] The collaboration module is further configured to send a first stop recording instruction to the first recording module when the current time reaches the end time of the current unit recording period.
[0090] The first recording module is further configured to, in response to the first recording stop instruction, stop recording after writing all video frames of the video frame group in which the video frame acquired at the current moment belongs to the video file of the current unit recording period.
[0091] The collaboration module is further configured to send a second stop recording instruction to the second recording module when the current time reaches the end time of the next unit recording period.
[0092] The second recording module is further configured to, in response to the second recording stop instruction, stop recording after writing all video frames of the video frame group in which the video frame acquired at the current moment belongs to the video file of the next unit recording period.
[0093] Each module in the above-mentioned device for reviewing billion-pixel fusion video can be implemented in whole or in part through software, hardware, or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in hardware form, or can be stored in a computer device memory in software form, so that the processor can call and execute the corresponding operations of each module.
[0094] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, a memory and a network interface connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data required or generated for executing the above-mentioned method for reviewing billion-level pixel fusion videos. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, a method for reviewing billion-level pixel fusion videos is implemented.
[0095] Those skilled in the art will understand that Figure 5The structure shown in the figure is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than shown in the figure, or combine certain components, or have a different component arrangement.
[0096] In one embodiment, a billion-pixel computational imaging system is provided. The billion-pixel computational imaging system includes an array camera and a server, wherein the array camera includes multiple local cameras, wherein:
[0097] The array camera is used to capture local videos through each local camera and send each local video stream to the server;
[0098] The server is configured to, in response to a recording instruction of the fused video, store the acquired partial video streams as multiple groups of recording files corresponding to multiple unit recording periods; wherein a group of recording files corresponding to one unit recording period is obtained by recording, for each partial video stream corresponding to the fused video, each video frame after the first key frame after the start time of the unit recording period and before the last reference frame of the video frame group corresponding to the end time of the unit recording period; of two adjacent unit recording periods, the start time of the next unit recording period is earlier than the end time of the previous unit recording period, and the time difference between the start time of the next unit recording period and the end time of the previous unit recording period is not less than the frame group length of the partial video stream;
[0099] The server is also used to respond to the playback instruction for the fused video, determine the video playback time, and obtain several groups of target video files corresponding to the video playback time from each group of video files corresponding to each unit recording time period stored; and perform splicing and fusion based on each group of target video files to obtain a fused video for playback.
[0100] In one embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor implements the steps in the above method embodiments when executing the computer program.
[0101] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above-mentioned method embodiments are implemented.
[0102] In one embodiment, a computer program product is provided, including a computer program, which implements the steps in the above method embodiments when executed by a processor.
[0103] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.
[0104] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiment methods can be implemented by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, database or other media used in the embodiments provided in this application may include at least one of non-volatile and volatile memory. Non-volatile memory may include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory may include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in the various embodiments provided herein may include at least one of a relational database and a non-relational database. Non-relational databases may include, but are not limited to, distributed databases based on blockchains. The processor involved in the various embodiments provided herein may be, but are not limited to, a general-purpose processor, a central processing unit, a graphics processing unit, a digital signal processor, a programmable logic unit, a data processing logic unit based on quantum computing, and the like.
[0105] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0106] The above-described embodiments are merely illustrative of several embodiments of the present application, and the description is relatively specific and detailed, but should not be understood as a limitation on the scope of the patent. It should be noted that for those skilled in the art, without departing from the concept of the present application, a number of modifications and improvements can be made, which are all within the scope of the present application. Therefore, the scope of protection of the present application should be subject to the appended claims.
Claims
1. A method for reviewing billion-pixel fusion video, characterized in that: The method comprises: In response to a review instruction for the fused video, determining a video review time, and obtaining a plurality of target video files corresponding to the video review time from each group of video files corresponding to each unit recording time period stored; The set of video files corresponding to a unit recording period is obtained by recording, for each local video stream corresponding to the fused video, the video frames after the first key frame after the start time of the unit recording period and before the last reference frame of the video frame group corresponding to the end time of the unit recording period; of two adjacent unit recording periods, the start time of the next unit recording period is earlier than the end time of the previous unit recording period, and the time difference between the start time of the next unit recording period and the end time of the previous unit recording period is not less than the frame group length of the local video stream; Perform splicing and fusion based on each group of target video files to obtain the fused video for review; The time difference is n times the frame group length of the local video stream, where n is a positive integer not less than 1; and the splicing and fusion based on each group of the target video files to obtain the fused video for playback includes: For a first group of target video files in each group of target video files, determining a first timestamp of a first video frame and a second timestamp of a last video frame of each target video file in the group, using the largest first timestamp among the first timestamps as a start time reference for splicing the group of target video files, and using the smallest second timestamp among the second timestamps as an end time reference for splicing the group of target video files; For the target video files in the other groups except the first group of target video files, determine the third timestamp of the first video frame and the fourth timestamp of the last video frame of each target video file, use the timestamp of the (n+1)th key frame in the target video file with the largest third timestamp in the group as the start time reference for splicing the target video files in the group, and use the smallest fourth timestamp among the fourth timestamps as the end time reference for splicing the target video files in the group; Based on the splicing start time reference and the splicing end time reference of each group of target video files, each group of target video files is spliced and fused to obtain the fused video for playback.
2. The method according to claim 1, characterized in that The time difference is the frame group duration of the local video stream; Before responding to the review instruction for the fused video, the method further includes: In response to a recording instruction of the fused video, determining the start time and end time of each unit recording period according to the recording start time, the unit recording duration, and the frame group duration of the partial video stream, wherein, of two adjacent unit recording periods, the start time of the next unit recording period is earlier than the end time of the previous unit recording period, and the time difference between the start time of the next unit recording period and the end time of the previous unit recording period is the frame group duration; When the current time reaches the start time of the current unit recording period, for each of the local video streams, the video frames after the first key frame after the start time of the current unit recording period and before the last reference frame of the video frame group corresponding to the end time of the current unit recording period are recorded as the recording files corresponding to the current unit recording period, and the recording files of each of the local video streams in the same unit recording period are stored as a group of recording files.
3. The method according to claim 2, characterized in that The method further comprises: The camera timestamp of the first key frame obtained after the start time of the video recording is used as the reference timestamp; The difference between the camera timestamp of each video frame in each video file and the reference timestamp is used as the video timestamp of each video frame, and the video timestamp is written into the video file.
4. The method according to claim 1, wherein The method is applied to a server, the server including a collaboration module and two recording modules, the time difference being the duration of a frame group of the local video stream; Before responding to the review instruction for the fused video, the method further includes: The collaborative module sends a first start recording instruction to the first recording module in response to the recording instruction of the fused video. The first recording module obtains each local video stream corresponding to the fused video in response to the first start recording instruction. For each local video stream, starting from the first key frame of the local video stream obtained after the current moment, the collaborative module records the first key frame and subsequent video frames as a video file for the current unit recording period. When the current time reaches the start time of the next unit recording period, the collaboration module sends a second start recording instruction to the second recording module. In response to the second start recording instruction, the second recording module obtains each local video stream corresponding to the fused video, and for each local video stream, starting from the first key frame of the local video stream obtained after the current time, records the first key frame and subsequent video frames as a video file for the next unit recording period. When the current time reaches the end time of the current unit recording period, the collaboration module sends a first stop recording instruction to the first recording module. In response to the first stop recording instruction, the first recording module writes all video frames of the video frame group in which the video frame acquired at the current moment belongs to the video file of the current unit recording period and then stops recording. When the current time reaches the end time of the next unit recording period, the collaboration module sends a second stop recording instruction to the second recording module. In response to the second stop recording instruction, the second recording module writes all video frames of the video frame group in which the video frame acquired at the current moment belongs to the video file of the next unit recording period and then stops recording. The start time of the next unit recording period is earlier than the end time of the current unit recording period, and the time difference between the start time of the next unit recording period and the end time of the current unit recording period is the frame group length of the local video stream.
5. A device for reviewing billion-pixel fusion video, characterized in that: The device comprises: An acquisition module is configured to determine a video review time in response to a review instruction for a fused video, and to obtain a plurality of target video files corresponding to the video review time from the respective groups of video files corresponding to the respective stored unit recording time periods; wherein a group of video files corresponding to a unit recording time period is obtained by recording, for each local video stream corresponding to the fused video, the video frames after the first key frame after the start time of the unit recording time period and before the last reference frame of the video frame group corresponding to the end time of the unit recording time period; of two adjacent unit recording time periods, the start time of the next unit recording time period is earlier than the end time of the previous unit recording time period, and the time difference between the start time of the next unit recording time period and the end time of the previous unit recording time period is not less than the frame group length of the local video stream; A fusion module is configured to perform splicing and fusion based on each group of target video files to obtain the fused video for playback; wherein the time difference is n times the frame group length of the local video stream, and n is a positive integer not less than 1; the splicing and fusion based on each group of target video files to obtain the fused video for playback includes: for the first group of target video files in each group of target video files, determining the first timestamp of the first video frame and the second timestamp of the last video frame of each target video file in the group, taking the largest first timestamp of each first timestamp as the splicing start time reference of the group of target video files, and taking the smallest second timestamp of each second timestamp as the splicing start time reference of the group of target video files. the target video files in the other groups except the first group are spliced together; for each target video file, the third timestamp of the first video frame and the fourth timestamp of the last video frame are determined; the timestamp of the n+1th key frame in the target video file with the largest third timestamp in the group is used as the splicing start time reference of the target video files in the group; the smallest fourth timestamp among the fourth timestamps is used as the splicing end time reference of the target video files in the group; based on the splicing start time reference and the splicing end time reference of each group, the target video files in each group are spliced and fused to obtain the fused video for playback.
6. A billion-pixel computational imaging system, characterized in that: The billion-pixel computational imaging system includes an array camera and a server, wherein the array camera includes multiple local cameras, wherein: The array camera is used to shoot local videos through each of the local cameras and send each local video stream to the server; The server is configured to store each of the acquired partial video streams as a plurality of groups of video files corresponding to a plurality of unit recording time periods in response to a recording instruction of the fused video; wherein a group of video files corresponding to a unit recording time period is obtained by recording, for each partial video stream corresponding to the fused video, each video frame after the first key frame after the start time of the unit recording time period and before the last reference frame of the video frame group corresponding to the end time of the unit recording time period; of two adjacent unit recording time periods, the start time of the next unit recording time period is earlier than the end time of the previous unit recording time period, and the time difference between the start time of the next unit recording time period and the end time of the previous unit recording time period is not less than the frame group length of the partial video stream; The server is also used to respond to the review instruction for the fused video, determine the video review time, and obtain several groups of target video files corresponding to the video review time in each group of video files corresponding to each unit recording time period stored; perform splicing and fusion based on each group of the target video files to obtain the fused video for review; wherein, the time difference is n times the frame group length of the local video stream, and n is a positive integer not less than 1; the splicing and fusion based on each group of the target video files to obtain the fused video for review includes: for the first group of target video files in each group of the target video files, determine the first timestamp of the first video frame and the second timestamp of the last video frame of each target video file in the group, and use the largest first timestamp among the first timestamps as the target video file of the group The splicing start time reference of each target video file is determined, and the smallest second timestamp among each second timestamp is used as the splicing end time reference of the group of target video files; for the target video files in other groups except the first group of target video files, the third timestamp of the first video frame and the fourth timestamp of the last video frame of each target video file are determined, and the timestamp of the n+1th key frame in the target video file with the largest third timestamp in the group is used as the splicing start time reference of the group of target video files, and the smallest fourth timestamp among each fourth timestamp is used as the splicing end time reference of the group of target video files; based on the splicing start time reference and the splicing end time reference of each group of target video files, the target video files in each group are spliced and fused to obtain the fused video for playback.
7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Multi-video automatic splicing method and system for track and field project, and medium
CN114401378A
Using worker nodes in a distributed video encoding system
US20170078376A1