A multi-picture space-time synchronous frame-by-frame cooperative playback method

CN122845870APending Publication Date: 2026-09-29XINAOTE (NANJING) VIDEO TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611310627.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-27
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

因此,在逐帧回放过程中,当播放位置由当前回放步序切换至另一回放步序时,不同视点画面所对应的数据获取过程和图像处理过程可能具有不同的处理路径及完成时序

Benefits of technology

1.通过建立虚拟视点帧与其生成所采用的真实视点源帧之间的关联关系,使系统在回放步序发生变化时能够根据真实视点源帧的变化情况确定对应虚拟视点帧的当前状态,从而避免在回放步序切换后显示与目标回放步序不对应的虚拟视点帧,减少因源帧变化未同步更新导致的多画面时序不一致,并使虚拟视点帧与其源真实视点帧之间的生成对应关系可追溯。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122845870A_ABST
    Figure CN122845870A_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video processing, in particular to a multi-picture space-time synchronous frame-by-frame cooperative playback method, which comprises the following steps: establishing a time sequence corresponding relationship and a unified space reference of multi-path real viewpoint videos; acquiring a real viewpoint frame corresponding to a current playback step sequence; generating a virtual viewpoint frame; and simultaneously establishing an association record between the virtual viewpoint frame and a source real viewpoint frame. When a frame-by-frame playback operation is switched to a target playback step sequence, consistent judgment is performed on the virtual viewpoint frame according to a source frame identifier of a target real viewpoint frame, invalid virtual viewpoint frames are determined, corresponding virtual viewpoint frames are regenerated based on the target real viewpoint frame, and the association record is updated. A target cooperative frame group is further constructed and complete state checking is performed, and cooperative display is performed after the time sequence consistency requirement is met. Through source frame association updating and state checking, the application realizes multi-viewpoint synchronous cooperative playback.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video processing technology, specifically to a collaborative playback method based on multi-screen spatiotemporal synchronization frame by frame. Background Technology

[0002] With the development of multi-view video acquisition, 3D scene reconstruction, and virtual viewpoint image generation technologies, the practice of capturing multiple perspectives of the same scene from multiple cameras and generating virtual viewpoint images beyond the actual camera positions based on the video images corresponding to different perspectives has been gradually applied to scenarios such as sports replay, video production, scene analysis, and interactive viewing. Multi-view video systems typically synchronize video data acquired by multiple cameras in time, enabling video frames from different real viewpoints to establish a correspondence according to their respective acquisition times or frame times. Furthermore, virtual viewpoint images can be generated based on the video data from multiple real viewpoints at corresponding frame times.

[0003] During the playback and rewind of multi-view videos, in addition to continuous playback, the playback position can be controlled according to viewing, analysis, or editing needs. For example, in a paused state, the video can be forwarded frame by frame, rewound frame by frame, or positioned to a specific playback position to observe the same scene at different times and from different viewpoints. For application scenarios that simultaneously display multiple viewpoint images, the displayed images can include real viewpoint images captured by different actual cameras, or virtual viewpoint images generated based on video frames from one or more real viewpoints.

[0004] In this context, real-viewpoint images can be obtained based on the playback position of the corresponding video sequence, while virtual-viewpoint images require image processing or scene rendering based on their corresponding source image data. Therefore, during frame-by-frame playback, when the playback position changes from the current playback sequence to another, the data acquisition and image processing processes corresponding to different viewpoint images may have different processing paths and completion sequences. In multi-screen collaborative display scenarios, it is necessary to ensure that each real-viewpoint and virtual-viewpoint image participating in the same display corresponds to the currently observed playback sequence, so that each viewpoint can collectively represent the scene state corresponding to the same target playback sequence from different spatial locations.

[0005] In summary, how to ensure that each viewpoint corresponds to the same target playback sequence after switching playback sequences, thereby maintaining temporal consistency among multiple screens, is a technical problem that urgently needs to be solved in this field.

[0006] To address this, a collaborative playback method based on multi-screen spatiotemporal synchronization frame by frame is proposed. Summary of the Invention

[0007] The purpose of this invention is to provide a multi-view synchronous frame-by-frame collaborative playback method based on multi-screen spatiotemporal synchronization. This method achieves multi-viewpoint synchronous collaborative playback through source frame association updates and state verification. The method includes establishing the temporal correspondence and unified spatial reference of multiple real-viewpoint videos, obtaining the real-viewpoint frame corresponding to the current playback step, generating virtual-viewpoint frames, and simultaneously establishing an association record between the virtual-viewpoint frames and the source real-viewpoint frames. When switching to the target playback step in response to a frame-by-frame playback operation, the method performs a consistency judgment on the virtual-viewpoint frames based on the source frame identifier of the target real-viewpoint frame, identifies invalid virtual-viewpoint frames, regenerates the corresponding virtual-viewpoint frames based on the target real-viewpoint frames, and updates the association record. Furthermore, a target collaborative frame group is constructed and a complete state check is performed. Collaborative display is then performed after the temporal consistency requirements are met.

[0008] To achieve the above objectives, the present invention provides the following technical solution: A collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame includes: Establish the temporal correspondence and unified spatial reference of multiple real viewpoint videos, obtain the current real viewpoint frame under the current playback step, and associate it with the source frame identifier; A virtual viewpoint frame is generated based on the current real viewpoint frame, and a virtual frame source association record containing the corresponding source frame identifier is established. In response to the frame-by-frame step-by-step operation, switch to the target playback sequence and obtain the corresponding target real viewpoint frame set; compare the source frame identifier of each frame in the target real viewpoint frame set with the virtual frame source association record of each virtual viewpoint frame. If they do not match, the corresponding virtual viewpoint frame is determined to be a failed virtual viewpoint frame. Based on the corresponding target real viewpoint frame, the failed virtual viewpoint frame is regenerated, and the virtual frame source association record is updated to restore it to a valid state; according to the collaborative display viewpoint set, the target real viewpoint frame and the virtual viewpoint frame in a valid state are constructed into a target collaborative frame group, and a complete state check is performed; When the complete state check passes, the target collaborative frame group is submitted for collaborative display; when it fails, the current display state is maintained until the target collaborative frame group passes the complete state check.

[0009] Preferably, the process of obtaining the current real viewpoint frame in the current playback sequence and associating it with the source frame identifier includes: obtaining the multiple real viewpoint videos captured for the same target scene, and the camera spatial parameters corresponding to each real viewpoint; establishing a temporal correspondence relationship for frame-by-frame playback of the multiple real viewpoint videos, and establishing a unified spatial reference between each real viewpoint based on the camera spatial parameters; determining the current real viewpoint frame corresponding to the current playback sequence from each real viewpoint video according to the current playback sequence and the temporal correspondence relationship, forming a current real viewpoint frame set; and associating each current real viewpoint frame in the current real viewpoint frame set with a real viewpoint identifier, a source frame identifier, and a playback sequence identifier.

[0010] Preferably, the process of establishing a virtual frame source association record containing a corresponding source frame identifier includes: generating at least one virtual viewpoint frame located between real viewpoints based on at least two current real viewpoint frames in the current real viewpoint frame set and the camera spatial parameters corresponding to the current real viewpoint frames; establishing a corresponding virtual frame source association record when generating the virtual viewpoint frame, the virtual frame source association record including: virtual frame identifier, real viewpoint identifiers of each source real viewpoint actually called to generate the virtual viewpoint frame, source frame identifiers corresponding to each source real viewpoint, and playback step sequence identifiers corresponding to the virtual viewpoint frame.

[0011] Preferably, the process of obtaining the corresponding target real viewpoint frame set includes: in response to the frame-by-frame stepping operation performed on the multiple real viewpoint videos, switching from the current playback sequence to the target playback sequence; and determining the target real viewpoint frames corresponding to each real viewpoint under the target playback sequence according to the temporal correspondence, thereby obtaining the target real viewpoint frame set.

[0012] Preferably, the process of comparing the source frame identifiers of each frame in the target real viewpoint frame set with the virtual frame source association records of each virtual viewpoint frame includes: for any virtual viewpoint frame, comparing the source frame identifiers of each source real viewpoint recorded in its virtual frame source association record with the source frame identifiers of the target real viewpoint frames corresponding to the same real viewpoint in the target real viewpoint frame set; if the source frame identifier of at least one source real viewpoint recorded in the virtual frame source association record is inconsistent with the source frame identifier corresponding to the same real viewpoint in the target real viewpoint frame set, then the virtual viewpoint frame is determined as the failed virtual viewpoint frame, and all the failed virtual viewpoint frames are collected to form a failed virtual viewpoint frame set; if all the source frame identifiers recorded in the virtual frame source association record are consistent with the source frame identifiers of the corresponding real viewpoints in the target real viewpoint frame set, then the valid state of the virtual viewpoint frame is maintained.

[0013] Preferably, the process of updating the virtual frame source association record includes: for each failed virtual viewpoint frame in the failed virtual viewpoint frame set, determining each source real viewpoint corresponding to the failed virtual viewpoint frame based on the corresponding virtual frame source association record; calling the target real viewpoint frame corresponding to the source real viewpoint in the target playback sequence from the target real viewpoint frame set; regenerating the updated virtual viewpoint frame corresponding to the target playback sequence based on the called target real viewpoint frame and the corresponding camera spatial parameters; updating the source frame identifier of the target real viewpoint frame actually used by the virtual viewpoint frame, replacing the original source frame identifier in the virtual frame source association record, and updating the playback sequence identifier in the virtual frame source association record to the target playback sequence, so that the updated virtual viewpoint frame is restored to the valid state.

[0014] Preferably, the process of performing the complete state check includes: obtaining a preset set of collaborative display viewpoints, the set of collaborative display viewpoints being used to limit the viewpoints to be displayed simultaneously required for a multi-screen collaborative playback, the viewpoints to be displayed including real viewpoints and virtual viewpoints; based on the set of collaborative display viewpoints, establishing each of the target real viewpoint frames belonging to the target playback sequence and each of the virtual viewpoint frames in the effective state as a target collaborative frame group corresponding to the target playback sequence; performing the complete state check on the target collaborative frame group, the complete state check including: verifying whether each of the viewpoints to be displayed has a frame corresponding to the target playback sequence; and verifying whether the source frame identifiers recorded in the virtual frame source association record of each of the virtual viewpoint frames are consistent with the source frame identifiers of the corresponding real viewpoints in the target real viewpoint frame set.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. By establishing the association between virtual viewpoint frames and the real viewpoint source frames used for their generation, the system can determine the current state of the corresponding virtual viewpoint frame based on the changes in the real viewpoint source frames when the playback sequence changes. This avoids displaying virtual viewpoint frames that do not correspond to the target playback sequence after the playback sequence is switched, reduces the inconsistency in the timing of multiple screens caused by the failure to update the source frame synchronously, and makes the generation correspondence between virtual viewpoint frames and their source real viewpoint frames traceable.

[0016] 2. By determining the status of virtual viewpoint frames associated with changed source frames based on the real viewpoint frames corresponding to the target playback sequence, and regenerating virtual viewpoint frames that need to be updated, the update process of the virtual viewpoint screen corresponds to the actual changed source frames. This eliminates the need to process all current virtual viewpoint screens uniformly, thus establishing a correspondence between the update range of different viewpoint screens and the changes in the playback sequence, and providing corresponding frame data for the subsequent construction of collaborative frame groups under the same playback sequence.

[0017] 3. By performing a complete state check on the effective state of each real viewpoint frame and virtual viewpoint frame and its corresponding playback sequence before the target collaborative frame group enters the display state, the viewpoint screens participating in the same collaborative display have a consistent target playback sequence. At the same time, since the generation process of the target collaborative frame group is independent of the current display state maintenance process, the virtual viewpoint frame update process is avoided from directly affecting the continuity of the current collaborative display, so that multiple viewpoint screens corresponding to a frame-by-frame operation are organized and presented according to the same target playback sequence. Attached Figure Description

[0018] Figure 1 A schematic diagram of a collaborative playback method based on multi-screen spatiotemporal synchronization frame by frame provided by the present invention; Figure 2 This is a schematic diagram of the process for comparing associated records provided by the present invention. Figure 3 This is a schematic diagram of the execution complete status check process provided by the present invention. Detailed Implementation

[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0020] Please see Figures 1 to 3 This invention provides a collaborative playback method based on multi-screen spatiotemporal synchronization frame by frame, the technical solution of which is as follows: A collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame, the specific process of which is as follows: Figure 1 As shown, it includes: Establish the temporal correspondence and unified spatial reference of multiple real viewpoint videos, obtain the current real viewpoint frame under the current playback step, and associate it with the source frame identifier; A virtual viewpoint frame is generated based on the current real viewpoint frame, and a virtual frame source association record containing the corresponding source frame identifier is established. In response to the frame-by-frame step-by-step operation, switch to the target playback sequence and obtain the corresponding target real viewpoint frame set; compare the source frame identifier of each frame in the target real viewpoint frame set with the virtual frame source association record of each virtual viewpoint frame. If they do not match, the corresponding virtual viewpoint frame is determined to be a failed virtual viewpoint frame. Based on the corresponding target real viewpoint frame, the failed virtual viewpoint frame is regenerated, and the virtual frame source association record is updated to restore it to a valid state; according to the collaborative display viewpoint set, the target real viewpoint frame and the virtual viewpoint frame in a valid state are constructed into a target collaborative frame group, and a complete state check is performed; When the complete state check passes, the target collaborative frame group is submitted for collaborative display; when it fails, the current display state is maintained until the target collaborative frame group passes the complete state check.

[0021] Example 1: Furthermore, the process of obtaining the current real viewpoint frame under the current playback step sequence and associating it with the source frame identifier includes: obtaining the multiple real viewpoint videos captured for the same target scene, and the camera spatial parameters corresponding to each real viewpoint; establishing a temporal correspondence relationship for frame-by-frame playback of the multiple real viewpoint videos, and establishing a unified spatial reference between each real viewpoint based on the camera spatial parameters; determining the current real viewpoint frame corresponding to the current playback step sequence from each real viewpoint video according to the current playback step sequence and the temporal correspondence relationship, forming a current real viewpoint frame set; associating each current real viewpoint frame in the current real viewpoint frame set with a real viewpoint identifier, a source frame identifier, and a playback step sequence identifier.

[0022] First, multiple real-viewpoint videos of the same target scene are acquired. Each real-viewpoint video corresponds to an actual acquisition viewpoint, and a unique real-viewpoint identifier is pre-assigned to each actual acquisition viewpoint. For each real-viewpoint video, the corresponding camera spatial parameters are also acquired. The camera spatial parameters may include camera intrinsic parameters and extrinsic parameters used to characterize the position and attitude of the camera in a unified scene coordinate system. The camera extrinsic parameters can be obtained through pre-calibration and are represented in the form of rotation matrices and translation vectors to show the coordinate transformation relationship between the camera coordinate system of each real-viewpoint and the unified scene coordinate system.

[0023] A time-series correspondence is established for frame-by-frame playback of the acquired multi-channel real-viewpoint videos. Specifically, based on the timestamps, synchronization trigger numbers, or video frame numbers recorded during the acquisition phase of each video channel, image frames in different real-viewpoint videos are mapped to a unified playback sequence. When all videos are acquired using a unified synchronization trigger method, image frames with the same synchronization frame number can be identified as real-viewpoint frames corresponding to the same playback sequence. When each video channel is recorded using timestamps, a unified time base is first established based on the timestamps of each video channel, and the time information in each video channel is mapped to a unified time axis. A corresponding playback time sequence is generated according to the preset playback time interval, and for each playback... At each playback time point, image frames matching the playback time point are searched in each real viewpoint video stream. The image frame with the smallest time difference from the playback time point and within the preset time tolerance range is identified as the corresponding frame. Only when all real viewpoint videos participating in collaborative playback can identify corresponding frames that meet the preset time tolerance requirements, the playback sequence corresponding to the playback time point is identified as a valid playback sequence. If no real viewpoint video participating in collaborative playback has an image frame that meets the preset time tolerance requirements, then the playback time point does not form a valid playback sequence. After the time sequence correspondence is established, playback sequence identifiers are created for each valid playback sequence in the playback order.

[0024] Based on the established temporal correspondence, a unified spatial reference is established according to the camera spatial parameters corresponding to each real viewpoint. Specifically, the target scene coordinate system is selected as the unified spatial coordinate system, and the transformation relationship from the camera coordinate system of each real viewpoint to the unified spatial coordinate system is determined according to the external parameters of each camera. For any real viewpoint, its spatial position in the camera coordinate system can be transformed to the unified spatial coordinate system according to the corresponding rotation matrix and translation vector. Thus, the camera position, optical axis direction, and subsequent scene spatial data corresponding to different real viewpoints are all represented using the same spatial reference.

[0025] During frame-by-frame playback, the established temporal correspondence is queried based on the playback step sequence identifier corresponding to the current playback step sequence. The image frame corresponding to the playback step sequence identifier is then determined from each real viewpoint video. For example, when the current playback step sequence is the kth playback step sequence, the image frame corresponding to the kth playback step sequence is obtained from the first real viewpoint video, the second real viewpoint video, and so on up to the nth real viewpoint video. The obtained image frames are then combined to form the current real viewpoint frame set. If the timing relationship is established using the synchronization frame sequence number, the image frame with the corresponding synchronization frame sequence number k in each video is directly selected. If the timing relationship is established using the timestamp, the image frame with the established correspondence in each video is determined based on the target time corresponding to the kth playback step sequence.

[0026] For each current real viewpoint frame in the current set of real viewpoint frames, its real viewpoint identifier, source frame identifier, and playback step sequence identifier are associated respectively. The real viewpoint identifier is used to indicate the actual acquisition viewpoint to which the frame belongs; the source frame identifier is used to uniquely identify the specific image frame in the real viewpoint video. The source frame identifier can be a video frame sequence number, a synchronous acquisition sequence number, or an identifier formed by a combination of the real viewpoint identifier and the frame sequence number; the playback step sequence identifier is used to indicate the frame-by-frame playback position corresponding to the current real viewpoint frame. Thus, for any image frame in the current set of real viewpoint frames, its source viewpoint can be determined according to the associated real viewpoint identifier, its specific source frame in the corresponding real viewpoint video can be determined according to the source frame identifier, and its current playback step sequence can be determined according to the playback step sequence identifier.

[0027] Furthermore, the process of establishing a virtual frame source association record containing the corresponding source frame identifier includes: generating at least two current real viewpoint frames in the current real viewpoint frame set, and the camera spatial parameters corresponding to the current real viewpoint frames, generating at least one virtual viewpoint frame located between real viewpoints; establishing the corresponding virtual frame source association record when generating the virtual viewpoint frame, the virtual frame source association record including: virtual frame identifier, real viewpoint identifiers of each source real viewpoint actually called to generate the virtual viewpoint frame, the source frame identifiers corresponding to each source real viewpoint, and the playback step sequence identifier corresponding to the virtual viewpoint frame.

[0028] After obtaining the set of current real viewpoint frames corresponding to the current playback step sequence, at least two current real viewpoint frames that are spatially adjacent to the virtual viewpoint are selected from the set of current real viewpoint frames as source real viewpoint frames according to the preset virtual viewpoint position. The virtual viewpoint is set between adjacent real viewpoints, and its position and viewing direction can be preset according to the positional relationship of the cameras corresponding to the adjacent real viewpoints, or it can be determined according to the interpolation position between the centers of the cameras of adjacent real viewpoints. For each selected source real viewpoint frame, the camera intrinsic parameters and camera extrinsic parameters of its corresponding real viewpoint are obtained at the same time to determine the spatial projection relationship between each source real viewpoint and the target virtual viewpoint.

[0029] Specifically, for at least two selected source real viewpoint frames, spatial correspondence is performed on the scene content in different source real viewpoint frames based on the camera spatial parameters corresponding to each source real viewpoint. For real viewpoint videos with depth information, the depth data corresponding to the current source real viewpoint frame can be directly obtained. For cases where depth information is not provided in advance, the depth of the corresponding scene point can be determined based on the image correspondence between at least two source real viewpoint frames and the relative positional relationship between cameras. Specifically, firstly, a correspondence is established based on the image features between the source real viewpoint frames to obtain the corresponding image positions of the same scene point in different real viewpoints; then, based on the disparity relationship between the corresponding image positions and the spatial positional relationship between the corresponding real viewpoint cameras, the spatial depth of the scene point relative to the corresponding camera is determined. The image correspondence can be obtained through feature matching, region matching, or other methods capable of determining the correspondence between different viewpoints. When a reliable correspondence cannot be established for some regions, it can be supplemented based on nearby effective depth information. Specifically, the disparity of the scene point can be determined based on the image correspondence between two or more source real viewpoint frames, and the depth value of the scene point can be obtained by combining it with the spatial positional relationship between the corresponding real viewpoint cameras. For regions where a correspondence cannot be obtained through the adopted image correspondence method, if there are nearby scene points with determined depths that can be used for depth supplementation, the depth information of the nearby scene points can be used for supplementation. Alternatively, depending on the actual virtual viewpoint generation method used, image information from existing real viewpoint frames can be used for subsequent fusion processing. Based on the depth and camera spatial parameters of each source real viewpoint, spatial backprojection is performed on the pixel positions in the source real viewpoint frames to convert the two-dimensional image positions into three-dimensional scene points in a unified scene coordinate system by combining the corresponding depth. Subsequently, based on the position of the target virtual viewpoint, the viewing direction, and the virtual camera imaging parameters, the three-dimensional scene points are projected onto the image plane corresponding to the target virtual viewpoint to obtain the corresponding pixel positions of the three-dimensional scene points in the target virtual viewpoint. The spatial backprojection and projection processes are completed based on camera intrinsic and extrinsic parameters. The camera intrinsic parameters are used to describe the transformation relationship between image coordinates and camera coordinates, and the camera extrinsic parameters are used to describe the position and pose transformation relationship between camera coordinates and unified scene coordinates.

[0030] For the overlapping region formed after different source real viewpoints complete spatial transformation based on the depth information and project onto the target virtual viewpoint image plane, due to depth estimation errors, camera imaging differences, and projection sampling errors, there may be slight differences in corresponding pixels within the same scene area. Therefore, the fusion ratio of the corresponding projected image is further determined based on the spatial position of the target virtual viewpoint between two adjacent source real viewpoints to smoothly fuse the overlapping region. Specifically, the line connecting the camera centers of the first and second source real viewpoints is used as the viewpoint interpolation path, and the distance from the camera center of the first source real viewpoint to the target virtual viewpoint is set as follows. The distance from the center of the second source real viewpoint camera to the target virtual viewpoint is Then the fusion weight of the projected image corresponding to the first source real viewpoint Fusion weights of the projected image corresponding to the second source real viewpoint It can be determined as follows: Therefore, after completing the geometric projection based on depth information, the fusion ratio is used to adjust the pixel contribution of the effective projection results of different source real viewpoints in the overlapping region, so that the projection results of source real viewpoints closer to the target virtual viewpoint have higher weights, and the effective projection pixels in the overlapping region are fused according to the fusion weights. For example, when the target virtual viewpoint is located in the middle position of two source real viewpoints, the fusion weight of both source real viewpoints is one-half; when the target virtual viewpoint is located at one-third of the distance from the first source real viewpoint to the second source real viewpoint, the fusion weight of the first source real viewpoint is two-thirds, and the fusion weight of the second source real viewpoint is one-third. For the same pixel position in the virtual viewpoint image, when both source real viewpoints form effective projections, the two projected pixel values ​​are weighted and combined according to the above fusion weights. The fusion operation only applies to the pixel region that has completed effective spatial projection and does not change the spatial position relationship of scene points determined based on depth information. When there is only one When a source real viewpoint forms an effective projection, the pixel value of that effective projection is used. When neither of the two source real viewpoints forms an effective projection, the corresponding region in the target virtual viewpoint image is determined as a hole region. For the hole region, when there are effective projected pixels within a preset neighborhood around the pixel to be supplemented, the spatial distance between the effective projected pixels and the pixel to be supplemented can be obtained, and the pixel values ​​of multiple effective projected pixels are weighted according to the distance weight to obtain the pixel value corresponding to the pixel to be supplemented. For continuous hole regions, the pixels to be supplemented that meet the above supplementation conditions can be filled in a way that expands gradually from the edge of the hole inward. When there are no effective projected pixels within a preset neighborhood around the pixel to be supplemented, the above neighborhood filling process is not performed on the pixel to be supplemented. When three or more real viewpoints participate in the generation of virtual viewpoint frames, the two real viewpoints with the closest distance on both sides of the target virtual viewpoint can be selected according to the position of the target virtual viewpoint, and the fusion weight is determined and the projection fusion is completed in the above manner. The above virtual viewpoint frame generation process is used to illustrate an implementation method of obtaining a virtual viewpoint frame based on the real viewpoint frame corresponding to the current playback step sequence. In other embodiments, other virtual viewpoint generation methods that can obtain virtual viewpoint frames based on one or more real viewpoint frames and corresponding spatial parameters may also be adopted. The subsequent source frame association, failure determination, regeneration and collaborative display processes performed on the virtual viewpoint frame are based on the source real viewpoint frame actually called when the virtual viewpoint frame is generated.

[0031] While generating each virtual viewpoint frame, a corresponding virtual frame identifier is assigned to the virtual viewpoint frame, and a virtual frame source association record is established corresponding to the virtual viewpoint frame. When establishing the virtual frame source association record, the real viewpoint identifiers of the source real viewpoints that actually participated in the image generation process during this virtual viewpoint frame generation are obtained one by one, and the source frame identifiers that the source real viewpoint frame has been associated with in the current real viewpoint frame set are further obtained. Thus, the virtual frame identifier, the real viewpoint identifiers of each actually invoked source real viewpoint, and the source frame identifiers that correspond one-to-one with each real viewpoint identifier are written into the virtual frame source association record.

[0032] Simultaneously, the current playback step sequence identifier corresponding to the generation of the virtual viewpoint frame is written into the virtual frame source association record. For example, when a virtual viewpoint frame is jointly generated by the source frames of the first real viewpoint in the current playback step sequence and the source frames of the second real viewpoint in the current playback step sequence, the virtual frame source association record shall at least record the virtual frame identifier of the virtual viewpoint frame, the identifier of the first real viewpoint and its corresponding source frame identifier, the identifier of the second real viewpoint and its corresponding source frame identifier, and the current playback step sequence identifier. When three or more source real viewpoints actually participate in the generation of the virtual viewpoint frame, the corresponding real viewpoint identifier and source frame identifier shall be recorded in sequence according to the number of source real viewpoints that actually participate in the generation. Real viewpoints that do not actually participate in the generation of this virtual viewpoint frame shall not be written into the virtual frame source association record.

[0033] Furthermore, the process of obtaining the corresponding target real viewpoint frame set includes: in response to the frame-by-frame stepping operation performed on the multiple real viewpoint videos, switching from the current playback sequence to the target playback sequence; and determining the target real viewpoint frames corresponding to each real viewpoint under the target playback sequence according to the temporal correspondence, thereby obtaining the target real viewpoint frame set.

[0034] When multiple real-viewpoint videos are in frame-by-frame playback mode, in response to the frame-by-frame stepping operation performed on the multiple real-viewpoint videos, a target playback stepping order is determined based on the current playback stepping order and the direction of the frame-by-frame stepping operation. Specifically, when performing a forward frame-by-frame stepping operation, the next valid playback stepping order after the current playback stepping order is determined as the target playback stepping order; when performing a backward frame-by-frame stepping operation, the previous valid playback stepping order before the current playback stepping order is determined as the target playback stepping order. The target playback stepping order is uniformly determined for each real-viewpoint video, and each real-viewpoint video determines its subsequent target real-viewpoint frames based on this target playback stepping order.

[0035] After determining the target playback sequence, based on the pre-established temporal correspondence between multiple real viewpoint videos, the target real viewpoint frames corresponding to each real viewpoint under the target playback sequence are determined. The temporal correspondence is used to establish the correspondence between the playback sequence and the source image frames in each real viewpoint video. When the temporal correspondence is established using the synchronization frame number, the synchronization frame number corresponding to the target playback sequence is obtained, and the image frames with the synchronization frame number are determined from each real viewpoint video. When the temporal correspondence is established using the timestamp, the target time corresponding to the target playback sequence is obtained, and the image frames corresponding to the target time in each real viewpoint video are determined based on the pre-completed time alignment results.

[0036] When establishing a time-series correspondence using timestamps, the image frame with the smallest time difference from the target time and within a preset time tolerance range in each real viewpoint video is determined based on the acquisition timestamp of each image frame. This correspondence is then pre-associated with the corresponding playback sequence. Therefore, after performing frame-by-frame stepping, the time-series correspondence can be directly queried based on the target playback sequence to obtain the target real viewpoint frames for each real viewpoint, without needing to independently determine the playback position for each real viewpoint video. If any real viewpoint video participating in collaborative playback does not have an image frame that meets the preset time tolerance requirement at the target time corresponding to the target playback sequence, then that playback sequence is not considered a valid playback sequence and does not form a complete set of target real viewpoint frames.

[0037] After the corresponding target real viewpoint frames are determined for each real viewpoint, the target real viewpoint frames are collected to form a target real viewpoint frame set. For each target real viewpoint frame in the set, its corresponding real viewpoint identifier and source frame identifier are retained and associated with the target playback step sequence identifier. The real viewpoint identifier is used to determine the real acquisition viewpoint to which the target real viewpoint frame belongs, the source frame identifier is used to uniquely determine its source image frame in the corresponding real viewpoint video, and the target playback step sequence identifier is used to determine the target playback position of the frame in this frame-by-frame collaborative playback.

[0038] Further, the process of comparing the source frame identifiers of each frame in the target real viewpoint frame set with the virtual frame source association records of each virtual viewpoint frame includes: for any virtual viewpoint frame, comparing the source frame identifiers of each source real viewpoint recorded in its virtual frame source association record with the source frame identifiers of the target real viewpoint frames corresponding to the same real viewpoint in the target real viewpoint frame set; if the source frame identifier of at least one source real viewpoint recorded in the virtual frame source association record is inconsistent with the source frame identifier corresponding to the same real viewpoint in the target real viewpoint frame set, then the virtual viewpoint frame is determined as the invalid virtual viewpoint frame, and all invalid virtual viewpoint frames are collected to form an invalid virtual viewpoint frame set; if all the source frame identifiers recorded in the virtual frame source association record are consistent with the source frame identifiers of the corresponding real viewpoints in the target real viewpoint frame set, then the valid state of the virtual viewpoint frame is maintained, and the specific process is as follows. Figure 2 As shown.

[0039] After determining the target real viewpoint frame set, source frame identifier comparison is performed sequentially on each virtual viewpoint frame currently participating in collaborative playback. The comparison is based on the previously formed complete target real viewpoint frame set, so that each source real viewpoint recorded in the virtual frame source association record can determine the corresponding target real viewpoint frame in the target real viewpoint frame set.

[0040] For any virtual viewpoint frame to be compared, its corresponding virtual frame source association record is obtained, and the source real viewpoints recorded in the virtual frame source association record are matched item by item. Specifically, the real viewpoint identifier of each source real viewpoint is used as the matching basis, and target real viewpoint frames with the same real viewpoint identifier are searched in the target real viewpoint frame set. Thus, each source real viewpoint in the virtual frame source association record is established with the same real viewpoint in the target real viewpoint frame set, and no cross-comparison is performed between different real viewpoints.

[0041] After establishing the correspondence, the source frame identifier corresponding to the real viewpoint in the virtual frame source association record is obtained, and the source frame identifier associated with the target real viewpoint frame of the same real viewpoint in the target real viewpoint frame set is obtained. A consistency comparison is performed on the two source frame identifiers, using the same identifier format as when establishing the source frame identifiers. When the original frame sequence number, synchronization frame sequence number, or pre-assigned unique frame identifier is used as the source frame identifier, the corresponding identifier values ​​are compared for similarity. If the two source frame identifiers are the same, it is determined that the source image frame corresponding to the real viewpoint has not changed; if the two source frame identifiers are different, it is determined that the target source image frame corresponding to the real viewpoint in the target playback step sequence is different from the source image frame actually used when generating the current virtual viewpoint frame.

[0042] For a virtual viewpoint frame generated from two or more source real viewpoints, the above comparison is performed sequentially according to each source real viewpoint recorded in its virtual frame source association record. When the two source frame identifiers corresponding to any source real viewpoint are inconsistent, the virtual viewpoint frame is determined to be a failed virtual viewpoint frame. For virtual viewpoint frames where no source frame identifiers are inconsistent, the source frame identifiers corresponding to the remaining source real viewpoints are compared until all source real viewpoints recorded in the virtual frame source association record are compared. The virtual viewpoint frame is kept valid only when the source frame identifiers corresponding to all source real viewpoints are consistent with the source frame identifiers corresponding to the same real viewpoint in the target real viewpoint frame set.

[0043] The virtual viewpoint frames currently participating in collaborative playback are compared sequentially in the manner described above, and the virtual viewpoint frames that are determined to be invalid are collected to form a set of invalid virtual viewpoint frames; virtual viewpoint frames that remain valid are not added to the set of invalid virtual viewpoint frames.

[0044] Further, the process of updating the virtual frame source association record includes: for each failed virtual viewpoint frame in the failed virtual viewpoint frame set, determining each source real viewpoint corresponding to the failed virtual viewpoint frame based on the corresponding virtual frame source association record; calling the target real viewpoint frame corresponding to the source real viewpoint in the target playback sequence from the target real viewpoint frame set; regenerating the updated virtual viewpoint frame corresponding to the target playback sequence based on the called target real viewpoint frame and the corresponding camera spatial parameters; updating the source frame identifier of the target real viewpoint frame actually used by the virtual viewpoint frame, replacing the original source frame identifier in the virtual frame source association record, and updating the playback sequence identifier in the virtual frame source association record to the target playback sequence, so that the updated virtual viewpoint frame is restored to the valid state.

[0045] After obtaining the set of failed virtual viewpoint frames, for any one of the failed virtual viewpoint frames, obtain its corresponding virtual frame source association record, and determine the actual source real viewpoints used when generating the failed virtual viewpoint frame based on the real viewpoint identifiers recorded in the virtual frame source association record. Using the real viewpoint identifiers as the corresponding basis, determine the target real viewpoint frames corresponding to the same real viewpoint in the target playback sequence in the target real viewpoint frame set, so that each source real viewpoint of the original virtual viewpoint frame is respectively mapped to the new source image frame in the target playback sequence.

[0046] Based on the virtual frame identifier of the failed virtual viewpoint frame, the virtual viewpoint corresponding to the virtual viewpoint frame is determined, and the already determined position, viewing direction, and virtual camera parameters of the virtual viewpoint are obtained. During the regeneration process, the position, viewing direction, and virtual camera parameters of the virtual viewpoint remain unchanged, and the source real viewpoint frames used in the original generation process are replaced with the target real viewpoint frames corresponding to the same real viewpoint in the target real viewpoint frame set. Subsequently, based on the camera spatial parameters corresponding to each target real viewpoint frame and the spatial parameters of the virtual viewpoint, according to the aforementioned virtual viewpoint frame generation method, the scene content corresponding to each target real viewpoint frame is converted to a unified spatial reference and projected onto the image plane of the virtual viewpoint. The formed effective projections are fused to generate an updated virtual viewpoint frame corresponding to the target playback sequence.

[0047] After the updated virtual viewpoint frame is generated, the corresponding virtual frame source association record is updated according to the target real viewpoint frames actually used in this regeneration process. Specifically, the virtual frame identifier and the real viewpoint identifier of each source real viewpoint are retained in the virtual frame source association record. According to the correspondence between the real viewpoint identifiers, the source frame identifier associated with the target real viewpoint frame actually used in this process is obtained. The original source frame identifier of the same real viewpoint in the virtual frame source association record is replaced with the source frame identifier of the corresponding target real viewpoint frame. For the updated virtual viewpoint frame jointly generated by two or more source real viewpoints, the corresponding source frame identifier of each source real viewpoint that actually participated in the regeneration is updated.

[0048] After the source frame identifiers corresponding to each actual target real viewpoint frame have been updated, the playback step sequence identifier in the virtual frame source association record is updated to the target playback step sequence identifier, and the updated virtual frame source association record is kept in correspondence with the newly generated updated virtual viewpoint frame. When the updated virtual viewpoint frame has been generated and the source frame identifiers and target playback step sequence identifiers corresponding to each actual target real viewpoint frame have been recorded, the updated virtual viewpoint frame is updated from the invalid state to the valid state, and each invalid virtual viewpoint frame in the invalid virtual viewpoint frame set is processed in the above manner.

[0049] Further, the process of performing the complete state check includes: obtaining a preset set of collaborative display viewpoints, which is used to limit the viewpoints to be displayed simultaneously required for a multi-screen collaborative playback, including real viewpoints and virtual viewpoints; based on the set of collaborative display viewpoints, establishing each target real viewpoint frame belonging to the target playback sequence and each virtual viewpoint frame in the effective state as a target collaborative frame group corresponding to the target playback sequence; performing the complete state check on the target collaborative frame group, which includes: verifying whether each viewpoint to be displayed has a frame corresponding to the target playback sequence; and verifying whether the source frame identifiers recorded in the virtual frame source association record of each virtual viewpoint frame are consistent with the source frame identifiers of the corresponding real viewpoints in the target real viewpoint frame set, the specific process being as follows. Figure 3 As shown.

[0050] After obtaining the target set of real viewpoint frames and determining the valid state of each virtual viewpoint frame, a pre-defined set of collaborative display viewpoints is obtained. This set determines the viewpoints to be displayed that need to participate in a multi-screen collaborative playback. Each viewpoint to be displayed is associated with at least a corresponding viewpoint type and viewpoint identifier. The viewpoint type is used to distinguish between real and virtual viewpoints. For real viewpoints, the viewpoint identifier uses the corresponding real viewpoint identifier; for virtual viewpoints, the viewpoint identifier is used to determine the virtual viewpoint and its current corresponding virtual viewpoint frame. The collaborative display viewpoint set is determined before the frame-by-frame collaborative playback begins, according to the real and virtual viewpoints to be displayed.

[0051] Based on the collaborative display viewpoint set, the corresponding frames to be displayed are determined according to the viewpoint identifiers of each viewpoint to be displayed. For real viewpoints, their real viewpoint identifiers are used as the matching basis to obtain the target real viewpoint frame corresponding to the real viewpoint in the target playback sequence from the target real viewpoint frame set. For virtual viewpoints, their virtual viewpoint identifiers are used to determine the virtual viewpoint frame currently corresponding to the virtual viewpoint, and only the virtual viewpoint frames in a valid state are used as the corresponding frames to be displayed. The target real viewpoint frames and valid virtual viewpoint frames obtained in the above manner are combined according to their respective corresponding viewpoints to be displayed to form a target collaborative frame group for which a complete state check is to be performed.

[0052] The target collaborative frame group undergoes a frame existence verification. Specifically, the target collaborative frame group is checked item by item according to the viewpoints to be displayed in the collaborative display viewpoint set. For real viewpoints, it is checked whether there is a target real viewpoint frame corresponding to its real viewpoint identifier; for virtual viewpoints, it is checked whether there is a virtual viewpoint frame corresponding to its virtual viewpoint identifier and in a valid state. When a corresponding viewpoint to be displayed in the collaborative display viewpoint set can be identified as a viewpoint to be displayed, the target collaborative frame group is determined to have passed the frame existence verification; when no corresponding viewpoint to be displayed can be identified as a viewpoint to be displayed, the target collaborative frame group is determined to have failed the frame existence verification.

[0053] After the target collaborative frame group passes the frame existence verification, a source frame consistency verification is further performed on each virtual viewpoint frame within it. For any virtual viewpoint frame, its virtual frame source association record is obtained, and the real viewpoint identifiers of each source real viewpoint recorded therein are used as the matching basis to determine the target real viewpoint frame corresponding to the same real viewpoint in the target real viewpoint frame set. If any source real viewpoint cannot be matched with a corresponding target real viewpoint frame in the target real viewpoint frame set, then the virtual viewpoint frame fails the source frame consistency verification. If a corresponding target real viewpoint frame can be matched, then the source frame identifier corresponding to the source real viewpoint in the virtual frame source association record is compared item by item with the source frame identifier associated with the target real viewpoint frame, according to the aforementioned source frame identifier comparison method. Only when all source real viewpoints recorded in the virtual frame source association record can be matched with corresponding target real viewpoint frames, and all corresponding source frame identifiers are consistent, is the virtual viewpoint frame matched with the target real viewpoint frame.

[0054] The source frame consistency verification of all virtual viewpoint frames in the target collaborative frame group is completed in the above manner. When all viewpoints to be displayed in the collaborative display viewpoint set have corresponding frames to be displayed, and all virtual viewpoint frames in the target collaborative frame group pass the source frame consistency verification, the target collaborative frame group is determined to have passed the integrity check. If any viewpoint to be displayed does not have a corresponding frame, or any virtual viewpoint frame fails the source frame consistency verification, the target collaborative frame group is determined to have failed the integrity check.

[0055] Example 2: After switching to the target playback sequence, the method further includes: Based on the set of target real viewpoint frames, a target step sequence source frame snapshot corresponding to the target playback step sequence is established. The target step sequence source frame snapshot at least records the target playback step sequence identifier and the mapping relationship between each real viewpoint identifier and the corresponding source frame identifier, and assigns a snapshot identifier to the target step sequence source frame snapshot. For each of the failed virtual viewpoint frames to be regenerated, the corresponding source real viewpoint is determined based on its virtual frame source association record, and the target real viewpoint frame corresponding to the source real viewpoint is called from the target real viewpoint frame corresponding to the target step source frame snapshot to perform the regeneration of the failed virtual viewpoint frame. The snapshot identifier is further written into the virtual frame source association record corresponding to the regenerated updated virtual viewpoint frame, so that the target real viewpoint frame and the updated virtual viewpoint frame belonging to the same target cooperative frame group are associated with the same target step sequence source frame snapshot.

[0056] After switching from the current playback sequence to the target playback sequence and obtaining the target real viewpoint frame set, a target sequence source frame snapshot corresponding to the target playback sequence is established based on the target real viewpoint frame set. Specifically, the real viewpoint identifier and source frame identifier corresponding to each target real viewpoint frame in the target real viewpoint frame set are obtained. The real viewpoint identifier is used as the basis for distinguishing different acquisition viewpoints. A mapping relationship is established between each real viewpoint identifier and the source frame identifier corresponding to the real viewpoint under the target playback sequence. Each mapping relationship and the target playback sequence identifier are recorded together in the target sequence source frame snapshot. At the same time, a snapshot identifier is assigned to the target sequence source frame snapshot to distinguish it from other playback sequence source frame snapshots.

[0057] The target step sequence source frame snapshot is formed after the target real viewpoint frame set is determined. Before the virtual viewpoint frame corresponding to the target playback step sequence is regenerated and the target collaborative frame group is formed, the mapping relationship between each real viewpoint identifier and the source frame identifier in the target step sequence source frame snapshot remains unchanged. This mapping relationship is used as the frame selection basis when regenerating the virtual viewpoint frame under the current target playback step sequence. When the subsequent frame-by-frame step operation switches to another target playback step sequence, the corresponding target step sequence source frame snapshot is re-established according to the new target real viewpoint frame set, and a snapshot identifier different from the previous snapshot is assigned.

[0058] For any failed virtual viewpoint frame in the failed virtual viewpoint frame set, its virtual frame source association record is obtained, and the source real viewpoints required to regenerate the virtual viewpoint frame are determined according to the real viewpoint identifiers of each source real viewpoint recorded therein. For each source real viewpoint, its real viewpoint identifier is used to find the corresponding mapping in the current target step sequence source frame snapshot to obtain the source frame identifier corresponding to the real viewpoint in the target playback step sequence. Then, using the real viewpoint identifier and its corresponding source frame identifier as the matching basis, the corresponding target real viewpoint frame is determined in the target real viewpoint frame set. For each source real viewpoint involved in a failed virtual viewpoint frame, the target real viewpoint frame is determined in the above manner, and only the target real viewpoint frame determined by the current target step sequence source frame snapshot is used for regeneration.

[0059] If any source real viewpoint recorded in the virtual frame source association record of a failed virtual viewpoint frame cannot be identified in the current target step sequence source frame snapshot, or if the corresponding target real viewpoint frame cannot be identified in the target real viewpoint frame set based on the real viewpoint identifier and the source frame identifier, then no other real viewpoint frame corresponding to the playback step sequence will be used to replace it, and the failed state of the virtual viewpoint frame will be maintained; when each source real viewpoint can be identified in the current target step sequence source frame snapshot, an updated virtual viewpoint frame corresponding to the target playback step sequence will be generated according to the aforementioned virtual viewpoint frame regeneration method.

[0060] After the updated virtual viewpoint frame is generated, the source frame identifier and playback step sequence identifier in the corresponding virtual frame source association record are updated according to the actual target real viewpoint frames used this time. Furthermore, the snapshot identifier of the current target step sequence source frame snapshot is written into the virtual frame source association record. In the updated virtual frame source association record, the source frame identifier corresponding to each source real viewpoint should be consistent with the corresponding mapping relationship in the target step sequence source frame snapshot indicated by the snapshot identifier.

[0061] The mapping relationship between each real viewpoint identifier and the source frame identifier in the target step sequence source frame snapshot is used to determine the target real viewpoint frame that constitutes the snapshot. The updated virtual viewpoint frame is associated with the same target step sequence source frame snapshot through the snapshot identifier written in its virtual frame source association record. For each updated virtual viewpoint frame regenerated under the same target playback step sequence, the same snapshot identifier of the current target step sequence source frame snapshot is written.

[0062] The failed virtual viewpoint frame is regenerated using an asynchronous generation task, and the regeneration process further includes: For each of the asynchronous generation tasks, a task version identifier is established, which is associated with at least the virtual viewpoint identifier of the corresponding virtual viewpoint and the snapshot identifier when the asynchronous generation task is started; When any of the asynchronous generation tasks returns the corresponding updated virtual viewpoint frame, before using the updated virtual viewpoint frame to modify the virtual frame source association record, virtual viewpoint frame state, or the target collaborative frame group to be submitted, the task version identifier of the asynchronous generation task is obtained, and the associated virtual viewpoint identifier and snapshot identifier are determined based on the task version identifier; the snapshot identifier associated with the asynchronous generation task is compared with the snapshot identifier of the target step sequence source frame snapshot corresponding to the target collaborative frame group to be submitted, and the current valid task version identifier of the virtual viewpoint for the target step sequence source frame snapshot is determined based on the virtual viewpoint identifier; When the snapshot identifier associated with the asynchronous generation task is consistent with the snapshot identifier corresponding to the target collaborative frame group to be submitted, and the task version identifier of the asynchronous generation task is consistent with the current valid task version identifier of the corresponding virtual viewpoint, the updated virtual viewpoint frame is received, the corresponding virtual frame source association record is updated, and the updated virtual viewpoint frame is allowed to be added to the target collaborative frame group to be submitted. When the snapshot identifier associated with the asynchronous generation task is inconsistent with the snapshot identifier corresponding to the target collaborative frame group to be submitted, or when the task version identifier of the asynchronous generation task is inconsistent with the current valid task version identifier of the corresponding virtual viewpoint, the updated virtual viewpoint frame is determined to be an expired generation result and discarded. It is prohibited to use the expired generation result to modify the corresponding virtual frame source associated record, virtual viewpoint frame status, or the target collaborative frame group to be submitted.

[0063] For each failed virtual viewpoint frame in the set of failed virtual viewpoint frames that needs to be regenerated, an asynchronous generation task is established according to its corresponding virtual viewpoint. For any asynchronous generation task, when starting the asynchronous generation task, the virtual viewpoint identifier corresponding to the virtual viewpoint frame to be regenerated and the snapshot identifier of the target step sequence source frame snapshot corresponding to the current target playback step sequence are obtained, and the virtual viewpoint identifier and snapshot identifier are associated with the asynchronous generation task to form a corresponding task version identifier.

[0064] For the same virtual viewpoint, only one currently valid asynchronous generation task is retained for the same target step sequence source frame snapshot. Specifically, when creating an asynchronous generation task, a corresponding task version identifier is assigned to the task, and the task version identifier is associated with the virtual viewpoint identifier and the target step sequence source frame snapshot identifier. When a new asynchronous generation task is detected for the same virtual viewpoint corresponding to the same target step sequence source frame snapshot, the new task version identifier is updated to the currently valid version, and the old task is marked as an invalid task. After the asynchronous generation task is completed, only generation results with task version identifiers consistent with the currently valid version are allowed to be used to update the virtual viewpoint frame.

[0065] After establishing the asynchronous generation task, the corresponding target step sequence source frame snapshot is determined based on the snapshot identifier associated with the task version identifier. Then, based on the source real viewpoints recorded in the original virtual frame source association record of the failed virtual viewpoint frame, the source frame identifier corresponding to each source real viewpoint under the target playback step sequence is determined from the target step sequence source frame snapshot. Finally, the corresponding target real viewpoint frame is called according to the source frame identifier to regenerate the virtual viewpoint frame. From the start of the asynchronous generation task until the task returns the generation result, the asynchronous generation task always determines its frame retrieval basis according to the target step sequence source frame snapshot associated with the task version identifier at startup, and the target step sequence source frame snapshot used by the asynchronous generation task is not changed due to subsequent playback step sequence switching.

[0066] For the current target collaborative frame group to be submitted, it is associated with the target step sequence source frame snapshot corresponding to the current target playback step sequence, and the snapshot identifier of the target step sequence source frame snapshot is used as the snapshot identifier corresponding to the current target collaborative frame group to be submitted. When frame-by-frame stepping occurs again and a new target playback step sequence is switched, a new target step sequence source frame snapshot is established based on the new target real viewpoint frame set, and the currently formed target collaborative frame group to be submitted is associated with the new target step sequence source frame snapshot.

[0067] When any asynchronous generation task completes and returns an updated virtual viewpoint frame, before modifying the virtual frame source association record, virtual viewpoint frame status, or the current target collaborative frame group to be submitted using the updated virtual viewpoint frame, the task version identifier of the asynchronous generation task is first obtained, and the virtual viewpoint identifier associated with the asynchronous generation task and the snapshot identifier associated at startup are determined from the task version identifier; at the same time, the snapshot identifier of the target step sequence source frame snapshot corresponding to the current target collaborative frame group to be submitted is obtained.

[0068] First, the snapshot identifier associated with the asynchronous generation task is compared with the snapshot identifier corresponding to the current target collaborative frame group to be submitted. When the two snapshot identifiers are consistent, the current valid task version identifier of the virtual viewpoint for the current target step sequence source frame snapshot is further determined according to the virtual viewpoint identifier, and the task version identifier of the asynchronous generation task is compared with the current valid task version identifier.

[0069] The return result of the asynchronous generation task is determined to be usable for the current target collaborative frame group only if the snapshot identifier associated with the asynchronous generation task matches the snapshot identifier corresponding to the current target collaborative frame group to be submitted, and the task version identifier of the asynchronous generation task matches the current valid task version identifier of the corresponding virtual viewpoint. Based on the actual target viewpoint frames used in this asynchronous generation task, the source frame identifier and playback step sequence identifier in the virtual frame source association record corresponding to the virtual viewpoint are updated, and the snapshot identifier of the current target step sequence source frame snapshot is written into the virtual frame source association record. After completing the above record update, the updated virtual viewpoint frame is determined as the valid virtual viewpoint frame of the virtual viewpoint under the current target playback step sequence, and it is allowed to be added to the current target collaborative frame group to be submitted as the frame to be displayed corresponding to the virtual viewpoint.

[0070] When the snapshot identifier associated with the asynchronous generation task is inconsistent with the snapshot identifier corresponding to the current target collaborative frame group to be submitted, or when the task version identifier of the asynchronous generation task is inconsistent with the current valid task version identifier of the corresponding virtual viewpoint, the updated virtual viewpoint frame returned by it is determined as an expired generation result. For expired generation results, they are not used to replace the current virtual viewpoint frame, nor are they used to modify the corresponding virtual frame source association record and validity status, nor are they allowed to be added to the current target collaborative frame group to be submitted, thus completing the discarding of the expired generation results.

[0071] If the virtual viewpoint corresponding to the expired generation result still needs to be regenerated under the current target playback step sequence, then the target step sequence source frame snapshot associated with the current target collaborative frame group to be submitted is used as the frame retrieval basis. The corresponding updated virtual viewpoint frame is obtained through the asynchronous generation task corresponding to the current snapshot. When the asynchronous generation task returns, it is re-verified according to the above-mentioned double verification process of snapshot identifier and task version identifier to determine whether it can be used for the current target collaborative frame group.

[0072] The collaborative display of the target collaborative frame group is implemented using mutually isolated current display frame group storage areas and candidate frame group storage areas, and the process includes: The collaborative frame group currently being displayed is stored in the current display frame group storage area, and the target real viewpoint frame belonging to the target playback step sequence and the virtual viewpoint frame in the effective state are written into the candidate frame group storage area to construct the target collaborative frame group. Before the target collaborative frame group passes the complete state check, it is prohibited to replace any frame in the candidate frame group storage area with the frame of the corresponding viewpoint in the current display frame group storage area. When the complete state check passes, in the same display state switch, the display object will be switched from the entire collaborative frame group in the current display frame group storage area to the target collaborative frame group in the candidate frame group storage area. When the complete state check fails, the display state corresponding to the current display frame group storage area remains unchanged.

[0073] During multi-view collaborative playback, a mutually isolated current display frame group storage area and candidate frame group storage area are set up. The current display frame group storage area is used to maintain the collaborative frame groups that have been submitted for collaborative display, and the candidate frame group storage area is used to construct the target collaborative frame group to be submitted for display next. Each frame in the current display frame group storage area corresponds to a viewpoint to be displayed in the collaborative display viewpoint set. Before the target collaborative frame group in the candidate frame group storage area passes the complete state check and completes the overall switch, the current collaborative display always uses the collaborative frame group in the current display frame group storage area.

[0074] When the frame-by-frame stepping operation switches the playback state to the target playback step sequence, the target playback step sequence is used as the step sequence benchmark for constructing the target collaborative frame group in the candidate frame group storage area. If there are candidate frames corresponding to other previous playback step sequences in the candidate frame group storage area, the previous candidate frames are excluded from the target collaborative frame group before constructing the target collaborative frame group, so that the frames used for the complete state check in the candidate frame group storage area all correspond to the current target playback step sequence.

[0075] Subsequently, according to each viewpoint to be displayed in the collaborative display viewpoint set, the corresponding frame to be displayed is written to the candidate frame group storage area. For real viewpoints, the target real viewpoint frame belonging to the target playback step sequence is determined from the target real viewpoint frame set based on its real viewpoint identifier, and written to the position corresponding to the real viewpoint. For virtual viewpoints, the corresponding virtual viewpoint frame is determined based on its virtual viewpoint identifier, and only virtual viewpoint frames that are currently in a valid state and meet the version requirements of the current target playback step sequence are selected as candidate frames to be displayed. In the case of using target step sequence source frame snapshots, if the snapshot identifier in the virtual frame source association record corresponding to the virtual viewpoint frame is consistent with the snapshot identifier of the target step sequence source frame snapshot corresponding to the current target collaborative frame group, it is determined that it meets the version requirements of the current target playback step sequence.

[0076] For real viewpoints in the collaborative display viewpoint set that have not yet obtained a corresponding target real viewpoint frame, or virtual viewpoints that do not yet have a valid virtual viewpoint frame that meets the above requirements, the viewpoint to be displayed remains in an unready state in the current candidate target collaborative frame group. The original display frame in the current display frame group storage area is not used, nor are candidate frames from other previous playback steps used for filling. When a corresponding target real viewpoint frame is subsequently obtained, or when an asynchronous generation task returns an updated virtual viewpoint frame that is acceptable after version verification, the obtained frame is then written to the corresponding position of the viewpoint to be displayed in the candidate frame group storage area.

[0077] During the process of constructing the target collaborative frame group in the candidate frame group storage area, changes in the frames in the candidate frame group storage area do not change the frames in the current display frame group storage area. For any viewpoint frame newly written or updated in the candidate frame group storage area, before the target collaborative frame group passes the complete state check, the frame is not allowed to replace the frame currently being displayed at the same viewpoint in the current display frame group storage area. Therefore, when only some viewpoints to be displayed in the candidate target collaborative frame group have been updated, the original collaborative frame group in the current display frame group storage area continues to be displayed.

[0078] Once the target collaborative frame group is formed in the candidate frame group storage area, it is checked according to the aforementioned complete state check method. Specifically, firstly, each viewpoint to be displayed is checked item by item according to the collaborative display viewpoint set to see if there is a corresponding frame belonging to the current target playback sequence. If all viewpoints to be displayed have corresponding frames, then the source frame consistency check is performed on the virtual frame source association record of each virtual viewpoint frame. If any viewpoint to be displayed is missing a corresponding frame, or if any virtual viewpoint frame fails the source frame consistency check, it is determined that the current candidate target collaborative frame group has not passed the complete state check.

[0079] When a target cooperative frame group passes the complete state check, it is designated as the target cooperative frame group to be switched. From the time it is designated as the target cooperative frame group to be switched until the current display switch is completed, the frame composition of this target cooperative frame group remains unchanged, and no subsequent frames are used to replace any frame corresponding to the viewpoint to be displayed. Subsequently acquired frames are used to construct subsequent candidate frame groups according to their corresponding playback sequence and version.

[0080] For the target collaborative frame group to be switched, the entire collaborative frame group is taken as the unit of display submission. In the same display state switch, each viewpoint to be displayed in the collaborative display viewpoint set is switched from the original display frame corresponding to the current display frame group storage area to the corresponding frame in the target collaborative frame group to be switched. Before the overall display switch is completed, no display frame replacement is performed for any viewpoint to be displayed individually.

[0081] After the overall display switch is completed, the target collaborative frame group to be switched is determined as the new current display collaborative frame group, and the storage area carrying the target collaborative frame group takes on the role of the current display frame group storage area. The original current display collaborative frame group exits the current display state, and its corresponding storage area can be used as the candidate frame group storage area for the subsequent target playback sequence. When a new frame-by-frame step operation occurs, the construction of the candidate target collaborative frame group, the complete state check, and the overall display switch are re-executed according to the new target playback sequence.

[0082] When the target collaborative frame group in the candidate frame group storage area fails the complete state check, the above-mentioned overall display switch is not performed. The collaborative frame group in the current display frame group storage area continues to be the current display object. Subsequently, when the missing frame is obtained or the corresponding updated virtual viewpoint frame is written to the candidate frame group storage area after version verification, the complete state check is re-performed on the updated candidate target collaborative frame group. Only after the re-check passes will the target collaborative frame group be submitted as a whole for collaborative display in the above manner.

[0083] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.

Claims

1. A collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame, characterized in that, include: Establish the temporal correspondence and unified spatial reference of multiple real viewpoint videos, obtain the current real viewpoint frame under the current playback step, and associate it with the source frame identifier; A virtual viewpoint frame is generated based on the current real viewpoint frame, and a virtual frame source association record containing the corresponding source frame identifier is established. In response to the frame-by-frame step-by-step operation, switch to the target playback sequence and obtain the corresponding target real viewpoint frame set; compare the source frame identifier of each frame in the target real viewpoint frame set with the virtual frame source association record of each virtual viewpoint frame. If they do not match, the corresponding virtual viewpoint frame is determined to be a failed virtual viewpoint frame. Based on the corresponding target real viewpoint frame, the failed virtual viewpoint frame is regenerated, and the virtual frame source association record is updated to restore it to a valid state; according to the collaborative display viewpoint set, the target real viewpoint frame and the virtual viewpoint frame in a valid state are constructed into a target collaborative frame group, and a complete state check is performed; When the complete state check passes, the target collaborative frame group is submitted for collaborative display; when it fails, the current display state is maintained until the target collaborative frame group passes the complete state check.

2. The collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame as described in claim 1, characterized in that, The process of obtaining the current real viewpoint frame under the current playback step sequence and associating it with the source frame identifier includes: obtaining the multiple real viewpoint videos captured for the same target scene, and the camera spatial parameters corresponding to each real viewpoint; establishing a temporal correspondence relationship for frame-by-frame playback of the multiple real viewpoint videos, and establishing a unified spatial reference between each real viewpoint based on the camera spatial parameters; determining the current real viewpoint frame corresponding to the current playback step sequence from each real viewpoint video according to the current playback step sequence and the temporal correspondence relationship, forming a current real viewpoint frame set; associating each current real viewpoint frame in the current real viewpoint frame set with a real viewpoint identifier, a source frame identifier, and a playback step sequence identifier.

3. The collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame as described in claim 2, characterized in that, The process of establishing a virtual frame source association record containing a corresponding source frame identifier includes: generating at least two virtual viewpoint frames located between real viewpoints based on at least two current real viewpoint frames in the current real viewpoint frame set and the camera spatial parameters corresponding to the current real viewpoint frames; establishing a corresponding virtual frame source association record when generating the virtual viewpoint frame, the virtual frame source association record including: virtual frame identifier, real viewpoint identifiers of each source real viewpoint actually called to generate the virtual viewpoint frame, source frame identifiers corresponding to each source real viewpoint, and playback step sequence identifiers corresponding to the virtual viewpoint frame.

4. The collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame as described in claim 1, characterized in that, The process of obtaining the corresponding target real viewpoint frame set includes: in response to the frame-by-frame stepping operation performed on the multiple real viewpoint videos, switching from the current playback sequence to the target playback sequence; and determining the target real viewpoint frames corresponding to each real viewpoint under the target playback sequence according to the temporal correspondence, thereby obtaining the target real viewpoint frame set.

5. The collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame as described in claim 1, characterized in that, The process of comparing the source frame identifiers of each frame in the target real viewpoint frame set with the virtual frame source association records of each virtual viewpoint frame includes: for any virtual viewpoint frame, comparing the source frame identifiers of each source real viewpoint recorded in its virtual frame source association record with the source frame identifiers of the target real viewpoint frames corresponding to the same real viewpoint in the target real viewpoint frame set; if the source frame identifier of at least one source real viewpoint recorded in the virtual frame source association record is inconsistent with the source frame identifier corresponding to the same real viewpoint in the target real viewpoint frame set, then the virtual viewpoint frame is determined as the failed virtual viewpoint frame, and all the failed virtual viewpoint frames are collected to form a failed virtual viewpoint frame set; if all the source frame identifiers recorded in the virtual frame source association record are consistent with the source frame identifiers of the corresponding real viewpoints in the target real viewpoint frame set, then the valid state of the virtual viewpoint frame is maintained.

6. The collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame as described in claim 2, characterized in that, The process of updating the virtual frame source association record includes: for each failed virtual viewpoint frame in the failed virtual viewpoint frame set, determining the source real viewpoints corresponding to the failed virtual viewpoint frames based on the corresponding virtual frame source association record; calling the target real viewpoint frame corresponding to the source real viewpoint in the target playback sequence from the target real viewpoint frame set; regenerating the updated virtual viewpoint frame corresponding to the target playback sequence based on the called target real viewpoint frame and the corresponding camera spatial parameters; updating the source frame identifier of the target real viewpoint frame actually used by the virtual viewpoint frame, replacing the original source frame identifier in the virtual frame source association record, and updating the playback sequence identifier in the virtual frame source association record to the target playback sequence, so that the updated virtual viewpoint frame is restored to the valid state.

7. The collaborative playback method based on multi-screen spatiotemporal synchronization frame-by-frame as described in claim 1, characterized in that, The process of performing a complete state check includes: obtaining a preset set of collaborative display viewpoints, which is used to limit the viewpoints to be displayed simultaneously for a multi-screen collaborative playback, including real viewpoints and virtual viewpoints; based on the set of collaborative display viewpoints, establishing each target real viewpoint frame belonging to the target playback sequence and each virtual viewpoint frame in the effective state as a target collaborative frame group corresponding to the target playback sequence; performing a complete state check on the target collaborative frame group, which includes: verifying whether each viewpoint to be displayed has a frame corresponding to the target playback sequence; and verifying whether the source frame identifiers recorded in the virtual frame source association record of each virtual viewpoint frame are consistent with the source frame identifiers of the corresponding real viewpoints in the target real viewpoint frame set.