Data processing method, system, related device and storage medium
By generating and synthesizing virtual information images with augmented reality effects in multi-angle free-view videos, the problem of balancing low-latency playback and rich visual experience is solved, achieving precise integration of AR effects and low-latency playback, thus improving the user experience.
Patent Information
- Application Number
- CN202010522454.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-06-10
- Publication Date
- 2026-01-13
- Estimated Expiration
- 2040-06-10
AI Technical Summary
Existing technologies struggle to balance low-latency playback and a rich visual experience in multi-angle, free-view videos, especially in live or broadcast scenarios where users cannot freely switch viewpoints to obtain a rich visual experience.
By acquiring target objects from video frames of multi-angle free-viewpoint videos, virtual information images based on augmented reality effects are generated and synthesized with video frames. A distributed system architecture is used to extract synchronous video frames at specified times from multiple synchronous video streams for reconstruction and synthesis, reducing data processing and transmission volume and achieving low-latency AR effect implantation.
It enables precise and rapid integration of AR effects into multi-angle free-view videos, meeting users' needs for low latency and rich visual experience, reducing data processing and transmission latency, and improving users' immersion and interactive experience.
Smart Images

Figure CN113784148B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present specification relate to the technical field of data processing, and in particular to a data processing method, system, related equipment and storage medium. BACKGROUND
[0002] With the continuous development of interconnection technology, more and more video platforms continue to provide higher-definition or higher-smoothness videos to improve the viewing experience of users. However, for videos with strong live experience, such as a video of a sports game, users can only watch the game from one viewpoint during the viewing process and cannot freely switch the viewpoint to watch the game pictures or game processes at different viewpoint positions, so they cannot experience the feeling of moving the viewpoint while watching the game in the live.
[0003] 6 Degree of Freedom (6DoF) technology is a technology for providing high-freedom viewing experience. Users can adjust the viewing angle of the video by interacting during the viewing process to watch from the free viewpoint angle they want to watch, thereby greatly improving the viewing experience.
[0004] To further enhance the viewing experience of 6DoF videos, there is currently an augmented reality (AR) special effect implantation scheme based on multi-angle free-view technology. However, the existing scheme for implanting AR special effects into multi-angle free-view videos cannot achieve low-latency playback, so it cannot meet the needs of users for rich visual experience and low latency during the video viewing process. SUMMARY
[0005] To meet the needs of users for rich visual experience during the video viewing process, embodiments of the present specification provide a data processing method, system, related equipment and storage medium.
[0006] Embodiments of the present specification provide a data processing method, comprising:
[0007] obtaining a target object in a video frame of a multi-angle free-view video;
[0008] obtaining a virtual information image generated based on augmented reality special effect input data of the target object;
[0009] synthesizing and displaying the virtual information image and the corresponding video frame.
[0010] Optionally, the multi-angle free-view video is formed by combining image groups based on a plurality of synchronous video frames of a specified frame time taken from a plurality of synchronous video streams, corresponding parameter data, pixel data and depth data of preset frame images in the image groups, and frame image reconstruction of a preset virtual viewpoint path, wherein the plurality of synchronous video frames contain frame images of different shooting angles.
[0011] Optionally, the virtual information image generated based on the augmented reality special effect input data of the target object comprises:
[0012] Based on the position of the target object in the video frame of the multi-angle free-view video obtained by three-dimensional calibration, a virtual information image matched with the position of the target object is obtained.
[0013] Optionally, the virtual information image and the corresponding video frame are synthesized and displayed, comprising: according to frame time sorting and virtual viewpoint position of the corresponding frame time, the virtual information image of the corresponding frame time is synthesized and displayed with the video frame of the corresponding frame time.
[0014] Optionally, the virtual information image and the corresponding video frame are synthesized and displayed, comprising at least one of the following:
[0015] The virtual information image and the corresponding video frame are fused to obtain a fused video frame, and the fused video frame is displayed.
[0016] The virtual information image is superimposed on the corresponding video frame to obtain a superimposed and synthesized video frame, and the superimposed and synthesized video frame is displayed.
[0017] Optionally, the display of the fused video frame comprises: inserting the fused video frame into a video stream to be played for playing and displaying.
[0018] Optionally, the target object in the video frame of the multi-angle free-view video comprises: in response to a special effect generation interaction control instruction, the target object in the video frame of the multi-angle free-view video is obtained.
[0019] Optionally, the virtual information image generated based on the augmented reality special effect input data of the target object comprises: based on the augmented reality special effect input data of the target object, the virtual information image corresponding to the target object is generated according to a preset special effect generation mode.
[0020] The embodiments of the present specification also provide another data processing method, comprising:
[0021] Receiving a plurality of synchronous video frames at a specified frame moment intercepted from a plurality of synchronous video streams as an image combination, the plurality of synchronous video frames containing frame images of different shooting perspectives;
[0022] Determining parameter data corresponding to the image combination;
[0023] Determining depth data of each frame image in the image combination;
[0024] Based on the parameter data corresponding to the image combination, the pixel data and the depth data of the preset frame image in the image combination, performing frame image reconstruction on a preset virtual viewpoint path to obtain video frames of a corresponding multi-angle free-view video;
[0025] In response to a special effect generation instruction, obtaining a target object in a video frame specified by the special effect generation instruction, obtaining augmented reality special effect input data of the target object, and generating a corresponding virtual information image based on the augmented reality special effect input data of the target object;
[0026] Synthesizing the virtual information image with the specified video frame to obtain a synthesized video frame;
[0027] Displaying the synthesized video frame.
[0028] Optionally, the generating of the corresponding virtual information image based on the augmented reality special effect input data of the target object comprises:
[0029] Taking the augmented reality special effect input data of the target object as input, based on the position of the target object in the video frame of the multi-angle free-view video obtained by three-dimensional calibration, a preset first special effect generation method is used to generate a virtual information image matching the target object in the corresponding video frame.
[0030] Optionally, the obtaining of the target object in the video frame specified by the special effect generation instruction and the obtaining of the augmented reality special effect input data of the target object comprise:
[0031] According to a service end special effect generation interactive control instruction, determining a special effect output type;
[0032] Obtaining historical data of the target object, processing the historical data according to the special effect output type to obtain augmented reality special effect input data corresponding to the special effect output type.
[0033] Optionally, the generating of the corresponding virtual information image based on the augmented reality special effect input data of the target object comprises at least one of the following:
[0034] Input the augmented reality special effect input data of the target object into a preset three-dimensional model, and output a virtual information image matched with the target object based on the position of the target object in the video frames of the multi-angle free-view video obtained through three-dimensional calibration.
[0035] Input the augmented reality special effect input data of the target object into a preset machine learning model, and output a virtual information image matched with the target object based on the position of the target object in the video frames of the multi-angle free-view video obtained through three-dimensional calibration.
[0036] Optionally, the synthesizing processing of the virtual information image and the specified video frame to obtain a synthesized video frame comprises:
[0037] Fusing processing of the virtual information image and the specified video frame based on the position of the target object in the specified video frame obtained through three-dimensional calibration to obtain a fused video frame.
[0038] Optionally, the displaying of the synthesized video frame comprises inserting the synthesized video frame into a video stream to be played of a playing control device for playing through a playing terminal.
[0039] Optionally, the method further comprises:
[0040] Generating a spliced image corresponding to the image combination based on the pixel data and the depth data of the image combination, the spliced image comprising a first field and a second field, wherein the first field comprises the pixel data of a preset frame image in the image combination, and the second field comprises the depth data of the image combination.
[0041] Storing the spliced image of the image combination and the parameter data corresponding to the image combination.
[0042] In response to an image reconstruction instruction from an interaction terminal, determining interaction frame time information of an interaction time, obtaining the spliced image of a preset frame image in an image combination corresponding to the interaction frame time and the parameter data corresponding to the image combination, and sending the spliced image and the parameter data to the interaction terminal, so that the interaction terminal selects corresponding pixel data and depth data in the spliced image and corresponding parameter data based on virtual viewpoint position information determined through an interaction operation, combines and renders the selected pixel data and depth data, and reconstructs and plays a video frame of a multi-angle free-view video corresponding to the virtual viewpoint position at the interaction frame time.
[0043] Optionally, the method further comprises:
[0044] In response to the service end special effect generation interaction control instruction, a virtual information image corresponding to the spliced image of the preset video frame indicated by the service end special effect generation interaction control instruction is generated;
[0045] The virtual information image corresponding to the spliced image of the preset video frame is stored.
[0046] Optionally, after receiving the image reconstruction instruction, the method further includes:
[0047] In response to a user end special effect generation interaction instruction from the interaction terminal, a virtual information image corresponding to the spliced image of the preset video frame is obtained;
[0048] The virtual information image corresponding to the spliced image of the preset video frame is sent to the interaction terminal, so that the interaction terminal performs synthesis processing on a video frame of a multi-angle free perspective video corresponding to a virtual viewpoint position at the interaction frame time and the virtual information image, obtains a synthesis video frame, and displays the synthesis video frame.
[0049] Optionally, the method further includes: in response to a user end special effect exit interaction instruction, the obtaining of the virtual information image corresponding to the spliced image of the preset video frame is stopped.
[0050] Optionally, the obtaining of the virtual information image corresponding to the spliced image of the preset video frame in response to the user end special effect generation interaction instruction from the interaction terminal includes:
[0051] Based on the user end special effect generation interaction instruction, a target object corresponding to the spliced image of the preset video frame is determined;
[0052] A virtual information image matching the target object in the preset video frame is obtained.
[0053] Optionally, the obtaining of the virtual information image matching the target object in the preset video frame includes:
[0054] A virtual information image matching the target object, which is generated based on a position of the target object in the preset video frame obtained by three-dimensional calibration in advance, is obtained.
[0055] Optionally, the sending of the virtual information image corresponding to the spliced image of the preset video frame to the interaction terminal, so that the interaction terminal performs synthesis processing on a video frame of a multi-angle free perspective video corresponding to a virtual viewpoint position at the interaction frame time and the virtual information image to obtain a synthesis video frame, includes:
[0056] The virtual information image corresponding to the spliced image of the preset video frame is sent to the interactive terminal, so that the interactive terminal superimposes the virtual information image on a video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the interactive frame moment, to obtain a superimposed and synthesized video frame.
[0057] The embodiment of the present specification further provides another data processing method, comprising:
[0058] In response to an image reconstruction instruction from the interactive terminal, determine interactive frame moment information of an interactive moment, obtain a spliced image of a preset frame image in image combination corresponding to the interactive frame moment and parameter data corresponding to the image combination and send to the interactive terminal, so that the interactive terminal selects corresponding pixel data and depth data in the spliced image and corresponding parameter data according to a preset rule based on virtual viewpoint position information determined by the interactive operation, combines and renders the selected pixel data and depth data, reconstructs a video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the interactive frame moment and plays it;
[0059] In response to a special effect generation interactive control instruction, obtain a virtual information image corresponding to a spliced image of a preset video frame indicated by the special effect generation interactive control instruction;
[0060] The virtual information image corresponding to the spliced image of the preset video frame is sent to the interactive terminal, so that the interactive terminal combines and processes the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the interactive frame moment and the virtual information image to obtain a synthesized video frame;
[0061] The synthesized video frame is displayed.
[0062] Optionally, the spliced image of the preset video frame is generated based on pixel data and depth data of the image combination at the interactive frame moment, and the spliced image comprises a first field and a second field, wherein the first field comprises pixel data of the preset frame image in the image combination, and the second field comprises depth data of the image combination.
[0063] The image combination at the interactive frame moment is obtained based on a plurality of synchronous video frames of a specified frame moment intercepted from a plurality of synchronous video streams, and the plurality of synchronous video frames contain frame images of different shooting angles.
[0064] Optionally, the virtual information image corresponding to the spliced image of the preset video frame indicated by the special effect generation interactive control instruction is obtained in response to the special effect generation interactive control instruction, comprising:
[0065] In response to a special effect generation interactive control instruction, obtain a target object in a video frame indicated by the special effect generation interactive control instruction;
[0066] obtaining a virtual information image generated in advance based on the augmented reality special effect input data of the target object.
[0067] The embodiment of the present specification also provides another data processing method, comprising:
[0068] real-time display of video frames of the multi-angle free-view video;
[0069] in response to a trigger operation on a special effect display identifier in the video frame of the multi-angle free-view video, obtaining a virtual information image of a video frame corresponding to a specified frame moment of the special effect display identifier;
[0070] synthesizing and displaying the virtual information image and the corresponding video frame.
[0071] Optionally, the method further comprises:
[0072] obtaining a virtual information image of a target object in a video frame corresponding to a specified frame moment of the special effect display identifier.
[0073] Optionally, the method further comprises:
[0074] based on a position of the target object in the video frame at the specified frame moment determined by three-dimensional calibration, superimposing the virtual information image on the video frame at the specified frame moment to obtain a superimposed and synthesized video frame and displaying the superimposed and synthesized video frame.
[0075] The embodiment of the present specification provides a data processing system, comprising:
[0076] a target object obtaining unit adapted to obtain a target object in a video frame of a multi-angle free-view video;
[0077] a virtual information image obtaining unit adapted to obtain a virtual information image generated based on augmented reality special effect input data of the target object;
[0078] an image synthesizing unit adapted to synthesize the virtual information image and the corresponding video frame to obtain a synthesized video frame;
[0079] a display unit adapted to display the obtained synthesized video frame.
[0080] The embodiment of the present specification provides another data processing system, comprising: a data processing device, a server, a play control device and a play terminal, wherein:
[0081] The data processing device is adapted to capture multiple synchronous video frames at a specified frame moment from multiple video data streams collected in real time from different positions in the live collection area based on a video frame capture instruction, and upload the multiple synchronous video frames at the specified frame moment obtained to the server;
[0082] The server is adapted to receive the multiple synchronous video frames uploaded by the data processing device as an image combination, determine parameter data corresponding to the image combination and depth data of each frame image in the image combination, and perform frame image reconstruction on a preset virtual viewpoint path based on the parameter data corresponding to the image combination, pixel data and depth data of a preset frame image in the image combination, to obtain video frames of a corresponding multi-angle free-view video; and in response to a special effect generation instruction, obtain a target object in a video frame specified by the special effect generation instruction, obtain augmented reality special effect input data of the target object, and generate a corresponding virtual information image based on the augmented reality special effect input data of the target object, synthesize the virtual information image with the specified video frame to obtain a synthesized video frame, and input the synthesized video frame to a play control device.
[0083] The play control device is adapted to insert the synthesized video frame into a video stream to be played.
[0084] The play terminal is adapted to receive the video stream to be played from the play control device and play in real time.
[0085] Optionally, the system further comprises an interaction terminal; wherein:
[0086] The server is further adapted to generate a spliced image corresponding to the image combination based on pixel data and depth data of the image combination, the spliced image comprising a first field and a second field, wherein the first field comprises pixel data of a preset frame image in the image combination, and the second field comprises depth data of the image combination; store the spliced image of the image combination and the parameter data corresponding to the image combination; and in response to an image reconstruction instruction from the interaction terminal, determine interaction frame moment information at an interaction moment, obtain the spliced image of the preset frame image in the image combination at the corresponding interaction frame moment and the parameter data corresponding to the image combination, and send them to the interaction terminal.
[0087] The interaction terminal is adapted to send the image reconstruction instruction to the server based on an interaction operation, and select corresponding pixel data and depth data in the spliced image and corresponding parameter data according to preset rules based on virtual viewpoint position information determined by the interaction operation, combine and render the selected pixel data and depth data, reconstruct video frames of a multi-angle free-view video corresponding to the virtual viewpoint position at the interaction frame moment, and play the video frames.
[0088] Optionally, the server is further adapted to generate an interactive control instruction according to a server-side special effect, and store a virtual information image corresponding to a spliced image of a preset video frame indicated by the server-side special effect generation interactive control instruction.
[0089] Optionally, the server is further adapted to, in response to a user-side special effect generation interactive instruction from an interactive terminal, acquire a virtual information image corresponding to a spliced image of a preset video frame, and send the virtual information image corresponding to the spliced image of the preset video frame to the interactive terminal.
[0090] The interactive terminal is adapted to synthesize a video frame of a multi-angle free-view video corresponding to a virtual viewpoint position at the interactive frame moment and the virtual information image to obtain a synthesized video frame and perform playing and display.
[0091] The embodiments of the present specification provide a server, comprising:
[0092] A data receiving unit is adapted to receive a plurality of synchronous video frames at a specified frame moment intercepted from a plurality of synchronous video streams as an image combination, wherein the plurality of synchronous video frames contain frame images of different shooting angles.
[0093] A parameter data calculation unit is adapted to determine parameter data corresponding to the image combination.
[0094] A depth data calculation unit is adapted to determine depth data of each frame image in the image combination.
[0095] A video data obtaining unit is adapted to perform frame image reconstruction on a preset virtual viewpoint path based on the parameter data corresponding to the image combination, pixel data and depth data of a preset frame image in the image combination, to obtain a video frame of a multi-angle free-view video.
[0096] A first virtual information image generation unit is adapted to, in response to a special effect generation instruction, acquire a target object in a video frame specified by the special effect generation instruction, acquire augmented reality special effect input data of the target object, and generate a corresponding virtual information image based on the augmented reality special effect input data of the target object.
[0097] An image synthesis unit is adapted to synthesize the virtual information image and the specified video frame to obtain a synthesized video frame.
[0098] A first data transmission unit is adapted to output the synthesized video frame to be inserted into a to-be-played video stream.
[0099] Optionally, the first virtual information image generation unit is adapted to input augmented reality special effect input data of the target object as input, based on the position of the target object in the video frames of the multi-angle free perspective video obtained by three-dimensional calibration, and generate a virtual information image matched with the target object in a corresponding video frame by using a preset first special effect generation manner.
[0100] The embodiment of the present specification provides another server, comprising:
[0101] The image reconstruction unit is adapted to determine the interaction frame time information of the interaction time, acquire the spliced image of the preset frame image in the image combination corresponding to the interaction frame time and the corresponding parameter data of the image combination in response to the image reconstruction instruction from the interaction terminal;
[0102] The virtual information image generation unit is adapted to generate a virtual information image corresponding to the spliced image of the image combination of the video frame indicated by the special effect generation interaction control instruction in response to the special effect generation interaction control instruction;
[0103] The data transmission unit is adapted to perform data interaction with the interaction terminal, comprising: transmitting the spliced image of the preset video frame in the image combination corresponding to the interaction frame time and the corresponding parameter data of the image combination to the interaction terminal, so that the interaction terminal selects the corresponding pixel data and depth data in the spliced image and the corresponding parameter data according to a preset rule based on the virtual viewpoint position information determined by the interaction operation, combines and renders the selected pixel data and depth data, reconstructs the image of the multi-angle free perspective video corresponding to the virtual viewpoint position at the interaction frame time, and plays the image; and transmitting the virtual information image corresponding to the spliced image of the preset frame image indicated by the special effect generation interaction control instruction to the interaction terminal, so that the interaction terminal synthesizes the video frame of the multi-angle free perspective video corresponding to the virtual viewpoint position at the interaction frame time with the virtual information image, obtains a multi-angle free perspective synthesis video frame, and plays the multi-angle free perspective synthesis video frame.
[0104] The embodiment of the present specification further provides an interaction terminal, comprising:
[0105] The first display unit is adapted to display the image of the multi-angle free perspective video in real time, wherein the image of the multi-angle free perspective video is reconstructed by the parameter data of the image combination, the pixel data and the depth data of the image combination of a plurality of synchronous video frame images at a specified frame time, and the plurality of synchronous video frames include frame images of different shooting angles.
[0106] The special effect data acquisition unit is adapted to acquire a virtual information image corresponding to a specified frame time of a special effect display identifier in the multi-angle free perspective video image in response to a triggering operation of the special effect display identifier.
[0107] The second display unit is adapted to overlay the virtual information image onto the video frame of the multi-angle free-view video.
[0108] This specification provides an electronic device including a memory and a processor. The memory stores computer instructions that can be executed on the processor. When the processor executes the computer instructions, it performs the steps of the method described in any of the foregoing embodiments.
[0109] This specification provides a computer-readable storage medium storing computer instructions that, when executed, perform the steps of the method described in any of the foregoing embodiments.
[0110] Compared with the prior art, the technical solutions of the embodiments in this specification have the following beneficial effects:
[0111] Using the data processing schemes in some embodiments of this specification, during the real-time playback of multi-angle free-viewpoint videos, the target object in the video frame of the multi-angle free-viewpoint video is obtained, and then a virtual information image generated based on the augmented reality effect input data of the target object is obtained. The virtual information image is then synthesized with the corresponding video frame and displayed. Through this process, only the video frame requiring AR effects needs to be synthesized with the virtual information image corresponding to the target object in the video frame during the playback of the multi-angle free-viewpoint video to obtain a video frame incorporating AR effects. It is not necessary to pre-generate all the video frames incorporating AR effects for a multi-angle free-viewpoint video before playback. Therefore, it is possible to accurately and quickly embed AR effects in multi-angle free-viewpoint videos, meeting users' needs for low-latency video viewing and rich visual experiences.
[0112] Furthermore, since the multi-angle free-viewpoint video is reconstructed based on the parameter data of the image combination formed by multiple synchronous video frames at a specified frame time from multiple synchronous video streams, along with the pixel data and depth data of the preset frame time in the image combination, a preset virtual viewpoint path is obtained. This eliminates the need for reconstruction based on all video frames in the multiple synchronous video streams, thus reducing data processing and transmission volume, and lowering the transmission latency of the multi-angle free-viewpoint video.
[0113] Furthermore, based on the position of the target object in the video frame of the multi-angle free-view video obtained by 3D calibration, a virtual information image matching the position of the target object is obtained. This makes the obtained virtual information image more consistent with the position of the target object in 3D space, and thus the displayed virtual information image is more in line with the real state in 3D space. As a result, the displayed synthetic video frame is more realistic and vivid, which can enhance the user's visual experience.
[0114] Furthermore, as the virtual viewpoint changes, the target object dynamically changes in the multi-angle free-view video. Therefore, by sorting the frames according to their times and the virtual viewpoint positions at those times, the virtual information image at the corresponding frame time is synthesized with the video frame at the object frame time and then displayed. In the resulting synthesized video frame, the virtual information image can change synchronously with the target object in the image frame of the multi-angle free-view video, making the synthesized video frame more realistic and vivid, enhancing the user's immersion in watching the multi-angle free-view video, and further improving the user experience.
[0115] Using the data processing schemes in some embodiments of this specification, for an image combination formed by multiple synchronous video frames at a specified frame time captured from multiple video streams, by determining the corresponding parameter data of the image combination and the depth data of each frame image in the image combination, on the one hand, based on the corresponding parameter data of the image combination, the pixel data and depth data of the preset frame images in the image combination, frame image reconstruction is performed on the preset virtual viewpoint path to obtain the video frame of the corresponding multi-angle free-view video image; on the other hand, in response to the special effects generation instruction, the target object in the video frame specified by the special effects generation instruction is obtained, the augmented reality special effects input data of the target object is obtained, and based on the augmented reality special effects input data of the target object, a corresponding virtual information image is generated, and the virtual information image is synthesized with the specified video frame to obtain a synthesized video frame and display it. In this data processing process, since only synchronous video frames at specified times are extracted from multiple synchronous video streams to reconstruct multi-angle free-view video, and virtual information images corresponding to the target objects in the video frames specified by the special effects generation instructions are generated, there is no need to upload massive amounts of synchronous video stream data. This distributed system architecture can save a lot of transmission and server processing resources. Moreover, under the condition of limited network transmission bandwidth, it can realize the real-time generation of synthetic video frames with augmented reality effects. Therefore, it can achieve low-latency playback of videos with multi-angle free-view augmented reality effects, thus meeting the dual needs of users for rich visual experience and low latency during video viewing.
[0116] Furthermore, the capture of synchronous video frames, the reconstruction of multi-angle free-view video, the generation of virtual information images, and the synthesis of multi-angle free-view video and virtual information images are all completed by different devices. This distributed system architecture can avoid a large amount of data processing on the same device, thus improving data processing efficiency and reducing transmission latency.
[0117] By employing some data processing schemes in the embodiments of this specification, in response to the special effects generation interactive control command, the virtual information image corresponding to the spliced image of the preset video frame indicated by the special effects generation interactive control command is obtained and sent to the interactive terminal. This allows the interactive terminal to synthesize the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the time of the interactive frame with the virtual information image, thereby obtaining and displaying the synthesized video frame. This can meet the user's needs for rich visual experience and real-time interaction, and enhance the user's interactive experience. Attached Figure Description
[0118] Figure 1 This specification shows a schematic diagram of the structure of a data processing system in a specific application scenario according to an embodiment of the present specification;
[0119] Figure 2 A flowchart of a data processing method according to an embodiment of this specification is shown;
[0120] Figure 3 A schematic diagram of the structure of a data processing system according to an embodiment of this specification is shown;
[0121] Figure 4 A flowchart of another data processing method in an embodiment of this specification is shown;
[0122] Figure 5 This document illustrates a schematic diagram of a video frame image in an embodiment of this specification.
[0123] Figure 6 A schematic diagram of a three-dimensional calibration method in an embodiment of this specification is shown;
[0124] Figure 7 A flowchart of another data processing method in an embodiment of this specification is shown;
[0125] Figures 8 to 12 A schematic diagram of the interactive interface of an interactive terminal according to an embodiment of this specification is shown;
[0126] Figure 13 A schematic diagram of the interactive interface of another interactive terminal in an embodiment of this specification is shown;
[0127] Figure 14 A flowchart of another data processing method is shown in an embodiment of this specification;
[0128] Figure 15 A schematic diagram of another data processing system in an embodiment of this specification is shown;
[0129] Figure 16 A schematic diagram of another data processing system in an embodiment of this specification is shown;
[0130] Figure 17 A schematic diagram of a server cluster architecture is shown in one embodiment of this specification;
[0131] Figures 18 to 20 This diagram illustrates the video effect of a playback interface of a playback terminal according to an embodiment of this specification.
[0132] Figure 21 A schematic diagram of the structure of another interactive terminal in an embodiment of the present invention is shown;
[0133] Figure 22 A schematic diagram of the structure of another interactive terminal in an embodiment of the present invention is shown;
[0134] Figures 23 to 26 A video effect diagram of the display interface of an interactive terminal in an embodiment of this specification is shown;
[0135] Figure 27 A schematic diagram of the structure of a server according to an embodiment of this specification is shown;
[0136] Figure 28 A schematic diagram of the structure of a server according to an embodiment of this specification is shown;
[0137] Figure 29 A schematic diagram of another server structure is shown in an embodiment of this specification. Detailed Implementation
[0138] In traditional live, rebroadcast, and recorded broadcast scenarios, users can only watch the game from one viewpoint and cannot freely switch viewpoints to view the game footage or process from different angles. Therefore, they cannot experience the feeling of watching the game while moving their viewpoints at the scene.
[0139] The use of 6 degrees of freedom (6DoF) technology can provide a highly flexible viewing experience. Users can adjust the viewing angle of the video through interactive means during the viewing process, and watch from the free viewpoint they want, thereby greatly improving the viewing experience.
[0140] With users' growing demand for richer visual experiences, the need has emerged to embed AR effects into videos. Currently, there are solutions for embedding AR effects into 2D or 3D videos. However, because multi-angle, free-viewpoint videos and AR effect data involve a large amount of image processing, rendering operations, and the transmission of massive amounts of video data, and because people are highly sensitive to latency in their video viewing experience, such as in live or near-live streaming scenarios where low-latency video playback is required, it is difficult to simultaneously meet users' needs for low-latency video playback and a rich visual experience.
[0141] To help those skilled in the art better understand the playback scenario of low-latency, multi-angle free-view video, a data processing system capable of realizing multi-angle free-view video playback is described below. Using this data processing system, low-latency playback of multi-angle free-view video is possible, and it can be applied to live streaming, broadcasting, and user-interactive video playback scenarios.
[0142] See Figure 1 The diagram illustrates the structure of a data processing system in a specific application scenario, showing the setup of a data processing system for a basketball game. The data processing system 10 includes a data acquisition array 11 composed of multiple acquisition devices, a data processing device 12, a cloud server cluster 13, a playback control device 14, a playback terminal 15, and an interactive terminal 16. Using the data processing system 10, multi-angle free-view video reconstruction can be achieved, allowing users to watch low-latency multi-angle free-view video.
[0143] Specifically, refer to Figure 1 Using the basketball hoop on the left as the focal point, and centering on the focal point, a fan-shaped area on the same plane as the focal point serves as the preset multi-angle free viewing angle range. Each acquisition device in the acquisition array 11 can be positioned in a fan shape at different locations within the on-site acquisition area according to the preset multi-angle free viewing angle range, allowing for real-time synchronous acquisition of video data streams from corresponding angles.
[0144] The data processing device 12 can send streaming commands to each acquisition device in the acquisition array 11 via a wireless local area network. Based on the streaming commands sent by the data processing device 12, each acquisition device in the acquisition array 11 transmits the obtained video data stream to the data processing device 12 in real time.
[0145] When the data processing device 12 receives a video frame capture instruction, it captures multiple synchronous video frames from the video frame at a specified frame time in the received multi-channel video data stream, and uploads the multiple synchronous video frames at the specified frame time to the server cluster 13 in the cloud.
[0146] Accordingly, the cloud server cluster 13 takes the received multiple synchronous video frames as an image combination, determines the corresponding parameter data of the image combination and the depth data of each frame image in the image combination, and performs frame image reconstruction on the preset virtual viewpoint path based on the corresponding parameter data of the image combination, the pixel data and depth data of the preset frame images in the image combination, to obtain the corresponding video frames of the multi-angle free viewpoint video.
[0147] In a specific implementation, the cloud-based server cluster 13 can store the pixel data and depth data of the image combination in the following manner:
[0148] Based on the pixel and depth data of the image combination, a stitched image corresponding to the frame time is generated. The stitched image includes a first field and a second field, wherein the first field includes the pixel data of the preset frame image in the image combination, and the second field includes a second field of the depth data of the preset frame image in the image combination. The acquired stitched image and corresponding parameter data can be stored in a data file. When it is necessary to obtain the stitched image or parameter data, it can be read from the corresponding storage space according to the corresponding storage address in the header file of the data file.
[0149] Then, the playback control device 14 can insert the received video frames of the multi-angle free-view video into the data stream to be played, and the playback terminal 15 receives the data stream to be played from the playback control device 14 and plays it in real time. The playback control device 14 can be a manual playback control device or a virtual playback control device. In a specific implementation, a dedicated server capable of automatically switching video streams can be set up as a virtual playback control device to control the data source. A broadcast control device such as a broadcast console can be used as a playback control device in this embodiment of the invention.
[0150] When the server cluster 13 in the cloud receives an image reconstruction instruction from the interactive terminal 16, it can extract the spliced image of the preset video frame in the corresponding image combination and the corresponding parameter data of the image combination and transmit it to the interactive terminal 16.
[0151] Based on the trigger operation, the interactive terminal 16 determines the interaction frame time information, sends an image reconstruction instruction containing the interaction frame time information to the server cluster 13, receives the spliced image of the preset video frame and the corresponding parameter data from the image combination of the corresponding interaction frame time returned from the server cluster 13 in the cloud, determines the virtual viewpoint position information based on the interaction operation, selects the corresponding pixel data and depth data and the corresponding parameter data in the spliced image according to the preset rules, combines and renders the selected pixel data and depth data, reconstructs the video frame of the multi-angle free viewpoint video corresponding to the virtual viewpoint position of the interaction frame time, and plays it.
[0152] Generally speaking, entities in a video are not completely static. For example, using the data processing system described above, during a basketball game, the entities captured by the acquisition array, such as athletes, basketballs, and referees, are mostly in motion. Correspondingly, the texture and pixel data in the image combination of the captured video frames also change continuously over time.
[0153] Using the above data processing system, on the one hand, users can directly watch videos with inserted multi-angle free-view video frames through the playback terminal 15, such as watching a live basketball game; on the other hand, while watching the video through the interactive terminal 16, users can view multi-angle free-view video at the interactive frame moment through interactive operations. It is understood that the above data processing system 10 may also include only the playback terminal 15 or only the interactive terminal 16, or the same terminal device may be used as both the playback terminal 15 and the interactive terminal 16.
[0154] Those skilled in the art will understand that multi-angle free-view video has a relatively large data volume, and the virtual information image data corresponding to AR effects is usually also large in volume. In addition, as can be seen from the working mechanism of the above data processing system, if AR effects are to be embedded into the reconstructed multi-angle free-view video at the same time as the reconstruction of the video, it will involve the processing of a large amount of data and the coordination of multiple devices. The complexity and data processing volume are difficult to achieve in terms of network data processing and transmission bandwidth resources. Therefore, how to embed AR effects to meet the user's visual experience needs during the playback of multi-angle free-view video has become a difficult problem to solve.
[0155] In view of this, the embodiments of this specification provide a solution, referring to Figure 2 The flowchart of the data processing method shown may specifically include the following steps:
[0156] S21, Obtain the target object in the video frame of the multi-angle free-view video.
[0157] In specific implementation, the video frames of the multi-angle free-view video can be obtained by reconstructing the frame images of the preset virtual viewpoint path based on the parameter data of the image combination formed by multiple synchronous video frames at a specified frame time extracted from multiple synchronous video streams, the pixel data and depth data of the preset frame images in the image combination, and the parameter data of the corresponding parameter data. The multiple synchronous video frames include frame images from different shooting angles.
[0158] In practical implementation, certain objects in the multi-angle free-view video can be identified as target objects based on certain indicator information (such as special effects display identifiers). This indicator information can be generated based on user interaction or obtained based on certain preset trigger conditions or third-party instructions. For example, in response to a special effects generation interaction control instruction, the target object in the video frame of the multi-angle free-view video can be obtained, and the indicator information can be set in the interaction control instruction. Specifically, the indicator information can be the identification information of the target object. As a specific example, the specific form of the indicator information corresponding to the target object can be determined based on the multi-angle free-view video frame structure.
[0159] In practice, the target object can be a specific entity in a video frame or sequence of video frames from multiple free-viewpoint videos, such as a specific person, animal, object, light beam, or environmental field / space. The embodiments in this specification do not limit the specific form of the target object.
[0160] In some embodiments of this specification, the multi-angle free-view video may be a 6DoF video.
[0161] S22, acquire the virtual information image generated based on the augmented reality effects input data of the target object.
[0162] In the embodiments described in this specification, the implanted AR effects are presented in the form of virtual information images. These virtual information images can be generated based on augmented reality effect input data of the target object. After determining the target object, a virtual information image generated based on the augmented reality effect input data of the target object can be obtained.
[0163] In the embodiments of this specification, the virtual information image corresponding to the target object can be generated in advance or generated in real time in response to the special effects generation command.
[0164] In specific implementation, the position of the target object in the video frame of the multi-angle free-view video can be obtained based on the three-dimensional calibration, and a virtual information image matching the position of the target object can be obtained. This makes the obtained virtual information image more consistent with the position of the target object in three-dimensional space, and the displayed virtual information image more in line with the real state in three-dimensional space. Therefore, the displayed synthetic video frame is more realistic and vivid, enhancing the user's visual experience.
[0165] In practice, based on the augmented reality effects input data of the target object, a virtual information image corresponding to the target object can be generated according to a preset effects generation method.
[0166] S23, The virtual information image is synthesized with the corresponding video frame and then displayed.
[0167] In practice, the synthesized video frames obtained after the synthesis process can be displayed on the terminal side.
[0168] The synthesized video frame, based on the video frame corresponding to the virtual information image, can be a single frame or multiple frames. If it is multiple frames, the virtual viewpoint image at the corresponding frame time can be synthesized with the video frame at the corresponding frame time according to the frame time order and the virtual viewpoint position at the corresponding frame time, and then displayed.
[0169] Since a virtual information image matching the virtual viewpoint position can be generated based on the virtual viewpoint position at the corresponding frame time, and then the virtual information image at the corresponding frame time is combined with the video frame at the corresponding frame time according to the frame time order and the virtual viewpoint position at the corresponding frame time, a composite video frame matching the virtual viewpoint position at the corresponding frame time can be automatically generated as the virtual viewpoint changes. This makes the augmented reality effects of the resulting composite video frame more realistic and vivid, thus further enhancing the user's visual experience.
[0170] In practice, there are multiple ways to synthesize and display the virtual information image with the corresponding video frame. Two specific examples are given below:
[0171] Example 1: The virtual information image is fused with the corresponding video frame to obtain a fused video frame, and the fused video frame is then displayed.
[0172] Example 2: The virtual information image is superimposed on the corresponding video frame to obtain a superimposed composite video frame, and the superimposed composite video frame is displayed.
[0173] In practice, the resulting composite video frames can be displayed directly, or they can be inserted into the video stream to be played for playback. For example, the fused video frames can be inserted into the video stream to be played for playback.
[0174] Using the embodiments of this specification, during the real-time playback of multi-angle free-viewpoint videos, the target object in the video frames of the multi-angle free-viewpoint video is acquired, and then a virtual information image generated based on the augmented reality effect input data of the target object is obtained. This virtual information image is then synthesized with the corresponding video frame and displayed. Through this process, only the video frame requiring AR effects needs to be synthesized with the virtual information image corresponding to the target object in the video frame during the playback of the multi-angle free-viewpoint video to obtain a video frame incorporating AR effects. It is unnecessary to pre-generate all the video frames incorporating AR effects for a single multi-angle free-viewpoint video before playback. Therefore, it is possible to accurately and quickly embed AR effects into multi-angle free-viewpoint videos, meeting users' needs for low-latency video viewing and a rich visual experience.
[0175] As mentioned above, embedding virtual information images corresponding to AR effects into multi-angle free-viewpoint videos is applicable to various application scenarios. To enable those skilled in the art to better understand and implement the embodiments of this specification, the following will elaborate on two application scenarios: interactive and non-interactive.
[0176] In non-interactive application scenarios, users can watch multi-angle free-view videos with embedded AR effects without user interaction. The timing, location, and content of the embedded AR effects can be controlled on the server side. Users can then automatically view the multi-angle free-view videos with embedded AR effects as the video stream plays on their devices. For example, during live or near-live broadcasts, by embedding AR effects into multi-angle free-view videos, composite video frames with embedded AR effects can be generated, meeting users' needs for low-latency video playback and a rich visual experience.
[0177] In interactive application scenarios, users can actively trigger the insertion of AR effects while watching videos from multiple free-viewpoint angles. By adopting the solution in the embodiments of this specification, AR can be quickly inserted into videos from multiple free-viewpoint angles, avoiding stuttering during video playback due to the long generation process. This enables the generation of multi-angle free-viewpoint videos with embedded AR effects based on user interaction, thus meeting users' needs for low-latency video playback and rich visual experience.
[0178] In practical implementation, corresponding to interactive scenarios, the system can respond to user-side special effects generation interactive control commands to obtain the target object from the video frames of the multi-angle free-view video. Then, a virtual information image generated based on the augmented reality special effects input data of the target object can be obtained, and the virtual information image can be composited with the corresponding video frames of the multi-angle free-view video and displayed.
[0179] The virtual information image corresponding to the target object can be generated in advance or in real time. For example, in non-interactive scenarios, it can be generated in response to server-side special effects generation instructions; for interactive scenarios, it can be generated in advance in response to server-side special effects generation instructions, or in real time in response to special effects generation interaction control instructions from the interactive terminal.
[0180] In some embodiments of this specification, the target object can be a specific entity in an image, such as a specific person, animal, object, or environmental space. Then, based on the target object indication information (e.g., an effect display identifier) in the special effects generation interactive control command, augmented reality effect input data for the target object is obtained. Based on this input data, a virtual information image corresponding to the target object is generated according to a preset special effects generation method. Specific special effects generation methods can be found in examples in subsequent embodiments, and will not be described in detail here.
[0181] In specific implementation, in order to process the data by compositing the video frames of the multi-angle free-view video with the virtual information image corresponding to the target image in the video frame, all or part of the data such as the data for generating the multi-angle free-view video and the input data for augmented reality effects can be pre-downloaded to the interactive terminal. The interactive terminal can perform some or all of the following operations: reconstruction of the multi-angle free-view video, generation of virtual information image, rendering of the video frames of the multi-angle free-view video and superimposition rendering of the virtual information image. Alternatively, the multi-angle free-view video and virtual information image can be generated on the server (such as a cloud server), and the compositing operation of the video frames of the multi-angle free-view video and the corresponding virtual information image can be performed only on the interactive terminal.
[0182] Furthermore, in non-interactive scenarios, the synthesized video frames from the multi-angle free-viewpoint video can be inserted into the data stream to be played. Specifically, a multi-angle free-viewpoint video containing synthesized video frames can be used as one of multiple data streams to be played, serving as the video stream to be selected for playback. For example, the video stream containing multi-angle free-viewpoint video frames can be used as an input video stream for a playback control device (such as a director control device), allowing the playback control device to select and use it.
[0183] It should be noted that in some cases, the same user may have a need to watch multi-angle free-view videos with embedded AR effects in both non-interactive and interactive scenarios. For example, while watching a live stream, a user might rewind to watch a replay of a particularly exciting scene or a segment of video they missed, thus fulfilling their interactive needs. Correspondingly, there will be composite video frames of multi-angle free-view videos with embedded AR effects obtained in both non-interactive and interactive scenarios.
[0184] To enable those skilled in the art to more clearly understand and implement the embodiments of this specification, the technical solutions in the embodiments of this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this specification.
[0185] The following section, with reference to the accompanying drawings, details the solutions for non-interactive application scenarios in the embodiments of this specification through specific examples.
[0186] In some embodiments of this specification, a data processing system employing a distributed system architecture, for a received image combination formed by multiple synchronous video frames at a specified frame time extracted from multiple video streams, determines the corresponding parameter data of the image combination and the depth data of each video frame in the image combination. On one hand, based on the corresponding parameter data of the image combination, the pixel data and depth data of the preset video frames in the image combination, frame image reconstruction is performed on a preset virtual viewpoint path to obtain the corresponding video frames of the multi-angle free-view video. On the other hand, in response to the special effects generation command, the target object in the video frame specified by the special effects generation command can be obtained, the augmented reality special effects input data of the target object can be obtained, and based on the augmented reality special effects input data of the target object, a corresponding virtual information image can be generated. The virtual information image is then synthesized with the specified video frame to obtain a synthesized video frame, which is then displayed for reference. Figure 3 The diagram shows the structure of a data processing system for an application scenario. The data processing system 30 includes: a data processing device 31, a server 32, a playback control device 33, and a playback terminal 34.
[0187] The data processing device 31 can extract video frames (including individual frame images) captured by the acquisition array in the field acquisition area. By extracting video frames to be generated as multi-angle free-view images, a large amount of data transmission and processing can be avoided. Subsequently, the server 32 generates video frames of the multi-angle free-view video, generates virtual information images in response to special effects generation instructions, and synthesizes the virtual information images and the video frames of the multi-angle free-view video to obtain a composite multi-angle free-view video frame. This fully utilizes the powerful computing capabilities of the server 32 to quickly generate composite multi-angle free-view video frames, which can then be promptly inserted into the data stream to be played in the playback control device 33. This achieves the playback of multi-angle free-view videos with AR special effects at a low cost, meeting users' needs for low-latency video playback and a rich visual experience.
[0188] Reference Figure 4 The flowchart shown illustrates the data processing method. To meet users' needs for low-latency video playback and a rich visual experience, video data can be processed through the following steps:
[0189] S41, receive multiple synchronous video frames at a specified frame time extracted from multiple synchronous video streams as an image combination, wherein the multiple synchronous video frames contain frame images from different shooting angles.
[0190] In practice, the data processing device can extract multiple video frames at a specified frame time from multiple synchronous video streams according to the received video frame extraction instructions and upload them, for example, to a cloud server or service cluster.
[0191] As a specific scenario example: In the on-site acquisition area, an acquisition array consisting of multiple acquisition devices can be deployed at different locations. This acquisition array can synchronously acquire multiple video data streams in real time and upload them to the data processing device. When the data processing device receives a video frame extraction command, it can extract a video frame at the corresponding frame time from the multiple video data streams based on the specified frame time information contained in the video frame extraction command. The specified frame time can be in units of frames, using frames N to M as the specified frame time, where N and M are both integers not less than 1, and N ≤ M; or, the specified frame time can be in units of time, using seconds X to Y as the specified frame time, where X and Y are both positive numbers, and X ≤ Y. Therefore, multiple synchronized video frames can include all frame-level synchronized video frames corresponding to the specified frame time, and the pixel data of each video frame forms a corresponding frame image.
[0192] For example, a data processing device can obtain the second frame of a multi-channel video data stream at a specified frame time according to the received video frame capture instruction. The data processing device then captures the second frame of each video data stream, and the captured second frames of each video data stream are synchronized at the frame level to obtain multiple synchronized video frames.
[0193] For example, assuming the capture frame rate is set to 25fps, that is, 25 frames are captured per second, the data processing device can obtain the video frames within the first second of multiple video data streams at the specified frame time according to the received video frame capture instruction. Then the data processing device can capture 25 video frames within the first second of each video data stream, and the first video frames within the first second of each video data stream are synchronized at the frame level, the second video frames within the first second of each video data stream are synchronized at the frame level, and so on, until the 25th video frames within the first second of each video data stream are synchronized at the frame level, which are then used as multiple synchronized video frames.
[0194] For example, based on the received video frame capture instruction, the data processing device can obtain the second and third frames of multiple video data streams at a specified frame time. Then, the data processing device can capture the video frames of the second and third frames of each video data stream respectively, and the captured video frames of the second and third frames of each video data stream are synchronized at the frame level to form multiple synchronized video frames.
[0195] In practice, the multiple video data streams can be either compressed or uncompressed.
[0196] S42, determine the corresponding parameter data for the image combination.
[0197] In practical implementation, the corresponding parameter data of the image combination can be obtained through a parameter matrix, which may include intrinsic parameter matrices, extrinsic parameter matrices, rotation matrices, and translation matrices. Thus, the relationship between the three-dimensional geometric position of a specified point on the surface of a spatial object and its corresponding point in the image combination can be determined.
[0198] In embodiments of the present invention, a Structure From Motion (SFM) algorithm can be employed to perform feature extraction, feature matching, and global optimization on the acquired image combination based on the parameter matrix. The obtained parameter estimates serve as the corresponding parameter data for the image combination. The feature extraction algorithm can include any of the following: Scale-Invariant Feature Transform (SIFT), Speeded-Up Robust Features (SURF), or Features from Accelerated Segment Test (FAST). The feature matching algorithm can include: Euclidean distance calculation methods, Random Sample Consensus (RANSC), etc. The global optimization algorithm can include: Bundle Adjustment (BA), etc.
[0199] S43, determine the depth data of each frame in the image combination.
[0200] In practical implementation, depth data for each frame can be determined based on multiple frames in the image set. This depth data can include depth values corresponding to the pixels of each frame in the image set. The distance from the acquisition point to various points in the scene can be used as these depth values, which directly reflect the geometry of the visible surface in the area to be viewed. For example, using the origin of the shooting coordinate system as the optical center, the depth value can be the distance from each point in the scene along the shooting optical axis to the optical center. Those skilled in the art will understand that these distances can be relative values, and multiple frames can use the same reference.
[0201] In one embodiment of the present invention, a binocular stereo vision algorithm can be used to calculate the depth data of each frame of the image. Alternatively, the depth data can also be indirectly estimated by analyzing features such as luminance and brightness characteristics of the frame images.
[0202] In another embodiment of the present invention, a multi-view stereo (MVS) algorithm can be used for frame image reconstruction. During the reconstruction process, all pixels can be used for reconstruction, or pixels can be downsampled and only a subset of pixels can be used for reconstruction. Specifically, pixels in each frame image can be matched to reconstruct the 3D coordinates of each pixel, obtaining points with image consistency, and then the depth data of each frame image can be calculated. Alternatively, pixels in selected frame images can be matched to reconstruct the 3D coordinates of pixels in each selected frame image, obtaining points with image consistency, and then the depth data of the corresponding frame images can be calculated. The pixel data of the frame images corresponds to the calculated depth data. The method of selecting frame images can be set according to the specific scenario; for example, a subset of frame images can be selected based on the distance between the frame images for which depth data needs to be calculated and other frame images.
[0203] S44. Based on the parameter data of the image combination, the pixel data and depth data of the preset frame image in the image combination, the preset virtual viewpoint path is reconstructed into a frame image to obtain the video frame of the corresponding multi-angle free viewpoint video.
[0204] In specific implementation, the pixel data of the frame image can be any one of YUV data or RGB data, or other data that can represent the frame image; the depth data can include depth values that correspond one-to-one with the pixel data of the frame image, or it can be a selection of values from the set of depth values that correspond one-to-one with the pixel data of the frame image, and the specific selection method depends on the specific scenario; the virtual viewpoint is selected from a multi-angle free viewing angle range, which is the range that supports switching the viewpoint of the area to be viewed.
[0205] In practice, the preset frame images can be all the frame images in the image combination, or a selected subset of frame images. The selection method can be set according to the specific scenario. For example, based on the positional relationship between the acquisition points, a subset of frame images at the corresponding position in the image combination can be selected; or, based on the desired frame time or frame time period, a subset of frame images at the corresponding frame time in the image combination can be selected.
[0206] Since the preset frame images can correspond to different frame times, each virtual viewpoint in the virtual viewpoint path can be mapped to a specific frame time. Based on the frame time corresponding to each virtual viewpoint, the corresponding frame image is obtained. Then, based on the image combination, the corresponding parameter data, the depth data, and pixel data of the frame image corresponding to each virtual viewpoint's frame time, frame image reconstruction is performed on each virtual viewpoint to obtain the corresponding multi-angle free-view video frames. Therefore, in specific implementations, in addition to realizing multi-angle free-view images at a specific moment, it is also possible to realize temporally continuous or non-continuous multi-angle free-view videos.
[0207] In one embodiment of the present invention, the image combination includes A synchronous video frames, wherein a1 synchronous video frames correspond to the first frame time, a2 synchronous video frames correspond to the second frame time, and a1+a2=A; and a virtual viewpoint path composed of B virtual viewpoints is preset, wherein b1 virtual viewpoints correspond to the first frame time, b2 virtual viewpoints correspond to the second frame time, and b1+b2≤2B. Based on the corresponding parameter data of the image combination, the pixel data and depth data of the frame images of the a1 synchronous video frames at the first frame time, the path composed of b1 virtual viewpoints is reconstructed into a first frame image. Based on the corresponding parameter data of the image combination, the pixel data and depth data of the frame images of the a2 synchronous video frames at the second frame time, the path composed of b2 virtual viewpoints is reconstructed into a second frame image, and finally, the corresponding multi-angle free-view video frame is obtained.
[0208] Understandably, the specified frame time and virtual viewpoint can be further divided, thereby obtaining more synchronized video frames and virtual viewpoints corresponding to different frame times, realizing the free switching of viewpoints over time, and improving the smoothness of multi-angle free-view video viewpoint switching.
[0209] It is understood that the above embodiments are merely illustrative examples and are not intended to limit the specific implementation methods.
[0210] In the embodiments of this specification, a depth image based rendering (DIBR) algorithm can be used to combine and render the pixel data and depth data of a preset frame image based on the corresponding parameter data of the image combination and a preset virtual viewpoint path, thereby realizing frame image reconstruction based on the preset virtual viewpoint path and obtaining the video frames of the corresponding multi-angle free viewpoint video.
[0211] S45, in response to the special effects generation instruction, obtain the target object in the video frame specified by the special effects generation instruction, obtain the augmented reality special effects input data of the target object, and generate the corresponding virtual information image based on the augmented reality special effects input data of the target object.
[0212] In specific implementation, in response to the special effects generation instruction, the augmented reality special effects input data of the target object can be used as input. Based on the position of the target object in the video frame of the multi-angle free-view video obtained by three-dimensional calibration, and a preset first special effects generation method can be used to generate a virtual information image in the corresponding video frame that matches the target object.
[0213] To accurately locate the position of the target object corresponding to the special effects generation instruction, in a specific implementation, for the video frame to which the AR special effects are to be implanted, a preset number of pixels can be selected. Based on the parameter data of the video frame and the real physical space parameters corresponding to the video frame, the spatial position of the preset number of pixels is determined, thereby determining the accurate position of the target object in the video frame.
[0214] Reference Figure 5 and Figure 6 , Figure 5 The video frame P50 shown depicts an image of a basketball game in progress, with multiple basketball players on the court, one of whom is making a shooting motion. To determine the position of the target object within the video frame, as... Figure 6 As shown, the pixel points A, B, C, and D corresponding to the four vertices of the restricted area of the basketball court are selected. Combined with the parameters of the real basketball court, the calibration can be completed using the parameters of the camera corresponding to a video frame. Then, the three-dimensional position information of the court in the corresponding virtual camera can be obtained according to the parameters of the virtual camera, thereby achieving accurate calibration of the three-dimensional spatial position relationship of the video frame containing the basketball court.
[0215] It is understandable that other pixels in the video frame can also be selected for 3D calibration to determine the position of the target object corresponding to the special effects generation instruction in the video frame. In specific implementations, to ensure more accurate 3D spatial relationships of specific objects in the image, pixels corresponding to stationary objects in the image are preferentially selected for 3D calibration. One or more pixels can be selected. To reduce the amount of data computation, contour points or vertices of regular objects in the image can be preferentially selected for 3D calibration.
[0216] Through 3D calibration, the generated virtual 3D information image can be accurately integrated with the multi-angle free-view video describing the real world at any position, any angle, and any viewpoint in 3D space. This enables seamless integration of virtual and reality, and achieves dynamic synchronization and harmonious unity of the video frames of the virtual information image and the multi-angle free-view video during playback. Therefore, the multi-angle free-view synthesized video frames obtained by the synthesis process are more natural and realistic, which can greatly enhance the user's visual experience.
[0217] In practical implementation, the server (such as a cloud server) can automatically generate special effects generation instructions, or it can respond to server-side user interaction operations and generate corresponding server-side special effects generation interaction control instructions. For example, the cloud server can automatically select the image combination to be embedded with AR special effects as the image combination specified by the special effects generation instruction through a preset AI recognition algorithm, and obtain the virtual information image corresponding to the specified image combination. Alternatively, the server-side user can specify an image combination through interactive operations. When the server receives a server-side special effects generation interaction control instruction triggered by the server-side special effects generation interaction control operation, it can obtain the specified image combination from the server-side special effects generation interaction instruction, and then obtain the virtual information image corresponding to the image combination specified by the special effects generation instruction.
[0218] In practice, a virtual information image corresponding to the image combination specified by the special effects generation instruction can be directly obtained from a preset storage space, or a matching virtual information image can be generated in real time according to the image combination specified by the special effects generation instruction.
[0219] To generate the virtual information image, in a specific implementation, the target object can be taken as the center, the target object in the video frame can be identified first, the augmented reality effect input data of the target object can be obtained, and then the augmented reality effect input data can be used as input to generate a virtual information image that matches the target object in the video frame using a preset first effect generation method.
[0220] In some embodiments of this specification, target objects in the video frame can be identified using image recognition technology. For example, the target object in the special effects area can be identified as a person (such as a basketball player), an object (such as a basketball or a scoreboard), an animal (such as a cat or a lion), etc.
[0221] In practical implementation, the augmented reality effect input data of the target object can be obtained in response to the server-side effect generation interaction control command. For example, if a server-side user selects a player in a live basketball game video through interactive operation, a server-side effect generation interaction control command corresponding to the interactive operation can be generated. Based on the server-side effect generation interaction control command, the augmented reality effect input data associated with the player can be obtained, such as name, position name in the basketball game (which can be a specific number or position type: such as center, forward, guard, etc.), and shooting percentage, etc.
[0222] In practical implementation, interactive control instructions can be generated based on the server-side special effects to determine the special effects output type. Then, historical data of the target object is acquired, and this historical data is processed according to the special effects data type to obtain augmented reality special effects input data corresponding to the special effects output type. For example, for a live basketball game, if the server-side user wants to obtain the shooting percentage of the target object's location based on the server-side special effects generated interactive control instructions, the distance between the target object's location and the ground projection of the net's center can be calculated. Historical shooting data within this distance can then be used as the augmented reality special effects input data for the target object.
[0223] In practice, server-side users can perform interactive control operations through corresponding interactive control devices. Based on the server-side users' interactive control operations, corresponding server-side interactive control commands for generating special effects can be obtained. In practice, server-side users can select the target object for the special effect to be generated through interactive operations. Furthermore, users can also select the augmented reality effect input data for the target object, such as the data type and data range (which can be selected based on time or geographic space).
[0224] It is understood that the server-side special effects generation interactive control instructions can also be automatically generated by the server. The server can make autonomous decisions through machine learning, selecting the image combination of the video frames to be implanted with special effects, the target object, and the augmented reality special effects input data of the target object.
[0225] The following describes how to generate a virtual information image that matches the target object in the video frame using a preset first special effects generation method through some specific implementation methods.
[0226] In one specific implementation of this specification, the augmented reality effect input data can be input into a preset three-dimensional model for processing to obtain a virtual information image that matches the target object in the video frame.
[0227] For example, after inputting the augmented reality effects input data into a preset 3D model, 3D graphic elements that match the augmented reality effects input data can be obtained and combined, and the display metadata in the augmented reality effects data and the 3D graphic element data can be output as virtual information images that match the target object in the video frame.
[0228] The three-dimensional model can be a three-dimensional model obtained by scanning an actual object, or it can be a constructed virtual model. The virtual model can include virtual object models and virtual character models. Virtual objects can be virtual magic wands or other items that do not exist in the real world. Virtual character models can be imagined figures or animal shapes, such as a three-dimensional model of the legendary Nezha, or a three-dimensional model of a virtual unicorn, dragon, etc.
[0229] In another specific implementation of this specification, the augmented reality effect input data can be used as input data and fed into a preset machine learning model for processing to obtain a virtual information image that matches the target object in the video frame.
[0230] In specific implementation, the preset machine learning model can be a supervised learning model, an unsupervised learning model, or a semi-supervised learning model (a combination of supervised and unsupervised learning models). The specific model used is not limited in the embodiments of this specification.
[0231] The virtual information image is generated using a machine learning model, which includes two stages: the model training stage and the model application stage.
[0232] During the model training phase, training sample data can be used as input data and fed into a preset machine learning model for training. The parameters of the machine learning model are adjusted, and after training, the model can be used as the preset machine learning model. The training sample data can include various images and videos collected from real physical spaces, or virtual images or videos generated by artificial modeling. After training, the machine learning model can automatically generate corresponding 3D images, 3D videos, and corresponding sound effects based on the input data.
[0233] In the model application stage: the augmented reality effect input data is used as input data and fed into the trained machine learning model, which can automatically generate an augmented reality effect model that matches the input data, that is, a virtual information image that matches the target object in the video frame.
[0234] In the embodiments of this specification, the form of the generated virtual information image varies depending on the 3D model used or the machine learning model used. Specifically, the generated virtual information image can be a static image, a dynamic video frame such as an animation, or even a video frame containing audio data.
[0235] S46, The virtual information image is combined with the specified video frame to obtain a composite video frame.
[0236] In practice, the virtual information image and the specified video frame can be fused together to obtain a fused video frame with embedded AR effects.
[0237] S47, The synthesized video frame is displayed.
[0238] The synthesized video frame is inserted into the video stream to be played on the playback control device for playback via the playback terminal.
[0239] In practical implementation, the playback control device can use multiple video streams as inputs. These video streams can come from various acquisition devices in the acquisition array or from other acquisition devices. The playback control device can select one input video stream as the video stream to be played as needed. Specifically, it can select a composite video frame of the multi-angle free-view video obtained in step S46 above and insert it into the video stream to be played, or switch the video stream from other input interfaces to the input interface containing the composite video frame of the multi-angle free-view video. The playback control device outputs the selected video stream to be played to the playback terminal, which can then play it. Therefore, in addition to viewing the multi-angle free-view video frames through the playback terminal, users can also view the composite video frames of the multi-angle free-view video with embedded AR effects through the playback terminal.
[0240] The playback terminal can be a video playback device such as a television, mobile phone, tablet, or computer, or other types of electronic devices that include a display screen or projection device.
[0241] In practice, the multi-angle free-viewpoint video composite video frames of the video stream to be played inserted into the playback control device can be retained in the playback terminal so that users can perform time-shift viewing. Time-shifting can be operations such as pausing, rewinding, and fast-forwarding to the current moment performed by the user while watching.
[0242] As can be seen from the above steps, for an image combination formed by multiple synchronous video frames at a specified frame time captured from multiple video streams, by determining the corresponding parameter data of the image combination and the depth data of each frame image in the image combination, on the one hand, based on the corresponding parameter data of the image combination, the pixel data and depth data of the preset frame images in the image combination, the preset virtual viewpoint path is reconstructed to obtain the corresponding video frame of the multi-angle free-view video; on the other hand, in response to the special effects generation instruction, the target object in the video frame obtains the augmented reality special effects input data of the target object, and based on the augmented reality special effects input data of the target object, a corresponding virtual information image is generated, and the virtual information image is synthesized with the specified video frame to obtain a synthesized video frame. Then, the synthesized video frame is inserted into the video stream to be played on the playback control device for playback through the playback terminal, which can realize a multi-angle free-view video with AR special effects.
[0243] Using the above data processing method, multiple synchronous video frames at a specified time are extracted from multiple synchronous video streams to reconstruct multi-angle free-view video and generate virtual information images corresponding to the target objects in the video frames specified by the special effects generation instructions. Therefore, there is no need to upload a huge amount of synchronous video stream data. This distributed system architecture can save a lot of transmission and server processing resources. Moreover, under the condition of limited network transmission bandwidth, it can realize the real-time or near-real-time generation of synthetic video frames with augmented reality effects. Thus, it can realize low-latency playback of multi-angle free-view synthetic video frames with embedded AR effects, thereby meeting the dual needs of users for rich visual experience and low latency during video viewing.
[0244] In practice, the steps described above, such as capturing synchronous video frames from multiple video streams, generating video frames of multi-angle free-view video based on image combinations formed by multiple synchronous video frames, obtaining virtual information images corresponding to the image combinations specified by the special effects generation instructions, and synthesizing the virtual information images and the specified image combinations to obtain synthesized video frames, can all be completed collaboratively by different hardware devices, i.e., a distributed processing architecture is adopted.
[0245] Continue to refer to Figure 4 In step S44, the depth data of the preset video frames in the image combination can be mapped to the corresponding virtual viewpoints according to the relationship between the virtual parameter data of each virtual viewpoint in the preset virtual viewpoint path and the corresponding parameter data of the image combination. Based on the pixel data and depth data of the preset video frames mapped to the corresponding virtual viewpoints and the preset virtual viewpoint path, frame image reconstruction is performed to obtain the video frames of the corresponding multi-angle free-view video.
[0246] The virtual parameter data of the virtual viewpoint may include virtual viewing position data and virtual viewing angle data; the corresponding parameter data of the image combination may include acquisition position data and shooting angle data, etc. A forward mapping followed by a reverse mapping method can be used to obtain the reconstructed video frames.
[0247] In practice, the location data and shooting angle data can be referred to as external parameter data. Parameter data can also include internal parameter data, which may include the attribute data of the acquisition device, thereby allowing for a more accurate determination of the mapping relationship. For example, internal parameter data may include distortion data; by taking distortion factors into account, the mapping relationship can be determined more accurately spatially.
[0248] Next, with reference to the accompanying drawings, the interactive application scenarios in the embodiments of this specification will be described in detail through specific examples.
[0249] like Figure 7 The flowchart of the data processing method shown is illustrated in some embodiments of this specification. On an interactive terminal, based on user interaction, the following steps can be used to obtain multi-angle free-viewpoint video frames with embedded AR effects:
[0250] The S71 displays video frames from multiple free-viewpoint angles in real time.
[0251] In specific implementation, the video frames of the multi-angle free-viewpoint video are reconstructed based on the parameter data of an image combination formed by multiple synchronous video frames at a specified frame time, the pixel data of the image combination, and the depth data of the image combination. The multiple synchronous video frames include frame images from different shooting angles. The reconstruction method of the multi-angle free-viewpoint video frames can be found in the description of the foregoing embodiments, and will not be described in detail here.
[0252] S72, in response to the triggering operation of the special effects display identifier in the video frame of the multi-angle free-view video, obtain the virtual information image of the video frame corresponding to the specified frame time of the special effects display identifier.
[0253] S73, The virtual information image is synthesized with the corresponding video frame and then displayed.
[0254] In practice, the superposition position of the virtual information image in the video frame of the multi-angle free-view video can be determined based on the special effects display identifier. Then, the virtual information image can be superimposed and displayed at the determined superposition position.
[0255] To enable those skilled in the art to better understand and implement this process, the following detailed explanation is provided through an image demonstration of an interactive terminal. (Refer to...)Figures 8 to 12 The diagram shows a video playback screen on the interactive terminal. The interactive terminal T80 plays the video in real time, wherein, as described in step S71, referring to... Figure 8 The interactive terminal then displays video frame P80. Next, video frame P81 contains several special effects display icons, including the special effects display icon I1. These are represented in video frame P80 by an inverted triangle pointing to the target object, such as... Figure 9 As shown. It is understood that other methods can also be used to display the special effects display identifier. When a terminal user touches or clicks the special effects display identifier I1, the system automatically acquires the virtual information image corresponding to the special effects display identifier I1, and overlays the virtual information image onto the video frame P81 of the multi-angle free-view video, as shown. Figure 10 As shown, a three-dimensional ring R1 is rendered with the athlete Q1 standing at the center of the playing field. Next, as... Figure 11 and Figure 12 As shown, when a terminal user touches and clicks on the special effects display identifier I2 in the video frame P81 of the multi-angle free-view video, the system automatically obtains the virtual information image corresponding to the special effects display identifier I2, and overlays the virtual information image onto the video frame P81 of the multi-angle free-view video, resulting in the multi-angle free-view video overlay video frame P82, which displays the hit rate information display board M0. The hit rate information display board M0 displays the target object, i.e., athlete Q1's position, name, and hit rate information.
[0256] like Figures 8 to 12 As shown, end users can continue to click on other special effects display icons shown in the video frame to watch videos displaying the corresponding AR effects of each special effects display icon.
[0257] Understandably, different types of embedded special effects can be distinguished by different types of special effects display icons.
[0258] In practice, special effects display indicators can be shown not only on the playback screen but also elsewhere. For example, for video frames that can display AR effects, special effects display indicators can be set at the progress position corresponding to the frame on the playback progress bar to inform the end user. Figure 13The diagram shows the interactive interface of the interactive terminal. The interactive terminal T130 displays the playback interface Sr131 and the position of the currently playing video frame in the entire progress bar L131. As can be seen from the information displayed on the progress bar L131, based on the position of the currently playing video frame in the entire video, the progress bar L131 is divided into a played segment L131a and an unplayed segment L131b. In addition, special effect display indicators D1 to D4 are displayed on the progress bar L131. Among them, special effect display indicator D1 is located in the played segment L131a, special effect display indicator D2 is the current video frame, located at the intersection of the played segment L131a and the unplayed segment L131b, and special effect display indicators D3 and D4 are located in the unplayed segment L131b. The terminal user can use the special effect display indicators on the progress bar L131 to rewind or fast forward to the corresponding video frame and watch the scene corresponding to the multi-angle free-view composite video frame with embedded AR special effects.
[0259] Reference Figure 14 The flowchart of the data processing method shown is illustrated in an interactive scenario of an embodiment of this specification. To realize the display of multi-angle free-viewpoint video composite video frames with AR effects embedded in the interactive terminal, the following steps can be used for data processing:
[0260] S141, in response to the image reconstruction command from the interactive terminal, the interaction frame time information at the interaction time is determined, the stitched image of the preset frame image in the image combination corresponding to the interaction frame time and the corresponding parameter data of the image combination are obtained and sent to the interactive terminal, so that the interactive terminal selects the corresponding pixel data and depth data and the corresponding parameter data in the stitched image according to the virtual viewpoint position information determined by the interaction operation and according to the preset rules, combines and renders the selected pixel data and depth data, reconstructs the video frame of the multi-angle free viewpoint video corresponding to the virtual viewpoint position at the interaction frame time, and plays it.
[0261] In a specific implementation, the stitched image of the preset frame image is generated based on the pixel data and depth data of the image combination at the time of the interaction frame. The stitched image includes a first field and a second field, wherein the first field includes the pixel data of the preset frame image in the image combination, and the second field includes the depth data of the image combination.
[0262] In a specific implementation, the image combination at the interactive frame moment is obtained based on multiple synchronous video frames extracted from multiple synchronous video streams at a specified frame moment, and the multiple synchronous video frames contain frame images from different shooting angles.
[0263] S142, in response to the special effects generation interactive control command, obtain the virtual information image corresponding to the spliced image of the preset video frame indicated by the special effects generation interactive control command.
[0264] In some embodiments of this specification, in response to an effects generation interactive control command, a target object in a preset video frame indicated by the effects generation interactive control command can be read; based on the target object, a virtual information image generated in advance based on augmented reality effects input data of the target object can be obtained.
[0265] In practical implementation, virtual information images matching the target object can be generated in various ways. Two examples are given below:
[0266] Example 1: The augmented reality effects data of the target object is used as input data and processed into a preset 3D model to obtain a virtual information image that matches the target object;
[0267] Example 2: The augmented reality effects data of the target object is used as input data and fed into a preset machine learning model for processing to obtain a virtual information image that matches the target object.
[0268] For specific implementation examples of the two examples above, please refer to the foregoing embodiments.
[0269] S143, the virtual information image corresponding to the spliced image of the preset video frame is sent to the interactive terminal, so that the interactive terminal will synthesize the video frame of the multi-angle free view video corresponding to the virtual viewpoint position at the time of the interactive frame with the virtual information image to obtain a synthesized video frame and display it.
[0270] To enable those skilled in the art to better understand and implement the embodiments of this specification, a data processing system suitable for interactive scenarios is provided below.
[0271] Reference Figure 15 In some embodiments of this specification, the data processing system 150 may include a server 151 and an interactive terminal 152, wherein:
[0272] The server 151 can respond to the image reconstruction command from the interactive terminal 152, determine the interaction frame time information at the interaction time, obtain the spliced image of the preset video frame in the image combination at the corresponding interaction frame time and the corresponding parameter data of the image combination and send it to the interactive terminal 152, and respond to the special effects generation interactive control command to generate the virtual information image corresponding to the spliced image of the preset video frame indicated by the special effects generation interactive control command.
[0273] The interactive terminal 152, based on the virtual viewpoint position information determined by the interactive operation, selects the corresponding pixel data, depth data, and corresponding parameter data in the stitched image according to preset rules, combines and renders the selected pixel data and depth data, reconstructs and plays the multi-angle free-view video image corresponding to the virtual viewpoint position at the time of the interactive frame; and synthesizes the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the time of the interactive frame with the virtual information image to obtain a synthesized video frame and plays it.
[0274] In a specific implementation, the server 151 may store a virtual information image corresponding to the stitched image of the preset frame image, or obtain the virtual information image corresponding to the stitched image of the preset frame image from a third party based on the augmented reality effect input data of the stitched image of the preset frame image, or generate the virtual information image corresponding to the stitched image of the preset frame image in real time.
[0275] In a specific implementation, the image combination at the interactive frame moment is obtained based on multiple synchronous video frames extracted from multiple synchronous video streams at a specified frame moment, and the multiple synchronous video frames contain frame images from different shooting angles.
[0276] The data processing system may further include a data processing device 153. As described in the previous embodiment, the data processing device 153 can extract video frames from the video frames acquired by the acquisition array in the field acquisition area. By extracting video frames to generate multi-angle free-view video, a large amount of data transmission and processing can be avoided. The acquisition devices in the field acquisition array can simultaneously acquire frame images from different shooting angles, and the data processing device can extract multiple synchronous video frames at a specified frame time from multiple synchronous video streams.
[0277] Subsequently, the data processing device 153 can upload the captured frame images to the server 151. The server 151 can store a stitched image of a combination of preset video frames and parameter data of the image combination.
[0278] In practice, the data processing system applicable to non-interactive scenarios and the data processing system applicable to interactive scenarios can be integrated.
[0279] Continue to refer to Figure 3As a specific example, in addition to obtaining video frames of multi-angle free-view video and the virtual information image, the server 32 can also generate a stitched image corresponding to the image combination formed by multiple synchronous video frames at a specified frame time, based on the pixel data and depth data of the image combination, in order to facilitate subsequent data acquisition. The stitched image may include a first field and a second field, wherein the first field includes the pixel data of the image combination and the second field includes the depth data of the image combination. Then, the stitched image corresponding to the image combination and the parameter data corresponding to the image combination are stored.
[0280] To save storage space, a stitched image corresponding to a preset video frame in the image combination can be generated based on the pixel data and depth data of the preset video frame. The stitched image corresponding to the preset video frame may include a first field and a second field, wherein the first field includes the pixel data of the preset video frame and the second field includes the depth data of the preset video frame. Then, only the stitched image corresponding to the preset video frame and the corresponding parameter data need to be stored.
[0281] The first field corresponds to the second field. The stitched image can be divided into an image region and a depth map region. The pixel field of the image region stores the pixel data of the multiple frame images, and the pixel field of the depth map region stores the depth data of the multiple frame images. The pixel field storing the pixel data of the frame images in the image region serves as the first field, and the pixel field storing the depth data of the frame images in the depth map region serves as the second field. The stitched image of the obtained image combination and the corresponding parameter data of the image combination can be stored in a data file. When it is necessary to obtain the stitched image or the corresponding parameter data, it can be read from the corresponding storage space according to the storage address contained in the header file of the data file.
[0282] Furthermore, the image combination can be stored in a video format, and there can be multiple image combinations. Each image combination can be an image combination corresponding to different frame times after the video has been de-encapsulated and decoded.
[0283] In practical implementation, in addition to watching multi-angle free-view videos through the playback terminal, users can also actively select to play multi-angle free-view videos through interactive operations during video viewing to further enhance the interactive experience. In some embodiments of this specification, the following methods are used for implementation:
[0284] In response to an image reconstruction command from an interactive terminal, the interaction frame time information at the moment of interaction is determined. The stitched image of a preset video frame in the image combination at the corresponding interaction frame time and the corresponding parameter data of the image combination are obtained and sent to the interactive terminal. The interactive terminal, based on the virtual viewpoint position information determined by the interaction operation, selects the corresponding pixel data, depth data and corresponding parameter data in the stitched image according to preset rules, combines and renders the selected pixel data and depth data, reconstructs the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the moment of interaction, and plays it.
[0285] The preset rules can be set according to specific scenarios. For example, based on the virtual viewpoint position information determined by the interactive operation, the position information of W neighboring virtual viewpoints closest to the virtual viewpoint at the interaction time can be selected by sorting them by distance. The pixel data and depth data corresponding to the above W+1 virtual viewpoints, including the virtual viewpoint at the interaction time, that satisfy the interaction frame time information can be obtained in the stitched image.
[0286] The interaction frame timing information is determined based on a trigger operation from the interactive terminal. This trigger operation can be input by the user or automatically generated by the interactive terminal. For example, the interactive terminal can automatically initiate a trigger operation when it detects the presence of a multi-angle free-viewpoint data frame. When manually triggered by the user, the trigger can be based on the user's selection of the trigger time after the interactive terminal displays an interaction prompt, or it can be based on historical timing information received by the interactive terminal from the user's actions. This historical timing information can be timing information prior to the current playback moment.
[0287] In a specific implementation, the interactive terminal 35 can combine and render the pixel data and depth data of the spliced image of the preset video frames in the image combination at the time of the acquired interactive frame, based on the spliced image and corresponding parameter data of the preset video frames in the image combination at the time of the acquired interactive frame, the interactive frame time information, and the virtual viewpoint position information at the time of the interactive frame, using the same method as in step S44 above, to obtain the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position of the interaction, and start playing the multi-angle free-view video at the virtual viewpoint position of the interaction.
[0288] Using the above scheme, video frames of multi-angle free-view video corresponding to the virtual viewpoint position of the interaction can be generated in real time based on the image reconstruction instructions from the interactive terminal, which can further enhance the user's interactive experience.
[0289] In practice, the interactive terminal and the playback terminal can be the same terminal device.
[0290] In practice, to facilitate subsequent data acquisition, a virtual information image corresponding to the spliced image of the preset frame image indicated by the server-side special effects generation interactive control command can be generated and stored in response to the server-side special effects generation interactive control command.
[0291] Subsequently, during the playback of the multi-angle free-view video corresponding to the spliced image of the preset frame image, the virtual information image can be superimposed and rendered on the spliced image of the preset frame image to obtain a multi-angle free-view video superimposed video frame with AR effects. Specifically, this can be implemented in scenarios such as multi-angle free-view video recording or on-demand playback. The virtual information image can be triggered by pre-setting or by user interaction.
[0292] Taking user interaction scenarios as an example, to further enhance the richness of the user's visual experience while watching multi-angle free-view videos, AR effects can be embedded in the videos. In some embodiments of this specification, this can be implemented in the following ways:
[0293] After receiving the image reconstruction instruction, it can also respond to the user-end special effects generation interaction instruction from the interactive terminal, obtain the virtual information image corresponding to the stitched image of the preset video frame, and send the virtual information image corresponding to the stitched image of the preset video frame to the interactive terminal, so that the interactive terminal overlays and renders the virtual information image on the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the time of the interaction frame, and obtains the multi-angle free-view overlay video frame with embedded AR special effects and plays it.
[0294] As a specific example, during video viewing, if the user's first interactive operation triggers the playback of a multi-angle free-view video, during playback, an interactive command is generated based on the user's second interactive operation corresponding to the user-side special effects. This allows the acquisition of a virtual information image corresponding to the stitched image of the preset frame images, i.e., the AR special effects image of the multi-angle free-view video to be embedded in the preset video frame. The preset video frame can be the video frame indicated by the user's second interactive operation, such as the frame image clicked by the user, or the frame sequence corresponding to the user's swipe operation.
[0295] In specific implementation, in response to the user's special effects exit interaction command, the acquisition of the virtual information image corresponding to the spliced image of the preset frame image can be stopped. Accordingly, during the rendering process of the interactive terminal, there is no need to overlay the virtual information image, and only the multi-angle free-view video is played.
[0296] Continuing with the example above, if during the playback of a multi-angle free-viewpoint overlay video frame that incorporates AR effects data, the user-side effects exit interaction command corresponding to the user's third interaction operation stops the acquisition and rendering display of the virtual information image corresponding to the stitched image of subsequent video frames.
[0297] In specific implementation, as a continuous video stream, some video streams may contain multi-angle free-view video data. In one or more multi-angle free-view video sequences, one or more sequences correspond to the virtual information image. Therefore, when the user-end effect exit interaction command is detected, the implantation of all subsequent AR effects in the video stream can be exited, or only the display of subsequent AR effects in one multi-angle free-view video sequence can be exited.
[0298] Similar to the method used to generate the aforementioned virtual information images, virtual information images can be generated based on server-side special effects generation instructions. In specific implementations, special effects generation instructions can be automatically generated by the server (such as a cloud server), or corresponding server-side special effects generation interactive control instructions can be generated in response to server-side user interaction operations.
[0299] Similarly, to generate the virtual information image, firstly, a stitched image of a preset frame corresponding to the virtual information image is determined, and then, a virtual information image matching the stitched image of the preset frame is generated.
[0300] There are several ways to determine the stitched image of the preset video frames corresponding to the virtual information image in specific implementations. For example, the cloud server can automatically select the stitched image of the preset video frames using a preset AI recognition algorithm as the stitched image for the AR effect data to be implanted. Alternatively, the server-side user can specify the stitched image of the preset video frames through interactive operations. When the server receives a server-side effect generation interactive control command triggered by a server-side effect generation interactive control operation, it can obtain the specified stitched image of the preset video frames from the server-side effect generation interactive control command, and then generate a virtual information image corresponding to the stitched image of the preset video frames specified by the effect generation command.
[0301] In some embodiments of this specification, objects in the video frame can be identified using image recognition technology as target objects to be matched with the AR effects to be implanted. For example, the target object can be identified as a person (such as a basketball player), an object (such as a basketball or a scoreboard), an animal (such as a cat or a lion), etc.
[0302] In practical implementation, the augmented reality effects input data of the target object can be obtained in response to server-side effects generation interaction control commands. For example, if a server-side user selects a player in a live basketball game video through interactive operation, a server-side effects generation interaction control command corresponding to the interactive operation can be generated. Based on the server-side effects generation interaction control command, the athlete data and goal data can be obtained. The athlete data can include basic data associated with the player, such as name, position name in the basketball game (specific position number, or position name such as center, forward, guard, etc.), and the goal data can include shooting percentage, all of which can be used as augmented reality effects input data.
[0303] In practice, interactive control instructions can be generated based on the server-side special effects first to determine the special effects output type. Then, historical data of the target object can be obtained, and the historical data can be processed according to the special effects data type to obtain augmented reality special effects input data corresponding to the special effects output type.
[0304] For example, in a live basketball game, based on the interactive control instructions generated by the server-side special effects, if the server-side user wants to obtain the shooting accuracy of the target object located within the special effects area, then the distance between the target object's location and the ground projection position of the center of the net can be calculated, and the historical shooting data of the target object within this distance can be obtained as the augmented reality special effects input data for the target object.
[0305] The method for generating special effects for virtual information images can be selected and set as needed. In a specific implementation of this specification, the augmented reality effects input data can be used as input data and processed into a preset 3D model to obtain a virtual information image that matches the target object in the stitched image of the preset video frame.
[0306] For example, by inputting the augmented reality effects input data into a preset 3D model, 3D graphic elements matching the input data can be obtained and combined. The display metadata from the input data and the 3D graphic element data can then be output as a virtual information image matching the target object in the video frame. The specific implementation of the 3D model can be found in the foregoing embodiments.
[0307] In another specific implementation of this specification, the augmented reality effect input data can be used as input data and fed into a preset machine learning model for processing to obtain a virtual information image that matches the target object in the video frame. In specific implementations, the preset machine learning model can be a supervised learning model, an unsupervised learning model, or a semi-supervised learning model (a combination of supervised and unsupervised learning models). The specific model used in this embodiment is not limited. For details on how to generate the virtual information image using a machine learning model, please refer to the foregoing embodiments; further details will not be repeated here.
[0308] In the embodiments of this specification, the generated virtual information image can be a static image, a dynamic image, or a dynamic image containing audio effects. The dynamic image or the dynamic image containing audio effects can be matched with one or more video frames based on the target object.
[0309] In specific implementations, the server can also directly save the virtual information images obtained during the live or quasi-live broadcast process as virtual information images obtained through the interactive terminal during the user interaction process.
[0310] It should be noted that, in the embodiments of this specification, the synthesized video frames displayed on the playback terminal and the synthesized video frames displayed on the interactive terminal are not fundamentally different. They can actually use the same virtual information image or different virtual information images. Correspondingly, the corresponding special effects generation methods can be the same or different. Similarly, the 3D model or machine learning model used in the special effects generation process can be the same model, the same type of model, or completely different models.
[0311] Furthermore, the playback terminal and the interactive terminal can be the same device. Users can directly stream or quasi-stream multi-angle free-view video via the terminal device, which can automatically play multi-angle free-view composite video frames with embedded AR effects. Users can also interact with the terminal device, playing multi-angle free-view video data and multi-angle free-view composite video frames with embedded AR effects based on their interactive operations. Through interaction, users can independently select which target objects' AR effects, i.e., virtual information images, to watch in recorded, rebroadcast, or on-demand videos.
[0312] The data processing methods described in the above embodiments can achieve low-latency playback of multi-angle free-view videos with embedded AR effects. To enable those skilled in the art to better understand and implement the embodiments of this specification, the following provides a corresponding description of the systems and key devices that can implement the above methods.
[0313] In some embodiments of this specification, reference is made toFigure 16 The schematic diagram of the data processing system shown indicates that the data processing system 160 may include: a target object acquisition unit 161, a virtual information image acquisition unit 162, an image synthesis unit 163, and a display unit 164, wherein:
[0314] The target object acquisition unit 161 is adapted to acquire the target object in the video frame of a multi-angle free-view video.
[0315] The virtual information image acquisition unit 162 is adapted to acquire a virtual information image generated based on the augmented reality effects input data of the target object;
[0316] The image synthesis unit 163 is adapted to synthesize the virtual information image with the corresponding video frame to obtain a synthesized video frame.
[0317] The display unit 164 is adapted to display the obtained synthetic video frames.
[0318] In practice, the units may be distributed in different devices, or some units may be located in the same device. Depending on the specific application scenario, the implementation schemes may vary.
[0319] Those skilled in the art will understand that each unit can be implemented by corresponding hardware or a combination of hardware and software. For example, a processor (specifically a CPU or FPGA, etc.) can be used as the target object acquisition unit 161, the virtual information image acquisition unit 162, and the image synthesis unit 163, etc., and a display can be used as the display unit 164.
[0320] The following will illustrate this through some specific application scenarios.
[0321] Reference Figure 3 The schematic diagram of the data processing system shown is illustrated in this embodiment of the invention. Figure 3 As shown, the data processing system 30 may include: a data processing device 31, a server 32, a playback control device 33, and a playback terminal 34, wherein:
[0322] The data processing device 31 is adapted to extract multiple synchronous video frames from the video frames at a specified frame time from the multi-channel video data streams that are synchronously collected in real time from different locations in the field collection area based on video frame extraction instructions, and upload the multiple synchronous video frames at the specified frame time to the server 12.
[0323] The server 32 is adapted to receive multiple synchronized video frames uploaded by the data processing device 31 as an image combination, determine the corresponding parameter data of the image combination and the depth data of each frame image in the image combination, and reconstruct the frame image of the preset virtual viewpoint path based on the corresponding parameter data of the image combination, the pixel data and depth data of the preset frame image in the image combination to obtain the video frame of the corresponding multi-angle free viewpoint video; and in response to the special effects generation instruction, obtain the target object in the video frame specified by the special effects generation instruction, obtain the augmented reality special effects input data of the target object, generate the corresponding virtual information image based on the augmented reality special effects input data of the target object, perform composite processing on the virtual information image and the specified video frame to obtain the composite video frame, and input the composite video frame to the playback control device 34;
[0324] The playback control device 33 is adapted to insert the synthesized video frame data into the video stream to be played;
[0325] The playback terminal 34 is adapted to receive the video stream to be played from the playback control device 33 and play it in real time.
[0326] In practice, the playback control terminal 33 can output the video stream to be played based on control commands.
[0327] As an optional example, the playback control device 33 can select one of multiple data streams as the video stream to be played, or continuously switch between multiple video streams to continuously output the video stream to be played. A broadcast control device can be used as one type of playback control device in this embodiment of the invention. The broadcast control device can be a manual or semi-manual broadcast control device that performs playback control based on external input control commands, or a virtual broadcast control device that can automatically perform broadcast control based on artificial intelligence, big data learning, or preset algorithms.
[0328] By employing the aforementioned data processing system, since it only extracts synchronous video frames at specified times from multiple synchronous video streams to reconstruct multi-angle free-view video and generates virtual information images corresponding to the image combinations specified by the special effects generation instructions, it eliminates the need to upload massive amounts of synchronous video stream data. This distributed system architecture can save significant transmission and server processing resources. Furthermore, under conditions of limited network bandwidth, it can achieve real-time generation of multi-angle free-view composite video frames with augmented reality effects. Therefore, it can achieve low-latency playback of multi-angle free-view augmented reality videos, thus meeting the dual needs of users for a rich visual experience and low latency during video viewing.
[0329] Furthermore, the data processing device 31 performs synchronous video frame capture, the server performs multi-angle free-view video reconstruction, virtual information image acquisition, and multi-angle free-view video and virtual information image synthesis processing (such as fusion processing), the playback control device selects the video stream to be played, and the playback device plays the video. This distributed system architecture can avoid a large amount of data processing on the same device, thus improving data processing efficiency and reducing transmission latency.
[0330] In practical implementation, the server 32 can be implemented through a server cluster consisting of multiple servers. This server cluster can include multiple homogeneous or heterogeneous individual server devices or server clusters. If a heterogeneous server cluster is used, the server devices within the heterogeneous server cluster can be configured according to the characteristics of the different data to be processed.
[0331] Reference Figure 17 The schematic diagram of the server cluster architecture shown in this specification illustrates that, in one embodiment, the heterogeneous server cluster 170 consists of a 3D depth reconstruction service cluster 171 and a cloud-based augmented reality effects generation and rendering server cluster 172, wherein:
[0332] The three-dimensional depth reconstruction service cluster 171 is adapted to reconstruct corresponding multi-angle free-view video based on multiple synchronous video frames extracted from multiple synchronous video streams.
[0333] The cloud-based augmented reality effects generation and rendering server cluster 172 is adapted to respond to effects generation instructions, obtain virtual information images corresponding to the image combination specified by the effects generation instructions, and perform fusion processing on the specified image combination and the virtual information images to obtain multi-angle free-viewpoint fused video frames.
[0334] Based on the different data processing mechanisms, the 3D depth reconstruction service cluster 171 and the cloud-based augmented reality effects generation and rendering server cluster 172 can each include multiple server sub-clusters or server groups. Different server clusters or server groups perform different functions and work together to complete the reconstruction of multi-angle free video frames.
[0335] In a specific implementation, the heterogeneous server cluster 170 may further include an augmented reality effects input data storage database 173, which is suitable for storing augmented reality effects input data that matches target objects in a specified image combination.
[0336] In one embodiment of this specification, a cloud service system composed of a cloud server cluster obtains the first multi-angle free-viewpoint fused video frame based on multiple uploaded synchronous video frames. The cloud service system employs a heterogeneous server cluster. The following will continue to use...Figure 1 The example shown illustrates how to implement a specific application scenario.
[0337] Reference Figure 1 The schematic diagram of the data processing system shown illustrates the deployment scenario of a data processing system for a basketball game. The data processing system 10 includes: a data acquisition array 11 composed of multiple acquisition devices, a data processing device 12, a cloud-based server cluster 13, a playback control device 14, and a playback terminal 15.
[0338] Reference Figure 1 Using the basketball hoop on the left as the focal point, and centering on the focal point, a fan-shaped area on the same plane as the focal point serves as the preset multi-angle free viewing angle range. Each acquisition device in the acquisition array 11 can be positioned in a fan shape at different locations within the on-site acquisition area according to the preset multi-angle free viewing angle range, allowing for real-time synchronous acquisition of video streams from corresponding angles.
[0339] In practical implementation, the acquisition devices in the acquisition array 11 can also be installed in the ceiling area of a basketball court, on basketball hoops, etc. Each acquisition device can be arranged in a straight line, fan shape, arc, circle, or irregular shape. The specific arrangement can be set according to one or more factors such as the specific site environment, the number of acquisition devices, the characteristics of the acquisition devices, and the imaging effect requirements. The acquisition devices can be any device with camera functionality, such as ordinary video cameras, mobile phones, professional video cameras, etc.
[0340] To avoid interfering with the operation of the acquisition equipment, the data processing device 12 can be placed in a non-acquisition area and can be considered as a field server. The data processing device 12 can send streaming commands to each acquisition device in the acquisition array 11 via a wireless local area network. Based on the streaming commands sent by the data processing device 12, each acquisition device in the acquisition array 11 transmits the acquired video data stream to the data processing device 12 in real time. Alternatively, each acquisition device in the acquisition array 11 can transmit the acquired video stream to the data processing device 12 in real time via a switch 17.
[0341] When the data processing device 12 receives a video frame capture instruction, it captures multiple synchronous video frames from the video frame at a specified frame time in the received multi-channel video data stream, and uploads the multiple synchronous video frames at the specified frame time to the server cluster 13 in the cloud.
[0342] Accordingly, the cloud-based server cluster 13 takes the received multiple synchronized video frames as an image combination, determines the corresponding parameter data of the image combination and the depth data of each frame image in the image combination, and performs frame image reconstruction on the preset virtual viewpoint path based on the corresponding parameter data of the image combination, the pixel data and depth data of the preset frame images in the image combination, to obtain the image data of the corresponding multi-angle free-view video; and in response to the special effects generation instruction, obtains the virtual information image corresponding to the image combination specified by the special effects generation instruction, and performs fusion processing on the specified image combination and the virtual information image to obtain the multi-angle free-view fused video frame.
[0343] Servers can be located in the cloud, and in order to process data in parallel more quickly, a cloud server cluster can be formed by multiple different servers or server groups according to the different types of data being processed.
[0344] For example, the cloud server cluster 13 may include: a first cloud server 131, a second cloud server 132, a third cloud server 133, a fourth cloud server 134, and a fifth cloud server 135.
[0345] The first cloud server 131 can be used to determine the corresponding parameter data of the image combination; the second cloud server 132 can be used to determine the depth data of each frame image in the image combination; the third cloud server 133 can use the Depth Image Based Rendering (DIBR) algorithm to reconstruct frame images of a preset virtual viewpoint path based on the corresponding parameter data of the image combination, the pixel data and depth data of the image combination; the fourth cloud server 134 can be used to generate multi-angle free-view video; and the fifth cloud server 135 can be used to respond to the special effects generation instruction, obtain the virtual information image corresponding to the image combination specified by the special effects generation instruction, and fuse the image combination with the virtual information image to obtain a multi-angle free-view fused video frame.
[0346] It is understood that the first cloud server 131, the second cloud server 132, the third cloud server 133, the fourth cloud server 134, and the fifth cloud server 135 may also be a server group composed of server arrays or server sub-clusters, and the embodiments of the present invention do not impose any restrictions.
[0347] Depending on the data processing and specific data processing mechanisms, each cloud server or cloud server cluster can use devices with different hardware configurations. For example, devices such as the fourth cloud server 134 and the fifth cloud server 135 that need to process a large number of images can use devices including graphics processing units (GPUs) or GPU groups.
[0348] In some embodiments of this specification, the GPU may employ a Compute Unified Device Architecture (CUDA) parallel programming architecture to render pixels from corresponding groups of texture maps and depth maps in a selected image composition. CUDA is a novel hardware and software architecture for allocating and managing computations on the GPU as data-parallel computing devices without mapping them to a graphics application programming interface (API).
[0349] When programming with CUDA, a GPU can be viewed as a computing device capable of executing a large number of threads in parallel. It runs as the main Central Processing Unit (CPU) or a coprocessor of the host machine; in other words, the data-parallel, computationally intensive parts of applications running on the host machine are offloaded to the GPU.
[0350] In a specific implementation, the server cluster 13 in the cloud can store the pixel data and depth data of the image combination in the following manner:
[0351] Based on the pixel and depth data of the image combination, a stitched image corresponding to the frame time is generated. The stitched image includes a first field and a second field, wherein the first field includes pixel data of a preset frame image in the image combination, and the second field includes a second field of depth data of the preset frame image in the image combination; and stores the stitched image of the image combination and the corresponding parameter data of the image combination. The acquired stitched image and corresponding parameter data can be stored in a data file. When it is necessary to obtain the stitched image or parameter data, it can be read from the corresponding storage space according to the corresponding storage address in the header file of the data file.
[0352] Then, the playback control device 14 can insert the received multi-angle free-viewpoint video fusion video frame data into the video stream to be played, and the playback terminal 15 receives the video stream to be played from the playback control device 14 and plays it in real time. The playback control device 14 can be a manual playback control device or a virtual playback control device. In specific implementations, a dedicated server capable of automatically switching video streams can be set up as a virtual playback control device to control the data source. A director's control device, such as a director's console, can be used as a playback control device in this embodiment of the invention.
[0353] It is understood that the data processing device 12 can be placed in a non-collection area on-site or in the cloud, depending on the specific scenario. The server (cluster) and playback control device can be placed in a non-collection area on-site, in the cloud, or on the terminal access side, depending on the specific scenario. The above embodiments are not intended to limit the specific implementation and protection scope of the present invention.
[0354] The data processing system used in the embodiments of this specification can not only play multi-angle free-view video in low-latency scenarios such as live streaming and quasi-live streaming, but also play multi-angle free-view video in scenarios such as recording and rebroadcasting based on user interaction.
[0355] Continue to refer to Figure 3 In a specific implementation, the data processing system 30 may also include an interactive terminal 35. The server 32 may respond to the image reconstruction command from the interactive terminal 35, determine the interaction frame time information of the interaction time, and send the spliced image of the corresponding image combination preset frame image stored at the corresponding interaction frame time and the parameter data corresponding to the corresponding image combination to the interactive terminal 35.
[0356] The interactive terminal 35 sends the image reconstruction instruction to the server based on the interactive operation, and selects the corresponding pixel data, depth data and corresponding parameter data in the stitched image according to the virtual viewpoint position information determined by the interactive operation and according to the preset rules. The selected pixel data and depth data are combined with the parameter data for rendering to reconstruct the video frame of the multi-angle free viewpoint video corresponding to the virtual viewpoint position to be interacted with and plays it.
[0357] The preset rules can be set according to specific scenarios, as detailed in the foregoing method embodiments.
[0358] Furthermore, the interaction frame timing information can be determined based on a trigger operation from the interactive terminal 35. This trigger operation can be a user-input trigger operation or an automatically generated trigger operation by the interactive terminal. For example, the interactive terminal can automatically initiate a trigger operation when it detects the presence of a multi-angle free viewpoint data frame. When the user manually triggers the interaction, it can be based on the timing information of the user selecting to trigger the interaction after the interactive terminal displays interactive prompts, or it can be based on historical timing information received by the interactive terminal from user actions. This historical timing information can be timing information prior to the current playback moment.
[0359] In a specific implementation, the interactive terminal 35 can combine and render the pixel data and depth data of the stitched image of the preset frame image in the image combination of the acquired interactive frame time, based on the stitched image and corresponding parameter data of the image combination of the acquired interactive frame time, the interactive frame time information, and the virtual viewpoint position information of the interactive frame time, using the same method as in step S44 above, to obtain the image of the multi-angle free-view video corresponding to the virtual viewpoint position of the interaction, and start playing the multi-angle free-view video at the virtual viewpoint position of the interaction.
[0360] Using the above solution, multi-angle free-view video corresponding to the virtual viewpoint position of the interaction can be generated in real time based on the image reconstruction instructions from the interactive terminal, which can further enhance the user's interactive experience.
[0361] In some data processing systems described in this manual, please refer to... Figure 3 The server 32 can also generate interactive control commands based on server-side special effects, generate virtual information images corresponding to the spliced images of the preset video frames indicated by the server-side special effects generation interactive control commands, and store them. Through this solution, by pre-generating virtual information images corresponding to the spliced images of the preset frame images, subsequent rendering and playback can be performed directly when playback is required, thereby reducing time delay, further enhancing the user's interactive experience, and improving the user's visual experience.
[0362] In terms of specific application scenarios, the data processing system can not only be used in live and near-live streaming scenarios to achieve low-latency playback of multi-angle free-view videos with AR effects, but also, based on user interaction, enable playback of multi-angle free-view videos with AR effects in any video playback scenario, such as pre-recorded or rebroadcast videos. As an example, users can interact with the server through an interactive terminal to obtain virtual information images corresponding to the stitched images of preset video frames, and render them on the interactive terminal, thereby achieving playback of multi-angle free-view composite video frames with AR effects. The following describes some application scenarios in detail.
[0363] Based on referenceFigure 3 The server 32 is also adapted to respond to a user-end special effects generation interaction command from the interactive terminal, obtain the virtual information image corresponding to the spliced image of the preset video frame, and send the virtual information image corresponding to the spliced image of the preset video frame to the interactive terminal 35.
[0364] The interactive terminal 35 is adapted to combine the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the time of the interactive frame with the virtual information image to obtain a synthesized video frame and play it.
[0365] The specific methods for the server to acquire and generate the virtual information image and the virtual information image can be found in the foregoing method embodiments, and will not be described in detail here.
[0366] To enable those skilled in the art to better understand and implement this method, the following first introduces a schematic diagram of the video effect displayed by the playback terminal in the embodiments of this specification through a specific application scenario.
[0367] Reference Figures 18 to 20 The diagram shown illustrates the video effect of the display interface of the playback terminal. Assuming... Figure 18 The playback interface Sr1 of the playback terminal T1 shows the (T-1)th video frame, which depicts the athlete sprinting towards the finish line from the athlete's right-hand perspective. Assume the data processing device captures multiple synchronized video frames from frame T to frame T+1 of the video stream and uploads them to the server. The server uses the received synchronized video frames from frame T to T+1 as an image combination. On one hand, based on the corresponding parameter data of the image combination, the pixel data and depth data of the preset frame images in the image combination, the server performs frame image reconstruction on a preset virtual viewpoint path to obtain the corresponding multi-angle free-viewpoint video frames. On the other hand, in response to the special effects generation command from the server user, the server obtains the virtual information image corresponding to the image combination specified by the special effects generation command. Then, the virtual information image is superimposed and rendered on the specified image combination to obtain the multi-angle free-viewpoint fused video frames corresponding to frames T to T+1, and the display effect on the playback terminal T1 is as follows: Figure 19 and Figure 20 As shown, where, Figure 19 The playback interface Sr2 displays the effect of the T-frame video, with the perspective switching to the athlete's front. It's evident that AR effects images are embedded on top of the real-world image. These images show the athlete sprinting towards the finish line, along with the embedded AR effects, including the athlete's basic information board M1 and two virtually generated footprints M2 matching the athlete's steps. This is to distinguish the virtual information images corresponding to the AR effects from the real images corresponding to the multi-angle free-view video frames. Figure 19 and Figure 20Solid lines represent real images, while dashed lines represent virtual information images corresponding to AR effects. The basic information panel M1 displays information such as the athlete's name, nationality, competition number, and best historical performance. Figure 20 The image shown is a rendering of the T+1th frame of the video. The perspective has shifted further to the left of the athlete. As can be seen from the image displayed on the Sr3 playback interface, the athlete has crossed the finish line. The specific information contained in the basic information panel M1 can be updated in real time over time. Figure 19 As can be seen, the athlete's current result was added, the position and shape of the footprint M2 changed with the athlete's steps, and a pattern M3 was added to indicate that the athlete won first place.
[0368] The playback terminal in the embodiments of this specification can be any one or more types of terminal devices such as television, computer, mobile phone, in-vehicle equipment, and projection equipment.
[0369] To enable those skilled in the art to better understand and implement the operating principle of the interactive terminal in the embodiments of the present invention, the following detailed description is provided with reference to the accompanying drawings and specific application scenarios.
[0370] Reference Figure 21 The schematic diagram of the interactive terminal shown is illustrated in some embodiments of this specification, such as... Figure 21 As shown, the interactive terminal 210 may include a first display unit 211, a virtual information image acquisition unit 212, and a second display unit 213, wherein:
[0371] The first display unit 211 is adapted to display images of multi-angle free-view video in real time. The images of the multi-angle free-view video are reconstructed by parameter data of an image combination formed by multiple synchronous video frame images at a specified frame time, pixel data and depth data of the image combination, and the multiple synchronous video frames include frame images from different shooting angles.
[0372] The virtual information image acquisition unit 212 is adapted to acquire a virtual information image corresponding to a specified frame moment of the special effects display mark in the multi-angle free-view video image in response to a trigger operation of the special effects display mark.
[0373] The second display unit 213 is adapted to overlay the virtual information image onto the video frame of the multi-angle free-view video.
[0374] Using the aforementioned interactive terminal, end users can interact with and view multi-angle free-view video images with embedded AR effects, which can enrich the user's visual experience.
[0375] Reference Figure 22The schematic diagram of another interactive terminal shown in this specification may be used in other embodiments where the interactive terminal 220 may include:
[0376] The video stream acquisition unit 221 is adapted to acquire a video stream to be played from a playback control device in real time. The video stream to be played includes video data and an interaction identifier, and the interaction identifier is associated with a specified frame time of the video stream to be played.
[0377] The playback and display unit 222 is adapted to play and display the video and interactive icons of the video stream to be played in real time.
[0378] The interactive data acquisition unit 223 is adapted to acquire interactive data corresponding to the specified frame time in response to the trigger operation of the interactive identifier. The interactive data includes a virtual information image corresponding to the spliced image of multi-angle free-view video frames and the preset video frames.
[0379] The interactive display unit 224 is adapted to display a composite video frame with multiple free-viewpoint perspectives at a specified frame time based on the interactive data.
[0380] The switching unit 225 is adapted to trigger a switch to the video stream to be played, which is acquired in real time from the playback control device by the video stream acquisition unit 221 and displayed in real time by the playback display unit 222, when an interaction end signal is detected.
[0381] The interactive data can be generated by the server and transmitted to the interactive terminal, or it can be generated by the interactive terminal itself.
[0382] During video playback, the interactive terminal can obtain the data stream to be played in real time from the playback control device and display corresponding interactive indicators at the appropriate frame moments. For example, the interactive indicator can be displayed on the progress bar, or it can be displayed directly on the display screen.
[0383] Reference Figure 3 and Figure 23 The interactive terminal T2 displays an interactive identifier V1 on its display interface Sr20. When the user does not select a trigger, the interactive terminal T2 can continue to read subsequent video data. When the user slides to select a trigger in the direction indicated by the arrow on the interactive identifier V1, the interactive terminal T2 receives feedback, generates an image reconstruction instruction for the specified frame time of the corresponding interactive identifier, and sends it to the server 32.
[0384] For example, when a user selects to trigger the currently displayed interactive identifier V1, the interactive terminal T2 receives feedback and generates an image reconstruction instruction for the corresponding specified frame times Ti to Ti+2 of the interactive identifier V1, and sends it to the server 32. The server 32 can send multiple frame images corresponding to the specified frame times Ti to Ti+1 according to the image reconstruction instruction.
[0385] Furthermore, at frame Ti+1, as Figure 24 As shown, the display interface Sr20 displays the interactive icon Ir. When the user clicks the interactive icon Ir, the interactive terminal T2 can obtain the corresponding virtual information image from the server.
[0386] Then, the multi-angle free-viewpoint fused image corresponding to frame Ti+2 can be displayed on the interactive terminal T2, such as... Figure 25 and Figure 26 The image shown is a video rendering of the interactive interface of the interactive terminal. Figure 25 The interactive interface Sr20 displays the effect of AR integration on the Ti+1 frame image. The perspective switches to the athlete's front, and it can be seen that virtual information images corresponding to AR effects are integrated onto the real image. The Ti+1 frame image displayed in the interactive interface Sr20 shows the real scene of the athlete sprinting towards the finish line, as well as the virtual information images, including the athlete's basic information board M4 and footprints M5 matching the athlete's steps. This is to distinguish between AR effects and real images. Figure 25 and Figure 26 Solid lines are used to identify real images, while dashed lines represent virtual information images. The basic information panel M4 displays information such as the athlete's name, nationality, competition number, and best historical performance. Figure 26 The image shown is a rendering of the (Ti+2)th frame of the video. The perspective shifts further to the left of the athlete, showing that the athlete has crossed the finish line. The information displayed on the M4 basic information panel can be updated in real time. Figure 26 As can be seen, the athlete's current result was added, the position and shape of the footprint M5 changed with the athlete's steps, and a pattern M6 was added to indicate that the athlete won first place.
[0387] The interactive terminal T2 can generate interactive data for interaction based on the multiple video frames, and can use image reconstruction algorithms to process the multi-angle free-view data of the interactive data, and obtain virtual information images from the server, and then play the multi-angle free-view video at the specified frame time, and play the multi-angle free-view synthesized video frame with AR effects embedded in the specified frame.
[0388] In specific implementations, the interactive terminal of this invention can be any one or more of the following types: an electronic device with touch screen functionality, a head-mounted virtual reality (VR) terminal, an edge node device connected to a display, or an IoT (Internet of Things) device with display functionality.
[0389] As described in the previous embodiment, to more accurately generate virtual information images that match video frames of multi-angle free-viewpoint videos, the target object corresponding to the stitched image of the preset video frame image can be identified, and the augmented reality effect input data of the target object can be obtained. In specific implementations, the interactive data may also include augmented reality effect input data of the target object, which may include at least one of the following: on-site analysis data, information data of the collected target object, information data of equipment associated with the collected target object, information data of items deployed on-site, and information data of logos displayed on-site. Based on the interactive data, the virtual information image can be generated, and then the multi-angle free-viewpoint composite video frame can be generated, thereby making the implanted AR effects richer and more targeted. As a result, end users can gain a deeper, more comprehensive, and more professional understanding of the content they are watching, further enhancing the user's visual experience.
[0390] This specification also provides corresponding server implementation schemes, which can be found in the embodiments. Figure 27 The diagram shown illustrates the structure of a server. In some embodiments of this specification, such as... Figure 27 As shown, server 270 may include: image reconstruction unit 271, virtual information image generation unit 272, and data transmission unit 273, wherein:
[0391] The image reconstruction unit 271 is adapted to respond to an image reconstruction command from an interactive terminal, determine the interaction frame time information at the interaction time, and obtain the spliced image of the preset frame images in the image combination at the corresponding interaction frame time and the corresponding parameter data of the image combination.
[0392] The virtual information image generation unit 272 is adapted to generate a virtual information image corresponding to the spliced image of the video frame indicated by the special effects generation interactive control command in response to the special effects generation interactive control command.
[0393] The data transmission unit 273 is adapted to interact with an interactive terminal, including: transmitting the spliced image of a preset video frame in the image combination at the corresponding interactive frame time and the corresponding parameter data of the image combination to the interactive terminal, so that the interactive terminal, based on the virtual viewpoint position information determined by the interactive operation, selects the corresponding pixel data, depth data and corresponding parameter data in the spliced image according to preset rules, combines and renders the selected pixel data and depth data, reconstructs and plays the multi-angle free-view video image corresponding to the virtual viewpoint position at the interactive frame time; and transmitting the virtual information image corresponding to the spliced image of the preset frame image indicated by the special effects generation interactive control command to the interactive terminal, so that the interactive terminal synthesizes the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the interactive frame time with the virtual information image to obtain a multi-angle free-view synthesized video frame and plays it.
[0394] This specification also provides another server in its embodiments, see below. Figure 28 The schematic diagram of the server shown indicates that server 280 may include:
[0395] The data receiving unit 281 is adapted to receive multiple synchronous video frames at a specified frame time extracted from multiple synchronous video streams as an image combination, wherein the multiple synchronous video frames contain frame images from different shooting angles.
[0396] The parameter data calculation unit 282 is adapted to determine the corresponding parameter data of the image combination;
[0397] The depth data calculation unit 283 is adapted to determine the depth data of each frame in the image combination;
[0398] The video data acquisition unit 284 is adapted to perform frame image reconstruction on the preset virtual viewpoint path based on the parameter data of the image combination, the pixel data and depth data of the preset frame image in the image combination, and to obtain the video frames of the corresponding multi-angle free viewpoint video.
[0399] The first virtual information image generation unit 285 is adapted to respond to a special effects generation instruction, acquire a target object in a video frame specified by the special effects generation instruction, acquire augmented reality special effects input data of the target object, and generate a corresponding virtual information image based on the augmented reality special effects input data of the target object;
[0400] The image synthesis unit 286 is adapted to synthesize the virtual information image with the specified video frame to obtain a synthesized video frame;
[0401] The first data transmission unit 287 is adapted to output a composite video frame to be inserted into a video stream to be played.
[0402] Reference Figure 29 This specification also provides another server in the embodiments. Server 290 differs from server 280 in that server 290 may further include: a stitched image generation unit 291 and a first data storage unit 292, wherein:
[0403] The image stitching generation unit 291 is adapted to generate a stitched image corresponding to the image combination based on the pixel data and depth data of the image combination. The stitched image includes a first field and a second field, wherein the first field includes the pixel data of a preset frame image in the image combination, and the second field includes the depth data of the image combination.
[0404] The first data storage unit 292 is adapted to store the stitched image of the image combination and the corresponding parameter data of the image combination.
[0405] In some embodiments of this specification, reference continues to be made to... Figure 29 The server 290 may further include: a data extraction unit 293 and a second data transmission unit 294, wherein:
[0406] The data extraction unit 293 is adapted to respond to the image reconstruction command from the interactive terminal, determine the interaction frame time information at the interaction time, and obtain the spliced image of the preset frame images in the image combination at the corresponding interaction frame time and the corresponding parameter data of the image combination.
[0407] The second data transmission unit 294 is adapted to send the spliced image of the corresponding image combination preset frame image at the corresponding interaction frame time and the corresponding parameter data of the corresponding image combination to the interactive terminal, so that the interactive terminal selects the corresponding pixel data and depth data and the corresponding parameter data in the spliced image according to the virtual viewpoint position information determined by the interactive operation and according to the preset rules, combines and renders the selected pixel data and depth data, reconstructs the video frame of the multi-angle free viewpoint video corresponding to the virtual viewpoint position at the interaction frame time, and plays it.
[0408] In specific implementations, the server described in some embodiments of this specification can also be used to generate and store augmented reality effect input data corresponding to the stitched image of the preset frame image, so as to facilitate the subsequent generation of virtual information images, improve the user's visual experience, and make effective use of data resources. (Continue to refer to...) Figure 29 The server 290 may further include: a second virtual information image generation unit 295 and a second data storage unit 296, wherein:
[0409] The second virtual information image generation unit 295 is adapted to generate a virtual information image corresponding to the spliced image of the preset frame image indicated by the server-side special effects generation interactive control command in response to the server-side special effects generation interactive control command.
[0410] The second data storage unit 296 is adapted to store virtual information images corresponding to the spliced images of preset frame images.
[0411] In specific implementation, we will continue to refer to Figure 29 The server 290 may further include: a second virtual information image acquisition unit 297 and a third data transmission unit 298, wherein:
[0412] The second virtual information image acquisition unit 297 is adapted to acquire the virtual information image corresponding to the spliced image of the preset frame image in response to the user terminal special effect generation interaction instruction from the interactive terminal after receiving the image reconstruction instruction;
[0413] The third data transmission unit 298 is adapted to send the virtual information image corresponding to the spliced image of the preset frame image to the interactive terminal, so that the interactive terminal can synthesize the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the time of the interactive frame with the virtual information image to obtain a multi-angle free-view synthesized video frame for playback.
[0414] It should be noted that the augmented reality effects input data in the embodiments of this specification can be, for example, athlete effects data and goal effects data in the basketball game scene described above. It is understood that the augmented reality effects input data in the embodiments of this specification is not limited to the above example types. In the case of a basketball game scene, corresponding augmented reality effects input data can also be generated based on various target objects contained in the on-site images collected from images of coaches, advertising signs, etc.
[0415] In practice, corresponding virtual information images can be generated based on one or more factors such as the specific application scenario, the characteristics of the target object, the related objects of the target object, and the specific special effects generation model (such as a preset 3D model, a preset machine learning model, etc.).
[0416] Those skilled in the art will understand that the specific units in each electronic device in the embodiments of this specification can be implemented by corresponding circuits. For example, the data acquisition units involved in the above embodiments can be implemented by processors, CPUs, input interfaces, etc.; the data storage units involved in the above embodiments can be implemented by various storage devices such as disks, EPROMs, and ROMs; and the data transmission units involved in the above embodiments can be implemented by communication interfaces, communication lines (wired / wireless), etc., which will not be listed here one by one.
[0417] This specification also provides a computer-readable storage medium storing computer instructions that, when executed, can perform the steps of the depth map processing method or the video reconstruction method described in any of the foregoing embodiments. Specific steps can be found in the descriptions of the foregoing embodiments, and will not be repeated here.
[0418] In specific implementations, the computer-readable storage medium may include, for example, any suitable type of memory cell, memory device, memory article, memory medium, storage device, storage article, storage medium and / or storage unit, such as memory, removable or non-removable medium, erasable or non-erasable medium, writable or rewritable medium, digital or analog medium, hard disk, floppy disk, optical disc read-only memory (CD-ROM), recordable optical disc (CD-R), rewritable optical disc (CD-RW), optical disc, magnetic medium, magneto-optical medium, removable memory card or disk, various types of digital universal optical disc (DVD), magnetic tape, cassette tape, etc.
[0419] Computer instructions may include any suitable type of code implemented using any appropriate high-level, low-level, object-oriented, visual, compiled, and / or interpreted programming language, such as source code, compiled code, interpreted code, executable code, static code, dynamic code, encrypted code, etc.
[0420] For details on the specific implementation, working principle, function, and effect of each device, system, equipment, or system in the embodiments of this specification, please refer to the specific description in the corresponding method embodiments.
[0421] While the embodiments disclosed in this specification are as described above, the present invention is not limited thereto. Any person skilled in the art can make various modifications and alterations without departing from the spirit and scope of the embodiments described in this specification; therefore, the scope of protection of the present invention should be determined by the scope defined in the claims.
Claims
1. A data processing method, comprising: The target object in the video frame of a multi-angle free-view video is obtained. The video frame of the multi-angle free-view video is obtained by reconstructing a preset virtual viewpoint path based on the corresponding parameter data of the image combination formed by multiple synchronous video frames at a specified frame time extracted from multiple synchronous video streams, the pixel data and depth data of the preset frame image in the image combination, and the pixel data and depth data of the image combination. The pixel data and depth data of the image combination are stored using the stitched image at the corresponding frame time. The stitched image includes a first field and a second field, wherein the first field includes the pixel data of the preset frame image in the image combination, and the second field includes the depth data of the preset frame image in the image combination. The multiple synchronous video frames contain frame images from different shooting angles. Acquiring a virtual information image generated from augmented reality effects input data based on the target object includes: obtaining the position of the target object in the video frame of the multi-angle free-view video based on 3D calibration, and obtaining a virtual information image that matches the position of the target object; The virtual information image is synthesized with the corresponding video frame and then displayed so that the virtual information image in the resulting synthesized video frame changes synchronously with the target object in the video frame of the multi-angle free-view video.
2. The data processing method according to claim 1, wherein the step of synthesizing the virtual information image with the corresponding video frame and displaying it includes: Based on the frame time sequence and the virtual viewpoint position of the corresponding frame time, the virtual information image of the corresponding frame time is synthesized with the video frame of the corresponding frame time and then displayed.
3. The data processing method according to any one of claims 1 to 2, wherein the step of synthesizing the virtual information image with the corresponding video frame and displaying it comprises at least one of the following: The virtual information image is fused with the corresponding video frame to obtain a fused video frame, and the fused video frame is displayed. The virtual information image is superimposed on the corresponding video frame to obtain a superimposed composite video frame, which is then displayed.
4. The data processing method according to claim 3, characterized in that, The display of the fused video frames includes: The fused video frames are inserted into the video stream to be played and displayed.
5. The data processing method according to any one of claims 1 to 2, wherein acquiring the target object in the video frame of the multi-angle free-view video includes: In response to the special effects generation interactive control command, the target object in the video frame of the multi-angle free-view video is obtained.
6. The data processing method according to claim 5, wherein acquiring the virtual information image generated based on the augmented reality effects input data of the target object comprises: Based on the augmented reality effects input data of the target object, a virtual information image corresponding to the target object is generated according to a preset effects generation method.
7. A data processing method, comprising: Receive multiple synchronous video frames at a specified frame time from multiple synchronous video streams as an image combination, wherein the multiple synchronous video frames contain frame images from different shooting angles; Determine the corresponding parameter data for the image combination; Determine the depth data of each frame in the image combination; Based on the parameter data of the image combination, the pixel data and depth data of the preset frame images in the image combination, the preset virtual viewpoint path is reconstructed to obtain the video frames of the corresponding multi-angle free viewpoint video. In response to an effects generation command, a target object in a video frame specified by the effects generation command is obtained, augmented reality effects input data of the target object is obtained, and a corresponding virtual information image is generated based on the augmented reality effects input data of the target object. The virtual information image is combined with the specified video frame to obtain a composite video frame. The virtual information image in the composite video frame changes synchronously with the target object in the video frame of the multi-angle free-view video. Display the synthesized video frames; The step of generating a corresponding virtual information image based on the augmented reality effects input data of the target object includes: taking the augmented reality effects input data of the target object as input, and using a preset first effects generation method to generate a virtual information image matching the target object in the corresponding video frame based on the position of the target object in the video frame of the multi-angle free-view video obtained by three-dimensional calibration; the method further includes: generating a corresponding stitched image of the image combination based on the pixel data and depth data of the image combination, wherein the stitched image includes a first field and a second field, wherein the first field includes the pixel data of the preset frame image in the image combination, and the second field includes the depth data of the image combination.
8. The data processing method according to claim 7, wherein the step of responding to an effect generation instruction, acquiring a target object in a video frame specified by the effect generation instruction, and acquiring augmented reality effect input data of the target object, comprises: Generate interactive control commands based on server-side special effects, and determine the type of special effects output; The historical data of the target object is obtained, and the historical data is processed according to the special effect output type to obtain augmented reality special effect input data corresponding to the special effect output type.
9. The data processing method according to claim 7, wherein generating a corresponding virtual information image based on the augmented reality effects input data of the target object comprises at least one of the following: The augmented reality effects input data of the target object is input into a preset three-dimensional model. Based on the position of the target object in the video frame of the multi-angle free-view video obtained by three-dimensional calibration, a virtual information image matching the target object is output. The augmented reality effects input data of the target object is input into a preset machine learning model. Based on the position of the target object in the video frames of the multi-angle free-view video obtained by 3D calibration, a virtual information image matching the target object is output.
10. The data processing method according to claim 7, wherein the step of combining the virtual information image with the specified video frame to obtain a combined video frame comprises: Based on the position of the target object in the specified video frame obtained by 3D calibration, the virtual information image is fused with the specified video frame to obtain a fused video frame.
11. The data processing method according to claim 7, characterized in that, The step of displaying the synthesized video frames includes: The synthesized video frame is inserted into the video stream to be played on the playback control device for playback via the playback terminal.
12. The data processing method according to claim 7, further comprising: Store the stitched image of the image combination and the corresponding parameter data of the image combination; In response to an image reconstruction command from an interactive terminal, the interaction frame time information at the interaction time is determined, and the stitched image of the preset frame images in the image combination at the corresponding interaction frame time and the corresponding parameter data of the image combination are obtained and sent to the interactive terminal. This allows the interactive terminal to select the corresponding pixel data, depth data and corresponding parameter data in the stitched image according to preset rules based on the virtual viewpoint position information determined by the interaction operation, and to combine and render the selected pixel data and depth data to reconstruct the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the interaction frame time and play it.
13. The data processing method according to claim 12, further comprising: In response to the server-side special effects generation interactive control command, a virtual information image corresponding to the spliced image of the preset video frame indicated by the server-side special effects generation interactive control command is generated; The virtual information image corresponding to the stitched image of the preset video frames is stored.
14. The data processing method according to claim 13, further comprising, upon receiving the image reconstruction instruction: In response to a user-side special effects generation interaction command from an interactive terminal, a virtual information image corresponding to the spliced image of the preset video frames is obtained; The virtual information image corresponding to the stitched image of the preset video frame is sent to the interactive terminal, so that the interactive terminal can synthesize the video frame of the multi-angle free view video corresponding to the virtual viewpoint position at the time of the interactive frame with the virtual information image to obtain a synthesized video frame and display it.
15. The data processing method according to claim 14, further comprising: In response to the user's exit interaction command for special effects, the acquisition of the virtual information image corresponding to the spliced image of the preset video frame is stopped.
16. The data processing method according to claim 14, wherein obtaining the virtual information image corresponding to the stitched image of the preset video frames in response to a user-terminal special effects generation interaction command from an interactive terminal includes: Based on the user-side special effects, generate interactive instructions to determine the target object in the spliced image of the preset video frames; Obtain a virtual information image that matches the target object in the preset video frame.
17. The data processing method according to claim 16, wherein acquiring the virtual information image matching the target object in the preset video frame comprises: A virtual information image matching the target object is generated based on the position of the target object in the preset video frame obtained in advance based on 3D calibration.
18. The data processing method according to any one of claims 14 to 17, wherein sending the virtual information image corresponding to the stitched image of the preset video frame to the interactive terminal, such that the interactive terminal performs a composite processing on the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the time of the interactive frame and the virtual information image to obtain a composite video frame, includes: The virtual information image corresponding to the stitched image of the preset video frame is sent to the interactive terminal, so that the interactive terminal overlays the virtual information image on the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the time of the interactive frame, thereby obtaining the overlaid composite video frame.
19. A data processing method, comprising: In response to an image reconstruction command from an interactive terminal, the interaction frame time information at the moment of interaction is determined, and the stitched image of the preset frame images in the image combination at the corresponding interaction frame time and the corresponding parameter data of the image combination are obtained and sent to the interactive terminal. This allows the interactive terminal to select the corresponding pixel data, depth data and corresponding parameter data in the stitched image according to preset rules based on the virtual viewpoint position information determined by the interaction operation, and to combine and render the selected pixel data and depth data to reconstruct the video frame of the multi-angle free viewpoint video corresponding to the virtual viewpoint position at the moment of interaction and play it. In response to an interactive control command for generating special effects, a virtual information image corresponding to a stitched image of a preset video frame indicated by the interactive control command for generating special effects is obtained; wherein, in response to the interactive control command for generating special effects, obtaining the virtual information image corresponding to a stitched image of a preset video frame indicated by the interactive control command for generating special effects includes: in response to the interactive control command for generating special effects, obtaining a target object in the video frame indicated by the interactive control command for generating special effects; obtaining a virtual information image pre-generated based on augmented reality special effects input data of the target object; the stitched image of the preset video frame is generated based on pixel data and depth data of the image combination at the interactive frame time, the stitched image includes a first field and a second field, the first field includes pixel data of the preset frame image in the image combination, and the second field includes depth data of the image combination; the image combination at the interactive frame time is obtained based on multiple synchronous video frames extracted from multiple synchronous video streams at a specified frame time, the multiple synchronous video frames containing frame images from different shooting angles; The virtual information image corresponding to the stitched image of the preset video frame is sent to the interactive terminal, so that the interactive terminal will synthesize the video frame of the multi-angle free view video corresponding to the virtual viewpoint position at the time of the interactive frame with the virtual information image to obtain a synthesized video frame. The synthesized video frames are displayed, and the virtual information images in the synthesized video frames change synchronously with the target objects in the video frames of the multi-angle free-view video.
20. A data processing method, comprising: The system displays video frames from multiple free-viewpoint videos in real time. These frames are based on parameter data from an image combination formed by multiple synchronous video frames captured at a specified frame time from multiple synchronous video streams, pixel data from a preset frame image in the image combination, and depth data from the preset frame image. Frame image reconstruction is performed on a preset virtual viewpoint path. The pixel data and depth data of the image combination are stored using a stitched image at the corresponding frame time. The stitched image includes a first field and a second field. The first field includes pixel data from the preset frame image in the image combination, and the second field includes depth data from the preset frame image in the image combination. The multiple synchronous video frames contain frame images from different shooting angles. In response to a triggering operation of a special effects display identifier in a video frame of the multi-angle free-view video, a virtual information image of a video frame corresponding to a specified frame moment of the special effects display identifier is obtained, including: obtaining a virtual information image of a target object in a video frame at a specified frame moment corresponding to the special effects display identifier; The virtual information image is synthesized with the corresponding video frame and then displayed, including: Based on the position of the target object in the video frame at the specified time determined by 3D calibration, the virtual information image is superimposed on the video frame at the specified time to obtain a superimposed composite video frame, which is then displayed. The virtual information image changes synchronously with the target object in the video frame of the multi-angle free-view video.
21. The data processing method according to claim 20, wherein the step of obtaining a virtual information image of a video frame corresponding to a specified frame time of the special effects display identifier in the image of the multi-angle free-view video in response to a triggering operation of the special effects display identifier includes: Obtain the virtual information image of the target object in the video frame at the specified frame time corresponding to the special effect display identifier.
22. A data processing system, comprising: The target object acquisition unit is adapted to acquire target objects in video frames of multi-angle free-view video. The video frames of the multi-angle free-view video are obtained by reconstructing frame images of a preset virtual viewpoint path based on the parameter data of an image combination formed by multiple synchronous video frames at a specified frame time extracted from multiple synchronous video streams, the pixel data and depth data of a preset frame image in the image combination, and the pixel data and depth data of the image combination. The pixel data and depth data of the image combination are stored using a stitched image at the corresponding frame time. The stitched image includes a first field and a second field, wherein the first field includes the pixel data of the preset frame image in the image combination, and the second field includes the depth data of the preset frame image in the image combination. The multiple synchronous video frames contain frame images from different shooting angles. A virtual information image acquisition unit is adapted to acquire a virtual information image generated based on augmented reality effect input data of the target object, including: obtaining the position of the target object in the video frame of the multi-angle free-view video based on the three-dimensional calibration, and obtaining a virtual information image that matches the position of the target object; The image synthesis unit is adapted to synthesize the virtual information image with the corresponding video frame to obtain a synthesized video frame; The display unit is adapted to display the obtained synthetic video frame, wherein the virtual information image in the synthetic video frame changes synchronously with the target object in the video frame of the multi-angle free-view video.
23. A data processing system, comprising: Data processing equipment, server, playback control equipment, and playback terminal, including: The data processing device is adapted to extract multiple synchronous video frames from multiple video data streams that are synchronously collected in real time from different locations in the field collection area based on video frame extraction instructions, and upload the multiple synchronous video frames at the specified frame time to the server. The server is adapted to receive multiple synchronized video frames uploaded by the data processing device as an image combination, determine the corresponding parameter data of the image combination and the depth data of each frame image in the image combination, and reconstruct frame images of a preset virtual viewpoint path based on the corresponding parameter data of the image combination, the pixel data and depth data of the preset frame images in the image combination, to obtain video frames of the corresponding multi-angle free-view video; and in response to a special effects generation command, acquire a target object in the video frame specified by the special effects generation command, acquire augmented reality special effects input data of the target object, and generate a corresponding virtual special effects based on the augmented reality special effects input data of the target object. The virtual information image is synthesized with the specified video frame to obtain a synthesized video frame, which is then input to a playback control device. The pixel data and depth data of the image combination are stored using a stitched image at the corresponding frame time. The stitched image includes a first field and a second field. The first field includes pixel data of a preset frame image in the image combination, and the second field includes depth data of the preset frame image in the image combination. The virtual information image is obtained based on the position of the target object in the video frame of the multi-angle free-view video, obtained through 3D calibration, and the virtual information image matches the position of the target object. The playback control device is adapted to insert the synthesized video frame into the video stream to be played, wherein the virtual information image in the synthesized video frame changes synchronously with the target object in the video frame. The playback terminal is adapted to receive the video stream to be played from the playback control device and play it in real time.
24. The data processing system according to claim 23 further includes an interactive terminal; wherein: The server is also adapted to generate a stitched image corresponding to the image combination based on the pixel data and depth data of the image combination; and to store the stitched image of the image combination and the corresponding parameter data of the image combination. In response to the image reconstruction command from the interactive terminal, the system determines the interaction frame time information at the interaction time, obtains the spliced image of the preset frame images in the image combination at the corresponding interaction frame time and the corresponding parameter data of the image combination, and sends it to the interactive terminal. The interactive terminal is adapted to send the image reconstruction instruction to the server based on the interactive operation, and select the corresponding pixel data, depth data and corresponding parameter data in the stitched image according to the virtual viewpoint position information determined by the interactive operation and according to the preset rules. The selected pixel data and depth data are combined and rendered to reconstruct the video frame of the multi-angle free viewpoint video corresponding to the virtual viewpoint position at the time of the interactive frame and play it.
25. The data processing system according to claim 24, wherein the server is further adapted to generate and store a virtual information image corresponding to the spliced image of the preset video frame indicated by the server-side special effects generation interactive control instruction.
26. The data processing system according to claim 25, wherein the interactive terminal is adapted to synthesize the video frame of the multi-angle free-view video corresponding to the virtual viewpoint position at the time of the interactive frame with the virtual information image to obtain a synthesized video frame and display it.
27. A server, comprising: The data receiving unit is adapted to receive multiple synchronous video frames at a specified frame time extracted from multiple synchronous video streams as an image combination, wherein the multiple synchronous video frames contain frame images from different shooting angles. The image stitching generation unit is adapted to generate a stitched image corresponding to the image combination based on the pixel data and depth data of the image combination. The stitched image includes a first field and a second field, wherein the first field includes the pixel data of a preset frame image in the image combination, and the second field includes the depth data of the image combination. A parameter data calculation unit is adapted to determine the corresponding parameter data of the image combination; A depth data calculation unit is adapted to determine the depth data of each frame in the image combination; The video data acquisition unit is adapted to reconstruct the frame images of the preset virtual viewpoint path based on the parameter data of the image combination, the pixel data and depth data of the preset frame images in the image combination, and to obtain the video frames of the corresponding multi-angle free viewpoint video. The first virtual information image generation unit is adapted to respond to a special effects generation instruction, acquire a target object in a video frame specified by the special effects generation instruction, acquire augmented reality special effects input data of the target object, and generate a corresponding virtual information image based on the augmented reality special effects input data of the target object; wherein, the first virtual information image generation unit is adapted to take the augmented reality special effects input data of the target object as input, and based on the position of the target object in the video frame of the multi-angle free-view video obtained by three-dimensional calibration, generate a virtual information image in the corresponding video frame that matches the target object using a preset first special effects generation method; An image synthesis unit is adapted to synthesize the virtual information image with the specified video frame to obtain a synthesized video frame, wherein the virtual information image in the synthesized video frame and the target object in the video frame change synchronously. The first data transmission unit is adapted to output the synthesized video frame to be inserted into the video stream to be played.
28. A server, comprising: The image reconstruction unit is adapted to respond to an image reconstruction command from an interactive terminal, determine the interaction frame time information at the interaction time, acquire the stitched image of the preset frame images in the image combination at the corresponding interaction frame time and the corresponding parameter data of the image combination and send it to the interactive terminal, so that the interactive terminal selects the corresponding pixel data and depth data and the corresponding parameter data in the stitched image according to the virtual viewpoint position information determined by the interaction operation and according to the preset rules, combines and renders the selected pixel data and depth data, reconstructs the video frame of the multi-angle free viewpoint video corresponding to the virtual viewpoint position at the interaction frame time and plays it. A virtual information image generation unit is adapted to generate a virtual information image corresponding to a stitched image of a combination of images of video frames indicated by the special effects generation interactive control command in response to the special effects generation interactive control command, including: in response to the special effects generation interactive control command, acquiring a target object in the video frame indicated by the special effects generation interactive control command; and acquiring a virtual information image pre-generated based on augmented reality special effects input data of the target object; The data transmission unit, adapted to interact with an interactive terminal, includes: transmitting a stitched image of a preset video frame from the image combination at the corresponding interactive frame moment, along with corresponding parameter data of the image combination, to the interactive terminal. This allows the interactive terminal to select corresponding pixel data, depth data, and corresponding parameter data from the stitched image according to preset rules based on the virtual viewpoint position information determined by the interactive operation. The selected pixel data and depth data are then combined and rendered to reconstruct an image of the multi-angle free-view video corresponding to the virtual viewpoint position at the interactive frame moment, and the image is then played. Additionally, the unit transmits a virtual information image corresponding to the stitched image of the preset frame image indicated by the special effects generation interactive control command to the interactive terminal, allowing the interactive terminal to transmit the virtual information image of the virtual viewpoint position at the interactive frame moment. The corresponding video frames of the multi-angle free-view video are synthesized with the virtual information image to obtain a multi-angle free-view composite video frame, which is then played. The stitched image of the preset video frame is generated based on the pixel data and depth data of the image combination at the interaction frame time. The stitched image includes a first field and a second field. The first field includes the pixel data of the preset frame image in the image combination, and the second field includes the depth data of the image combination. The image combination at the interaction frame time is obtained by extracting multiple synchronous video frames at a specified frame time from multiple synchronous video streams. These multiple synchronous video frames contain frame images from different shooting angles. The virtual information image in the composite video frame changes synchronously with the target object in the video frames of the multi-angle free-view video.
29. An interactive terminal, comprising: The first display unit is suitable for displaying images of multi-angle free-view video in real time. The images of the multi-angle free-view video are reconstructed from parameter data of an image combination formed by multiple synchronous video frame images at a specified frame time, pixel data and depth data of the image combination. The pixel data and depth data of the image combination are stored using a stitched image at the corresponding frame time. The stitched image includes a first field and a second field. The first field includes pixel data of a preset frame image in the image combination, and the second field includes depth data of the preset frame image in the image combination. The multiple synchronous video frames include frame images from different shooting angles. The special effects data acquisition unit is adapted to respond to a triggering operation of a special effects display identifier in the multi-angle free-view video image, and acquire a virtual information image corresponding to a specified frame moment of the special effects display identifier, including: acquiring a virtual information image of a target object in a video frame at a specified frame moment corresponding to the special effects display identifier; The second display unit is adapted to overlay the virtual information image onto the video frame of the multi-angle free-view video, wherein the virtual information image changes synchronously with the target object in the video frame of the multi-angle free-view video.
30. An electronic device comprising a memory and a processor, the memory storing computer instructions executable on the processor, the processor performing the steps of the method according to any one of claims 1 to 21 when executing the computer instructions.
31. A computer storage medium, characterized in that, The storage medium records computer instructions that are executed by a processor to implement the method according to any one of claims 1 to 21.
Citation Information
Patent Citations
System and method for multi-angle real-time rebroadcasting of shot targets
CN103051830A
Three-dimensional environment information display method and device
CN108629830A
Augmented reality shooting method, apparatus, electronic device and storage medium
CN109089038A
Method and apparatus for providing 3-dimension image to head mount display
CN109361913A