Video processing method and apparatus, and device and storage medium
By constructing a visual shell and multi-viewpoint video information, the problem of insufficient immersive experience in existing free-viewpoint videos is solved, achieving better immersive playback and free-viewpoint rendering effects, and enhancing the audience's interactive experience.
Patent Information
- Application Number
- PCT/CN2025/087989
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-12
- Filing Date
- 2025-04-09
- Publication Date
- 2026-01-15
AI Technical Summary
Existing free-viewpoint video technology only has the ability to view 3D information during interaction, and its immersive experience is limited.
By acquiring the video frames to be processed and their acquisition attribute information, a visual shell is constructed and multi-view video information is determined. The video content is compressed, decompressed and rendered on the terminal, and combined with interactive operations to generate haptic feedback.
It achieves a better immersive playback experience, with free-viewpoint video rendering and realistic spatial perception effects, enhancing the audience's immersion and interactive experience.
Smart Images

Figure CN2025087989_15012026_PF_FP_ABST
Abstract
Description
Video processing methods, apparatus, equipment and storage media
[0001] This application claims priority to Chinese Patent Application No. 202410939490.5, filed on July 12, 2024, the disclosure of which is incorporated herein by reference in its entirety. Technical Field
[0002] This disclosure relates to a video processing method, apparatus, device, and storage medium. Background Technology
[0003] Currently, free-viewpoint video technology has emerged in video playback direction. This technology breaks through the traditional passive video viewing experience. Viewers can rotate the viewing angle of the video by swiping the screen, thereby enjoying an immersive experience in the whole scene.
[0004] However, existing free-viewpoint videos primarily achieve the adjustment of the viewing angle by shooting videos from discrete multiple perspectives to present the video from the adjusted perspective. For viewers, free-viewpoint videos based on existing technologies only offer the ability to view 3D information during interaction, and the immersive experience they can achieve is limited. Summary of the Invention
[0005] This disclosure provides a video processing method, apparatus, device, and storage medium that enables processed videos to have a better immersive playback experience.
[0006] In a first aspect, embodiments of this disclosure provide a video processing method, which is applied to a first terminal and includes:
[0007] Acquire the video frames to be processed and their acquisition attribute information. The video frames to be processed are the video frames in the received video stream to be processed. The acquisition attribute information includes acquisition based on monocular or binocular lenses, as well as acquisition based on multi-view lenses.
[0008] Based on the collected attribute information, construct a visual shell for the video to be processed and determine the visual shell information of the visual shell; and determine the multi-view video information of the video frame to be processed.
[0009] Video content compression is performed based on visual shell information and multi-view video information to generate video compression information for the video frames to be processed.
[0010] The video compression information is transmitted to the second terminal, so that the second terminal can perform video rendering based on the decompressed content of the video compression information to form displayable video frames.
[0011] Secondly, this disclosure also provides a video processing method, which is applied to a second terminal and includes:
[0012] The received video compression information is decompressed to obtain decoded multi-view video information and visual shell information. The video compression information is obtained from the video compression information stream sent by the first terminal. The video compression information is formed by the first terminal compressing the video content based on the visual shell information and multi-view video information of the video frame to be processed. The video frame to be processed is obtained from the video stream to be processed.
[0013] Based on the multi-view video information and the pose information of the second terminal, determine the video frame to be rendered that matches the orientation of the second terminal, render the video frame to be rendered to form a displayable video frame and display it.
[0014] When an interactive operation is received that acts on a displayable video frame, if the target of the interactive operation is determined to be a set target object in the displayable video frame based on the visual shell information, then the first haptic feedback corresponding to the interactive operation is generated.
[0015] Thirdly, embodiments of this disclosure also provide a video processing apparatus, which is configured in a first terminal and includes:
[0016] The video acquisition module is used to acquire the video frames to be processed and their acquisition attribute information. The video frames to be processed are the video frames in the received video stream. The acquisition attribute information includes acquisition based on monocular or binocular lenses, as well as acquisition based on multi-view lenses.
[0017] The information determination module is used to construct a visual shell of the video to be processed based on the collected attribute information, and to determine the visual shell information of the visual shell; and to determine the multi-view video information of the video frame to be processed.
[0018] The video compression module is used to compress video content based on visual shell information and multi-view video information, and generate video compression information for the video frames to be processed.
[0019] The information transmission module is used to transmit video compression information to the second terminal, so that the second terminal can perform video rendering based on the decompressed content of the video compression information to form displayable video frames.
[0020] Fourthly, embodiments of this disclosure also provide a video processing apparatus, which is configured in a second terminal and includes:
[0021] The video decoding module is used to decompress the received video compression information to obtain the decoded multi-view video information and visual shell information. The video compression information is obtained from the video compression information stream sent by the first terminal. The video compression information is formed by the first terminal compressing the video content based on the visual shell information and multi-view video information of the video frame to be processed. The video frame to be processed is obtained from the video stream to be processed.
[0022] The video rendering module is used to determine the video frame to be rendered that matches the orientation of the second terminal based on the multi-view video information and the pose information of the second terminal, render the video frame to be rendered to form a displayable video frame and display it.
[0023] The first feedback module is used to generate the first haptic feedback corresponding to the interactive operation when it receives an interactive operation on a displayable video frame and determines that the operation object of the interactive operation is a set target object in the displayable video frame based on the visual shell information.
[0024] Fifthly, embodiments of this disclosure also provide a computer device, the computer device comprising:
[0025] One or more processors;
[0026] Storage device for storing one or more programs.
[0027] When one or more programs are executed by one or more processors, the one or more processors implement the video processing method provided in any embodiment of this disclosure.
[0028] Sixthly, embodiments of this disclosure also provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the video processing method provided in any embodiment of this disclosure.
[0029] In a seventh aspect, embodiments of this disclosure also provide a computer program product, including a computer program that, when executed by a processor, implements the video processing method provided in any embodiment of this disclosure. Attached Figure Description
[0030] To more clearly illustrate the technical solutions of the exemplary embodiments of this disclosure, the accompanying drawings used in describing the embodiments are briefly introduced below. Obviously, the accompanying drawings described are only a portion of the embodiments to be described in this disclosure, and not all of them. For those skilled in the art, other drawings can be obtained from these drawings without any creative effort.
[0031] Figure 1a shows a schematic flowchart of a video processing method provided in an embodiment of this disclosure;
[0032] Figure 1b shows the effect of the video frame to be processed formed based on multi-lens capture in the video processing method provided in the embodiments of this disclosure;
[0033] Figure 1c shows the effect of rendering the generated multi-view video frame in the three-dimensional scene corresponding to the video frame to be processed in the video processing method provided in the embodiments of this disclosure.
[0034] Figure 1d shows the effect of rendering a multi-view video frame generated in a three-dimensional scene corresponding to another video frame to be processed in the video processing method provided in the embodiments of this disclosure.
[0035] Figure 1e shows the effect of compressing a group of video frames after compression encoding in the video processing method provided in the embodiment of this disclosure;
[0036] Figure 2 shows a schematic flowchart of a video processing method provided in another embodiment of this disclosure;
[0037] Figure 3a shows the effect of the displayable video frame formed by the video processing method provided by the present disclosure embodiment on the display terminal side when the video frame to be processed is a monocular video frame.
[0038] Figure 3b shows the effect of displaying a video frame formed by the video processing method provided in this embodiment of the present disclosure on the display terminal side when the video frame to be processed is a multi-lens captured video frame.
[0039] Figure 3c illustrates the effect of displaying a displayable video frame on the display terminal side using the video processing method provided in this embodiment of the present disclosure when the video frame to be processed is a multi-lens captured video frame.
[0040] Figure 4a shows the effect of haptic feedback when a click operation is performed on a displayable video frame formed by the video processing method provided in this embodiment of the present disclosure, and the haptic feedback conditions are met.
[0041] Figure 4b shows the effect of haptic feedback generated when the target object in the video frame formed by the video processing method provided in the embodiments of this disclosure meets the haptic feedback conditions;
[0042] Figure 5a shows a flowchart of one implementation of video processing using the video processing method provided in this embodiment of the present disclosure in a live streaming scenario;
[0043] Figure 5b shows another implementation flowchart of video processing using the video processing method provided in this embodiment of the present disclosure in a live streaming scenario;
[0044] Figure 6 shows a schematic diagram of the structure of a video processing apparatus provided in an embodiment of the present disclosure;
[0045] Figure 7 shows a schematic diagram of the structure of a video processing apparatus provided in another embodiment of the present disclosure;
[0046] Figure 8 shows a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Detailed Implementation
[0047] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0048] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0049] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0050] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules, or units, and are not used to limit the order of functions performed by these devices, modules, or units or their interdependencies. It should also be noted that the modifications of "a" and "a plurality of" mentioned in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0051] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0052] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0053] It is understood that before using the technical solutions disclosed in the various embodiments of this disclosure, users should be informed of the types, scope of use, and usage scenarios of the personal information involved in this disclosure in an appropriate manner in accordance with relevant laws and regulations, and user authorization should be obtained.
[0054] For example, upon receiving a user's active request, a prompt message is sent to the user to explicitly inform them that the requested operation will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the software or hardware, such as the electronic device, application, server, or storage medium performing the operations of this disclosed technical solution, based on the prompt message.
[0055] As an optional but non-limiting implementation, in response to a user's active request, sending a prompt message to the user can be done via a pop-up window, where the prompt message can be presented in text format. Furthermore, the pop-up window can also include a selection control allowing the user to choose "agree" or "disagree" to provide personal information to the electronic device.
[0056] It is understood that the above notification and user authorization process are merely illustrative and do not constitute a limitation on the implementation of this disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of this disclosure.
[0057] Figure 1a shows a flowchart of a video processing method provided in an embodiment of the present disclosure. This embodiment is applicable to the situation of video effects. The method can be executed by a video processing device, which can be implemented by software and / or hardware, and can be configured in a terminal and / or server to implement the video processing method in the embodiment of the present disclosure. In this embodiment, the terminal and / or server executing the video processing method can be referred to as the first terminal.
[0058] It should be noted that one application scenario of this embodiment can be a live streaming scenario. In existing live streaming scenarios, the broadcaster can use image acquisition devices with monocular / dual-lens or even multiple lenses to capture the live video stream. Generally, the live streaming link directly encodes the captured raw live video stream and transmits it to the viewer. After decoding, if the live video stream includes monocular / dual-lens video frames, only the existing content of the monocular / dual-lens video frames can be rendered, and free-viewpoint playback of the video cannot be achieved.
[0059] If the live video stream includes video frames from multiple perspectives captured by multiple cameras at the same time frame, it can only select the video frames to be rendered from the multiple perspective video frames. The video perspectives that can be rendered are limited, and a good 3D immersive viewing experience cannot be achieved.
[0060] Based on this, the video processing method provided in this embodiment can be executed through a first terminal to first construct a 3D scene and determine multi-view video from the video stream acquired by the image acquisition device. Specifically, as shown in Figure 1a, the video processing method provided in this embodiment may include:
[0061] S101. Obtain the video frame to be processed and the acquisition attribute information of the video frame to be processed. The video frame to be processed is a video frame in the received video stream to be processed. The acquisition attribute information includes acquisition based on monocular or binocular lenses, and acquisition based on multi-lens lenses.
[0062] It should be noted that this embodiment mainly proposes the implementation logic of a video processing method, without specifically limiting the execution subject. Different execution devices can be used as the execution subject to implement the logic execution of the video processing method, depending on the application scenario in which the video processing method is used. For example, this embodiment can set the application scenario as a live streaming scenario, and the first terminal as the execution subject of the video processing method provided in this embodiment can be the server in the live streaming link.
[0063] In a scenario where video playback is achieved through multi-terminal collaboration, the method provided in this embodiment allows the first terminal to act as a receiver of the acquired video stream to be processed, enabling the construction of a 3D scene and the generation of multi-view video frames from the acquired video stream. For example, in a live streaming scenario, the first terminal can act as a server in the live streaming link, establishing communication with the broadcaster's client to receive the video stream to be processed sent by the broadcaster's client. The video stream contains video frames to be processed. It can be assumed that the broadcaster's client is connected to an image acquisition device, or that an image acquisition device is installed on the broadcaster's client.
[0064] This step obtains the video frames to be processed and their acquisition attribute information from the video stream to be processed. The video stream to be processed generates different video frames depending on the acquisition attributes of the image acquisition device. For example, when the image acquisition device includes monocular / dual-lens lenses, the video frames to be processed are monocular / dual-lens video frames. In this case, the acquisition attribute information of the video frames to be processed can also be considered as being acquired based on monocular or dual-lens lenses. When the image acquisition device has multi-lens lenses set with different shooting angles, multiple video frames from different angles can be acquired at the same time frame, and all of these multiple video frames can be used as video frames to be processed at the same time frame and obtained through this step. In this case, the acquisition attribute information of the video frames to be processed can be that acquired by multi-lens lenses.
[0065] S102. Based on the acquired attribute information, construct a visual shell for the video to be processed and determine the visual shell information of the visual shell; and determine the multi-view video information of the video frame to be processed.
[0066] The video processing method provided in this embodiment can be considered as enabling the video frame to be processed to possess three-dimensional spatial information and to have video content from more different perspectives. That is, it is equivalent to constructing a three-dimensional scene for the video frame to be processed and determining the video content from multiple viewpoints. This processing provides the basic support for achieving free-viewpoint rendering of the video. It is understood that the display effects of a video that satisfies free-viewpoint rendering can be described as follows: its display perspective can be adjusted according to the position of the display device, thereby making the video display more spatial. Existing free-viewpoint rendering of videos requires that the video first possess three-dimensional spatial information and also scene information from different perspectives in three-dimensional space.
[0067] Therefore, if you want the video frame to be processed to be rendered into a displayable video frame with a sense of space and can be displayed from different perspectives during playback, you can first construct a three-dimensional scene for the video frame to be processed through this step. Specifically, you can construct a visual shell for the scene content contained in the video frame to be processed, and determine the video content corresponding to different perspectives in the three-dimensional scene of the video frame to be processed. Specifically, you can determine the multi-view video information of the video frame to be processed.
[0068] In this embodiment, the visual shell can be understood as a 3D contour formed by reconstructing the shape of a geometric entity in a 3D scene, similar to adding a contour shell to the geometric entity. The visual shell can be represented by triangular facets, and the visual shell information can be formed by the triangular facets and their corresponding vertex information, texture information, and material information. Multi-view video information can be understood as the frame information formed by rendering the same object content in the captured video from different perspectives. It can be a direct summary of the multi-view frame information, or a record of the information required to form the multi-view frame information. Here, multi-viewpoints can correspond to multiple capture perspectives of a virtual camera in a 3D scene.
[0069] In this embodiment, different methods can be used to construct the visual shell and determine the multi-view video information based on the different acquisition attribute information corresponding to the video frame to be processed.
[0070] Specifically, when the video frame to be processed acquired by the image acquisition device is a monocular / binocular video frame, that is, when the corresponding acquisition attribute information is acquired based on monocular or binocular lenses, the spatial depth information of the video frame to be processed can be obtained by performing depth estimation and foreground / background segmentation. Then, the foreground mask frame and the background mask frame are determined by the spatial depth information. The visual shell of the video frame to be processed in three-dimensional space can be constructed based on the foreground mask frame and the background mask frame. Furthermore, the triangular facet information, texture information, and material information that characterize the visual shell can be determined as the visual shell information of the visual shell.
[0071] Meanwhile, the determination of multi-view video content can also be achieved by filling the background content of the background mask frame with the background filled image. For example, the background area can be expanded and the content filled by the background filled image, thereby segmenting the background image area corresponding to multiple different viewpoints from the expanded background image. The background image area of different viewpoints and the foreground mask frame can constitute multiple video frames of different viewpoints of the video frame to be processed, and finally multi-view video information containing multiple video frames of different viewpoints can be formed.
[0072] Similarly, as another implementation description, when the video frame to be processed acquired by the image acquisition device is multiple video frames from different perspectives, that is, when the corresponding acquisition attribute information is based on multi-view lens acquisition, it is equivalent to the obtained video frame to be processed itself having three-dimensional characteristics. However, considering that the number of video frames captured from different perspectives is the same as the number of multi-view lenses used, directly rendering the video frame to be processed cannot truly and naturally achieve free-view rendering of the video. Therefore, this embodiment can reconstruct a three-dimensional scene from multiple video frames at the same time frame that are the video frames to be processed through this step. It can also construct a visual shell for the scene content in the reconstructed three-dimensional scene, and render more video frames from different perspectives in the reconstructed three-dimensional scene. The visual shell information based on the constructed visual shell can better ensure the authenticity and integrity of the rendered video frames, thereby obtaining a larger number of multi-view video frames, and the obtained multi-view video frames can be determined as the multi-view video information of the video frame to be processed.
[0073] Following the above description, for video frames with multiple different perspectives as the video frames to be processed, the method of determining three-dimensional point cloud data for the three-dimensional scene can also be used to form multi-view video information of the video frames to be processed. For example, the three-dimensional point cloud data used to form multi-view video frames can be directly used as multi-view video information.
[0074] S103. Based on the visible shell information and the multi-viewpoint video information, compress the video content to generate the video compression information of the video frame to be processed.
[0075] It's important to understand that the terminal acting as the display device may not be the same as the first terminal; for example, the display device could be a second terminal. In this case, it's necessary to transmit the visual shell information and multi-view video information used for video rendering, specifically for the video frame to be processed, to the second terminal. This allows the second terminal to render the video frame under free-view conditions using the visual shell information and multi-view video information. Furthermore, it's known that the transmission of video-related data requires encoding before transmission; video content compression can be considered one form of encoding.
[0076] In this embodiment, the first terminal and the second terminal are used as examples of different terminals. For example, in a live broadcast scenario, the second terminal can be the viewing terminal on the audience side. Therefore, before transmitting the visual shell information and multi-view video information to the second terminal, this step can be used to compress the video content of the visual shell information and multi-view video information so as to achieve the encoding processing of video-related data information through compression.
[0077] It is understandable that the video compression method will differ depending on the method used to determine the visual shell information and multi-view video information. In one implementation, if the video frame to be processed corresponding to the determined visual shell information and multi-view video information is a monocular / binocular video frame, then the corresponding video content compression method can be described as follows: video content compression can be achieved by stitching together the foreground mask image frame and the multi-view background image frame contained in the multi-view video information with the video frame to be processed to form a stitched video frame. Alternatively, compression of the visual shell information can be achieved by extracting only the mesh data (such as triangle markers and vertex information) from the visual shell information while omitting texture information. Finally, the stitched video frame and mesh data can be used as the video compression information for the video frame to be processed.
[0078] In another implementation, if the determined visual shell information and multi-view video information correspond to video frames to be processed from different perspectives, the visual shell information can also be compressed into grid data. However, for multi-view video information, when the multi-view video information consists of multiple video frames from different perspectives, the multi-view video frames can be stitched and compressed in groups to form stitched video frames with the same number of groups. Alternatively, when the multi-view video information is 3D point cloud data, a specific 3D point cloud data compression strategy can be used for compression. Finally, the resulting network data and multiple stitched video frames or point cloud compressed data can be used as the video compression information of the video frames to be processed.
[0079] S104. The video compression information is transmitted to the second terminal, so that the second terminal performs video rendering based on the decompressed content of the video compression information to form a displayable video frame.
[0080] In this embodiment, this step transmits the video compression information to the second terminal in the form of a bitstream. The second terminal can be considered a device capable of playing video.
[0081] In this embodiment, regardless of whether the first terminal and the second terminal are different terminals, considering the desire for the video frame to be processed to ultimately render and display a spatial visual effect from a free-viewpoint perspective, the video compression information transmitted to the second terminal (or the video compression information transmitted to the first terminal for use in the video rendering processing module, in which case the first terminal can also be considered the second terminal) needs to be rendered again by the second terminal based on the video compression information from a free-viewpoint perspective, thereby forming a displayable video frame with a spatial visual effect that can be played on the second terminal. Here, a displayable video frame can be considered as a video frame that corresponds to a viewing angle adapted to the viewing posture and orientation of the second terminal.
[0082] The technical solution described in this embodiment, regardless of whether the video frame to be processed is a monocular / binocular video frame or a multi-view video frame acquired in a multi-view format, can achieve the conversion of the video frame to three-dimensional space and the determination of the corresponding multi-view video frame by adding a visual shell and performing secondary processing to determine the multi-view video. Based on the constructed visual shell and multi-view video information, in a video playback scenario, the video frame to be processed can achieve free-view video rendering according to different viewing angles, thus providing a basic information guarantee for effectively expanding the application scope of free-view video viewing. This technical solution ensures that the video to be processed can have a more realistic free-view viewing effect when rendered.
[0083] Based on the above description, to better ensure an immersive viewing experience for viewers when watching displayable video frames, further haptic feedback can be generated after determining that the displayable video frame meets the haptic feedback conditions, in addition to rendering the displayable video frame with a perspective matching the pose of the second terminal. In this embodiment, the determination of whether the haptic feedback conditions are met can be achieved through the determined visual shell information. For example, scenarios where haptic feedback may occur include: clicking (touching), translating (dragging), or even rotating an object in the scene involved in the displayable video frame. Thus, when the second terminal receives interactive operations such as clicking, translating (dragging), and rotating, it can perform collision detection through the visual shell information to determine whether these operations have acted on the target object. If they have acted on the target object, the haptic feedback conditions can be considered met, and corresponding vibration or sound effects can be presented as haptic feedback.
[0084] Therefore, it can be seen that this technical solution not only ensures a more realistic and free viewing experience when the video to be processed is rendered, but also provides basic information to guarantee the haptic feedback of the rendered video frames. Thus, based on the visual shell information and multi-viewpoint video information provided by this technical solution, it can achieve a better perception of the physical force feedback of people or other entities in the video scene during the video rendering stage.
[0085] As a first optional embodiment of this example, based on the above embodiment, the steps of constructing a visual shell of the video to be processed according to the acquired attribute information and determining the visual shell information of the visual shell, and determining the multi-view video information of the video frame to be processed, can be specified as follows:
[0086] a1) When the attribute information is acquired based on monocular or binocular lens acquisition, determine whether the acquired video frame to be processed is a monocular video frame or a binocular video frame, and perform depth estimation and foreground processing on the video frame to be processed to obtain spatial depth information, foreground masking frame and foreground missing frame.
[0087] It is understood that the step description provided in this first optional embodiment mainly addresses the determination of visual shell information and multi-viewpoint video information when the acquisition attribute information of the video frame to be processed is based on monocular or binocular lens acquisition.
[0088] Therefore, in this optional embodiment, the acquired video frame to be processed is a monocular video frame or a binocular video frame, which is equivalent to a two-dimensional video frame. To achieve a free-viewpoint playback effect in three-dimensional space for the two-dimensional video frame, this optional embodiment requires three-dimensional conversion processing of the video frame to be processed.
[0089] For example, this step can be used to first perform depth estimation and foreground matting on the video frame to be processed, and obtain the depth map frame, the foreground mask map frame after separating the foreground and background, and the foreground missing map frame after foreground matting of the video frame to be processed. This adds spatial depth information to the video frame to be processed, and the added spatial depth information can better determine more scene information for the video frame to be processed.
[0090] For example, a depth map frame records the corresponding depth value for each pixel in the video frame to be processed. A trained depth estimation network model can be used to estimate the depth of the video frame to be processed, thereby obtaining a depth map frame with the same size as the video frame to be processed, using depth values as pixel information. Similarly, a trained background segmentation model can be used to segment the foreground and background of the video frame to be processed by foreground matting.
[0091] b1) Based on the spatial depth information and the missing foreground frame, determine the background mask frame of the video frame to be processed, and construct the first visual shell of the video frame to be processed based on the foreground mask frame and the background mask frame. The triangular facet information, texture information and material information representing the first visual shell are determined as the first visual shell information.
[0092] In this embodiment, to achieve the goal of rendering a free-view video with a sense of space in the final video frame to be processed, the missing foreground image frame can be filled with information by combining the obtained spatial depth information with the pixel value information to be processed of the video frame. In this way, the background mask image frame of the video frame to be processed can also be obtained. In this embodiment, the above steps are equivalent to obtaining the foreground mask image frame and the background mask image frame of the video frame to be processed.
[0093] In this context, the foreground mask frame can be considered as a frame containing only the foreground object content. The background mask frame, on the other hand, can be considered as a background frame with a depth effect formed by filling in the missing foreground content with spatial depth information.
[0094] In this embodiment, one specific implementation of determining the visual shell can be described as follows: Binarizing the foreground mask frame and the background mask frame yields the processed foreground contour image and background contour image. These foreground and background contour images can be considered as projections of image content from a 3D scene onto a 2D scene. Then, using the capture parameters of a given virtual camera at different viewpoints, contour images (such as foreground and background contour images) are determined from different viewpoints. Based on these contour images, different contour cones can be formed for the foreground object and the background object. Finally, the contour corresponding to the intersection of the contour cones from different viewpoints can be called the visual shell of the corresponding object (foreground object or background object). In this embodiment, the determined visual shell can be denoted as the first visual shell, and the triangular facet information, texture information, and material information representing the first visual shell can be determined as the first visual shell information of the visual shell.
[0095] c1) Extract image regions from the background masking frame from different viewpoints, determine the extracted image regions as background frames from different viewpoints, and construct the first multi-viewpoint video information of the video frame to be processed based on the background frame and the foreground masking frame.
[0096] In this embodiment, the background mask frame can be considered as a frame with a sense of space formed by expanding the original background area and filling it with background content. With the screen size of the display device remaining unchanged, this is equivalent to adding an extra displayable area. The background mask frame, including both the expanded area and the original background area, can simulate image capture from different viewpoints by cropping the image area.
[0097] For example, the captured image region can be a region with the same size as the video frame to be processed. In the implementation of image region cropping, it can start from the upper left corner of the background mask frame, and crop from left to right and from top to bottom according to a set cropping step size. The number of image regions formed by cropping can be determined by the set cropping step size. If the cropping step size value is small, the number of image regions formed will be larger. Correspondingly, if the cropping step size value is large, the number of image regions formed will be smaller. In this embodiment, each image region formed by cropping can be regarded as a background frame corresponding to a viewpoint. A viewpoint can be regarded as a capture angle of a virtual camera, and the viewpoint information of each viewpoint can be determined by the cropping angle when cropping the background mask frame.
[0098] In this embodiment, after determining the background frames from different viewpoints, the foreground mask frame and multiple background frames can be regarded as multi-view video information formed by the video frames to be processed. In this embodiment, the multi-view video information is recorded as the first multi-view video information.
[0099] The above-described technical solution of this first optional embodiment provides a method for determining the visual shell information and multi-view video information when the video frame to be processed is a monocular / binocular video frame. Under the premise that the processing objective is to convert the video frame to be processed into a video frame that can be displayed from a free-viewpoint perspective, the visual shell information and multi-view video information implemented relative to monocular / binocular video frames in this first optional embodiment provide basic information guarantees for the free-viewpoint playback of the video frame.
[0100] As a second optional embodiment of this embodiment, based on the first optional embodiment, the video compression information for generating the video frame to be processed by compressing video content based on the visual shell information and the multi-viewpoint video information can be specified as follows:
[0101] a2) Combine the foreground mask frame and the multi-view background frame contained in the first multi-view video information with the video frame to be processed to perform image stitching to form the first stitched video frame.
[0102] It is understandable that, depending on the formation method of the visual shell information and the multi-view video information, the compression method used when compressing and encoding the visual shell information and the multi-view video information will also be different. This second optional embodiment provides the compression method corresponding to the determination of visual shell information and multi-view video information using the first optional embodiment described above.
[0103] It is understood that the first multi-view video information formed through the above method specifically consists of related image information of a foreground mask frame and a multi-view background frame. For example, in this step of the second optional embodiment, the foreground mask frame and the multi-view background frame constituting the first multi-view video information can be combined with the video frame to be processed for image stitching, thereby forming a first stitched video frame. The foreground mask frame mainly includes a grayscale image representing the outline of the foreground object. The purpose of also including the video frame to be processed in the stitching is to restore the pixel information of the foreground object in the foreground frame.
[0104] b2) The first spliced video frame is video encoded and compressed to form an encoded video frame, and the vertex coordinates of the triangular facet information in the first visible shell information are extracted to form the first grid data.
[0105] In this embodiment, the first spliced video frame can be encoded and compressed using a common video encoding method, thereby forming an encoded video frame. This step can also extract mesh data, which includes only the vertex coordinates of the triangular facets, from the determined first visible shell information; in this embodiment, this is denoted as the first mesh data.
[0106] c2) Video compression information of the video frame to be processed is constructed based on the encoded video frame and the first grid data.
[0107] In this embodiment, the encoded video frame and the first network data can be regarded as video compression information of the video frame to be processed through this step, and used as part of the transmittable bitstream for transmission to the second terminal, so that the second terminal can perform free-viewpoint video rendering based on the video compression information.
[0108] The technical solution described in this second optional embodiment provides an encoding and compression implementation of the visual shell information and multi-view video information formed in the first optional embodiment. This compression implementation provides the prerequisite for the transmission of the visual shell information and multi-view video information to other terminals, ensuring the effective transmission of the visual shell information and multi-view video information to other terminals.
[0109] As a third optional embodiment of this example, another implementation method is provided. Specifically, this third optional embodiment can construct a visual shell of the video to be processed based on the acquired attribute information, and determine the visual shell information of the visual shell; and determine the multi-view video information of the video frame to be processed, specifically as follows:
[0110] a3) When the acquisition attribute information is based on multi-camera acquisition, determine that the acquired video frames to be processed include video frames captured by different cameras at the same time frame.
[0111] It is understood that the steps described in this third optional embodiment are mainly for determining the visible shell information and multi-viewpoint video information when the image acquisition device's acquisition attribute information is based on multi-lens acquisition.
[0112] Therefore, in this optional embodiment, the acquired video frame to be processed is equivalent to video frames captured from different perspectives by multiple lenses at the same time frame. In this case, the video frame to be processed can be considered a relatively coarse 3D video frame. To achieve the effect of playing video content from a more refined 3D perspective, this optional embodiment requires performing multi-view video rendering processing on the video frame to be processed, which involves a richer number of perspectives.
[0113] This step can obtain video frames from different perspectives as video frames to be processed. Different video frames can be considered as video frames captured at the same time frame by different lenses in a multi-view camera at corresponding capture times.
[0114] For example, Figure 1b shows the effect of a video frame to be processed formed based on multi-lens capture in the video processing method provided in this embodiment of the present disclosure. As shown in Figure 1b, it can be assumed that 12 lenses are used for image capture on the image acquisition device side. Thus, for the same scene object 110, it can be captured by 12 lenses in the same time frame, and video frames 120 can be formed after different lenses capture the scene object from different capture angles.
[0115] b3) Based on multiple video frames to be processed, construct a three-dimensional scene of the video frames to be processed, and based on the constructed three-dimensional scene, construct a second visual shell of the video frames to be processed, and determine the triangular facet information, texture information and material information representing the second visual shell as the second visual shell information.
[0116] In this embodiment, multiple video frames captured from different perspectives by different lenses can be used to reconstruct a 3D scene of the video frame to be processed. In the constructed 3D scene, a visual shell can be reconstructed for the entity objects included in the video frame to be processed. In this embodiment, the constructed visual shell can be referred to as the second visual shell, and the triangular facet information, texture information, and material information representing the second visual shell can be determined as the second visual shell information of the second visual shell.
[0117] Specifically, for the construction of the 3D scene corresponding to the video frame to be processed, the 3D scene can be reconstructed by combining the relative spatial position information of the lens used in the video frame capture with a motion structure reconstruction algorithm, thereby obtaining a 3D scene that simulates the real spatial scene where the video frame to be processed is located. Simultaneously, after determining the 3D scene of the video frame to be processed, the entity objects contained in the video frame can be identified from the 3D scene. A visual shell can be constructed for the entity objects by determining the shape of contour cones, thereby obtaining the triangular facet information (such as the vertex coordinates of the triangular facets), texture information, and material information of the constructed visual shell as the visual shell information of the constructed visual shell.
[0118] For example, constructing a 3D scene of the video frames to be processed based on multiple video frames, and constructing a second visual shell of the video frames to be processed based on the constructed 3D scene, may include:
[0119] b31) Determine the relative pose information between pairs of video frames based on the video frame association information between multiple video frames.
[0120] In this embodiment, for multiple video frames to be processed, a 3D scene needs to be constructed first. This requires determining the video content association information between the video frames, which can be referred to as video frame association information. One way to implement video frame association information is as follows: feature extraction is performed on each video frame, and feature matching of video content in different video frames is performed based on the extracted features. Then, the matched feature pairs can be verified based on geometric features. Finally, a search tree of video frames can be built using the verified feature pairs. The search tree can contain the association information between video frames, and the relative pose information between any two video frames in 3D space can be determined based on the search tree.
[0121] b32) Based on the relative pose information, the motion structure reconstruction algorithm is used to determine the three-dimensional coordinate information for scene construction and the absolute pose information of the video frame.
[0122] In this embodiment, based on the relative pose information of the pairs of video frames determined above, combined with a pre-set motion result reconstruction algorithm for 3D scene construction, the 3D coordinate information of key points in the video frame in 3D space can be determined, and at the same time, the absolute pose information of the lens used for video frame capture in 3D space can be determined.
[0123] b33) Based on the three-dimensional coordinate information and the absolute pose information of the video frame, combined with the three-dimensional scene reconstruction model, construct the three-dimensional scene of the video frame to be processed and the visual shell of the entity objects in the three-dimensional scene, and record the visual shell as the second visual shell of the video frame to be processed.
[0124] In this embodiment, the three-dimensional coordinate information and the absolute pose information of the video frame determined in the above steps can be used as the input information of the given three-dimensional scene reconstruction model. Finally, the three-dimensional scene of the real physical space where the video frame to be processed is located can be output through the three-dimensional scene reconstruction module.
[0125] In this embodiment, in the constructed 3D scene, this step can construct a visual shell for all or some of the entity objects contained in the 3D scene. The entity objects contained in the video frame to be processed also exist in the 3D scene. Therefore, this step is equivalent to constructing a visual shell for the entity objects in the video frame to be processed.
[0126] In one implementation of constructing a visual shell, the silhouette contour information of the entity object can be obtained when it is captured from multiple perspectives. The silhouette contours formed by the entity object under all perspectives are intersected, and the silhouette contour at the intersection position can be determined as the visual shell of the entity object. In this embodiment, the visual shell of the entity object can be drawn, and the triangular facet information, texture information and material information of the drawn visual shell can be obtained as the visual shell information of the visual shell.
[0127] In this embodiment, the entity objects identified in the relative three-dimensional scene can be regarded as the second visual shell of the video frame to be processed, and the corresponding second visual shell information can be obtained.
[0128] The above-described technical solution in this optional embodiment provides another way to determine the visible shell information and multi-viewpoint video information, which also provides basic information assurance for video frames to be played from a free perspective.
[0129] c3) Using the selected multi-view rendering strategy combined with the second visual shell information, multi-view content rendering is performed on the 3D scene to obtain the second multi-view video information of the video frame to be processed.
[0130] In this embodiment, the visual shell information describes the outer contour of an entity object in a 3D scene. This embodiment incorporates the determined visual shell information into a multi-view rendering strategy to render multi-view content, thereby obtaining more realistic new-view video frames or 3D point cloud information that can be used to construct more realistic new-view videos. The multi-view rendering strategy can be neural radiation rendering using a neural radiation field model, or it can use a 3D Gaussian splash model to generate 3D point cloud data and directly participate in the video rendering on the playback side as multi-view video information.
[0131] This embodiment can directly construct multi-view video information of the video frame to be processed based on multiple new perspective video frames obtained. The number of new perspective video frames obtained is much greater than the number of lenses used in the video frame acquisition stage. For example, the number of multiple lenses involved in the acquisition of the video frame to be processed may not be greater than 10, but the number of new perspective video frames formed by rendering can be at least 72 frames, which is equivalent to a more refined panoramic rendering of the three-dimensional scene in 360 degrees.
[0132] One implementation method for using a selected multi-view rendering strategy combined with second visual shell information to render multi-view content of a 3D scene and obtain the second multi-view video information of the video frame to be processed can be described as follows:
[0133] c31) Input the second visible shell information, the three-dimensional coordinate information of the object in the three-dimensional scene, the ray angle information, the color information and the density information into the given target neural radiation field model.
[0134] In this embodiment, a target neural radiation field model is specifically used to render video frames from new perspectives in a 3D scene. Specifically, the second visible shell information, the 3D coordinate information of the object, the ray angle information, the color information, and the density information determined in the 3D scene can be directly used as the input information for the target neural radiation field, and input into the target neural radiation field model in this step.
[0135] For example, the target neural radiation field model can be obtained through pre-training. This target neural radiation field model can be regarded as a multilayer perceptron network model, which can have the network structure of the multilayer perceptron network model. The historical network parameters of the multilayer perceptron network model can be reused in the initial neural radiation field model. Through iterative training of the initial neural radiation field model, network parameters suitable for new perspective video rendering in this embodiment can be obtained.
[0136] It should be noted that, in the training phase of the initial neural radiation field model to train the target neural radiation field model, compared with the existing training methods, this embodiment adds sampling visual shell information of scene objects under the sampling perspective to each set of sample data in the sample training set used for model training. At the same time, the sample data also includes spatial position information, angle information, color information and density information of the sampling perspective.
[0137] As described above, by adding the sampled visual shell information, the training of the neural radiation field model can be supervised, so that the target neural radiation field model formed after training has a higher precision rendering effect, and can also ensure that the target neural radiation field model formed after training can generate highly realistic new perspective videos more quickly.
[0138] c32) The multiple video frames with different viewpoint information output by the target neural radiation field model are used as multi-view video frames to be processed, and the second multi-view video information is constructed based on the multi-view video frames. The video frames are obtained by rendering the three-dimensional scene under different viewpoints.
[0139] In this embodiment, each new perspective video frame output by the target neural radiation field model is equivalent to outputting multiple video frames containing different viewpoint information. These video frames can be used as re-rendered multi-view video frames with more refined viewpoint division. This step can directly regard the obtained multi-view video frames as the second multi-view video information of the video frames to be processed.
[0140] For example, Figure 1c shows an effect demonstration diagram of multi-view video frames generated by rendering a video frame to be processed in a three-dimensional scene in the video processing method provided in the embodiments of this disclosure. As shown in Figure 1c, the four sub-figures shown therein can be regarded as four new perspective video frames in the multi-view video frames formed by rendering scene objects in the video frame to be processed from different perspectives in the three-dimensional scene after the video frame to be processed shown in Figure 1b is formed into a three-dimensional scene.
[0141] Similarly, Figure 1d shows the effect of rendering a multi-view video frame generated in a 3D scene corresponding to another video frame to be processed in the video processing method provided in this embodiment of the present disclosure. Figure 1d also includes four sub-figures, each of which is equivalent to a new perspective video frame formed by rendering the entity object 11 in the 3D scene from different perspectives. It can be seen that the entity objects contained in each sub-figure of Figure 1d present different display effects from different perspectives.
[0142] The above steps of this optional embodiment provide a specific implementation method for the second multi-view video information. The second multi-view video information can be regarded as being composed of a larger number of rendered video frames. Through this video frame rendering method, video frames from different perspectives with a stronger sense of realism can be rendered.
[0143] Furthermore, another approach to obtaining the second multi-view video information of the video frame to be processed by using a selected multi-view rendering strategy combined with the second visual shell information can be described as follows:
[0144] c33) takes the three-dimensional coordinate information of objects in the three-dimensional scene as input information and inputs it into the given target three-dimensional Gaussian splash model.
[0145] It is understood that this optional embodiment also provides another implementation of the second multi-view video information, specifically by using the three-dimensional coordinate information of objects in the three-dimensional scene as input information through a trained three-dimensional Gaussian splash model.
[0146] c34) Obtain the 3D point cloud data output by the 3D Gaussian splash model, and determine the 3D point cloud data as the second multi-view video information of the video frame to be processed.
[0147] In this embodiment, the three-dimensional point cloud data output by the three-dimensional Gaussian splash model can be directly regarded as the second multi-view video information of the video frame to be processed.
[0148] It is known that the 3D Gaussian splash model does not render video frames from different new perspectives based on the object's 3D coordinate information. Instead, it directly obtains denser 3D point cloud data by Gaussian processing of the more critical 3D coordinate information. The 3D point cloud data can be represented as a point cloud disk to achieve a more detailed description of the 3D scene.
[0149] As a fourth optional embodiment of this example, based on the third optional embodiment described above, the video compression information for generating the video frame to be processed by compressing video content based on the visual shell information and the multi-viewpoint video information can be specified as follows:
[0150] a4) Extract the vertex coordinates of the triangular facets from the second visible shell information to form the second mesh data.
[0151] Understandably, given the above-mentioned approach of constructing a visual shell based on the side contour of the entity object, determining the visual shell information, and determining multi-view video information based on the target neural radiation field model or three-dimensional Gaussian splash model, corresponding compression methods will also be used to compress and encode the visual shell information and multi-view video information.
[0152] In this embodiment, the second visible shell information determined by this third optional embodiment can also be used to extract only the vertex coordinates in the triangular facet information to form mesh data, which is used as the second mesh data formed by the compression of the visible shell information in this embodiment.
[0153] b4) When the second multi-view video information consists of video frames from multiple different viewpoints, the video frames from different viewpoints are grouped to obtain at least one group of grouped video frames, and the grouped video frames are compressed and encoded according to the aggregation coding strategy to form corresponding compressed video frames.
[0154] In this embodiment, when multiple video frames from different viewpoints are generated using the aforementioned target neural radiation field model as the second multi-view video information, considering the large number of newly obtained video frames, the video frames from different viewpoints can be grouped and compressed. Thus, after grouping the video frames from different viewpoints, at least one set of video frames can be obtained, and the video frames in this set can preferably be denoted as grouped video frames.
[0155] For example, when the target neural radiation field model outputs 72 video frames from different viewpoints, these frames can be divided into groups of four, resulting in 18 groups of video frames. These groups can be formed by selecting four video frames sequentially according to a certain rotational order, such as clockwise or counterclockwise.
[0156] In this embodiment, for each group of video frames, compression encoding can be performed according to an aggregation coding strategy, thereby generating corresponding compressed video frames. The implementation process of compression encoding using the aggregation coding strategy can be described as follows:
[0157] For video frames from different perspectives within a group, effective region detection can be performed first (e.g., detecting only the contained entity objects), and the effective region can be cropped using image cropping methods. Then, the effective region corresponding to each video frame can be extracted using different attributes, such as extracting the corresponding texture frames, depth frames, and transparency frames. The physical range of the depth frames can be quantized a second time to form a broader depth frame, and the image content in the transparency frames can be processed to generate transparency frames containing brightness and chromaticity information of the image content. Finally, by stitching together the texture frames, depth frames, and transparency frames of each group of video frames included in a group, and encoding the stitched frame, a compressed video frame corresponding to that group is formed.
[0158] For example, Figure 1e shows the effect of compressed video frames formed after compressing and encoding a group of grouped video frames in the video processing method provided in the embodiments of this disclosure. As shown in Figure 1e, specifically, the four figures in Figure 1d can be regarded as a group of grouped video frames after grouping multi-viewpoint video frames, and the compressed video frames shown can be considered as formed after compressing the group of grouped video frames according to the aggregation encoding method.
[0159] c4) When the second multi-view video information is three-dimensional point cloud data, it is compressed and encoded according to the set point cloud data compression strategy to form the corresponding point cloud compressed information.
[0160] In this embodiment, when generating 3D point cloud data as the second multi-view video information using the aforementioned 3D Gaussian splash model, a compression strategy for compressing and encoding point cloud data can be directly applied to compress the 3D movie data, thereby obtaining the corresponding point cloud compression information. The point cloud data compression strategy that can be used can be a video-based point cloud compression strategy, or a geometry-based point cloud compression strategy, etc.
[0161] d4) Based on the compressed data of the second grid data and the compressed video frame or point cloud compression information, the video compression information of the video frame to be processed is constituted.
[0162] In this embodiment, the compressed video frame or point cloud compression information can be combined with the second grid data in this step as the video compression information of the video frame to be processed. It can also be used as part of the transmittable bitstream for transmission to the second terminal, so that the second terminal can perform free-view video rendering based on the video compression information.
[0163] The above-described technical solution of this fourth optional embodiment provides an encoding and compression implementation of the visual shell information and multi-view video information formed in contrast to the third optional embodiment. This compression implementation provides the prerequisite for the transmission of the visual shell information and multi-view video information to other terminals, ensuring the effective transmission of the visual shell information and multi-view video information to other terminals.
[0164] In this specific embodiment, Figure 2 shows a flowchart of another video processing method provided by this disclosure embodiment. The method provided by this embodiment is also applicable to video playback. The method can be executed by a video processing device, which can be implemented by software and / or hardware, and can be configured in a terminal and / or server to implement the video processing method in this disclosure embodiment.
[0165] It is understood that the terminal or server executing the video processing method provided in this embodiment can preferably be referred to as the second terminal. The video processing method executed by the second terminal in this embodiment can be considered as a method for video rendering processing based on the video compression information transmitted by the first terminal. The video rendering achieved by the video processing method provided in this embodiment can determine a video with a free-viewpoint playback effect for the acquired video to be processed.
[0166] Specifically, as shown in Figure 2, a video processing method provided in this embodiment may include:
[0167] S201. Decompress the received video compression information to obtain decoded multi-view video information and visual shell information. The video compression information is obtained from the video compression information stream sent by the first terminal. The video compression information is formed by the first terminal compressing video content based on the visual shell information and multi-view video information of the video frame to be processed. The video frame to be processed is obtained from the video stream to be processed.
[0168] In this embodiment, video rendering can be considered to be performed frame by frame. Therefore, the video processing method provided in this embodiment realizes the rendering of the video frame corresponding to a time frame from the perspective of time frame. First, this step can be used to decompress the video compression information received in the time frame. Specifically, the video compression information can be decompressed by the decompression method corresponding to the compression method used for the visual shell information and multi-view video information, thereby obtaining the multi-view video information and visual shell information used for video rendering in the time frame.
[0169] It should be noted that the video compression information in this embodiment can be considered as being determined by the first terminal and transmitted by the first terminal in the form of a bitstream to the second terminal, which is the execution subject of the method provided in this embodiment. The bitstream transmitted can be a video compression information stream, and the video compression information can be considered as information obtained from the video compression information stream at the time frame level.
[0170] In this embodiment, the video compression information determined by the first terminal corresponds to a video frame to be processed in the video stream to be processed. It can form visual shell information and multi-view video information in relation to the video frame to be processed by the video processing method given in the above embodiment.
[0171] The visual shell information can be considered as the visual shell constructed from the entity objects in the video frame to be processed within a 3D scene. The multi-view video information can be considered as a collection of video frames corresponding to the video frame to be processed from different viewpoints. Specifically, video frames from different viewpoints can be a combination of a background image frame (formed by cropping image regions from different viewpoints) and a foreground image frame of the video frame to be processed; additionally, video frames from different viewpoints can be video frames formed by rendering the video frame to be processed within a 3D scene from different viewpoints; or they can be 3D point cloud data determined by 3D Gaussian splashing from a 3D scene constructed based on the video frame to be processed.
[0172] S202. Based on the multi-view video information and the pose information of the second terminal, determine the video frame to be rendered that matches the orientation of the second terminal, render the video frame to be rendered to form a displayable video frame and display it.
[0173] It is understood that free-viewpoint rendering of video is mainly affected by the pose of the display device; that is, the rendering perspective during video rendering will adjust as the pose of the display device is adjusted. The second terminal executing the method provided in this embodiment can serve as the display device for video playback.
[0174] For the second terminal, which is the execution subject in this embodiment, to render the video from different perspectives, it is necessary to first obtain the pose information of the second terminal. When performing video rendering in units of time frames, it is necessary to obtain the pose information of the second terminal in each time frame, as well as the orientation information of the second terminal.
[0175] In this embodiment, this step can determine the video frame to be rendered that matches the orientation of the second terminal based on the multi-view video information obtained above and the pose information of the second terminal obtained, and form a displayable video frame by rendering the video frame to be rendered.
[0176] It is known that the rendering of the video frame to be rendered depends on the specific information content contained in the multi-view video information. The process of determining the video frame to be rendered also varies depending on the information content contained.
[0177] Specifically, as one method for determining the video frame to be rendered, when the multi-view video information includes foreground mask frames and background frames from different perspectives, the video frame to be rendered can be determined by combining the foreground mask frames and background frames from different perspectives with the orientation of the terminal. For example, a matching background frame corresponding to the pose information of the second terminal can be determined from the background frames of different videos, and the video frame to be rendered can be formed by fusing the matching background frame and the foreground mask frame.
[0178] In another implementation of determining the video frame to be rendered, when the multi-view video information contains multiple multi-view video frames with different perspectives, the spatial position information of each multi-view video frame can be determined. Then, based on the second spatial position information and the pose information of the second terminal, as well as the orientation of the terminal, the video frame to be rendered can be determined from the multi-view video frames.
[0179] Similarly, in another implementation of determining the video frame to be rendered, when the multi-view video information contains 3D point cloud data, a 3D scene can be constructed using the 3D point cloud data. The rendering viewpoint of the time frame can be determined based on the pose information of the second terminal, and the corresponding 3D coordinate points in the 3D scene under the rendering viewpoint can be determined. Based on each 3D coordinate point, a 3D video frame can be constructed, and the projected video frame of the 3D video frame can be obtained. Based on the projected video frame and the orientation of the terminal, the video frame to be rendered can be determined.
[0180] For example, to better understand the display effect of the displayable video frame achieved by the method provided in this embodiment, Figure 3a shows the effect of the displayable video frame formed by the video processing method provided in this embodiment on the display terminal side when the video frame to be processed is a monocular video frame. As shown in Figure 3a, the given screen content can be considered as a displayable video frame presented by combining the pose information of the display terminal side with the three-dimensional scene construction of the monocular video frame and the determination of multi-viewpoint video frames. It can be seen that compared with the traditional monocular video frame, this displayable video frame has a sense of space, and the displayed video content is also richer.
[0181] It should be noted that when the video frame to be processed is a monocular / binocular video frame, since the video frame itself is a two-dimensional video frame, even after the construction of a three-dimensional scene and the determination of multi-view video frames, the range of view it can display is limited. It is impossible to achieve a 360-degree adjustment of the displayable viewpoint according to the pose adjustment of the display terminal. After the pose adjustment of the display terminal exceeds the displayable viewpoint range of the determined multi-view video frames, the video frames corresponding to the two ends of the displayable viewpoint range can be rendered as displayable video.
[0182] For example, Figure 3b shows an effect diagram of a displayable video frame formed by the video processing method provided in this embodiment on the display terminal side when the video frame to be processed is a multi-camera captured video frame. As shown in Figure 3b, the given screen content can be considered as a displayable video frame presented based on the three-dimensional scene construction of the multi-camera captured video frames and the determination of multi-viewpoint video frames, combined with the pose information of the display terminal side. This embodiment can realize the effect of displaying the video content contained in the video frame to be processed according to the adjustment of the pose of the display terminal with video frames corresponding to any viewpoint. It can be seen that the displayable video frame shown in Figure 3b mainly presents the display effect corresponding to adjusting the viewing angle to behind the person or object.
[0183] Similarly, Figure 3c illustrates the effect of displaying a displayable video frame on the display terminal side using the video processing method provided in this embodiment when the video frame to be processed is a multi-lens capture video frame. The content presented in Figure 3c can be considered as an illustration of the effect of processing the video frame to be processed based on multi-lens capture in Figure 1b using the video processing method provided in this embodiment on the first terminal to form the multi-view video frame shown in Figure 1c, and finally displaying the multi-view video frame on the display device side (such as the second terminal). It is equivalent to realizing a 360-degree panoramic view of the entity object shown in Figure 3c by adjusting the pose on the display device side.
[0184] S203. When an interactive operation is received that acts on the displayable video frame, if the target of the interactive operation is determined to be a set target object in the displayable video frame based on the visual shell information, then a first haptic feedback corresponding to the interactive operation is generated.
[0185] In this embodiment, after rendering the determined video frame to be rendered into a displayable video frame through the above steps and displaying the displayable video frame, the interactive operations on the displayable video frame can be monitored. If an interactive operation is detected on the displayable video frame, the interactive operation can be received through this step. The interactive operation can be dragging, rotating, or clicking on the displayable video frame.
[0186] This embodiment can determine the specific interactive event (such as drag event, rotation event, and click event) corresponding to the received interactive operation and the corresponding operation position information in three-dimensional space. By combining the operation position information with the visual shell information, collision detection can be performed in the video scene involved in the interactive operation within the displayable video frame. When the collision detection shows that the operation object touched by the interactive operation in the three-dimensional scene is a target object set in the three-dimensional scene, it can be determined that the execution condition of the first haptic feedback is met, thereby generating the first haptic feedback relative to the interactive operation. The generation logic of the first haptic feedback can be described as: generating a force feedback waveform that matches the operation event involved in the interactive operation, and using the force feedback waveform to form a vibration. At the same time, the sound effect satisfied by the operation event involved in the interactive operation can also be determined and played.
[0187] The vibration and sound effects generated after interacting with the displayable video frame can be considered as the first tactile feedback of the second terminal relative to the displayable video frame.
[0188] To facilitate better understanding, this embodiment also provides an example of the effect of haptic feedback. Figure 4a shows the effect of haptic feedback when a click operation is performed on a displayable video frame formed by the video processing method provided in this embodiment, and the haptic feedback conditions are met. As shown in Figure 4a, it can be seen that the click operation 41 is applied to the pipa player 42 in the displayable video frame. The pipa player 42 is equivalent to an entity object that the interactive operation collides with in the three-dimensional scene involved in the displayable video frame. When the pipa player 42 is set as the target object, it can be considered that the interactive operation at this time is equivalent to meeting the haptic feedback execution conditions, thereby generating a force feedback waveform for generating vibration, as shown in the figure as the first waveform 43 presented in the left sensing area of the head-mounted device.
[0189] The video processing method provided in this embodiment can determine the video compression information corresponding to a unit time frame from the video compression information stream sent by the first terminal in the form of a bitstream. By decompressing the video compression information and combining it with the pose information of the second terminal, a video frame to be rendered that corresponds to a certain capture perspective in that unit time frame can be determined. By rendering the video frame to be rendered, displayable video frames under different free perspectives can be realized. Through the above technical solution, regardless of whether the initially acquired video frame to be processed is a mono / dual-view video frame or a multi-view video frame acquired in a multi-view format, free perspective video rendering can be realized on the second terminal for video rendering according to different viewing perspectives, effectively expanding the application scope of free perspective video viewing. At the same time, while ensuring a more realistic free perspective viewing effect during rendering, this technical solution also adds haptic feedback to the displayed video. This technical solution can better perceive the force feedback of people or other entities in the video scene, realize human-computer force perception and auditory interaction, better enhance the immersion in video viewing, and also ensure the naturalness and realism of the interactive experience.
[0190] As a fifth optional embodiment of this example, based on the implementation of S201 to S203 above, a further implementation method is proposed for determining the video frame to be rendered that matches the orientation of the second terminal based on the multi-viewpoint video information and the pose information of the second terminal. The implementation steps are described as follows:
[0191] a5) When the multi-view video information consists of foreground masking frames and background frames from different viewpoints, the foreground masking frames are fused with the corresponding background frames from different viewpoints to obtain multi-view fused frames.
[0192] This optional embodiment describes how to determine the video frame to be rendered when the multi-view video information consists of foreground mask frames and background frames from different viewpoints. Specifically, this step first obtains the foreground mask frames and background frames from different viewpoints in the multi-view video information. The foreground mask frame includes the foreground content of the corresponding video frame to be processed, and each background frame includes the background content of the corresponding video frame to be processed from different viewpoints. This step can fuse the foreground content and background content, thereby forming a fused frame relative to different viewpoints. The fused frame from different viewpoints can be denoted as the multi-view fused frame.
[0193] b5) Based on the first spatial position information corresponding to the multi-view fusion frames and the pose information of the second terminal, determine the first orientation vector and the first display size of the multi-view fusion frames relative to the virtual camera in the second terminal.
[0194] In this embodiment, the multi-view fusion frame can be considered to possess spatial position information in three-dimensional space. This embodiment can record the spatial position information corresponding to each multi-view fusion frame as the first spatial position information. Simultaneously, this embodiment can also obtain the pose information corresponding to the second terminal. Through the first spatial position information and pose information of the fusion frame at each viewpoint, the first orientation vector and first display size of each multi-view fusion frame relative to the virtual camera in the second terminal can be determined.
[0195] For example, the virtual camera can be considered to be located at the position of the first terminal display screen, which is equivalent to using the pose of the second terminal as the capture pose of the virtual camera in the 3D scene. The first orientation vector can be considered as the orientation information of the multi-view video frame relative to the virtual camera. The first display size can be considered to be determined based on the distance of the virtual camera relative to foreground objects in the 3D scene. The amount of image information corresponding to the multi-view video frame when it is actually presented on the display screen can be determined based on the distance of the virtual camera relative to foreground objects in the multi-view video frame.
[0196] When the virtual camera is far from the foreground object, the multi-view video frame can present more content information, which can form a first display size that can contain the determined content information size. This first display size is in a reduced state relative to the screen size. When the virtual camera is close to the foreground object, the multi-view video frame can present less content information, which can form a first display size that can contain the determined content information size. This first display size is in a magnified state relative to the screen size.
[0197] c5) Determine the target viewpoint fusion frame from the multi-viewpoint fusion frames whose first orientation vector is parallel to the orientation of the second terminal, and determine the video area determined by the target viewpoint fusion frame relative to the first display size as the video frame to be rendered.
[0198] In this embodiment, the orientation of the second terminal can also be considered as the capture orientation of the virtual camera. This step requires determining the target viewpoint fusion frame whose first orientation vector is parallel to the orientation from the video frames corresponding to different viewpoints, i.e., multi-viewpoint video frames. Then, the video area corresponding to the first display size can be determined from the target viewpoint fusion frame, and finally, the video area is used to construct the video frame to be rendered.
[0199] The fifth optional embodiment described above provides an implementation for determining the video frame to be rendered when the multi-view video information consists of foreground mask frames corresponding to monocular / dual-view video frames and background frames from different viewpoints. This demonstrates the diversification of video rendering forms and thus reflects the expansion of the application scope of free-time video.
[0200] As a sixth optional embodiment of this example, based on the implementation of S201 to S203 described above, an implementation method can also be proposed for determining the video frame to be rendered that matches the orientation of the second terminal based on the multi-viewpoint video information and the pose information of the second terminal. The implementation steps are described as follows:
[0201] a6) When the multi-view video information is composed of multi-view video frames, obtain the second spatial location information corresponding to each multi-view video frame.
[0202] This sixth optional embodiment provides a method for determining the video to be rendered when the multi-view video information is composed of multi-view video frames. This step can obtain the second spatial location information corresponding to each multi-view video frame.
[0203] Unlike the description of the fifth optional embodiment above, the multi-view video information can be considered to be determined by the video frames to be processed captured by multiple lenses. The multi-view video frames included in the multi-view video information can be considered to be obtained by rendering the video frames to be processed captured by multiple lenses in the constructed 3D scene.
[0204] b6) Based on the second spatial location information and the pose information of the second terminal, determine the second orientation vector and the second display size of the multi-view video frame relative to the virtual camera in the second terminal.
[0205] In this embodiment, as described in the fifth optional embodiment above, the second orientation vector and the second display size of the multi-view video frame relative to the virtual camera in the two terminals can also be determined by the second spatial location information and the pose information of the second terminal.
[0206] c6) Determine the target viewpoint video frame from the multi-viewpoint video frames whose second orientation vector is parallel to the orientation of the second terminal, and determine the video area of the target viewpoint video frame relative to the second display size as the video frame to be rendered.
[0207] In this embodiment, similar to the description of the first optional embodiment above, after determining the second orientation vector, a target viewpoint video frame parallel to the orientation of the second terminal can be found from the multi-viewpoint video frames, and the video area corresponding to the second display size can be determined from the target viewpoint video frame as the video frame to be rendered.
[0208] As a seventh optional embodiment of this example, based on the implementation of S201 to S203 above, another implementation method is proposed for determining the video frame to be rendered that matches the orientation of the second terminal based on the multi-view video information and the pose information of the second terminal. The implementation steps are described as follows:
[0209] a7) When the multi-view video information is 3D point cloud data, render the 3D point cloud data to form a 3D scene.
[0210] This seventh optional embodiment provides a method for determining the video frames to be rendered when the multi-view video information is 3D point cloud data. This step can render a more realistic 3D scene using 3D point cloud data.
[0211] b7) Based on the pose information of the second terminal, determine the third display size and capture view of the virtual camera in the second terminal, and determine the three-dimensional coordinates of the target in the capture view in the three-dimensional scene.
[0212] In this embodiment, the pose information of the second terminal is used to determine the capture pose of the virtual camera. The third display size and capture view of the virtual camera can be determined by the distance of the pose information relative to the entity objects in the three-dimensional scene. Thus, the target three-dimensional coordinate points of the scene objects that can be captured from the three-dimensional scene under the capture view can be determined.
[0213] c7) Construct a three-dimensional video frame based on the target three-dimensional coordinate points and project the three-dimensional video frame onto the target plane to form a target projection video frame, and determine the video area of the target projection video frame relative to the third display size as the video frame to be rendered, wherein the target projection plane is a plane parallel to the orientation of the second terminal.
[0214] In this embodiment, a 3D video frame in a 3D scene can be constructed using the target 3D coordinate points, and this 3D video frame can be projected onto the target plane to form a target projected video frame. This step can also determine a video region matching the third display size from the target projected video frame as the video frame to be rendered.
[0215] It is known that the target projection plane for projecting three-dimensional video frames can be considered as a plane parallel to the orientation of the second terminal.
[0216] As an eighth optional embodiment of this example, based on the implementation of S201 to S203 above, when an interactive operation is received acting on the displayable video frame, if it is determined according to the visual shell information that the interactive operation satisfies the first haptic feedback condition, then the first haptic feedback corresponding to the interactive operation is specifically generated as follows:
[0217] a8) When an interactive operation is received that acts on a displayable video frame, determine the operation event corresponding to the interactive operation and the operation position information corresponding to the operation event in three-dimensional space. The operation events include click events, drag events and rotation events.
[0218] If an interactive action is triggered within a displayable video frame, this step can receive the action and determine the specific type of event it corresponds to (e.g., drag, click, or rotation) by analyzing the screen coordinates on the display screen. This step can also retrieve the specific 3D location of the operation based on the screen coordinates.
[0219] b8) Based on the operation location information and the visible shell information, determine the operation object corresponding to the operation event in the three-dimensional scene where the displayable video frame is located.
[0220] In this embodiment, after obtaining the visual shell information through the above steps, the spatial position of the entity object in the displayable video frame within the 3D scene can be determined using the visual shell information. This step performs collision detection to determine whether the operation event touches the entity object, based on the operation position information of the operation event and the visual shell information that characterizes the spatial position of the entity object in the 3D scene. Furthermore, when the operation event touches the entity object, that entity object can be identified as the operation object of the operation event.
[0221] c8) If the object being operated on belongs to the set target object in the 3D scene, then generate a force feedback waveform that matches the operation event, generate a vibration sensation on the stress feedback waveform and play the corresponding sound effect of the operation event.
[0222] In this embodiment, if the identified operation object belongs to a target object pre-defined in the 3D scene, the interactive operation can be considered to have met the execution conditions for haptic feedback, thus generating vibration through this step. Specifically, a force feedback waveform matching the operation event can be generated, and the generated force feedback waveform forms the corresponding vibration.
[0223] It should be noted that this embodiment often presents not only haptic feedback, but also auditory and even visual feedback. Different auditory feedback can correspond to different sound effects for different operation events. Specifically, the force feedback waveform used for haptic feedback, the sound effects for auditory feedback, and even the visual reminders for visual feedback all vary depending on the operation event.
[0224] This embodiment, as described in the optional embodiments above, provides a specific implementation of haptic feedback generated through interactive operations. The provided haptic feedback allows for better perception of force feedback from people or other entities in the video scene, enabling human-computer force and auditory interaction, thus enhancing the immersive experience of video viewing and ensuring the naturalness and realism of the interactive experience.
[0225] As a ninth optional embodiment of this example, based on the implementation of S201 to S203 described above, the following steps may be further included:
[0226] Based on the previously displayed video frames and the displayable video frames, when it is determined that there is a dynamic object in the displayable video frame and the action behavior of the dynamic object meets the second haptic feedback condition, the second haptic feedback for the corresponding displayable video frame is generated.
[0227] It should be noted that this ninth optional embodiment provides another implementation of haptic feedback for displaying physical objects in a video frame. As we know, physical objects such as people and animals in the real world are generally capable of movement, and this movement provides haptic feedback in the real world. For example, a boxer feels the wind when throwing a punch, and there is a stomping sound when feet touch the ground. This embodiment can simulate the haptic feedback given by actions in the real world through the added second haptic feedback implementation logic.
[0228] For example, this embodiment first needs to determine which objects in the video content included in the displayable video frame can be considered as dynamic objects capable of making movements. Then, it can analyze the movement behavior of the dynamic objects. After the movement behavior satisfies the second somatosensory feedback condition (such as the appearance of punching, stomping, or other real-world movements that can bring about changes in body sensation), a second somatosensory feedback associated with the dynamic object in the displayable video frame can be formed. For example, providing the sound effect of punching and simulating the vibration of the punch's force.
[0229] For example, in this embodiment, the presence of a dynamically changing target object can be determined by combining the displayable video frames with the previously identified displayable video frames over a period of time. If such a target object exists, it can be considered to be a dynamic object.
[0230] For example, Figure 4b illustrates the effect of haptic feedback generated when a target object in a displayable video frame formed by the video processing method provided in this embodiment of the present disclosure meets the haptic feedback conditions. As shown in Figure 4b, the displayable video frame includes a person 44 performing boxing exercises or a boxing performance. Therefore, haptic feedback can be given to the action behavior of the person 44, such as a force feedback waveform simulating the wind of a punch, thereby forming a vibration. The vibration can be fed back through the second waveform 45 of the sensing area on the left side of the head-mounted device shown in the figure.
[0231] Furthermore, to facilitate a better understanding of the video processing method provided in this embodiment, an example is given for description. Specifically, this example takes a live streaming scenario as an example, where the broadcaster's terminal uses a single-lens camera to capture the video stream to be processed. Based on this, Figure 5a shows a flowchart of a video processing implementation using the video processing method provided in this embodiment in a live streaming scenario. As shown in Figure 5, the live streaming link includes the broadcaster's terminal 51, the live streaming cloud server 52, and the viewer's terminal 53. The process of implementing video processing in the entire live streaming link can be described as follows:
[0232] S1. The broadcaster terminal captures images through a monocular camera, forms a video stream to be processed, encodes the video stream to be processed, and transmits it to the live streaming cloud server.
[0233] S2. The live cloud server decodes the video stream to be processed and obtains the spatial depth information determined by depth estimation and background segmentation of the video frames to be processed in the video stream. It also constructs the first visual shell of the video frames to be processed based on the foreground mask frame and the background mask frame, and obtains the first visual shell information.
[0234] S3. The live streaming cloud server fills the background mask frame with background content to obtain the background filled image, and generates a multi-view image from the background filled image, as well as the first multi-view video information of the video frame to be processed based on the background frame and the foreground mask frame.
[0235] S4. The live streaming cloud server combines the foreground masking frame and the multi-view background frame contained in the first multi-view video information with the video frame to be processed to perform image stitching to form the first stitched video frame.
[0236] S5. The live streaming cloud server performs video encoding and compression on the first spliced video frame to form an encoded video frame, and extracts the first visible shell information to form the first grid data.
[0237] S6. The live streaming cloud server uses the compressed data of the encoded video frames and the first grid data to form the video compression information of the video frames to be processed.
[0238] S7, the live streaming cloud server transmits the video compression information to the viewer's terminal.
[0239] S8. The viewer-side terminal decompresses the received video compression information to obtain the decoded multi-view video information and visual shell information.
[0240] S9. The viewer-side terminal merges the foreground masking frames and background frames from different viewpoints in the multi-view video information to obtain a multi-view fused frame.
[0241] S10. The viewer-side terminal determines the first orientation vector and the first display size of the multi-view fusion frame relative to the virtual camera in the second terminal based on the first spatial position information corresponding to the multi-view fusion frame and the pose information of the second terminal.
[0242] S11. The viewer-side terminal determines the target viewpoint fusion frame from the multi-viewpoint fusion frame whose first orientation vector is parallel to the orientation of the second terminal, and determines the video area determined by the target viewpoint fusion frame relative to the first display size as the video frame to be rendered.
[0243] S12. The viewer-side terminal renders the video frames to be rendered, forming displayable video frames and then displays them.
[0244] S13. When the viewer-side terminal receives an interactive operation that acts on a displayable video frame, if it determines that the interactive operation meets the first haptic feedback condition based on the visual shell information, then the first haptic feedback for the corresponding interactive operation is generated.
[0245] S14. The viewer-side terminal determines, based on the displayable video frames and the displayable video frames in front, that there are dynamic objects in the displayable video frames and that the actions of the dynamic objects meet the second haptic feedback conditions, and then generates the second haptic feedback for the corresponding displayable video frame.
[0246] Similarly, to better understand the video processing method provided in this embodiment, an example is given for description. Specifically, this example takes a live streaming scenario as an example, where the broadcaster's terminal uses multiple cameras to capture the video stream to be processed. Based on this, Figure 5b shows another implementation flowchart of video processing using the video processing method provided in this embodiment in a live streaming scenario. As shown in Figure 5b, it also includes the broadcaster's terminal 51, the live streaming cloud server 52, and the viewer's terminal 53 in the live streaming link. The process of implementing video processing in the entire live streaming link in this example can be described as follows:
[0247] S01. The broadcaster terminal captures images through a multi-camera system, forming a video stream to be processed that includes multiple video frames to be processed within the same time frame. The video stream to be processed is then encoded and transmitted to the live streaming cloud server.
[0248] S02. The live streaming cloud server determines the relative pose information between pairs of video frames based on the video frame association information between multiple video frames that are to be processed.
[0249] S03. The live streaming cloud server determines the three-dimensional coordinate information and the absolute pose information of the video frame for scene construction based on the relative pose information and the motion structure reconstruction algorithm, and constructs the three-dimensional scene of the video frame to be processed based on the three-dimensional coordinate information and the absolute pose information.
[0250] S04. The live streaming cloud server determines the side silhouette contour information of objects in the 3D scene from multiple perspectives, determines the construction of a visual shell relative to the object based on the side silhouette contour information, and obtains the corresponding second visual shell information as the second visual shell of the video frame to be processed.
[0251] S05. The live streaming cloud server uses a target neural radiation field model or a three-dimensional Gaussian splash model, combined with the second visible shell information, to perform multi-view content rendering on the three-dimensional scene and obtain the second multi-view video information of the video frame to be processed.
[0252] S06. When multi-view video information consists of video frames from multiple different viewpoints, the live streaming cloud server groups the video frames from different viewpoints to obtain at least one group of grouped video frames, and compresses and encodes the grouped video frames according to the aggregation coding strategy to form corresponding compressed video frames.
[0253] S07. When the multi-view video information is three-dimensional point cloud data, the live streaming cloud server performs compression encoding according to the set point cloud data compression strategy to form corresponding point cloud compressed information.
[0254] S08. The live streaming cloud server uses the second grid data formed by the second visual shell information and the compressed video frame or point cloud compression information to form the video compression information of the video frame to be processed.
[0255] S09. The live streaming cloud server transmits the video compression information to the viewer's terminal.
[0256] S010, The viewer-side terminal decompresses the received video compression information to obtain the decoded multi-view video information and visual shell information.
[0257] S011. When the multi-view video information is composed of multi-view video frames, the viewer-side terminal obtains the second spatial location information corresponding to each of the multi-view video frames.
[0258] S012. The viewer-side terminal determines the second orientation vector and the second display size of the multi-view video frames relative to the virtual camera in the second terminal based on the second spatial position information and the pose information of the second terminal.
[0259] S013. The viewer-side terminal determines the target viewpoint video frame from the multi-viewpoint video frames whose second orientation vector is parallel to the orientation of the second terminal, and determines the video area of the target viewpoint video frame relative to the second display size as the video frame to be rendered.
[0260] S014. When the multi-view video information is three-dimensional point cloud data, the viewer-side terminal renders the three-dimensional point cloud data to form a three-dimensional scene.
[0261] S015. The audience-side terminal determines the third display size and capture angle of the virtual camera in the second terminal based on the pose information of the second terminal, and determines the three-dimensional coordinate point of the target in the capture angle in the three-dimensional scene.
[0262] S016. The viewer-side terminal constructs a three-dimensional video frame based on the target three-dimensional coordinate point and projects the three-dimensional video frame onto the target plane to form a target projection video frame, and determines the video area of the target projection video frame relative to the third display size as the video frame to be rendered, wherein the target projection plane is a plane parallel to the orientation of the second terminal.
[0263] S017. The viewer-side terminal renders the video frames to be rendered, forming displayable video frames and then displays them.
[0264] S018. When the viewer-side terminal receives an interactive operation that acts on a displayable video frame, if it determines that the interactive operation meets the first haptic feedback condition based on the visual shell information, then the first haptic feedback for the corresponding interactive operation is generated.
[0265] S019. The viewer-side terminal determines, based on the displayable video frames and the displayable video frames in front, that there are dynamic objects in the displayable video frames and that the actions of the dynamic objects meet the second haptic feedback conditions, and then generates the second haptic feedback for the corresponding displayable video frame.
[0266] S018. When an interactive operation is received that acts on a displayable video frame, if it is determined from the visual shell information that the interactive operation meets the first haptic feedback condition, then the first haptic feedback for the corresponding interactive operation is generated.
[0267] S019. Given a displayable video frame and a displayable video frame, when it is determined that there is a dynamic object in the displayable video frame and the action behavior of the dynamic object meets the second somatosensory feedback condition, a second somatosensory feedback is formed for the corresponding displayable video frame.
[0268] Figure 6 shows a schematic diagram of a video processing device provided in an embodiment of this disclosure. This embodiment is applicable to video playback. The device can be implemented by software and / or hardware and can be configured in a terminal and / or server to implement the video processing method in this disclosure. The device can be configured on a first terminal and may specifically include: a video acquisition module 61, an information determination module 62, a video compression module 63, and an information transmission module 64.
[0269] Among them, the video acquisition module 61 is used to acquire the video frame to be processed and the acquisition attribute information of the video frame to be processed. The video frame to be processed is the video frame in the received video stream to be processed. The acquisition attribute information includes acquisition based on monocular or binocular lens and acquisition based on multi-view lens.
[0270] The information determination module 62 is used to construct a visual shell of the video to be processed based on the collected attribute information, and determine the visual shell information of the visual shell; and to determine the multi-view video information of the video frame to be processed.
[0271] The video compression module 63 is used to compress video content based on visual shell information and multi-view video information, and generate video compression information of the video frame to be processed.
[0272] The information transmission module 64 is used to transmit video compression information to the second terminal, so that the second terminal can perform video rendering based on the decompressed content of the video compression information to form displayable video frames.
[0273] This embodiment provides a video processing device that, regardless of whether the video frame to be processed is a monocular / binocular video frame or a multi-view video frame acquired in a multi-view format, can perform secondary processing by adding a visual shell and determining multi-viewpoint video. This enables the conversion of the video frame to three-dimensional space and the determination of the corresponding multi-viewpoint video frame. Based on the constructed visual shell and multi-viewpoint video information, in a video playback scenario, the video frame to be processed can achieve free-viewpoint video rendering according to different viewing angles, thus providing a fundamental information guarantee for effectively expanding the application scope of free-viewpoint video viewing. This technical solution ensures that the rendered video has a more realistic free-viewpoint viewing effect.
[0274] Furthermore, the information determination module 62 can specifically be used for:
[0275] When the attribute information is acquired based on monocular or binocular lens acquisition, the acquired video frame to be processed is determined to be a monocular video frame or a binocular video frame, and depth estimation and foreground processing are performed on the video frame to be processed to obtain spatial depth information, foreground masking frame and foreground missing frame.
[0276] Based on spatial depth information and foreground missing frames, the background mask frame of the video frame to be processed is determined, and the first visual shell of the video frame to be processed is constructed based on the foreground mask frame and the background mask frame. The triangular facet information, texture information and material information representing the first visual shell are determined as the first visual shell information.
[0277] The background masking frame is cropped from different viewpoints, and the cropped image area is determined as the background frame under different viewpoints. The first multi-viewpoint video information is formed by the background frame and the foreground masking frame to be processed.
[0278] Furthermore, the video compression module 63 can specifically be used for:
[0279] The foreground masking frame and the multi-view background frame contained in the first multi-view video information are combined with the video frame to be processed to perform image stitching to form the first stitched video frame.
[0280] The first spliced video frame is video encoded and compressed to form an encoded video frame, and the vertex coordinates of the triangular facet information extracted from the first visible shell information are used to form the first grid data.
[0281] Video compression information is composed of encoded video frames and first grid data to form the video frames to be processed.
[0282] Furthermore, the information determination module 62 may specifically include:
[0283] The frame determination unit is used to determine, when the acquisition attribute information is based on multi-view lens acquisition, that the acquired video frames to be processed include video frames captured by different lenses in the same time frame.
[0284] The information construction unit is used to construct a three-dimensional scene of the video frame to be processed based on multiple video frames as video frames to be processed, and to construct a second visual shell of the video frame to be processed based on the constructed three-dimensional scene, and to determine the triangular facet information, texture information and material information representing the second visual shell as the second visual shell information.
[0285] The information acquisition unit is used to perform multi-view content rendering on the 3D scene by adopting a selected multi-view rendering strategy combined with the second visual shell information, and to obtain the second multi-view video information of the video frame to be processed.
[0286] Furthermore, the information building block can be specifically used for:
[0287] Based on the video frame association information between multiple video frames, determine the relative pose information between each pair of video frames;
[0288] Based on the relative pose information, the motion structure reconstruction algorithm determines the three-dimensional coordinate information and the absolute pose information of the video frame for scene construction, and constructs the three-dimensional scene of the video frame to be processed based on the three-dimensional coordinate information and the absolute pose information.
[0289] Determine the silhouette information of entity objects in a 3D scene from multiple perspectives, construct a visual shell relative to the entity object based on the silhouette information, and use it as the second visual shell of the video frame to be processed to obtain the corresponding second visual shell information.
[0290] Furthermore, the information acquisition unit can specifically be used for:
[0291] The second visual shell information, the three-dimensional coordinate information, ray angle information, color information and density information of the objects in the three-dimensional scene are used as input information and input into the given target neural radiation field model.
[0292] Multiple video frames with different viewpoint information output by the target neural radiation field model are used as multi-view video frames to be processed, and second multi-view video information is constructed based on the multi-view video frames. The video frames are obtained by rendering the three-dimensional scene from different viewpoints.
[0293] In the training of the target neural radiation field, in addition to the spatial position information, angle information and color information of the sampling viewpoint, a set of sample data in the sample training set also includes the sampling visual shell information of the scene objects under the sampling viewpoint, so as to enhance the geometric representation of the scene objects in the sampling viewpoint through the sampling visual shell information.
[0294] Furthermore, the information acquisition unit can also be used specifically for:
[0295] The three-dimensional coordinate information of objects in the three-dimensional scene is used as input information and input into the given target three-dimensional Gaussian splash model;
[0296] The three-dimensional point cloud data output by the three-dimensional Gaussian splash model is obtained, and the three-dimensional point cloud data is determined as the second multi-view video information of the video frame to be processed.
[0297] Furthermore, the video compression module 63 can also be used specifically for:
[0298] The vertex coordinates of the triangular facets in the second visible shell information are extracted to form the second mesh data;
[0299] When the second multi-view video information consists of video frames from multiple different viewpoints, the video frames from different viewpoints are grouped to obtain at least one group of grouped video frames, and the grouped video frames are compressed and encoded according to the aggregation coding strategy to form the corresponding compressed video frames.
[0300] When the second multi-view video information is three-dimensional point cloud data, it is compressed and encoded according to the set point cloud data compression strategy to form the corresponding point cloud compressed information.
[0301] The compressed data based on the second grid data, along with the compressed video frame or point cloud compression information, constitutes the video compression information of the video frame to be processed.
[0302] The above-described apparatus can execute the methods provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the methods.
[0303] Figure 7 shows a schematic diagram of a video processing apparatus according to another embodiment of the present disclosure. This embodiment is applicable to video playback. The apparatus can be implemented by software and / or hardware and can be configured in a terminal and / or server to implement the video processing method of the present disclosure. The apparatus can be configured on a second terminal and may specifically include: a video decoding module 71, a video rendering module 72, and a first feedback module 73.
[0304] The video decoding module 71 is used to decompress the received current video compression information to obtain the decoded current multi-view video information and current visual shell information. The current video compression information is the video compression information in the video compression information stream sent by the first terminal. The video compression information is formed by the first terminal compressing the video content based on the visual shell information and multi-view video information determined by the first terminal based on the video frame to be processed in the video stream to be processed.
[0305] The video rendering module 72 is used to determine the video frame to be rendered that matches the orientation of the second terminal based on the multi-view video information and the pose information of the second terminal, render the video frame to be rendered to form the currently displayable video frame and display it.
[0306] The first feedback module 73 is used to generate first haptic feedback for the corresponding interactive operation when it receives an interactive operation applied to the currently displayable video frame and determines that the interactive operation meets the first haptic feedback condition based on the current visible shell information.
[0307] This embodiment provides a video processing device that can determine the video compression information corresponding to a unit time frame from the video compression information stream sent by a first terminal in the form of a bitstream. By decompressing the video compression information and combining it with the pose information of a second terminal, it determines the video frame to be rendered that corresponds to a certain capture perspective in that unit time frame. By rendering the video frame to be rendered, it can display video frames under different free perspectives. Through the above technical solution, regardless of whether the initially acquired video frame to be processed is a mono / dual-view video frame or a multi-view video frame acquired in a multi-view format, free perspective video rendering can be achieved on the second terminal for video rendering according to different viewing perspectives, effectively expanding the application scope of free perspective video viewing. At the same time, while ensuring a more realistic free perspective viewing effect during rendering, this technical solution also adds haptic feedback to the displayed video. This technical solution can better perceive the force feedback of people or other entities in the video scene, realize human-computer force perception and auditory interaction, better enhance the immersion in video viewing, and also ensure the naturalness and realism of the interactive experience.
[0308] Furthermore, the video rendering module 72 can specifically be used for:
[0309] When multi-view video information consists of foreground masking frames and background frames from different viewpoints, the foreground masking frames are fused with the corresponding background frames from different viewpoints to obtain multi-view fused frames.
[0310] Based on the first spatial position information corresponding to the multi-view fusion frames and the pose information of the second terminal, the first orientation vector and the first display size of the multi-view fusion frames relative to the virtual camera in the second terminal are determined.
[0311] From the multi-view fusion frames, determine the target view fusion frame whose first orientation vector is parallel to the orientation of the second terminal, and determine the video area of the target view fusion frame relative to the first display size as the video frame to be rendered.
[0312] Furthermore, the video rendering module 72 can also be used specifically for:
[0313] Given that the current multi-view video information is composed of multi-view video frames, obtain the second spatial location information corresponding to each multi-view video frame;
[0314] Based on the second spatial location information and the current pose information of the second terminal, determine the second orientation vector and the second current display size of the multi-view video frames relative to the virtual camera in the second terminal.
[0315] From the multi-view video frames, determine the target view video frame whose second orientation vector is parallel to the current orientation of the second terminal, and determine the video area of the target view video frame relative to the second current display size as the video frame to be rendered.
[0316] Furthermore, the video rendering module 72 can also be specifically used for:
[0317] Given that the current multi-view video information is 3D point cloud data, the 3D point cloud data is rendered to form the current 3D scene.
[0318] Based on the current pose information of the second terminal, determine the third current display size and current capture view of the virtual camera in the second terminal, and determine the target 3D coordinate point in the current 3D scene that is in the current capture view.
[0319] A three-dimensional video frame is constructed based on the target three-dimensional coordinate points, and the three-dimensional video frame is projected onto the target plane to form a target projection video frame. The video area determined by the target projection video frame relative to the third display size is determined as the video frame to be rendered. The target projection plane is a plane parallel to the current orientation of the second terminal.
[0320] Furthermore, the first feedback module 73 can specifically be used for:
[0321] Based on the current visual shell information, determine the object visual shell information of the scene object corresponding to the video content of the currently displayable video frame in the 3D scene;
[0322] When an interactive operation is received that is applied to the currently displayable video frame, the corresponding operation event and the operation position information in the three-dimensional space are determined.
[0323] Based on the operation location information and the object's visible shell information, determine the collision detection of the operation event relative to the scene object;
[0324] If the scene object that caused the collision is detected to be the set target object, a force feedback waveform that matches the operation event is generated, producing a vibration sensation on the stress feedback waveform while playing the corresponding sound effect of the operation event.
[0325] Furthermore, the device also includes a second feedback module, which is used to determine, based on the previous displayable video frame and the current displayable video frame, that there is a dynamic object in the current displayable video frame and the action behavior of the dynamic object satisfies the second haptic feedback condition, and then generate a second haptic feedback corresponding to the current displayable video frame.
[0326] The above-described apparatus can execute the methods provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the methods.
[0327] It is worth noting that the various units and modules included in the above-mentioned device are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be realized; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of this disclosure.
[0328] Figure 8 illustrates a schematic diagram of the structure of a computer device provided in an embodiment of this disclosure. Referring now to Figure 8, a schematic diagram of the structure of a computer device (e.g., the terminal device or server in Figure 8) 80 suitable for implementing embodiments of this disclosure is shown. The terminal device in the embodiments of this disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. The computer device shown in Figure 8 is merely an example and should not impose any limitations on the functionality and scope of use of the embodiments of this disclosure.
[0329] As shown in Figure 8, the computer device 80 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 81, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 82 or a program loaded from a storage device 88 into a random access memory (RAM) 83. The RAM 83 also stores various programs and data required for the operation of the computer device 80. The processing unit 81, ROM 82, and RAM 83 are interconnected via a bus 85. An edit / output (I / O) interface 84 is also connected to the bus 85.
[0330] Typically, the following devices can be connected to I / O interface 84: input devices 86 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 87 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 88 including, for example, magnetic tapes, hard disks, etc.; and communication devices 89. Communication device 89 allows computer device 80 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 8 shows a computer device 80 with various devices, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0331] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 89, or installed from a storage device 88, or installed from a ROM 82. When the computer program is executed by a processing device 81, it performs the functions defined in the methods of embodiments of this disclosure.
[0332] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0333] The computer device provided in this embodiment and the video processing method provided in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0334] This disclosure provides a computer storage medium storing a computer program that, when executed by a processor, implements the video processing method provided in the above embodiments.
[0335] It should be noted that the computer-readable medium described above in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof.
[0336] In this disclosure, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0337] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol), and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0338] The aforementioned computer-readable medium may be included in the aforementioned computer device; or it may exist independently and not assembled into the computer device.
[0339] The aforementioned computer-readable medium carries one or more programs that, when executed by the computer device, cause the computer device to:
[0340] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0341] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0342] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".
[0343] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0344] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0345] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0346] Furthermore, although the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while some specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0347] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.
Claims
1. A video processing method applied to a first terminal, the method comprising: Acquire the video frame to be processed and the acquisition attribute information of the video frame to be processed. The video frame to be processed is a video frame in the received video stream to be processed. The acquisition attribute information includes acquisition based on monocular or binocular lenses, as well as acquisition based on multi-view lenses. Based on the acquired attribute information, a visual shell of the video frame to be processed is constructed, and the visual shell information of the visual shell is determined; In addition, determine the multi-view video information of the video frame to be processed; Based on the visual shell information and the multi-view video information, video content compression is performed to generate video compression information for the video frame to be processed. The video compression information is transmitted to the second terminal, so that the second terminal can perform video rendering based on the decompressed content of the video compression information to form displayable video frames.
2. The method according to claim 1, wherein, The step involves constructing a visual shell for the video frame to be processed based on the acquired attribute information, and determining the visual shell information of the visual shell. And, determining the multi-view video information of the video frame to be processed, including: When the acquired attribute information is acquired based on monocular or binocular lens acquisition, the acquired video frame to be processed is determined to be a monocular video frame or a binocular video frame, and depth estimation and foreground processing are performed on the video frame to be processed to obtain spatial depth information, foreground masking frame and foreground missing frame. Based on the spatial depth information and the foreground missing frame, a background mask frame of the video frame to be processed is determined, and a first visual shell of the video frame to be processed is constructed based on the foreground mask frame and the background mask frame. The triangular facet information, texture information and material information that characterize the first visual shell are determined as the first visual shell information. The background masking frame is cropped from different viewpoints, and the cropped image regions are determined as background frames from different viewpoints. The first multi-viewpoint video information of the video frame to be processed is constructed based on the background frames and the foreground masking frame.
3. The method according to claim 2, wherein, The step of compressing video content based on the visible shell information and the multi-view video information to generate video compression information for the video frame to be processed includes: The foreground masking frame and the multi-view background frame contained in the first multi-view video information are combined with the video frame to be processed to perform image stitching to form the first stitched video frame. The first spliced video frame is video encoded and compressed to form an encoded video frame, and the vertex coordinates of the triangular facet information in the first visible shell information are extracted to form the first grid data. The video compression information of the video frame to be processed is constructed based on the encoded video frame and the first grid data.
4. The method according to claim 1, wherein, The step involves constructing a visual shell for the video frame to be processed based on the acquired attribute information, and determining the visual shell information of the visual shell. And, determining the multi-view video information of the video frame to be processed, including: When the acquired attribute information is acquired based on multi-camera acquisition, it is determined that the acquired video frames to be processed include video frames captured by different cameras in the same time frame. Based on multiple video frames that are the video frames to be processed, a three-dimensional scene of the video frames to be processed is constructed, and based on the constructed three-dimensional scene, a second visual shell of the video frames to be processed is constructed, and the triangular facet information, texture information and material information that characterize the second visual shell are determined as the second visual shell information. By employing a selected multi-view rendering strategy in conjunction with the second visual shell information, multi-view content rendering is performed on the three-dimensional scene to obtain the second multi-view video information of the video frame to be processed.
5. The method according to claim 4, wherein, The process of constructing a 3D scene of the video frame to be processed based on multiple video frames, and constructing a second visual shell of the video frame to be processed based on the constructed 3D scene, includes: Based on the video frame association information between the multiple video frames, the relative pose information between each pair of video frames is determined; Based on the relative pose information, the motion structure reconstruction algorithm is used to determine the three-dimensional coordinate information for scene construction and the absolute pose information of the video frame. Based on the three-dimensional coordinate information and the absolute pose information of the video frame, combined with the three-dimensional scene reconstruction model, a three-dimensional scene of the video frame to be processed and a visual shell of the entity objects in the three-dimensional scene are constructed, and the visual shell is denoted as the second visual shell of the video frame to be processed.
6. The method according to claim 4 or 5, wherein, The step of employing a selected multi-view rendering strategy combined with the second visual shell information to perform multi-view content rendering on the 3D scene, and obtaining the second multi-view video information of the video frame to be processed, includes: The second visible shell information, the three-dimensional coordinate information, ray angle information, color information and density information of the objects in the three-dimensional scene are used as input information and input into the given target neural radiation field model; The multiple video frames with different viewpoint information output by the target neural radiation field model are used as the multi-view video frames of the video frame to be processed, and the second multi-view video information is constructed based on the multi-view video frames. The video frames output by the target neural radiation field model are obtained by rendering the three-dimensional scene under different viewpoints. In the training of the target neural radiation field, in addition to the spatial position information, angle information and color information of the sampling viewpoint, a set of sample data in the sample training set also includes the sampling visual shell information of the scene objects under the sampling viewpoint, so as to enhance the geometric representation of the scene objects in the sampling viewpoint through the sampling visual shell information.
7. The method according to claim 4 or 5, wherein, The step of employing a selected multi-view rendering strategy combined with the second visual shell information to perform multi-view content rendering on the 3D scene, and obtaining the second multi-view video information of the video frame to be processed, includes: The three-dimensional coordinate information of the objects in the three-dimensional scene is used as input information and input into the given target three-dimensional Gaussian splash model; The three-dimensional point cloud data output by the three-dimensional Gaussian splash model is obtained, and the three-dimensional point cloud data is determined as the second multi-view video information of the video frame to be processed.
8. The method according to any one of claims 4 to 7, wherein, The step of compressing video content based on the visible shell information and the multi-view video information to generate video compression information for the video frame to be processed includes: Extract the vertex coordinates of the triangular facets from the second visible shell information to form the second mesh data; When the second multi-view video information consists of video frames from multiple different viewpoints, the video frames from different viewpoints are grouped to obtain at least one group of grouped video frames, and the grouped video frames are compressed and encoded according to the aggregation coding strategy to form corresponding compressed video frames. When the second multi-view video information is three-dimensional point cloud data, it is compressed and encoded according to the set point cloud data compression strategy to form corresponding point cloud compressed information; The compressed data based on the second grid data and the compressed video frame or the point cloud compression information constitute the video compression information of the video frame to be processed.
9. A video processing method applied to a second terminal, the method comprising: The received video compression information is decompressed to obtain decoded multi-view video information and visual shell information. The video compression information is obtained from the video compression information stream sent by the first terminal. The video compression information is formed by the first terminal compressing the video content based on the visual shell information and multi-view video information of the video frame to be processed. The video frame to be processed is obtained from the video stream to be processed. Based on the multi-view video information and the pose information of the second terminal, determine the video frame to be rendered that matches the orientation of the second terminal, render the video frame to be rendered to form a displayable video frame and display it. When an interactive operation is received that acts on the displayable video frame, if the target of the interactive operation is determined to be a set target object in the displayable video frame based on the visual shell information, then a first haptic feedback corresponding to the interactive operation is generated.
10. The method according to claim 9, wherein, The step of determining the video frame to be rendered that matches the orientation of the second terminal based on the multi-view video information and the pose information of the second terminal includes: When the multi-view video information consists of a foreground masking frame and background frames from different viewpoints, the foreground masking frame is fused with the corresponding background frame from different viewpoints to obtain a multi-view fused frame. Based on the first spatial position information corresponding to the multi-view fusion frames and the pose information of the second terminal, the first orientation vector and the first display size of the multi-view fusion frames relative to the virtual camera in the second terminal are determined. From the multi-viewpoint fusion frames, a target viewpoint fusion frame whose first orientation vector is parallel to the orientation of the second terminal is determined, and the video area determined by the target viewpoint fusion frame relative to the first display size is determined as the video frame to be rendered.
11. The method according to claim 9, wherein, The step of determining the video frame to be rendered that matches the orientation of the second terminal based on the multi-view video information and the pose information of the second terminal includes: When the multi-view video information is composed of multi-view video frames, the second spatial location information corresponding to each of the multi-view video frames is obtained; Based on the second spatial location information and the pose information of the second terminal, the second orientation vector and the second display size of the multi-view video frame relative to the virtual camera in the second terminal are determined respectively; From the multi-view video frames, determine the target view video frame whose second orientation vector is parallel to the orientation of the second terminal, and determine the video area of the target view video frame relative to the second display size as the video frame to be rendered.
12. The method according to claim 9, wherein, The step of determining the video frame to be rendered that matches the orientation of the second terminal based on the multi-view video information and the pose information of the second terminal includes: When the multi-view video information is three-dimensional point cloud data, the three-dimensional point cloud data is rendered to form a three-dimensional scene; Based on the pose information of the second terminal, the third display size and capture angle of the virtual camera in the second terminal are determined, and the target three-dimensional coordinate point in the three-dimensional scene within the capture angle is determined; A three-dimensional video frame is constructed based on the target three-dimensional coordinate points, and the three-dimensional video frame is projected onto the target plane to form a target projection video frame. The video area determined by the target projection video frame relative to the third display size is determined as the video frame to be rendered. The target projection plane is a plane parallel to the orientation of the second terminal.
13. The method according to any one of claims 9 to 12, wherein, When an interactive operation is received acting on the displayable video frame, if the target of the interactive operation is determined to be a set target object in the displayable video frame based on the visual shell information, then a first haptic feedback corresponding to the interactive operation is generated, including: When an interactive operation is received that acts on the video frame that can be displayed, the operation event corresponding to the interactive operation and the operation position information corresponding to the operation event in three-dimensional space are determined. The operation event includes click event, drag event and rotation event. Based on the operation location information and the visible shell information, determine the operation object corresponding to the operation event in the three-dimensional scene where the displayable video frame is located; If the object being operated on belongs to a set target object in the three-dimensional scene, a force feedback waveform matching the operation event is generated, producing a vibration corresponding to the force feedback waveform while playing a sound effect corresponding to the operation event.
14. The method according to any one of claims 9 to 13, further comprising: Based on the previously displayable video frames and the displayable video frames, when it is determined that there is a dynamic object in the displayable video frames and the action behavior of the dynamic object satisfies the second haptic feedback condition, a second haptic feedback corresponding to the displayable video frames is generated.
15. A video processing apparatus, configured on a first terminal, the apparatus comprising: The video acquisition module is configured to acquire a video frame to be processed and the acquisition attribute information of the video frame to be processed. The video frame to be processed is a video frame in the received video stream to be processed. The acquisition attribute information includes acquisition based on monocular or binocular lenses, as well as acquisition based on multi-lens lenses. The information determination module is configured to construct a visual shell of the video frame to be processed based on the acquired attribute information, and determine the visual shell information of the visual shell. In addition, determine the multi-view video information of the video frame to be processed; The video compression module is configured to compress video content based on the visual shell information and the multi-view video information, and generate video compression information of the video frame to be processed. The information transmission module is configured to transmit the video compression information to a second terminal, so that the second terminal can perform video rendering based on the decompressed content of the video compression information to form displayable video frames.
16. A video processing apparatus, configured in a second terminal, the apparatus comprising: The video decoding module is configured to decompress the received video compression information to obtain decoded multi-view video information and visual shell information. The video compression information is obtained from the video compression information stream sent by the first terminal. The video compression information is formed by the first terminal compressing the video content based on the visual shell information and multi-view video information of the video frame to be processed. The video frame to be processed is obtained from the video stream to be processed. The video rendering module is configured to determine a video frame to be rendered that matches the orientation of the second terminal based on the multi-view video information and the pose information of the second terminal, render the video frame to be rendered to form a displayable video frame and display it. The first feedback module is configured to, when receiving an interactive operation applied to the displayable video frame, generate a first haptic feedback corresponding to the interactive operation if the target of the interactive operation is determined to be a set target object in the displayable video frame based on the visual shell information.
17. A computer device, comprising: One or more processors; Storage device, configured to store one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the video processing method as described in any one of claims 1-14.
18. A computer-readable storage medium storing a computer program, wherein, When the computer program is executed by a processor, it implements the video processing method as described in any one of claims 1-14.
19. A computer program product comprising a computer program that, when executed by a processor, implements the video processing method according to any one of claims 1-14.
Citation Information
Patent Citations
Video processing method and device, equipment and storage medium
CN121334402A
Video reconstruction method, system and device and computer readable storage medium
CN111667438A
Image processing device, encoding device, decoding device, image processing method, program, encoding method, and decoding method
CN111788601A
5G strong-interaction remote special delivery teaching system based on holographic terminal and working method of 5G strong-interaction remote special delivery teaching system
CN112562433A
Modeling 3D objects with opacity hulls
US20030231173A1