Cross-scene VR3D live broadcast method and system based on fixed specially-like seat
By capturing video from fixed premium seats and generating 3D signals using a lightweight deep learning model, the VR3D live streaming method solves the problems of high cost and poor experience in VR live streaming, realizes a low-cost and easy-to-deploy VR3D live streaming system, and enhances the user's stereoscopic experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 北京市艺值科技有限公司
- Filing Date
- 2025-11-25
- Publication Date
- 2026-05-01
AI Technical Summary
Existing VR live streaming technology is costly and complex to operate, making it difficult for ordinary users and organizers of small and medium-sized events to afford. Off-site users cannot share a unified VIP viewpoint, resulting in poor experience consistency and insufficient 3D effect.
We adopt a cross-scene VR3D live streaming method based on fixed premium seats. We build a four-level architecture through 5G, WiFi and CDN network connections, use lightweight deep learning models to generate 3D signals, and combine them with lightweight VR glasses to provide a low-cost and easy-to-deploy VR3D live streaming system.
It enables low-cost and easy-to-deploy VR3D live streaming, allowing users in the back row of the venue and those not present to share a unified VIP 3D perspective, enhancing the user's stereoscopic experience.
Smart Images

Figure CN121967732A_ABST
Abstract
Description
A method and system for cross-scene VR 3D live streaming based on fixed premium seats. Technical Field
[0001] This invention relates to the fields of virtual reality (VR) and live streaming technology, and in particular, to a technology for providing fixed VIP viewing angles for back-row and off-site users in the cultural and entertainment industry (such as concerts and sporting events); specifically, it relates to a cross-scene VR 3D live streaming method and system based on fixed VIP camera positions. Background Technology
[0002] Most existing live video streams are two-dimensional. In live events such as concerts and sporting events, VIP seats, due to their superior view and closer proximity, typically only account for 5% of the total seating capacity. Most audience members can only watch from the back rows. Those in the back rows are often obstructed by crowds or have poor viewing angles, making it impossible to enjoy the close-up experience of VIP seats. Furthermore, non-live viewers are limited to a fixed, two-dimensional view from devices like phones and televisions. This flat image fails to convey the spatial sense of the event (such as the layering of the singer and stage background, the distance between players and the audience), lacking a sense of depth and failing to provide an immersive experience.
[0003] In 2024, the market size of my country's cultural and entertainment live streaming industry exceeded 600 billion yuan. With the increasing demand for "immersive experiences," over 70% of users hope to break through the limitations of traditional flat video perspectives, and 45% of these users are explicitly willing to pay for "high-quality 3D viewing experiences." However, existing technological solutions have not yet achieved large-scale adoption due to high costs and complex operations. The main problems include:
[0004] 1. Existing VR live streaming technology requires multiple dedicated panoramic cameras and complex data stitching equipment, with a single set costing over 50,000 yuan. Deployment requires a professional team for debugging, and users need to purchase high-performance VR headsets costing over 1,000 yuan each. Existing multi-camera VR panoramic live streaming systems are too expensive for ordinary users and small and medium-sized event organizers due to their over-reliance on multi-camera data stitching and dedicated hardware.
[0005] Taking Tencent Cloud's VR live streaming solution as an example, this solution includes three or more panoramic cameras, one dedicated data stitching server, one streaming media distribution server, and a dedicated VR headset. The panoramic cameras must have high resolution (8K or higher) and panoramic shooting capabilities, the stitching server must have multi-view geometric algorithm processing capabilities, and the VR headset must support panoramic image display and head motion tracking. This solution achieves VR panoramic live streaming through the following steps:
[0006] ① Deploy 3 panoramic cameras at different locations at the event site to ensure 360° field of view coverage;
[0007] ②Each camera collects video data of its location in real time and transmits it to a dedicated splicing server via a wired network;
[0008] ③ The server runs a "multi-view geometric stitching algorithm" to synthesize video data from multiple cameras into a 360° panoramic 3D video stream;
[0009] ④ The streaming media distribution server distributes the panoramic 3D video stream to the user's dedicated VR headset via the network;
[0010] ⑤ Users wearing VR headsets can switch between different viewing angles by rotating their heads to achieve "panoramic viewing";
[0011] This solution has the following drawbacks:
[0012] ① The hardware cost is high. The total cost of 3 panoramic cameras + stitching server exceeds 80,000 yuan, which is only suitable for large-scale events. Organizers of small and medium-sized events cannot afford it.
[0013] ②VR headsets typically weigh over 350g, and prolonged use can cause facial pressure and fatigue for users.
[0014] ③ Non-on-site users need to manually turn their heads to adjust the viewing angle, and cannot share the unified "premium view of the VIP seats". Furthermore, the 3D effect is not optimized for fixed viewing angles, resulting in insufficient stereoscopic effect of the picture.
[0015] ④ The deployment is complex and requires professional technicians to adjust the camera position and stitching parameters; ordinary personnel cannot operate it.
[0016] 2. Although some VR live streaming solutions currently support 3D effects, they require the camera to have head motion tracking capabilities, which adds to the complexity and cost of the equipment. Furthermore, non-on-site users cannot share a unified "premium view" and need to manually adjust the view, resulting in a poor experience consistency. Summary of the Invention
[0017] Therefore, the purpose of this invention is to develop and design a cross-scene VR 3D live streaming method and system based on a fixed VIP camera position. It designs a dedicated 3D conversion scheme for a fixed, high-quality viewpoint, constructing a four-level architecture consisting of a fixed VIP camera position, a cloud processing platform, a mobile app, and lightweight VR glasses. Each level of the architecture is connected via 5G, WiFi, and CDN networks, forming a complete closed loop of video acquisition, data processing, 3D conversion, and image display. This conveys the spatial hierarchy of the scene and enables cross-scene VR 3D live streaming. The system provides a low-cost, easily deployable solution. It uses a single ordinary fixed camera to capture the VIP viewpoint, which is then converted into a 3D signal in real time via a mobile app. Combined with lightweight VR glasses, this allows users in the back rows and those not present to share a unified VIP 3D viewpoint, improving the stereoscopic effect of the image and enhancing the user viewing experience.
[0018] This invention provides a cross-scene VR3D live streaming method based on fixed premium seating positions, comprising the following steps:
[0019] S1. Video data from the perspective of the VIP seats is collected by a panoramic camera set in a fixed VIP seat position, and the video data is transmitted to the cloud processing platform via a 5G network.
[0020] S2. A lightweight deep learning model is built based on the MobileDepthNet architecture. After scene feature enhancement training, it is deployed on a cloud processing platform. The deep learning model performs deep analysis on the video data to generate a depth map of the image. Based on the depth map, parallax is calculated to generate data containing 3D information. The data containing 3D information is distributed to the user's mobile APP through the CDN network.
[0021] Specifically, this invention targets fixed-viewpoint videos (without motion tracking or multi-camera data) and employs a monocular depth estimation scheme of "lightweight deep learning model + scene feature enhancement". It simulates the logic of the human eye judging the distance of objects through image details. It utilizes four inherent image features in video frames: "edge contour, color contrast, object occlusion relationship, and texture density", combined with a pre-trained scene model (such as concert / ball game) to infer the "depth value" of each pixel in the image (i.e., the distance between the real object corresponding to that point and the camera), and finally generates a global depth map.
[0022] For example, the correlation logic between image features and depth is shown in Table 1.
[0023] Table 1
[0024]
[0025] In one embodiment of the present invention, a deep learning model is designed based on the MobileDepthNet architecture for lightweight scenarios on mobile devices / cloud platforms. The following three optimizations are made for "fixed-view live streaming" to ensure a balance between real-time performance (processing ≤30ms per frame) and accuracy:
[0026] Input layer optimization: The input resolution is fixed at 1920×1080 (to match the accuracy after downsampling of 4K video and reduce computing power consumption), and "denoising preprocessing" (Gaussian filtering) is performed on the video frames to reduce the interference of light flicker and on-site noise on feature extraction;
[0027] Feature extraction layer optimization: A new "scene feature attention module" has been added, with preset weights for concerts / sports games. For example, in concert scenes, the feature weights of "human outlines and high-saturation areas" are strengthened (weight value 1.2-1.5), while the weights of "irrelevant background areas" are weakened (weight value 0.6-0.8).
[0028] Output layer optimization: The output is a "pixel-level depth value matrix" (1920×1080 resolution), with the depth value range limited to 0.3-20m (covering the viewing range of the VIP seats). "Depth smoothing processing" (mean filtering) is used to eliminate abrupt changes in depth between adjacent pixels to avoid "discontinuities" in the depth map.
[0029] S3. Receive the data containing 3D information via a mobile APP and convert it into a 3D image split for the left and right eyes, then transmit the 3D image to lightweight VR glasses via WiFi;
[0030] S4. The 3D image is displayed in a split-screen format for the left and right eyes using lightweight VR glasses to obtain the same 3D viewing experience as the VIP seats.
[0031] Lightweight VR glasses refer to VR devices that retain only the core functions of dual-screen display, optical magnification, and wireless reception, without complex motion tracking chips or high-performance processors, thus lowering the barrier to entry for users.
[0032] Furthermore, the method for generating a depth map of the image by performing depth analysis on the video data in step S2 includes:
[0033] S21. Perform preprocessing on video frames, including frame cropping and resolution adaptation, noise reduction and enhancement, and ROI region marking.
[0034] S22. The feature extraction module backbone (optimized MobileNetV3) of the deep learning model performs hierarchical feature extraction on the preprocessed video frames, and outputs four types of feature maps: edge feature map, color contrast feature map, occlusion relationship feature map and texture density feature map.
[0035] S23. Input the four types of feature maps into the feature fusion module, and combine them with the pre-trained scene depth mapping model to calculate the depth value of each pixel;
[0036] S24. Generate a depth map based on the depth value of each pixel, and perform post-processing on the depth map, including filtering, depth value cropping and quantization, and alignment with the original video frame, to optimize accuracy and practicality.
[0037] Furthermore, the method for calculating disparity based on the depth map in step S2 includes:
[0038] Based on the pinhole imaging model, the imaging height of an object on the camera's imaging plane is inversely proportional to its depth. Combining this with the difference in visual angle between the human eye's pupillary distance, the basic parallax calculation formula is obtained as follows:
[0039] ;
[0040] Where B is the interpupillary distance, f is the camera focal length, D is the object depth, and P is the pixel size;
[0041] Substitute the depth map pixel by pixel into the basic disparity calculation formula to calculate the initial disparity. ;
[0042] Specifically, molecules It is determined by the interpupillary distance B and the camera focal length f, and is an inherent coefficient reflecting the difference in binocular visual angle. For a fixed camera position and standard interpupillary distance, this value is a constant (e.g., when B = 0.065m and f = 0.015m). ); denominator The parallax is determined by the object's depth D. The closer the object (smaller D), the smaller the pixel size (smaller P), the smaller the denominator, and the larger the parallax d, following the parallax rule of near-large and far-small.
[0043] Due to slight wide-angle distortion at the premium seating positions and a wide depth range (0.3-20m), the initial parallax may exceed the fusion range of the human eye (e.g., 812 pixels in the foreground far exceeds the maximum parallax that the two eyes can fuse). To ensure reasonable parallax, the initial parallax is corrected (to avoid excessive foreground parallax leading to ghosting, and excessive distance parallax leading to a lack of stereoscopic effect). The initial parallax is adjusted to a scene-appropriate parallax that is actually usable through threshold constraints and distortion correction. (To address edge distortion and depth range constraints), the calculation formula for the scene adaptation parallax is as follows:
[0044] ;
[0045] Where k is the distortion correction coefficient. The minimum disparity threshold, This is the maximum disparity threshold.
[0046] When the human eye views an object, the left and right eyes form different perspectives on the same object due to the difference in distance (interpupillary distance). This "perspective difference" is called parallax. Parallax has a fixed inverse relationship with the object's depth: the closer the object (the smaller the depth value), the greater the parallax; the farther the object (the greater the depth value), the smaller the parallax.
[0047] Specifically, unlike existing solutions that directly acquire 3D video using binocular cameras, this invention uses monocular video to convert 3D parallax data. It acquires 2D video using only one camera and calculates the positional deviation (parallax) of the left and right eyes when viewing the same object based on image features (edges, contrast, occlusion relationships). This eliminates the need for two cameras to simulate binocular shooting, reducing hardware costs. Based on the mapping logic of "monocular depth map → binocular parallax", this invention converts the "pixel-level depth value" in the depth map into "left and right eye pixel offset (parallax)" according to the pinhole imaging model and human visual physiological parameters, ensuring that the generated 3D image conforms to the stereoscopic vision habits of the human eye.
[0048] Furthermore, the method for performing layered feature extraction on the preprocessed video frames in step S22 to output four types of feature maps—edge feature map, color contrast feature map, occlusion relationship feature map, and texture density feature map—includes the following:
[0049] The Canny edge detection algorithm is used to calculate the "gradient value" (0-255) of each pixel to generate an edge feature map; for example, the gradient value of the pixels at the edge of the singer's face is concentrated in 60-100, while the gradient value of the pixels at the edge of the background is concentrated in 10-40.
[0050] Calculate the contrast ratio of each pixel in the YCbCr color space (Y is the brightness, and Cb and Cr represent the concentration offset of blue and red, respectively). The calculation formula is: contrast ratio = (maxY - minY) / (maxY + minY), and generate a contrast feature map (value range 0-1).
[0051] By analyzing pixel connectivity, occlusion edges in the image (such as the boundary between the singer and the background) are identified. Occlusion areas are labeled with an occlusion label (value 1), and unocclusion areas are labeled with an unocclusion label (value 0), generating an occlusion relationship feature map.
[0052] The texture richness of the 3×3 neighborhood around each pixel is calculated using Local Binary Pattern (LBP), the proportion of textured pixels is statistically analyzed (value range 0-1), and a texture density feature map is generated.
[0053] In one embodiment of the present invention, the feature processing method in fixed-viewpoint video depth estimation includes:
[0054] I. Edge Contour Processing: Accurately extracts objects and lays the foundation for depth judgment.
[0055] Edge contours are the core basis for distinguishing different objects and judging distance (near objects have clear edges, distant objects have blurred edges). This invention achieves accurate edge extraction through "multi-stage edge detection + adaptive filtering", specifically including:
[0056] 1. Preprocessing: Eliminate interference and enhance edge signals;
[0057] Gaussian noise reduction filtering: To address potential light flicker and camera noise (such as artifacts caused by strong stage lighting at a concert) in fixed-viewpoint videos, a 3×3 Gaussian filter kernel is used to smooth the video frames. The formula is as follows: ,in =1.2 (balancing denoising effect and edge preservation), high-frequency noise is eliminated through convolution operation to avoid noise being misidentified as edges;
[0058] Brightness normalization: Perform "global brightness normalization" on video frames, mapping pixel brightness values (0-255) to the range of [0.1, 0.9] to prevent the edges of overly bright / dark areas (such as backlighting on the stage or shadows in the audience seats) from being masked by extreme brightness values.
[0059] 2. Edge detection: Dual threshold segmentation to preserve valid edges.
[0060] An optimized version of the Canny edge detection algorithm is used to extract edges in three steps:
[0061] Gradient calculation: The gradient magnitude and direction of each pixel are calculated using the Sobel operator (x-direction and y-direction), with the following formula:
[0062] gradient magnitude gradient direction = ;
[0063] in (Horizontal gradient) reflects the vertical edge. (Vertical gradient) reflects the horizontal edge, ensuring that the horizontal (such as the edge of a singer's height) and vertical (such as the edge of the stage width) contours of an object are captured simultaneously;
[0064] Non-maximum suppression: "Local maxima preservation" is performed on the gradient magnitude along the gradient direction - if the gradient magnitude of a pixel is greater than that of its two adjacent pixels along the gradient direction, the pixel is preserved (determined as an edge), otherwise it is suppressed (determined as a non-edge), to avoid the edges from "widening" (such as compressing the edge of a singer's face from 1 pixel wide to 0.5 pixels wide).
[0065] Dual-threshold edge concatenation: Set a high threshold (e.g., 80) and a low threshold (e.g., 30):
[0066] Pixels with gradient magnitude greater than the high threshold are directly identified as "strong edges" (100% preserved, such as the clear facial contours of a singer).
[0067] Pixels with gradient magnitude less than the low threshold are directly classified as "non-edge" (discarded, such as textures with blurred backgrounds).
[0068] Pixels with gradient magnitudes between the two: if they are connected to a "strong edge", they are determined to be a "weak edge" (retained, such as the shallow edge of a singer's clothing folds), otherwise they are discarded, to ensure the continuity and integrity of the edges.
[0069] 3. Post-processing: Edge optimization, adapting to fixed viewpoint scenes.
[0070] Edge length filtering: For a fixed viewpoint (such as the center of a concert stage), set an "effective edge length threshold" (such as ≥10 pixels) and discard edges shorter than this threshold (such as the edges of tiny light spots in the background) to reduce the interference of invalid edges on depth calculation;
[0071] Edge direction weighting: Based on the "Region of Interest (ROI)" (such as the center of the stage) with a fixed viewpoint, the edges within the ROI are given a weight of 1.5 times, and the edges outside the ROI are given a weight of 0.8 times, giving priority to strengthening the edge features of the core area (such as the singer's outline).
[0072] II. Color Contrast Optimization: Enhance brightness / color differences and assist in depth layering:
[0073] Color contrast reflects the "visual prominence" of an object (near objects have vibrant colors and strong contrast, while distant objects have faded colors and weak contrast). This invention optimizes color contrast through "scene-based contrast enhancement + dynamic range adjustment," and the specific method is as follows:
[0074] 1. Color space conversion: from RGB to YCbCr, separating luminance and chrominance;
[0075] The video frames were converted from the RGB color space to the YCbCr color space (Y: luminance channel, Cb: blue difference channel, Cr: red difference channel) because:
[0076] The human eye is far more sensitive to changes in brightness than to changes in chromaticity. The luminance channel (Y) is the core carrier of contrast. After separation, the luminance channel can be optimized separately to avoid color distortion caused by excessive adjustment of the chromaticity channel (such as the preservation of the original color of the singer's costume in a concert).
[0077] Conversion formulas (based on BT.601 standard): Y = 0.299R + 0.587G + 0.114B Cb = -0.1687R - 0.3313G + 0.5B + 128 Cr = 0.5R - 0.4187G - 0.0813B + 128
[0078] 2. Enhanced brightness and contrast: Dynamically adjusts brightness and contrast in different areas to avoid overexposure or underexposure;
[0079] Adaptive histogram equalization (CLAHE) is used instead of traditional global histogram equalization to avoid local overexposure (such as in stage lighting areas) caused by global adjustments. The specific steps are as follows:
[0080] Region division: The luminance channel (Y) is divided into 8×8 sub-blocks (each sub-block is 240×144 pixels, adapted to 1920×1080 resolution), and the histogram of each sub-block is calculated independently;
[0081] Contrast Limit: Set a "contrast threshold" (e.g., 40). If the number of pixels at a certain gray level in the histogram of a sub-block exceeds this threshold, the excess will be evenly distributed to other gray levels to avoid excessive stretching of brightness within the sub-block (e.g., the strong light area on stage) and loss of detail.
[0082] Interpolation splicing: Bilinear interpolation is performed on the equalization results of adjacent sub-blocks to eliminate abrupt changes in brightness at the boundaries of sub-blocks (such as the boundary between the audience seating and the stage), ensuring a natural brightness transition.
[0083] Color contrast optimization: Enhances color differences and highlights foreground objects;
[0084] To perform "saturation enhancement" on the chroma channels (Cb, Cr), the formula is:
[0085] in: is the average value of the Cb and Cr channels of the video frame (reflecting the overall color tone); k is the saturation enhancement coefficient (scene adaptive: k=1.3 for concert scenes, k=1.1 for ball game scenes), ensuring that the colors of foreground objects (such as singers' costumes and players' jerseys) are more vivid, forming a clear difference from the background (such as the blurred audience seats);
[0086] The constraints are: Cb' and Cr' must be within the range of 0-255 to avoid color overflow and distortion.
[0087] 4. Post-processing: Color consistency verification, adaptation to fixed viewing angles.
[0088] Inter-frame color smoothing: Since the color tone of a fixed-view video is stable (such as the stage lighting at a concert remains unchanged), the color mean difference between the current frame and the previous 3 frames is calculated. If the difference is greater than 5 (pixel value), the current frame is corrected using the mean of the previous 3 frames to avoid sudden color interference (such as audience flashes).
[0089] ROI Color Weighting: For ROI areas with a fixed viewing angle (such as the center of the stage), an additional 1.1 times chromaticity enhancement is applied to further strengthen the color contrast of the core area (such as the color layering between the singer and the stage background).
[0090] III. Handling object occlusion relationships: Clearly define the front and back layers and accurately determine the depth order;
[0091] Object occlusion relationships are a "strong constraint" for depth determination (the occluding object must be closer than the occluded object). This invention achieves accurate processing of occlusion relationships through "occlusion edge detection + region marking + depth correction," as detailed below:
[0092] 1. Occlusion Edge Detection: Identifying the Boundary Line Between Foreground and Background Objects
[0093] Based on the edge detection results, occlusion edges are determined by "gradient direction consistency + pixel brightness abrupt change". The steps are as follows:
[0094] Edge orientation clustering: Statistically analyze the gradient direction of edge pixels (e.g., 0°, 45°, 90°, 135°). If the gradient direction of a certain edge is concentrated (e.g., 80%-90% of the pixels have a gradient direction of 90°±10°), and the difference in pixel brightness on both sides of the edge is >30 (pixel value), then it is initially determined to be a "potential occlusion edge" (e.g., the vertical boundary between the singer's body and the stage background).
[0095] Occlusion direction determination: For the regions on both sides of the potential occlusion edge (left side A, right side B), calculate the texture density (LBP value) of the 5×5 neighborhood:
[0096] If the texture density of region A is greater than the texture density of region B, then A is determined to be an occluded object (foreground) and B is an occluded object (background).
[0097] If the texture density of region A is less than the texture density of region B, then B is determined to be an occluder and A is the occluded object. For example, in a concert, the texture density of the singer's body area (A) is 50 (the clothing texture is clear), and the texture density of the background area (B) is 10 (blurred light spots). Then the singer is determined to be an occluder and the background is the occluded object.
[0098] 2. Occlusion area marking: Divide the area into "occluded area" and "occluded area";
[0099] The regions on both sides of the occlusion edge are marked using the "region growing algorithm". The steps are as follows:
[0100] Seed point selection: Select 1 seed point (the pixel with the highest texture density) on the "occluder side" (such as the singer side) of the occlusion edge, and select 1 seed point on the "occluded side" (such as the background side);
[0101] Regional growth rules:
[0102] Occluder region growth: Pixels in the 8-neighborhood around the seed point, if the texture density is ≥ 80% of the mean of the occluder side and the color difference with the seed point is < 15, are classified as occluder regions and marked as "O" (Occluder).
[0103] Occluded area growth: Pixels in the 8-neighborhood around the seed point, if the texture density is ≤ 120% of the mean of the occluded object and the color difference with the seed point is < 15, are classified as occluded areas and marked as "C" (Occluded).
[0104] Boundary correction: Align the boundary of the grown region with the occlusion edge to ensure that the boundary line between the occluded area and the occluded area is completely aligned with the occlusion edge (to avoid mark offset causing depth judgment errors).
[0105] 3. Depth value correction: Adjust the depth based on occlusion relationships to ensure logical consistency;
[0106] Based on the occlusion marker results, "constraint corrections" are applied to the pixel values after depth calculation, including:
[0107] Basic correction: The depth value of the obscured area (C) = the depth value of the obstructing area (O) + Δd (Δd is a fixed offset, set according to the scene: Δd = 0.5m for concerts, Δd = 1m for ball games), ensuring that the obscured object is always farther than the obstructing object; for example: the depth value of the singer (O area) = 1m → the depth value of the obscured stage prop (C area) = 1.5m;
[0108] Multi-layer occlusion handling: If multi-layer occlusion exists (e.g., singer A occludes prop B, prop B occludes background C), then increment Δd according to the occlusion level:
[0109] Depth value of zone A = 1m → Depth value of zone B = 1 + 0.5 = 1.5m → Depth value of zone C = 1.5 + 0.5 = 2m;
[0110] Edge transition correction: For the depth values within a 10-pixel range on both sides of the occluded edge, a "linear transition" is adopted (e.g., the edge pixel depth of area O is 1m, the edge pixel depth of area C is 1.5m, and the middle pixel depth gradually changes from 1m to 1.5m) to avoid "discontinuities" in the depth map (e.g., abrupt changes in depth between the singer and the background).
[0111] 4. Verification and optimization: Ensure that the occlusion relationship is consistent with reality;
[0112] Inter-frame occlusion consistency check: Since the position of objects in a fixed-view video is stable (such as a singer in the center of a concert stage not moving), compare the occlusion markers of the current frame with those of the previous 5 frames. If the marker of a certain area changes from "O" to "C" (or vice versa), it is judged as an incorrect marker and corrected by the majority of the markers of the previous 5 frames.
[0113] Depth validity check: If the depth value of the occluded area is less than the depth value of the occluding object area (violating occlusion logic), the depth value of the occluded area will be forcibly set to "the depth value of the occluding object area + Δd" to ensure that the depth order is absolutely correct.
[0114] Further, step S21 includes:
[0115] S211, Frame Extraction and Resolution Adaptation: From the 4K 60fps video stream transmitted from the fixed premium seat, extract one original video frame every 20ms (synchronized with the transmitted data packets), and downsample the original video frame to 1920×1080 resolution (reducing the computational load of the model while retaining core details).
[0116] S212, Denoising and Enhancement: Gaussian filtering (kernel size=3×3) is used to remove high-frequency noise in video frames (such as noise generated by stage lighting flicker), and histogram equalization is used to improve the contrast between light and dark areas of the image (especially for dark scenes, such as backlit areas in concerts), ensuring that edge and texture features can be effectively extracted.
[0117] S213, ROI Region Marking: Based on a preset region of interest (such as a 3m×2m area in the center of the stage in a concert) from a fixed VIP seat perspective, mark the ROI region of the image. The ROI region will be processed first during subsequent feature extraction (to reduce the computational power consumption of irrelevant areas).
[0118] The fixed VIP viewing angle refers to a fixed shooting angle that is consistent with the viewing height (1.2-1.5m) and viewing direction (directly facing the core area of the event, such as the center of the concert stage or the center line of the ball game field) of the VIP audience at the event. The coverage of the angle is based on the standard of being able to fully present the key information of the core area (such as covering the singer's performance area of 3m×2m in a concert, and covering the midfield area of a ball game of 10m×8m), which is different from the existing VR panoramic 360° perspective without a fixed focus.
[0119] Further, step S23 includes:
[0120] S231. Feature weight fusion: Based on the scene preset weights (edge features 30%, color contrast 25%, occlusion relationship 25%, texture density 20%), the four types of feature values of each pixel are weighted and summed to obtain a comprehensive feature value (value range 0-1).
[0121] S232. Depth Mapping: The comprehensive feature values are converted into actual depth values using a non-linear mapping function (an S-shaped curve fitted based on the training data); an example of the conversion formula is as follows:
[0122] Depth value (m) = 20 / (1 + e^(5×(comprehensive feature value - 0.5)))
[0123] (Explanation: The closer the comprehensive feature value is to 1, the closer the depth value is to 0.3m; the closer the comprehensive feature value is to 0, the closer the depth value is to 20m, matching the viewing range of the premium seats.)
[0124] S233, Occlusion area depth correction: For pixels marked with occlusion labels, force the depth value of the pixel = occlusion depth value + 0.5m (e.g., occlusion depth value 1m → occluded area depth value 1.5m) to avoid logical contradictions.
[0125] Further, step S24 includes:
[0126] S241, Depth Smoothing: Apply a 5×5 mean filter to the generated depth map to eliminate abrupt depth changes between adjacent pixels (such as depth discontinuities at the boundary between the singer and the background), making the depth transition more consistent with real physical laws.
[0127] S242, Depth Value Cropping and Quantization: Crops the depth value to 0.3-20m (pixels outside the range are forcibly set to 0.3m or 20m) and quantizes it into an 8-bit integer (value range 0-255, corresponding to a linear mapping of 0.3-20m), reducing data storage and subsequent transmission bandwidth;
[0128] S243. Align the depth map with the original video frame: Align the quantized depth map with the original 1920×1080 video frame at the pixel level (ensuring that each pixel in the depth map corresponds one-to-one with the pixel in the original video frame), providing accurate input for subsequent parallax generation.
[0129] Furthermore, the method for receiving the data containing 3D information via a mobile APP and converting it into a split-screen 3D image for the left and right eyes in step S3 includes:
[0130] The VR3D data format for fixed-viewpoint 3D live streaming is designed by encapsulating the raw video stream, pixel-level parallax data, and auxiliary control information to ensure that the mobile app can efficiently parse and generate 3D images that conform to human visual habits. The VR3D data structure adopts a frame-level encapsulation + block-level storage design to avoid overall parsing failure due to single-frame data corruption. The specific structure is shown in Table 2.
[0131] Table 2
[0132]
[0133]
[0134] The mobile app completes a five-step process: reading VR3D data, verifying the format, parsing parameters, extracting frame data, and restoring parallax. It accurately separates the video stream and parallax data from the .vr3d format, laying the foundation for generating 3D images that conform to human visual habits.
[0135] The steps involved in parsing VR3D format data are as follows:
[0136] Step 1, Data Reception and Initialization:
[0137] The app receives .vr3d format data from CDN nodes via the HTTP-FLV protocol, employing a "segmented reception + caching" strategy—every 1MB of data received is stored in the phone's local cache (default path: / Android / data / uni.yizhizh.com / cache / ) to avoid memory overflow;
[0138] Read the first 8 bytes of data. If it is not equal to "VR3D_2025", immediately trigger a "format error" prompt (such as a pop-up window "Unsupported file format, please update the APP"). If the format is correct, continue reading the "basic parameter block offset" in the file header to locate the starting address of the basic parameter block.
[0139] Step 2, Basic Parameter Analysis and Configuration:
[0140] Parameter reading: The parameters "video resolution (W×H), frame rate (F), parallax bit depth (B), maximum parallax threshold (D_max), and color space (C)" are parsed sequentially from the basic parameter block and stored in the "global configuration variables" of the APP.
[0141] In one embodiment of the present invention, W=1920, H=1080, F=60, B=8, D_max=35, C=0 (RGB) are obtained through parsing. The APP automatically sets the video rendering size to 1920×1080, and the parallax calculation range is constrained to 0-35 pixels.
[0142] Compatibility and Adaptation: If the current version of the APP does not support the parsed format (e.g., the APP supports V1.0 but parses V2.0), a pop-up message will appear saying "The APP needs to be updated to the latest version to support this format", and the parsing process will be terminated.
[0143] Step 3: Frame data block location and reading:
[0144] Frame index construction: The APP obtains the total number of frames N from the "Total Frames" field in the file header, traverses all frame data blocks, extracts the "frame number, timestamp, video data length, and parallax data length" of each frame, and constructs a "frame index table" (stored in memory) to facilitate quick location by frame number or timestamp in the future.
[0145] Frame data extraction: Based on the user's viewing progress (e.g., watching frame 100), find the starting address of frame 100 from the frame index table, first read the "video data length (L_v)", then extract L_v bytes of video data; next, read the "parallax data length (L_d)" and extract L_d bytes of parallax data;
[0146] Data verification: Calculate the CRC32 checksum of the captured frame data (video data + parallax data) and compare it with the "frame checksum" in the frame data block. If they match, proceed to the next step of parsing; if they do not match, the APP sends a "frame retransmission request" (carrying the frame sequence number) to the cloud until the correct data is obtained.
[0147] Step 4: Video data decoding and parallax data restoration:
[0148] Video decoding: The APP calls the phone's hardware decoder (such as MediaCodec for Android and AVFoundation for iOS) to decode the H.265 format video data into 2D video frames in RGB / YCbCr format (resolution consistent with the basic parameter block), and stores them as "raw video frame matrix" (dimension: H×W×3, where 3 represents the three RGB channels).
[0149] Parallax data restoration: If the parallax data bit depth B=8, convert the parallax data binary stream into an "8-bit unsigned integer matrix" (dimension: H×W), where each element represents the parallax value d of the corresponding pixel (0≤d≤D_max);
[0150] If a "parallax compression flag" exists (a new field added to the basic parameter block, 1 represents compression, 0 represents no compression), it is first decompressed by "run-length encoding (RLE)" - for example, compressed data "35,5" represents that the parallax value of 5 consecutive pixels is 35, which is restored to [35,35,35,35,35];
[0151] Data alignment: Ensure that the pixel positions of the "original video frame matrix" and the "disparity matrix" correspond one-to-one (i.e., the pixel in the i-th row and j-th column of the video frame corresponds to the disparity value in the i-th row and j-th column of the disparity matrix) to avoid misalignment during subsequent 3D generation.
[0152] After acquiring the "raw video frames + parallax matrix," the app simulates the difference in binocular vision in the human eye, converting the 2D video frames into 3D images suitable for VR glasses display through four steps: "parallax allocation → left and right eye frame generation → distortion correction → image output." The process of generating 3D images includes the following steps:
[0153] Step 1, Disparity Assignment (Determine the direction and amount of left and right eye offset):
[0154] Based on the principle of human binocular vision, the left eye observes objects from a perspective biased to the right, while the right eye observes them from a perspective biased to the left. Therefore, it is necessary to perform a "directional offset" on the pixels of the original video frame. The offset amount is proportionally allocated based on the parallax value, as shown in the following formula:
[0155] ,
[0156] ;
[0157] Where d is the disparity value of a pixel in the disparity matrix (0≤d≤35 pixels, from the parsed disparity matrix);
[0158] This is the rightward offset of the pixel in the left-eye frame (in pixels), rounded down (e.g., when d=35). =17);
[0159] This is the leftward offset of the pixel in the right-eye frame (unit: pixels), such as when d=35. =18;
[0160] In one embodiment of the present invention, the disparity value d=35 of the "singer's face pixel" in the original video frame is calculated to be d_left=17 and d_right=18. That is, the pixel is shifted 17 pixels to the right in the left eye frame and 18 pixels to the left in the right eye frame, ensuring that the total difference between the left and right eye viewing angles is 35 pixels, which meets the disparity requirement.
[0161] Step 2, Generation of left and right eye frames (pixel offset and boundary padding):
[0162] Left eye frame generation:
[0163] Initialize the "left eye frame matrix" (dimensions: H×W×3, consistent with the original video frames);
[0164] For each pixel (i,j) in the original video frame (where i is the row number and j is the column number), calculate its new column number in the left-eye frame: = j + ;
[0165] like If W (not exceeding the right boundary of the screen), then the RGB value of the original pixel (i,j) is assigned to the left eye frame (i,j_left);
[0166] like If the pixel value is greater than or equal to W (exceeding the right boundary), the "edge pixel copying" strategy is adopted - the RGB value of the rightmost pixel (i, W-1) of the original frame is assigned to the left eye frame (i, j_left - W) to avoid black borders;
[0167] Right eye frame generation:
[0168] Initialize the "Right Eye Frame Matrix" (dimensions: H×W×3);
[0169] For each pixel (i,j) in the original video frame, calculate its new column number in the right-eye frame: ;
[0170] like If the pixel does not exceed the left edge of the frame, then the RGB value of the original pixel (i,j) is assigned to the right eye frame (i,j_right).
[0171] like (Exceeding the left boundary), the same "edge pixel copying" method is used - the RGB value of the leftmost pixel (i,0) of the original frame is assigned to the right eye frame (i,j_right + W).
[0172] To avoid "holes" after adjacent pixels are shifted (such as when a pixel is shifted and no pixel is used to fill its original position), the mobile app uses "bilinear interpolation to fill the unassigned pixels in the left eye frame and the right eye frame. For example, when the left eye frame (i,j) is unassigned, the average RGB value of the four surrounding assigned pixels is used as the fill value to ensure a smooth image.
[0173] Step 3, Distortion Correction (Adapting to Lightweight VR Glasses Optical Systems):
[0174] Because the Fresnel lenses in VR glasses suffer from "barrel distortion" (edge pixels are stretched), if the left and right eye frames are output directly, users will see distorted 3D images. Therefore, a polynomial distortion correction algorithm is used to correct this distortion. The specific steps are as follows:
[0175] (1) Obtain glasses parameters: The APP reads the "distortion coefficients" (e.g., k1=-0.3, k2=0.1, k3=-0.05, representing third-order distortion coefficients) and "optical center" (e.g., (W / 2, H / 2), i.e., the center of the screen) of the currently connected VR glasses from the "VR device parameter library" (which contains distortion parameters of mainstream VR glasses models).
[0176] (2) Distortion coordinate calculation: For each pixel (x, y) (screen coordinates) in the left and right eye frames, calculate the ideal coordinates of the pixel (x, y) before distortion using the following formula.
[0177] ,
[0178] ,
[0179] ;
[0180] Where (x,y) are the pixel coordinates (screen coordinates, range 0≤x) of the corrected left and right eye frames. <W,0≤y<H);
[0181] ( ) represents the optical center coordinates of the VR glasses (usually the center of the screen, such as (960, 540) corresponding to a 1920×1080 resolution);
[0182] r is the normalized distance from the pixel to the optical center (dimensionless, range 0-1).
[0183] k1, k2, and k3 are distortion coefficients (determined by the VR glasses hardware; for example, k1 = -0.3 represents first-order barrel distortion correction).
[0184] ( The coordinates are the ideal pixel coordinates before distortion (which need to be mapped to the coordinate range of the original left and right eye frames);
[0185] The app calls OpenCV's remap function to remap the left and right eye frames according to the mapping relationship between "ideal coordinates → screen coordinates" to generate "distortion-corrected left and right eye frames" - for example, assigning the RGB values of the pixels at the (x_distort, y_distort) position in the original left and right eye frames to the (x, y) position in the corrected frame to ensure that the user sees a distortion-free 3D image through VR glasses.
[0186] Step 4: Synchronize 3D image output with VR glasses:
[0187] (1) Image format adaptation: The distortion-corrected left and right eye frames are spliced together according to the "3D display format" supported by the VR glasses (the present invention uses the "side-by-side" format by default) - that is, the left eye frame is placed on the left and the right eye frame is placed on the right to generate a "single frame 3D image" (resolution: H×(2W), such as 1080×3840);
[0188] (2) Low latency transmission: The APP transmits the "single frame 3D image" to the VR glasses via WiFi 6 (or Bluetooth 5.2, depending on the user's device). During the transmission process, "UDP protocol + frame sequence number marking" is used to ensure that each frame of data arrives in order, and the transmission latency is controlled within ≤20ms.
[0189] (3) Audio-visual synchronization: The APP reads the "timestamp" in the frame data block and compares it with the audio playback timestamp of the temple speaker. If the video frame timestamp is ahead of the audio timestamp by more than 50ms, the video transmission is paused for 10ms; if it is behind by more than 50ms, the video decoding speed is accelerated (such as skipping 1 non-key frame) to ensure that the audio-visual synchronization error is ≤30ms, so as to avoid the user's discomfort of "the picture and the sound are out of sync".
[0190] This invention optimizes the processing method for VR3D parsing and 3D image generation in weak network environments, mainly including:
[0191] If the app detects that the network bandwidth is less than 20Mbps (judged by the average rate of received data), it automatically requests "low bitrate .vr3d data" from the cloud (e.g., reducing the video resolution from 1920×1080 to 1280×720 and the parallax data bit depth from 8 bits to 6 bits) to dynamically reduce the bitrate; at the same time, it simplifies "bilinear interpolation" to "nearest neighbor interpolation" when generating 3D images to reduce the phone's computing power consumption;
[0192] When the app receives .vr3d data, it simultaneously caches the parsed "left and right eye frames" to the phone's local storage (the default maximum cache size is 1 hour of content). When the network is interrupted, it automatically switches to "offline mode" and reads the locally cached left and right eye frames to generate 3D images, thus avoiding viewing interruptions.
[0193] This invention optimizes the design of an abnormal scenario handling method, including:
[0194] If the parallax data parsing of a certain frame fails (e.g., the wrong data is still obtained after 3 retransmissions), the APP automatically enables "temporary parallax generation" - based on the parallax matrix of the previous 3 frames, the parallax matrix of the current frame is generated through "inter-frame interpolation" (e.g., current frame parallax = (parallax of the previous 1 frame + parallax of the previous 2 frames + parallax of the previous 3 frames) / 3), to ensure that the 3D effect is not interrupted;
[0195] If the app detects that the phone's CPU usage is greater than 90% (obtained through the system API), it will reduce the "distortion correction order" (e.g., from third-order distortion correction to first-order) and reduce the frame rate from 60fps to 30fps to prioritize the smoothness of 3D image generation and avoid screen stuttering.
[0196] The logical chain for .vr3d parsing and 3D generation in this invention is as follows: .vr3d format data (file header + basic parameter block + frame data block) → APP parsing (format verification → parameter reading → frame data extraction → video decoding + parallax restoration) → 3D generation (parallax allocation → left and right eye frame offset → distortion correction → split-screen splicing) → WiFi transmission to VR glasses → user watching 3D live broadcast. The entire process is designed around "low latency, high compatibility, and weak network adaptation" to ensure that both on-site back-row users and off-site users can obtain a stable and distortion-free 3D viewing experience.
[0197] This invention also provides a cross-scene VR3D live streaming system based on fixed premium-seat camera positions, used to execute the cross-scene VR3D live streaming method based on fixed premium-seat camera positions as described above, including:
[0198] Fixed VIP seat video data acquisition module: used to acquire video data from the VIP seat perspective through a panoramic camera set in the fixed VIP seat, and transmit the video data to the cloud processing platform through a 5G network;
[0199] The cloud processing platform's deep analysis and parallax calculation module is used to build a lightweight deep learning model based on the MobileDepthNet architecture, perform scene feature enhancement training, and then deploy it on the cloud processing platform. The deep learning model performs deep analysis on the video data to generate a depth map of the image, performs parallax calculation based on the depth map, generates data containing 3D information, and distributes the data containing 3D information to the user's mobile APP through the CDN network.
[0200] Mobile APP 3D data to 3D image module: used to receive the data containing 3D information through a mobile APP and convert it into a 3D image split for the left and right eyes, and transmit the 3D image to lightweight VR glasses via WiFi;
[0201] VR glasses 3D image display module: used to display the 3D image in a split-screen format for the left and right eyes through lightweight VR glasses, so as to obtain the same 3D viewing experience as the VIP seats.
[0202] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of the cross-scene VR3D live streaming method based on a fixed VIP seat as described above.
[0203] The present invention also provides a computer device, the computer device including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the steps of the cross-scene VR3D live streaming method based on fixed VIP seats as described above.
[0204] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0205] This invention provides a cross-scene VR 3D live streaming and system based on fixed VIP seating positions. It designs a dedicated 3D conversion scheme for fixed, high-quality viewpoints, constructing a four-level architecture: fixed VIP seating position—cloud processing platform—mobile APP—lightweight VR glasses. Each level is connected via 5G, WiFi, and CDN networks, forming a complete closed loop of video acquisition—data processing—3D conversion—image display. This conveys the spatial hierarchy of the scene and enables cross-scene VR 3D live streaming. A single ordinary fixed camera captures the VIP viewing angle, which is then converted into a 3D signal in real time via the mobile APP. Combined with lightweight VR glasses, this allows users in the back rows and those not present to share a unified VIP 3D viewpoint, improving the stereoscopic effect and enhancing the user viewing experience. Furthermore, it is cost-effective, easy to deploy, and highly adaptable, with broad application prospects. Attached Figure Description
[0206] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.
[0207] In the attached diagram:
[0208] Figure 1 is a flowchart of a cross-scene VR3D live streaming method based on a fixed VIP seat according to an embodiment of the present invention. Detailed Implementation
[0209] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and products consistent with some aspects of this disclosure as detailed in the appended claims.
[0210] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0211] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."
[0212] The embodiments of the present invention will be described in further detail below.
[0213] This invention provides a cross-scene VR3D live streaming method based on a fixed premium camera position, as shown in Figure 1, including the following steps:
[0214] S1. Video data from the perspective of the VIP seats is collected by a panoramic camera set in a fixed VIP seat position, and the video data is transmitted to the cloud processing platform via a 5G network.
[0215] S2. A lightweight deep learning model is built based on the MobileDepthNet architecture. After scene feature enhancement training, it is deployed on a cloud processing platform. The deep learning model performs deep analysis on the video data to generate a depth map of the image. Based on the depth map, parallax is calculated to generate data containing 3D information. The data containing 3D information is distributed to the user's mobile APP through the CDN network.
[0216] The method for generating a depth map of the video image by performing depth analysis on the video data includes:
[0217] S21. Preprocessing of the video frames includes frame cropping and resolution adaptation, noise reduction and enhancement, and ROI region marking; Step S21 includes:
[0218] S211. From the 4K 60fps video stream transmitted from the fixed VIP seat, extract one original video frame every 20ms, and downsample the original video frame to a resolution of 1920×1080.
[0219] S212. Use Gaussian filtering (kernel size=3×3) to remove high-frequency noise in video frames, and use histogram equalization to improve the contrast of the image, ensuring that edge and texture features can be effectively extracted.
[0220] S213. Based on the preset region of interest (a 3m×2m area in the center of the stage during a concert) from the fixed VIP seat perspective, mark the ROI region of the image, and prioritize the processing of the ROI region during subsequent feature extraction.
[0221] S22. The feature extraction module backbone of the deep learning model performs hierarchical feature extraction on the preprocessed video frames and outputs four types of feature maps: edge feature map, color contrast feature map, occlusion relationship feature map and texture density feature map.
[0222] Methods for performing hierarchical feature extraction on preprocessed video frames to output four types of feature maps: edge feature map, color contrast feature map, occlusion relationship feature map, and texture density feature map include:
[0223] The Canny edge detection algorithm is used to calculate the gradient value (0-255) of each pixel to generate an edge feature map;
[0224] Calculate the contrast ratio of each pixel in the YCbCr color space using the formula: contrast ratio = (maxY - minY) / (maxY + minY), and generate a contrast feature map (value range 0-1).
[0225] By analyzing pixel connectivity, occlusion edges in the image (such as the boundary between the singer and the background) are identified. Occlusion areas are labeled with an occlusion label (value 1), and unocclusion areas are labeled with an unocclusion label (value 0), generating an occlusion relationship feature map.
[0226] The texture richness of the 3×3 neighborhood around each pixel is calculated using Local Binary Pattern (LBP), the proportion of textured pixels is statistically analyzed (value range 0-1), and a texture density feature map is generated.
[0227] S23. Input the four types of feature maps into the feature fusion module, and combine them with a pre-trained scene depth mapping model (such as one trained using over 100,000 concert / sports game annotation data) to calculate the depth value of each pixel; Step S23 includes:
[0228] S231. Based on the scene preset weights (edge features 30%, color contrast 25%, occlusion relationship 25%, texture density 20%), the four types of feature values of each pixel are weighted and summed to obtain a comprehensive feature value (value range 0-1).
[0229] S232. The comprehensive feature values are converted into actual depth values through a nonlinear mapping function (an S-shaped curve fitted based on the training data).
[0230] S233. For pixels marked with occlusion tags, force the depth value of the pixel to be equal to the depth value of the occlusion object plus 0.5m to avoid logical contradictions.
[0231] S24. Generate a depth map based on the depth value of each pixel, and perform post-processing on the depth map including filtering, depth value cropping and quantization, and alignment with the original video frame. Step S24 includes:
[0232] S241. Apply a 5×5 mean filter to the generated depth map to eliminate abrupt changes in depth between adjacent pixels, making the depth transition more consistent with real physical laws.
[0233] S242. Crop the depth value to 0.3-20m (force pixels outside the range to be set to 0.3m or 20m) and quantize it into an 8-bit integer (value range 0-255, corresponding to a linear mapping of 0.3-20m) to reduce data storage and subsequent transmission bandwidth;
[0234] S243. Align the quantized depth map with the original 1920×1080 video frame at the pixel level to provide accurate input for subsequent parallax generation.
[0235] The method for calculating disparity based on the depth map includes:
[0236] Based on the pinhole imaging model, the imaging height of an object on the camera's imaging plane is inversely proportional to its depth. Combining this with the difference in visual angle between the human eye's pupillary distance, the basic parallax calculation formula is obtained as follows:
[0237] ;
[0238] Where B is the interpupillary distance, f is the camera focal length, D is the object depth, and P is the pixel size;
[0239] Substitute the depth map pixel by pixel into the basic disparity calculation formula to calculate the initial disparity. ;
[0240] Due to slight wide-angle distortion and a wide depth range in the premium seating position, the initial parallax may exceed the fusion range of the human eye. To ensure reasonable parallax, the initial parallax is corrected by applying threshold constraints and distortion correction to achieve a scene-appropriate parallax suitable for actual use. The formula for calculating the scene adaptation parallax is:
[0241] ;
[0242] Where k is the distortion correction coefficient. The minimum disparity threshold, This is the maximum disparity threshold.
[0243] This embodiment targets fixed-viewpoint videos (without motion tracking or multi-camera data). It employs a monocular depth estimation scheme combining a lightweight deep learning model and scene feature enhancement. This scheme simulates the human eye's logic of judging the distance of objects through image details. It utilizes four inherent image features in video frames—edge contours, color contrast, object occlusion relationships, and texture density—combined with a pre-trained scene model (such as a concert / sports game) to infer the "depth value" of each pixel in the image (i.e., the distance between the real object at that point and the camera), ultimately generating a global depth map.
[0244] The actual test verification of processing speed, depth accuracy, and anti-interference capability in this embodiment of the invention is as follows:
[0245] On a cloud server (single GPU: NVIDIA A100), the processing time for each 1920×1080 video frame is 22-28ms (average 25ms), which meets the real-time requirement of "≤30ms per frame processing" for 4K 60fps video.
[0246] In actual tests at concert scenes, the depth estimation error rate for near objects (0.5-3m) is ≤8% (e.g., a singer at an actual distance of 1m shows a depth of 0.92-1.08m on the depth map); the error rate for far objects (5-10m) is ≤15% (e.g., a background at an actual distance of 8m shows a depth of 6.8-9.2m on the depth map), which fully meets the accuracy requirements for 3D parallax generation.
[0247] To address common interference factors in live streaming scenarios, strong anti-interference capabilities are achieved through preprocessing and model optimization:
[0248] Lighting interference: Stage lighting flicker (frequency 50Hz) → Preprocessing stage "inter-frame smoothing" (taking the average of the previous and next 3 frames), depth value fluctuation ≤5%;
[0249] Dynamic interference: When the singer moves quickly (such as jumping), the model avoids "ghosting" in the depth map by "inter-frame depth prediction" (motion compensation based on the depth value of the previous frame). The depth error rate in dynamic scenes is ≤10%.
[0250] S3. Receive the data containing 3D information via a mobile APP and convert it into a 3D image split for the left and right eyes, then transmit the 3D image to lightweight VR glasses via WiFi;
[0251] S4. The 3D image is displayed in a split-screen format for the left and right eyes using lightweight VR glasses, providing a 3D viewing experience consistent with that of a premium seat. The lightweight VR glasses retain only the core functions of dual-screen display, optical magnification, and wireless reception. They do not contain complex motion tracking chips or high-performance processors, and weigh ≤130g with a cost ≤300 yuan. This distinguishes them from existing professional VR headsets that integrate multiple sensors, weigh over 350g, and cost over 1000 yuan, thus lowering the barrier to entry for users.
[0252] This invention also provides a cross-scene VR3D live streaming system based on fixed premium-seat camera positions, used to execute the cross-scene VR3D live streaming method based on fixed premium-seat camera positions as described above, including:
[0253] Fixed VIP seat video data acquisition module: used to acquire video data from the VIP seat perspective through a panoramic camera set in the fixed VIP seat, and transmit the video data to the cloud processing platform through a 5G network;
[0254] The cloud processing platform's deep analysis and parallax calculation module is used to build a lightweight deep learning model based on the MobileDepthNet architecture, perform scene feature enhancement training, and then deploy it on the cloud processing platform. The deep learning model performs deep analysis on the video data to generate a depth map of the image, performs parallax calculation based on the depth map, generates data containing 3D information, and distributes the data containing 3D information to the user's mobile APP through the CDN network.
[0255] Mobile APP 3D data to 3D image module: used to receive the data containing 3D information through a mobile APP and convert it into a 3D image split for the left and right eyes, and transmit the 3D image to lightweight VR glasses via WiFi;
[0256] VR glasses 3D image display module: used to display the 3D image in a split-screen format for the left and right eyes through lightweight VR glasses, so as to obtain the same 3D viewing experience as the VIP seats.
[0257] The complete workflow of this embodiment in a concert setting is as follows:
[0258] 1. Preliminary Deployment: Before the event begins, staff will install fixed VIP seats in the VIP area, complete the viewing angle adjustment and network connection tests to ensure normal video capture and transmission; the cloud processing platform will start all functional modules and complete the connection with the camera positions and CDN network;
[0259] 2. Live Stream Start: After the concert begins, fixed VIP seats will capture video streams and transmit them to the cloud; the cloud processing platform will generate a special format containing parallax data and distribute it to users' mobile phones via CDN;
[0260] 3. User Viewing: Users open the mobile app, select the corresponding live stream channel, complete the 3D conversion and depth-of-field adjustment, and then connect the VR glasses. After wearing the glasses, users can enjoy a 3D viewing experience as if they were "sitting in the first row of the VIP seats." Users in the back rows do not need to worry about being blocked, and users not present can enjoy the same experience remotely. 4. End of Live Stream: After the event ends, the fixed VIP seats will stop capturing data, the cloud processing platform will shut down the relevant modules, and users will disconnect the app from the VR glasses to complete the viewing.
[0261] This embodiment's architecture eliminates the need for complex motion tracking equipment, with clearly defined functional divisions across all modules, reducing hardware costs and deployment complexity. A single fixed camera captures the VIP seating view, which is then converted into a 3D signal in real-time via a mobile app. Combined with lightweight VR glasses, this allows both on-site and off-site users to share a unified VIP 3D view. Camera costs are reduced to the level of a standard camera (≤5000 RMB), and VR glasses cost ≤300 RMB and weigh ≤130g, significantly lowering the hardware barrier. The total live stream latency is ≤150ms, meeting real-time viewing requirements. All users (regardless of whether they are on-site or not) experience the same 3D stereoscopic effect as those in the VIP seating area. The operation is simple, requiring no specialized knowledge, and is suitable for users of all ages.
[0262] Key data regarding latency control, AI depth estimation accuracy, and hardware cost in this embodiment are as follows:
[0263] (1) Delay control: The total system delay is ≤150ms, of which “camera station → cloud (≤50ms)” depends on the actual transmission rate of 5G network (1Gbps+, UDP protocol packet loss rate ≤1%); “cloud → APP (≤80ms)” depends on CDN node coverage density (300+ nodes nationwide, access delay in third- and fourth-tier cities ≤80ms); “APP → VR glasses (≤20ms)” depends on the actual transmission delay of WiFi 6 protocol (≤15ms). The above data are all from the actual test results of laboratory simulated activity scenarios (200 people online at the same time).
[0264] (2) AI depth estimation accuracy: The lightweight model used has a depth estimation error rate of ≤5% on the public datasets KITTI (autonomous driving scenario) and NYU Depth V2 (indoor scenario). After optimization for concert and ball game scenarios, the depth recognition accuracy of "near subject (singer, player)" is ≥92%, which meets the 3D stereoscopic requirements. In actual tests, compared with the existing multi-camera stitching scheme, the difference in users' subjective stereoscopic sense score is ≤10%.
[0265] (3) Hardware cost: The market price of a regular fixed camera (4K resolution) is ≤4,000 yuan, the 5G transmission module is ≤1,000 yuan, and the material cost of lightweight VR glasses (dual LCD + WiFi) is ≤200 yuan. The overall hardware cost is 93% lower than the existing multi-camera VR solution (3 8K panoramic cameras + server, cost over 80,000 yuan). The cost data comes from supply chain quotations and small-batch trial production calculations.
[0266] Typical examples of scene adaptation for concerts are as follows:
[0267] For small and medium-sized concerts (venue capacity 5,000 people, 50 VIP seats), one fixed camera was deployed in the central area of the VIP seats, covering a stage width of 6m and a height of 3m. 4,950 audience members in the back row watched the concert through "mobile APP + VR glasses". Non-on-site users accessed the concert remotely through the APP. When the 3D stereoscopic effect was adjusted to level 3, users reported that "the spatial layers of singers, bands and audiences can be clearly distinguished", the delay was ≤150ms, and there was no obvious audio-visual asynchrony. (2) Football match scenario: For football leagues (small and medium-sized stadiums, capacity 10,000 people, 100 VIP seats), the camera was deployed in the center line of the VIP seats, covering a 20m×15m midfield area of the stadium. When users watched the concert, the depth of field was adjusted to level 2. The actual test showed that "the distance between the players and the ground and the audience seats when they ran can be clearly perceived". In a weak network environment (4G network, bandwidth 20Mbps), the bit rate was automatically reduced to 1080P, and the stuttering rate was ≤2%, which met the real-time viewing needs.
[0268] This embodiment has the following advantages compared to existing multi-camera VR live streaming systems:
[0269] 1. Cost advantage: Existing multi-camera VR live streaming systems (such as Tencent Cloud solutions) cost over 80,000 yuan per set. The total cost of the "fixed camera + lightweight VR glasses" in this embodiment is ≤5,500 yuan, a cost reduction of 93%, which can be easily afforded by small and medium-sized event organizers.
[0270] 2. Consistent Experience Advantage: Existing VR live streaming requires users to manually turn their heads to switch perspectives, resulting in significant differences in perspectives among different users. In this embodiment, all users share a unified premium 3D perspective, improving the consistency of the experience by 100%, and the stereoscopic effect is adjustable to adapt to different content viewing needs.
[0271] 3. User coverage advantage: Existing VR live streaming only covers on-site users or requires off-site users to pay separately. This embodiment supports both on-site back row users and off-site users, expanding the user coverage by more than 2 times. Moreover, users do not need to purchase high-performance equipment; ordinary mobile phones and lightweight VR glasses are sufficient for use.
[0272] 4. Advantages in ease of operation: Existing VR live streaming requires professional personnel for deployment and debugging, and the operation is complicated for users. In this embodiment, the camera deployment only requires manual adjustment of the bracket, and the user operation only requires "opening the APP + connecting VR". It also supports voice control, so even elderly users can quickly get started, greatly reducing the difficulty of operation.
[0273] 5. Real-time advantage: Existing VR live streaming often has a total latency of over 300ms due to multi-camera splicing. This embodiment has a total latency of ≤150ms, which meets the real-time requirements of live streaming such as concerts and sports games. Moreover, it can automatically reduce the bit rate in weak network environments, reducing the stuttering rate by 80%.
[0274] The technical solution of the present invention has been described in conjunction with preferred embodiments. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions resulting from these changes or substitutions will all fall within the scope of protection of the present invention.
[0275] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A cross-scene VR 3D live streaming method based on fixed premium seating positions, characterized in that, Includes the following steps: S1. Video data from the VIP seating area is captured by a panoramic camera positioned at a fixed VIP seating location, and transmitted to a cloud processing platform via a 5G network. S2. A lightweight deep learning model is built based on the MobileDepthNet architecture, trained for scene feature enhancement, and then deployed to the cloud processing platform. The deep learning model performs depth analysis on the video data to generate a depth map of the image. Parallax is calculated based on the depth map to generate data containing 3D information, which is then distributed to a user's mobile app via a CDN network. S3. The mobile app receives the data containing 3D information and converts it into a split-screen 3D image for both eyes. The 3D image is then transmitted to lightweight VR glasses via WiFi. S4. The lightweight VR glasses display the 3D image in a split-screen format for both eyes, providing a 3D viewing experience consistent with that of the VIP seating area.
2. The cross-scene VR3D live streaming method based on fixed premium seating positions according to claim 1, characterized in that, The method for generating a depth map of the video data through deep analysis in step S2 includes: S21, preprocessing the video frames by frame cropping and resolution adaptation, noise reduction and enhancement, and ROI region marking; S22, performing hierarchical feature extraction on the preprocessed video frames through the feature extraction module of a deep learning model, outputting four types of feature maps: edge feature map, color contrast feature map, occlusion relationship feature map, and texture density feature map; S23, inputting the four types of feature maps into the feature fusion module, combining them with a pre-trained scene depth mapping model, and calculating the depth value of each pixel; S24, generating a depth map based on the depth value of each pixel, and performing postprocessing on the depth map by filtering, depth value cropping and quantization, and alignment with the original video frames.
3. The cross-scene VR3D live streaming method based on fixed premium seating positions according to claim 2, characterized in that, The method for calculating disparity based on the depth map in step S2 includes: based on the pinhole imaging model, the imaging height of an object on the camera's imaging plane is inversely proportional to its depth. Combining this with the difference in visual angle between human pupillary distances, the basic disparity calculation formula is obtained as follows: Where B is the interpupillary distance, f is the camera focal length, D is the object depth, and P is the pixel size; the depth map is substituted pixel by pixel into the basic disparity calculation formula to calculate the initial disparity. Due to slight wide-angle distortion and a wide depth range in the premium seating position, when the initial parallax exceeds the fusion range of the human eye, it is corrected to ensure reasonable parallax. The initial parallax is adjusted to a scene-appropriate parallax through threshold constraints and distortion correction. The formula for calculating the scene adaptation parallax is: Where k is the distortion correction coefficient, The minimum disparity threshold, This is the maximum disparity threshold.
4. The cross-scene VR3D live streaming method based on fixed premium seating positions according to claim 2, characterized in that, The method for performing hierarchical feature extraction on the preprocessed video frames in step S22 to output four types of feature maps—edge feature map, color contrast feature map, occlusion relationship feature map, and texture density feature map—includes the following: Using the Canny edge detection algorithm, calculate the gradient value (0-255) of each pixel to generate an edge feature map; calculate the YCbCr color space contrast of each pixel using the formula: contrast = (maxY - minY) / (maxY + minY), where Y is brightness, and Cb and Cr represent the concentration offset of blue and red, respectively, to generate a contrast feature map; identify occlusion edges in the image through pixel connectivity analysis, label occluded areas with occlusion tags, and label unoccluded areas with non-occlusion tags, to generate an occlusion relationship feature map; and calculate the texture richness of the 3×3 neighborhood around each pixel using local binary mode, statistically analyze the proportion of textured pixels, and generate a texture density feature map.
5. The cross-scene VR3D live streaming method based on fixed premium seating positions according to claim 2, characterized in that, Step S21 includes: S211, extracting one original video frame every 20ms from the 4K 60fps video stream transmitted from the fixed VIP seat, and downsampling the original video frame to 1920×1080 resolution; S212, using Gaussian filtering to remove high-frequency noise in the video frame, and improving the contrast of the image through histogram equalization to ensure that edge and texture features can be effectively extracted; S213, marking the ROI region of the image based on the preset region of interest from the fixed VIP seat viewpoint, and prioritizing the processing of the ROI region during subsequent feature extraction.
6. The cross-scene VR3D live streaming method based on fixed premium seating positions according to claim 2, characterized in that, Step S23 includes: S231, according to the scene preset weights, performing a weighted summation of the four types of feature values of each pixel to obtain a comprehensive feature value; S232, converting the comprehensive feature value into an actual depth value through a non-linear mapping function; S233, for pixels marked with occlusion labels, forcing the depth value of the pixel = the depth value of the occlusion object + 0.5m to avoid logical contradictions.
7. The cross-scene VR3D live streaming method based on fixed premium seating positions according to claim 2, characterized in that, Step S24 includes: S241, applying a 5×5 mean filter to the generated depth map to eliminate abrupt depth changes between adjacent pixels, making the depth transition more consistent with real physical laws; S242, cropping the depth values to 0.3-20m, forcing pixels outside the range to be set to 0.3m or 20m, and quantizing them into 8-bit integers to reduce data storage and subsequent transmission bandwidth; S243, aligning the quantized depth map pixel-level with the original 1920×1080 video frame to provide accurate input for subsequent parallax generation.
8. A cross-scene VR3D live streaming system based on fixed premium-seat camera positions, used to execute the cross-scene VR3D live streaming method based on fixed premium-seat camera positions as described in any one of claims 1-7, characterized in that, include: The system includes: a fixed VIP seating area video data acquisition module (used to acquire video data from the VIP seat perspective using panoramic cameras positioned at fixed VIP seats, and transmit the video data to a cloud processing platform via a 5G network); a cloud processing platform depth analysis and parallax calculation module (used to build a lightweight deep learning model based on the MobileDepthNet architecture, perform scene feature enhancement training, deploy it on the cloud processing platform, perform depth analysis on the video data using the deep learning model to generate a depth map of the image, perform parallax calculation based on the depth map to generate data containing 3D information, and distribute the data containing 3D information to the user's mobile APP via a CDN network); a mobile APP 3D data to 3D image conversion module (used to receive the data containing 3D information via the mobile APP and convert it into a split-screen 3D image for the left and right eyes, and transmit the 3D image to lightweight VR glasses via WiFi); and a VR glasses 3D image display module (used to display the 3D image in a split-screen format for the left and right eyes using lightweight VR glasses, providing a 3D viewing experience consistent with that of the VIP seats).
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the cross-scene VR3D live streaming method based on a fixed VIP seat as described in any one of claims 1-7.
10. A computer device, the computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the cross-scene VR3D live streaming method based on a fixed VIP seat as described in any one of claims 1-7.