Television signal VR watching implementation system and method based on 2D-to-3D processing

By designing a separate architecture with independent transmitter and portable receiver, low-latency 3D signal transmission from set-top box to VR glasses is achieved, solving the problem that existing VR glasses cannot access cable TV signals and providing a high-quality immersive 3D viewing experience.

CN121888019APending Publication Date: 2026-04-17北京市艺值科技有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
北京市艺值科技有限公司
Filing Date
2025-11-25
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing VR glasses cannot directly access cable TV signals, lack 2D to 3D conversion algorithms and hardware support, cannot provide immersive 3D effects, and existing solutions are complex to operate and have high latency, failing to meet users' mobile viewing needs.

Method used

The design incorporates a 2D-to-3D TV signal VR viewing system, employing a separate architecture with an independent transmitter and a portable receiver. It transmits 2D signals from the set-top box to VR glasses via millimeter-wave wireless signals, enabling 3D display. The core processing function resides on the transmitter, supporting low-latency wireless transmission and portable power supply.

Benefits of technology

It achieves a high-quality 3D TV viewing experience in long-distance and mobile environments, reduces operational complexity and latency, has wide applicability, and meets users' diverse viewing needs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121888019A_ABST
    Figure CN121888019A_ABST
Patent Text Reader

Abstract

The invention provides a television signal VR watching realization system and method based on 2D-to-3D processing, and the system comprises a transmitting end and a portable receiving end which are connected with a set top box, the transmitting end comprises an HDMI signal access unit, a 2D-to-3D processing unit and a signal transmitting unit which are connected in sequence, and the 2D-to-3D processing unit comprises a quad-core processor, an NPU processing unit and an algorithm storage chip which are connected in sequence; the receiving end comprises a signal receiving unit and a signal output unit which are connected in sequence; the signal receiving unit is connected with the signal transmitting unit, and the signal output unit is connected with VR glasses. According to the invention, an independent transmitting end and a portable receiving end are constructed, a split type 2D-3D system is designed, a complete link of a set top box 2D signal-transmitting end 2D-3D-wireless transmission-receiving end output-VR glasses display is realized, a processing function is transferred to the independent transmitting end, and 3D video stream television watching out of the range of the set top box is realized by matching with the portable receiving end.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of cross-disciplinary technology of virtual reality (VR) display and video signal processing, and more specifically, to a system and method for realizing VR viewing of television signals based on 2D-to-3D conversion processing. Background Technology

[0002] Currently, VR display technology has expanded from professional gaming devices to home entertainment, and user demand for "large-screen" and "immersive" viewing of traditional television programs is gradually increasing, especially among groups accustomed to traditional television (such as many elderly users) who are used to watching cable TV content but also desire the immersive experience of VR. Therefore, the current user demand for "immersive 3D TVs" is constantly rising.

[0003] However, existing VR headsets largely rely on online streaming media or dedicated VR content. The VR content ecosystem is still dominated by online videos and games, which are disconnected from traditional cable TV signals. They lack decoding and conversion modules for HDMI cable signals, making direct connection to cable TV signals (such as HDMI output from set-top boxes) impossible. Furthermore, the 3D display of existing VR headsets depends on the 3D format (such as split-screen or top-bottom split-screen) video signals provided by the content source, while cable TV signals are 2D planar signals. Lacking real-time 2D-to-3D conversion algorithms and hardware support, they cannot be directly converted to VR-compatible 3D stereoscopic effects. Therefore, it is currently difficult for those accustomed to traditional TVs to use VR headsets to watch cable TV programs. Traditional 3D TVs also require dedicated 3D format content and have limited screen size, failing to provide a VR-level immersive field of view (typically, the VR field of view is >90°, far exceeding the 30-50° of a TV). Moreover, some existing VR headsets only support dedicated interfaces, requiring additional adapters when connecting to set-top boxes. Their hardware interface designs are not adapted to traditional home appliances such as TVs, making operation complex and susceptible to signal interference.

[0004] Meanwhile, users who want "immersive 3D TVs" also want to break through the limitations of traditional TV viewing scenarios and achieve "mobile" viewing. According to the 2024 Home Entertainment Devices Report, 65% of users want to watch TV content from their living room set-top boxes in their bedrooms.

[0005] Immersive 3D TVs require 2D-to-3D signal conversion. Existing 2D-to-3D technologies mainly fall into two categories: hardware-integrated (such as high-end VR headsets) and software-based (such as VR glasses connected to a mobile app). However, both existing VR TV viewing solutions lack low-latency, long-range wireless signal transmission designs, failing to meet users' mobile viewing needs. Currently, users must remain stationary near the set-top box, unable to move to bedrooms, balconies, or other locations farther from the set-top box, resulting in a poor viewing experience. The main reasons for this include the following:

[0006] Existing 2D-to-3D conversion functions are mostly integrated into VR glasses or mobile apps. Among them, the processing module integrated into VR glasses weighs over 300g, resulting in poor wearing comfort; due to the high degree of hardware integration, the cost of dedicated chips accounts for a large proportion of the cost, exceeding 1000 yuan, making the price threshold relatively high.

[0007] The 2D to 3D conversion solution of mobile apps relies on mobile phones as an intermediate node. Users need to operate the mobile phone and VR glasses at the same time. Because the mobile phone has to handle multiple tasks such as signal reception, 2D to 3D conversion and data transmission at the same time, the mobile phone battery is consumed quickly (continuous use ≤2 hours), and the battery life is poor, which cannot meet the needs of long-term viewing.

[0008] Existing VR headsets supporting HDMI input (such as some industrial VR displays) typically consist of: a VR display module (dual screens, 1920×1080 resolution), an HDMI input interface (supporting only 1080P 60Hz), a simplified processor (without 2D-to-3D conversion), and a single power supply interface. The immersive TV experience is achieved by receiving external video signals (such as from a computer or game console) via HDMI and displaying the flat image directly on the VR screen; the user must manually adjust the interpupillary distance. The limitations are: it can only display 2D flat signals and cannot convert 2D content from cable TV to 3D stereoscopic effects; it has poor interface compatibility (not supporting the HDCP protocol of some set-top boxes); and it lacks optimization for TV signals (such as fast channel switching response), failing to meet the core requirement of "watching 3D cable TV with VR."

[0009] Currently, there is also a solution of "HDMI wireless transmitter + VR glasses". Taking a certain brand's HDMI wireless screen projector, model WH-500, as an example, its structure includes: HDMI wireless transmitter, HDMI wireless receiver (which needs to be connected to VR glasses via an HDMI to Type-C adapter), and external power adapter.

[0010] The solution requires the following steps to achieve immersive TV functionality: the set-top box's HDMI signal is connected to the transmitter, and the 2D signal is transmitted to the receiver via 2.4GHz WiFi. The receiver outputs the signal to the VR glasses via an adapter, and the user needs to manually enable the "simulated 3D" function in the VR glasses.

[0011] However, this solution also has obvious drawbacks: the HDMI wireless transmitter can only transmit 2D signals and has no native 2D to 3D conversion function, resulting in poor 3D effects; the 2.4GHz WiFi transmission latency is ≥200ms, causing audio and video to be out of sync; the receiver requires a separate power supply, and there are many conversion steps, making operation complicated; the "simulated 3D" function only stretches the image horizontally and horizontally, without a realistic sense of depth, and cannot meet the core requirement of "high-quality 3D + mobile viewing". Summary of the Invention

[0012] Therefore, the purpose of this invention is to develop and design a VR viewing system and method for television signals based on 2D-to-3D conversion processing. It constructs a split architecture of "independent transmitter + portable receiver," designs a split 2D-to-3D system, and realizes a complete link of "set-top box 2D signal → transmitter 2D-to-3D conversion → wireless transmission → receiver output → VR glasses display." The core processing function is transferred to the independent transmitter, paired with a portable receiver, providing users with a 3D television viewing solution that is independent of the set-top box. Furthermore, it features independent HDMI input and power supply interfaces, making it compatible with VR 3D viewing of cable television signals, with wide applicability and satisfying users' diverse viewing experiences.

[0013] This invention provides a VR viewing system for television signals based on 2D-to-3D conversion processing, comprising: a transmitter connected to a television set-top box and a portable receiver. The transmitter includes an HDMI signal access unit, a 2D-to-3D processing unit, and a signal transmission unit connected in sequence. The 2D-to-3D processing unit includes a quad-core processor, an NPU processing unit, and an algorithm storage chip connected in sequence. The receiver includes a signal receiving unit and a signal output unit connected in sequence. The signal receiving unit is connected to the signal transmission unit, and the signal output unit is connected to VR glasses.

[0014] Specifically, the set-top box is connected to the transmitter via an HDMI cable, the transmitter can connect to the receiver via millimeter-wave wireless signal, and the receiver is connected to the VR glasses via a Type-C cable.

[0015] Preferably, the transmitter uses an ABS shell with a passive heat dissipation grille inside. The receiver uses a flexible printed circuit board (FPC) structure and is equipped with LED indicator lights: a solid blue light indicates normal reception; a flashing yellow light indicates a weak signal; and a red light indicates no signal. It is equipped with an adjustable mounting clip to be fixed in the desired position.

[0016] Specifically, the 2D-to-3D processing unit includes a main processor and a dedicated image processor (GPU). The 2D-to-3D algorithm module is built into the main processor firmware and includes submodules for depth estimation and left / right eye parallax generation. The main processor receives the video stream decoded from HDMI (1920×1080 pixels per frame) via the MIPI-CSI interface, processes it through the GPU, and outputs left and right eye 3D images via the MIPI-DSI interface. The 2D-to-3D algorithm module communicates with the processor via an internal bus, with a processing latency of ≤50ms.

[0017] Preferably, the VR display and optical system of the VR glasses consists of the following components: a Micro-OLED display screen, a Fresnel lens group, and an ambient light sensor. The display screen is connected to the main processor via a MIPI-DSI interface. The lens group is fixed in front of the display screen (10mm pitch), and the interpupillary distance adjustment knob (2cm in diameter) moves left and right through gear linkage with the lens. The ambient light sensor is soldered to the front of the frame and communicates with the main processor via an I2C interface. The left and right eye screens of the VR glasses display parallax-processed images respectively, allowing users to see a stereoscopic superposition effect through the lenses, simulating the immersive field of vision of a "100-inch 3D TV". The interpupillary distance adjustment adapts to different users, avoiding ghosting.

[0018] Preferably, the VR viewing system for television signals of the present invention includes an operation and auxiliary module, comprising the following components: physical buttons (3: power button, depth-of-field adjustment button + / -, channel switching button); an infrared remote control receiver (compatible with set-top box remote control, allowing direct channel changing using a television remote control); and a lightweight frame (ABS material, weight ≤300g, adjustable headband). The buttons and remote control receiver are connected to the main processor via a GPIO interface; the frame has built-in heat dissipation fins (in contact with the processor), providing passive cooling without a fan (reducing noise). Users can change channels using the original set-top box remote control or adjust the depth of field using the physical buttons on the glasses; the operation logic is consistent with traditional televisions; the lightweight design is suitable for prolonged wear (no pressure even after 2 hours of continuous use).

[0019] Furthermore, the HDMI signal access unit includes, in sequence, an HDMI 2.0 input interface for connecting a TV set-top box, an HDCP 2.2 decryption chip for ensuring the secure transmission of protected video content (such as 4K video, Blu-ray, etc.) and preventing unauthorized copying or tampering, and a signal detection circuit that automatically identifies the input resolution and frame rate (e.g., 1080P 60fps / 4K 30fps), wherein the signal detection circuit is connected to the quad-core processor.

[0020] Furthermore, the signal transmitting unit includes: a 60GHz millimeter-wave transmitting circuit (transmission rate 10Gbps) and a 360° rotatable directional adjustable antenna (optimizing the transmission direction) connected in sequence. The 60GHz millimeter-wave transmitting circuit is connected to an adaptive frequency hopping anti-interference module (avoiding interference in the 2.4GHz band).

[0021] Furthermore, the signal receiving unit includes a 60GHz millimeter-wave receiving circuit, a receiving antenna, and a signal demodulation chip that restores the millimeter-wave signal to a 3D video stream, connected in sequence; the signal output unit is provided with a Type-C female connector (10Gbps transmission rate, preferably supporting USB 3.2 Gen2 protocol) for connecting to VR glasses.

[0022] Furthermore, the transmitting end is equipped with a DC power interface (preferably 12V / 2A), which is connected to an external power adapter; the receiving end is equipped with a power supply unit, which includes a Type-C male connector (preferably supporting PD 2.0 protocol, input 5V / 3A, can be connected to a power bank) and a power distribution circuit connected in sequence. The power distribution circuit divides the power into two paths: one path powers the receiving end itself (power consumption ≤3W), and the other path powers the VR glasses (output 5V / 2A).

[0023] Specifically, a power bank can be used to connect to the receiver via a Type-C cable, forming a closed loop of "signal input-processing-transmission-display-power supply".

[0024] This invention also provides a method for VR viewing of television signals based on 2D-to-3D conversion processing, applied to the VR viewing system for television signals based on 2D-to-3D conversion processing described above, comprising the following steps:

[0025] S1. The 2D HDMI signal output from the set-top box is connected to the transmitter. After being decrypted by the HDCP 2.2 decryption chip, the signal parameters are identified by the signal detection circuit, and the 2D signal is transmitted to the quad-core processor.

[0026] S2. The quad-core processor calls the NPU processing unit to run the 2D to 3D algorithm, generates a depth map through the edge of the screen and color contrast, calculates the left and right eye parallax based on the depth map, and generates a 3D signal for left and right split screen.

[0027] S3. The signal transmitting unit packages the 3D signal into a 60GHz millimeter-wave data packet and transmits it to the receiving end;

[0028] Preferably, for the receiver placed in the bedroom, it can transmit towards the bedroom via a directional antenna, with a transmission delay of ≤25ms and a transmission distance of ≥15 meters in open environments;

[0029] S4. Power on the receiver (connect the receiver to a power bank via a Type-C connector; the indicator light will illuminate blue) and search for millimeter-wave signals from the transmitter; adjust the pairing signal receiving unit and signal transmitting unit (this can be determined by the indicator lights; a solid blue light indicates successful pairing); the signal demodulation chip will restore the received millimeter-wave signal to a 3D video stream.

[0030] S5. Transmit the 3D video stream (via Type-C female connector) to the VR glasses, power the VR glasses through the power distribution circuit, and watch the 3D video stream TV content while wearing the glasses.

[0031] In practical applications, this invention can be moved to any location in the bedroom within 8 meters of the set-top box without signal interruption.

[0032] Furthermore, the method for generating a depth map through image edges and color contrast in step S2 includes:

[0033] S21. Extract the edge contours of the 2D image through multi-stage edge detection and adaptive filtering (the edge contours of objects in the image are the basis for distinguishing different objects and judging distance), including:

[0034] To address potential light flicker and camera noise in 2D videos, a 3×3 Gaussian filter kernel is used to smooth and denoise the video frames, balancing denoising effectiveness with edge preservation. Convolution operations are employed to eliminate high-frequency noise, preventing noise from being misidentified as edges. The expression is as follows:

[0035] ;

[0036] Where (x, y) represents the pixel; =1.2;

[0037] Global brightness normalization is performed on video frames, mapping pixel brightness values ​​of 0-255 to the range of [0.1, 0.9] to prevent the edges of overly bright / dark areas from being masked by extreme brightness values;

[0038] An optimized Canny edge detection algorithm is used, and the gradient magnitude and direction of each pixel are calculated using the Sobel operator in the x and y directions:

[0039] gradient magnitude gradient direction = ;

[0040] Among them, horizontal gradient Reflecting vertical edges and vertical gradients Reflects horizontal edges, ensuring that both the horizontal and vertical contours of the object are captured simultaneously;

[0041] Local maxima are preserved along the gradient direction. If the gradient magnitude of a pixel is greater than that of its two adjacent pixels along the gradient direction, it is determined to be an edge and the pixel is preserved. If the gradient magnitude of a pixel is not greater than that of its two adjacent pixels along the gradient direction, it is determined to be a non-edge and non-maxima are suppressed to prevent the edges from widening.

[0042] Dual-threshold edge connectivity: Set a high threshold (e.g., 80) and a low threshold (e.g., 30). For pixels with a gradient magnitude greater than the high threshold, they are directly identified as strong edges (100% retained); for pixels with a gradient magnitude less than the low threshold, they are directly identified as non-edges (discarded, such as textures with blurred backgrounds); for pixels with a gradient magnitude between the high and low thresholds, if they are connected to strong edges, they are identified as weak edges (retained); if they are not connected to strong edges, they are discarded, ensuring the continuity and integrity of edges.

[0043] S22. Color contrast reflects the "visual prominence" of objects (near objects have vibrant colors and strong contrast, while distant objects have faded colors and weak contrast). Color contrast is optimized through "scene-based contrast enhancement + dynamic range adjustment," strengthening the differences in brightness / color and assisting in depth layering, including:

[0044] Convert the color space, transforming the video frames from the RGB color space to the YCbCr color space (Y: luminance channel, Cb: blue difference channel, Cr: red difference channel), separating luminance and chrominance;

[0045] The reason is that the human eye is much more sensitive to changes in brightness than to changes in chromaticity, and the luminance channel (Y) is the core carrier of contrast. After separation, the luminance channel can be optimized separately to avoid color distortion caused by excessive adjustment of the chromaticity channel.

[0046] The conversion formula (based on the BT.601 standard) is as follows:

[0047] Y =0.299R+0.587G + 0.114B Cb =-0.1687R-0.3313G + 0.5B + ​​128 Cr = 0.5R- 0.4187G - 0.0813B + 128;

[0048] Where Y is the luminance channel, Cb is the blue difference channel, and Cr is the red difference channel; R, G, and B are the red, green, and blue primary colors, respectively.

[0049] Adaptive histogram equalization (CLAHE) is used to enhance brightness and contrast, and local overexposure caused by global equalization is avoided through dynamic adjustment of different regions. This includes:

[0050] The luminance channel Y is divided into 8×8 sub-blocks, each sub-block is 240×144 pixels, adapted to 1920×1080 resolution, and the histogram of each sub-block is calculated independently.

[0051] Set a contrast threshold (e.g., 40). If the number of pixels at a certain gray level in the histogram of a sub-block exceeds the threshold, the excess will be evenly distributed to other gray levels to avoid excessive stretching of brightness within the sub-block, which would lead to loss of detail.

[0052] Bilinear interpolation is used to stitch together the equalization results of adjacent sub-blocks to eliminate abrupt brightness changes at the sub-block boundaries and make the brightness transition natural.

[0053] Saturation enhancement for chrominance channels Cb and Cr is expressed as follows:

[0054] ;

[0055] in: is the average value of the Cb and Cr channels of the video frame, reflecting the overall color tone; k is the saturation enhancement coefficient; , The value should be within the range of 0-255 to avoid color overflow and distortion.

[0056] Based on the logic that the human eye judges the distance of objects by the details of the image, the depth value of each pixel in the image (i.e. the distance between the pixel and the virtual camera) is inferred by using the image features of edge contours and color contrast in video frames, and a global depth map is generated.

[0057] Preferably, the depth estimation model is optimized for common cable TV content (such as news anchors and sports events), reducing the depth value of the anchor's face (to highlight the three-dimensionality) and increasing the background depth value (to enhance the sense of layering); when switching channels quickly (<0.5 seconds), the depth calculation is paused and the 2D to 3D transition frame is directly output to avoid screen stuttering.

[0058] Furthermore, the method for calculating the left and right eye disparity based on the depth map in step S2 includes:

[0059] Based on the pinhole imaging principle, the image height on the virtual camera's imaging plane is inversely proportional to the depth. Combining this with the difference in visual angle between the human eye's pupils, the basic parallax calculation formula is:

[0060] ;

[0061] Where B is the interpupillary distance, f is the camera focal length, D is the object depth, and P is the pixel size;

[0062] Substitute the depth map pixel by pixel into the basic disparity calculation formula to calculate the initial disparity. ;

[0063] When the initial disparity exceeds the range of human eye fusion, it is corrected to ensure that the disparity is reasonable. The initial disparity is adjusted to a usable adaptive disparity through threshold constraints and distortion correction. The formula for calculating the adaptation parallax is:

[0064] ;

[0065] Where k is the distortion correction coefficient. The minimum disparity threshold, This is the maximum disparity threshold.

[0066] Specifically, methods for generating 3D signals for left and right split-screen displays include:

[0067] The cable TV signal is set as a video frame stream in .vr3d data format. This video frame stream and pixel-level parallax data are encapsulated to generate a 3D image conforming to human visual perception at the transmitting end. The .vr3d data structure employs a frame-level encapsulation + block-level storage design to prevent overall parsing failure due to single-frame data corruption. The specific structure is shown in Table 1.

[0068] Table 1

[0069]

[0070] The logical chain for .vr3d parsing and 3D generation is as follows: .vr3d format data (file header + basic parameter block + frame data block) → data parsing (format verification → parameter reading → frame data extraction → video decoding + parallax restoration) → 3D generation (parallax allocation → left and right eye frame offset → distortion correction → split-screen stitching) → WiFi transmission to VR glasses → users watch 3D live broadcast. The entire process is designed around "low latency, high compatibility, and weak network adaptation" to ensure that both on-site back-row users and off-site users can obtain a stable and distortion-free 3D viewing experience.

[0071] This invention completes a five-step process: VR3D data reading → format verification → parameter parsing → frame data extraction → parallax restoration. It accurately extracts the video stream and parallax data from the .vr3d format to generate a 3D image that conforms to human visual perception. The .vr3d format data parsing includes the following steps:

[0072] Step 1, Data Reception and Initialization: Receive cable TV signals from the set-top box, set the data to .vr3d format, and adopt a "segmented reception + buffering" strategy - store every 1MB of data received in the local buffer to avoid memory overflow;

[0073] Read the first 8 bytes of data. If it is not equal to "VR3D_2025", immediately trigger a "format error" prompt. If the format is correct, continue reading the "basic parameter block offset" in the file header and locate the starting address of the basic parameter block.

[0074] Step 2, Basic Parameter Analysis and Configuration:

[0075] Parameter reading: The "video resolution (W×H), frame rate (F), parallax bit depth (B), maximum parallax threshold (D_max), color space (C)" are parsed sequentially from the basic parameter block and stored in the "global configuration variables" of the APP;

[0076] Step 3: Frame data block location and reading:

[0077] Frame index construction: Obtain the total number of frames N from the "Total Frames" field in the file header, traverse all frame data blocks, extract the "frame number, timestamp, video data length, and parallax data length" of each frame, and construct a "frame index table" (stored in memory) to facilitate quick location by frame number or timestamp later;

[0078] Frame data extraction: Based on the user's viewing progress (e.g., watching frame 100), find the starting address of frame 100 from the frame index table, first read the "video data length (Lv)", then extract Lv bytes of video data; next, read the "parallax data length (Ld)" and extract Ld bytes of parallax data.

[0079] Step 4: Video data decoding and parallax data restoration:

[0080] Call the video decoder to decode the H.265 format video data into 2D video frames in RGB / YCbCr format (resolution consistent with the basic parameter block), and store them as "raw video frame matrix" (dimension: H×W×3, 3 represents RGB three channels).

[0081] Parallax data restoration: If the parallax data bit depth B=8, convert the parallax data binary stream into an "8-bit unsigned integer matrix" (dimension: H×W), where each element represents the parallax value d of the corresponding pixel (0≤d≤D_max);

[0082] If a "parallax compression flag" exists (a new field added to the basic parameter block, 1 represents compression, 0 represents no compression), it is first decompressed by "run-length encoding (RLE)". For example, compressed data "35,5" represents that the parallax value of 5 consecutive pixels is 35, which is restored to [35,35,35,35,35];

[0083] Data alignment: Ensure that the pixel positions of the "original video frame matrix" and the "disparity matrix" correspond one-to-one (i.e., the pixel in the i-th row and j-th column of the video frame corresponds to the disparity value in the i-th row and j-th column of the disparity matrix) to avoid misalignment during 3D generation.

[0084] After obtaining the "original video frame matrix + disparity matrix", the difference between human binocular vision is simulated. Through the steps of "disparity allocation → left and right eye frame generation → distortion correction → image output", the 2D video frames are converted into 3D images suitable for VR glasses display.

[0085] The process of generating 3D images includes the following steps:

[0086] Step 1, Disparity Assignment (Determine the direction and amount of left and right eye offset):

[0087] Based on the principle of human binocular vision, the left eye observes objects from a perspective biased to the right, while the right eye observes them from a perspective biased to the left. Therefore, it is necessary to perform a "directional offset" on the pixels of the original video frame, with the offset amount proportionally allocated based on the parallax value:

[0088] ,

[0089] ;

[0090] Where d is the disparity value of a pixel in the disparity matrix (0≤d≤35 pixels, from the parsed disparity matrix);

[0091] This is the rightward offset of the pixel in the left-eye frame (in pixels), rounded down (e.g., when d=35). =17);

[0092] This is the leftward offset of the pixel in the right-eye frame (unit: pixels), such as when d=35. =18;

[0093] Step 2, Generation of left and right eye frames (pixel offset and boundary padding):

[0094] Left eye frame generation: Initialize the "left eye frame matrix" (dimension: H×W×3, consistent with the original video frame);

[0095] For each pixel (i,j) in the original video frame (where i is the row number and j is the column number), calculate its new column number in the left-eye frame: = j + ;

[0096] like If W (not exceeding the right boundary of the image), then the RGB value of the original pixel (i,j) is assigned to the left eye frame (i,j). );

[0097] like For pixels ≥ W (exceeding the right boundary), the "edge pixel copying" strategy is used—the RGB value of the rightmost pixel (i, W-1) in the original frame is assigned to the left eye frame (i, W-1). - W), to avoid black borders;

[0098] Right eye frame generation: Initialize the "right eye frame matrix" (dimension: H×W×3);

[0099] For each pixel (i,j) in the original video frame, calculate its new column number in the right-eye frame: ;

[0100] like If the pixel does not exceed the left edge of the frame, then the RGB value of the original pixel (i,j) is assigned to the right eye frame (i,j_right).

[0101] like (Exceeding the left boundary), the same "edge pixel copying" method is used—the RGB value of the leftmost pixel (i,0) of the original frame is assigned to the right eye frame (i, + W);

[0102] To avoid "holes" after adjacent pixels are shifted (such as when a pixel is shifted and there are no pixels filling the original position), "bilinear interpolation filling" is used for the "unassigned pixels" in the left eye frame and the "unassigned pixels" in the right eye frame. For example, when (i,j) in the left eye frame is unassigned, the average RGB value of the four surrounding assigned pixels is taken as the filling value to ensure smooth image.

[0103] Step 3, Distortion Correction: Due to the "barrel distortion" (stretched edge pixels) of the Fresnel lens of the VR glasses, if the left and right eye frames are directly output, the user will see a distorted 3D picture. Therefore, it is corrected by a polynomial distortion correction algorithm, including: (1) Obtaining glasses parameters: Read the "distortion coefficients" (such as k1=-0.3, k2=0.1, k3=-0.05, representing third-order distortion coefficients) and "optical center" [such as (W / 2, H / 2, i.e., the center of the picture] of the currently connected VR glasses from the "VR device parameter library" (built-in distortion parameters of mainstream VR glasses models);

[0104] (2) Distortion coordinate calculation: For each pixel (x, y) (screen coordinates) in the left and right eye frames, calculate the ideal coordinates of the pixel (x, y) before distortion using the following formula.

[0105] ,

[0106] ,

[0107] ;

[0108] Where (x,y) are the pixel coordinates (screen coordinates, range 0≤x) of the corrected left and right eye frames. <W,0≤y<H);

[0109] ( ) represents the optical center coordinates of the VR glasses (usually the center of the screen, such as (960, 540) corresponding to a 1920×1080 resolution);

[0110] r is the normalized distance from the pixel to the optical center (dimensionless, range 0-1).

[0111] k1, k2, and k3 are distortion coefficients (determined by the VR glasses hardware; for example, k1 = -0.3 represents first-order barrel distortion correction).

[0112] ( The coordinates are the ideal pixel coordinates before distortion (which need to be mapped to the coordinate range of the original left and right eye frames);

[0113] The OpenCV remap function is called to remap the left and right eye frames according to the mapping relationship between "ideal coordinates and screen coordinates" to generate distortion-corrected left and right eye frames, ensuring that users see distortion-free 3D images through VR glasses.

[0114] Step 4: Synchronize 3D image output with VR glasses:

[0115] (1) Image format adaptation: The distortion-corrected left and right eye frames are stitched together according to the "3D display format" supported by the VR glasses [the present invention uses the "side-by-side" format by default] - that is, the left eye frame is placed on the left and the right eye frame is placed on the right to generate a "single frame 3D image" [resolution: H×(2W), such as 1080×3840];

[0116] (2) Low-latency transmission: The "single frame 3D image" is transmitted to the VR glasses. During the transmission process, the "UDP protocol + frame sequence number mark" is used to ensure that each frame of data arrives in order and the transmission delay is controlled within ≤20ms;

[0117] (3) Audio-visual synchronization: Read the "timestamp" in the frame data block and compare it with the audio playback timestamp of the VR glasses. If the video frame timestamp is ahead of the audio playback timestamp by more than 50ms, then pause the video transmission for 10ms; if it is behind by more than 50ms, then speed up the video decoding speed (such as skipping 1 non-key frame) to ensure that the audio-visual synchronization error is ≤30ms and avoid the user's discomfort of "the picture and sound are out of sync".

[0118] The present invention also provides a computer-readable storage medium having a computer program stored thereon, wherein the program, when executed by a processor, implements the steps of the method for VR viewing of television signals based on 2D-to-3D processing as described above.

[0119] The present invention also provides a computer device, the computer device including a memory, a processor and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the steps of the method for VR viewing of television signals based on 2D to 3D processing as described above.

[0120] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0121] This invention provides a VR viewing implementation and system for television signals based on 2D-to-3D processing. It constructs a split architecture of "independent transmitter + portable receiver," designing a split 2D-to-3D system to achieve a complete link: "set-top box 2D signal → transmitter 2D-to-3D → wireless transmission → receiver output → VR glasses display." The core processing function is transferred to the independent transmitter, paired with a portable receiver, enabling 3D video streaming television viewing without the need for a set-top box. Furthermore, it features independent HDMI input and power supply interfaces, ensuring compatibility with VR 3D viewing of cable television signals. This provides users with diverse viewing experiences, offering good applicability and broad application prospects. Attached Figure Description

[0122] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention.

[0123] In the attached diagram:

[0124] Figure 1 This is a schematic diagram illustrating the basic process of VR viewing of television signals according to an embodiment of the present invention;

[0125] Figure 2 This is a flowchart illustrating a method for VR viewing of television signals based on 2D-to-3D conversion processing, according to an embodiment of the present invention. Detailed Implementation

[0126] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and products consistent with some aspects of this disclosure as detailed in the appended claims.

[0127] The terminology used in this disclosure is for the purpose of describing particular embodiments only and is not intended to be limiting of the disclosure. The singular forms “a,” “the,” and “the” as used in this disclosure and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.

[0128] It should be understood that although the terms first, second, third, etc., may be used in this disclosure to describe various information, such information should not be limited to these terms. These terms are used only to distinguish information of the same type from one another. For example, without departing from the scope of this disclosure, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to determination."

[0129] The embodiments of the present invention will be described in further detail below.

[0130] This invention provides a VR viewing system for television signals based on 2D-to-3D conversion processing, comprising: a transmitter connected to a television set-top box and a portable receiver. The transmitter includes an HDMI signal access unit, a 2D-to-3D processing unit, and a signal transmission unit connected in sequence. The 2D-to-3D processing unit includes a quad-core processor, an NPU processing unit, and an algorithm storage chip connected in sequence. The receiver includes a signal receiving unit and a signal output unit connected in sequence. The signal receiving unit is connected to the signal transmission unit, and the signal output unit is connected to VR glasses.

[0131] The set-top box connects to the transmitter via an HDMI cable, the transmitter connects to the receiver via millimeter-wave wireless signal, and the receiver connects to the VR glasses via a Type-C cable.

[0132] The transmitter uses an ABS shell with an internal passive heat dissipation grille. The receiver uses a flexible printed circuit board (FPC) structure and features LED indicator lights: a solid blue light indicates normal reception; a flashing yellow light indicates a weak signal; and a red light indicates no signal. An adjustable mounting clip allows for secure installation in the desired location.

[0133] The 2D to 3D processing unit includes a main processor and a dedicated image processor (GPU); the 2D to 3D algorithm module is built into the main processor firmware and includes depth estimation and left and right eye parallax generation sub-modules.

[0134] The main processor receives the video stream decoded from HDMI (1920×1080 pixels per frame) through the MIPI-CSI interface, and after processing by the GPU, outputs the left and right eye 3D images through the MIPI-DSI interface; the 2D to 3D algorithm module communicates with the processor through the internal bus, with a processing latency of ≤50ms.

[0135] The HDMI signal access unit includes, in sequence, an HDMI 2.0 input interface for connecting a TV set-top box, an HDCP 2.2 decryption chip for ensuring the secure transmission of protected video content (such as 4K video, Blu-ray, etc.), an HDCP 2.2 decryption chip to prevent unauthorized copying or tampering, and a signal detection circuit that automatically identifies the input resolution and frame rate (e.g., 1080P 60fps / 4K 30fps). The signal detection circuit is connected to the quad-core processor.

[0136] The signal transmitting unit includes: a 60GHz millimeter-wave transmitting circuit (transmission rate 10Gbps) and a 360° rotatable directional adjustable antenna (optimizing the transmission direction) connected in sequence. The 60GHz millimeter-wave transmitting circuit is connected to an adaptive frequency hopping anti-interference module (avoiding interference in the 2.4GHz band).

[0137] The signal receiving unit includes a 60GHz millimeter-wave receiving circuit, a receiving antenna, and a signal demodulation chip that restores the millimeter-wave signal to a 3D video stream, connected in sequence; the signal output unit is equipped with a Type-C female connector (10Gbps transmission rate, preferably supporting USB 3.2 Gen2 protocol) for connecting VR glasses.

[0138] The transmitter is equipped with a DC power interface (preferably 12V / 2A), which is connected to an external power adapter. The receiver is equipped with a power supply unit, which includes a Type-C male connector (preferably supporting PD 2.0 protocol, input 5V / 3A, can be connected to a power bank) and a power distribution circuit connected in sequence. The power distribution circuit divides the power into two paths: one path powers the receiver itself (power consumption ≤3W), and the other path powers the VR glasses (output 5V / 2A).

[0139] In this embodiment, a power bank is used to connect to the receiver via a Type-C cable, forming a closed loop of "signal input-processing-transmission-display-power supply".

[0140] This embodiment employs an independent 2D-to-3D converter design at the transmitter, transferring the 2D-to-3D processing module from the existing VR glasses / phone to a separate transmitter. It utilizes a "processor + NPU" hardware architecture to achieve real-time conversion of 4K 60fps signals, solving the problems of heavy and costly VR glasses. The transmitter and receiver use 60GHz millimeter-wave wireless transmission technology, achieving a transmission rate of 10Gbps, latency ≤25ms, and a distance ≥8 meters. The signal remains stable after penetrating non-load-bearing walls, resolving the issues of high latency and short range found in existing WiFi transmissions. The receiver adopts an integrated "signal output + dual-channel power supply" design, simultaneously outputting 3D signals and powering the VR glasses via a Type-C interface. It supports power bank input, eliminating the need for an external power adapter and addressing portability and battery life concerns. The transmitter incorporates a signal detection circuit and an HDCP2.2 decryption chip, ensuring compatibility with set-top box signals of different resolutions (720P-4K), eliminating the need for manual user adjustments and lowering the operational threshold.

[0141] This invention also provides a method for VR viewing of television signals based on 2D-to-3D conversion processing, applicable to the VR viewing system for television signals based on 2D-to-3D conversion processing described above. (See also...) Figure 2 As shown, it includes the following steps:

[0142] S1. The 2D HDMI signal output from the set-top box is connected to the transmitter. After being decrypted by the HDCP 2.2 decryption chip, the signal parameters are identified by the signal detection circuit, and the 2D signal is transmitted to the quad-core processor.

[0143] S2. The quad-core processor calls the NPU processing unit to run the 2D to 3D algorithm, generates a depth map through the edge of the screen and color contrast, calculates the left and right eye parallax based on the depth map, and generates a 3D signal for left and right split screen.

[0144] Methods for generating depth maps based on image edges and color contrast include:

[0145] S21. Extract the edge contours of the 2D image through multi-stage edge detection and adaptive filtering (the edge contours of objects in the image are the basis for distinguishing different objects and judging distance), including:

[0146] To address potential light flicker and camera noise in 2D videos, a 3×3 Gaussian filter kernel is used to smooth and denoise the video frames, balancing denoising effectiveness with edge preservation. Convolution operations are employed to eliminate high-frequency noise, preventing noise from being misidentified as edges. The expression is as follows:

[0147] ;

[0148] Where (x, y) represents the pixel; =1.2;

[0149] Global brightness normalization is performed on video frames, mapping pixel brightness values ​​of 0-255 to the range of [0.1, 0.9] to prevent the edges of overly bright / dark areas from being masked by extreme brightness values;

[0150] An optimized Canny edge detection algorithm is used, and the gradient magnitude and direction of each pixel are calculated using the Sobel operator in the x and y directions:

[0151] gradient magnitude gradient direction = ;

[0152] Among them, horizontal gradient Reflecting vertical edges and vertical gradients Reflects horizontal edges, ensuring that both the horizontal and vertical contours of the object are captured simultaneously;

[0153] Local maxima are preserved along the gradient direction. If the gradient magnitude of a pixel is greater than that of its two adjacent pixels along the gradient direction, it is determined to be an edge and the pixel is preserved. If the gradient magnitude of a pixel is not greater than that of its two adjacent pixels along the gradient direction, it is determined to be a non-edge and non-maxima are suppressed to prevent the edges from widening.

[0154] Dual-threshold edge connectivity: Set a high threshold (e.g., 80) and a low threshold (e.g., 30). For pixels with a gradient magnitude greater than the high threshold, they are directly identified as strong edges (100% retained); for pixels with a gradient magnitude less than the low threshold, they are directly identified as non-edges (discarded, such as textures with blurred backgrounds); for pixels with a gradient magnitude between the high and low thresholds, if they are connected to strong edges, they are identified as weak edges (retained); if they are not connected to strong edges, they are discarded, ensuring the continuity and integrity of edges.

[0155] S22. Color contrast reflects the "visual prominence" of objects (near objects have vibrant colors and strong contrast, while distant objects have faded colors and weak contrast). Color contrast is optimized through "scene-based contrast enhancement + dynamic range adjustment," strengthening the differences in brightness / color and assisting in depth layering, including:

[0156] Convert the color space, transforming the video frames from the RGB color space to the YCbCr color space (Y: luminance channel, Cb: blue difference channel, Cr: red difference channel), separating luminance and chrominance;

[0157] The reason is that the human eye is much more sensitive to changes in brightness than to changes in chromaticity, and the luminance channel (Y) is the core carrier of contrast. After separation, the luminance channel can be optimized separately to avoid color distortion caused by excessive adjustment of the chromaticity channel.

[0158] The conversion formula (based on the BT.601 standard) is as follows:

[0159] Y =0.299R+0.587G + 0.114B Cb =-0.1687R-0.3313G + 0.5B + ​​128 Cr = 0.5R- 0.4187G - 0.0813B + 128;

[0160] Where Y is the luminance channel, Cb is the blue difference channel, and Cr is the red difference channel; R, G, and B are the red, green, and blue primary colors, respectively.

[0161] Adaptive histogram equalization (CLAHE) is used to enhance brightness and contrast, and local overexposure caused by global equalization is avoided through dynamic adjustment of different regions. This includes:

[0162] The luminance channel Y is divided into 8×8 sub-blocks, each sub-block is 240×144 pixels, adapted to 1920×1080 resolution, and the histogram of each sub-block is calculated independently.

[0163] Set a contrast threshold (e.g., 40). If the number of pixels at a certain gray level in the histogram of a sub-block exceeds the threshold, the excess will be evenly distributed to other gray levels to avoid excessive stretching of brightness within the sub-block, which would lead to loss of detail.

[0164] Bilinear interpolation is used to stitch together the equalization results of adjacent sub-blocks to eliminate abrupt brightness changes at the sub-block boundaries and make the brightness transition natural.

[0165] Saturation enhancement for chrominance channels Cb and Cr is expressed as follows:

[0166] ;

[0167] in: is the average value of the Cb and Cr channels of the video frame, reflecting the overall color tone; k is the saturation enhancement coefficient; , The value should be within the range of 0-255 to avoid color overflow and distortion.

[0168] Based on the logic that the human eye judges the distance of objects by the details of the image, the depth value of each pixel in the image (i.e. the distance between the pixel and the virtual camera) is inferred by using the image features of edge contours and color contrast in video frames, and a global depth map is generated.

[0169] The depth estimation model is optimized for common cable TV content (such as news anchors and sports events), reducing the depth value of the anchor's face (to highlight the three-dimensionality) and increasing the background depth value (to enhance the sense of layering); when switching channels quickly (<0.5 seconds), the depth calculation is paused and the 2D to 3D transition frame is directly output to avoid screen stuttering.

[0170] Methods for calculating left-right eye disparity based on depth maps include:

[0171] Based on the pinhole imaging principle, the image height on the virtual camera's imaging plane is inversely proportional to the depth. Combining this with the difference in visual angle between the human eye's pupils, the basic parallax calculation formula is:

[0172] ;

[0173] Where B is the interpupillary distance, f is the camera focal length, D is the object depth, and P is the pixel size;

[0174] Substitute the depth map pixel by pixel into the basic disparity calculation formula to calculate the initial disparity. ;

[0175] When the initial disparity exceeds the range of human eye fusion, it is corrected to ensure that the disparity is reasonable. The initial disparity is adjusted to a usable adaptive disparity through threshold constraints and distortion correction. The formula for calculating the adaptation parallax is:

[0176] ;

[0177] Where k is the distortion correction coefficient. The minimum disparity threshold, This is the maximum disparity threshold.

[0178] Specifically, methods for generating 3D signals for left and right split-screen displays include:

[0179] The cable TV signal is set as a video frame stream in .vr3d data format, and the video frame stream and pixel-level parallax data are encapsulated to generate a 3D image that conforms to human visual habits at the transmitting end. The .vr3d data structure adopts a frame-level encapsulation + block-level storage design to avoid overall parsing failure due to single frame data corruption.

[0180] The logical chain for .vr3d parsing and 3D generation is as follows: .vr3d format data (file header + basic parameter block + frame data block) → data parsing (format verification → parameter reading → frame data extraction → video decoding + parallax restoration) → 3D generation (parallax allocation → left and right eye frame offset → distortion correction → split-screen stitching) → WiFi transmission to VR glasses → users watch 3D live broadcast. The entire process is designed around "low latency, high compatibility, and weak network adaptation" to ensure that both on-site back-row users and off-site users can obtain a stable and distortion-free 3D viewing experience.

[0181] This invention completes a five-step process: VR3D data reading → format verification → parameter parsing → frame data extraction → parallax restoration. It accurately extracts the video stream and parallax data from the .vr3d format to generate a 3D image that conforms to human visual perception. The .vr3d format data parsing includes the following steps:

[0182] Step 1, Data Reception and Initialization: Receive cable TV signals from the set-top box, set the data to .vr3d format, and adopt a "segmented reception + buffering" strategy - store every 1MB of data received in the local buffer to avoid memory overflow;

[0183] Read the first 8 bytes of data. If it is not equal to "VR3D_2025", immediately trigger a "format error" prompt. If the format is correct, continue reading the "basic parameter block offset" in the file header and locate the starting address of the basic parameter block.

[0184] Step 2, Basic Parameter Analysis and Configuration:

[0185] Parameter reading: The "video resolution (W×H), frame rate (F), parallax bit depth (B), maximum parallax threshold (D_max), color space (C)" are parsed sequentially from the basic parameter block and stored in the "global configuration variables" of the APP;

[0186] Step 3: Frame data block location and reading:

[0187] Frame index construction: Obtain the total number of frames N from the "Total Frames" field in the file header, traverse all frame data blocks, extract the "frame number, timestamp, video data length, and parallax data length" of each frame, and construct a "frame index table" (stored in memory) to facilitate quick location by frame number or timestamp later;

[0188] Frame data extraction: Based on the user's viewing progress (e.g., watching frame 100), find the starting address of frame 100 from the frame index table, first read the "video data length (Lv)", then extract Lv bytes of video data; next, read the "parallax data length (Ld)" and extract Ld bytes of parallax data.

[0189] Step 4: Video data decoding and parallax data restoration:

[0190] Call the video decoder to decode the H.265 format video data into 2D video frames in RGB / YCbCr format (resolution consistent with the basic parameter block), and store them as "raw video frame matrix" (dimension: H×W×3, 3 represents RGB three channels).

[0191] Parallax data restoration: If the parallax data bit depth B=8, convert the parallax data binary stream into an "8-bit unsigned integer matrix" (dimension: H×W), where each element represents the parallax value d of the corresponding pixel (0≤d≤D_max);

[0192] If a "parallax compression flag" exists (a new field added to the basic parameter block, 1 represents compression, 0 represents no compression), it is first decompressed by "run-length encoding (RLE)". For example, compressed data "35,5" represents that the parallax value of 5 consecutive pixels is 35, which is restored to [35,35,35,35,35];

[0193] Data alignment: Ensure that the pixel positions of the "original video frame matrix" and the "disparity matrix" correspond one-to-one (i.e., the pixel in the i-th row and j-th column of the video frame corresponds to the disparity value in the i-th row and j-th column of the disparity matrix) to avoid misalignment during 3D generation.

[0194] After obtaining the "original video frame matrix + disparity matrix", the difference between human binocular vision is simulated. Through the steps of "disparity allocation → left and right eye frame generation → distortion correction → image output", the 2D video frames are converted into 3D images suitable for VR glasses display.

[0195] The process of generating 3D images includes the following steps:

[0196] Step 1, Disparity Assignment (Determine the direction and amount of left and right eye offset):

[0197] Based on the principle of human binocular vision, the left eye observes objects from a perspective biased to the right, while the right eye observes them from a perspective biased to the left. Therefore, it is necessary to perform a "directional offset" on the pixels of the original video frame, with the offset amount proportionally allocated based on the parallax value:

[0198] ,

[0199] ;

[0200] Where d is the disparity value of a pixel in the disparity matrix (0≤d≤35 pixels, from the parsed disparity matrix); This is the rightward offset of the pixel in the left-eye frame (in pixels), rounded down (e.g., when d=35). =17); This is the leftward offset of the pixel in the right-eye frame (unit: pixels), such as when d=35. =18;

[0201] Step 2, Generation of left and right eye frames (pixel offset and boundary padding):

[0202] Left eye frame generation: Initialize the "left eye frame matrix" (dimension: H×W×3, consistent with the original video frame);

[0203] For each pixel (i,j) in the original video frame (where i is the row number and j is the column number), calculate its new column number in the left-eye frame: = j + ;

[0204] like If W (not exceeding the right boundary of the image), then the RGB value of the original pixel (i,j) is assigned to the left eye frame (i,j). );

[0205] like For pixels ≥ W (exceeding the right boundary), the "edge pixel copying" strategy is used—the RGB value of the rightmost pixel (i, W-1) in the original frame is assigned to the left eye frame (i, W-1). - W), to avoid black borders;

[0206] Right eye frame generation: Initialize the "right eye frame matrix" (dimension: H×W×3);

[0207] For each pixel (i,j) in the original video frame, calculate its new column number in the right-eye frame: ;

[0208] like If the pixel does not exceed the left edge of the frame, then the RGB value of the original pixel (i,j) is assigned to the right eye frame (i,j_right).

[0209] like (Exceeding the left boundary), the same "edge pixel copying" method is used—the RGB value of the leftmost pixel (i,0) of the original frame is assigned to the right eye frame (i, + W);

[0210] To avoid "holes" after adjacent pixels are shifted (such as when a pixel is shifted and there are no pixels filling the original position), "bilinear interpolation filling" is used for the "unassigned pixels" in the left eye frame and the "unassigned pixels" in the right eye frame. For example, when (i,j) in the left eye frame is unassigned, the average RGB value of the four surrounding assigned pixels is taken as the filling value to ensure smooth image.

[0211] Step 3, Distortion Correction: Due to the "barrel distortion" (stretched edge pixels) of the Fresnel lens of the VR glasses, if the left and right eye frames are directly output, the user will see a distorted 3D picture. Therefore, it is corrected by a polynomial distortion correction algorithm, including: (1) Obtaining glasses parameters: Read the "distortion coefficients" (such as k1=-0.3, k2=0.1, k3=-0.05, representing third-order distortion coefficients) and "optical center" [such as (W / 2, H / 2, i.e., the center of the picture] of the currently connected VR glasses from the "VR device parameter library" (built-in distortion parameters of mainstream VR glasses models);

[0212] (2) Distortion coordinate calculation: For each pixel (x, y) (screen coordinates) in the left and right eye frames, calculate the ideal coordinates of the pixel (x, y) before distortion using the following formula.

[0213] ,

[0214] ,

[0215] ;

[0216] Where (x,y) are the pixel coordinates (screen coordinates, range 0≤x) of the corrected left and right eye frames. <W,0≤y<H);

[0217] ( ) represents the optical center coordinates of the VR glasses (usually the center of the screen, such as (960, 540) corresponding to a 1920×1080 resolution);

[0218] r is the normalized distance from the pixel to the optical center (dimensionless, range 0-1).

[0219] k1, k2, and k3 are distortion coefficients (determined by the VR glasses hardware; for example, k1 = -0.3 represents first-order barrel distortion correction).

[0220] ( The coordinates are the ideal pixel coordinates before distortion (which need to be mapped to the coordinate range of the original left and right eye frames);

[0221] The OpenCV remap function is called to remap the left and right eye frames according to the mapping relationship between "ideal coordinates and screen coordinates" to generate distortion-corrected left and right eye frames, ensuring that users see distortion-free 3D images through VR glasses.

[0222] Step 4: Synchronize 3D image output with VR glasses:

[0223] (1) Image format adaptation: The distortion-corrected left and right eye frames are stitched together according to the "3D display format" supported by the VR glasses [the present invention uses the "side-by-side" format by default] - that is, the left eye frame is placed on the left and the right eye frame is placed on the right to generate a "single frame 3D image" [resolution: H×(2W), such as 1080×3840];

[0224] (2) Low-latency transmission: The "single frame 3D image" is transmitted to the VR glasses. During the transmission process, the "UDP protocol + frame sequence number mark" is used to ensure that each frame of data arrives in order and the transmission delay is controlled within ≤20ms;

[0225] (3) Audio-visual synchronization: Read the "timestamp" in the frame data block and compare it with the audio playback timestamp of the VR glasses. If the video frame timestamp is ahead of the audio playback timestamp by more than 50ms, then pause the video transmission for 10ms; if it is behind by more than 50ms, then speed up the video decoding speed (such as skipping 1 non-key frame) to ensure that the audio-visual synchronization error is ≤30ms and avoid the user's discomfort of "the picture and sound are out of sync".

[0226] S3. The signal transmitting unit packages the 3D signal into a 60GHz millimeter-wave data packet and transmits it to the receiving end;

[0227] For the receiver placed in the bedroom, it can transmit towards the bedroom via a directional antenna, with a transmission delay of ≤25ms and a transmission distance of ≥15 meters in open environments;

[0228] S4. Power on the receiver (connect the receiver to the power bank via the Type-C male connector; the indicator light will turn blue) and search for the millimeter-wave signal from the transmitter; adjust the pairing signal receiving unit and signal transmitting unit (this can be determined by the indicator light; a solid blue light indicates successful pairing); the signal demodulation chip will restore the received millimeter-wave signal into a 3D video stream.

[0229] S5. Transmit the 3D video stream (via Type-C female connector) to the VR glasses, power the VR glasses through the power distribution circuit, and watch the 3D video stream TV content while wearing the glasses.

[0230] In practical applications, this embodiment can be moved to any location in the bedroom within 8 meters of the set-top box without signal interruption.

[0231] This invention, through a split architecture of "independent transmitter + portable receiver", achieves a complete link of "set-top box 2D signal → transmitter 2D to 3D conversion → wireless transmission → receiver output → VR glasses display". Taking home viewing in the bedroom and setting up the set-top box in the living room as an example, the set-top box is connected to the transmitter via an HDMI cable, the transmitter is connected to the receiver via a millimeter-wave wireless signal, the receiver is connected to the VR glasses via a Type-C cable, and the power bank is connected to the receiver via a Type-C cable, forming a complete closed loop of "signal input - processing - transmission - display - power supply".

[0232] The complete workflow of this embodiment in field application includes:

[0233] 1. Deployment phase: Fix the transmitter next to the set-top box in the living room, connect the HDMI cable and power adapter, and adjust the transmitting antenna to face the bedroom;

[0234] 2. Reception preparation: In the bedroom, the user connects the receiver to the power bank and VR glasses, and adjusts the receiving antenna until the blue light is constantly on;

[0235] 3. Signal processing and transmission: The transmitter converts the 2D signal from the set-top box into a 3D signal in real time and transmits it to the receiver via millimeter waves;

[0236] 4. Viewing stage: The VR glasses receive and display 3D signals. Users can watch while lying in bed. The power bank supports continuous use for 4 hours. A yellow light flashes to remind users when the battery is low.

[0237] This embodiment offers superior 3D effects compared to existing technologies. Existing HDMI wireless transmitters can only transmit 2D signals, while the transmitter of this invention can natively convert 2D to 3D, generating a realistic stereoscopic effect through depth estimation, improving the user's subjective stereoscopic perception score by 80%. It also boasts lower latency; existing solutions have a 2.4GHz WiFi latency ≥200ms, while this embodiment's millimeter-wave transmission latency is ≤25ms, with a total latency ≤50ms, eliminating audio-visual desynchronization issues. Furthermore, it offers better portability; existing receivers require separate power and adapters, while this embodiment's receiver weighs only 28g, supports power bank charging, and can be fixed in any location, freeing users from cable constraints during movement. Operation is also simpler; existing solutions require manually switching to a mobile app and adjusting VR glasses parameters, while this embodiment only requires two steps: "connect the set-top box to the transmitter → connect the receiver to the glasses + power bank," making it easy for the elderly and children to use quickly. Figure 1 This embodiment illustrates the basic process for implementing VR viewing of television signals.

[0238] A second embodiment of the present invention also provides a TV signal VR viewing system that connects to VR glasses via a wired connection. It directly connects to a set-top box via an HDMI interface, and the built-in system converts the 2D cable TV signal into a VR 3D stereoscopic signal in real time, enabling the viewing of 3D cable TV programs with VR glasses. Through an integrated design of "HDMI wired access module + 2D to 3D real-time processing unit + VR stereoscopic display system," VR 3D viewing of cable TV signals is achieved. The hardware structure's interface and signal receiving module consist of the following components: an HDMI 2.1 input interface, a DC 5V / 2A power supply interface, and a signal buffer chip. The HDMI interface connects to the input of the signal buffer chip via an HDMI signal cable, and the output of the signal buffer chip connects to the main processor. The DC 5V / 2A power supply interface powers the entire device through a power management chip and is electrically isolated from the HDMI signal circuit.

[0239] The HDMI 2.1 input interface in this second embodiment can be directly plugged into the HDMI output port of the set-top box via an HDMI signal cable, without the need for an adapter, and supports hot-swapping; it is compatible with mainstream brand set-top boxes (such as Skyworth, Huawei, and cable TV set-top boxes) and automatically recognizes the signal format (720P / 1080P / 4K). The process of realizing VR viewing of TV signals is as follows:

[0240] 1. Receive 2D video frames (such as news or TV dramas) output by the set-top box;

[0241] 2. Depth estimation: By analyzing image edges and color contrast, a depth map is generated (deep values ​​are low for near objects and high for distant objects).

[0242] 3. Parallax generation: Calculate the left and right eye parallax based on the depth map (the parallax is larger for near scenes and smaller for distant scenes), and generate split-screen images for the left and right eyes;

[0243] 4. Depth Adjustment: Users can adjust the depth of field level (1-5 levels, corresponding to parallax of 0-30 pixels) via buttons to adapt to different content (e.g., high depth of field for sports events, low depth of field for news).

[0244] The signal processing flow includes: set-top box HDMI signal → HDMI interface → buffer chip → main processor decoding → 2D to 3D algorithm processing → GPU distortion correction → dual-screen display → lens imaging for 3D effect.

[0245] This second embodiment requires no network connection and is directly compatible with mainstream set-top box HDMI outputs; 2D to 3D conversion latency is ≤100ms, and 3D stereoscopic effect is adjustable (depth of field 0-5 levels); it uses an independent HDMI input + DC power supply interface, ensuring stable connection without adapters, and is simple to operate, adapting to the usage habits of elderly users and traditional TV viewers. Its main advantages are as follows:

[0246] 1. It adopts a wired connection with an HDMI 2.1 interface, directly connects to the set-top box HDMI, is compatible with set-top box signals, requires no network or adapter, covers more than 90% of home cable TV users, and solves the compatibility problem between VR glasses and traditional TV signals;

[0247] 2. Based on a real-time 2D to 3D conversion algorithm for cable TV content, through dynamic depth estimation and parallax adjustment, ordinary TV signals are converted into VR stereoscopic effects. All cable TV programs can be viewed in 3D, which is different from existing 3D TVs that rely on dedicated content sources, making the 3D experience more flexible.

[0248] 3. It adopts dual interfaces (HDMI + DC power supply) and dual control methods (physical buttons + infrared remote control), retains the channel switching logic of traditional remote control, and adjusts the depth of field with physical buttons, making operation simpler. Users do not need to learn new operations, adapting to the operating habits of elderly users and solving the problem of complex operation of VR devices.

[0249] 4. The optical system is lightweight (100° field of view + adjustable interpupillary distance), weighing less than 300g, improving comfort; the 100° field of view far exceeds that of ordinary TVs (about 50°), and the 2560×2560 resolution avoids graininess, achieving a superior immersive 3D display, with a viewing experience close to a "private 3D cinema".

[0250] The technical solution of the present invention has been described in conjunction with preferred embodiments. However, it will be readily understood by those skilled in the art that the scope of protection of the present invention is obviously not limited to these specific embodiments. Without departing from the principles of the present invention, those skilled in the art can make equivalent changes or substitutions to the relevant technical features, and the technical solutions resulting from these changes or substitutions will all fall within the scope of protection of the present invention.

[0251] The above description is merely a preferred embodiment of the present invention and is not intended to limit the invention. Various modifications and variations can be made to the present invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.

Claims

1. A system for VR viewing implementation of a television signal based on 2D to 3D processing, characterized by, include: The device includes a transmitter connected to a TV set-top box and a portable receiver. The transmitter includes an HDMI signal input unit, a 2D-to-3D processing unit, and a signal transmission unit connected in sequence. The 2D-to-3D processing unit includes a quad-core processor, an NPU processing unit, and an algorithm storage chip connected in sequence. The receiver includes a signal receiving unit and a signal output unit connected in sequence. The signal receiving unit is connected to the signal transmission unit, and the signal output unit is connected to VR glasses.

2. The VR viewing system for television signals based on 2D-to-3D conversion processing according to claim 1, characterized in that, The HDMI signal access unit includes an HDMI 2.0 input interface for connecting a TV set-top box, an HDCP 2.2 decryption chip for ensuring the secure transmission of protected video content and preventing unauthorized copying or tampering, and a signal detection circuit for automatically identifying the input resolution and frame rate, all connected in sequence. The signal detection circuit is connected to the quad-core processor.

3. The VR viewing system for television signals based on 2D-to-3D conversion processing according to claim 2, characterized in that, The signal transmitting unit includes: a 60GHz millimeter-wave transmitting circuit and a 360° rotatable directional adjustable antenna connected in sequence. The 60GHz millimeter-wave transmitting circuit is connected to an adaptive frequency hopping anti-interference module.

4. The VR viewing system for television signals based on 2D-to-3D conversion processing according to claim 1, characterized in that, The signal receiving unit includes a 60GHz millimeter-wave receiving circuit, a receiving antenna, and a signal demodulation chip that restores the millimeter-wave signal to a 3D video stream, connected in sequence; the signal output unit is provided with a Type-C female connector for connecting the Type-C interface of VR glasses.

5. The VR viewing system for television signals based on 2D-to-3D conversion processing according to claim 1, characterized in that, The transmitter is equipped with a DC power interface, which is connected to an external power adapter. The receiver is equipped with a power supply unit, which includes a Type-C male connector and a power distribution circuit connected in sequence. The power distribution circuit divides the power into two paths: one path powers the receiver itself, and the other path powers the VR glasses.

6. A method for VR viewing of television signals based on 2D-to-3D conversion processing, applied to the VR viewing system for television signals based on 2D-to-3D conversion processing as described in any one of claims 1-5, characterized in that, Includes the following steps: S1. The 2D HDMI signal output from the set-top box is connected to the transmitter. After being decrypted by the HDCP 2.2 decryption chip, the signal parameters are identified by the signal detection circuit, and the 2D signal is transmitted to the quad-core processor. S2. The quad-core processor calls the NPU processing unit to run the 2D to 3D algorithm, generates a depth map through the edge of the screen and color contrast, calculates the left and right eye parallax based on the depth map, and generates a 3D signal for left and right split screen. S3. The signal transmitting unit packages the 3D signal into a 60GHz millimeter-wave data packet and transmits it to the receiving end; S4. The receiver is powered on and searches for millimeter-wave signals from the transmitter; the paired signal receiving unit and signal transmitting unit are adjusted, and the signal demodulation chip restores the received millimeter-wave signal into a 3D video stream; S5. Transmit the 3D video stream to the VR glasses, power the VR glasses through the power distribution circuit, and watch the 3D video stream TV content by wearing the glasses.

7. The method for VR viewing of television signals based on 2D-to-3D conversion processing according to claim 6, characterized in that, The method for generating a depth map through image edges and color contrast in step S2 includes: S21. Extract the edge contours of the 2D image through multi-stage edge detection and adaptive filtering, including: To address potential light flicker and camera noise in 2D videos, a 3×3 Gaussian filter kernel is used to smooth and denoise the video frames, balancing denoising effectiveness with edge preservation. Convolution operations are employed to eliminate high-frequency noise, preventing noise from being misidentified as edges. The expression is as follows: ; Where (x, y) represents the pixel; =1.2; Global brightness normalization is performed on video frames, mapping pixel brightness values ​​of 0-255 to the range of [0.1, 0.9] to prevent the edges of overly bright / dark areas from being masked by extreme brightness values; An optimized Canny edge detection algorithm is used, and the gradient magnitude and direction of each pixel are calculated using the Sobel operator in the x and y directions: gradient magnitude gradient direction = ; Among them, horizontal gradient Reflecting vertical edges and vertical gradients Reflects horizontal edges, ensuring that both the horizontal and vertical contours of the object are captured simultaneously; Local maxima are preserved along the gradient direction. If the gradient magnitude of a pixel is greater than that of its two adjacent pixels along the gradient direction, it is determined to be an edge and the pixel is preserved. If the gradient magnitude of a pixel is not greater than that of its two adjacent pixels along the gradient direction, it is determined to be a non-edge and non-maxima are suppressed to prevent the edges from widening. Dual-threshold edge connection: Set a high threshold and a low threshold. For pixels with a gradient magnitude greater than the high threshold, they are directly identified as strong edges. For pixels with a gradient magnitude less than the low threshold, they are directly identified as non-edges. For pixels with a gradient magnitude between the high threshold and the low threshold, if they are connected to strong edges, they are identified as weak edges. If they are not connected to strong edges, they are discarded to ensure the continuity and integrity of edges. S22. Optimize color contrast through "scene-based contrast enhancement + dynamic range adjustment", strengthen the difference between light and dark / color, and assist in depth layering, including: Convert the color space, transforming the video frames from the RGB color space to the YCbCr color space, separating luminance and chrominance. The conversion formula is as follows: Y =0.299R+0.587G + 0.114B Cb =-0.1687R-0.3313G + 0.5B + ​​128 Cr = 0.5R -0.4187G - 0.0813B + 128; Where Y is the luminance channel, Cb is the blue difference channel, and Cr is the red difference channel; R, G, and B are the red, green, and blue primary colors, respectively. Adaptive histogram equalization (CLAHE) is used to enhance brightness and contrast, and regional dynamic adjustments are used to avoid localized overexposure caused by global adjustments, including: The luminance channel Y is divided into 8×8 sub-blocks, each sub-block is 240×144 pixels, adapted to 1920×1080 resolution, and the histogram of each sub-block is calculated independently. Set a contrast threshold. If the number of pixels at a certain gray level in the histogram of a sub-block exceeds the threshold, the excess will be evenly distributed to other gray levels to avoid excessive stretching of brightness within the sub-block, which would lead to loss of detail. Bilinear interpolation is used to stitch together the equalization results of adjacent sub-blocks to eliminate abrupt brightness changes at the sub-block boundaries and make the brightness transition natural. Saturation enhancement for chrominance channels Cb and Cr is expressed as follows: ; in: is the average value of the Cb and Cr channels of the video frame, reflecting the overall color tone; k is the saturation enhancement coefficient; , The value should be within the range of 0-255 to avoid color overflow and distortion. Based on the logic that the human eye judges the distance of objects by the details in the image, the depth map of the entire field is generated by using the image features of edge contours and color contrast in video frames to infer the depth value of each pixel in the image.

8. The method for VR viewing of television signals based on 2D-to-3D conversion processing according to claim 6, characterized in that, The method for calculating the left and right eye disparity based on the depth map in step S2 includes: Based on the pinhole imaging principle, the image height on the virtual camera's imaging plane is inversely proportional to the depth. Combining this with the difference in visual angle between the human eye's pupils, the basic parallax calculation formula is: ; Where B is the interpupillary distance, f is the camera focal length, D is the object depth, and P is the pixel size; Substitute the depth map pixel by pixel into the basic disparity calculation formula to calculate the initial disparity. ; When the initial disparity exceeds the range of human eye fusion, it is corrected to ensure that the disparity is reasonable. The initial disparity is adjusted to a usable adaptive disparity through threshold constraints and distortion correction. The formula for calculating the adaptation parallax is: ; Where k is the distortion correction coefficient. The minimum disparity threshold, This is the maximum disparity threshold.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the steps of the method for VR viewing of television signals based on 2D-to-3D processing as described in any one of claims 6-8.

10. A computer device, the computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the steps of the method for VR viewing of television signals based on 2D-to-3D processing as described in any one of claims 6-8.