3D video coding plug stream generation method based on edge calculation

By adopting the 3D video encoding streaming generation method based on edge computing in VR live broadcast, using lidar to obtain three-dimensional signals and combining the technology of video processing module, the problem of poor user experience under bandwidth compression in the existing VR live broadcast technology is solved, and high-quality 3D video stream generation and user experience improvement is achieved.

CN120075419APending Publication Date: 2025-05-30FUSHI ZHITONG ELECTRONIC TECH (JINAN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411904306.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Under the cost of network bandwidth, existing VR live broadcast technology often compresses the transmission code rate, resulting in unclear user experience, dizziness, poor interactivity and lack of social elements.

Method used

The 3D video encoding streaming generation method based on edge computing is adopted, and 2D video and audio input is obtained through the video acquisition device, three-dimensional stereo signals are obtained using lidar technology, and processed and integrated through the video processing module and the edge computing processing video module to generate a 3D video stream that can be directly pushed.

Benefits of technology

It realizes the generation of high-quality 3D video streams without relying on a 360° panoramic camera, which improves the user experience of VR live broadcast, simplifies the process, and supports the splicing and fusion of multiple 3D video formats.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120075419A_ABST
    Figure CN120075419A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of video processing, in particular to a 3D video coding plug stream generation method based on edge calculation, which comprises the following steps: performing 2D video input and audio input by video acquisition equipment; when the left eye and the right eye of a person are simulated to observe the same object, due to the difference of parallax angles, the person can perceive the depth information of a picture so as to obtain a three-dimensional signal; image information of 2D video input and audio input is processed through a video processing module; the video processed by the video processing module is transmitted to an edge calculation processing video module, and a 2D plane and related three-dimensional information are integrated into a flow chart; the 3D video is subjected to stream pushing through the video stream pushing module, and a video stream which can be directly played by VR eyes or a VR live broadcast platform is produced. According to the invention, related processes of VR live broadcast are simplified, splicing of 3D video formats in various different modes can be carried out, and video images are successfully spliced, and videos are superposed and fused to a three-dimensional scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of video processing, and particularly to a 3D video encoding push stream generation method based on edge computing. Background Art

[0002] In the era of mobile Internet, whether it is text, pictures or videos, they can all be live-streamed through the network. Network live-streaming not only compresses time and space, enabling users to watch sports events, major activities and news in real time, but also can add functions such as text chat and gift-giving, converting the traditional one-way transmission into two-way interaction and increasing the interest of live-streaming. However, the lack of a sense of presence and the monotony of interaction are the biggest deficiencies of traditional 2D live-streaming, and the emergence of VR live-streaming technology has solved this problem, promoting the user live-streaming experience to a new level.

[0003] VR live-streaming has been widely used in scenarios such as sports events, hot news, concerts, press conferences, etc. By live-streaming through VR technology, users can feel as if they are on the scene. VR live-streaming requires a greater network bandwidth compared to flat 2D videos. However, due to cost factors of network bandwidth, many VR live-streaming platforms will compress the transmission bit rate to reduce costs, resulting in an unclear and dizzy feeling for users, greatly reducing the experience. Coupled with the single form of VR live-streaming experience, users can only watch passively and lack interaction, making the experience dull. However, with the rapid development of cloud computing, 5G, and gigabit broadband networks, the improvement of VR live-streaming technology, and the integration of social elements, the VR live-streaming experience has been greatly improved, providing feasibility for the commercial implementation of VR live-streaming. Through VR live-streaming shooting technology and near-eye display technology of VR terminals, users can obtain a more immersive viewing experience with a larger field of view. Currently, most VR live-streams on the market rely on 360° panoramic cameras for push streaming, which are relatively expensive and inconvenient. Summary of the Invention

[0004] The object of the present invention is to overcome the deficiencies of the prior art and propose a 3D video encoding push stream generation method based on edge computing.

[0005] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0006] A 3D video encoding push stream generation method based on edge computing, comprising:

[0007] The video acquisition device performs 2D video input and audio input to capture basic elements, provides coordinates in the X and Y directions for the test target, and provides a corrected orientation angle and depth for the stereoscopic video; according to the fact that when a person observes the same thing with the left and right eyes, due to the difference in the parallax angle, a person can perceive the depth information of the picture to obtain a three-dimensional stereo signal;

[0008] The video processing module processes the image information of the 2D video input and the audio input and removes noise; specifically, for several groups of pictures (GOPs), each GOP represents a group of consecutive video frames. When performing predictive coding and transform coding on the images, the images will be first divided. The division method is a quadtree. When dividing the quadtree, the entire video frame will be divided into several square coding tree blocks (CTBs). The CTB can be further divided into coding blocks (CBs), and the CB can also be divided into prediction blocks (PBs) and transform blocks (TBs).

[0009] The video processed by the video processing module is transmitted to the edge computing video processing module. The edge computing video processing module uses AI technology to integrate the video information and integrates the 2D plane and related three-dimensional information into a flowchart.

[0010] The 3D video is pushed through the video streaming module to produce a video stream that can be directly played by VR glasses or a VR live broadcast platform.

[0011] Furthermore, a lidar technology is used to detect the three-dimensional stereo signal. The specific process is as follows:

[0012] Analyze information such as the magnitude of the reflected energy on the surface of the target object, the amplitude, frequency, and phase of the reflected spectrum, and output a point cloud to present the precise three-dimensional structure information of the target object.

[0013] Multiple lasers are emitted at the set angles to obtain multiple reflected signals based on obstacles; combined with the time range, the scanning angle of the laser, and INS information, after data processing, this information is combined with the X, Y, Z coordinates to form a three-dimensional stereo signal with distance information, spatial position information, etc.

[0014] Furthermore, each video acquisition device can perform 2D video input and audio input. The method for video and audio synchronization between multiple video acquisition devices is: synchronize based on the I frame (the first frame of each GOP).

[0015] Furthermore, the processing flow of the edge computing video processing module is as follows:

[0016] Image distortion calibration: Establish corresponding mathematical model parameters according to the distortion method and restore the image.

[0017] Coordinate conversion: Process the models of the left and right eye perspectives, and convert the 2D coordinates and the X, Y, Z and other three-dimensional information of the three-dimensional view captured into coordinates for processing the 3D perspective.

[0018] Image registration: The areas collected by adjacent cameras will overlap. In the overlapping area, the best image registration is found by looking for suitable stitching points.

[0019] Video fusion: Multiple registered images are fused to form a complete image.

[0020] Audio and color calibration: Synchronize audio and video and calibrate translated subtitles, and compare and calibrate the original image color with the generated video color.

[0021] Technical effects of the present invention:

[0022] Compared with the prior art, a 3D video encoding push stream generation method based on edge computing of the present invention integrates 3D video into encoding based on AI processing of edge computing, simplifies the related processes of VR live broadcast; can simulate the binocular mode of human eyes to splice 3D video formats in multiple different modes, and can be directly applied to relevant platforms supporting VR live broadcast for viewing, successfully splicing, superimposing and fusing video images into a three-dimensional scene, without relying on a 360° panoramic camera for push stream, and is easy to use. Description of the drawings

[0023] Figure 1 It is a flowchart of the 3D video encoding push stream generation method based on edge computing of the present invention;

[0024] Figure 2 It is a flowchart of the video frame processing method of the present invention;

[0025] Figure 3 It is a flowchart of the processing method of the edge computing video processing module of the present invention. Detailed implementation manners

[0026] To make the objectives, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the specification.

[0027] Embodiment 1:

[0028] As Figure 1 shown, a 3D video encoding push stream generation method based on edge computing involved in this embodiment includes:

[0029] The 2D video input and audio input are performed by a video acquisition device. Its main functions are to input ordinary audio and video, capture the basic elements of the overall environment, avoid problems and deviations during edge operations to provide a calibrated direction; avoid information loss and missing important scenes when processing 3D video input; provide the coordinates in the X and Y directions for the test target object, and provide the corrected orientation angle and depth for the stereoscopic video.

[0030] When a person observes the same object with the left and right eyes, due to the difference in the parallax angle, the person can perceive the depth information of the picture to obtain a three-dimensional stereo signal; this process uses lidar technology to detect, analyze information such as the magnitude of the reflected energy on the surface of the target object, the amplitude, frequency, and phase of the reflected wave spectrum, and outputs a point cloud, thereby presenting the precise three-dimensional structure information of the target object; multiple lasers are emitted at a set angle to obtain multiple reflected signals based on obstacles; then, in combination with the time range, the scanning angle of the laser, and INS information, after data processing, these information are combined with the X, Y, and Z coordinates to form a three-dimensional stereo signal with distance information, spatial position information, etc.; specifically, time information: as a synchronization mechanism for simulating the binocular perspective, positioning, synchronization, and calibration are performed according to time, with an accuracy of up to milliseconds. The steps for processing according to the dual-perspective model embedded in edge computing are as follows: A. Synchronize the left and right perspectives according to time; B. The dual-perspective model processes the 3D information of the left and right eye perspectives respectively: collect the relevant information transmitted by the relevant radar; calculate the three-dimensional coordinates (X, Y, Z) through the model; C. Integrate these coordinate information into a three-dimensional stereo signal and provide it to the next step. The specific process for calculating the three-dimensional coordinate model using the parameters obtained through binocular calibration is as follows: 1. Data extraction: Extract the time information and integrate the data obtained from the front-end radar; 2. Model training: According to the AI capabilities of edge computing, train the perspective models of the left and right eyes; use the pre-trained perspective models, which can reduce the generation time; in this step, a preliminary 3D diagram is generated. First, the radar information is used to generate individual points, and then countless MARK points are connected to form an outline; 3. Feature extraction: In this step, based on data extraction and model training, a 3D wireframe is generated, and then the three-dimensional coordinate model is calculated using the parameters obtained through binocular calibration. Finally, the coordinate features of the left and right dual perspectives are extracted; 4. Integrate information: Integrate the coordinate information into a three-dimensional stereo signal.

[0031] The video processing module processes the image information of the 2D video input and audio input and removes noise; specifically, the video frames are processed, such as Figure 2As shown, for several groups of pictures (GOPs), each GOP represents a group of consecutive video frames. When performing predictive coding and transform coding on images, the images will first be divided. The division method is a quadtree. When dividing the quadtree, the entire video frame will be divided into several square coding tree blocks (CTBs). A CTB can be further divided into coding blocks (CBs), and a CB can also be divided into prediction blocks (PBs) and transform blocks (TBs), as Figure 2 shown; each video acquisition device can perform 2D video input and audio input. There is a possibility that video and audio may be out of sync among multiple video acquisition devices. For the phenomena of frame skipping and picture desynchronization that occur during this process, I-frame (the first frame of each GOP) synchronization can be performed based on the I-frame to avoid the influence of two different pictures on the follow-up; for image denoising: due to the complex and variable environment during video acquisition, noise inevitably exists in the acquired images. To avoid affecting the 3D video effect, the images need to be denoised; then the audio information, the original image, the PBs and TBs of the decomposed image frames, and the three-dimensional information of the relevant stereo coordinates are transmitted to the edge computing video processing module;

[0032] The edge computing video processing module uses AI technology to integrate the video information and integrates the 2D plane and the relevant three-dimensional information into a flowchart;

[0033] The 3D video is pushed through the video streaming module to generate a video stream that can be directly played by a VR headset or a VR live streaming platform. 5G and NET can be used for video transmission during transmission.

[0034] As Figure 3 shown, the processing flow of the edge computing video processing module is as follows:

[0035] Image distortion calibration: Establish corresponding mathematical model parameters according to the distortion method to restore the image, so as to solve the corresponding distortion problems caused by the camera for image acquisition or the shooting angle; Image distortion is divided into pincushion distortion and barrel distortion, that is, the indentation and protrusion of the planar graph. Image distortion calibration is to establish corresponding mathematical model parameters according to the distortion method and convert the distorted image into a flat picture; Three models are embedded: 1. Image detection model: used to check whether the graph is distorted; 2. Pincushion distortion correction model: process the indented picture into a planar image; 3. Barrel distortion: process the protruding picture into a planar image; The main reasons are the camera shooting angle and installation, resulting in distortion in some parts of the image; This process is to detect the distorted image and restore the distorted image; In this process, if the main image detection model detects a distorted image, it will also generate corresponding distortion coefficients. In the subsequent distortion correction process, the image will be corrected and calibrated according to the distortion coefficients to generate a normal planar image;

[0036] Coordinate transformation: Coordinate transformation of the left and right views. In this process, there are models for processing the left and right eye perspectives, which will convert the 2D coordinates and the X, Y, Z and other three-dimensional information of the three-dimensional view captured into the coordinates for processing the 3D perspective. This process uses the AI processing ability in edge computing to establish relevant perspective models, integrate relevant coordinates to generate 3D coordinates, and prepare for the next step of generating 3D videos;

[0037] Image registration: The areas collected by adjacent cameras will overlap, leaving enough redundancy at the picture splicing area for splicing. Image registration requires finding suitable splicing and stitching points in the overlapping area; Image registration involves the selection of feature space, similarity measure, search space and search strategy, and finding the splicing and stitching points is essentially a process of finding similar feature points within the search space according to the search strategy; Among them, feature space: When performing image registration, due to the large amount of original image data and too much redundant information, the original data is generally not processed, but the feature data of the image is extracted. Therefore, a suitable feature space needs to be selected, such as gray scale, brightness, contour, etc.; Search space: The search for similar feature points is targeted, and the position range where registration may occur needs to be locked; Search strategy: A suitable search strategy needs to be selected during the search process. For example, for landscape graphs, through intelligent recognition and corresponding landscape models, the process can be simplified, and suitable image registration can be found, etc., to minimize unnecessary searches, optimize the calculation amount, and improve the search efficiency;

[0038] Video fusion: Fuse multiple registered images to form a complete image; for image fusion, one image needs to be selected as the reference image, and the data of the overlapping areas of the other registered images are calculated through a fusion algorithm and unified into the coordinate space where the reference image is located; fusion algorithms include direct averaging method, weighted averaging method, distance weighting method, etc.

[0039] Audio and color calibration: Synchronize audio and video and calibrate translated subtitles, and compare and calibrate the colors extracted from the original image and the colors of the generated video.

[0040] The edge computing technology of the present invention is a distributed open platform that integrates the core capabilities of network, computing, storage, and application on the network edge side close to the object or data source, and provides edge intelligent services nearby. Simply put, edge computing directly analyzes the data collected from the terminal in the local device or network close to where the data is generated, without the need to transmit the data to the cloud data processing center; it can simulate the image acquisition of both eyes and perceive the depth information of the picture, and through video processing and AI edge computing, splice and recombine the three-dimensional stereoscopic information and 2D video frame pictures, and can fuse them into 3D videos for VR live broadcast; AI edge computing can not only perform video recombination information, but also perform speech recognition, subtitle translation, and related functions of AI-related training and intelligent recognition; support multiple 3D video formats: full-height up-and-down splicing format; half-height up-and-down splicing format; half-width left-and-right splicing format; full-width left-and-right splicing format; half-width left-and-right splicing format (binocular mode).

[0041] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any appropriate changes or modifications made by any person of ordinary skill in the art who meets the claims of the present invention shall fall within the patent protection scope of the present invention.

Claims

1. A 3D video encoding streaming generation method based on edge computing, characterized in that: include: The video acquisition device performs 2D video input and audio input to capture basic elements, provide X and Y coordinates for the test target, and provide corrected orientation angle and depth for the stereo video; based on the difference in parallax angle when the left and right eyes of a simulated person observe the same object, the person perceives the depth information of the picture to obtain a three-dimensional signal; The video processing module processes the image information of the 2D video input and the audio input and removes noise. Specifically, for a number of groups of images GOP, each GOP represents a group of continuous video frames. When performing predictive coding and transform coding on the image, the image is divided in a quadtree manner. When dividing the quadtree, the entire video frame is divided into a number of square coding tree blocks CTB, and the CTB is further divided into coding blocks CB, and the CB is divided into prediction blocks PB and transform blocks TB. The video processed by the video processing module is transmitted to the edge computing video processing module. The edge computing video processing module integrates the video information using AI technology and integrates the 2D plane and related three-dimensional information into a flow chart; The 3D video is pushed through the video streaming module to produce a video stream that can be directly played by VR glasses or VR live broadcast platforms.

2. The 3D video encoding streaming generation method based on edge computing according to claim 1 is characterized in that: The three-dimensional signal is detected by using laser radar technology, and the specific process is as follows: Analyze the reflected energy, amplitude, frequency and phase information of the reflected spectrum on the surface of the target object, and output a point cloud, thereby presenting accurate three-dimensional structural information of the target object; Multiple lasers are emitted at set angles to obtain multiple reflected signals based on obstacles; combined with the time range, laser scanning angle, INS information, and after data processing, combined with X, Y, and Z coordinates, a three-dimensional signal with distance information and spatial position information is formed.

3. The 3D video encoding streaming generation method based on edge computing according to claim 1 is characterized in that: Each video acquisition device can perform 2D video input and audio input. The method for synchronizing the video and audio between the multiple video acquisition devices is: synchronizing the I frame according to the visual I frame.

4. The 3D video encoding streaming generation method based on edge computing according to claim 1 is characterized in that: The processing flow of the edge computing video processing module is as follows: Image distortion calibration: Establish corresponding mathematical model parameters according to the distortion mode to restore the image; Coordinate conversion: After processing the left and right eye perspective models, convert the captured 2D coordinates and the X, Y, Z and other 3D information of the 3D view into coordinates for processing the 3D perspective. Image registration: The areas captured by adjacent cameras will overlap. In the overlapping areas, the best image registration can be found by finding suitable stitching points. Video fusion: fuse multiple registered images into a complete image; Audio and color calibration: audio and video synchronization and translation subtitle calibration, and comparison calibration of the original image color extraction and generated video color.