Real-time rendering method of virtual scene based on machine vision and related equipment

Through machine vision technology, feature point extraction and posture analysis are performed on the real-time collected physical space video stream. Combined with the perspective compensation and anti-distortion rendering of the LED screen, an output data set adapted to the LED display terminal is generated. This solves the problem of insufficient timeliness in the dynamic adjustment of virtual scenes in virtual shooting, and realizes efficient virtual scene and real-time fusion and high-quality picture output.

CN120495145BActive Publication Date: 2025-09-16SHENZHEN TECNON EXCO-VISION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510963455.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-14
Publication Date
2025-09-16
Estimated Expiration
2045-07-14

AI Technical Summary

Technical Problem

During the virtual shooting process, the dynamic adjustment of the virtual scene is not timely enough, and there is a lack of real-time parallax adjustment and scene lighting synchronization, which makes the post-production repair work cumbersome.

Method used

Machine vision technology is used to extract spatiotemporal feature points and analyze target poses from the real-time collected physical space video stream, map it to a virtual coordinate system, and drive perspective transformation. Combined with LED screen perspective compensation and anti-distortion rendering, a high dynamic range (HDR) virtual background image is generated, and color and geometry correction are performed to finally generate an output dataset suitable for LED display terminals.

Benefits of technology

It achieves seamless synchronization between camera movement and physical target movement, enhances the immersion and synchronization of virtual scenes and real-time fusion, outputs clear, non-jaggy foreground subjects that blend naturally with HDR backgrounds, and improves the clarity and visual expressiveness of the picture.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495145B_ABST
    Figure CN120495145B_ABST
Patent Text Reader

Abstract

The present application relates to the field of virtual scene shooting technology, and provides a method for real-time rendering of virtual scenes based on machine vision and related equipment. Six-degree-of-freedom posture data is obtained by extracting spatiotemporal feature points and analyzing target postures of physical space video streams, and the six-degree-of-freedom posture data is mapped to a virtual coordinate system and driven by perspective transformation to obtain virtual shooting cone parameters. According to the LED display unit parameters, the virtual shooting cone parameters are compensated for LED screen perspective to obtain an anti-distortion rendering instruction set. According to the anti-distortion rendering instruction set, real-time matting and super-resolution reconstruction are performed to generate an HDR virtual background image. The HDR virtual background image is color-space calibrated and geometrically distorted to generate an output data set. The present application improves the accuracy and image quality of real-time rendering of virtual scenes through the coordinated optimization of multi-level correction and high-performance reconstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of virtual scene shooting, and in particular to a real-time rendering method of a virtual scene based on machine vision and related equipment. Background Art

[0002] Virtual filming plays a pivotal role in modern film, television, and live performances because it transcends the limitations of physical space and time, transporting audiences into any imaginable environment and significantly enriching visual expression and narrative possibilities. Through the overlay and real-time interaction of virtual backgrounds, creative teams can preview the final effect on-set. This not only shortens the lead time for location scouting and post-production color grading and compositing, but also reduces the costs of transporting sets and building real-world locations. This creates an immersive "what you see is what you get" filming experience, thereby enhancing production efficiency and artistic expression.

[0003] In traditional filming, to create a virtual background, a common approach is to first use a green screen (or blue screen) as a base in front of the real scene. When filming the actors' performances, only the foreground movement is recorded. The background footage, pre-filmed or rendered offline in 3D software, is then applied to the final image using a flat surface or projection. Another approach is to use a large LED screen to create a "virtual stage," looping a scene video pre-rendered using a game engine for the actors to reference. However, the background perspective and camera position during camera movement rely on pre-set data scripts, lacking real-time parallax adjustment and scene lighting synchronization. Most approaches rely on additional color grading and keying in post-production to achieve the desired effect. Summary of the Invention

[0004] In view of this, the present application provides a real-time rendering method of a virtual scene based on machine vision and related equipment to solve the problem of insufficient timeliness of dynamic adjustment of the virtual scene during virtual shooting.

[0005] In a first aspect, the present application provides a method for real-time rendering of a virtual scene based on machine vision, the method comprising:

[0006] Extract spatiotemporal feature points and analyze target poses of the real-time collected physical space video stream to obtain the six-degree-of-freedom pose data of the preset target object;

[0007] Performing virtual coordinate system mapping and perspective transformation driving processing on the six-degree-of-freedom posture data according to a preset virtual scene model to obtain virtual shooting cone parameters;

[0008] Performing LED screen perspective compensation processing on the virtual shooting cone parameters according to preset LED display unit parameters to obtain an anti-distortion rendering instruction set;

[0009] Performing real-time matting and super-resolution reconstruction processing according to the anti-distortion rendering instruction set to generate an HDR virtual background image;

[0010] Color space calibration and geometric distortion correction are performed on the HDR virtual background image to generate an output data set adapted to the LED display terminal.

[0011] In an optional embodiment, the six-degree-of-freedom pose data includes three-dimensional space coordinates and rotation angles, and the performing spatiotemporal feature point extraction and target pose analysis processing on the real-time collected physical space video stream to obtain the six-degree-of-freedom pose data of the preset target object includes:

[0012] According to the preset FAST corner detection algorithm, feature points are detected between consecutive frames for each frame of the real-time collected physical space video stream to obtain a set of corner point coordinates;

[0013] According to the preset BRIEF feature point description algorithm and the preset target object key points, key point classification and cross-frame matching are performed on the coordinates of each corner point in the corner point coordinate set to obtain the spatiotemporal motion trajectory of each key point;

[0014] The PnP solution is performed on the spatiotemporal motion trajectory to obtain the three-dimensional spatial coordinates and rotation angle of the target object.

[0015] In an optional embodiment, performing virtual coordinate system mapping and perspective transformation driving processing on the six-degree-of-freedom pose data according to a preset virtual scene model to obtain virtual shooting cone parameters includes:

[0016] Performing virtual coordinate system mapping processing on the three-dimensional space coordinates according to a preset virtual scene model to obtain virtual position data of the target object in the current LED screen center coordinate system;

[0017] According to a preset virtual camera position, performing LookAt perspective driving processing on the virtual position data to obtain virtual camera orientation parameters;

[0018] A view cone boundary calculation is performed on the virtual camera orientation parameters to obtain virtual shooting view cone parameters.

[0019] In an optional embodiment, performing LED screen perspective compensation processing on the virtual shooting frustum parameters according to preset LED display unit parameters to obtain an anti-distortion rendering instruction set includes:

[0020] Calculating pixel position offsets based on preset LED display unit parameters to obtain horizontal pixel offsets and vertical pixel offsets of the LED display unit;

[0021] Boundary deformation processing is performed on the virtual shooting frustum parameters according to the horizontal pixel offset and the vertical pixel offset to obtain an anti-distortion rendering instruction set.

[0022] In an optional embodiment, performing real-time matting and super-resolution reconstruction according to the anti-distortion rendering instruction set to generate an HDR virtual background image includes:

[0023] Constructing a virtual shooting frustum according to the virtual shooting frustum parameters, and clipping and segmenting the virtual scene model and the virtual shooting frustum according to the anti-distortion rendering instruction set to obtain micro-polygonal geometric patches;

[0024] Performing dynamic lighting tracing processing on the geometric facets to obtain lighting information of each facet, and performing time-series sampling and interpolation processing on the lighting information to obtain high dynamic lighting information of each facet;

[0025] The geometric surface is subjected to super-resolution reconstruction processing according to high dynamic lighting information to generate an HDR virtual background image.

[0026] In an optional embodiment, performing color space calibration and geometric distortion correction on the HDR virtual background image to generate an output data set adapted to the LED display terminal includes:

[0027] Performing 3D-LUT mapping processing on the HDR virtual background image according to a preset color mapping table to obtain color data under a target color gamut;

[0028] Performing brightness curve calibration processing on the color data to obtain a corrected image that conforms to the brightness characteristics of the LED display terminal;

[0029] Gaussian filtering and softening processing is performed on the seam area of ​​adjacent corrected images to generate an output data set adapted to the LED display terminal.

[0030] In an optional embodiment, before sending the output data set to the LED display terminal, the method further includes:

[0031] Sending a preset detection signal to the LED display terminal according to the output data set to obtain a feedback signal corresponding to the LED display unit;

[0032] Performing display status detection on the LED display unit according to the feedback signal to obtain status information corresponding to the LED display unit;

[0033] When the status information is preset abnormal information, an alarm is issued according to a preset alarm method.

[0034] A second aspect of the present application provides a device for real-time rendering of a virtual scene based on machine vision, the device comprising:

[0035] The posture analysis module is used to extract spatiotemporal feature points and perform target posture analysis on the real-time collected physical space video stream to obtain the six-degree-of-freedom posture data of the preset target object;

[0036] a shooting analysis module, configured to perform virtual coordinate system mapping and perspective transformation driving processing on the six-degree-of-freedom posture data according to a preset virtual scene model, so as to obtain virtual shooting view cone parameters;

[0037] a rendering compensation module, configured to perform LED screen perspective compensation processing on the virtual shooting frustum parameters according to preset LED display unit parameters, so as to obtain an anti-distortion rendering instruction set;

[0038] A background reconstruction module is used to perform real-time matting and super-resolution reconstruction processing according to the anti-distortion rendering instruction set to generate an HDR virtual background image;

[0039] The background calibration module is used to perform color space calibration and geometric distortion correction on the HDR virtual background image to generate an output data set adapted to the LED display terminal.

[0040] The third aspect of the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for real-time rendering of a virtual scene based on machine vision as described above are implemented.

[0041] In a fourth aspect, the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the steps of the method for real-time rendering of a virtual scene based on machine vision as described above are implemented.

[0042] In summary, this application has at least the following beneficial technical effects:

[0043] 1. Mapping six-degree-of-freedom pose data to a virtual coordinate system and driving perspective transformation enables seamless synchronization of camera movement and physical target movement, greatly enhancing the immersiveness and synchronization of virtual-reality fusion in virtual scenes and meeting the real-time response requirements for highly dynamic interactive scenes.

[0044] 2. Real-time keying based on the anti-distortion rendering instruction set, combined with a super-resolution algorithm to finely reconstruct keying edges and details, can output clear, non-aliased foreground subjects at high frame rates, which blend naturally with HDR background images, improving the clarity and visual expression of the picture. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0046] Figure 1 This is a flow chart of a method for real-time rendering of a virtual scene based on machine vision provided in an embodiment of the present application;

[0047] Figure 2 This is a functional module diagram of a virtual scene real-time rendering device based on machine vision provided in an embodiment of the present application;

[0048] Figure 3 It is a structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0049] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0050] like Figure 1 FIG2 is a flowchart of a method for real-time rendering of a virtual scene based on machine vision according to an embodiment of the present application. The method for real-time rendering of a virtual scene based on machine vision according to an embodiment of the present application comprises the following steps.

[0051] Step S1: extracting spatiotemporal feature points and analyzing target posture of the real-time collected physical space video stream to obtain six-degree-of-freedom posture data of the preset target object.

[0052] It should be understood that the six-degree-of-freedom pose data is used to represent the position and attitude measurement results of the target object in three-dimensional space, including but not limited to the three-dimensional space coordinates (x, y, z) along the three orthogonal directions of X, Y, and Z, and the rotation angles (θ x ,θ y ,θ z ). The six-degree-of-freedom pose data not only describes the precise position of the target object in real space, but also reflects its orientation and tilt. It serves as the basis for driving the virtual camera perspective when rendering the virtual scene in real time, so that the virtual background presented on the screen is aligned with the actor's movements at the millimeter level.

[0053] First, after continuously capturing an actor's motion video (i.e., a physical video stream) with a high-speed industrial camera, it is necessary to quickly identify reproducibly salient corner points within each frame of the physical video stream to ensure stable tracking even at high frame rates (e.g., 120 frames per second). Specifically, the FAST corner detection algorithm strikes a balance between efficiency and stability: each frame is divided into small, equally sized grids. A candidate pixel is then selected at the center of each grid. A small ring is constructed around this pixel, and the brightness changes of multiple sampling points within the ring are measured. If the brightness difference between a preset number of consecutive sampling points and the center pixel reaches a threshold, the center pixel is identified as a corner. This method allows the most effective tracking markers to be quickly selected from thousands to tens of thousands of corner points on an actor's clothing, face, or props, much like quickly identifying people wearing different-colored hats in a crowd. The corner point coordinates serve as input for subsequent spatiotemporal matching, and their accuracy directly impacts the fusion of the virtual background and the actor's movements.

[0054] Subsequently, the detected corner coordinates in each frame need to be assigned a unique "feature ID" to facilitate cross-frame identification of the same landmark. The BRIEF feature point description algorithm generates a binary string by comparing the brightness of pixel pairs surrounding each corner point. For example, by comparing the brightness of two adjacent points and recording the results sequentially. This binary string effectively summarizes the local texture characteristics of the corner point, much like assigning a unique number to each face in a photograph, enabling rapid matching in subsequent frames. Simultaneously, the system classifies the set of corner coordinates using predefined key point information of the target object (e.g., collar buttons, facial landmarks, or prop corners), selecting corners that correspond to the target object structure and eliminating interference from background or other irrelevant objects. By calculating the Hamming distance between the corresponding corner descriptors in the current and previous frames and combining the key point classification results to determine whether the same corner point appears consecutively, the spatiotemporal motion trajectory of each key point is generated. These trajectories record the continuous displacement of key locations of the target object along the time axis, providing multiple sets of 2D and 3D correspondences for subsequent pose analysis.

[0055] Finally, the obtained keypoint spatiotemporal trajectory is input into the PnP (Perspective-n-Point) solution module to calculate the pose using at least four sets of 3D model coordinates and their 2D projection point coordinates on the image plane. The PnP solution minimizes the error between the real image projection point and the model projection position to calculate a set of translation vectors (i.e., 3D space coordinates) (t x ,t y ,t z ) and the rotation matrix or equivalently the Euler angles (θ x ,θ y,θ z ) to determine the precise six-degree-of-freedom pose of the target object in the world coordinate system. For example, multiple cameras can observe the position of a mobile phone in a room and infer its spatial coordinates and orientation. The resulting six-degree-of-freedom pose data is used to update the virtual camera position and viewing angle, ensuring that the virtual background displayed on the LED screen seamlessly integrates with the actor's movements.

[0056] Step S2: performing virtual coordinate system mapping and perspective transformation driving processing on the six-degree-of-freedom posture data according to a preset virtual scene model to obtain virtual shooting cone parameters.

[0057] During the mapping process, based on the reference frame information defined by the pre-built virtual scene model, the coordinates of the LED screen's center point in the virtual scene coordinate system are first extracted. The translation offset between the object's original position in the world coordinate system and this center point is then calculated. This offset not only reflects the actor's position relative to the screen center in the virtual world but also ensures precise spatial alignment between the rendered background and the actual filmed image. Subsequently, combined with the rotation angle parameters, the rotation matrix corresponding to the Euler angles or quaternions is used to convert the facing direction, originally referenced to the camera's optical center, into a virtual viewing direction with the LED screen's center as the origin. For example, when an actor looks up, a positive rotation around the X-axis will cause the mapped position to appear slightly higher in the upward direction. When the actor turns their head to the left, a rotation around the Z-axis will cause the virtual camera to tilt leftward, as if the real camera is moving in sync with the actor's gaze. The meticulous alignment of translation and rotation during the mapping process is like precisely matching real-world GPS coordinates with a virtual map, achieving a seamless transition between the two.

[0058] After mapping the target object's three-dimensional spatial coordinates to the virtual coordinate system, the virtual camera's perspective needs to be driven. The LookAt visual driving method used in the embodiments of this application is centered around three key elements: "virtual camera position," "observation target point," and "upward direction," enabling intuitive simulation of the focus and tracking mechanisms of a real camera. It should be understood that the virtual camera position represents the actor's current position in the virtual scene coordinate system after mapping. The observation target point can be the actor's facial center or a chest marker to ensure that the rendering engine accurately projects the light, shadow, and depth of field effects in the virtual scene onto the actor. The upward direction is generally set to the positive Z axis of the global coordinate system to keep the image vertical and not tilted. After inputting these three elements into the LookAt algorithm, the system automatically generates a set of orthogonal coordinate axes, with the first axis pointing to the target point, the second axis perpendicular to the first axis and aligned with the upward direction, and the third axis, together with the first two, forming a left-handed coordinate system. This coordinate axis matrix is ​​the transformation matrix of the virtual camera. By applying this transformation to all geometric objects and lighting parameters in the scene, the translation and rotation of the virtual lens can be achieved, simulating the dynamic following effect of the camera moving with the actor (that is, the virtual camera orientation parameter).

[0059] After obtaining the virtual camera's orientation parameters, the corresponding viewing frustum boundaries must be calculated based on the preset lens properties. These virtual shooting frustum parameters include, but are not limited to, the left, right, top, and bottom field of view boundaries, as well as the distances to the near and far clipping planes—a total of six key parameters. These parameters define the volume of space the virtual camera can "see" and also determine which virtual object surfaces enter the rendering pipeline for rasterization or ray tracing. Calculating the viewing frustum boundaries first considers the lens' equivalent focal length and sensor size. For example, a 50mm focal length equivalent to a full-frame lens corresponds to a diagonal viewing angle of approximately 46°. For different LED screen resolutions and rendering output ratios, this viewing angle must be mapped to the width and height dimensions of the near clipping plane. Using a simple trigonometric mapping relationship, the half-width of the field of view at the near clipping plane distance is determined as the near clipping plane distance multiplied by the ratio of focal length to sensor width. The coordinates of the left and right boundaries are then derived. Using the same method to obtain the upper and lower bounds, combined with a far clipping plane distance (e.g., 1000m) that is further than the near clipping plane, we can fully describe the virtual camera's visible volume. This volume definition not only optimizes rendering efficiency (i.e., calculations are performed only on the visible area), but also ensures that the virtual background maintains perspective consistency and atmospheric realism when the actor moves to the edge of the screen.

[0060] Step S3: performing LED screen perspective compensation processing on the virtual shooting cone parameters according to preset LED display unit parameters to obtain an anti-distortion rendering instruction set.

[0061] The LED display unit parameters are the physical properties of large-scale tiled LED screens, such as curvature radius and pixel density, which determine the deformation patterns when presenting a flat image on a curved or curved screen. The resulting anti-distortion rendering instruction set is a set of parameters that describes how to adjust vertex coordinates and the projection matrix within the virtual rendering engine. This eliminates edge stretching or compression caused by screen curvature, ensuring that the rendered image is displayed on the LED screen in a realistic, distortion-free manner.

[0062] Specifically, the arc height from the center of the screen to the edge of each splicing unit is first measured using a laser rangefinder or structured light scanning device to determine the screen's curvature radius. For example, if the measured curvature radius is 8 meters, the entire screen can be considered to be approximately an arc with a radius of 8 meters. At the same time, the corresponding number of pixels per millimeter length (i.e., pixel density) needs to be collected. For example, when the pixel density is 2.5 pixels / mm, the physical distance between two adjacent pixel centers is approximately 0.4 mm. The curvature radius and pixel density together constitute the parameters of the LED display unit.

[0063] Next, each field of view boundary, defined by the parameters of the virtual filming frustum, is projected onto the curved surface geometry model of the LED screen to calculate the positional offset of each boundary point on the actual surface. The frustum appears as a set of straight line boundaries in virtual space. When these lines are mapped onto a curved surface, geometric distortion occurs at the edges due to varying projection distances from the curved surface. To correct this distortion, compensation is applied to the coordinates of each vertex to be rendered or the mapping position of each pixel. By comparing the difference between the ideal plane projection and the actual curved surface projection in the same screen coordinate system, the so-called horizontal and vertical pixel offsets can be determined.

[0064] Specifically, taking the straight-line distance x from the center of the screen to a horizontal boundary point as an example, if the boundary point is x on an ideal plane, the actual arc length mapped to the curved surface should be R × arcsin(x / R). The difference between this arc length and the original straight-line distance is the pixel-level offset required for compensation. While the above relationship can be precisely expressed using a formula, in practice, the corresponding offset percentage can be obtained through table lookup or interpolation. Converting the offset percentage to a pixel-level value yields the horizontal offset. Similarly, the vertical offset is derived using the relationship for the vertical offset of the screen. The offset reflects that the closer the actual curved surface is to the edge, the greater the curvature difference, and the corresponding pixel compensation is more significant. Near the center of the screen, the offset approaches zero because the surface is approximately flat.

[0065] The calculated horizontal and vertical offsets are then applied to the coordinates of the four boundaries of the virtual viewing frustum, deforming the frustum. For example, if the horizontal offset corresponding to the right boundary is δr, simply multiplying the original right boundary value by (1+δr) will expand the ideal straight-line field of view to the actual compensated range. Similarly, multiplying the left boundary by (1-δl) and the top and bottom boundaries by (1+δt) and (1-δb), respectively, yields a frustum shape that perfectly matches the curved LED surface during virtual rendering. This scaling of the boundaries isn't a simple geometric scaling; rather, each edge is individually processed based on the offset direction and magnitude to ensure accurate correction of distortion all around.

[0066] It is worth noting that after the boundary deformation is completed, the corresponding parameters in the projection matrix must be adjusted synchronously so that the vertex shader can correctly map the three-dimensional vertices to the compensated screen coordinate system when performing vertex coordinate transformation. The resulting anti-distortion rendering instruction set essentially contains the updated view frustum boundary values ​​and the matching projection matrix configuration. It is loaded into the graphics card driver or rendering engine when the rendering pipeline is initialized to ensure that each frame of the image can automatically apply distortion compensation at the hardware or software level.

[0067] For example, at a large-scale immersive concert, the LED screen presents an 8-meter-radius curved surface. When a virtual camera in the center of the stage focuses on the performer and outputs a real-time rendering, ignoring the physical curvature of the screen will cause the image on either side of the stage to appear elongated, resulting in distorted background scenery and unnatural perspective when the singer walks to the edge of the stage. After calculating pixel position offsets, the rendering engine fine-tunes the left and right boundaries of the field of view outward and inward, respectively—expanding the right side by approximately 5% and contracting the left side by approximately 5%—while also correcting the upper and lower boundaries based on the vertical curvature of the screen. The resulting anti-distortion rendering instructions, when output to the LED controller, ensure that the virtual scene presented on the screen maintains its realistic geometry regardless of the singer's movement.

[0068] The obtained anti-distortion rendering instruction set can not only compensate for the surface distortion of a single LED splicing unit itself, but also unify and coordinate the geometric errors between multiple splicing screens to achieve seamless alignment of the overall picture over a span of hundreds of meters.

[0069] Step S4: performing real-time matting and super-resolution reconstruction processing according to the anti-distortion rendering instruction set to generate an HDR virtual background image.

[0070] After the anti-distortion rendering instruction set is generated, a pyramid is constructed within the 3D scene based on the virtual view frustum parameters, which exactly corresponds to the actual camera field of view. This pyramid, acting as a perspective projector in the virtual world, determines which scene content is subject to subsequent geometry clipping and rendering. The anti-distortion rendering instruction set provides the vertex coordinate deformation necessary for curved LED screens, ensuring that the virtual geometry is mapped to the display without stretching or compression due to the curved surface. The pre-loaded 3D scene model and the constructed virtual view frustum are fed into the clipping and segmentation engine. Frustum clipping is performed, eliminating all faces outside the frustum to reduce redundant computation. Micropolygon tiling is then used to further subdivide the remaining geometry to a size of less than one pixel. This generates hundreds of millions of micropolygons, large enough to maintain smooth contours and intact detail even with rapid scene motion and dramatic changes in perspective. Micropolygon geometry is the smallest rendering unit in a real-time rendering engine. Its significance lies in the ability to complete geometric filtering and local subdivision in the vertex shading stage of the rendering pipeline, providing sufficient geometric accuracy for subsequent lighting calculations and pixel-level anti-aliasing.

[0071] After obtaining a collection of clipped and segmented micropolygon patches, the rendering pipeline switches to the lighting calculation stage. The dynamic ray tracing algorithm emits several rays per micropolygon patch and simulates the interaction between light and scene materials and light sources, obtaining information about the intensity and color distribution of the emitted light for each patch under the current viewing angle and lighting conditions. Ray tracing performs global illumination calculations here. By tracing rays emitted from the camera's viewpoint and their multiple bounces, it captures complex phenomena such as diffuse and specular reflections within the scene, as well as ambient occlusion. This high-dynamic-range (HDR) lighting data faithfully reproduces the contrast between light and dark, showcasing the sheen of metallic materials, the refraction of glass, and the soft shadows of cloth. To balance real-time performance and image quality, the ray tracing depth and number of samples per patch are dynamically adjusted based on the current scene complexity. This ensures that even with fast-moving actors, at least hundreds of valid rays interact with the patch, preventing "blooming" of highlights or "collapse" of shadows.

[0072] After calculating the lighting information for each micropolygon patch, temporal sampling and interpolation are required to generate high-dynamic lighting information for the patch. Temporal sampling combines the motion vector data from the previous frames to accumulate and compare the multiple lighting calculation results for the same micropolygon on the timeline. Catmull-Rom splines or other time-domain interpolation algorithms are then used to smoothly transition the lighting values, preserving the real details brought about by dynamic lighting changes while eliminating flickering caused by inter-frame noise or sampling differences. The resulting high-dynamic lighting information includes not only the brightness and chromaticity of the patch at the current moment, but also the peak and minimum brightness values ​​in the previous and next frames. Each micropolygon geometric patch carries a set of multi-frame fused lighting data, which can preserve the details of the bright areas and the texture of the dark areas in high-contrast scenes, preventing edge compression or overflow in the picture.

[0073] Once the high-dynamic lighting information is ready, the rendering pipeline enters the super-resolution reconstruction phase, which focuses on fusing the original micropolygon-based geometric details with the time-interpolated lighting information to produce the final HDR virtual background image. The super-resolution reconstruction algorithm samples each micropolygon at the screen pixel level and, based on its position and lighting changes in previous frames, performs sub-pixel interpolation between adjacent pixels to increase the output resolution to 8K or higher. Motion compensation across multiple frames prevents blurring or ghosting caused by rapid object movement, while leveraging the inherent high geometric precision of the micropolygons ensures that edges and contours remain sharp even after magnification. The HDR compositing process maps the luminance information of the patches to floating-point or fixed-point high-precision formats, exceeding the traditional 8-bit range, to preserve lighting details at the thousands or even tens of thousands of nits. An alpha channel is generated in the final image for seamless compositing with the real-time foreground keying. This results in an HDR virtual background image that simultaneously displays extremely bright light sources and deep shadow detail.

[0074] For example, in an immersive virtual performance scene, when an actor waves a fluorescent prop and quickly crosses the center of the stage, through the above-mentioned continuous micropolygon clipping, dynamic light tracking, temporal sampling interpolation and super-resolution reconstruction process, it is possible to output a frame of background image in real time that has realistic mirror reflections, delicate shadow gradients, ultra-high resolution and wide dynamic range. After HDR synthesis and foreground matting results are superimposed, no matter how the actor moves or how the weapon flashes, the picture can maintain silky light and shadow transitions and exquisite details, thus perfectly meeting the immersive display needs of the LED curved screen.

[0075] Step S5: performing color space calibration and geometric distortion correction processing on the HDR virtual background image to generate an output data set adapted to the LED display terminal.

[0076] After completing super-resolution reconstruction to generate a high dynamic range (HDR) virtual background image, the image's color and geometric characteristics must be professionally calibrated to ensure that the image displays the desired visual effect on the LED display terminal while eliminating color casts and stitching artifacts caused by the physical characteristics of the hardware. It should be understood that an HDR virtual background image is a set of floating-point data that stores light intensity values ​​beyond the traditional 8-bit color range. This allows it to preserve the vast difference in brightness, from faint shadows like moonlight to intense highlights like spotlights. However, if this high fidelity is not calibrated, it is very likely to cause color burn-in, overexposure, or loss of detail in dark areas on a large LED screen. To this end, it is necessary to first use a pre-made color mapping table and, using a three-dimensional lookup table (i.e., 3D-LUT) mapping algorithm, convert the wide color gamut colors contained in the HDR virtual background image into the target color gamut that the LED screen can display.

[0077] Specifically, a color mapping table, typically based on the color gamut mapping relationship between BT.2020 and DCI-P3, lists the target output values ​​for each red, green, and blue combination in the input color space in three dimensions. The original color values ​​of each pixel in the HDR virtual background image are then looked up in a 3D-LUT and trilinearly interpolated to produce color data that retains the rich color gradation of the original image while being compressed within the display device's color gamut. The 3D-LUT mapping process prevents flat transitions in extremely saturated colors or those that exceed the target display's capabilities, ensuring visually coherent and natural colors. This is similar to smoothly adjusting the rich colors of an oil painting to realistic pigment hues that can be rendered on a ceramic plate. For example, if the red channel value of a pixel in the HDR virtual background image reaches 1.5 (exceeding the standard 1.0), it will be mapped to the maximum red value of 1.0 supported by the LED screen through a 3D-LUT lookup and interpolation. This process also preserves relative brightness and hue relationships, preventing bright areas from being clipped to pure white due to exceeding the hardware's range.

[0078] After the color gamut mapping is completed, the overall brightness curve needs to be calibrated to adapt to the electro-optical characteristics of the LED panel. LED display units often exhibit nonlinear responses in different brightness ranges, especially in extremely dark and bright areas, which are prone to "sinking" or "explosion" phenomena, that is, dark details are lost or highlight areas are whitened. To this end, the PQ (Perceptual Quantizer)-EOTF (Electro-Optical Transfer Function) brightness curve calibration algorithm is used to perform a nonlinear transformation on the color-mapped data. The algorithm maps the input brightness value to the output signal according to the following formula based on the maximum peak brightness of the LED panel (for example, 1500 nits) and the gamma value (for example, 2.4):

[0079]

[0080] Among them, L in Indicates the input brightness value after 3D-LUT mapping, L max Represents the peak brightness supported by the LED screen, and γ is used to adjust the parameters of the human eye's contrast perception. Through this nonlinear calibration, each pixel can maintain a smooth transition in the medium brightness range, and can also show sufficient levels of detail in the extremely bright and dark ranges, making the overall picture bright, full and layered. The above formula adopts the power-law EOTF model of physics and display engineering, and draws on the geometric relationship of classical mathematics to ensure that the picture can be both fidelity and invisible on the high-dynamic, high-curvature LED screen. For example, in a desert scene illuminated by sunlight, the highlights reflected by the sunlight will not be suppressed into pure white, but will retain subtle textures while not causing visual fatigue to the audience due to excessive brightness.

[0081] After color and brightness are adjusted, a key issue remains: the physical seams between large-scale LED splicing screens. Because each LED display unit is difficult to achieve a completely seamless connection, seams several pixels wide often remain between panels. Unprocessed, these seams appear abrupt and result in noticeable color or brightness jumps at the spliced ​​screen. To address this, a Gaussian filtering algorithm is used to smooth the seams between adjacent corrected images. Based on a two-dimensional normal distribution function, Gaussian filtering takes a weighted average of pixels within a certain width (e.g., every 10 pixels) on either side of the seam line, replacing the hard edge at the seam with a smooth transition. Specifically, the geometric location of the seam is first determined at a logical level. For each pixel to be processed, the values ​​of several Gaussian-weighted pixels in its neighborhood are accumulated, ultimately replacing the original pixel with the smoothed result.

[0082] After three-fold fine processing of 3D-LUT mapping, PQ-EOTF brightness calibration, and Gaussian softening, a dedicated output data set for LED display terminals is generated, which has realistic colors and high dynamic range characteristics and achieves seamless connection at the physical splicing. When the output data set is transmitted to the controller or video splicer of the LED display terminal, it will be pushed to each display unit in a point-by-point frame refresh mode, and synchronized with the SDI or HDMI interface to ensure that the control signal and image content correspond accurately. For example, in a real-time virtual performance, when the director switches to a close-up shot, the background presented on the screen can not only maintain the original strong atmosphere of the performance set, but also eliminate the visual disconnection caused by the panel seams, truly realizing the seamless integration between physical shooting and virtual rendering.

[0083] In an optional embodiment, before sending the output data set to the LED display terminal, the system can also actively detect the operating status of the display unit to prevent black screen, distorted screen or color deviation on the live screen due to equipment failure or splicing unit failure, thereby ensuring the continuity and stability of the performance or shooting.

[0084] It should be understood that the detection signal used in the embodiments of this application refers to a specially designed test instruction or pattern data that can not only trigger the self-test program within the LED display unit, but also provide information about the unit's current operating status via the feedback interface. Therefore, the feedback signal transmitted back contains the LED unit's response data to the detection instruction, including but not limited to panel brightness, current and voltage monitoring values, temperature sensor readings, communication link status, and driver chip error codes.

[0085] When the detection begins, the system first calls the output data set as a trigger reference, and embeds the detection signal into the first frame to be sent or the control channel according to the frame rate and resolution requirements of the display content. Once the detection signal reaches the controller of the LED display terminal, the firmware inside the controller will automatically recognize the signal and distribute it to the driver boards of each splicing unit. After receiving the detection instruction, each driver board will package its own working parameters and operating status into a feedback signal in a specific format within a short time, and return it to the central processing unit through a redundant two-way data link or a dedicated monitoring bus. This return process is similar to sending heartbeat packets in a large server cluster and collecting the operating indicators of each node. Its significance lies in that through real-time monitoring, abnormal phenomena such as dead lights, weak brightness, color temperature imbalance or communication packet loss that may occur in the unit can be discovered in a timely manner.

[0086] After the feedback signal is transmitted back to the system of the embodiment of the present application, it needs to be parsed and processed for status detection in order to obtain "status information" (used as a basis for judging whether each LED display unit is currently operating normally). First, the feedback data is compared according to the preset detection standards, where each indicator such as brightness value, current voltage and temperature corresponds to a normal operating range, and is marked as abnormal if it exceeds the range. For example, if the operating current of a unit is lower than the minimum threshold, there may be a short circuit in the driver chip or a fault in the screen; if the temperature sensor reading is much higher than the normal value of the environment, poor heat dissipation may occur or there may be a risk of burning; if the packet loss rate of the communication link message is higher than the set upper limit, the display content is very likely to be mosaic or flickering. By cross-validating multiple types of feedback data, an information set containing three levels of status: "normal", "warning" or "fault" can be obtained, thereby providing a basis for subsequent alarms or maintenance.

[0087] After acquiring the status information of the LED display units, anomalies must be identified and alarms triggered to promptly notify the on-site director, script supervisor, or backend operations and maintenance personnel, preventing interruptions to program recording and live broadcasts due to equipment issues. Alarms can be sent in a variety of ways, such as sending an alert message to the on-site display console and scrolling warning text on the side of the large screen. Alternatively, the backend management system can send SMS or in-app notifications to the director and script supervisor's mobile devices, illuminating a warning light in the control room and accompanied by an audible prompt to alert operations and maintenance personnel to the scene. Furthermore, to enhance response efficiency, the system can automatically generate a maintenance report containing the faulty unit number, anomaly type, timestamp, and a recommended action (e.g., "Please replace splicing unit number 5") and send it to the equipment maintenance team for prompt replacement or repair. Alarm handling and maintenance report generation are similar to the real-time monitoring and intelligent operation and maintenance of each workstation on the production line in a modern smart factory. Not only can the fault be pinpointed to a single screen, but it also streamlines and traces maintenance workflows, fundamentally improving the stability and audience experience of large-scale performances or film shoots.

[0088] For example, during a live broadcast of an immersive concert, when the performers have just finished a high-intensity interactive performance, the system will first send a predefined grayscale test signal and collect the brightness and temperature readings returned by each splicing unit before pushing the next virtual scene to the LED screen. If the returned data shows that the brightness of unit 7 is only 60% of the normal value, while the adjacent units remain within the range of 95% to 105%, it means that unit 7 may have pixel aging or loose wiring problems. At this time, the system will immediately scroll "Warning: Unit 07 brightness low" on the lower right corner of the background control screen and automatically send a notification "Unit 7 brightness abnormality, please arrange replacement" to the director and the script supervisor's handheld terminal. At the same time, the location of unit 7 will be highlighted on the operation and maintenance screen in the control room, so that technicians can quickly arrive to inspect or replace it, avoiding visible black blocks or color deviations when the next scene switches to this area.

[0089] This application is applied to the field of virtual scene shooting technology. It extracts spatiotemporal feature points and analyzes the target posture of the physical space video stream to obtain six-degree-of-freedom posture data. The six-degree-of-freedom posture data is mapped to a virtual coordinate system and driven by perspective transformation to obtain virtual shooting cone parameters. The virtual shooting cone parameters are compensated for LED screen perspective according to the LED display unit parameters to obtain an anti-distortion rendering instruction set. According to the anti-distortion rendering instruction set, real-time matting and super-resolution reconstruction are performed to generate an HDR virtual background image. The HDR virtual background image is color-space calibrated and geometrically corrected to generate an output data set. This application not only significantly improves the accuracy and image quality of real-time rendering of virtual scenes through a series of collaborative processes such as spatiotemporal feature point extraction, six-degree-of-freedom posture analysis, virtual coordinate system mapping, LED screen perspective compensation, real-time matting and super-resolution reconstruction, as well as color and geometric correction, but also takes into account the real-time, stability and scalability of the system. It has important engineering application value and market promotion prospects.

[0090] like Figure 2 , which is a functional module diagram of a virtual scene real-time rendering device based on machine vision provided in an embodiment of the present application.

[0091] In some embodiments, the virtual scene real-time rendering device 2 based on machine vision may include multiple functional modules composed of computer program segments. The computer program of each program segment in the virtual scene real-time rendering device 2 based on machine vision may be stored in the memory of the server and executed by at least one processor to perform (see Figure 1 (Describes) the functionality of a real-time rendering method for virtual scenes based on machine vision.

[0092] In this embodiment, the machine vision-based virtual scene real-time rendering device 2 can be divided into multiple functional modules according to the functions it performs. These functional modules may include: a pose analysis module 21, a shooting analysis module 22, a rendering compensation module 23, a background reconstruction module 24, a background calibration module 25, and a detection and alarm module 26. As used herein, a module refers to a series of computer program segments that can be executed by at least one processor and can perform fixed functions, and are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.

[0093] The posture analysis module 21 is used to extract spatiotemporal feature points and perform target posture analysis on the real-time collected physical space video stream to obtain six-degree-of-freedom posture data of the preset target object.

[0094] In an optional embodiment, the posture analysis module 21 is specifically used to:

[0095] According to the preset FAST corner detection algorithm, feature points are detected between consecutive frames for each frame of the real-time collected physical space video stream to obtain a set of corner point coordinates;

[0096] According to the preset BRIEF feature point description algorithm and the preset target object key points, key point classification and cross-frame matching are performed on the coordinates of each corner point in the corner point coordinate set to obtain the spatiotemporal motion trajectory of each key point;

[0097] The PnP solution is performed on the spatiotemporal motion trajectory to obtain the three-dimensional spatial coordinates and rotation angle of the target object.

[0098] The shooting analysis module 22 is used to perform virtual coordinate system mapping and perspective transformation driving processing on the six-degree-of-freedom posture data according to a preset virtual scene model to obtain virtual shooting cone parameters.

[0099] In an optional embodiment, the shooting analysis module 22 is specifically configured to:

[0100] Performing virtual coordinate system mapping processing on the three-dimensional space coordinates according to a preset virtual scene model to obtain virtual position data of the target object in the current LED screen center coordinate system;

[0101] According to a preset virtual camera position, performing LookAt perspective driving processing on the virtual position data to obtain virtual camera orientation parameters;

[0102] A view cone boundary calculation is performed on the virtual camera orientation parameters to obtain virtual shooting view cone parameters.

[0103] The rendering compensation module 23 is used to perform LED screen perspective compensation processing on the virtual shooting cone parameters according to preset LED display unit parameters to obtain an anti-distortion rendering instruction set.

[0104] In an optional embodiment, the rendering compensation module 23 is specifically configured to:

[0105] Calculating pixel position offsets based on preset LED display unit parameters to obtain horizontal pixel offsets and vertical pixel offsets of the LED display unit;

[0106] Boundary deformation processing is performed on the virtual shooting frustum parameters according to the horizontal pixel offset and the vertical pixel offset to obtain an anti-distortion rendering instruction set.

[0107] The background reconstruction module 24 is used to perform real-time matting and super-resolution reconstruction according to the anti-distortion rendering instruction set to generate an HDR virtual background image.

[0108] In an optional embodiment, the background reconstruction module 24 is specifically configured to:

[0109] Constructing a virtual shooting frustum according to the virtual shooting frustum parameters, and clipping and segmenting the virtual scene model and the virtual shooting frustum according to the anti-distortion rendering instruction set to obtain micro-polygonal geometric patches;

[0110] Performing dynamic lighting tracing processing on the geometric facets to obtain lighting information of each facet, and performing time-series sampling and interpolation processing on the lighting information to obtain high dynamic lighting information of each facet;

[0111] The geometric surface is subjected to super-resolution reconstruction processing according to high dynamic lighting information to generate an HDR virtual background image.

[0112] The background calibration module 25 is used to perform color space calibration and geometric distortion correction processing on the HDR virtual background image to generate an output data set adapted to the LED display terminal.

[0113] In an optional embodiment, the background calibration module 25 is specifically configured to:

[0114] Performing 3D-LUT mapping processing on the HDR virtual background image according to a preset color mapping table to obtain color data under a target color gamut;

[0115] Performing brightness curve calibration processing on the color data to obtain a corrected image that conforms to the brightness characteristics of the LED display terminal;

[0116] Gaussian filtering and softening processing is performed on the seam area of ​​adjacent corrected images to generate an output data set adapted to the LED display terminal.

[0117] In an optional embodiment, the machine vision-based virtual scene real-time rendering device 2 further includes a detection alarm module 26, and the detection alarm module 26 is specifically used to:

[0118] Sending a preset detection signal to the LED display terminal according to the output data set to obtain a feedback signal corresponding to the LED display unit;

[0119] Performing display status detection on the LED display unit according to the feedback signal to obtain status information corresponding to the LED display unit;

[0120] When the status information is preset abnormal information, an alarm is issued according to a preset alarm method.

[0121] It should be understood that the various variations and specific embodiments of the methods provided in the above embodiments are also applicable to the real-time rendering device for virtual scenes based on machine vision in this embodiment. Through the above detailed description of the real-time rendering method for virtual scenes based on machine vision, those skilled in the art can clearly understand the implementation method of the real-time rendering device for virtual scenes based on machine vision in this embodiment. For the sake of brevity of the specification, it will not be described in detail here.

[0122] like Figure 3 , which is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application.

[0123] In a preferred embodiment of the present invention, the electronic device 3 may include, but is not limited to: a memory 31 , at least one processor 32 and at least one communication bus 33 .

[0124] Those skilled in the art should understand that Figure 3 The structure of the electronic device 3 shown does not constitute a limitation of the embodiment of the present invention. The electronic device 3 may also include more or less other hardware or software than shown in the figure, or a different component arrangement.

[0125] In some embodiments, the electronic device 3 is a device that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes but is not limited to microprocessors, application-specific integrated circuits, programmable gate arrays, digital processors and embedded devices.

[0126] It should be noted that the electronic device 3 is only an example. Other existing or future electronic products that are suitable for this application should also be included in the scope of protection of this application and included here by reference.

[0127] In some embodiments, the memory 31 stores a computer program that, when executed by the at least one processor 32, implements all or part of the steps in the method for real-time rendering of a virtual scene based on machine vision. The memory 31 includes a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), a one-time programmable read-only memory (OTPROM), an electronically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM), or other optical disk storage, magnetic disk storage, magnetic tape storage, or any other computer-readable medium capable of carrying or storing data. Furthermore, the computer-readable storage medium may primarily include a program storage area and a data storage area, wherein the program storage area may store an operating system, at least one application required for a function, and the like.

[0128] In some embodiments, the at least one processor 32 is the control core (Control Unit) of the electronic device 3. It connects the various components of the entire electronic device 3 using various interfaces and lines. It executes or runs programs or modules stored in the memory 31 and calls data stored in the memory 31 to perform various functions of the electronic device 3 and process data. For example, when the at least one processor 32 executes the computer program stored in the memory 31, it implements all or part of the steps of the real-time rendering method of a virtual scene based on machine vision described in the embodiments of the present application; or it implements all or part of the functions of the real-time rendering device of a virtual scene based on machine vision. The at least one processor 32 can be composed of an integrated circuit, such as a single packaged integrated circuit, or a plurality of packaged integrated circuits with the same or different functions, including a combination of one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and various control chips.

[0129] In some embodiments, the at least one communication bus 33 is configured to enable communication between the memory 31 and the at least one processor 32. Although not shown, the electronic device 3 may also include a power supply (e.g., a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 32 via a power management device, thereby enabling the power management device to manage charging, discharging, and power consumption. The power supply may also include one or more DC or AC power supplies, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, and other components. The electronic device 3 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be detailed here.

[0130] The above-mentioned integrated unit implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module stored in a storage medium includes a number of instructions for causing an electronic device (which can be a personal computer, electronic device, or network device, etc.) or a processor to execute portions of the methods described in various embodiments of the present application.

[0131] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is only a logical function division, and other division methods may be used in actual implementation.

[0132] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, and may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of this embodiment based on actual needs.

[0133] The above are all preferred embodiments of the present application, and are not intended to limit the scope of protection of the present application. Therefore, any equivalent changes made based on the structure, shape, and principle of the present application should be included in the scope of protection of the present application.

Claims

1. A method for real-time rendering of virtual scenes based on machine vision, characterized in that: The method comprises: Extract spatiotemporal feature points and analyze target poses of the real-time collected physical space video stream to obtain the six-degree-of-freedom pose data of the preset target object; Performing virtual coordinate system mapping and perspective transformation driving processing on the six-degree-of-freedom posture data according to a preset virtual scene model to obtain virtual shooting cone parameters; Performing LED screen perspective compensation processing on the virtual shooting cone parameters according to preset LED display unit parameters to obtain an anti-distortion rendering instruction set; Performing real-time matting and super-resolution reconstruction processing according to the anti-distortion rendering instruction set to generate an HDR virtual background image; Performing color space calibration and geometric distortion correction on the HDR virtual background image to generate an output data set adapted to the LED display terminal; The step of performing LED screen perspective compensation processing on the virtual shooting cone parameters according to preset LED display unit parameters to obtain an anti-distortion rendering instruction set includes: Calculating pixel position offsets based on preset LED display unit parameters to obtain horizontal pixel offsets and vertical pixel offsets of the LED display unit; performing boundary deformation processing on the virtual shooting frustum parameters according to the horizontal pixel offset and the vertical pixel offset to obtain an anti-distortion rendering instruction set; The performing of real-time matting and super-resolution reconstruction according to the anti-distortion rendering instruction set to generate an HDR virtual background image includes: Constructing a virtual shooting frustum according to the virtual shooting frustum parameters, and clipping and segmenting the virtual scene model and the virtual shooting frustum according to the anti-distortion rendering instruction set to obtain micro-polygonal geometric patches; Performing dynamic lighting tracing processing on the geometric facets to obtain lighting information of each facet, and performing time-series sampling and interpolation processing on the lighting information to obtain high dynamic lighting information of each facet; The geometric surface is subjected to super-resolution reconstruction processing according to high dynamic lighting information to generate an HDR virtual background image.

2. The method for real-time rendering of a virtual scene based on machine vision according to claim 1, characterized in that: The six-degree-of-freedom pose data includes three-dimensional space coordinates and rotation angles, and the extraction of spatiotemporal feature points and target pose analysis processing of the real-time collected physical space video stream to obtain the six-degree-of-freedom pose data of the preset target object includes: According to the preset FAST corner detection algorithm, feature points are detected between consecutive frames for each frame of the real-time collected physical space video stream to obtain a set of corner point coordinates; According to the preset BRIEF feature point description algorithm and the preset target object key points, key point classification and cross-frame matching are performed on the coordinates of each corner point in the corner point coordinate set to obtain the spatiotemporal motion trajectory of each key point; The PnP solution is performed on the spatiotemporal motion trajectory to obtain the three-dimensional spatial coordinates and rotation angle of the target object.

3. The method for real-time rendering of virtual scenes based on machine vision according to claim 2, characterized in that: The performing virtual coordinate system mapping and perspective transformation driving processing on the six-degree-of-freedom posture data according to a preset virtual scene model to obtain virtual shooting cone parameters includes: Performing virtual coordinate system mapping processing on the three-dimensional space coordinates according to a preset virtual scene model to obtain virtual position data of the target object in the current LED screen center coordinate system; According to a preset virtual camera position, performing LookAt perspective driving processing on the virtual position data to obtain virtual camera orientation parameters; A view cone boundary calculation is performed on the virtual camera orientation parameters to obtain virtual shooting view cone parameters.

4. The method for real-time rendering of virtual scenes based on machine vision according to claim 1, characterized in that: The performing color space calibration and geometric distortion correction processing on the HDR virtual background image to generate an output data set adapted to the LED display terminal includes: Performing 3D-LUT mapping processing on the HDR virtual background image according to a preset color mapping table to obtain color data under a target color gamut; Performing brightness curve calibration processing on the color data to obtain a corrected image that conforms to the brightness characteristics of the LED display terminal; Gaussian filtering and softening processing is performed on the seam area of ​​adjacent corrected images to generate an output data set adapted to the LED display terminal.

5. The method for real-time rendering of virtual scenes based on machine vision according to claim 1, characterized in that: Before sending the output data set to the LED display terminal, the method further includes: Sending a preset detection signal to the LED display terminal according to the output data set to obtain a feedback signal corresponding to the LED display unit; Performing display status detection on the LED display unit according to the feedback signal to obtain status information corresponding to the LED display unit; When the status information is preset abnormal information, an alarm is issued according to a preset alarm method.

6. A virtual scene real-time rendering device based on machine vision, applied to the virtual scene real-time rendering method based on machine vision according to claim 1, characterized in that: The device comprises: The posture analysis module is used to extract spatiotemporal feature points and perform target posture analysis on the real-time collected physical space video stream to obtain the six-degree-of-freedom posture data of the preset target object; a shooting analysis module, configured to perform virtual coordinate system mapping and perspective transformation driving processing on the six-degree-of-freedom posture data according to a preset virtual scene model, so as to obtain virtual shooting view cone parameters; a rendering compensation module, configured to perform LED screen perspective compensation processing on the virtual shooting frustum parameters according to preset LED display unit parameters, so as to obtain an anti-distortion rendering instruction set; A background reconstruction module is used to perform real-time matting and super-resolution reconstruction according to the anti-distortion rendering instruction set to generate an HDR virtual background image; The background calibration module is used to perform color space calibration and geometric distortion correction on the HDR virtual background image to generate an output data set adapted to the LED display terminal.

7. An electronic device, characterized in that: The electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the method for real-time rendering of a virtual scene based on machine vision are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for real-time rendering of a virtual scene based on machine vision according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Display method and device and electronic equipment

    CN117455974A

  • Augmented reality guidance for imaging systems

    US20220287676A1