A naked-eye 3D image automatic generation display system

By using intelligent brushes and preset templates for interactive design, the system dynamically divides the focus and non-focus areas. Combined with real-time rendering and asynchronous processing, it solves the problems of monotonous interaction and lag in naked-eye 3D display technology, achieving efficient naked-eye 3D image generation and an optimized user experience.

CN121284215BActive Publication Date: 2026-04-17SHENZHEN HUARUIAN TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511852221.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-04-17
Estimated Expiration
2045-12-10

AI Technical Summary

Technical Problem

In existing glasses-free 3D display technologies, the depth map generation methods have limited interaction, making it difficult for users to intuitively and accurately control the image depth levels. This results in a high operational threshold, frequent stuttering during real-time preview and rendering, and an inability to optimize the allocation of computing resources, thus restricting the promotion of glasses-free 3D technology in real-time applications.

Method used

The interactive design combines intelligent brushes and preset templates. The guidance module receives user interaction commands, generates a global depth map, and dynamically divides the focus and non-focus areas in the rendering module. The focus area is rendered in real time, while the non-focus area is downgraded. The asynchronous background processing optimizes the allocation of computing resources.

Benefits of technology

It significantly improves the efficiency of naked-eye 3D image generation and user experience, lowers the operational threshold, improves the efficiency of complex edge processing, solves the problem of interaction lag, ensures smooth real-time preview and the precision of the final output, and promotes the popularization of naked-eye 3D technology in real-time interactive applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121284215B_ABST
    Figure CN121284215B_ABST
Patent Text Reader

Abstract

This invention discloses a naked-eye 3D image automatic generation and display system, relating to the field of computer vision technology. It includes a guidance module, a generation module, and a rendering module. The guidance module receives user interaction commands. The generation module generates a global depth map of the original 2D image based on the interaction commands. The rendering module dynamically divides the original 2D image into focus and non-focus areas based on the user's brush position, performs real-time rendering of the focus areas, and performs downgrading and asynchronous background processing on the non-focus areas. By combining the global depth map, the final naked-eye 3D image is generated. This invention, through innovative interactive design and intelligent rendering mechanism, combines multiple modes of intelligent brushes with preset templates. Furthermore, by dynamically dividing the focus and non-focus areas based on the user's brush position, it achieves intelligent allocation of computing resources, effectively promoting the popularization of naked-eye 3D technology in real-time interactive applications.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a naked-eye 3D image automatic generation and display system. Background Technology

[0002] In recent years, naked-eye 3D display technology, as a cutting-edge technology that can present stereoscopic visual effects without wearing auxiliary devices, has shown broad application prospects in the fields of consumer electronics and advertising media. The production process of traditional naked-eye 3D content relies heavily on professional software and complex depth map drawing. It not only requires operators to have professional graphics knowledge, but also the process is cumbersome and time-consuming, which greatly limits the creation of ordinary users and the popularization of high-quality naked-eye 3D images.

[0003] In existing technologies, depth map generation methods have a single interactive mode, making it difficult for users to intuitively and accurately control the depth levels of different elements in the image. When dealing with complex edges and main objects, the operation threshold is high and the efficiency is low. In the real-time preview and rendering stages, existing technologies often adopt a uniform rendering strategy for the entire image in order to ensure smoothness, which easily causes lag when users perform interactive operations, seriously affecting the user experience. Existing technologies lack intelligent region division and resource scheduling mechanisms, which prevents the optimal allocation of computing resources. They cannot ensure the real-time interactive response speed while taking into account the fineness of the final output image, thus limiting the promotion of naked-eye 3D technology in real-time applications. Summary of the Invention

[0004] The technical problem solved by this invention is that the interaction mode of depth map generation methods is singular, making it difficult for users to intuitively and accurately control the depth level of different elements in the image. When dealing with complex edges and main objects, the operation threshold is high and the efficiency is low. In the real-time preview and rendering stage, in order to ensure smoothness, existing technologies often adopt a uniform rendering strategy for the entire image, which easily causes lag when users perform interactive operations, seriously affecting the user experience. Existing technologies lack intelligent region division and resource scheduling mechanisms, which makes it impossible to optimize the allocation of computing resources and ensure the real-time interactive response speed while taking into account the fineness of the final output image, thus limiting the promotion of naked-eye 3D technology in real-time applications.

[0005] To solve the above-mentioned technical problems, the present invention provides the following technical solution: a naked-eye 3D image automatic generation and display system, comprising a guidance module, a generation module and a rendering module;

[0006] The guidance module is used to receive user interaction commands, which include smearing operations on the original 2D image of the naked-eye 3D image using a smart brush. The smearing operations include pushing forward, pulling back, and selecting a preset template.

[0007] The generation module is used to generate a global depth map of the original 2D image according to the interaction instructions;

[0008] The rendering module is used to dynamically divide the original 2D image into focus and non-focus areas based on the user's brush position, render the focus area in real time, and perform downgrade processing and asynchronous background processing on the non-focus area, and generate the final naked-eye 3D image by combining the global depth map.

[0009] As a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the intelligent brush includes different interaction modes;

[0010] The different interaction modes include depth adjustment mode, edge feathering mode, and intelligent recognition mode;

[0011] The depth adjustment mode is used to set the depth value of the corresponding pixel region in the original 2D image according to the smearing operation;

[0012] The edge feathering mode is used to feather the edge region of the original 2D image after setting a depth value. The feathering operation is used to perform weighted interpolation calculation on the depth value of the edge region according to the trajectory intensity of the smearing operation to generate a smooth edge region.

[0013] The intelligent recognition mode is used to identify the main object of the original 2D image through an image segmentation algorithm, generate a pixel mask of the main object, and guide the user to assign a depth value to the main object.

[0014] The logic for performing deep assignment on the main object includes:

[0015] Render a depth control control that is linked to the main object on the user interface; the depth control control control is a linear slider.

[0016] The system monitors the user's sliding operation on the linear slider in real time, maps the physical position of the sliding linear slider to a preset depth range, and obtains the target depth value.

[0017] Set the depth value of all pixels within the pixel mask area to the target depth value.

[0018] As a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the preset template includes portrait template, landscape template, architectural template and still life template;

[0019] Each preset template includes an initial depth map layout that matches the content type of the preset template and a depth assignment range for the main object.

[0020] As a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the logic for generating the global depth map of the original 2D image includes:

[0021] Based on the interaction instructions, a global depth map is generated using a depth propagation algorithm;

[0022] The logic for generating a global depth map using the depth propagation algorithm includes:

[0023] Each pixel in the original 2D image is used as a node in the depth map to construct an initial depth map;

[0024] Based on the interaction instructions, the nodes corresponding to the pixels painted by the user through the smearing operation or the key areas defined by the selected preset template are used as seed nodes, and the depth value of the seed nodes is set to a preset first depth value, and the depth values ​​of the remaining nodes are set to a preset second depth value.

[0025] The key areas include the key areas of the portrait template, the key areas of the landscape template, the key areas of the architectural template, and the key areas of the still life template;

[0026] The key areas of the portrait template include the face area, hair area, and body area;

[0027] The key areas of the landscape template include the background area, the middle ground area, and the foreground area;

[0028] The key areas of the building formwork include the wall area, door and window area, and eaves area;

[0029] The key areas of the still life template include the still life body area, the still life shadow area, and the still life support plane area.

[0030] Among them, the weight of the preset first depth value is the preset first weight, and the weight of the preset second depth value is the preset second weight;

[0031] Calculate the edge weight between two nodes in the initial depth graph, and use a random walk algorithm to traverse all nodes in the initial depth graph along the path with the highest edge weight, taking the preset first depth value of the seed node, to obtain the global depth graph.

[0032] As a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the generation module further includes a detection unit and a smoothing unit;

[0033] The detection unit is used to perform convolution calculation on the global depth map by using an edge detection operator to obtain the depth gradient magnitude of each pixel in the global depth map;

[0034] Pixels with a depth gradient magnitude greater than a preset mutation threshold are identified as mutation points, and spatially continuous sets of mutation points are identified as mutation regions.

[0035] The smoothing unit is used to smooth the abrupt change region using a bilateral filtering algorithm.

[0036] As a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the logic of dynamically dividing the focal region and non-focal region according to the user's brush position includes:

[0037] The current position of the user's brush is taken as the center point of the focus area. At the same time, the radius of the focus area is dynamically calculated based on the real-time movement speed of the brush and the preset focus sensitivity.

[0038] Based on the center point and the radius of the range, a circular region is divided on the original 2D image, and the circular region is taken as the focal region. The region on the original 2D image other than the circular region is taken as the non-focal region.

[0039] The mathematical expression for dynamically calculating the radius of the focal region is: ;

[0040] in, This represents the radius of the focal region. This indicates the preset focus sensitivity of the brush. This indicates the real-time movement speed of the brush. This indicates the preset base radius coefficient. It represents a positive infinitesimal.

[0041] As a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the logic for downgrading non-focal areas includes:

[0042] The image data of the non-focal area is downsampled to obtain low-resolution data. The downsampling process is used to reduce the resolution of the non-focal area to a preset ratio of the original resolution.

[0043] The low-resolution data is processed by 3D rendering to obtain a first image, and the image of the non-focal area is used as a second image;

[0044] The first and second images are upsampled and then fused to obtain a preview output image after downsampling the non-focal areas.

[0045] The logic for upsampling and fusion includes:

[0046] A high-quality interpolation algorithm is used to enlarge the first image to the same original resolution as the second image to obtain an intermediate image. At the same time, a low-frequency second image is obtained by applying Gaussian blur to the second image. The second image and the low-frequency second image are then subtracted at the pixel level to obtain a detail image.

[0047] The intermediate image and the detail image are weighted and fused to obtain a preview output image.

[0048] As a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the logic for asynchronous background processing of non-focal areas includes:

[0049] After downgrading the non-focal area, the background process is started, and the non-focal area in the original 2D image is rendered with high precision through the background process to obtain a high precision non-focal area image.

[0050] The high-precision non-focus area image is used to replace the corresponding area of ​​the preview output image in the original 2D image to obtain an intermediate rendering image.

[0051] In a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the rendering module further includes a cache unit;

[0052] The cache unit is used to store the high-precision non-focus area image generated by the asynchronous background processing. When the user's brush enters the stored high-precision non-focus area image, the high-precision non-focus area image is retrieved from the cache unit for display.

[0053] As a preferred embodiment of the naked-eye 3D image automatic generation and display system of the present invention, the logic of generating the final naked-eye 3D image by combining the global depth map includes:

[0054] Based on the global depth map, a raster rendering algorithm is used to perform multi-viewpoint rendering on the intermediate rendering map to generate N initial viewpoint images. According to the preset viewing distance, the parallax intensity between the N initial viewpoint images is dynamically adjusted, and the N initial viewpoint images with adjusted parallax intensity are synthesized to generate the final naked-eye 3D image.

[0055] The beneficial effects of this invention are as follows: Through innovative interactive design and intelligent rendering mechanism, this invention significantly improves the generation efficiency and user experience of naked-eye 3D images. It combines intelligent brushes with multiple modes and preset templates, allowing users to intuitively and accurately control image depth, greatly reducing the operational threshold and improving the efficiency of complex edge processing. Simultaneously, by dynamically dividing the focus area and non-focus area according to the user's brush position, it performs real-time high-definition rendering of the focus area while downgrading and asynchronously processing the non-focus area in the background, achieving intelligent allocation of computing resources. This not only effectively solves the problem of interactive lag and ensures smooth real-time preview but also takes into account the precision of the final output, powerfully promoting the popularization of naked-eye 3D technology in real-time interactive applications. Attached Figure Description

[0056] Figure 1 This is a schematic diagram of the basic process of an automatic naked-eye 3D image generation and display system provided in one embodiment of the present invention. Detailed Implementation

[0057] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments.

[0058] Example, refer to Figure 1 As an embodiment of the present invention, a naked-eye 3D image automatic generation and display system is provided, including a guidance module, a generation module and a rendering module;

[0059] The guidance module is used to receive user interaction commands, which include the ability to smear the original 2D image of the naked-eye 3D image using a smart brush. The smearing operations include pushing forward, pulling back, and selecting a preset template.

[0060] The generation module is used to generate a global depth map of the original 2D image based on interactive instructions;

[0061] The rendering module dynamically divides the original 2D image into focus and non-focus areas based on the user's brush position. It renders the focus areas in real time and performs downgraded processing and asynchronous background processing on the non-focus areas. The final naked-eye 3D image is generated by combining the global depth map.

[0062] Through the collaborative work of the guidance module, generation module, and rendering module, end-to-end intelligent conversion from ordinary 2D images to high-quality naked-eye 3D images is achieved. The guidance module, as the entry point for user intent, is responsible for capturing intuitive interactive commands. Based on the interactive commands, the generation module uses advanced algorithms to construct a global depth map that reflects the spatial hierarchy of the image. On this basis, the rendering module uses an innovative partitioned rendering strategy to fuse depth information with the original image, ultimately presenting a stereoscopic effect. This design seamlessly connects user interaction, depth calculation, and real-time rendering, forming an organic whole. This ensures the continuity and efficiency of the entire creation process and avoids the problems of inefficiency and poor results caused by the disconnect between various links.

[0063] The smart brush includes different interaction modes;

[0064] Different interaction modes include depth adjustment mode, edge feathering mode, and intelligent recognition mode;

[0065] The depth adjustment mode is used to set the depth value of the corresponding pixel area in the original 2D image based on the smearing operation;

[0066] The edge feathering mode is used to feather the edge regions of the original 2D image after setting a depth value. The feathering operation is used to perform weighted interpolation calculation on the depth value of the edge region based on the intensity of the smearing operation trajectory to generate a smooth edge region.

[0067] The intelligent recognition mode is used to identify the main object of the original 2D image through image segmentation algorithm, generate the pixel mask of the main object, and guide the user to assign a depth value to the main object;

[0068] The logic for deep assignment of the main object includes:

[0069] Render a depth control control on the user interface that is linked to the main object; the depth control control is a linear slider.

[0070] The system monitors the user's sliding operation on the linear slider in real time, maps the physical position of the linear slider to a preset depth range, and obtains the target depth value.

[0071] Set the depth value of all pixels within the pixel mask area to the target depth value.

[0072] Different interaction modes can be switched intuitively by clicking the icon buttons in the interface toolbar.

[0073] Edge feathering mode smooths depth transitions by combining distance and Gaussian weighting. It extends outward from the user's smearing trajectory as the core, and the depth value of the feathered pixel is weighted by a Gaussian function based on its distance from the core area. This weight is then smoothly blended with the original depth value, thus naturally eliminating harsh depth boundaries.

[0074] The intelligent recognition mode uses the advanced Mas R-CNN deep learning model for instance segmentation, which can automatically identify and accurately delineate common subjects in the image. The recognition results are presented in the form of a mask, and the depth value is assigned after the user clicks to confirm. The Mas R-CNN deep learning model has high accuracy in most scenarios. If the recognition fails, the user can switch back to manual mode at any time to make corrections.

[0075] The depth adjustment mode is used to set the depth value of the corresponding pixel area in the original 2D image based on the smearing operation. The depth adjustment mode allows users to set the depth directly by smearing. In the depth adjustment mode, the brush is like a depth sculpting tool. The area smeared by the brush will be assigned a specific depth value. The specific depth value is linked with the depth slider on the user interface to achieve the effect that the deeper the press, the more significant the depth change, thereby accurately converting the user's intuitive operation into pixel depth information.

[0076] By defining multiple interaction modes for the intelligent brush, a powerful and flexible toolset for creating depth maps is provided to users. The depth adjustment mode allows users to directly shape depth undulations on the image by pushing forward or pulling back the brush. The edge feathering mode intelligently smooths the boundaries of depth changes and uses weighted interpolation calculations to make the transition between different depth levels smooth and natural, eliminating harsh jaggedness. The intelligent recognition mode uses image segmentation algorithms to automatically identify the main object in the image and generate a depth control slider that is linked to the main object. Users only need to slide the slider to assign a uniform depth value to the entire subject. This multi-mode fusion design simplifies the complex depth editing work to intuitive smearing and sliding, greatly reducing the operation threshold and significantly improving the efficiency and accuracy when processing complex edges and main objects.

[0077] The preset templates include portrait templates, landscape templates, architectural templates, and still life templates;

[0078] Each preset template includes an initial depth map layout that matches the content type of the preset template and a depth assignment range for the main object.

[0079] Each preset template includes an initial depth map layout that matches the content type. The initial depth map layout defines the hierarchical relationship between different regions in the image and sets a depth value range for the main object, ensuring that the core element is placed in a prominent and reasonable depth position in three-dimensional space. When a user selects a template, this preset depth structure and value range will be automatically applied to the current image, providing the user with a scientific and efficient starting point for creating depth maps. The user only needs to make minor adjustments based on this.

[0080] It introduces preset templates for specific content types, providing users with a professional starting point for creation and greatly simplifying the depth map construction process in specific scenarios. These templates are not simple image filters, but rather embed initial depth map layouts that match their respective content types and reasonable depth assignment ranges for the main objects. After selecting the corresponding template, a depth framework that conforms to the visual logic of the real world is automatically constructed. Users only need to make minor adjustments or emphasis on this framework to quickly generate realistic 3D effects. This provides non-professional users with a professional-level creative shortcut.

[0081] The logic for generating a global depth map of the original 2D image includes:

[0082] Based on interactive commands, a global depth map is generated using a depth propagation algorithm;

[0083] The logic for generating a global depth map using the depth propagation algorithm includes:

[0084] The initial depth map is constructed by treating each pixel in the original 2D image as a node in the depth map.

[0085] Based on interactive commands, the nodes corresponding to the pixels painted by the user through the smearing operation or the key areas defined by the selected preset template are used as seed nodes, and the depth value of the seed nodes is set to the preset first depth value, and the depth values ​​of the remaining nodes are set to the preset second depth value.

[0086] Key areas include key areas for portrait templates, key areas for landscape templates, key areas for architectural templates, and key areas for still life templates;

[0087] The key areas of a portrait template include the face, hair, and body.

[0088] The key areas of a landscape template include the background, middle ground, and foreground.

[0089] Key areas for building formwork include wall areas, door and window areas, and eaves areas;

[0090] The key areas of a still life template include the still life body area, the still life shadow area, and the still life support plane area;

[0091] Among them, the weight of the preset first depth value is the preset first weight, and the weight of the preset second depth value is the preset second weight;

[0092] Calculate the edge weight between two nodes in the initial depth graph. Use a random walk algorithm to traverse all nodes in the initial depth graph along the path with the highest edge weight, based on the preset first depth value of the seed node, to obtain the global depth graph.

[0093] The first depth value is set as the target depth value specified by the user for the main area through a smearing operation or a preset template. The preset first weight is set to a high fixed value close to 1.0 as a hard constraint in the algorithm, ensuring that the depth value is forcibly retained during the propagation process, forming the basic skeleton of the depth map.

[0094] The second depth value is set as a reference depth value automatically inferred from the secondary region. The preset second weight is set to a low fixed value between 0.5 and 0.7 as a soft constraint in the algorithm, so that the second depth value can be smoothly corrected by the neighboring strong depth signal during propagation, so as to achieve a natural depth transition.

[0095] Key regions include key regions for portrait templates, landscape templates, architectural templates, and still life templates. The division of each key region is not based on fixed physical dimensions, but is defined on a two-dimensional image plane using pixels as the unit. The input 2D image is analyzed pixel by pixel, and each pixel is classified into a certain key region. The extent of a region refers to the set of all pixels that make up that region, and its size and shape depend entirely on the image content itself. All key region divisions and extent definitions use pixels as the basic unit.

[0096] The implementation details of the depth propagation algorithm lie in its multi-stage, coarse-to-fine iterative optimization process. The input to the depth propagation algorithm is a sparse depth map generated by the user through smearing or stencils, i.e., seed points containing only a few pixel depth values. The algorithm first constructs a full-resolution multi-scale pyramid, starting from the coarsest top layer. Using these seed points as constraints, it propagates depth by solving a global energy function. The global energy function contains two terms: a data term, ensuring that the propagated depth value remains consistent with the user input at the seed point location; and a smoothing term, based on the image's color, gradient, and texture information, favoring smoothing. By assigning similar depth values ​​to regions with similar colors or gentle gradients, the boundaries of objects are respected. After a coarse-scale propagation, the result is upsampled and used as the initial value for the next finer-scale layer. At the same time, higher-resolution image details are introduced for constraint. This process iterates down layer by layer until the original resolution is restored. The depth propagation algorithm performs a boundary optimization by detecting the alignment between the depth gradient and the image gradient and sharpening the edges to ensure that the boundaries of the depth map accurately match the actual contours of the objects. This generates a dense depth map that smoothly, naturally, and accurately fills the entire image from the sparse input.

[0097] When calculating the edge weight between two nodes, pixel color difference and spatial distance are mainly considered. Color difference is measured by calculating the Euclidean distance between pixels in the RGB color space. The smaller the difference, the greater the weight. Spatial distance is measured by calculating the Euclidean distance between pixels. The closer the distance, the greater the weight. The final edge weight is the weighted sum of these two factors. The weight of color difference is set much higher than that of spatial distance to ensure that depth propagation prioritizes following the object boundary rather than a smooth transition.

[0098] The random walk algorithm marks the first depth value point set by the user as the absorbing state, and the other points as transient states. At the beginning of the iteration, the depth values ​​of all transient nodes are initialized. In each iteration, each transient node distributes its own depth value to its neighbors according to the weight of its edge with the neighboring nodes, while receiving depth values ​​from its neighbors. This process is repeated until the change in the depth values ​​of all transient nodes is less than a very small threshold. At this point, the algorithm is considered to have converged, and the final global depth map is generated.

[0099] This paper proposes a specific and efficient implementation path for generating global depth maps based on the depth propagation algorithm. The core of the depth propagation algorithm lies in simulating the natural diffusion process of depth information. Each pixel of the original 2D image is regarded as a node in the depth map, and a network structure is constructed. Pixels painted by the user with a smart brush or nodes corresponding to key areas defined by a preset template are marked as seed nodes. The seed nodes are assigned user-defined depth values. A random walk algorithm is used to allow the depth values ​​to propagate from these seed nodes along the path with the highest relevance to the image content to all surrounding unknown nodes until the entire image is covered. This depth propagation method based on content relevance can intelligently cross texture and color boundaries. The generated global depth map is extremely accurate in terms of detail preservation and hierarchical transition, and can better restore the complex spatial structure contained in the original 2D image.

[0100] The generation module also includes a detection unit and a smoothing unit;

[0101] The detection unit is used to perform convolution calculation on the global depth map by using an edge detection operator to obtain the depth gradient magnitude of each pixel in the global depth map;

[0102] Pixels with a depth gradient magnitude greater than a preset mutation threshold are identified as mutation points, and spatially continuous sets of mutation points are identified as mutation regions.

[0103] The smoothing unit is used to smooth abrupt regions using a bilateral filtering algorithm.

[0104] Edge detection is performed using the Canny operator. During convolution calculation, a Gaussian filter kernel is used, with the kernel size dynamically set to 5x5 based on the image resolution to smooth noise. The high and low thresholds are automatically determined by calculating the statistical distribution of the image gradient, rather than fixed empirical values, to adapt to different image contrasts.

[0105] The preset mutation threshold is achieved by first calculating the gradient magnitude histogram of the entire depth map, and then selecting a gradient value with a high percentile as the mutation threshold. This method can adaptively identify depth discontinuities that are statistically significant jumps, avoiding the limitations of fixed thresholds in different scenarios.

[0106] The smoothing unit uses a bilateral filtering algorithm. The spatial domain standard deviation of the bilateral filtering algorithm is fixed at 3.0 pixels, which determines the neighborhood range of the smoothing effect. The color domain standard is set to 0.1 times the standard deviation of the depth value.

[0107] By adding detection and smoothing units, the generated global depth map is refined through post-processing optimization to ensure the purity and naturalness of the final 3D effect. The detection unit uses edge detection operators to perform convolution calculations on the global depth map, accurately identifying pixels where the depth value changes drastically, i.e., abrupt change points, and marking spatially continuous sets of abrupt change points as abrupt change regions. These regions are often noise or unnatural jumps in the depth map. The smoothing unit specifically targets these abrupt change regions, using a bilateral filtering algorithm for smoothing. The unique feature of bilateral filtering is that while smoothing pixel values, it considers the spatial distance and color similarity between pixels. It can effectively eliminate depth noise while perfectly protecting the sharpness of object edges, avoiding the loss of details caused by overall image blurring, and ensuring that the boundaries between different objects in the final 3D image are clear and distinct.

[0108] The logic for dynamically dividing the focus area and non-focus area based on the user's brush position includes:

[0109] The current position of the user's brush is taken as the center point of the focus area. At the same time, the radius of the focus area is dynamically calculated based on the real-time movement speed of the brush and the preset focus sensitivity.

[0110] Based on the center point and the radius of the range, a circular region is divided on the original 2D image. The circular region is used as the focal region, and the region on the original 2D image other than the circular region is used as the non-focal region.

[0111] The mathematical expression for dynamically calculating the radius of the focal region is: ;

[0112] in, Indicates the radius of the focal region. Indicates the preset focus sensitivity of the brush. This indicates the real-time movement speed of the brush. This indicates the preset base radius coefficient. It represents a positive infinitesimal.

[0113] Real-time brush movement speed Measurements are performed based on continuous sampling points, recording the timestamp and screen coordinates of each sampling point, and the pixel distance between two consecutive sampling points is calculated. Divide by time interval ,Right now This allows us to obtain the real-time movement speed of the brush, thus achieving real-time speed response.

[0114] The preset focus sensitivity is a user-adjustable parameter. The system provides a slider control that allows users to freely set the preset focus sensitivity within the range of low-speed precision to high-speed coverage. The higher the preset focus sensitivity, the more significant the impact of brush speed on the radius of the focus area, giving users complete control over the operation feel.

[0115] The preset base radius coefficient is set to a range of 0.5 to 2.0.

[0116] A positive infinitesimal is an extremely small positive number used to prevent the denominator from being zero, ensuring that the brush remains stationary. The formula remains valid.

[0117] A dynamic and intelligent focus area division mechanism was established, achieving precise matching of computing resources and user visual attention. The current brush position is used as the center point of the focus area, and a reasonable radius is dynamically calculated based on the brush's real-time movement speed and preset focus sensitivity. The mathematical logic is as follows: This means that when the brush moves slowly, the focus area shrinks to concentrate resources on fine rendering, and when the brush moves quickly, the focus area expands accordingly to cover a wider prediction range. This adaptive focus calculation method can intelligently predict the user's focus and allocate the most valuable computing resources to where they are most needed in real time, thereby optimizing rendering performance while ensuring smooth interaction.

[0118] The logic for downgrading non-focus areas includes:

[0119] Image data in non-focal areas is downsampled to obtain low-resolution data. The sampling process is used to reduce the resolution of non-focal areas to a preset ratio of the original resolution.

[0120] The first image is obtained by performing 3D rendering on low-resolution data, and the image of the non-focal area is used as the second image.

[0121] The first and second images are upsampled and then fused to obtain a preview output image after downsampling the non-focal areas.

[0122] The logic for upsampling and fusion includes:

[0123] A high-quality interpolation algorithm is used to enlarge the first image to the same original resolution as the second image to obtain an intermediate image. At the same time, a low-frequency second image is obtained by applying Gaussian blur to the second image. The second image and the low-frequency second image are then subtracted at the pixel level to obtain a detail image.

[0124] The intermediate image and the detail image are weighted and fused to obtain the preview output image.

[0125] The downsampling process uses a bilinear interpolation algorithm to achieve a balance between computational efficiency and anti-aliasing effect. The preset ratio is fixed at one-quarter of the original resolution, that is, the length and width are reduced to 1 / 2 of the original, so as to significantly reduce the amount of computation in subsequent processing.

[0126] High-quality interpolation algorithms employ bicubic interpolation, which reconstructs edges and textures more smoothly during the magnification process and provides clearer details compared to bilinear interpolation.

[0127] The intermediate image and the detail image are weighted and fused to obtain the preview output image. The intermediate image, which has been enlarged by a high-quality interpolation algorithm, is aligned pixel-wise with the detail image extracted from the original depth map. Fixed fusion weights are assigned to these two images, with the intermediate image given a dominant weight of 0.7 and the detail image given a secondary weight of 0.3. During fusion, the final depth value of each pixel in the image is obtained by calculating the weighted average of the depth values ​​of the intermediate image and the detail image. In this way, the fused preview output image retains the smooth and coherent main depth structure generated by multi-scale propagation, while also overlaying high-frequency textures and edge details from the original image, thus providing users with a preview effect that is both consistent with the overall depth logic and rich in realism.

[0128] This paper presents an efficient and high-fidelity de-focusing strategy for non-focal areas, balancing the performance of real-time rendering with the quality of the preview image. The strategy involves downsampling the image data of non-focal areas to generate low-resolution data, which is then rapidly rendered in 3D to obtain the first image, compensating for the detail loss caused by downsampling. Simultaneously, the original non-focal area image is processed by separating low-frequency components through Gaussian blur, and then subtracting this from the original image to obtain a detail image containing details. The first image, upsampled to its original resolution, is then weighted and fused with the detail image to obtain the final preview output image. This process ensures that even in de-focusing mode, the preview image retains key visual details, reducing the perceived difference between the preview and the final result during editing.

[0129] The logic for asynchronous background processing of non-focus areas includes:

[0130] After downgrading the non-focal area, the background process is started to perform high-precision rendering on the non-focal area of ​​the original 2D image to obtain a high-precision non-focal area image.

[0131] Replace the corresponding area in the original 2D image with a high-precision non-focus area image to obtain an intermediate rendering image.

[0132] The background workflow utilizes the GPU's asynchronous computing queue. By submitting high-precision rendering tasks as independent rendering channels to the GPU's dedicated queue, they are executed in parallel with the UI rendering on the main thread, thus avoiding blocking foreground interactions and achieving a smooth user experience.

[0133] High-precision rendering uses full resolution, complex shading models, and high-quality texture filtering, while downgraded rendering uses quarter resolution, simplified shaders, and low-quality filtering, solely for providing fast, real-time interactive previews. The two differ by orders of magnitude in terms of computational cost and output quality.

[0134] An asynchronous background processing mechanism was introduced to resolve the inherent contradiction between real-time interactive response and final high-precision rendering. When degrading non-focus areas to ensure foreground smoothness, an independent background worker thread is started. The background worker thread uses idle CPU resources to perform complete high-precision 3D rendering processing on the original non-focus area image data. Once the background processing is complete, the corresponding area in the foreground preview image is seamlessly replaced with this newly generated high-precision non-focus area image to form an intermediate rendering image. This dual-thread parallel processing mode ensures that the user interface always maintains a smooth interactive experience, while silently building the final high-quality result in the background, achieving a dual guarantee of user experience and output quality.

[0135] The rendering module also includes a cache unit;

[0136] The cache unit is used to store high-precision non-focus area images generated by asynchronous background processing. When the user's brush enters the stored high-precision non-focus area image, the high-precision non-focus area image is retrieved from the cache unit for display.

[0137] The cache unit is managed using the LRU algorithm, and the upper limit of the cache unit size is fixed. When space is insufficient, the cached image that has not been accessed for the longest time will be discarded first. The call response time is in the millisecond level to ensure smooth view switching. When the user modifies the critical editing state such as the global depth map, the entire cache will be invalidated immediately to ensure data consistency.

[0138] By adding a caching unit, the system's response speed and computing resource utilization efficiency are further improved, forming an intelligent performance optimization closed loop. The caching unit is specifically used to store high-precision non-focus area images generated by asynchronous background processing and continuously monitors the user's brush movement trajectory. Once the brush enters an area where a high-precision image has been cached, the image will be directly retrieved from the caching unit for display, skipping the real-time calculation and rendering steps. This intelligent caching mechanism effectively avoids repeated rendering operations in the same area by the user, and can bring significant performance improvements for large high-resolution images or scenarios where users habitually make repeated modifications.

[0139] The logic for generating the final naked-eye 3D image by combining the global depth map includes:

[0140] Based on the global depth map, a raster rendering algorithm is used to render the intermediate rendering image from multiple viewpoints, generating N initial viewpoint images. According to the preset viewing distance, the parallax intensity between the N initial viewpoint images is dynamically adjusted, and the N initial viewpoint images with adjusted parallax intensity are synthesized to generate the final naked-eye 3D image.

[0141] The preset viewing distance is set to 60 centimeters based on the display screen size, serving as the optimal viewing reference.

[0142] The parallax intensity between N initial viewpoint images is dynamically adjusted. An infrared ranging sensor detects the user's actual viewing distance in real time and compares it with a preset value. A nonlinear algorithm is used for dynamic adjustment. When the user moves away from the screen, the parallax intensity decreases exponentially to prevent eye fatigue. When the user moves closer, it is moderately enhanced, but does not exceed the preset viewing distance.

[0143] The N initial viewpoint images, after adjusting the parallax intensity, are synthesized. A sub-pixel interleaving algorithm that precisely matches the lenticular lens grating is used to render N viewpoint images at different horizontal positions based on the dynamically adjusted parallax intensity. According to the grating angle of the display, the corresponding sub-pixels of these N viewpoint images are resampled and interleaved. The resulting composite image has each column of pixels precisely corresponding to a specific viewpoint. When light passes through the lenticular lens grating, it can accurately guide the images of different viewpoints to the user's left and right eyes, thus forming a continuous, crosstalk-free naked-eye 3D visual effect.

[0144] The method for generating the final naked-eye 3D image was clearly defined, ensuring the realism, comfort, and universality of the stereoscopic effect. Based on a carefully generated global depth map, a raster rendering algorithm was used to render the intermediate rendering image from multiple viewpoints, generating N initial viewpoint images with different perspectives. The system can dynamically adjust the parallax intensity between these N initial viewpoint images according to the preset viewing distance. When the viewing distance is far, the parallax is appropriately reduced to avoid visual fatigue, and when the viewing distance is close, the parallax is increased to enhance the stereoscopic effect. Finally, the N viewpoint images with adjusted parallax intensity are synthesized to generate the final naked-eye 3D image. This dynamic parallax adjustment function enables the generated 3D content to adapt to different viewing environments and display devices, providing users with a stereoscopic visual experience that is always in the optimal comfort range.

[0145] This invention significantly improves the generation efficiency and user experience of naked-eye 3D images through innovative interactive design and intelligent rendering mechanisms. It combines intelligent brushes with multiple modes and preset templates, allowing users to intuitively and precisely control image depth, greatly reducing the operational threshold and improving the efficiency of complex edge processing. Simultaneously, by dynamically dividing the focus and non-focus areas based on the user's brush position, it performs real-time high-definition rendering of the focus area while downgrading and asynchronously processing the non-focus area in the background, achieving intelligent allocation of computing resources. This not only effectively solves the problem of interactive lag and ensures smooth real-time preview but also takes into account the precision of the final output, powerfully promoting the popularization of naked-eye 3D technology in real-time interactive applications.

[0146] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0147] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.

Claims

1. A naked-eye 3D image automatic generation and display system, characterized in that, This includes a bootstrapping module, a generation module, and a rendering module; The guidance module is used to receive user interaction commands, which include smearing operations on the original 2D image of the naked-eye 3D image using a smart brush. The smearing operations include pushing forward, pulling back, and selecting a preset template. The smart brush includes different interaction modes. The preset templates include portrait templates, landscape templates, architectural templates, and still life templates. Each preset template includes an initial depth map layout that matches the content type of the preset template and a depth assignment range for the main object. The initial depth map layout includes the hierarchical relationship of each region in the image. The generation module is used to generate a global depth map of the original 2D image according to the interaction command using a depth propagation algorithm. The logic of generating a global depth map using a depth propagation algorithm includes: using each pixel in the original 2D image as a node of the depth map to construct an initial depth map. Based on the interaction instructions, the nodes corresponding to the pixels painted by the user through the smearing operation or the key areas defined by the selected preset template are used as seed nodes, and the depth value of the seed nodes is set to a preset first depth value, and the depth values ​​of the remaining nodes are set to a preset second depth value. Calculate the edge weight between two nodes in the initial depth graph, and use a random walk algorithm to traverse all nodes in the initial depth graph along the path with the highest edge weight, taking the preset first depth value of the seed node, to obtain the global depth graph. The rendering module is used to dynamically divide the original 2D image into focus and non-focus areas based on the user's brush position, render the focus area in real time, and perform downgrading and asynchronous background processing on the non-focus areas. The final naked-eye 3D image is generated by combining the global depth map. The logic for generating the final naked-eye 3D image by combining the global depth map includes: based on the global depth map, using a raster rendering algorithm to perform multi-viewpoint rendering on the intermediate rendering image after asynchronous background processing, generating N initial viewpoint images; dynamically adjusting the parallax intensity between the N initial viewpoint images according to a preset viewing distance; and synthesizing the N initial viewpoint images after adjusting the parallax intensity to generate the final naked-eye 3D image.

2. The naked-eye 3D image automatic generation and display system as described in claim 1, characterized in that, The different interaction modes include depth adjustment mode, edge feathering mode, and intelligent recognition mode; The depth adjustment mode is used to set the depth value of the corresponding pixel region in the original 2D image according to the smearing operation; The edge feathering mode is used to feather the edge region of the original 2D image after setting a depth value. The feathering operation is used to perform weighted interpolation calculation on the depth value of the edge region according to the trajectory intensity of the smearing operation to generate a smooth edge region. The intelligent recognition mode is used to identify the main object of the original 2D image through an image segmentation algorithm, generate a pixel mask of the main object, and guide the user to assign a depth value to the main object. The logic for performing deep assignment on the main object includes: Render a depth control control that is linked to the main object on the user interface; the depth control control control is a linear slider. The system monitors the user's sliding operation on the linear slider in real time, maps the physical position of the sliding linear slider to a preset depth range, and obtains the target depth value. Set the depth value of all pixels within the pixel mask area to the target depth value.

3. The naked-eye 3D image automatic generation and display system as described in claim 2, characterized in that, The key areas include the key areas of the portrait template, the key areas of the landscape template, the key areas of the architectural template, and the key areas of the still life template; The key areas of the portrait template include the face area, hair area, and body area; The key areas of the landscape template include the background area, the middle ground area, and the foreground area; The key areas of the building formwork include the wall area, door and window area, and eaves area; The key areas of the still life template include the still life body area, the still life shadow area, and the still life support plane area. Among them, the weight of the preset first depth value is the preset first weight, and the weight of the preset second depth value is the preset second weight.

4. The naked-eye 3D image automatic generation and display system as described in claim 3, characterized in that, The generation module also includes a detection unit and a smoothing unit; The detection unit is used to perform convolution calculation on the global depth map by using an edge detection operator to obtain the depth gradient magnitude of each pixel in the global depth map; Pixels with a depth gradient magnitude greater than a preset mutation threshold are identified as mutation points, and spatially continuous sets of mutation points are identified as mutation regions. The smoothing unit is used to smooth the abrupt change region using a bilateral filtering algorithm.

5. The naked-eye 3D image automatic generation and display system as described in claim 4, characterized in that, The logic for downgrading non-focus areas includes: The image data of the non-focal area is downsampled to obtain low-resolution data. The downsampling process is used to reduce the resolution of the non-focal area to a preset ratio of the original resolution. The low-resolution data is processed by 3D rendering to obtain a first image, and the image of the non-focal area is used as a second image; The first and second images are upsampled and then fused to obtain a preview output image after downsampling the non-focal areas. The logic for upsampling and fusion includes: A high-quality interpolation algorithm is used to enlarge the first image to the same original resolution as the second image to obtain an intermediate image. At the same time, a low-frequency second image is obtained by applying Gaussian blur to the second image. The second image and the low-frequency second image are then subtracted at the pixel level to obtain a detail image. The intermediate image and the detail image are weighted and fused to obtain a preview output image.

6. The naked-eye 3D image automatic generation and display system as described in claim 5, characterized in that, The logic for asynchronous background processing of non-focus areas includes: After downgrading the non-focal area, the background process is started, and the non-focal area in the original 2D image is rendered with high precision through the background process to obtain a high-precision non-focal area image. The high-precision non-focus area image is used to replace the corresponding area of ​​the preview output image in the original 2D image to obtain an intermediate rendering image.

7. The naked-eye 3D image automatic generation and display system as described in claim 6, characterized in that, The rendering module also includes a cache unit; The cache unit is used to store the high-precision non-focus area image generated by the asynchronous background processing. When the user's brush enters the stored high-precision non-focus area image, the high-precision non-focus area image is retrieved from the cache unit for display.

Citation Information

Patent Citations

  • Method and equipment for generating three-dimensional image on touch screen

    CN102307308A

  • Naked eye 3D image processing method, device and equipment

    CN109522866A