Grating type 3D display method and tablet computer
Through the raster 3D display method, deep neural networks are used to process light field image sequences to generate high-quality three-dimensional images, solving the problems of complexity and limited image resolution of three-dimensional information capture devices in the prior art, and achieving high-quality three-dimensional displays that operate efficiently on ordinary computing devices.
Patent Information
- Application Number
- CN202510205209.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, three-dimensional information capture has problems such as complex camera connections, large equipment, and limited image resolution.
Using the raster 3D display method, the light field image sequence is obtained through the camera array and input it into the pre-trained fully connected deep neural network, a grid with multiple resolution levels is generated for hash encoding, and an element image array is generated for 3D display through the pixel rearrangement method.
The light field image acquisition process is simplified, the number and complexity of hardware devices is reduced, the efficiency and quality of image synthesis is improved, and it can operate efficiently on ordinary computing devices.
Smart Images

Figure CN120111205A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of 3D display, and in particular to a grating type 3D display method and a tablet computer. Background Art
[0002] With the continuous development of three-dimensional display technology, how to efficiently acquire, process and display high-quality three-dimensional images has become a hot topic of research. Existing light field imaging methods mainly rely on multi-camera arrays to capture view information from different angles to reconstruct the three-dimensional image of the object. These methods usually use camera array-based technology to capture the three-dimensional information of the object by taking images from multiple perspectives. However, the traditional multi-camera array method has many challenges, such as complex connections between cameras, bulky equipment, limited image resolution, and lens distortion during the capture process. In addition, how to effectively process and synthesize these view images for high-quality three-dimensional visual display remains an urgent problem to be solved. Summary of the invention
[0003] In view of the above technical problems, the present invention provides a grating 3D display method and a tablet computer to solve the problems of complex camera connection, bulky equipment and limited image resolution in the prior art in capturing three-dimensional information.
[0004] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0005] According to one aspect of the present invention, a grating 3D display method is disclosed, the method comprising: Acquire a sequence of light field images, wherein the sequence of light field images is obtained by shooting at preset step intervals when a camera array moves horizontally, wherein the camera array includes a plurality of digital cameras arranged along a vertical axis, wherein the digital camera located in the center focuses on the photographed object itself, and the remaining digital cameras rotate upward or downward at a given angle according to the focus position of the central digital camera to focus on the photographed object; Inputting the light field image sequence into a pre-trained fully connected deep neural network, the fully connected deep neural network determines the camera pose of each image in the light field image sequence based on predefined view parameters, generates a grid with multiple resolution levels based on the determined camera pose, hash codes the grid, eliminates hash conflicts by linear interpolation, calculates the encoded hash values to obtain a corrected camera pose, and obtains N*N visual images, where N represents the number of cameras required in the horizontal and vertical directions; Based on the pixel rearrangement method, the N*N visual images are rearranged according to the determined camera posture to obtain an element image array, and the element image array is input into a grating type 3D display screen for 3D display.
[0006] Furthermore, after the camera array acquires the light field image sequence, an average histogram of the images captured by the central digital camera in the light field image sequence is calculated, and the images captured by other digital cameras are matched with the average histogram to perform brightness and color correction to eliminate color and brightness differences between different digital cameras.
[0007] Furthermore, during hash coding, the grid of the input vision is created from wide to fine, and the level of refinement increases from small to large. The number of levels determines the number of types of grids created. Different grids have different angular index values, and the average value of the angular index value of each grid cannot be equal to the angular index value of other grids.
[0008] Furthermore, when eliminating hash conflicts by linear interpolation, it includes: The hash integer coordinates corresponding to the angle index values are stored in a corresponding lookup table, and linear interpolation is performed according to the corresponding position alignment in the input light field image sequence. During the linear interpolation, a linear polynomial is used to perform curve fitting to create an image of a new angle in a discrete set of the light field image sequences.
[0009] Furthermore, when the element image array is obtained based on the pixel rearrangement method, it specifically includes: The number of virtual cameras, the distance between adjacent virtual cameras, and the distance between the virtual cameras and the image plane are set, and each element image in the element image array is calculated based on the following formula: ; in, is a perspective image set, which includes images from different perspectives, are the coordinates of the element image, and Respectively represent the horizontal and vertical positions of the virtual camera, and Respectively represent the number of horizontal and vertical pixels of each element image.
[0010] According to another aspect of the present invention, a grating 3D display tablet computer is disclosed, the tablet computer comprising a display screen having a grating lens array, and a processor, the processor being configured to: Acquire a sequence of light field images, wherein the sequence of light field images is obtained by shooting at preset step intervals when a camera array moves horizontally, wherein the camera array includes a plurality of digital cameras arranged along a vertical axis, wherein the digital camera located in the center focuses on the photographed object itself, and the remaining digital cameras rotate upward or downward at a given angle according to the focus position of the central digital camera to focus on the photographed object; Inputting the light field image sequence into a pre-trained fully connected deep neural network, the fully connected deep neural network determines the camera pose of each image in the light field image sequence based on predefined view parameters, generates a grid with multiple resolution levels based on the determined camera pose, hash codes the grid, eliminates hash conflicts by linear interpolation, calculates the encoded hash values to obtain a corrected camera pose, and obtains N*N visual images, where N represents the number of cameras required in the horizontal and vertical directions; Based on the pixel rearrangement method, the N*N visual images are rearranged according to the determined camera posture to obtain an element image array, and the element image array is input into the display screen for 3D display.
[0011] The technical solution disclosed in this disclosure has the following beneficial effects: Compared with the traditional multi-camera array system, the present invention can simplify the light field image acquisition process by using only a small number of digital camera arrays. This simplified method not only reduces the number of hardware devices, but also reduces the complexity and cost of the equipment, and is suitable for a wider range of application scenarios; Based on the deep neural network model, high-quality 3D images are generated by efficiently synthesizing the perspectives of captured images, which improves the efficiency and quality of image synthesis. By using a simplified light field image acquisition method and efficient neural network perspective synthesis, the present invention reduces hardware costs while reducing dependence on high-performance computing resources. Compared with traditional three-dimensional reconstruction methods that require a large number of cameras and high computing power, the present invention can run efficiently on ordinary computing devices while providing high-quality three-dimensional visual effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 A flowchart of a grating 3D display method in an embodiment of this specification; Figure 2 This is a structural block diagram of a grating 3D display tablet computer in an embodiment of this specification. DETAILED DESCRIPTION
[0013] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as being limited to the examples set forth herein; on the contrary, these embodiments are provided so that the present disclosure will be more comprehensive and complete, and the concepts of the example embodiments are fully conveyed to those skilled in the art. The described features, structures, or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced while omitting one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known technical solutions are not shown or described in detail to avoid obscuring various aspects of the present disclosure.
[0014] In addition, the accompanying drawings are only schematic illustrations of the present disclosure. The same reference numerals in the figures represent the same or similar parts, and thus their repeated description will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities, which do not necessarily correspond to physically or logically independent entities. These functional entities can be implemented in software form, or implemented in one or more hardware modules or integrated circuits, or implemented in different networks and / or processor devices and / or microcontroller devices.
[0015] In one embodiment, if Figure 1 As shown, this specification provides a raster 3D display method, and the execution subject of the method can be a computer, a server, a mobile phone, a tablet computer, etc. The method can specifically include steps S101 to S103: In step S101, a light field image sequence is acquired. The light field image sequence is obtained by shooting a camera array at a preset step interval when the camera array moves horizontally. The camera array includes a plurality of digital cameras arranged along a vertical axis. The digital camera located in the center focuses on the photographed object itself, and the remaining digital cameras rotate upward or downward at a given angle according to the focus position of the central digital camera to focus on the photographed object.
[0016] Among them, the light field image sequence refers to a series of images acquired from different perspectives. These images contain different angle information of the object and are used to construct a three-dimensional view. These images are taken by a camera array. The camera array consists of multiple digital cameras, and these cameras are arranged along the vertical axis. The camera array is a vertical layout, in which each camera can capture different angles of the object. The camera array moves horizontally, that is, the camera array moves horizontally as a whole, and the step length of each movement is pre-set (called the preset step length interval). This horizontal movement allows the camera array to capture different perspectives of the object at different positions. In this way, by moving horizontally, the camera can cover a wider observation area and capture more perspective information of the object. In the camera sequence, the digital camera in the center is responsible for focusing on the object itself, that is, it is aimed at and clearly captures the front view of the object; for other cameras (cameras not in the center), their focus positions will be adjusted according to the focus position of the central camera, that is, rotate up or down: these cameras will rotate up and down according to the given angle, thereby changing their shooting angles to ensure that they focus on different parts of the object. These rotation angles are pre-set so that each camera can capture different viewpoints.
[0017] In step S102, the light field image sequence is input into a pre-trained fully connected deep neural network, and the fully connected deep neural network determines the camera pose of each image in the light field image sequence based on predefined view parameters, generates a grid with multiple resolution levels based on the determined camera pose, hash codes the grid, eliminates hash conflicts through linear interpolation, calculates the encoded hash values, obtains the corrected camera pose, and obtains N*N visual images, where N represents the number of cameras required in the horizontal and vertical directions.
[0018] Among them, the camera pose refers to the position and direction of each camera. In the present invention, the network infers and determines the camera pose of each image based on the input light field image sequence and predefined view parameters. The view parameters include information such as the viewing angle, the interval between cameras, and the rotation angle. Through these parameters, the neural network can calculate the specific position and direction of each camera (i.e., the camera pose). Based on the determined camera pose, the network uses this information to generate a multi-resolution grid. The different resolution levels of the grid mean that the grid is divided at different levels of detail. For example, a coarse grid is used to represent a wide range of viewing angles, while a fine grid represents a finer viewing angle. In this way, the network can process image information at different levels to ensure that the three-dimensional scene is constructed from coarse to fine. Hash coding is to convert the spatial information of each grid into a concise numerical representation. Through hash coding, each camera view (or grid) has a unique hash value to help identify different viewing angles. Linear interpolation is used to eliminate hash conflicts, that is, to avoid repeated mapping of hash values of different viewing angles. Linear interpolation ensures that each grid can obtain a unique hash value by smoothing the hash values of different grids. Through these interpolation processes, the network can more accurately calculate and correct the pose of each camera and eliminate any errors caused by hash conflicts. Finally, through the corrected camera poses, the network can generate N×N visual images, where N represents the number of cameras required in the horizontal and vertical directions. Through these images, a more complete three-dimensional view can be formed, and the perspective captured by each camera provides a part of the three-dimensional data for the final synthesis. In summary, the process of step 102 is to use a fully connected deep neural network to calculate the camera position and direction (camera pose) corresponding to each image based on the input light field image sequence, and to efficiently process the image by generating a grid with multiple resolution levels. Through hash coding and linear interpolation methods, the network can eliminate conflicts and correct camera poses, and finally generate N×N high-quality visual images, which will be used for the synthesis and display of three-dimensional images.
[0019] In step S103, based on the pixel rearrangement method, the N*N visual images are rearranged according to the determined camera posture to obtain an element image array, and the element image array is input into a grating 3D display screen for 3D display.
[0020] Among them, the pixel rearrangement method is a technology that rearranges the pixels in an image according to specific rules. In the present invention, the purpose of pixel rearrangement is to rearrange the N×N visual images calculated by the fully connected deep neural network so that they meet the display requirements and can form a correct three-dimensional effect on the final 3D display screen. Based on the known camera posture, the visual image will be rearranged according to the spatial position and viewing angle to ensure that the actual position corresponding to the pixel of each image is consistent with its position in the three-dimensional space. This rearrangement is to ensure that the display device can correctly present three-dimensional images from different viewing angles. Each image in the element image array represents a view at a certain angle. When the audience watches from different angles, they can see different images and feel the three-dimensional effect of the object. The grating 3D display screen uses a grating lens array to guide the pixels of each element image to different directions, so that each audience can see different views of the object from different angles, thereby producing a realistic three-dimensional effect.
[0021] In one embodiment, after the camera array acquires the light field image sequence, an average histogram of the images captured by the central digital camera in the light field image sequence is calculated, and the images captured by other digital cameras are matched with the average histogram to perform brightness and color correction to eliminate color and brightness differences between different digital cameras.
[0022] Among them, the camera array captures a sequence of light field images of the object through multiple cameras. Among these acquired light field image sequences, the image taken by the digital camera in the center is usually considered to be the most standard image, that is, it takes a "frontal" shot of the object. In order to correct the color and brightness differences of the images of other cameras, it is necessary to first calculate the average histogram of the image taken by the central camera. The histogram is a statistical graph of the brightness or color distribution of the image, which describes the frequency of different brightness or color values in the image. By calculating the average histogram of the image taken by the central camera, a reference color and brightness distribution can be obtained. The images taken by other digital cameras (i.e., the images of non-central cameras) will have differences in brightness and color due to factors such as camera configuration and lighting environment. Therefore, before subsequent processing, these images need to be matched with the average histogram of the image taken by the central camera. Histogram matching adjusts the brightness and color of these images to make them closer to the average brightness and color distribution of the image taken by the central camera. Through this matching, the color and brightness inconsistency problems caused by differences in settings and environments between different cameras can be eliminated. In the process of histogram matching, the brightness and color distribution of the image are adjusted to ensure that the visual effects of all images are more consistent. This correction process helps eliminate color and brightness differences between cameras, so that the color tone, brightness, etc. of all images remain uniform during the synthesis or display process, thereby improving the quality of the final 3D image.
[0023] In one embodiment, during hash coding, the grid of the input vision is created from wide to fine, and the level of refinement increases from small to large. The number of levels determines the number of types of grids created, different grids have different angular index values, and the average value of the angular index value of each grid cannot be equal to the angular index value of other grids.
[0024] And, when eliminating hash collisions by linear interpolation, including: The hash integer coordinates corresponding to the angle index values are stored in a corresponding lookup table, and linear interpolation is performed according to the corresponding position alignment in the input light field image sequence. During the linear interpolation, a linear polynomial is used to perform curve fitting to create an image of a new angle in a discrete set of the light field image sequences.
[0025] Among them, in the hash coding process, grids are first created based on the input visual data (such as images obtained from different perspectives). These grids gradually change from wide to fine. A wide grid refers to a larger division of the input perspective, which is suitable for processing a large range of perspectives, while a fine grid is used to divide the image more finely, which is suitable for capturing more details and smaller perspective differences. In this process, the refinement level of the grid is carried out from a lower level to a higher level. That is, a coarser grid is first used to make a preliminary division of the perspective, and then the grid division is refined by increasing the level to obtain higher resolution perspective data. The refinement level from small to large means that as the level increases, the grid division will become finer and finer, and each grid will contain more perspective information. The level of the grid (the degree of refinement from low to high) determines the number of types of the final grid. Each level represents a different grid size. The more levels there are, the higher the degree of refinement of the grid, the more perspective data can be obtained, and more types of grids can be generated. Specifically, the more levels there are, the more grids of different sizes and resolutions can be created, so that perspective information can be encoded and processed at different resolutions. Each grid has a unique angular index value, which is used to identify the specific location of the grid in space. Different grids have different angular index values to ensure that each grid can be represented independently and uniquely. The average value of the angular index value of each grid needs to remain unique and cannot be the same as the average angular index value of other grids. This is to prevent hash conflicts, that is, two different grids are incorrectly mapped to the same hash position, resulting in information loss or overlap. The different average values of the angular index values ensure that each grid has a unique identifier during the hash encoding process, thereby avoiding confusion or repeated mapping between different grids and ensuring the accuracy of the hash encoding and the integrity of the data.
[0026] In one embodiment, when the element image array is obtained based on the pixel rearrangement method, it specifically includes: The number of virtual cameras, the distance between adjacent virtual cameras, and the distance between the virtual cameras and the image plane are set, and each element image in the element image array is calculated based on the following formula: ; in, is a perspective image set, which includes images from different perspectives, are the coordinates of the element image, and Respectively represent the horizontal and vertical positions of the virtual camera, and Respectively represent the number of horizontal and vertical pixels of each element image.
[0027] In one embodiment, the fully connected deep neural network, when trained, includes: The light field image sequence acquired by the camera array is input into a fully connected deep neural network. The light field image sequence contains image data from different perspectives, and each image represents an image of an object taken from a different angle. The neural network determines the camera pose of each image based on the input light field image sequence and predefined view parameters (such as viewing angle, spacing between cameras, and rotation angle). According to the determined pose of each camera, a multi-resolution grid is generated. The resolution of these grids gradually changes from coarse to fine. The grids are used to represent different levels of perspective information in order to process different levels of image details. For each grid, the neural network performs hash encoding, which converts the spatial information of the grid into a compact hash value. Through linear interpolation, the network eliminates possible hash conflicts. During the training process, the network parameters are gradually optimized by feeding back the error between the input image and the generated view. Through this iterative optimization process, the network can eventually generate high-quality view synthesis based on the input view. At the end of the training, N×N view images are generated based on the determined camera poses, where N represents the number of cameras required in the horizontal and vertical directions. These view images will be used for subsequent image synthesis and 3D reconstruction.
[0028] Based on the same idea, Figure 2 As shown, an exemplary embodiment of the present disclosure further provides a grating 3D display tablet computer, the tablet computer comprising a display screen 201 having a grating lens array, and a processor 202, the processor 202 being configured to: Acquire a sequence of light field images, wherein the sequence of light field images is obtained by shooting at preset step intervals when a camera array moves horizontally, wherein the camera array includes a plurality of digital cameras arranged along a vertical axis, wherein the digital camera located in the center focuses on the photographed object itself, and the remaining digital cameras rotate upward or downward at a given angle according to the focus position of the central digital camera to focus on the photographed object; Inputting the light field image sequence into a pre-trained fully connected deep neural network, the fully connected deep neural network determines the camera pose of each image in the light field image sequence based on predefined view parameters, generates a grid with multiple resolution levels based on the determined camera pose, hash codes the grid, eliminates hash conflicts by linear interpolation, calculates the encoded hash values to obtain a corrected camera pose, and obtains N*N visual images, where N represents the number of cameras required in the horizontal and vertical directions; Based on the pixel rearrangement method, the N*N visual images are rearranged according to the determined camera posture to obtain an element image array, and the element image array is input into the display screen for 3D display.
[0029] Compared with the traditional multi-camera array system, the present invention can simplify the light field image acquisition process by using only a small number of digital camera arrays. This simplified method not only reduces the number of hardware devices, but also reduces the complexity and cost of the equipment, and is suitable for a wider range of application scenarios; Based on the deep neural network model, high-quality 3D images are generated by efficiently synthesizing the perspectives of captured images, which improves the efficiency and quality of image synthesis. By using a simplified light field image acquisition method and efficient neural network perspective synthesis, the present invention reduces hardware costs while reducing dependence on high-performance computing resources. Compared with traditional three-dimensional reconstruction methods that require a large number of cameras and high computing power, the present invention can run efficiently on ordinary computing devices while providing high-quality three-dimensional visual effects.
[0030] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the exemplary implementation of the present disclosure.
[0031] In addition, the above-mentioned figures are only schematic illustrations of the processes included in the method according to the exemplary embodiment of the present disclosure, and are not intended to be limiting. It is easy to understand that the processes shown in the above-mentioned figures do not indicate or limit the time sequence of these processes. In addition, it is also easy to understand that these processes can be performed synchronously or asynchronously, for example, in multiple modules.
[0032] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary technical means in the art that are not disclosed in the present disclosure. The specification and embodiments are to be considered exemplary only, and the true scope and spirit of the present disclosure are indicated by the claims.
Claims
1. A grating 3D display method, characterized in that: The method comprises: Acquire a sequence of light field images, wherein the sequence of light field images is obtained by shooting at preset step intervals when a camera array moves horizontally, wherein the camera array includes a plurality of digital cameras arranged along a vertical axis, wherein the digital camera located in the center focuses on the photographed object itself, and the remaining digital cameras rotate upward or downward at a given angle according to the focus position of the central digital camera to focus on the photographed object; Inputting the light field image sequence into a pre-trained fully connected deep neural network, the fully connected deep neural network determines the camera pose of each image in the light field image sequence based on predefined view parameters, generates a grid with multiple resolution levels based on the determined camera pose, hash codes the grid, eliminates hash conflicts by linear interpolation, calculates the encoded hash values to obtain a corrected camera pose, and obtains N*N visual images, where N represents the number of cameras required in the horizontal and vertical directions; Based on the pixel rearrangement method, the N*N visual images are rearranged according to the determined camera posture to obtain an element image array, and the element image array is input into a grating type 3D display screen for 3D display.
2. The grating 3D display method according to claim 1, characterized in that: After the camera array acquires the light field image sequence, an average histogram of the images captured by the central digital camera in the light field image sequence is calculated, and the images captured by other digital cameras are matched with the average histogram to perform brightness and color correction to eliminate color and brightness differences between different digital cameras.
3. The grating 3D display method according to claim 1, characterized in that: During hash coding, the grid of the input vision is created from wide to fine, and the level of refinement increases from small to large. The number of levels determines the number of types of the grids created. Different grids have different angular index values, and the average value of the angular index value of each grid cannot be equal to the angular index value of other grids.
4. The grating 3D display method according to claim 3, characterized in that: When eliminating hash collisions by linear interpolation, including: The hash integer coordinates corresponding to the angle index values are stored in a corresponding lookup table, and linear interpolation is performed according to the corresponding position alignment in the input light field image sequence. During the linear interpolation, a linear polynomial is used to perform curve fitting to create an image of a new angle in a discrete set of the light field image sequences.
5. The grating 3D display method according to claim 1, characterized in that: When the element image array is obtained based on the pixel rearrangement method, it specifically includes: The number of virtual cameras, the distance between adjacent virtual cameras, and the distance between the virtual cameras and the image plane are set, and each element image in the element image array is calculated based on the following formula: ; in, is a perspective image set, which includes images from different perspectives, are the coordinates of the element image, and Respectively represent the horizontal and vertical positions of the virtual camera, and Respectively represent the number of horizontal and vertical pixels of each element image.
6. A grating 3D display tablet computer, characterized in that: The tablet computer includes a display screen having a lenticular lens array, and a processor, wherein the processor is configured to: Acquire a sequence of light field images, wherein the sequence of light field images is obtained by shooting at preset step intervals when a camera array moves horizontally, wherein the camera array includes a plurality of digital cameras arranged along a vertical axis, wherein the digital camera located in the center focuses on the photographed object itself, and the remaining digital cameras rotate upward or downward at a given angle according to the focus position of the central digital camera to focus on the photographed object; Inputting the light field image sequence into a pre-trained fully connected deep neural network, the fully connected deep neural network determines the camera pose of each image in the light field image sequence based on predefined view parameters, generates a grid with multiple resolution levels based on the determined camera pose, hash codes the grid, eliminates hash conflicts by linear interpolation, calculates the encoded hash values to obtain a corrected camera pose, and obtains N*N visual images, where N represents the number of cameras required in the horizontal and vertical directions; Based on the pixel rearrangement method, the N*N visual images are rearranged according to the determined camera posture to obtain an element image array, and the element image array is input into the display screen for 3D display.