Imaging method, imaging apparatus, head-mounted display, and storage medium
By modulating spherical waves and decoding high-dimensional reconstruction models, and integrating image and depth information acquisition, the problems of large size and high cost of existing imaging systems are solved, and miniaturized and low-cost multidimensional information acquisition is achieved.
Patent Information
- Application Number
- PCT/CN2025/119830
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-12
- Filing Date
- 2025-09-08
- Publication Date
- 2026-03-19
AI Technical Summary
Existing imaging systems require additional depth cameras to acquire object depth information, resulting in large system size and high cost.
By using a mask to modulate the phase or amplitude of spherical waves emitted from a point light source, a grayscale image is generated. A high-dimensional reconstruction model is then used to decode depth information, integrating image and depth information acquisition functions.
The size and weight of the imaging system have been reduced, lowering production and usage costs, while enabling multi-dimensional information detection of the scene.
Smart Images

Figure CN2025119830_19032026_PF_FP_ABST
Abstract
Description
Imaging method, imaging device, head-mounted display, and storage medium
[0001] Priority information
[0002] This application claims priority to and the benefit of the filing date of the patent application with the China National Intellectual Property Office on September 12, 2024, with the patent application number of “2024112817708” and incorporates it by reference in its entirety. TECHNICAL FIELD
[0003] The present application relates to the technical field of image processing, and more particularly, to an imaging method, an imaging device, a head-mounted display, and a non-transitory computer readable storage medium. BACKGROUND
[0004] At present, when collecting depth information of objects in a scene, in an imaging system, it is usually necessary to set a depth camera in addition to setting an imaging camera for capturing image information to collect depth information, resulting in a large volume and high cost of the imaging system. SUMMARY
[0005] The embodiments of the present application provide an imaging method, an imaging device, a head-mounted display, and a non-transitory computer readable storage medium. The volume of the imaging system required for imaging can be reduced, and the cost can be saved.
[0006] To this end, one object of the present application is to provide an imaging method, comprising: phase modulating a spherical wave emitted by a point light source through a mask to change a propagation parameter of the spherical wave, the propagation parameter comprising at least one of a phase and an amplitude, the parameters of the spherical waves of different depths being different after being modulated by the mask, the photographed scene comprising a plurality of the point light sources; generating a grayscale image based on the spherical wave with the changed propagation parameter; and performing high-dimensional reconstruction on the grayscale image based on a preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information.
[0007] Another object of the present application is to provide an imaging device, comprising a modulation module, a generation module, and a reconstruction module, wherein the modulation module is configured to phase modulate a spherical wave emitted by a point light source through a mask to change a propagation parameter of the spherical wave, the propagation parameter comprising at least one of a phase and an amplitude, the parameters of the spherical waves of different depths being different after being modulated by the mask, the photographed scene comprising a plurality of the point light sources; the generation module is configured to generate a grayscale image based on the spherical wave with the changed propagation parameter; and the reconstruction module is configured to perform high-dimensional reconstruction on the grayscale image based on a preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information.
[0008] Still another object of the present application is to provide a head-mounted display comprising a processor and a memory; the memory stores a computer program, and the processor implements the imaging method according to any one of the above embodiments when executing the program.
[0009] Still another object of the present application is to provide a non-transitory computer-readable storage medium having a computer program stored thereon, and the computer program, when executed by a processor, implements the imaging method according to any one of the above embodiments.
[0010] The imaging method, the imaging device, the head-mounted display and the non-transitory computer-readable storage medium of the embodiments of the present application change the propagation parameters of the spherical wave emitted by the point light source by phase modulation of the spherical wave by the mask, wherein the propagation parameters include at least one of the phase and the amplitude, and the parameters of the spherical wave at different depths are different after being modulated by the mask, the captured scene includes a plurality of point light sources, that is, the depth information of the spherical wave can be captured by the mask to reduce the volume and weight of the hardware required when capturing the depth information for imaging, and to save production costs; then the depth information is encoded based on the spherical wave with changed propagation parameters, and a grayscale image is generated; and then the grayscale image is high-dimensional reconstructed based on a preset high-dimensional reconstruction model, and the depth information corresponding to each pixel of the grayscale image is decoded to generate a high-dimensional image containing the depth information, that is, the high-dimensional image can not only display the scene content of the captured scene, but also accurately reflect the depth information of the captured scene and the objects in the captured scene, and realize the detection of multi-dimensional information of the scene.
[0011] Additional aspects and advantages of the embodiments of the present application will be in part apparent and in part pointed out hereinafter in the description of the embodiments of the present application. BRIEF DESCRIPTION OF DRAWINGS
[0012] The above and / or additional aspects and advantages of the present application will become apparent and be readily appreciated from the description of the embodiments, taken in conjunction with the following drawings in which:
[0013] FIG. 1 is a schematic diagram of an application scenario of the imaging method according to some embodiments of the present application;
[0014] FIG. 2 is a flowchart of the imaging method according to some embodiments of the present application;
[0015] FIG. 3 is a schematic diagram of a scene of the imaging method according to some embodiments of the present application;
[0016] FIG. 4 is a flowchart of the imaging method according to some embodiments of the present application;
[0017] FIG. 5 is a flowchart of the imaging method according to some embodiments of the present application;
[0018] FIG. 6 is a schematic diagram of a scene of the imaging method according to some embodiments of the present application;
[0019] FIG. 7 is a flowchart of an imaging method according to some embodiments of the present application;
[0020] FIG. 8 is a flowchart of an imaging method according to some embodiments of the present application;
[0021] FIG. 9 is a flowchart of an imaging method according to some embodiments of the present application;
[0022] FIG. 10 is a scenario diagram of an imaging method according to some embodiments of the present application;
[0023] FIG. 11 is a flowchart of an imaging method according to some embodiments of the present application;
[0024] FIG. 12 is a flowchart of an imaging method according to some embodiments of the present application;
[0025] FIG. 13 is a flowchart of an imaging method according to some embodiments of the present application;
[0026] FIG. 14 is a scenario diagram of an imaging method according to some embodiments of the present application;
[0027] FIG. 15 is a block diagram of an imaging device according to some embodiments of the present application;
[0028] FIG. 16 is a connection state diagram of a non-volatile computer readable storage medium and a processor according to some embodiments of the present application. DETAILED DESCRIPTION
[0029] Embodiments of the present application are described in detail below with reference to the accompanying drawings, in which examples of embodiments are shown, wherein the same or similar reference numerals represent the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are examples for explaining the embodiments of the present application and should not be construed as limiting the embodiments of the present application.
[0030] For the convenience of understanding the present application, the terms appearing in the present application are explained as follows:
[0031] Extended Reality (XR) refers to combining reality and virtuality through a computer to create a virtual environment that can be interacted with by a human. XR includes Augmented Reality (AR), Virtual Reality (VR), and Mixed Reality (MR).
[0032] Augmented Reality (AR): also known as augmented reality, is a technology that skillfully combines virtual information with the real world, widely uses multimedia, high-dimensional modeling, real-time tracking and registration, intelligent interaction, sensing and other technical means, simulates and emulates computer-generated text, images, high-dimensional models, music, video and other virtual information, and applies it to the real world, so that the two kinds of information complement each other, thereby realizing the "enhancement" of the real world.
[0033] Virtual Reality (VR): also known as virtual reality or spiritual reality technology. Virtual reality technology includes computer, electronic information, simulation technology, its basic implementation is mainly based on computer technology, using and integrating high-dimensional graphics technology, multimedia technology, simulation technology, display technology, servo technology and other latest developments of various high-tech technologies, with the help of the graphics processing unit (GPU) in the VR device to process the image in the current scene to produce a virtual world with realistic high-dimensional vision, tactile, olfactory and other sensory experiences, so that people in the virtual world have a sense of being there.
[0034] Mixed Reality (MR): refers to a new visualization environment generated by merging the real world and the virtual world, in which physical and digital objects coexist and can interact with the real world in real time and obtain information in time.
[0035] Detecting multi-dimensional information (such as depth information and time dimension information) of an object is of great significance to many application fields. For example, in the XR field, capturing depth information and time dimension information of a scene can help users accurately interact with objects in the virtual environment and the real environment; in the field of self-moving robots, detecting multi-dimensional information of an object can help self-moving robots to avoid obstacles and perform grasping operations; in the field of automatic driving of vehicles, detecting multi-dimensional information of the surrounding environment can help vehicles to identify and analyze the surrounding environment and plan paths.
[0036] In the XR field, taking the detection of depth information of an object as an example, when collecting the depth information of a scene, an imaging camera for capturing image information is usually set up, and a depth camera is additionally set up to collect depth information; and in traditional lens imaging technology, the size and volume of the lens are usually proportional to the focal length of the lens, that is, the longer the focal length, the larger the lens, resulting in a large size and volume of the final imaging system, which is relatively limited in application in various fields.
[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following takes the imaging method of the present application applied to a head-mounted display (AR head-mounted display) of augmented reality technology (Augmented Reality, AR) as an example to introduce the imaging method of the present application.
[0038] It can be understood that the imaging method of the present application can also be widely applied to the fields of autonomous driving, self-moving robots, Internet of Things, and human interaction, which need to collect depth information of a scene, and the implementation principles thereof are basically similar to those of the following embodiments, which will not be described here. It needs to be emphasized that this is only an example and is not a specific limitation on the applicable scope of the embodiments of the present disclosure.
[0039] Please refer to FIG. 1, which is an application scenario diagram of an imaging method provided by an embodiment of the present application. The application scenario provided by the present application includes a head-mounted display 100 (such as an AR head-mounted display), which includes a body 10, a display (such as a display 20 and a display 30 shown in FIG. 1), and a mask 40, which can be arranged on the display (as shown in FIG. 1, arranged on the display 20 and the display 30).
[0040] Based on the above introduction of the related scene, the present application provides an imaging method, which will be described in detail as follows.
[0041] Please refer to FIG. 2, which is a flowchart of the imaging method provided by an embodiment of the present application. The imaging method provided by the present application is implemented by steps 011, 012, and 013, which will be described in detail as follows.
[0042] Step 011: Modulating the phase of the spherical wave emitted by the point light source through the mask to change the propagation parameter of the spherical wave, the propagation parameter including at least one of the phase and the amplitude, and the spherical wave of different depths has different parameters after being modulated by the mask, and the photographed scene includes a plurality of point light sources.
[0043] Wherein, the mask can be an element capable of changing the propagation parameter of the spherical wave. For example, the mask can be a phase mask (which can be used to modulate the phase of the spherical wave); or it can also be an amplitude mask (which can be used to modulate the amplitude of the spherical wave), etc.
[0044] Wherein, the spherical wave can be a light wave with a concentric spherical wave front.
[0045] Wherein, the depth can be the distance between the position of the point light source and the position of the spherical wave emitted by the point light source when captured; it can also be the distance between the position of the point light source and the position of the spherical wave when the propagation parameter of the spherical wave changes, etc.
[0046] The propagation parameter can include a phase, or the propagation parameter can include an amplitude, or the propagation parameter can include a phase and an amplitude, which are not limited in the present application.
[0047] Specifically, a user can view a three-dimensional (3D) scene through a head-mounted display, and the 3D scene can be a projection of a real-world scene in which the user is located, or a projection after the real world and the virtual world are fused, etc. The photographed scene can be a scene region in the scene located within the field of view of the user. An object in the scene can be regarded as an object composed of countless point light sources, and the light waves emitted by each point light source are spherical waves. Referring to FIG. 3, a depth can be a distance Z between a point light source 60 and a mask 40 by setting the mask 40 for the head-mounted display and a sensor 50 (for example, the sensor 50 can be a charge-coupled device (CCD), a polarization encoding image sensor, etc.) that captures the spherical wave signal.
[0048] The spherical wave emitted by the point light source reaches the mask after propagating a distance Z, and the mask can perform phase modulation on a propagation parameter (at least one of a phase and an amplitude of the spherical wave) of the spherical wave. After the spherical wave of different depths passes through the mask, the propagation parameter will change differently, that is, there will be a phase difference after the spherical wave of different depths passes through the mask. For example, taking an amplitude mask as an example, the amplitude mask can modulate the amplitude of the spherical wave passing through the amplitude mask, and the amplitudes of the spherical waves of different depths will change differently (wherein the specific change and the physical characteristics of the amplitude mask are related).
[0049] Optionally, the mask includes a phase mask, the phase mask includes a plurality of mask units, and a phase distribution of the phase mask is determined based on heights of the mask units.
[0050] The phase mask includes a plurality of mask units, the modulation of each mask unit on the spherical wave is independent, the height of the mask unit affects the optical path difference of the spherical wave, thereby modulating the phase of the spherical wave, and therefore the height of each mask unit determines the phase distribution of the phase mask.
[0051] For example, taking a phase mask with a unit structure of microns as an example, assuming that the phase distribution of the phase mask is φ(x, y), a mapping relationship between the phase distribution φ(x, y) and the height h(x, y) of the mask unit is established to determine the phase distribution φ(x, y) of the mask based on the height h(x, y) of each mask unit.
[0052] Step 012: generating a grayscale image based on the spherical wave with the changed propagation parameter.
[0053] The gray-scale image can be a two-dimensional image (2D image), and the gray-scale image can include image information (such as light intensity information) capable of showing the content of the real scene.
[0054] Specifically, please continue to refer to the parameter diagram 3, the sensor receives the spherical wave modulated by the mask, and the spherical wave is converted into an optical signal on the sensor. Because the mask can modulate the spherical wave according to the depth of the spherical wave, that is, the spherical wave with different depths can exhibit different characteristics after passing through the mask, and the characteristics correspond to the depth. The sensor receives the spherical wave and converts the spherical wave with different characteristics into an electrical signal (so the electrical signal also contains the depth information of the spherical wave), and then generates a gray-scale image according to the electrical signal to encode the depth information into the pixels of the gray-scale image.
[0055] Step 013: based on the preset high-dimensional reconstruction model, the gray-scale image is reconstructed in high dimension to generate a high-dimensional image containing depth information.
[0056] The depth information can be a depth value, for example, please continue to refer to FIG. 3, the depth information can be the size of Z.
[0057] The preset high-dimensional reconstruction model can be a depth learning model, a neural network model, etc. The high-dimensional reconstruction model can reconstruct each pixel on the gray-scale image to obtain a reconstructed high-dimensional image, and can also decode the depth information included in each pixel to determine the depth information corresponding to each pixel.
[0058] The dimension range of the high-dimensional reconstruction can be three dimensions, four dimensions, or five dimensions, etc.
[0059] Specifically, taking the dimension range of the high-dimensional reconstruction as three dimensions as an example, by inputting the gray-scale image into the preset high-dimensional reconstruction model, the high-dimensional reconstruction model analyzes the depth information of each pixel on the gray-scale image to generate a high-dimensional image. The high-dimensional image not only includes two-dimensional image information but also includes third-dimensional depth information, which embodies the multi-dimensional data of the real scene (for example, the depth distance between a certain object and the user in the real scene), and realizes the targeted detection of the multi-dimensional data of the real scene.
[0060] The current RGB imaging camera (i.e. visible light imaging camera) is usually based on an optical-mechanical module (including a lens) to capture light in the scene to image, and the camera thickness is usually 10-20 millimeters (mm), and the weight of the optical-mechanical module is usually greater than 50 grams (g). However, in the actual application of the imaging method of the present application, no lens is needed, and imaging can be completed by capturing the spherical wave of a point light source (such as a sensor). In production, the thickness of the hardware can be controlled to be less than 2 mm, and the weight is less than or equal to 0.5 g. Therefore, in actual application, only imaging, the imaging method of the present application is more miniaturized and lightweight, and has less limitations in various fields of application.
[0061] In addition, compared with the current method of collecting image information and depth information by setting an RGB imaging camera and a depth camera, aligning the RGB imaging camera and the depth camera and simultaneously working, the imaging method of the present application generates a gray-scale image from the spherical wave modulated by the mask, encodes the depth information, and decodes the depth information by high-dimensional reconstruction of the gray-scale image based on a high-dimensional reconstruction model, thereby obtaining a high-dimensional image including depth information. The imaging method can integrate the functions of collecting image information and depth information, and does not need to additionally set a depth camera, thereby saving production and use costs.
[0062] In this way, the spherical wave emitted by the point light source is phase-modulated by the mask to change the propagation parameters of the spherical wave, wherein the propagation parameters include at least one of the phase and the amplitude, and the spherical wave of different depths has different parameters after being modulated by the mask. The scene to be photographed includes multiple point light sources, i.e. the depth information of the spherical wave is captured by the mask; the depth information is encoded based on the spherical wave with changed propagation parameters, and a gray-scale image is generated; the gray-scale image is high-dimensional reconstructed based on a preset high-dimensional reconstruction model, and the depth information corresponding to each pixel of the gray-scale image is decoded to generate a high-dimensional image including depth information. In other words, the high-dimensional image can not only display the scene content of the scene to be photographed, but also accurately reflect the depth information between the scene to be photographed and the objects in the scene to be photographed and the user, thereby realizing the detection of multi-dimensional information of the scene.
[0063] Compared with the current method of simultaneously setting an RGB imaging camera and a depth camera to collect image and depth information, the imaging method of the present application can integrate the functions of collecting image information and depth information, does not need to additionally set a depth camera, can be more miniaturized and lightweight, and can save production and use costs.
[0064] In the current imaging technology, if the time dimension information is to be captured, a time coding element (such as a digital micromirror device (DMD)) can be arranged in the camera module, but the cost of the DMD is high, resulting in high hardware cost when capturing the time dimension information.
[0065] In the case where the dimension range includes three dimensions, and the third dimension is the time dimension; or the dimension range includes four dimensions, the third dimension and the fourth dimension are the depth and time dimensions, the preset high-dimensional reconstruction model can also generate a high-dimensional image containing time information.
[0066] Referring to FIG. 4, in some embodiments, the imaging method further comprises:
[0067] Step 014: acquiring the acquisition time of each frame of grayscale image;
[0068] Step 013: performing high-dimensional reconstruction on the grayscale image based on the preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information, comprising:
[0069] Step 0131: performing high-dimensional reconstruction on the grayscale image based on the preset high-dimensional reconstruction model and the acquisition time of the grayscale image to generate a high-dimensional image containing depth information and time information.
[0070] Specifically, the acquisition time of each frame of grayscale image can be acquired, and the high-dimensional image containing depth information and time information can be obtained by decoding the grayscale image based on the acquisition time and the high-dimensional reconstruction model. For example, taking the case of modulating a spherical wave by a phase mask, collecting the spherical wave by a super high-speed sensor (for example, a super high-speed camera), and generating a grayscale image as an example, the super high-speed sensor can determine the acquisition time of the grayscale image, and the high-dimensional reconstruction model is used to perform high-dimensional reconstruction on the grayscale image. Since the acquisition time of each frame of grayscale image can be determined, the generated high-dimensional image can include depth information and time information, so that the high-dimensional image can not only exhibit the depth information of the scene and the objects in the scene, but also show the information in the time dimension (for example, reflecting the dynamic changes of the scene), further improving the detection capability of multi-dimensional information.
[0071] Referring to FIG. 5, optionally, step 012: generating a grayscale image based on the spherical wave with the changed propagation parameter, comprising:
[0072] Step 0121: based on a preset exposure function, controlling the image sensor to expose in time and in area within an image output period to obtain continuous multiple frames of grayscale images;
[0073] Step 0122: synthesizing and outputting the continuous multiple frames of grayscale images;
[0074] Step 014: acquiring the acquisition time of each frame of gray image, comprising:
[0075] Step 0141: determining the acquisition time of the plurality of frames of gray image based on the output time of the synthesized image.
[0076] Wherein, the preset exposure function can be a rolling exposure function, which can be used to represent the exposure area of the image sensor at each time point, and S(t|x, y) represents the preset exposure function, wherein t is the time point, x is the row index of the pixel of the gray image, and y is the column index of the pixel of the gray image.
[0077] Wherein, in the image output period, the area of the image sensor receiving the spherical wave is exposed once; or the image output period can also be the time for the image sensor to output an image.
[0078] Optionally, the image sensor comprises a direction from both sides to the center, and is divided into a plurality of sub-regions in sequence along the direction from both sides to the center, and each sub-region of the image sensor is exposed in sequence along the direction from both sides to the center.
[0079] Wherein, the area of the image sensor receiving light can be divided into a plurality of sub-regions in sequence along the direction from both sides (such as the upper and lower sides or the left and right sides) of the area to the center of the area, and the image sensor exposes each sub-region in sequence along the direction when dividing the area (such as rolling exposure) during exposure, for example, referring to FIG. 6, the area A is divided into sub-region a and sub-region b in sequence along the direction from both sides of the area A to the center of the area A, wherein sub-region a includes sub-region a1 and sub-region a2, and the image sensor can expose sub-region a at t1 and expose sub-region b at t2 during exposure, t1 and t2 are continuous, and the area A is exposed.
[0080] Specifically, the image sensor exposes different regions according to a preset exposure function S(t|x, y) when performing exposure. For example, by dividing the region of the image sensor receiving the spherical wave into multiple sub-zones and dividing one image output period into multiple time points, the image sensor exposes different sub-zones at different time points, i.e., each sub-zone receives the spherical wave at different time points; the image sensor collects the gray-scale images at each time point, synthesizes the continuous multiple frames of images in one image output period, and then outputs the synthesized image (for example, please refer to FIG. 6, taking one image output period including t1 moment and t2 moment as an example, by synthesizing the continuous two frames of images (i.e., the image at t1 moment and the image at t2 moment), a synthesized image is obtained, and the synthesized image is output at T moment), the synthesized image can include the light information at different time points (the time point corresponding to each frame of gray-scale image) in one time period, thus, based on the synthesized image and the output time of the synthesized image, the collection time of the multiple frames of gray-scale images can be determined. When the high-dimensional reconstruction model and the collection time of the gray-scale image are used to perform high-dimensional reconstruction on the gray-scale image, the time information of the gray-scale image can be obtained, and the cost is low.
[0081] In this way, by obtaining the collection time of each frame of gray-scale image, i.e., by encoding the collection time of the gray-scale image, the depth information and the time information of the image can be captured in the image output period (i.e., in the single complete exposure time) in which the image sensor exposes the region receiving the spherical wave once, the multi-dimensional information of the photographed scene is obtained, and when the imaging method is implemented, the volume and weight of the device are small, and the cost is low.
[0082] It can be understood that when the time information is captured, the external time encoding element can also be used to replace the rolling exposure.
[0083] When the multi-dimensional information is detected, it is also important to capture the information in the wavelength (color) dimension.
[0084] Please refer to FIG. 7, in some embodiments, the propagation parameters of the spherical waves of different depths and wavelengths are different after being modulated by the mask,
[0085] Step 013: based on a preset high-dimensional reconstruction model, performing high-dimensional reconstruction on the gray-scale image to generate a high-dimensional image containing depth information, including:
[0086] Step 0132: based on a preset high-dimensional reconstruction model, performing high-dimensional reconstruction on the gray-scale image to generate a high-dimensional image containing depth information and color information.
[0087] The color information can include the RGB information of the image.
[0088] Specifically, the imaging method of the present application is applied to a gray-scale sensor, and the mask includes a phase mask as an example. Since the phase mask has different wavefront modulation effects on the light of R, G and B channels, i.e., different phases during modulation, the spherical waves of R, G and B channels can be respectively named as φR(x, y), φG(x, y) and φB(x, y) corresponding to the modulated phases. After the spherical waves of R, G and B channels pass through the phase mask, the changed propagation parameters are different. Moreover, since the wavelengths of R, G and B lights are different, the different wavelengths of light will exhibit different gray-scale responses on the image sensor, i.e., the R, G and B lights will exhibit different gray-scale values on the image sensor after passing through the phase mask. Therefore, the gray-scale image can also be reconstructed in high dimensions based on a preset high-dimensional reconstruction model, and a high-dimensional image containing depth information and color information can be generated by decoding the modulation phase and the gray-scale value. Compared with the process of obtaining an RGB image by a traditional camera, the present application can replace the traditional camera with a gray-scale sensor, maintain the accuracy of the image color, greatly reduce the volume and weight of the camera, and reduce the cost of the traditional camera.
[0089] Referring to FIG. 8, in some embodiments, the method further comprises:
[0090] Step 015: obtaining a high-dimensional reconstruction model based on a preset training set, the training set including high-dimensional image samples containing at least one of depth information, color information and time information.
[0091] For example, in the case of training a high-dimensional reconstruction model that can output a high-dimensional image containing depth information, the training set used at least needs to include high-dimensional image samples containing depth information; for another example, in the case of training a high-dimensional reconstruction model that can output a high-dimensional image containing depth information and time information, the training set used at least needs to include high-dimensional image samples containing depth information and time information; for another example, in the case of training a high-dimensional reconstruction model that can output a high-dimensional image containing depth information, color information and time information, the training set used at least needs to include high-dimensional image samples containing depth information, color information and time information.
[0092] Specifically, the high-dimensional reconstruction model can be trained by obtaining a training set. The training set can include high-dimensional image samples. By inputting the high-dimensional image samples into the high-dimensional reconstruction model, the high-dimensional reconstruction model outputs training samples containing at least one of depth information, color information and time information, i.e., the high-dimensional reconstruction model is trained by the training set.
[0093] To obtain color information, when the high-dimensional reconstruction model is trained according to the high-dimensional image sample containing image information, the high-dimensional image sample, the spherical wave of the R channel, the G channel and the B channel of the same pixel can be respectively input into the high-dimensional reconstruction model, and the reconstruction results corresponding to the R channel, the G channel and the B channel are respectively added as the result of the pixel, so as to obtain the high-dimensional reconstruction result, and generate the high-dimensional image containing color information.
[0094] It should be noted that when the high-dimensional reconstruction model is used to perform high-dimensional reconstruction on the grayscale image to capture multi-dimensional information, the multi-dimensional information includes at least one of depth information, color information and time information, and when the high-dimensional reconstruction model is trained, the high-dimensional image sample containing the depth information, the color information and the time information can be used, so that the depth information, the color information and the time information can be obtained at the same time according to the grayscale image.
[0095] In other words, through the preset high-dimensional reconstruction model, the capture of information in five dimensions of two-dimensional information, depth information, time information and color information of the image can be realized at the same time.
[0096] Referring to FIG. 9, optionally, the high-dimensional reconstruction model includes an encoder and a decoder.
[0097] Step 015: obtaining the high-dimensional reconstruction model based on the preset training set, including:
[0098] Step 0151: encoding the high-dimensional image sample through the encoder to generate a grayscale image sample, the high-dimensional image sample including one or more continuous high-dimensional images;
[0099] Step 0152: performing high-dimensional reconstruction on the grayscale image sample through the decoder to generate a high-dimensional training image;
[0100] Step 0153: adjusting at least one of the high-dimensional reconstruction model and the mask based on the high-dimensional training image and the high-dimensional image sample until the high-dimensional reconstruction model converges.
[0101] Step 013: performing high-dimensional reconstruction on the grayscale image based on the preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information, including:
[0102] Step 0132: performing high-dimensional reconstruction on the grayscale image based on the converged decoder to generate a high-dimensional image containing depth information.
[0103] The high-dimensional reconstruction model includes an encoder and a decoder. The encoder can be a forward propagation network (Forward Net) based on a differentiable model. The encoder can encode multi-dimensional information (such as depth information and time information) to obtain a gray image containing encoded information. The decoder can be a network capable of inverse processing and reconstruction of the gray image to recover multi-dimensional information (such as depth information and time information). For example, the decoder can be a residual network (Rec Net) or a convolutional neural network (CNN). For another example, the architecture of the decoder can be a Transformer architecture (a neural network architecture based on a self-attention mechanism).
[0104] The high-dimensional image sample includes one or more continuous high-dimensional images. For example, the continuous high-dimensional images can be high-dimensional images of a same high-dimensional object at different time points in an image output period (including multiple continuous time sequences).
[0105] Specifically, referring to FIG. 10, the encoder includes a Forward Net, and the decoder includes a Rec Net. First, the encoder can encode the high-dimensional image sample (including one high-dimensional image or multiple high-dimensional images corresponding to different time points in an image output period T) according to the training set, to convert the high-dimensional image sample v(t, x, y, z) and depth information, time information, or color information contained in the high-dimensional image sample into a two-dimensional gray image L(x, y) sample containing the depth information, time information, or color information. The gray image L(x, y) sample contains corresponding depth information, time information, or color information. Then, the decoder can decode the corresponding depth information, time information, or color information (i.e., multi-dimensional information) contained in the gray image L(x, y) sample, and restore the gray image L(x, y) to the high-dimensional space to reconstruct a high-dimensional training image close to the high-dimensional sample image.
[0106] In the case of generating high-dimensional training images, at least one of the high-dimensional reconstruction model and the mask can be adjusted by a back propagation algorithm (for example, by setting a loss function, calculating the loss value between the high-dimensional sample image and the high-dimensional training image, adjusting the model parameters (such as the parameters of the decoder and the encoder) of the high-dimensional reconstruction model according to the loss value; for another example, by adjusting the height of the mask unit on the surface of the mask to adjust the modulation phase of the spherical wave by the mask, etc.) to optimize the model, so that the depth information, time information or color information contained in the reconstructed high-dimensional training image is more accurate, that is, the depth information, time information or color information contained in the high-dimensional training image is more similar to the high-dimensional sample image, so that the display effect of the high-dimensional training image in the high-dimensional space is closer to the high-dimensional sample image. By continuously adjusting the high-dimensional reconstruction model to convergence (both the encoder and the decoder converge), a preset high-dimensional reconstruction model is obtained. When the high-dimensional reconstruction model is applied, the gray image output high-dimensional reconstruction model, and the converged decoder can perform high-dimensional reconstruction on the gray image to generate a high-dimensional image containing depth information.
[0107] For example, a sub-wavelength phase super surface as a phase mask, the required optimal phase distribution φ1(x, y) can be determined according to the high-dimensional training image and the high-dimensional image sample, that is, in the case that the phase distribution of the phase mask is φ1(x, y), the output high-dimensional training image (and the contained depth information) and the high-dimensional image sample are the closest or even the same. In the simulation stage, the parameters of the phase mask can be adjusted by electromagnetic field simulation such as finite difference time domain (FDTD), rigorous coupled wave analysis (RCWA) and the like to obtain a suitable phase mask to construct a high-dimensional reconstruction model.
[0108] Referring to FIG. 11, optionally, the imaging method further comprises:
[0109] Step 016: establishing a first image model of the gray image and a second image model of the reconstructed high-dimensional image;
[0110] Step 017: based on the wave function of the spherical wave on the sensor, establishing a mapping model of the gray image and the reconstructed high-dimensional image; or,
[0111] Step 018: based on the wave function of the spherical wave on the sensor and the preset exposure function, establishing a mapping model of the gray image and the reconstructed high-dimensional image;
[0112] Step 019: based on the mapping model, constructing the encoder and the decoder.
[0113] Specifically, by the first image model, the received spherical wave is acquired to obtain a gray image; then by establishing a second image model, high-dimensional information of the gray image is recovered according to the gray image to reconstruct a high-dimensional image. According to the wave function of the spherical wave at the sensor (the wave function can include depth information, time information, etc. of the spherical wave at the sensor), the depth information of the spherical wave is captured according to the propagation parameter of the spherical wave, and a gray image including the depth information is obtained, and then a mapping model of the gray image and the reconstructed high-dimensional image is established to analyze the depth information of the gray image; or according to the wave function of the spherical wave at the sensor and a preset exposure function (a mapping relationship between an exposure time of the sensor receiving the spherical wave and an exposure area of the sensor receiving the spherical wave), the time information and the depth information of the spherical wave are captured, and a gray image including the time information and the depth information is obtained, and then a mapping model of the gray image and the reconstructed high-dimensional image is established to analyze the time information and the depth information of the gray image, so as to construct an encoder and a decoder according to the mapping model.
[0114] Since the wave function of the spherical wave to the sensor is related to the depth of the point light source (the distance from the point light source to the mask), that is, the light intensity distribution on the sensor is associated with the depth information of the point light source, the phase distribution of the spherical wave can be modulated by the mask to capture and encode the depth information, and the object parameters of the mask are optimized by reverse propagation to realize the reconstruction of the gray image based on the high-dimensional reconstruction model (to decode the depth information). The imaging method of the application has a more compact hardware structure, smaller volume and lower cost when applied.
[0115] Referring to FIG. 12, the imaging method further includes:
[0116] Step 019: determining a point spread function based on the wave function of the spherical wave at the sensor, the point spread function and the modulus of the wave function of the spherical wave at the sensor being in a positive correlation relationship;
[0117] Step 017: establishing a mapping model of the gray image and the reconstructed high-dimensional image based on the wave function of the spherical wave at the sensor, including:
[0118] Step 0171: establishing a mapping model of the gray image and the reconstructed high-dimensional image based on the point spread function;
[0119] Step 018: establishing a mapping model of the gray image and the reconstructed high-dimensional image based on the wave function of the spherical wave at the sensor and a preset exposure function, including:
[0120] Step 0181: establishing a mapping model of the gray image and the reconstructed high-dimensional image based on the point spread function and the preset exposure function.
[0121] Specifically, since a 3D object can be regarded as an object composed of an infinite number of point light sources, under Fresnel diffraction (i.e., diffraction of light waves in a near-field region), a point spread function (PSF) has translational invariance in a planar space and has a magnifying or reducing characteristic in a depth direction. Therefore, an imaging process of a point light source to a sensor can be modeled by establishing a convolution model, and the imaging process of the point light source to the sensor is modeling the PSF.
[0122] A point spread function PSF can represent how a point light source is imaged in a sensor. Referring again to FIG. 3, the point spread function is represented as PSF(x, y, z), x is a row index of a pixel of a gray-scale image, y is a column index of the pixel of the gray-scale image, z is a depth, and the mask is a phase mask. Since the point spread function and a modulus (matrix) of a wave function U of a spherical wave in the sensor are in a positive correlation, i.e., PSF(x, y, z) ∝ |U(x, y, z)| 2
[0123] Therefore, the point spread function PSF(x, y, z) can be determined according to the wave function U of the spherical wave in the sensor, and the point light source is mapped to the sensor according to the PSF(x, y, z) by convolution to obtain a gray-scale image L(x, y), and the convolution process can be represented as: L(x, y) = ∫0 ∞ v(x, y, z) * PSF(x, y, z) dz
[0124] wherein v(x, y, z) represents information of the spherical wave of the point light source at the position (x, y) and the depth z.
[0125] In the case of considering time information, the point light source is mapped to the sensor according to the PSF(x, y, z) by convolution, and the convolution process can be represented as:
[0126] wherein L(x, y, t) represents an encoding result (a gray-scale image including time information), and v(x, y, z, t) represents information of the spherical wave of the point light source at the position (x, y) and the depth z at the time t.
[0127] In other words, the above convolution process includes encoding the depth z first and then encoding the encoding result in a time domain. In the case of rolling exposure, the time t is encoded according to a preset exposure function S(t|x, y) to obtain the gray-scale image L(x, y), and the convolution process can be represented as:
[0128] Therefore, according to the point light source, the point spread function PSF(x, y, z), and the exposure function S(t|x, y), the convolution process of encoding the depth information and the time information into the gray-scale image can be represented as:
[0129] It can be understood that, when collecting the gray-scale image by the hyperspeed sensor, since the collection time can be determined, the convolution process of encoding the depth information and the time information into the gray-scale image can be expressed as: L(x, y) = ∫0 ∞ v(x, y, z, t) * PSF(x, y, z) dzdt
[0130] In this way, by the point spread function and the point light source, the gray-scale image of the point light source to the sensor can be established, and the depth information is contained in the point spread function and is encoded into the gray-scale image. Then, by analyzing the gray-scale image, the depth information can be obtained. The exposure function includes the time information, and therefore, when establishing the gray-scale image according to the point spread function and the exposure function, the depth information and the time information can be encoded. Similarly, by analyzing the gray-scale image, the depth information and the time information can be obtained, in other words, the mapping model of the gray-scale image and the reconstructed high-dimensional image can be established.
[0131] Referring to FIG. 13, optionally, the imaging method further comprises:
[0132] Step 020: determining a first wave function of the spherical wave at the incident surface of the mask based on the distance between the point light source and the incident surface of the mask;
[0133] Step 021: determining a second wave function of the spherical wave at the exit surface of the mask based on the phase distribution of the mask and the first wave function;
[0134] Step 022: obtaining a wave function of the spherical wave at the sensor based on the distance between the mask and the sensor, the wavelength of the point light source, and the second wave function.
[0135] Specifically, taking the mask as a phase mask as an example, according to the distance z between the point light source and the incident surface of the mask, the first wave function U1 of the spherical wave at the incident surface of the mask is determined under the Fresnel diffraction condition, that is:
[0136] Wherein, A is the amplitude of the spherical wave, i is an imaginary number, k is the wave number, and x and y represent the position of the spherical wave at the incident surface of the phase mask. ′ and y ′ represent the position of the spherical wave at the incident surface of the phase mask.
[0137] After the spherical wave passes through the phase mask, the phase will change, and the second wave function U2 of the spherical wave at the exit surface of the mask can be determined according to the phase distribution of the mask and the first wave function U1, that is:
[0138] Wherein, represents the phase modulation process of the phase mask to the spherical wave.
[0139] Finally, according to the phase mask and the distance d of the sensor, the wavelength λ of the point light source and the second wave function U2, according to the Fourier optics theory, the wave function U of the spherical wave at the sensor is obtained, that is:
[0140] where is the Fourier transform, fx is the spectral coordinate of x', and fy is the spectral coordinate of y'.
[0141] For another example, taking the mask as the phase mask, in the case of capturing the depth information, the time information and the phase information of the spherical wave, the wave function U of the spherical wave at the sensor can be:
[0142] At this time, the imaging result of the sensor also includes the integral of the response of the spherical wave of all wavebands, and therefore the convolution process of encoding the depth information, the time information and the wavelength information into the gray-scale image can be represented as:
[0143] where PSF(x, y, z, λ) is the point spread function of the spherical wave modulated by the mask. Similarly, the physical parameters of the mask can be optimized by back propagation, and the gray-scale image can be decoded by the high-dimensional reconstruction model to obtain the depth information, the time information and the wavelength information.
[0144] It can be understood that more information (such as amplitude information and polarization information) can be introduced when the spherical wave is modulated, that is, the spherical wave is modulated by using a composite mask (the composite mask can be composed of an amplitude mask and a phase mask, etc.), and the corresponding information is introduced when the high-dimensional reconstruction model is trained, so that the amplitude information and the polarization information can be obtained after the gray-scale image generated according to the modulated spherical wave is decoded.
[0145] For example, referring to FIG. 14, taking the imaging method of the application applied to the VR field and capturing the time information and the depth information as an example, the implementation process of the imaging method of the application in the VR head-mounted display can be as shown in FIG. 14. The spherical wave of the 3D object (which can be regarded as a point light source) is encoded by the mask to capture the depth information; then the sensor receives the spherical wave based on the rolling exposure, synthesizes the continuous multiple frames of images in the image output period to obtain a gray-scale image; and then the gray-scale image is input into the high-dimensional reconstruction model to perform high-dimensional reconstruction on the gray-scale image to generate a high-dimensional image containing the depth information and the time information. In the training process of the high-dimensional reconstruction model, the physical parameters of the mask can be optimized by the back propagation method to improve the accuracy of the obtained high-dimensional image.
[0146] Referring to FIG. 15, to better implement the imaging method of the embodiments of the present application, the embodiments of the present application further provide an imaging device 300. The imaging device 300 comprises a modulation module 301, a generation module 302, and a reconstruction module 303. The modulation module 301 is configured to modulate a spherical wave emitted by a point light source by a mask to change a propagation parameter of the spherical wave, the propagation parameter comprising at least one of a phase and an amplitude, and the spherical wave at different depths has different parameters after being modulated by the mask, and the photographed scene comprises a plurality of point light sources. The generation module 302 is configured to generate a gray-scale image based on the spherical wave with the changed propagation parameter. The reconstruction module 303 is configured to perform high-dimensional reconstruction on the gray-scale image based on a preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information.
[0147] In some embodiments, the imaging device 300 further comprises an acquisition module 304 configured to acquire an acquisition time of each frame of the gray-scale image. The reconstruction module 303 is specifically further configured to perform high-dimensional reconstruction on the gray-scale image based on the preset high-dimensional reconstruction model and the acquisition time of the gray-scale image to generate a high-dimensional image containing depth information and time information.
[0148] In some embodiments, the generation module 302 is specifically further configured to control the image sensor to expose in time and in area within one image output period based on a preset exposure function to obtain continuous multiple frames of the gray-scale image, and to synthesize and output the continuous multiple frames of the gray-scale image. The acquisition module 304 is specifically further configured to determine the acquisition time of the multiple frames of the gray-scale image based on an output time of the synthesized image.
[0149] In some embodiments, the propagation parameters of the spherical waves at different depths and wavelengths after being modulated by the mask are different. The reconstruction module 303 is specifically further configured to perform high-dimensional reconstruction on the gray-scale image based on the preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information and color information.
[0150] In some embodiments, the acquisition module 304 is specifically further configured to acquire the high-dimensional reconstruction model based on a preset training set, and the training set comprises high-dimensional image samples containing at least one of depth information, color information, and time information.
[0151] In some embodiments, the high-dimensional reconstruction model comprises an encoder and a decoder. The acquisition module 304 is specifically further configured to encode the high-dimensional image samples by the encoder to generate gray-scale image samples, the high-dimensional image samples comprising one or more continuous high-dimensional images; to perform high-dimensional reconstruction on the gray-scale image samples by the decoder to generate high-dimensional training images; to adjust at least one of the high-dimensional reconstruction model and the mask based on the high-dimensional training images and the high-dimensional image samples until the high-dimensional reconstruction model converges; and the reconstruction module 303 is specifically further configured to perform high-dimensional reconstruction on the gray-scale image based on the converged decoder to generate the high-dimensional image containing the depth information.
[0152] In some embodiments, the imaging apparatus 300 further comprises a constructing module 305 configured to establish a first image model of the grayscale image and a second image model of the reconstructed high-dimensional image; establish a mapping model of the grayscale image and the reconstructed high-dimensional image based on a wave function of the spherical wave at the sensor, or establish a mapping model of the grayscale image and the reconstructed high-dimensional image based on the wave function of the spherical wave at the sensor and a preset exposure function; and construct an encoder and a decoder based on the mapping model.
[0153] In some embodiments, the imaging apparatus 300 further comprises a determining module 306 configured to determine a point spread function based on the wave function of the spherical wave at the sensor, the point spread function and the wave function of the spherical wave at the sensor being in positive correlation; and the constructing module 305 is specifically configured to establish the mapping model of the grayscale image and the reconstructed high-dimensional image based on the point spread function; and establish the mapping model of the grayscale image and the reconstructed high-dimensional image based on the point spread function and the preset exposure function.
[0154] In some embodiments, the determining module 306 is further configured to determine a first wave function of the spherical wave at an incident surface of the mask based on a distance between the point light source and the incident surface of the mask; determine a second wave function of the spherical wave at an exit surface of the mask based on a phase distribution of the mask and the first wave function; and obtain the wave function of the spherical wave at the sensor based on a distance between the mask and the sensor, a wavelength of the point light source and the second wave function.
[0155] The imaging apparatus 300 is described above from the perspective of functional modules, which can be implemented in the form of hardware, in the form of instructions of software, or in the form of a combination of hardware and software modules. Specifically, each step of the method embodiments in the embodiments of the present application can be completed by integrated logic circuits of hardware in a processor and / or instructions in the form of software, and the steps of the method disclosed in the embodiments of the present application can be directly embodied as hardware coding processors to complete, or be completed by a combination of hardware and software modules in the coding processor. Alternatively, the software module can be located in a mature storage medium in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and combines the hardware thereof to complete the steps in the above method embodiments.
[0156] In some embodiments, referring to FIG. 1, which is a structural schematic diagram of a head-mounted display provided in embodiments of the present application. The head-mounted display 100 comprises a processor 70 and a memory 80. The memory 80 stores a computer program 81 executable on the processor 70, which, when executed by the processor 70, implements the processes of the embodiments of the display method described above and achieves the same technical effects. To avoid repetition, details are not described herein.
[0157] Referring to FIG. 16, the present application also provides a computer readable storage medium 600, which stores a computer program 610. When the computer program 610 is executed by a processor 620, the steps of the imaging method of any of the embodiments described above are implemented. To avoid repetition, details are not described herein.
[0158] In the description of the present application, the description of the terms "some embodiments", "in an example", "exemplarily", etc. means that the specific features, structures, materials or characteristics described in connection with the embodiments or examples are included in at least one embodiment or example of the present application. In the present application, the illustrative description of the above terms does not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, the person skilled in the art can combine and combine the different embodiments or examples described in the present application and the features of the different embodiments or examples without contradiction.
[0159] Any process or method descriptions in flow charts or described herein in other ways can be understood as representing code modules, segments, or portions of code that include one or more executable instructions for performing specific logic functions (or steps) in the process, and that the various embodiments of the application include the additional implementation of the described processes in which the order of execution of the executable instructions can be changed, additional or fewer processes are performed, or the functions of the described processes are performed in different sequences, as appropriate, and that the described processes can be performed in substantially simultaneous with other processes, or in a reverse order, depending upon the functionality involved.
[0160] Although the embodiments of the present application have been shown and described above, it should be understood that the above embodiments are exemplary and should not be construed as limiting the present application, and those of ordinary skill in the art can make changes, modifications, replacements and variations to the above embodiments within the scope of the present application.
Claims
1. An imaging method, wherein, The method comprises: Phase modulation is performed on a spherical wave emitted by a point light source through a mask to change a propagation parameter of the spherical wave, the propagation parameter comprising at least one of a phase and an amplitude, the spherical wave at different depths being modulated by the mask to have different parameters, and the photographed scene comprising a plurality of the point light sources; A grayscale image is generated based on the spherical wave with the changed propagation parameter; High-dimensional reconstruction is performed on the grayscale image based on a preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information.
2. The imaging method of claim 1, wherein, The method further comprises: Acquiring a capture time of each frame of the grayscale image; The high-dimensional reconstruction performed on the grayscale image based on the preset high-dimensional reconstruction model to generate the high-dimensional image containing the depth information comprises: The high-dimensional reconstruction performed on the grayscale image based on the preset high-dimensional reconstruction model and the capture time of the grayscale image to generate a high-dimensional image containing depth information and time information.
3. The imaging method of claim 1 or 2, wherein, The grayscale image generated based on the spherical wave with the changed propagation parameter comprises: In one image output period, a time-division and region-division exposure of an image sensor is controlled based on a preset exposure function to obtain a plurality of continuous frames of the grayscale image; The plurality of continuous frames of the grayscale image are synthesized and outputted; The acquisition of the capture time of each frame of the grayscale image comprises: The capture time of the plurality of frames of the grayscale image is determined based on an output time of a synthesized image.
4. The imaging method of claim 3, wherein, The method comprises: The image sensor comprises a plurality of sub-regions sequentially divided in a direction from both sides to a center, and each of the sub-regions is sequentially exposed in the direction from both sides to the center.
5. The imaging method of any one of claims 1-3, wherein, The propagation parameters of the spherical wave at different depths and wavelengths modulated by the mask are different, The high-dimensional reconstruction performed on the grayscale image based on the preset high-dimensional reconstruction model to generate the high-dimensional image containing the depth information comprises: The high-dimensional reconstruction performed on the grayscale image based on the preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information and color information.
6. The imaging method of any one of claims 1, 3, and 5, wherein, The method further comprises: The high-dimensional reconstruction model is acquired based on a preset training set, the training set comprising high-dimensional image samples containing at least one of depth information, color information and time information.
7. The imaging method of claim 6, wherein, The high-dimensional reconstruction model comprises an encoder and a decoder, and the acquisition of the high-dimensional reconstruction model based on the preset training set comprises: The high-dimensional image samples are encoded by the encoder to generate grayscale image samples, the high-dimensional image samples comprising one or more continuous high-dimensional images; The grayscale image samples are high-dimensionally reconstructed by the decoder to generate high-dimensional training images; At least one of the high-dimensional reconstruction model and the mask is adjusted based on the high-dimensional training images and the high-dimensional image samples until the high-dimensional reconstruction model converges; The high-dimensional reconstruction performed on the grayscale image based on the preset high-dimensional reconstruction model to generate the high-dimensional image containing the depth information comprises: The high-dimensional reconstruction performed on the grayscale image based on the converged decoder to generate the high-dimensional image containing the depth information.
8. The imaging method of claim 7, wherein, The method comprises: A first image model of the grayscale image and a second image model of the reconstructed high-dimensional image are established; establish a mapping model of the gray image and the high-dimensional image after reconstruction based on the wave function of the spherical wave at the sensor, or establish a mapping model of the gray image and the high-dimensional image after reconstruction based on the wave function of the spherical wave at the sensor and a preset exposure function; construct the encoder and the decoder based on the mapping model.
9. The imaging method of claim 8, wherein, Further comprising: determine a point spread function based on the wave function of the spherical wave at the sensor, the point spread function and a modulus of the wave function of the spherical wave at the sensor are in positive correlation; the establishing of the mapping model of the gray image and the high-dimensional image after reconstruction based on the wave function of the spherical wave at the sensor comprises: establish the mapping model of the gray image and the high-dimensional image after reconstruction based on the point spread function; the establishing of the mapping model of the gray image and the high-dimensional image after reconstruction based on the wave function of the spherical wave at the sensor and a preset exposure function comprises: establish the mapping model of the gray image and the high-dimensional image after reconstruction based on the point spread function and a preset exposure function.
10. The imaging method of claim 8 or 9, wherein, the method further comprises: determine a first wave function of the spherical wave at an incident surface of the mask based on a distance between the point light source and the incident surface of the mask; determine a second wave function of the spherical wave at an exit surface of the mask based on a phase distribution of the mask and the first wave function; obtain the wave function of the spherical wave at the sensor based on a distance between the mask and the sensor, a wavelength of the point light source and the second wave function.
11. The imaging method of claim 1, wherein, the mask comprises a phase mask, the phase mask comprises a plurality of mask units, and a phase distribution of the phase mask is determined based on heights of the mask units.
12. An imaging device, wherein, comprise: a modulation module configured to modulate a spherical wave emitted by a point light source by a mask to change a propagation parameter of the spherical wave, the propagation parameter comprising at least one of a phase and an amplitude, the parameters of the spherical wave at different depths being different after being modulated by the mask, and the photographed scene comprising a plurality of the point light sources; a generation module configured to generate a gray image based on the spherical wave with the changed propagation parameter; a reconstruction module configured to perform high-dimensional reconstruction on the gray image based on a preset high-dimensional reconstruction model to generate a high-dimensional image containing depth information.
13. A head-mounted display comprising a memory and a processor; the memory storing a computer program, wherein, The processor executes the program to implement the imaging method of any one of claims 1-11.
14. A non-transitory computer readable storage medium having stored thereon a computer program, wherein, The computer program is executed by the processor to implement the imaging method of any one of claims 1-11.
Citation Information
Patent Citations
Holographic AR three-dimensional display method and module and near-to-eye display system
CN113885209A
Three-dimensional imaging parameter adjustment method and device, electronic equipment and storage medium
CN117095123A
Hyperspectral and deep fusion imaging method, device and system based on multi-dimensional acquisition
CN117391979A
Imaging method, imaging device, head-mounted display and storage medium
CN119316581A
Optical Element for Deconvolution
US20230119549A1