Image preprocessing neural network model, training method, near-eye display method, device, equipment and medium
By combining color shift correction and pre-correction branches in the image preprocessing neural network model with optical characteristics, the color shift and aberration of near-eye display systems are corrected, solving the problems of visual clarity and color shift and improving the user experience.
Patent Information
- Application Number
- CN202511380055.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-25
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-09-25
AI Technical Summary
Existing near-eye display systems are limited by size when using complex optical mechanisms, making it difficult to improve visual clarity and reduce aberrations. At the same time, single-channel or single-network methods are prone to color shift in high-frequency edge scenarios, affecting user experience.
An image preprocessing neural network model is adopted, which includes a color shift correction branch and a pre-correction branch. Color shift is corrected by the brightness channel of the color space image, and pre-correction is performed by combining optical characteristics. The two correction results are fused by a fusion module to reduce color shift and improve visual effect.
Significantly reduces color shift in high-frequency edge scenarios, improves the user experience of near-eye display systems, ensures pre-correction effect and color consistency, and improves visual quality.
Smart Images

Figure CN120852207A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of near-eye display technology, and in particular to an image preprocessing neural network model, training method, near-eye display method, device, equipment and medium. Background Technology
[0002] Current near-eye display systems are often limited by size, making it impossible to use complex optical mechanisms to improve visual acuity and reduce aberrations.
[0003] Traditional methods employ real-time image pre-correction to correct aberrations. However, users encounter numerous scenarios with sharp edges when using near-eye display systems. Because these scenarios feature clear and sharp edges, ordinary image pre-correction algorithms are prone to color shifts when processing the RGB (Red, Green, Blue) channels of the displayed image. Such color shifts are easily detected by the human eye in sharp-edge scenarios, impacting the user experience. In other words, existing single-channel or single-network methods struggle to simultaneously address high-frequency edge distortion and optical correction effects.
[0004] In summary, how to ensure the pre-correction effect while reducing the degree of color cast is a technical problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0005] In view of this, the purpose of this application is to provide an image preprocessing neural network model, training method, near-eye display method, device, equipment and medium, which can both ensure the pre-correction effect and reduce the degree of color deviation.
[0006] To achieve the above objectives, this application provides the following technical solution: An image preprocessing neural network model includes: a color shift correction branch configured to correct the color shift of an image to be displayed based on the features of independent luminance channels in a color space image, to obtain a first corrected image; wherein the color space image is converted from the image to be displayed, and the color space image contains at least one independent luminance channel feature and at least one color channel feature; a pre-correction branch configured to pre-correct the image to be displayed based on optical properties, to obtain a second corrected image; and a fusion module configured to fuse the first corrected image and the second corrected image to obtain a fused image.
[0007] Optionally, both the image to be displayed and the first corrected image are RGB images. The image preprocessing neural network model further includes a conversion module, which is configured to convert the image to be displayed into a first color space image through inverse gamma transformation and matrix operations. The color shift correction branch is configured to adjust the characteristics of the luminance channel of the first color space image to increase the luminance contrast of the image to be displayed, thereby obtaining a first color space corrected image. The conversion module is also configured to convert the first color space corrected image into the first corrected image through matrix operations and gamma transformation.
[0008] Optionally, both the first color space image and the first color space corrected image are Yxy images.
[0009] Optionally, the fusion module includes: a channel connection unit configured to stitch the first corrected image and the second corrected image along the channel dimension to obtain a channel stitched feature map; a channel attention mechanism unit configured to perform channel weight allocation on the channel stitched feature map to obtain a first processing result; a spatial attention mechanism unit configured to perform spatial position weight allocation on the first processing result to obtain a second processing result; and a channel dimensionality reduction unit configured to perform channel dimensionality reduction processing on the second processing result to obtain the fused image.
[0010] A method for training an image preprocessing neural network model includes: acquiring a sample set, the sample set including multiple sample images; inputting any one of the sample images in the sample set into the image preprocessing neural network model described above to obtain a processed image; performing optomechanical imaging simulation on any one of the processed images to obtain a corresponding simulated image; constructing a loss function based on any one of the sample images in the sample set and the corresponding simulated image; and updating the parameters of the image preprocessing neural network model based on the loss function.
[0011] Optionally, the processed image includes the first corrected image, and the simulated image includes a color-shift correction simulated image, which is generated by optical-mechanical imaging simulation of the first corrected image. The step of constructing a loss function based on any sample image in the sample set and the corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, includes: converting the sample image into a second color space image, and converting the color-shift correction simulated image into a third color space image; constructing a color point shift loss function for the color-shift correction branch based on the color channel features of the second color space image and the color channel features of the third color space image; and updating the parameters of the color-shift correction branch based on the color point shift loss function of the color-shift correction branch.
[0012] Optionally, the processed image includes the fused image, and the simulated image includes the fused simulated image, which is generated by optical-mechanical imaging simulation of the fused image. The step of constructing a loss function based on any sample image in the sample set and the corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, includes: constructing an error loss function for the fusion module based on the sample image and the fused simulated image; converting the sample image into a second color space image and the fused simulated image into a fourth color space image; constructing a color point shift loss function for the fusion module based on the color channel features of the second color space image and the color channel features of the fourth color space image; and updating the parameters of the fusion module based on the error loss function and the color point shift loss function.
[0013] Optionally, updating the parameters of the fusion module based on the error loss function and the color point offset loss function of the fusion module includes: updating the parameters of the fusion module based on the error loss function and the color point offset loss function of the fusion module, as well as the weight ratio.
[0014] Optionally, the processed image includes the second corrected image, and the simulated image includes a pre-corrected simulated image, which is generated by optical-mechanical imaging simulation of the second corrected image. The step of constructing a loss function based on any sample image in the sample set and the corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, includes: constructing an error loss function of the pre-corrected branch based on the sample image and the pre-corrected simulated image; and updating the parameters of the pre-corrected branch based on the error loss function of the pre-corrected branch.
[0015] Optionally, performing optomechanical imaging simulation on the processed image includes: convolving the processed image using a pre-calibrated point spread function, or inputting the processed image into a pre-trained display network to obtain a corresponding simulated image.
[0016] A near-eye display method includes: inputting an image to be displayed into an image preprocessing neural network model as described in any of the above claims to obtain the fused image; and displaying the fused image on a display.
[0017] A training apparatus for an image preprocessing neural network model includes: an acquisition module for acquiring a sample set, the sample set including multiple sample images; an input module for inputting any one sample image from the sample set into the image preprocessing neural network model as described above to obtain a processed image; an imaging simulation module for performing optomechanical imaging simulation on any one of the processed images to obtain a corresponding simulated image; and a construction and update module for constructing a loss function based on any one sample image from the sample set and the corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function.
[0018] An electronic device includes: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of a training method for an image preprocessing neural network model as described above, or to implement the steps of a near-eye display method as described above.
[0019] A readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a training method for an image preprocessing neural network model as described above, or the steps of a near-eye display method as described above.
[0020] This application provides an image preprocessing neural network model, training method, near-eye display method, apparatus, device, and medium. The image preprocessing neural network model includes: a color shift correction branch configured to correct the color shift of the image to be displayed based on the features of independent luminance channels in a color space image, to obtain a first corrected image; wherein the color space image is converted from the image to be displayed, and the color space image contains at least one independent luminance channel feature and at least one color channel feature; a pre-correction branch configured to pre-correct the image to be displayed based on optical characteristics, to obtain a second corrected image; and a fusion module configured to fuse the first corrected image and the second corrected image to obtain a fused image.
[0021] The technical solution disclosed in this application includes an image preprocessing neural network model comprising a color shift correction branch, a pre-correction branch, and a fusion module connected to the color shift correction branch and the pre-correction branch. The color shift correction branch is configured to correct the color shift of the image to be displayed based on the characteristics of the luminance channel in the color space image converted from the image to be displayed. This corrects high-frequency characteristic color shift phenomena and avoids cross-coupling errors caused by direct processing of the RGB three channels by performing color shift correction on the luminance channel of the color space image, thereby significantly reducing color shift in sharp edge scenes. The pre-correction branch is configured to pre-correct the image to be displayed based on optical characteristics. This achieves an image pre-correction function corresponding to the optical characteristics, thereby offsetting or compensating for optical characteristic defects in the optomechanical system of the near-eye display system. The fusion module is configured to fuse the first corrected image obtained from the color shift correction branch and the second corrected image obtained from the pre-correction branch to obtain a fused image that both maintains the pre-correction effect and reduces the degree of color shift. This makes the image entering the human eye in the near-eye display system closer to the image to be displayed in the near-eye display system, so that even in scenes with sharp edges, the human eye can hardly detect the color shift, thereby improving the user experience of the near-eye display system. In other words, compared with existing single-channel or single-network methods that are difficult to balance high-frequency edge and optical correction effects, this application ensures the consistency of pre-correction effect and color through a dual-branch architecture and fusion module, significantly improving the user's subjective experience.
[0022] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0023] Figure 1 This is a schematic diagram of the structure of an image pre-correction neural network model provided in an embodiment of this application; Figure 2 A flowchart of image preprocessing provided in this application embodiment; Figure 3 A schematic diagram of a multi-scale codec cascade structure provided in an embodiment of this application; Figure 4 A schematic diagram of another image preprocessing neural network model provided in this application embodiment; Figure 5 Another flowchart of image preprocessing provided in this application embodiment; Figure 6 This application provides another flowchart of image preprocessing for display. Figure 7 This is a schematic diagram of the structure of another image preprocessing neural network model provided in an embodiment of this application; Figure 8This is a schematic diagram of the structure of another image preprocessing neural network model provided in an embodiment of this application; Figure 9 This is a schematic diagram of another image preprocessing neural network provided in an embodiment of this application; Figure 10 A flowchart illustrating a training method for an image preprocessing neural network model provided in this application embodiment; Figure 11 A schematic diagram illustrating the construction of a loss function during the training of an image preprocessing neural network model, provided in an embodiment of this application; Figure 12 A flowchart illustrating another method for training an image preprocessing neural network model provided in this application embodiment; Figure 13 A flowchart illustrating another method for training an image preprocessing neural network model provided in an embodiment of this application; Figure 14 This is an overall data roadmap for training the image preprocessing neural network model provided in the embodiments of this application; Figure 15 A flowchart illustrating another method for training an image preprocessing neural network model provided in an embodiment of this application; Figure 16 A flowchart illustrating the forward inference of the image to be displayed, provided in an embodiment of this application. Figure 17 The diagram illustrates the image to be displayed provided in the embodiments of this application, the image obtained by correcting the image to be displayed using an existing pre-correction algorithm, and the image obtained by correcting the image using an image preprocessing neural network model. Detailed Implementation
[0024] With the introduction of low-cost, readily available virtual reality (VR) and augmented reality (AR) hardware and systems, technologies such as VR and AR are rapidly changing the way we work, interact, and socialize. Near-eye display systems are among the most widely used applications of VR and AR. Taking VR as an example, HMD (Head-Mounted Display) is one of the more well-known VR devices, providing users with a vivid and immersive visual experience through high-performance head tracking technology and real-time rendering based on a GPU (Graphics Processing Unit).
[0025] While near-eye display systems offer many advantages, their production process requires a balance between image quality and device size and manufacturing costs. Current near-eye display systems are limited by size, preventing the design of overly complex optical structures. This makes it difficult to use multi-lens elements to improve visual clarity and reduce aberrations. As a result, the visual quality of traditional near-eye display systems often falls short of optimal standards. Users may see images with significant geometric distortion, spherical aberration, and field curvature defects. These aberrations and imperfections can lead to visual fatigue and diminish the immersive experience of VR and AR applications.
[0026] Currently, real-time image pre-correction technology (i.e., purely software-based) is used to correct the aforementioned aberrations, allowing users to comfortably enjoy the experience without wearing glasses. However, users encounter many scenes with sharp edges when using near-eye display systems, such as text scenes. Because the characteristic edges in such scenes are clear and sharp, ordinary image pre-correction algorithms can easily cause subtle color shifts in the processing of the RGB three-color channels of the displayed image. These color shifts are not easily perceived by the human eye in ordinary scene images, but in scenes with sharp edges, the human eye is more sensitive to the slight color shifts caused by the output of the image pre-correction algorithm. In other words, in scenes with sharp edges, slight color shifts are easily detected by the human eye, affecting the user experience. That is, existing single-channel or single-network methods struggle to simultaneously address high-frequency edge issues and optical correction effects.
[0027] To address this, this application provides an image preprocessing neural network model, training method, near-eye display method, device, and medium. The image preprocessing neural network model includes a color shift correction branch, a pre-correction branch, and a fusion module connected to the color shift correction branch and the pre-correction branch. The color shift correction branch corrects the color shift of the image to be displayed based on the characteristics of the luminance channel in the color space image converted from the image to be displayed. This corrects high-frequency characteristic color shift phenomena and avoids cross-coupling errors caused by direct processing of the RGB three channels by performing color shift correction in the luminance channel of the color space image, thereby displaying and retrieving color shifts even in sharp edge scenes. The pre-correction branch pre-corrects the image to be displayed based on optical characteristics, achieving image pre-correction functions corresponding to the optical characteristics, i.e., correcting aberrations such as geometric distortion, spherical aberration, and field curvature. The fusion module merges the first corrected image obtained from the color shift correction branch and the second corrected image obtained from the pre-correction branch to obtain an image that both ensures the pre-correction effect and reduces the degree of color shift. This makes the image displayed by the near-eye display system closer to the image to be displayed, thereby improving the user's experience of the near-eye display system. In other words, this application ensures the consistency of pre-correction effect and color through a dual-branch architecture and fusion module, and greatly improves the user's subjective experience.
[0028] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0029] See Figure 1 and Figure 2 ,in, Figure 1 This is a schematic diagram of the structure of an image pre-correction neural network model provided in an embodiment of this application. Figure 2 This application provides a flowchart of a preprocessing process for an image to be displayed, as illustrated in an embodiment of this application. The image preprocessing neural network model provided in this application may include: A color shift correction branch is configured to correct the color shift of the image to be displayed based on the features of the independent luminance channels in the color space image, to obtain a first corrected image; wherein the color space image is converted from the image to be displayed, and the color space image contains at least one independent luminance channel feature and at least one color channel feature; The pre-correction branch is configured to pre-correct the image to be displayed based on optical characteristics to obtain a second corrected image; The fusion module is configured to fuse the first corrected image and the second corrected image to obtain a fused image.
[0030] The image preprocessing neural network model provided in this application embodiment may include a color shift correction branch, a pre-correction branch, and a fusion module. The fusion module can be connected to the color shift correction branch and the pre-correction branch respectively. The color shift correction branch is responsible for correcting high-frequency characteristic color shift phenomena, the pre-correction branch is responsible for implementing the pre-correction function, and the fusion module is responsible for fusing the output results of the color shift correction branch and the pre-correction branch to obtain an output that both reduces the degree of color shift and ensures the pre-correction effect, thereby improving the user experience of the near-eye display system.
[0031] The color shift correction branch can correct the color shift of the image to be displayed based on the characteristics of the independent luminance channels in the color space image converted from the image to be displayed, thus obtaining a first corrected image. That is, the color shift correction branch can adjust the characteristics of the independent luminance channels in the color space image converted from the image to be displayed to correct the color shift of the image to be displayed. The image to be displayed is a rendered image (i.e., an ideal digital image) generated or acquired by the near-eye display system. After preprocessing, this image is sent to the display and presented to the user's eye through the optical engine (i.e., the optical system; the optical engine in the near-eye display system can be of various types, such as: Pancake (folded optical path), aspherical, spherical, Fresnel, etc.). For example, the image to be displayed can be an RGB image, or it can be an RGBA image (which adds an alpha channel to RGB to represent transparency), etc. A color space image is converted from an image to be displayed. A color space image contains at least one independent luminance channel feature and at least one chrominance channel feature. By separating luminance and chrominance in a color space image, the color shift correction branch can adjust the features of the independent luminance channel in the color space image, thereby correcting the color shift of the image to be displayed. For example, a color space image can be a Yxy image, where the Y channel represents luminance (i.e., the Y channel independently represents the brightness of a color, directly related to the human eye's perception of light and dark, and is the only luminance information carrier in the color space; that is, the Y component in Yxy clearly separates luminance information), and the x and y channels are collectively called chrominance coordinates (i.e., the x and y components in Yxy jointly define chrominance information, that is, jointly define the color itself, excluding luminance), used to describe the hue and saturation of the color. The x and y channels are completely independent of the Y channel. Alternatively, a Lab image (Y is the luminance channel, a and b are the chrominance channels) or a YUV image (Y is the luminance channel, U and V are the chrominance channels), etc., are images containing at least one independent luminance channel feature and at least one chrominance channel feature. By converting the image to be displayed into a color space image with separate luminance and chrominance channels, color deviation correction of the image to be displayed can be performed based on the characteristics of the independent luminance channel in the color space image, thus avoiding the cross-coupling error caused by direct processing of the three RGB channels (because the red, green and blue components of RGB jointly determine the brightness and chrominance of the color, so changing any one component will simultaneously affect the brightness and chrominance of the color).
[0032] The pre-correction branch can pre-correct the image to be displayed based on the optical characteristics of the optomechanical system in the near-eye display system to obtain a second corrected image. That is, the pre-correction branch can achieve image pre-correction functions corresponding to optical characteristics to counteract or compensate for optical defects in the optomechanical system of the near-eye display system (such as geometric distortion, spherical aberration, and field curvature). The pre-correction algorithm used in the pre-correction branch can be a common Wiener filter, LR (Low-Rank Filtering) filter, or various iterative solution algorithms, or it can be a neural network structure such as CNN (Convolutional Neural Networks) or Transformer (a deep learning model network structure based on an attention mechanism). Furthermore, the main architecture of the pre-correction branch can be a multi-scale encoder-decoder cascade structure. The input of this branch is the image to be displayed, which is responsible for pre-correcting the image to address optical blurring in subsequent optical paths. The output of this branch is the corrected image (i.e., the second corrected image). The multi-scale codec cascade structure refers to cascading multiple codec modules, each processing features at different scales. Each module can contain multiple convolutional layers for encoding and multiple deconvolutional layers for decoding. It should be noted that there are many basic codec modules, such as CNNs, Transformers, and other neural network structures. For example, see [link to example]. Figure 3 This is a schematic diagram of a multi-scale codec cascade structure provided in an embodiment of this application. It should be noted that... Figure 3 Taking a 3×288×288 image as an example, the input 3×288×288 image is first downsampled to obtain a 3×144×144 image. Then, it is processed by multiple codec modules to process features at different scales. After that, it is upsampled to obtain a 3×288×288 image. The upsampled 3×288×288 image is added to the input 3×288×288 image (or fused through other methods). The resulting image is then processed by multiple codec modules to process features at different scales, finally obtaining a second corrected 3×288×288 image. In the codec module, conv2d is a convolutional layer used to compress the image in the spatial dimension and expand it in the channel dimension, and up_conv is a reverse convolutional layer used to expand the image in the spatial dimension and compress it in the channel dimension.
[0033] The fusion module fuses the first corrected image obtained from the color shift correction branch and the second corrected image obtained from the pre-correction branch to obtain a fused image. By fusing the first corrected image from the color shift correction branch and the second corrected image from the pre-correction branch, the color-shifted image and the pre-corrected image based on optical characteristics can be merged to obtain a fused image that reduces the degree of color shift while ensuring the pre-correction effect. This ensures the consistency between the pre-correction effect and the color, thereby improving the user's subjective experience. The fused image output by the fusion module (which is also the output of the image preprocessing neural network model) can be input to the display of the near-eye display system for display. The image displayed on the display can be directly perceived by the human eye after passing through the optomechanical system of the near-eye display system.
[0034] As described above, the preprocessing flow for the image to be displayed can be as follows: After the near-eye display system generates or acquires the rendered image to be displayed (i.e., an ideal digital image), it can input the image to be displayed into an image preprocessing neural network model. The image preprocessing neural network model receives the image to be displayed as input and converts it into a color space image with separated luminance and chrominance (i.e., containing at least one independent luminance channel feature and at least one color). Alternatively, after the near-eye display system generates or acquires the rendered image to be displayed, it can convert the image to be displayed into a color space image. The image to be displayed and the color space image are then input into the image preprocessing neural network model. The image preprocessing neural network model receives the image to be displayed as input and converts it into a color space image (i.e., containing at least one independent luminance channel feature and at least one color). The display image and color space image are used as inputs; the color space image converted from the image to be displayed is input into the color shift correction branch, so that the color shift correction branch corrects the color shift of the image to be displayed based on the characteristics of the independent luminance channels in the color space image to obtain a first corrected image; the image to be displayed is input into the pre-correction branch, so that the pre-correction branch pre-corrects the image to be displayed based on optical characteristics to obtain a second corrected image; the first corrected image obtained by the color shift correction branch and the second corrected image obtained by the pre-correction branch are input into the fusion module, and the fusion module fuses the first corrected image and the second corrected image to obtain a fused image that can reduce the degree of color shift and ensure the pre-correction effect.
[0035] As can be seen from the above, the embodiments of this application can solve the color shift problem of high-frequency details in the pre-correction algorithm of the current near-eye display system during correction, and at the same time, the pre-correction effect is guaranteed by the dual-branch fusion method.
[0036] The technical solution disclosed in this application includes an image preprocessing neural network model comprising a color shift correction branch, a pre-correction branch, and a fusion module connected to the color shift correction branch and the pre-correction branch. The color shift correction branch is configured to correct the color shift of the image to be displayed based on the characteristics of the luminance channel in the color space converted from the image to be displayed. This corrects high-frequency characteristic color shift phenomena and avoids cross-coupling errors caused by direct processing of the RGB three channels by performing color shift correction on the luminance channel of the color space image, thereby significantly reducing color shift in sharp-edge scenes. The pre-correction branch is configured to pre-correct the image to be displayed based on optical characteristics. This achieves an image pre-correction function corresponding to the optical characteristics, thereby offsetting or compensating for optical characteristic defects in the optomechanical system of the near-eye display system. The fusion module is configured to fuse the first corrected image obtained from the color shift correction branch and the second corrected image obtained from the pre-correction branch to obtain a fused image that both maintains the pre-correction effect and reduces the degree of color shift. This makes the image entering the human eye in the near-eye display system closer to the image to be displayed in the near-eye display system, so that even in scenes with sharp edges, the human eye can hardly detect the color shift, thereby improving the user experience of the near-eye display system. In other words, the embodiments of this application ensure the consistency of pre-correction effect and color through a dual-branch architecture and a fusion module, significantly improving the user experience.
[0037] See Figure 4 and Figure 5 ,in, Figure 4 This is a schematic diagram of another image preprocessing neural network model provided in an embodiment of this application. Figure 5 This application provides another flowchart for image preprocessing in an embodiment of the present application. The embodiment of the present application provides an image preprocessing neural network model where both the image to be displayed and the first corrected image are RGB images. The image preprocessing neural network model may further include a conversion module, configured to convert the image to be displayed into a first color space image through inverse gamma transformation and matrix operations. The color shift correction branch is configured to adjust the characteristics of the luminance channel of the first color space image to increase the luminance contrast of the image to be displayed, thereby obtaining a first color space corrected image. The conversion module is also configured to convert the first color space corrected image into a first corrected image through matrix operations and gamma transformation.
[0038] In this embodiment, considering that RGB images are the mainstream and commonly used image format in near-eye display systems, both the image to be displayed and the first calibration image can be RGB images to better adapt to the display, reduce driving complexity, and provide excellent color reproduction capabilities, matching human vision and offering a superior visual experience. Furthermore, the second calibration image can also be an RGB image; that is, the pre-calibration branch can pre-calibrate the RGB format image to be displayed based on optical characteristics to obtain the RGB format second calibration image. Correspondingly, the fusion module can fuse the RGB format first calibration image and the RGB format second calibration image to obtain a fused RGB format image.
[0039] Furthermore, the image preprocessing neural network model may also include a conversion module, which can convert the RGB format image to be displayed into a first color space image through inverse gamma transformation and matrix operations. Specifically, the conversion module can first convert the RGB format image to be displayed into a linear domain RGB image through inverse gamma transformation, and then convert the linear domain RGB image to be displayed into a first color space image through matrix operations. The first color space image contains at least one independent luminance channel feature and at least one color channel feature. The first color space image serves as the input to the color shift correction branch (specifically, the features of the independent luminance channels in the first color space image serve as the input to the color shift correction branch). For example, the first color space image can be a Yxy image, etc.
[0040] The color shift correction branch specifically adjusts the features of independent luminance channels in the first color space image to increase the luminance contrast of the image to be displayed, thus obtaining a first color space corrected image. In other words, the color shift correction branch can increase the luminance contrast of the image to be displayed by adjusting the features of independent luminance channels in the first color space image, thereby correcting the color shift of the image. The main architecture of the color shift correction branch can be a multi-scale codec cascade structure. The input is the features of independent luminance channels in the first color space image, and the output is the processed luminance channel features. The processed luminance channel features are combined with the features of the color channels in the first color space image to obtain the first color space corrected image. It should be noted that the multi-scale codec cascade structure in the color shift correction branch can be the same as the multi-scale codec cascade structure in the pre-correction branch, such as... Figure 3 The difference lies in that, for the color shift correction branch, its input is a 1×288×288 image (where one dimension represents the characteristics of an independent luminance channel in the first color space image), and the final result is also a 1×288×288 image. Of course, the multi-scale codec cascade structure in the color shift correction branch can be different from the multi-scale codec cascade structure in the pre-correction branch, and this application does not limit this.
[0041] Based on the above, the conversion module in the image preprocessing neural network model can also convert the first color space corrected image obtained by the color shift correction branch into a first corrected image through matrix operations and gamma transformation. Specifically, the conversion module can first convert the first color space corrected image into a linear domain RGB corrected image through matrix operations, and then perform a gamma transformation on the linear domain RGB corrected image to obtain a gamma domain RGB corrected image (i.e., the first corrected image).
[0042] See Figure 6 This is another flowchart of image preprocessing provided in this application embodiment. The image preprocessing neural network model provided in this application embodiment uses a first color space image and a first color space corrected image, both of which are Yxy images.
[0043] In the embodiments of this application, both the first color space image and the first color space corrected image can be Yxy images. The Y channel represents luminance and is the only luminance information carrier in the color space. The x and y channels are collectively called chromaticity coordinates. That is, the x and y components in Yxy jointly define the color itself, excluding luminance, and are used to describe the hue and saturation of the color. The x and y channels are completely independent of the Y channel.
[0044] Based on the above, the conversion module converts the image to be displayed into an image in the first color space through inverse gamma transformation and matrix operations, which can be specifically described as follows: (11) The conversion module first transforms the RGB format image to be displayed into a linear domain RGB image to be displayed through inverse gamma transformation; (12) Use the matrix operation of Formula 1 below to convert the RGB image to be displayed in the linear domain obtained in step (11) to the XYZ domain: (Formula 1) Wherein, R in Formula 1 linear G linear B linear These represent the values of the red, green, and blue channels in the RGB image to be displayed in the linear domain, respectively. In Formula 1, X, Y, and Z represent the X (red primary color stimulus), Y (green primary color stimulus), and Z (blue primary color stimulus) values of the RGB image to be displayed after conversion to the XYZ domain, respectively. It should be noted that the X, Y, and Z values represent tristimulus values, which are quantitative indicators describing the degree of stimulation of the three primary colors by which the human retina perceives color.
[0045] (13) Use the following formula 2 to convert the XYZ domain obtained in step (12) to the Yxy domain to obtain the first color space image in Yxy format: (Formula 2) In Formula 2, Y, x, and y in Yxy represent the Y, x, and y values in the first color space image of Yxy format, respectively. The Y value represents the brightness of the color, and the x and y values represent the chromaticity.
[0046] The conversion module transforms the first color space corrected image into the first corrected image through matrix operations and gamma transformation. Specifically, this process can be described as follows: (21) The conversion module first converts the first color space corrected image in Yxy format to the XYZ domain according to the following formula 3: (Formula 3) Among them, in formula 3 , , These represent the red, green, and blue primary color stimuli levels, respectively, after the first color space corrected image in Yxy format is converted to the XYZ domain.
[0047] (22) Use the matrix operation of Formula 4 below to convert the XYZ domain obtained in step (21) into a linear domain RGB corrected image: (Formula 4) Among them, in formula 4 , , These represent the values of the red, green, and blue channels in the linear domain RGB-corrected image, respectively.
[0048] (23) Perform a gamma transform on the RGB corrected image in the linear domain to obtain the RGB corrected image in the gamma domain (i.e., the first corrected image in RGB format).
[0049] When the image to be displayed, the first corrected image, and the second corrected image are all RGB images, and the first color space image and the first color space corrected image are both Yxy images, the preprocessing flow of the image to be displayed is as follows: Figure 6 As shown, R 0 G 0 B 0 R represents the image to be displayed. 1 G 1 B 1 R represents the first corrected image. 2 G 2 B 2 Y represents the second corrected image, Yxy represents the first color space image, Y 1 This represents the adjusted Y-channel feature obtained by adjusting the Y-channel characteristics in Yxy using the color shift correction branch. 1xy represents the first color space calibrated image.
[0050] See Figure 7 and Figure 8 ,in, Figure 7 This is a schematic diagram of the structure of another image preprocessing neural network model provided in an embodiment of this application. Figure 8 This is a schematic diagram of another image preprocessing neural network model provided in an embodiment of this application. The image preprocessing neural network model provided in this application includes a fusion module that may include: The channel connection unit is configured to stitch the first corrected image and the second corrected image along the channel dimension to obtain a channel stitching feature map; The channel attention mechanism unit is configured to perform channel weight allocation on the channel splicing feature map to obtain the first processing result; The spatial attention mechanism unit is configured to assign spatial location weights to the first processing result to obtain the second processing result. The channel dimensionality reduction unit is configured to perform channel dimensionality reduction processing on the second processing result to obtain a fused image.
[0051] In this embodiment, the fusion module in the image preprocessing neural network model can mainly consist of a channel attention mechanism and a spatial attention mechanism module. It performs weight redistribution and fusion on the two inputs (i.e., the first corrected image and the second corrected image) to obtain the final output result (i.e., the fused image). For details, please refer to... Figure 9 This is a schematic diagram of another image preprocessing neural network provided in the embodiments of this application. Specifically, the fusion module can dynamically adjust the weight distribution of the two inputs and effectively fuse them, thereby generating a fused image with better pre-correction and less color cast. Specifically, the fusion module may include a channel connection unit, a channel attention mechanism unit, a spatial attention mechanism unit, and a channel dimensionality reduction unit.
[0052] The channel connection unit is configured to stitch the first corrected image and the second corrected image along the channel dimension to obtain a channel stitching feature map. For example, if both the first corrected image and the second corrected image are 3×288×288 images, the channel connection unit can stitch the first corrected image and the second corrected image into a 6×288×288 image.
[0053] The channel attention mechanism unit is configured to redistribute channel weights on the channel stitching feature map to obtain the first processing result. Channel weight redistribution refers to automatically learning the importance of channels in the first and second corrected images and dynamically adjusting the weights to amplify channels that are relevant to pre-correction and minimize color cast, suppress background noise or redundant channels, and enhance sensitivity to key features. In other words, channel weight redistribution dynamically learns the importance of different feature channels, scales channel features by generating weight coefficients, thereby increasing the contribution of important channels and suppressing irrelevant or noisy channels, so that the fusion module focuses on key channels that are relevant to pre-correction and minimize color cast.
[0054] The spatial attention mechanism unit is configured to assign spatial location weights to the first processing result to obtain the second processing result. Here, spatial location weight assignment refers to automatically learning the importance of the spatial locations of the first and second corrected images and dynamically adjusting the weights to amplify spatial regions related to pre-correction effects and minimize color cast, suppress background noise or redundant channels, and enhance sensitivity to key features. In other words, spatial location weight redistribution refers to dynamically learning the importance of each spatial location in the feature map, generating a two-dimensional weight map to weight the spatial locations, thereby strengthening key region features and suppressing irrelevant background regions.
[0055] The channel dimensionality reduction unit is configured to perform channel dimensionality reduction on the second processing result to obtain a fused image. The channel dimensionality reduction unit can be a 1×1 convolutional layer, which linearly combines channels along the channel dimension to compress the high-channel number of the second processing result into a low-channel number while maintaining the spatial resolution. For example, the channel dimensionality reduction unit can compress 6 dimensions in a 6×288×288 image into 3 dimensions, while keeping 288×288 unchanged, to obtain a 3×288×288 fused image.
[0056] This application also provides a method for training an image preprocessing neural network model, see [link to relevant documentation]. Figure 10 The flowchart of a training method for an image preprocessing neural network model provided in this application embodiment may include: S101: Obtain a sample set, which may include multiple sample images.
[0057] It should be noted that the execution entity of the image preprocessing neural network model training method provided in this application embodiment can be a GPU server, cloud computing platform, or other device independent of the near-eye display system. After these devices have trained the relevant parameters of the image preprocessing neural network model, they can integrate the image preprocessing neural network model and its relevant parameters into the near-eye display system. This allows the near-eye display system to preprocess the image to be displayed using the integrated image preprocessing neural network model. In other words, the image preprocessing neural network model can be trained offline to shorten the training cycle, improve training efficiency, and reduce the demand for computing resources and cost of the near-eye display system. Of course, the image preprocessing neural network model can also be trained online, or it can be trained by combining offline training of the base model with online fine-tuning and optimization. This application embodiment does not limit the training method of the image preprocessing neural network model.
[0058] When training an image preprocessing neural network model, a sample set can first be obtained, which contains multiple sample images. These sample images can all be images generated and rendered by a near-eye display system and are to be displayed. The format of the sample images can be the same as the format of the images to be displayed.
[0059] S102: Input any sample image from the sample set into any of the above image preprocessing neural network models to obtain the processed image.
[0060] Input any sample image from the sample set into any of the above image preprocessing neural network models to obtain the processed image.
[0061] It should be noted that the relevant description of the image preprocessing neural network model can be found in the description of the relevant part of the image preprocessing neural network model provided in the embodiments of this application, and will not be repeated here.
[0062] S103: Perform optomechanical imaging simulation on any processed image to obtain the corresponding simulated image.
[0063] To improve the accuracy of parameter determination for the image preprocessing neural network model, thereby enhancing its pre-correction and color cast correction effects on the displayed image, after obtaining the processed image using the model, an optomechanical imaging simulation can be performed on any processed image to obtain a corresponding simulated image. This simulation simulates the optomechanical imaging results in a near-eye display system during the training of the image preprocessing neural network model. This facilitates the construction of a loss function based on the simulated and sample images, and the updating of the model's parameters based on the constructed loss function. This completes the training loop, improving the accuracy of parameter determination and enhancing the pre-correction and color cast correction effects. The result is an image that maintains good pre-correction while exhibiting near-zero color cast, ensuring the image entering the human eye closely resembles the displayed image and improving the user experience.
[0064] S104: Construct a loss function based on any sample image in the sample set and its corresponding simulated image, and update the parameters of the image preprocessing neural network model based on the loss function.
[0065] Based on the above, a loss function can be constructed using any sample image in the sample set and its corresponding simulated image. This loss function helps the image preprocessing neural network model to be trained better. Then, backpropagation can be performed based on the constructed loss function to optimize and update the parameters of the image preprocessing neural network model. After updating the parameters of the image preprocessing neural network model according to the loss function, the process can return to step S102, i.e., steps S102 and S104 can be executed until the training termination condition is met. The training termination condition mentioned here may include reaching a training iteration threshold, convergence of performance indicators such as loss or accuracy, or the loss function reaching a threshold. Training can be terminated as long as any one of these conditions is met.
[0066] Training the image preprocessing neural network model as described above can improve its reliability and accuracy, thereby enhancing its preprocessing effect on the image to be displayed in the near-eye display system. This results in a fused image that guarantees both pre-correction and virtually eliminates color cast, ensuring that the image entering the human eye in the near-eye display system is almost identical to the image to be displayed. This makes color cast virtually imperceptible even in scenes with sharp edges, thus improving the user experience of the near-eye display system. In other words, this embodiment solves the color cast problem in high-frequency details during correction in current pre-correction algorithms for near-eye display systems, while ensuring pre-correction effectiveness through a dual-branch fusion method.
[0067] See Figure 11 and Figure 12 ,in, Figure 11 This is a schematic diagram illustrating the construction of a loss function during the training of an image preprocessing neural network model, as provided in an embodiment of this application. Figure 12 This is a flowchart illustrating another training method for an image preprocessing neural network model provided in this application embodiment. The image preprocessing neural network model training method provided in this application embodiment includes a processed image that may include a first corrected image, and a simulated image that may include a color-shift correction simulated image. The color-shift correction simulated image is generated by optical-mechanical imaging simulation from the first corrected image. A loss function is constructed based on any sample image in the sample set and the corresponding simulated image. The parameters of the image preprocessing neural network model are updated based on the loss function, and the method may include: The sample image is converted into a second color space image, and the color-shifted correction simulation image is converted into a third color space image. The color point offset loss function of the color offset correction branch is constructed based on the color channel features of the second color space image and the color channel features of the third color space image. Update the parameters of the color shift correction branch based on the color point offset loss function of the color shift correction branch.
[0068] In this embodiment, the processed image may include a first corrected image, which can be the first corrected image obtained from the color shift correction branch in the image preprocessing neural network model. Correspondingly, the simulated image may include a color shift correction simulated image, which is generated by optical-mechanical imaging simulation of the first corrected image. The format of the color shift correction simulated image is the same as that of the first corrected image, and can also be the same as that of the sample image, for example, both can be in RGB format.
[0069] Based on the above, the process of constructing a loss function based on any sample image in the sample set and its corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, can specifically include: Step S201: Convert the sample image into a second color space image, and convert the color-shifted correction simulation image into a third color space image. The second and third color space images have the same format, each containing at least one independent luminance channel feature and at least one color channel feature. Furthermore, the second and third color space images have the same format as the first color space image described above; for example, both can be Yxy images.
[0070] Step S202: Construct the color point shift loss function of the color shift correction branch based on the color channel features of the second color space image and the third color space image. The type of the color point shift loss function of the color shift correction branch can be MSE (Mean Square Error), MAE (Mean Absolute Error), Euclidean distance, or cosine similarity loss, etc. Alternatively, the type of the color point shift loss function of the color shift correction branch can also be RMSE (Root Mean Square Error) or MRE (Mean Relative Error). The specific color point shift loss function of the color shift correction branch can be set according to actual needs.
[0071] Step S203: Update the parameters of the color shift correction branch according to the color point offset loss function of the color shift correction branch constructed in step S203.
[0072] By constructing the color channel features of the second color space image converted from the sample image and the color channel features of the third color space image converted from the color-shift correction simulation image, the color point offset loss function of the color shift correction branch can be constructed. The reliability and accuracy of the parameters of the color shift correction branch can be improved by updating the parameters of the color shift correction branch based on the loss function, thereby improving the color shift correction effect of the color shift correction branch.
[0073] As described above, in this embodiment of the application, color shift is constrained simultaneously by changes in color points during the training of the luminance channel, thereby effectively controlling the color shift phenomenon. That is, this embodiment of the application uses a dual-branch neural network to constrain the color of the input image and perform image pre-correction on the input image, respectively, and fuses the branch outputs of both to obtain an output that ensures both no color shift and the effectiveness of the pre-correction algorithm.
[0074] This application provides a method for training an image preprocessing neural network model. The processed image may include a fused image, and the simulated image may include a fused simulated image. The fused simulated image is generated by performing optomechanical imaging simulation on the fused image. A loss function is constructed based on any sample image in the sample set and the corresponding simulated image. The parameters of the image preprocessing neural network model are updated based on the loss function. This method may include: An error loss function for the fusion module is constructed based on sample images and fused simulated images; The sample image is converted into a second color space image, and the fused simulated image is converted into a fourth color space image; The color point offset loss function of the fusion module is constructed based on the color channel features of the second color space image and the color channel features of the fourth color space image. The parameters of the fusion module are updated based on the error loss function and color point offset loss function of the fusion module.
[0075] In this embodiment, the processed image may include a fused image, that is, a fused image obtained by the fusion module in the image preprocessing neural network model. Correspondingly, the simulated image may include a fused simulated image, which is generated by performing optomechanical imaging simulation on the fused image. The format of the fused simulated image is the same as that of the fused image, and may also be the same as that of the sample image, for example, both may be in RGB format.
[0076] Based on the above, the process of constructing a loss function based on any sample image in the sample set and its corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, can specifically include: Step S301: Construct the error loss function of the fusion module based on the sample image and the fused simulated image. The error loss function of the fusion module can be of type MSE, MAE, RMSE, or MRE, etc., and the specific type can be set according to actual needs.
[0077] Step S302: Convert the sample image into a second color space image and convert the fused simulated image into a fourth color space image. The second and fourth color space images have the same format, each containing at least one independent luminance channel feature and at least one color channel feature. Furthermore, the second and fourth color space images have the same format as the first color space image described above; for example, they can both be Yxy images.
[0078] Step 303: Construct the color point shift loss function for the fusion module based on the color channel features of the second color space image and the fifth color space image obtained in step S302. The type of the color point shift loss function for the fusion module can be MSE, Euclidean distance, or cosine similarity loss, or it can be RMSE or MRE. The specific type of the color point shift loss function for the fusion module can be set according to actual needs.
[0079] It should be noted that there is no restriction on the execution order between steps S301 and steps 302-303.
[0080] Step 304: Update the parameters of the fusion module according to the error loss function of the fusion module constructed in step S301 and the color point offset loss function of the fusion module constructed in steps 302-303.
[0081] By constructing an error loss function for the fusion module based on the sample images and the fused simulation images, and incorporating this error loss function into the parameter updates of the fusion module, the pre-correction effect can be guaranteed. Similarly, by constructing a color point offset loss function for the fusion module based on the color channel features of the second color space image converted from the sample images and the fourth color space image converted from the fused simulation images, and incorporating this color point offset loss function into the fusion module updates, color cast correction can be achieved, reducing the degree of color cast and ensuring that color cast in high-frequency details is not severe. In other words, by constructing the error loss function and the color point offset loss function for the fusion module, and updating the fusion module based on these two loss functions, the fusion module can ensure both the pre-correction effect and that color cast in high-frequency details is not severe (i.e., guarantee the color cast correction effect) during fusion.
[0082] This application provides a training method for an image preprocessing neural network model, which updates the parameters of the fusion module based on the error loss function and color point offset loss function of the fusion module, and may include: The parameters of the fusion module are updated based on the error loss function and color point offset loss function of the fusion module, as well as the weight ratio.
[0083] In this embodiment, when updating the parameters of the fusion module based on the error loss function and the color point offset loss function of the fusion module, in order to ensure that the fusion module can guarantee both the pre-correction effect and the color shift of high-frequency details during fusion, different weight ratios can be assigned to the error loss function and the color point offset loss function of the fusion module. The parameters of the fusion module are updated based on the error loss function and the color point offset loss function of the fusion module and the weight ratios (the weight ratio α of the error loss function of the fusion module, the weight ratio β of the color point offset loss function of the fusion module, α+β=1).
[0084] Specifically, the sum of the error loss and color point offset loss of the fusion module can be calculated according to different weights. That is, the loss function of the fusion module can be calculated by weighted summation based on the error loss function and its weight ratio α of the fusion module, and the color point offset loss function and its weight ratio β of the fusion module. That is: Loss function of fusion module = α × Error loss function of fusion module + β × Color point offset loss function of fusion module, where α + β = 1. The values of α and β are set in advance through experiments or experience. Taking the error loss function of the fusion module as MES as an example, Loss = α × MES + β × ΔColor, where Loss is the loss function of the fusion module, and ΔColor is the color point offset loss function of the fusion module. Based on parameter search, the weight ratio α of the MES item can be given as 0.75, and the weight ratio β of the color point offset loss function can be given as 0.75.
[0085] This application provides a method for training an image preprocessing neural network model. The processed image may include a second corrected image, and the simulated image may include a pre-corrected simulated image. The pre-corrected simulated image is generated by optomechanical imaging simulation from the second corrected image. A loss function is constructed based on any sample image in the sample set and the corresponding simulated image. The parameters of the image preprocessing neural network model are updated based on the loss function. This method may include: The error loss function of the pre-calibration branch is constructed based on the sample image and the pre-calibrated simulated image; The parameters of the pre-corrected branch are updated based on the error loss function of the pre-corrected branch.
[0086] In this embodiment, the processed image may include a second corrected image, which can be the second corrected image obtained from the pre-correction branch in the image preprocessing neural network model. Correspondingly, the simulated image may include a pre-corrected simulated image, which is generated by performing optomechanical imaging simulation on the second corrected image. The format of the pre-corrected simulated image is the same as that of the second corrected image, and can also be the same as that of the sample image, for example, both can be in RGB format.
[0087] Based on the above, the process of constructing a loss function based on any sample image in the sample set and its corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, can specifically include: Step S401: Construct the error loss function of the pre-calibration branch based on the sample image and the pre-calibration simulated image. The error loss function of the pre-calibration branch can be MSE, MAE, RMSE, or MRE, etc. (the specific type of error loss function of the pre-calibration branch can be set according to actual needs), and edge ringing can be constrained by TV (Total variation) regularization.
[0088] Step S402: Update the parameters of the pre-correction branch according to the error loss function of the pre-correction branch.
[0089] By constructing the error loss function of the pre-calibration branch based on the sample image and the pre-calibration simulation image, and updating the pre-calibration branch based on the error loss function, the reliability and accuracy of the parameters of the pre-calibration branch can be improved, thereby improving the pre-calibration effect of the pre-calibration branch.
[0090] For example, see Figure 13 This is a flowchart illustrating another method for training an image preprocessing neural network model provided in this application embodiment, wherein... Figure 13Taking the sample image, first corrected image, color cast correction simulation image, second corrected image, pre-corrected simulation image, fused image, and fused simulation image as examples where all are RGB images, and the second, third, and fourth color space images are Yxy images, as an example, R0G0B0 represents the sample image, Y1x1y1 represents the second color space image, Y'1 represents the adjusted Y channel features obtained by adjusting the Y1 channel features in the second color space image through the color cast correction branch, R1G1B1 represents the first corrected image, R2G2B2 represents the color cast correction simulation image, and Y2x2y2 represents... The third color space image, ΔColor1 represents the color point offset loss function of the color shift correction branch, R3G3B3 represents the second corrected image, R4G4B4 represents the pre-corrected simulated image, R5G5B5 represents the fused image, R6G6B6 represents the fused simulated image, ELoss1 represents the error loss function of the pre-correction branch, Y3x3y3 represents the fourth color space image, ELoss2 represents the error loss function of the fusion module, ΔColor2 represents the color point offset loss function of the fusion module, and Loss represents the loss function of the fusion module. Loss = α × ELoss2 + β × ΔColor2. Figure 13 It can be seen that R0G0B0 is converted to Y1x1y1. The color shift correction branch in the image preprocessing neural network model adjusts Y1 in Y1x1y1 to Y'1 to increase the brightness and contrast of the sample image. Y'1 and x1y1 are combined to form Y'1x1y1 (the second color space corrected image). Y'1x1y1 is reconstructed to R1G1B1. The optomechanical imaging simulation of R1G1B1 is R2G2B2. R2G2B2 is converted to Y2x2y2. ΔColor1 is constructed using x1 and y1 in Y1x1y1 and x2 and y2 in Y2x2y2. The parameters of the color shift correction branch are updated using ΔColor1. R0G0B0 is processed by the pre-correction branch in the image preprocessing neural network model to obtain R3G3B3. The optomechanical imaging simulation of R3G3B3 is R4G4B4. R0G0B0 and R4G4B4 construct ELoss1. The parameters of the pre-correction branch are updated using ELoss1. The fusion module in the image preprocessing neural network model fuses R1G1B1 and R3G3B3 into R5G5B5. The optomechanical imaging simulation of R5G5B5 is R6G6B6. R6G6B6 is converted into Y3x3y3. ELoss2 is constructed using R0G0B0 and R6G6B6. ΔColor2 is constructed using x1 and y1 in Y1x1y1 and x3 and y3 in Y3x3y3. Loss = α × ELoss2 + β × ΔColor2. The parameters of the fusion module are updated using Loss.
[0091] This application provides a method for training an image preprocessing neural network model, which performs optomechanical imaging simulation on the processed image, and may include: The processed image is convolved using a pre-calibrated point spread function, or the processed image is input into a pre-trained display network to obtain the corresponding simulated image.
[0092] In this embodiment, optomechanical imaging simulation of the processed image can be performed in the following two ways: 1. Convolve the processed image using a pre-calibrated point spread function to obtain the corresponding simulated image; that is, optomechanical imaging simulation can be achieved by forward convolution of the processed image using the point spread function. 2. Pre-train a display network to represent the input and output process of the optomechanical system. Input the processed image into the pre-trained display network to obtain the simulated image corresponding to the processed image.
[0093] The above method enables accurate simulation of optomechanical imaging, thereby improving the accuracy of the trained image preprocessing neural network model.
[0094] To further explain the training process of the image preprocessing neural network model, please refer to [link / reference]. Figure 14 and Figure 15 ,in, Figure 14 This is the overall data roadmap for training the image preprocessing neural network model provided in the embodiments of this application. Figure 15 A flowchart illustrating another training method for an image preprocessing neural network model provided in this application embodiment. The overall training process data transmission route of the image preprocessing neural network model can be... Figure 14 The image preprocessing neural network model is composed of various modules. The main training process provided in this embodiment is as follows: 1. The computational control module controls the entire training process, sending rendered sample images to the image preprocessing neural network model. 2. The image preprocessing neural network model receives the sample images as input and inputs them into the color shift correction branch and the pre-correction branch, respectively. The input to the color shift correction branch is the feature of the independent luminance channel in the color space after the sample image is converted to a color space image, and the input to the pre-correction branch is the sample image. The output of the image preprocessing neural network model is a fused image, representing the final pre-correction result. 3. For the training process, an imaging simulation module is constructed to simulate the process of a display image passing through an optical engine and entering the human eye. 4. The computational control module controls the output of the imaging simulation module and the sample images to solve for the loss function. Based on the obtained loss function, backpropagation is performed to optimize the parameters of the image preprocessing neural network model.
[0095] In this system, the computational control module is responsible for controlling the startup of all other modules and their corresponding digital computation stages. It also manages the overall rendering pipeline of the near-eye display system and is typically a highly integrated chip. Specifically, the computational control module is responsible for the forward inference calculations of the image preprocessing neural network model and the imaging simulation module, the backpropagation calculations based on the loss function results, and the rendering process of the sample images and the images to be displayed.
[0096] For a detailed description of the image preprocessing neural network model, please refer to the corresponding section above, and it will not be repeated here.
[0097] The loss function calculation module is used to construct the loss function for the image preprocessing neural network model. Specifically, to better constrain the performance of the color shift correction branch, the pre-correction branch, and the fusion module, separate loss functions are designed for the outputs of each branch. For more details, please refer to [link to relevant documentation]. Figure 15 The detailed explanations of the corresponding parts mentioned above will not be repeated here.
[0098] The imaging simulation module helps the image preprocessing neural network model simulate optomechanical imaging results during training, constructing a complete training loop. Specifically, the imaging simulation module can perform optomechanical imaging simulation using a pre-calibrated point spread function or a pre-trained display network.
[0099] This application also provides a near-eye display method, which may include: Input the image to be displayed into any of the above image preprocessing neural network models to obtain the fused image; The merged image is displayed on the monitor.
[0100] This application also provides a near-eye display method, which inputs the image to be displayed into any of the aforementioned image preprocessing neural network models (i.e., image preprocessing neural network models trained by the training method) to obtain a fused image that guarantees pre-correction and has almost no color cast. After obtaining the fused image, it can be displayed on the monitor in the near-eye display system. Subsequently, the image displayed on the monitor can pass through the optical engine of the near-eye display system. If the user is wearing the near-eye display system, the image can enter the user's eye after passing through the optical engine, so that the image entering the user's eye is almost identical to the image to be displayed by the near-eye display system. This makes it so that even in scenes with sharp edges, the human eye can hardly detect color cast, thereby improving the user's experience with the near-eye display system.
[0101] For example, see Figure 16 and Figure 17 ,in, Figure 16 This is a flowchart illustrating the forward inference of the image to be displayed, provided in an embodiment of this application. Figure 17The diagram illustrates the image to be displayed provided in the embodiments of this application, the image obtained by correcting the image to be displayed using an existing pre-correction algorithm, and the image obtained by correcting the image using an image preprocessing neural network model. Figure 16 The overall flow of the forward inference process is demonstrated. The computational control module controls the input of the image to be displayed into the image preprocessing neural network model. The fused image output by the image preprocessing neural network model is input into the display for display (the display module is responsible for illuminating the fused image on the display). The image displayed on the display is directly perceived by the human eye after passing through the optical engine inside the near-eye display system. During the forward inference process, the image preprocessing neural network model is loaded with trained parameters. Figure 17 In the image, (a) is an example image to be displayed, (b) is an example image obtained by correcting the image to be displayed in (a) using an existing pre-correction algorithm, and (c) is an example image obtained by correcting the image to be displayed in (a) using an image preprocessing neural network model provided in some embodiments of this application. From (a) to (c), it can be seen that the image obtained by correcting the image using a simple RGB channel pre-correction algorithm has obvious color shift, while the image obtained by correcting the image using an image preprocessing neural network model has significantly improved color shift and is very close to the image to be displayed.
[0102] This application embodiment also provides a training device for an image preprocessing neural network model, which may include: an acquisition module for acquiring a sample set, which may include multiple sample images; an input module for inputting any one sample image from the sample set into any of the above-mentioned image preprocessing neural network models to obtain a processed image; an imaging simulation module for performing optomechanical imaging simulation on any one of the processed images to obtain a corresponding simulated image; and a construction and update module for constructing a loss function based on any one sample image from the sample set and the corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function.
[0103] This application provides a training device for an image preprocessing neural network model. The processed image may include a first corrected image, and the simulated image may include a color-shift correction simulated image. The color-shift correction simulated image is generated by optical-mechanical imaging simulation of the first corrected image. If the image preprocessing neural network model does not include a conversion module, the construction and update module may include: a first conversion submodule, used to convert the sample image into a second color space image and convert the color-shift correction simulated image into a third color space image; a first construction submodule, used to construct a color point offset loss function for the color-shift correction branch based on the color channel features of the second color space image and the color channel features of the third color space image; and a first update submodule, used to update the parameters of the color-shift correction branch based on the color point offset loss function of the color-shift correction branch. If the image preprocessing neural network model includes a conversion module, the conversion module can be used to convert the sample image into a second color space image and the color-shift correction simulation image into a third color space image. The construction and update module can include: a first construction submodule, used to construct the color point offset loss function of the color shift correction branch based on the color channel features of the second color space image and the color channel features of the third color space image; and a first update submodule, used to update the parameters of the color shift correction branch based on the color point offset loss function of the color shift correction branch.
[0104] This application provides a training device for an image preprocessing neural network model. The processed image may include a fused image, and the simulated image may include a fused simulated image. The fused simulated image is generated by optical-mechanical imaging simulation of the fused image. If the image preprocessing neural network model does not include a conversion module, the construction and update module may include: a second construction submodule, used to construct an error loss function for the fusion module based on the sample image and the fused simulated image; a second conversion submodule, used to convert the sample image into a second color space image and the fused simulated image into a fourth color space image; a third construction submodule, used to construct a color point offset loss function for the fusion module based on the color channel features of the second color space image and the color channel features of the fourth color space image; and a second update submodule, used to update the parameters of the fusion module based on the error loss function and the color point offset loss function of the fusion module. If the image preprocessing neural network model includes a transformation module, the transformation module can be used to transform the sample image into a second color space image and the fused simulation image into a fourth color space image. The construction and update module can include: a second construction submodule, used to construct the error loss function of the fusion module based on the sample image and the fused simulation image; a third construction submodule, used to construct the color point offset loss function of the fusion module based on the color channel features of the second color space image and the color channel features of the fourth color space image; and a second update submodule, used to update the parameters of the fusion module based on the error loss function and the color point offset loss function of the fusion module.
[0105] This application provides a training device for an image preprocessing neural network model. The second update submodule may include an update unit for updating the parameters of the fusion module based on the error loss function and color point offset loss function of the fusion module, as well as the weight ratio.
[0106] This application provides a training device for an image preprocessing neural network model. The processed image may include a second corrected image, and the simulated image may include a pre-corrected simulated image. The pre-corrected simulated image is generated by optical-mechanical imaging simulation of the second corrected image. The construction and update module may include: a fourth construction submodule, used to construct the error loss function of the pre-corrected branch based on the sample image and the pre-corrected simulated image; and a third update submodule, used to update the parameters of the pre-corrected branch according to the error loss function of the pre-corrected branch.
[0107] This application provides a training device for an image preprocessing neural network model. The imaging simulation module may include an imaging simulation unit, which is used to convolve the processed image using a pre-calibrated point spread function, or to input the processed image into a pre-trained display network to obtain a corresponding simulated image.
[0108] For a description of the relevant parts of the training device for an image preprocessing neural network model provided in this application embodiment, please refer to the detailed description of the corresponding parts in the training method for an image preprocessing neural network model provided in this application embodiment, and will not be repeated here.
[0109] This application also provides an electronic device, which may include: Memory, used to store computer programs; The processor is used to execute computer programs to implement the steps of training any of the above-described image preprocessing neural network models or to implement the steps of the above-described near-eye display method.
[0110] This application also provides a readable storage medium storing a computer program. When the computer program is executed by a processor, it implements the steps of the training method for any of the above-described image preprocessing neural network models or the steps of the above-described near-eye display method.
[0111] For a description of the relevant parts of the electronic device and readable storage medium provided in the embodiments of this application, please refer to the detailed description of the corresponding parts in the training method of the image preprocessing neural network model and the near-eye display method provided in the embodiments of this application, which will not be repeated here.
[0112] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0113] It should be understood that various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0114] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0115] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0116] In this application, unless otherwise expressly specified and limited, the terms "installation," "connection," "joining," and "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise expressly limited. Those skilled in the art can understand the specific meaning of the above terms in this application according to the specific circumstances.
[0117] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. An image preprocessing neural network model, characterized in that, include: A color shift correction branch is configured to correct the color shift of the image to be displayed based on the features of independent luminance channels in a color space image, to obtain a first corrected image; wherein the color space image is converted from the image to be displayed, and the color space image contains at least one independent luminance channel feature and at least one color channel feature; The pre-correction branch is configured to pre-correct the image to be displayed based on optical characteristics to obtain a second corrected image; The fusion module is configured to fuse the first corrected image and the second corrected image to obtain a fused image.
2. The image preprocessing neural network model according to claim 1, characterized in that, Both the image to be displayed and the first corrected image are RGB images. The image preprocessing neural network model further includes a conversion module, which is configured to convert the image to be displayed into a first color space image through inverse gamma transformation and matrix operations. The color shift correction branch is configured to adjust the characteristics of the luminance channel of the first color space image to increase the luminance contrast of the image to be displayed, thereby obtaining a first color space corrected image. The conversion module is also configured to convert the first color space corrected image into the first corrected image through matrix operations and gamma transformation.
3. The image preprocessing neural network model according to claim 2, characterized in that, Both the first color space image and the first color space corrected image are Yxy images.
4. The image preprocessing neural network model according to any one of claims 1 to 3, characterized in that, The fusion module includes: The channel connection unit is configured to stitch the first corrected image and the second corrected image in the channel dimension to obtain a channel stitching feature map; The channel attention mechanism unit is configured to perform channel weight allocation on the channel splicing feature map to obtain a first processing result; The spatial attention mechanism unit is configured to assign spatial location weights to the first processing result to obtain the second processing result; The channel dimensionality reduction unit is configured to perform channel dimensionality reduction processing on the second processing result to obtain the fused image.
5. A training method for an image preprocessing neural network model, characterized in that, include: Obtain a sample set, which includes multiple sample images; Input any sample image from the sample set into the image preprocessing neural network model as described in any one of claims 1 to 4 to obtain the processed image; Perform an optomechanical imaging simulation on any of the processed images to obtain the corresponding simulated image; A loss function is constructed based on any sample image in the sample set and the corresponding simulated image, and the parameters of the image preprocessing neural network model are updated based on the loss function.
6. The training method for the image preprocessing neural network model according to claim 5, characterized in that, The processed image includes the first corrected image, and the simulated image includes a color-shifted corrected simulated image, which is generated by optical-mechanical imaging simulation of the first corrected image. The step of constructing a loss function based on any sample image in the sample set and the corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, includes: The sample image is converted into a second color space image, and the color shift correction simulation image is converted into a third color space image; The color point offset loss function of the color offset correction branch is constructed based on the color channel features of the second color space image and the color channel features of the third color space image. The parameters of the color shift correction branch are updated based on the color point offset loss function of the color shift correction branch.
7. The training method for the image preprocessing neural network model according to claim 5 or 6, characterized in that, The processed image includes the fused image, and the simulated image includes the fused simulated image. The fused simulated image is generated by performing optomechanical imaging simulation on the fused image. The step of constructing a loss function based on any sample image in the sample set and the corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, includes: The error loss function of the fusion module is constructed based on the sample images and the fused simulation images; The sample image is converted into a second color space image, and the fused simulated image is converted into a fourth color space image; The color point offset loss function of the fusion module is constructed based on the color channel features of the second color space image and the color channel features of the fourth color space image. The parameters of the fusion module are updated based on the error loss function and color point offset loss function of the fusion module.
8. The training method for the image preprocessing neural network model according to claim 7, characterized in that, The step of updating the parameters of the fusion module based on the error loss function and the color point offset loss function of the fusion module includes: The parameters of the fusion module are updated based on the error loss function and color point offset loss function of the fusion module, as well as the weight ratio.
9. The training method for the image preprocessing neural network model according to claim 7, characterized in that, The processed image includes the second corrected image, and the simulated image includes a pre-corrected simulated image, which is generated by optical-mechanical imaging simulation of the second corrected image. The step of constructing a loss function based on any sample image in the sample set and the corresponding simulated image, and updating the parameters of the image preprocessing neural network model based on the loss function, includes: The error loss function of the pre-corrected branch is constructed based on the sample image and the pre-corrected simulated image; The parameters of the pre-correction branch are updated based on the error loss function of the pre-correction branch.
10. The training method for the image preprocessing neural network model according to claim 5 or 6, characterized in that, Performing optomechanical imaging simulation on the processed image includes: The processed image is convolved using a pre-calibrated point spread function, or the processed image is input into a pre-trained display network to obtain a corresponding simulated image.
11. A near-eye display method, characterized in that, include: The image to be displayed is input into the image preprocessing neural network model as described in any one of claims 1 to 4 to obtain the fused image; The fused image is displayed on a monitor.
12. A training device for an image preprocessing neural network model, characterized in that, include: The acquisition module is used to acquire a sample set, which includes multiple sample images; An input module is used to input any one of the sample images in the sample set into the image preprocessing neural network model as described in any one of claims 1 to 4 to obtain a processed image; An imaging simulation module is used to perform optomechanical imaging simulation on any of the processed images to obtain a corresponding simulated image. An update module is used to construct a loss function based on any sample image in the sample set and the corresponding simulated image, and to update the parameters of the image preprocessing neural network model based on the loss function.
13. An electronic device, characterized in that, include: Memory, used to store computer programs; A processor, configured to execute the computer program to implement the steps of the training method for the image preprocessing neural network model as described in any one of claims 5 to 10, or to implement the steps of the near-eye display method as described in claim 11.
14. A readable storage medium, characterized in that, The readable storage medium stores a computer program that, when executed by a processor, implements the steps of the training method for the image preprocessing neural network model as described in any one of claims 5 to 10, or the steps of the near-eye display method as described in claim 11.
Citation Information
Patent Citations
Binocular thermal imaging system and super-resolution image acquisition method
CN112927139A
Color gamut correction method of near-to-eye display equipment based on color characterization
CN117376492A
Zero-sample low-light image enhancement method based on dynamic feature aggregation and color correction
CN118941483A
Underwater image enhancement method based on color correction and wavelet fusion
CN120070293A
Chroma perception multi-resolution image fusion network framework and method based on state space model
CN120526266A