Training method, display control method and device of view optimization model

By training a view optimization model to process multi-viewpoint views of naked-eye 3D displays and using convolutional neural networks for crosstalk compensation, the crosstalk problem of naked-eye 3D displays is solved, thereby improving display quality and user experience.

CN121509637APending Publication Date: 2026-02-10BOE TECHNOLOGY GROUP CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511785213.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Naked-eye 3D displays are prone to imaging crosstalk between the left and right views during the display process, which leads to ghosting and visual fatigue, and existing technologies are difficult to solve effectively.

Method used

By training the view optimization model, using convolutional neural networks to process pixel data of multi-view views, generating predicted multi-view views, and adjusting model parameters by calculating model loss, crosstalk compensation is achieved to counteract the crosstalk effect of the target display device.

Benefits of technology

It significantly reduces the user-perceived crosstalk rate, improves the display quality and viewing comfort, and effectively suppresses ghosting.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121509637A_ABST
    Figure CN121509637A_ABST
Patent Text Reader

Abstract

A view optimization model training method, display control method and apparatus, the training method comprising: obtaining training data, the training data comprising pixel data of a plurality of groups of first multi-view views; inputting the pixel data of the first multi-view view into a to-be-trained view optimization model to obtain pixel data of a second multi-view view; according to the pixel data of the second multi-view view and a crosstalk rate parameter of a target display device, generating pixel data of a predicted multi-view view of the second multi-view view on the target display device; calculating the model loss of a view optimization model according to the pixel data of the first multi-view view, the pixel data of the second multi-view view and the pixel data of the predicted multi-view view, and adjusting the model parameters of the view optimization model based on the calculated model loss to obtain a trained view optimization model.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure relates to, but is not limited to, the technical field of display, and in particular to a training method of a view optimization model, a display control method and device. BACKGROUND

[0002] Naked-eye 3D display technology has been widely used in advertising media, exhibition display, consumer electronics and other fields due to its advantage of presenting stereoscopic visual effect without wearing special glasses.

[0003] However, due to process limitations, imaging crosstalk phenomenon of left and right views is prone to occur in the display process of naked-eye 3D display. Crosstalk is a situation in which light that one eye of a viewer sees to some extent is seen by the other eye, i.e. the picture originally entering the left eye is seen by the right eye, and the picture originally entering the right eye is seen by the left eye. This phenomenon can seriously degrade image quality, causing the viewer to perceive ghosting and causing visual fatigue. SUMMARY

[0004] The embodiments of the present disclosure provide a training method of a view optimization model, a display control method and device, which can significantly reduce the crosstalk rate perceived by the user and improve the display quality of the naked-eye 3D display and the viewing comfort of the user.

[0005] The embodiments of the present disclosure provide a training method of a view optimization model, comprising: obtaining training data, the training data comprising pixel data of a plurality of first multi-viewpoint views; inputting the pixel data of the first multi-viewpoint views into a to-be-trained view optimization model to obtain a second multi-viewpoint view; generating pixel data of a predicted multi-viewpoint view of the second multi-viewpoint view on a target display device according to the pixel data of the second multi-viewpoint view and a crosstalk rate parameter of the target display device; calculating a model loss of the view optimization model according to the pixel data of the first multi-viewpoint view, the pixel data of the second multi-viewpoint view and the pixel data of the predicted multi-viewpoint view, and adjusting a model parameter of the view optimization model based on the calculated model loss to obtain a trained view optimization model.

[0006] The embodiments of the present disclosure also provide a training device of a view optimization model, comprising a memory and a processor connected to the memory, the memory is used to store instructions, and the processor is configured to execute the steps of the training method of the view optimization model according to the instructions stored in the memory.

[0007] This disclosure also provides a computer-readable storage medium storing executable instructions that, when executed by a processor, can implement the view optimization model training method as described in any embodiment of this disclosure.

[0008] This disclosure also provides a computer program product including instructions that, when executed by a computer, perform a training method for a view optimization model as described in any embodiment of this disclosure.

[0009] This disclosure also provides a display control method, including: The pixel data of the initial multi-view view is input into a pre-trained view optimization model for crosstalk compensation to obtain the pixel data of the second multi-view view. The view optimization model is trained by the view optimization model training method as described in any embodiment of this disclosure. The pixel data of the second multi-viewpoint view is subjected to image interleaving processing, and the pixel data after image interleaving processing is input to the target display device for display. The target display device is a naked-eye 3D display.

[0010] This disclosure also provides a display control device, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute steps of the display control method as described in any embodiment of this disclosure based on the instructions stored in the memory.

[0011] This disclosure also provides a computer-readable storage medium storing executable instructions that, when executed by a processor, can implement the display control method as described in any embodiment of this disclosure.

[0012] This disclosure also provides a computer program product including instructions that, when executed by a computer, perform a display control method as described in any embodiment of this disclosure.

[0013] The view optimization model training method, display control method, and apparatus of this disclosure involve inputting pixel data of a first multi-view view into the view optimization model to obtain pixel data of a second multi-view view. Based on the pixel data of the second multi-view view and the crosstalk rate parameter of the target display device, pixel data of a predicted multi-view view of the second multi-view view on the target display device is generated. A model loss is calculated based on the pixel data of the first, second, and predicted multi-view views. The model parameters of the view optimization model are adjusted based on the calculated model loss. This ensures that crosstalk compensation processing is performed on the multi-view view before the image is output to the target display device. This crosstalk compensation process, after passing through the crosstalk effect of the target display device itself, is precisely canceled out, resulting in a near-ideal image perceived by the human eye. This significantly reduces the user-perceived crosstalk rate, effectively suppresses ghosting, and fundamentally improves the display quality and viewing comfort of naked-eye 3D displays. It is adaptable to various naked-eye 3D display devices with different crosstalk rate levels.

[0014] Other features and advantages of this disclosure will be set forth in the following description, and will be apparent in part from the description, or may be learned by practicing the disclosure. Other advantages of this disclosure may be realized and obtained by means of the methods described in the description and the accompanying drawings. Attached Figure Description

[0015] The accompanying drawings are used to provide an understanding of the technical solutions of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the technical solutions of this disclosure and do not constitute a limitation on the technical solutions of this disclosure.

[0016] Figure 1A This is a schematic diagram illustrating an ideal image (without crosstalk) seen by the left and right eyes as an exemplary embodiment of the present disclosure; Figure 1B This is a schematic diagram of the actual image (with crosstalk) seen by the left and right eyes as an exemplary embodiment of the present disclosure; Figure 2 This is a flowchart illustrating a method for training a view optimization model, which is an exemplary embodiment of this disclosure. Figure 3 This is a schematic diagram illustrating the display principle of a multi-viewpoint display panel, which is an exemplary embodiment of the present disclosure. Figure 4 This is a schematic diagram of an actual image seen by the left and right eyes after crosstalk compensation by a view optimization model, as an exemplary embodiment of the present disclosure. Figure 5 This is a flowchart illustrating an exemplary embodiment of a display control method according to the present disclosure; Figure 6 A schematic diagram of the structure of a training device for a view optimization model, which is an exemplary embodiment of the present disclosure; Figure 7 This is a schematic diagram of the structure of a display control device as an exemplary embodiment of the present disclosure. Detailed Implementation

[0017] The specific embodiments of this disclosure will be described in further detail below with reference to the accompanying drawings and examples. The following examples are used to illustrate this disclosure, but are not intended to limit its scope. It should be noted that, unless otherwise specified, the embodiments and features described in this disclosure can be arbitrarily combined with each other.

[0018] The display principle of glasses-free 3D displays is as follows: a lens cylinder or parallax grating is placed in front of the display panel, causing the image seen by the left eye to differ from that seen by the right eye, thus creating a 3D visual effect. However, due to the small area of ​​pixels, and the limitations of manufacturing processes, the beam-splitting cylinder or grating cannot be made to the same level as the pixels. Therefore, during the display process, it is inevitable that pixels intended for the left eye will be projected onto the right eye, and vice versa, resulting in crosstalk.

[0019] like Figure 1A As shown, taking a 3D left and right viewpoint image of a red-green map as an example, the left view is red and the right view is green. After image interlacing processing, when displayed on a naked-eye 3D display device, ideally without crosstalk, if you cover your right eye and only use your left eye, the image you see should be entirely red; if you cover your left eye and only use your right eye, the image you see should be entirely green. Figure 1B As shown, in reality, the red image seen by the left eye contains green parts, and the green image seen by the right eye contains red parts. That is, the image from the left eye leaks into the right eye's view area, and the image from the right eye leaks into the left eye's view area, which is the phenomenon of image crosstalk.

[0020] The main methods for solving image crosstalk problems using related technologies include: 1) Optical component optimization: Crosstalk can be reduced by improving the optical structure of parallax barriers or cylindrical lenses. However, this method is limited by the manufacturing precision of physical components and often comes at the cost of resolution.

[0021] 2) Image post-processing techniques: Traditional image processing methods such as image filtering, edge enhancement, and pixel blurring in crosstalk areas are used to reduce ghosting. However, these methods cannot fundamentally solve the crosstalk problem and have limited effectiveness.

[0022] like Figure 2 As shown, this disclosure provides a method for training a view optimization model, including: Step 201: Obtain training data, wherein the training data includes pixel data of multiple sets of first multi-viewpoint views; Step 202: Input the pixel data of the first multi-view view into the view optimization model to be trained to obtain the second multi-view view; Step 203: Generate the pixel data of the predicted multiview view of the second multiview view on the target display device based on the pixel data of the second multiview view and the crosstalk rate parameter of the target display device; Step 204: Calculate the model loss of the view optimization model based on the pixel data of the first multi-view view, the pixel data of the second multi-view view, and the pixel data of the predicted multi-view view. Adjust the model parameters of the view optimization model based on the calculated model loss to obtain the trained view optimization model.

[0023] The training method for the view optimization model in this embodiment involves inputting a first multi-view view into the view optimization model to be trained to obtain a second multi-view view. Based on the pixel data of the second multi-view view and the crosstalk rate parameters of the target display device, pixel data of the predicted multi-view view on the target display device is generated. The model loss of the view optimization model is calculated based on the pixel data of the first, second, and predicted multi-view views. The model parameters of the view optimization model are adjusted based on the calculated model loss, so that the crosstalk compensation of the view optimization model for the multi-view view exactly cancels out the crosstalk effect of the target display device itself. This results in a near-ideal image for the human eye, significantly reducing the user-perceived crosstalk rate, effectively suppressing ghosting, and improving image contrast and detail clarity. This fundamentally improves the display quality and viewing comfort of naked-eye 3D display devices and is adaptable to various naked-eye 3D display devices with different crosstalk rate levels.

[0024] In this embodiment of the disclosure, the target display device can be any kind of naked-eye 3D display device. This embodiment of the disclosure does not limit the structure, type, size, etc. of the target display device, nor does it limit the naked-eye 3D technology used by the target display device.

[0025] In this embodiment of the disclosure, the target display device can be a 2-viewpoint display device or any n-viewpoint display device, where n is a natural number greater than 2. For example, the target display device can be a 5-viewpoint, 9-viewpoint, 24-viewpoint, or 48-viewpoint display device, etc. By setting multiple viewpoints, users can see the 3D display image from multiple locations.

[0026] like Figure 3As shown, the display panel 10 is provided with five viewpoints: viewpoint 1, viewpoint 2, viewpoint 3, viewpoint 4, and viewpoint 5. At this time, the grating 11 located in front of the display panel 10 allows the user's eyes at a certain position to see the display images corresponding to two adjacent viewpoints among the five viewpoints. For example, the user's left eye can see the display image corresponding to viewpoint 3, and the user's right eye can see the display image corresponding to viewpoint 2. At this time, the user can see a 3D display image.

[0027] In this embodiment of the disclosure, each group of first multi-viewpoint views includes images corresponding to the number of viewpoints. For example, when the number of viewpoints is 2, each group of first multi-viewpoint views includes 2 images, one of which is a left viewpoint image and the other is a right viewpoint image. As another example, when the number of viewpoints is 5, each group of first multi-viewpoint views includes 5 images, which can be sequentially labeled as the first viewpoint image, the second viewpoint image, the third viewpoint image, the fourth viewpoint image, and the fifth viewpoint image.

[0028] In this embodiment of the disclosure, a multi-view view can be a collection of images obtained by multiple cameras taking pictures of the same scene from different angles (such as horizontal surround or vertical distribution). However, this disclosure does not limit this, and in other examples, a multi-view view can also be multiple images with unrelated content.

[0029] In this embodiment of the disclosure, the pixel data corresponding to each group of first multi-view views is uncompensated crosstalk data. This embodiment of the disclosure inputs the pixel data of the first multi-view views into a view optimization model, and obtains a crosstalk-compensated second multi-view view based on the output of the view optimization model. That is, the pixel data corresponding to each group of second multi-view views is data for which crosstalk compensation has been applied to the pixel data of the first multi-view views using the view optimization model. The pixel data corresponding to each group of predicted multi-view views is data for crosstalk prediction of the second multi-view views based on the crosstalk rate parameter of the target display device.

[0030] In this embodiment, the pixel data corresponding to each group of multi-view views (including a first multi-view view, a second multi-view view, and a predicted multi-view view) may include pixel values ​​of multiple pixels, wherein the pixel value of each pixel may be luminance data. In other embodiments, the pixel data corresponding to each group of multi-view views may also include at least one of the following: RGB values, grayscale values, contrast, color saturation, or other image data. The following description will use luminance data as an example for the pixel data corresponding to each group of multi-view views.

[0031] In this embodiment of the disclosure, the view optimization model can be a convolutional neural network (CNN); however, this disclosure does not limit it.

[0032] In some exemplary embodiments, the number of viewpoints corresponding to the multi-viewpoint is 2. In step 203, based on the pixel data of the second multi-viewpoint view and the crosstalk rate parameter of the target display device, pixel data of the predicted multi-viewpoint view of the second multi-viewpoint view on the target display device is generated, including: The pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the predicted pixel value of each pixel in the multi-view view: = (1 - c) * + c * Formula (1); = c * + (1 - c) * Formula (2); in, To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in a multi-view view; To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in a multi-viewpoint view; This is the pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the second multi-viewpoint view; is the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the second multi-view view, where i and j are both natural numbers, and c is the crosstalk rate of the target display device, and c is between 0 and 1.

[0033] In this embodiment of the disclosure, the value ranges of i and j are determined according to the number of pixels contained in each view. For example, assuming that the view contains 3840*2160 pixels, the value range of i can be between 0 and 3839, or between 1 and 3840; the value range of j can be between 0 and 2159, or between 1 and 2160.

[0034] In this embodiment of the disclosure, the crosstalk rate c of the target display device can be the average crosstalk rate. That is, assuming that the crosstalk rate c is uniform across the entire display screen of the target display device, the crosstalk rate c is a real number between 0 and 1 (for example, c = 0.05 represents 5% crosstalk). When c = 0, the image actually seen by the left and right eyes is the ideal image.

[0035] In practical use, the crosstalk rate *c* of the target display device can be obtained through standard optical device testing methods. Typically, the crosstalk rate of the target display device is not uniform (e.g., ...). Figure 1B As shown in the figure, the crosstalk rate of the target display device is related to the viewing position and the position of the pixels on the screen. The crosstalk rate is small or almost non-existent at the center position and large at the edge. In this embodiment, for the sake of simplicity, the crosstalk rate c of the target display device is represented by the average crosstalk rate.

[0036] In this embodiment of the disclosure, when generating the pixel data of the predicted multiview view of the second multiview view on the target display device based on the pixel data of the second multiview view and the crosstalk rate parameter of the target display device, the calculation is performed pixel by pixel according to the aforementioned formulas (1) and (2). Since each view can be represented as a pixel matrix in the computer, the aforementioned formulas (1) and (2) are used independently to calculate each element in the pixel matrix to obtain the predicted multiview view.

[0037] Below is a very simplified 3 3. An example of a grayscale image will be used for illustration. Assume that in the second multi-viewpoint view, the left view L_opt (the actual image is a black-and-white grid image) is represented in the computer as a pixel matrix as follows: .

[0038] The right view R_opt (the actual image is a black and white grid image) is represented in the computer as a pixel matrix as follows: .

[0039] The pixel-by-pixel calculation process is as follows: Assuming a crosstalk rate of c = 0.1 (10% crosstalk rate), apply the aforementioned formulas (1) to (2) to calculate the predicted left view L_ in the multi-view view. The first pixel: = (1–c) * + c * =(1-0.1)*255 + 0.1*0 = 0.9*255 +0.1*0 = 229.5 ≈ 230.

[0040] Similarly, the prediction of the right view in a multi-viewpoint view is calculated. The first pixel: = c* + (1–c)* = 0.1*255 +(1-0.1) *0 = 0.1*255 +0.9*0 = 25.5 ≈ 26.

[0041] Similarly, the calculations are performed on the other pixels sequentially to obtain the left view L_ in the complete predicted multi-view view. : .

[0042] This image matrix, when applied to the actual image, means that areas that were originally pure white have turned gray (mixed with black), and areas that were originally pure black have also turned gray (mixed with white), resulting in ghosting. This can also be seen from formulas (1) and (2). It contains c * This ingredient shouldn't have been seen. It contains c * This unseen component is the ghosting caused by crosstalk.

[0043] Formulas (1) and (2) can be written in matrix form as follows: Formula (3).

[0044] Formula (3) can be simplified to: A = M * I, where A = I= M= M is the crosstalk matrix. When the crosstalk rate c is 0.1, the crosstalk matrix M = It should be noted that this explanation uses a 2-viewpoint example. When the number of viewpoints is n (n>2), the crosstalk matrix M is an n*n matrix.

[0045] In some exemplary embodiments, step 204, calculating the model loss of the view optimization model based on the pixel data of the first multi-view view, the pixel data of the second multi-view view, and the pixel data of the predicted multi-view view, includes: Substitute the pixel value of each pixel in the first multiview, the second multiview, and the predicted multiview into the following formula to calculate the total loss for each pixel. Then, sum the total losses for all pixels to obtain the model loss of the view optimization model: = α * + β * Formula (4); =|| - || + || - || Formula (5); = || - || + || - || Formula (6); in, The total loss of the pixel in the i-th row and j-th column is... The first loss is for the pixel in the i-th row and j-th column. The second loss is for the pixel in the i-th row and j-th column, where α is the first weighting coefficient, β is the second weighting coefficient, and α+β=1. To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in a multi-view view, To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in a multi-view view, This is the pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the second multi-viewpoint view; This represents the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the second multi-view view. This is the pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the first multi-viewpoint view; This is the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the first multi-viewpoint view.

[0046] In this embodiment, || ... || represents the norm, which can be either the L1 norm (sum of absolute errors) or the L2 norm (sum of squared errors). The L1 norm represents the sum of the absolute values ​​of the differences. It is insensitive to outliers and tends to maintain sharp edge images, making it suitable for suppressing ghosting when clear object boundaries are desired. The L2 norm represents the sum of the squared differences. It is more sensitive to outliers and tends to produce smoother images, but can easily lead to blurred details. Since crosstalk in naked-eye 3D applications mainly manifests as image crosstalk at boundaries, using the L1 norm can better preserve edge information and avoid over-smoothing.

[0047] If the image optimized by the view optimization model is perfect, then predict the pixel in the i-th row and j-th column of the multi-view view. , [The pixel in row i and column j of the original first multi-view view should be consistent with the pixel in row i and column j of the original multi-view view.] , They are completely identical. Therefore, the difference between them can be used to guide the learning of the view optimization model. As shown in Equation (5), the first loss is defined. (Crosstalk simulation loss) is the loss for predicting the pixel in the i-th row and j-th column of a multi-view view. , [The pixel in row i and column j of the original first multi-view view] , The difference between them.

[0048] However, using only crosstalk simulation loss as the loss function may cause the view optimization model to generate some mathematically correct but visually unnatural and noisy optimized images. Therefore, it is necessary to add a constraint to ensure that the optimized image itself looks normal. As shown in Equation (6), the second loss is defined. (Image fidelity loss) is the pixel in the i-th row and j-th column of the second multi-view view. , [The pixel in row i and column j of the original first multi-view view] , The difference between them.

[0049] As shown in formula (4), the first loss Second loss The weighted combination yields the total loss for the pixel in the i-th row and j-th column. The values ​​of the first weighting coefficient α and the second weighting coefficient β can be adjusted according to the relative importance of the two losses. For example, the value of α ranges from 0.5 to 0.8, and the value of β ranges from 0.2 to 0.5. In actual use, the values ​​can be adjusted according to the actual target display device.

[0050] In this embodiment, the total loss of all pixels is summed, and the summation result is used as the final model loss. This summation can be performed by directly summing the total losses of all pixels, or by performing a weighted summation. The result of either direct summation or weighted summation is used as the final model loss. In the weighted summation, the weight corresponding to each pixel can be set as needed.

[0051] The model parameters are adjusted using the backpropagation algorithm to minimize the total loss across all pixels. Or the weighted sum of the total loss of all pixels ,in, Let be the weight coefficient for the pixel in the i-th row and j-th column. During training, the model attempts to find the optimal model parameters to minimize the overall model loss. The resulting view optimization model learns how to generate a second multi-view that effectively suppresses ghosting. In the early stages of model training, the image quality of the second multi-view generated by the view optimization model is poor, resulting in a first loss. Second loss Both are relatively large, and the model simultaneously learns to compensate for crosstalk and maintain image quality; in the middle stage of model training, the view optimization model begins to find effective compensation patterns, with the first loss... Significant decline, second loss To ensure that compensation does not excessively distort the image; during the convergence phase of model training, the view optimization model finds the optimal balance point, the first loss. Second loss All parameters are at a low level, effectively suppressing crosstalk while maintaining a natural image quality. During model training, automatic balancing algorithms can be incorporated, such as adaptive weights based on the loss magnitude, uncertainty-based weighting, and dynamic gradient balancing. Dynamic balancing of the first loss is also possible. Second loss .

[0052] In other exemplary embodiments, the model loss of the view optimization model is calculated based on pixel data of the first multi-view view, pixel data of the second multi-view view, and pixel data of the predicted multi-view view, including: Substitute the pixel value of each pixel in the first multiview, the second multiview, and the predicted multiview into the following formula to calculate the total loss for each pixel. Then, sum the total losses for all pixels to obtain the model loss of the view optimization model: = α * + β * + γ *Loss_perceptual formula (7); =|| - || + || - || Formula (8); = || - || + || - || Formula (9); Loss_perceptual = || Φ(L_opt) - Φ(L_ideal) || + || Φ(R_opt) - Φ(R_ideal) || Formula (10).

[0053] in, The total loss of the pixel in the i-th row and j-th column is... The first loss is for the pixel in the i-th row and j-th column. Let be the second loss for the pixel in the i-th row and j-th column, and let Loss_perceptual be the feature-perceptual loss. Let α be the first weight coefficient, β be the second weight coefficient, and γ be the third weight coefficient, with α + β + γ = 1. To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in a multi-view view, To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in a multi-view view, This is the pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the second multi-viewpoint view; This represents the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the second multi-view view. This is the pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the first multi-viewpoint view; Let Φ be the pixel value of the pixel in the i-th row and j-th column corresponding to the right view in the first multi-view view, Φ represent the feature extractor in the pre-trained network, Φ(L_opt) is the image quality feature value of the left view in the second multi-view view, Φ(R_opt) is the image quality feature value of the right view in the second multi-view view, Φ(L_ideal) is the image quality feature value of the left view in the first multi-view view, and Φ(Rideal) is the image quality feature value of the right view in the first multi-view view.

[0054] In this embodiment of the disclosure, when calculating the total loss for each pixel, in order to further improve visual quality, in addition to the first loss... Second loss In addition, a feature-perceptual loss, Loss_perceptual, is introduced, as shown in Equation (10). Loss_perceptual can be obtained by extracting image quality features of the first and second multi-view views through a pre-trained deep network and comparing the differences in the image quality features of the first and second multi-view views in the feature space. The type of the pre-trained deep network can be set as needed, for example, it can be a VGG (Visual Geometry Group) network; however, this disclosure does not limit it. By adding the feature-perceptual loss, the generated image can be semantically similar to the original image, thereby better preserving visual quality.

[0055] As shown in formula (7), the first loss Second loss The total loss for each pixel is obtained by combining the feature-perceptual loss with the feature-perceptual loss. The values ​​of the first weighting coefficient α, the second weighting coefficient β, and the third weighting coefficient γ can be adjusted according to the relative importance of the three losses; for example, α = 0.6, β = 0.3, and γ = 0.1. However, this disclosure does not limit this, and adjustments can be made according to the actual target display device in practical use.

[0056] In other exemplary embodiments, the crosstalk rate parameter of the target display device can be defined as a function related to the row and column positions of the pixels. Taking a multi-viewpoint view with two viewpoints as an example, the pixel value of each pixel in the second multi-viewpoint view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the predicted pixel value of each pixel in the multi-viewpoint view: = (1- ) * + * Formula (11); = * + (1- ) * Formula (12).

[0057] In practical glasses-free 3D displays, crosstalk rate is not uniform and is typically related to the viewing position and the pixel's location on the screen (e.g., center versus edge). By defining the crosstalk rate parameter of the target display device as a function related to the row and column positions of the pixels, a spatially varying crosstalk rate distribution can be obtained. This method, when implemented, more accurately describes the crosstalk characteristics of different screen regions.

[0058] However, even if the crosstalk rate parameters of the target display device are trained using a uniform crosstalk rate, the view optimization model of this disclosure embodiment can still effectively handle non-uniform crosstalk in real-world systems. The reasons are as follows: (1) The local characteristics of the view optimization model enable it to adapt to crosstalk compensation in different image regions. As mentioned earlier, the view optimization model can employ a CNN, which features local connectivity and weight sharing. This means that each convolutional kernel in the network slides across the entire image and independently processes each local region. Therefore, even if the actual crosstalk rate varies in different regions of the image (e.g., less crosstalk in the center of the screen and more crosstalk at the edges), the CNN can adaptively apply different processing methods to different regions through the learned convolutional kernels. In other words, the network can learn to apply stronger compensation to regions with high crosstalk and weaker compensation to regions with low crosstalk.

[0059] (2) The view optimization model learns a general compensation strategy during training, rather than a specific numerical result. The training objective of the view optimization model is to minimize the crosstalk simulation loss, that is, to teach the model to generate a set of optimized views that approximate the ideal image after uniform crosstalk simulation. In this process, the view optimization model does not simply learn a fixed mathematical transformation, but rather learns a general strategy for adjusting the output based on the input image content and crosstalk rate. This general strategy can be extended to non-uniform crosstalk because the view optimization model has learned how to adjust the output to compensate for crosstalk based on image features (such as edges and textures). Therefore, even if the actual crosstalk rate distribution is uneven, the view optimization model can adaptively compensate based on local image features.

[0060] For example, such a general strategy may include one or more of the following: stronger compensation is needed at object boundaries, consistency is needed in complex texture areas, compensation is needed in flat areas, and color balance is maintained.

[0061] (3) By using training data that includes diversity, the view optimization model can learn a more comprehensive crosstalk compensation capability. During training, the crosstalk rate parameter of the target display device can use various crosstalk rate conditions. After training, the view optimization model will learn to handle various crosstalk scenarios. The view optimization model's processing capability will be stronger if the training data includes samples with both uniform crosstalk (the crosstalk rate is the same at different pixel positions on the screen) and non-uniform crosstalk (the crosstalk rate is different at different pixel positions on the screen). Even if the target display device's crosstalk rate parameter is trained using only uniform crosstalk rate, the view optimization model can still learn a robust compensation strategy by using a large amount of data with different scenarios and crosstalk rates. This strategy can generalize to some extent to non-uniform crosstalk scenarios.

[0062] For example, different scenarios may include different viewing angles (simulating spatial variation crosstalk), different image types, and different display characteristics. For instance, different viewing angles may include left, center, right, top, and bottom; different image types may include high-frequency textures, smoothed gradients, sharpened edges, complex scenes, and face images; and different display characteristics may include equalized crosstalk, edge-enhanced crosstalk, and center optimization.

[0063] After the view optimization model is trained, it can be deployed to the target display device (a glasses-free 3D display device). During the display process on the target display device, the initial multi-viewpoint view to be displayed is input into the view optimization model. The second multi-viewpoint view output by the view optimization model undergoes image interleaving processing. The image-interleaved image is then input to the target display device for display. At this time, due to the physical crosstalk of the target display device itself, the "pre-distortion" in the second multi-viewpoint view precisely cancels out the crosstalk effect of the screen, ultimately allowing the human eye to see a clear left and right view image with almost no ghosting. That is, the view optimization model achieves pre-emptive and active compensation for crosstalk, effectively suppressing the ghosting phenomenon. Figure 4 As shown, by comparing the red and green views, it can be seen that the crosstalk seen by the left and right eyes after the crosstalk optimization of the present disclosure embodiment is greatly reduced.

[0064] In some exemplary embodiments, the number of viewpoints corresponding to multiple viewpoints is n, where n is greater than 2.

[0065] Based on the pixel data of the second multi-view view and the crosstalk rate parameter of the target display device, pixel data of the predicted multi-view view of the second multi-view view on the target display device is generated, including: The pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the predicted pixel value of each pixel in the multi-view view: Formula (13); in, To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in a multi-view view, is the pixel value of the pixel in the i-th row and j-th column of the k-th viewpoint in the second multi-view view, where k is a natural number between 1 and n; Let be the crosstalk rate between the y-th viewpoint and the x-th viewpoint in the target display device, and Let x be a real number between 0 and 1, and let x and y be natural numbers between 1 and n with x ≠ y. Let c11 = c12 + c13 + ... + c1n, caa = ca1 + ... + ca(a-1) + ca(a+1) + ... + can, cnn = cn1 + cn2 + ... + cn(n-1), and a be a natural number between 2 and n-1.

[0066] The training method for the view optimization model and the display control method of this disclosure are not limited to the size of the number of viewpoints n. Optimization can be performed for any number of viewpoints, such as 4, 9, 24, or 48. For example, taking n=9 as an example, the pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the predicted pixel value of each pixel in the multi-view view: Formula (14).

[0067] Since crosstalk typically occurs between adjacent views, the above formula (13) can be simplified. Specifically, formula (13) can be simplified to: Formula (15).

[0068] Taking n=9 as an example, the pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the pixel value of each pixel in the predicted multi-view view: Formula (16).

[0069] Furthermore, in practical applications, crosstalk between adjacent viewpoints can be considered symmetrical and uniform. Therefore, the above formula (15) can be further simplified. Specifically, formula (15) can be simplified as follows: Formula (17).

[0070] in, Let c be the crosstalk rate of the target display device, and c is a real number between 0 and 1.

[0071] Taking n=9 as an example, the pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the pixel value of each pixel in the predicted multi-view view: Formula (18).

[0072] When there are more than 2 viewpoints, the model loss can be calculated using any of the following methods.

[0073] In some exemplary embodiments, calculating the model loss of the view optimization model based on pixel data of a first multi-view view, pixel data of a second multi-view view, and pixel data of a predicted multi-view view includes: Substitute the pixel value of each pixel in the first multiview, the second multiview, and the predicted multiview into the following formula to calculate the total loss for each pixel. Then, sum the total losses for all pixels to obtain the model loss of the view optimization model: = α* + β* Formula (19); =α1*|| - ||+α2*|| - ||+……+αn*|| - || Formula (20); =β1*|| - ||+β2*|| - ||+……+βn*|| - || Formula (21).

[0074] in, The total loss of the pixel in the i-th row and j-th column is... The first loss is for the pixel in the i-th row and j-th column. The second loss is for the pixel in the i-th row and j-th column, where α, β, α1 to αn, and β1 to βn are all weighting coefficients. To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in a multi-view view, This is the pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the second multi-viewpoint view; is the pixel value of the pixel in the i-th row and j-th column of the k-th viewpoint in the first multi-view view, where k is a natural number between 1 and n, and n is greater than 2.

[0075] In other exemplary embodiments, the model loss of the view optimization model is calculated based on pixel data of the first multi-view view, pixel data of the second multi-view view, and pixel data of the predicted multi-view view, including: Substitute the pixel value of each pixel in the first multiview, the second multiview, and the predicted multiview into the following formula to calculate the total loss for each pixel. Then, sum the total losses for all pixels to obtain the model loss of the view optimization model: = α* + β* + γ*Loss_perceptual formula (22); =α1*|| - ||+α2*|| - ||+……+αn*|| - || Formula (23); =β1*|| - ||+β2*|| - ||+……+βn*|| - || Formula (24); Loss_perceptual=γ1*||Φ(V1_opt)-Φ(V1_ideal)||+γ2*||Φ(V2_opt)-Φ(V2_ideal)||+…+γn*||Φ(Vn_opt) - Φ(Vn_ideal) || Formula (25).

[0076] in, The total loss of the pixel in the i-th row and j-th column is... The first loss is for the pixel in the i-th row and j-th column. The second loss is for the pixel in the i-th row and j-th column, where Loss_perceptual is the feature-perceptual loss, and α, β, α1 to αn, β1 to βn, and γ1 to γn are all weight coefficients. To predict the pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in a multi-view view, This is the pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the second multi-viewpoint view; Φ(Vk_opt) is the pixel value of the pixel in the i-th row and j-th column of the k-th viewpoint in the first multi-view view, Φ(Vk_ideal) is the image quality feature value of the k-th view in the second multi-view view, and Φ(Vk_ideal) is the image quality feature value of the k-th view in the first multi-view view. k is a natural number between 1 and n, and n is greater than 2.

[0077] For multiple viewpoint norms, the importance of different viewpoints may vary in practical applications. Therefore, this embodiment introduces viewpoint weights α1 to αn, β1 to βn, and γ1 to γn to perform weighted summation of the three losses for different viewpoints. Specifically, α1 to αn, β1 to βn, and γ1 to γn can be set according to the weight of the viewing position or the weight of the severity of crosstalk. For example, some viewpoints may correspond to more common positions, so they have higher weights, while some viewpoints may correspond to less common positions, so they have lower weights; or, some viewpoints with severe crosstalk need more focus on optimization, so they have higher weights, while some viewpoints with less severe crosstalk do not need to be focused on optimization, so they have lower weights.

[0078] like Figure 5 As shown in the embodiments of this disclosure, a display control method is also provided, including: Step 501: Input the pixel data of the initial multi-view view into the pre-trained view optimization model for crosstalk compensation to obtain the pixel data of the second multi-view view. The view optimization model is trained by the view optimization model training method as described in any embodiment of this disclosure. Step 502: Perform image interleaving processing on the pixel data of the second multi-viewpoint view, and input the pixel data after image interleaving processing to the target display device for display. The target display device is a naked-eye 3D display.

[0079] In this embodiment, when the view optimization model is trained and deployed on a real glasses-free 3D device, only the view optimization model that has learned "crosstalk compensation" needs to be run. That is, when using the view optimization model for display control, we only need to input the pixel data of the initial multi-view view into the pre-trained view optimization model for crosstalk compensation to obtain the pixel data of the second multi-view view. It is not necessary to generate the predicted multi-view pixel data of the second multi-view view on the target display device based on the pixel data of the second multi-view view and the crosstalk rate parameter of the target display device. This is because when image interleaving is performed on the pixel data of the second multi-view view and the interleaved pixel data is input to the target display device for display, the target display device will automatically generate corresponding crosstalk. In this embodiment, through pre-training, the crosstalk compensation processing performed by the view optimization model is precisely canceled out by the crosstalk effect of the screen itself, thus allowing the human eye to see an image close to ideal.

[0080] For details on how to train the view optimization model, please refer to the training method for the view optimization model described above, which will not be repeated here.

[0081] This disclosure also provides a training apparatus for a view optimization model, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute the steps of a view optimization model training method as described in any embodiment of this disclosure based on the instructions stored in the memory.

[0082] like Figure 6 As shown, in one example, the training device for the view optimization model may include: a first processor 610, a first memory 620, a first bus system 630, and a first transceiver 640, wherein the first processor 610, the first memory 620, and the first transceiver 640 are connected through the first bus system 630, the first memory 620 is used to store instructions, and the first processor 610 is used to execute the instructions stored in the first memory 620 to control the first transceiver 640 to transmit and receive signals. Specifically, the first transceiver 640, under the control of the first processor 610, acquires training data, which includes pixel data of multiple sets of first multi-view views. The first processor 610 inputs the first multi-view views into the view optimization model to be trained to obtain pixel data of the second multi-view views. Based on the pixel data of the second multi-view views and the crosstalk rate parameter of the target display device, the processor generates pixel data of the predicted multi-view view of the second multi-view views on the target display device. Based on the pixel data of the first multi-view views, the pixel data of the second multi-view views, and the pixel data of the predicted multi-view views, the processor calculates the model loss of the view optimization model and adjusts the model parameters of the view optimization model based on the calculated model loss to obtain the trained view optimization model.

[0083] It should be understood that the first processor 610 may be a central processing unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0084] The first memory 620 may include read-only memory and random access memory, and provides instructions and data to the first processor 610. A portion of the first memory 620 may also include non-volatile random access memory. For example, the first memory 620 may also store device type information.

[0085] In addition to a data bus, the first bus system 630 may also include a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 6 The general designated all buses as the first bus system 630.

[0086] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the first processor 610 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the first memory 620. The first processor 610 reads information from the first memory 620 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, further details are omitted here.

[0087] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the training method for the view optimization model as described in any embodiment of this disclosure. The method for quantifying gene expression in multiple cell types by executing executable instructions is essentially the same as the training method for the view optimization model provided in the above embodiments of this disclosure, and will not be described in detail here.

[0088] In some possible implementations, various aspects of the view optimization model training method provided in this disclosure can also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps in the view optimization model training method according to various exemplary embodiments of this disclosure as described above. For example, the computer device can execute the view optimization model training method described in the embodiments of this disclosure.

[0089] This disclosure also provides a display control device, including a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to execute steps of the display control method as described in any embodiment of this disclosure based on the instructions stored in the memory.

[0090] like Figure 7As shown, in one example, the display control device may include: a second processor 710, a second memory 720, a second bus system 730, and a second transceiver 740. The second processor 710, the second memory 720, and the second transceiver 740 are connected via the second bus system 730. The second memory 720 stores instructions, and the second processor 710 executes the instructions stored in the second memory 720 to control the second transceiver 740 to transmit and receive signals. Specifically, the second transceiver 740, under the control of the second processor 710, acquires pixel data of an initial multi-view view. The second processor 710 inputs the pixel data of the initial multi-view view into a pre-trained view optimization model for crosstalk compensation to obtain pixel data of a second multi-view view. This view optimization model is trained using the view optimization model training method described in any embodiment of this disclosure. Image interleaving processing is performed on the pixel data of the second multi-view view, and the image-interleaved pixel data is input to a target display device for display. The target display device is a naked-eye 3D display.

[0091] It should be understood that the second processor 710 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0092] The second memory 720 may include read-only memory and random access memory, and provides instructions and data to the second processor 710. A portion of the second memory 720 may also include non-volatile random access memory. For example, the second memory 720 may also store device type information.

[0093] In addition to a data bus, the second bus system 730 may also include a power bus, a control bus, and a status signal bus. However, for clarity, in... Figure 7 The general designated all buses as the second bus system 730.

[0094] In implementation, the processing performed by the processing device can be accomplished through integrated logic circuits in the hardware of the second processor 710 or through software instructions. That is, the method steps of this embodiment can be executed by a hardware processor, or by a combination of hardware and software modules within the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other storage media. This storage medium is located in the second memory 720. The second processor 710 reads information from the second memory 720 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0095] This disclosure also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the display control method as described in any embodiment of this disclosure. The method for quantifying gene expression in multiple cell types by executing executable instructions is essentially the same as the display control method provided in the above embodiments of this disclosure, and will not be described in detail here.

[0096] In some possible implementations, various aspects of the display control method provided in this disclosure may also be implemented as a program product comprising program code that, when run on a computer device, causes the computer device to perform the steps of the display control method according to various exemplary embodiments of this disclosure as described above. For example, the computer device may execute the display control method described in the embodiments of this disclosure.

[0097] The program product may employ any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: an electrical connection having one or more wires, a portable disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.

[0098] It will be understood by those skilled in the art that all or some of the steps, systems, or apparatuses disclosed above, and their functional modules / units, can be implemented as software, firmware, hardware, or suitable combinations thereof. In hardware implementations, the division between functional modules / units mentioned above does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed collaboratively by several physical components. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or microprocessor, or as hardware, or as an integrated circuit, such as an application-specific integrated circuit (ASIC). Such software may be distributed on a computer-readable medium, which may include computer storage media (or non-transitory media) and communication media (or transient media). As is known to those skilled in the art, the term computer storage media includes volatile and non-volatile, removable and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technologies, CD-ROM, digital versatile disc (DVD) or other optical disc storage, magnetic cartridges, magnetic tape, disk storage or other magnetic storage devices, or any other medium that can be used to store desired information and can be accessed by a computer. Furthermore, it is well known to those skilled in the art that communication media typically contain computer-readable instructions, data structures, program modules, or other data in modulated data signals such as carrier waves or other transmission mechanisms, and may include any information delivery medium.

[0099] It should be noted that the above embodiments or implementation methods are merely exemplary and not restrictive. Therefore, this disclosure is not limited to the content specifically shown and described herein. Various modifications, substitutions, or omissions can be made to the form and details of the implementations without departing from the scope of this disclosure.

Claims

1. A method for training a view optimization model, characterized in that, include: Acquire training data, which includes pixel data from multiple sets of first multi-viewpoint views; The pixel data of the first multi-view view is input into the view optimization model to be trained to obtain the pixel data of the second multi-view view. Based on the pixel data of the second multi-view view and the crosstalk rate parameter of the target display device, pixel data of the predicted multi-view view of the second multi-view view on the target display device is generated; The model loss of the view optimization model is calculated based on the pixel data of the first multi-view view, the pixel data of the second multi-view view, and the pixel data of the predicted multi-view view. The model parameters of the view optimization model are adjusted based on the calculated model loss to obtain a trained view optimization model.

2. The method according to claim 1, characterized in that, The number of viewpoints corresponding to the multi-viewpoint is 2; the step of generating the pixel data of the predicted multi-viewpoint view on the target display device based on the pixel data of the second multi-viewpoint view and the crosstalk rate parameter of the target display device includes: The pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the pixel value of each pixel in the predicted multi-view view: = (1 - c) * + c * ; = c * + (1 - c) * ; in, The pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the predicted multi-view view; The pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the predicted multi-viewpoint view; The pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the second multi-viewpoint view; is the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the second multi-viewpoint view, where i and j are both natural numbers, and c is the crosstalk rate of the target display device, and c is a real number between 0 and 1.

3. The method according to claim 1, characterized in that, The number of viewpoints corresponding to the multi-viewpoint is n, where n is greater than 2; the step of generating the pixel data of the predicted multi-viewpoint view on the target display device based on the pixel data of the second multi-viewpoint view and the crosstalk rate parameter of the target display device includes: The pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the pixel value of each pixel in the predicted multi-view view: ; in, The pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the predicted multi-view view. , is the pixel value of the pixel in the i-th row and j-th column of the k-th viewpoint in the second multi-viewpoint view, where k is a natural number between 1 and n, and i and j are both natural numbers; Let be the crosstalk rate between the y-th viewpoint and the x-th viewpoint in the target display device, and Let x be a real number between 0 and 1, and let x and y be natural numbers between 1 and n with x ≠ y. Let c11 = c12 + c13 + ... + c1n, caa = ca1 + ... + ca(a-1) + ca(a+1) + ... + can, cnn = cn1 + cn2 + ... + cn(n-1), and a be a natural number between 2 and n-1.

4. The method according to claim 1, characterized in that, The number of viewpoints corresponding to the multi-viewpoint is n, where n is greater than 2; the step of generating the pixel data of the predicted multi-viewpoint view on the target display device based on the pixel data of the second multi-viewpoint view and the crosstalk rate parameter of the target display device includes: The pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the pixel value of each pixel in the predicted multi-view view: ; in, The pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the predicted multi-view view. , is the pixel value of the pixel in the i-th row and j-th column of the k-th viewpoint in the second multi-viewpoint view, where k is a natural number between 1 and n, and i and j are both natural numbers; Let be the crosstalk rate between the y-th viewpoint and the x-th viewpoint in the target display device, and Let x be a real number between 0 and 1, and let x and y be natural numbers between 1 and n, with x ≠ y.

5. The method according to claim 1, characterized in that, The number of viewpoints corresponding to the multi-viewpoint is n, where n is greater than 2; the step of generating the pixel data of the predicted multi-viewpoint view on the target display device based on the pixel data of the second multi-viewpoint view and the crosstalk rate parameter of the target display device includes: The pixel value of each pixel in the second multi-view view and the crosstalk rate parameter of the target display device are substituted into the following formula to calculate the pixel value of each pixel in the predicted multi-view view: ; in, The pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the predicted multi-view view. , is the pixel value of the pixel in the i-th row and j-th column of the k-th viewpoint in the second multi-viewpoint view, where k is a natural number between 1 and n, and i and j are both natural numbers; Let c be the crosstalk rate of the target display device, and c is a real number between 0 and 1.

6. The method according to any one of claims 2 to 5, characterized in that, The crosstalk rate parameter of the target display device is the average crosstalk rate, or the crosstalk rate parameter of the target display device is a function related to the row and column positions of the pixels.

7. The method according to claim 1, characterized in that, The step of calculating the model loss of the view optimization model based on the pixel data of the first multi-view view, the pixel data of the second multi-view view, and the pixel data of the predicted multi-view view includes: Substitute the pixel value of each pixel in the first multi-view view, the second multi-view view, and the predicted multi-view view into the following formula to calculate the total loss for each pixel. Then, sum the total losses of all pixels to obtain the model loss of the view optimization model: = α * + β * ; =|| - || + || - ||; = || - || + || - ||; in, The total loss of the pixel in the i-th row and j-th column is... The first loss is for the pixel in the i-th row and j-th column. The second loss is for the pixel in the i-th row and j-th column, where α is the first weighting coefficient and β is the second weighting coefficient. The pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the predicted multi-view view. The pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the predicted multi-view view. The pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the second multi-viewpoint view; This refers to the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the second multi-viewpoint view. The pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the first multi-viewpoint view; is the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the first multi-viewpoint view, where i and j are both natural numbers.

8. The method according to claim 1, characterized in that, The step of calculating the model loss of the view optimization model based on the pixel data of the first multi-view view, the pixel data of the second multi-view view, and the pixel data of the predicted multi-view view includes: Substitute the pixel value of each pixel in the first multi-view view, the second multi-view view, and the predicted multi-view view into the following formula to calculate the total loss for each pixel. Then, sum the total losses of all pixels to obtain the model loss of the view optimization model: = a * + b * + c *Loss_perceptual; =|| - || + || - ||; = || - || + || - ||; Loss_perceptual = || Φ(L_opt) - Φ(L_ideal) || + || Φ(R_opt) - Φ(R_ideal) || in, The total loss of the pixel in the i-th row and j-th column is... The first loss is for the pixel in the i-th row and j-th column. Let be the second loss for the pixel in the i-th row and j-th column, let Loss_perceptual be the feature-perceptual loss, let α be the first weight coefficient, β be the second weight coefficient, and γ be the third weight coefficient. The pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the predicted multi-view view. The pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the predicted multi-view view. The pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the second multi-viewpoint view; This refers to the pixel value of the pixel in the i-th row and j-th column corresponding to the right viewpoint in the second multi-viewpoint view. The pixel value of the pixel in the i-th row and j-th column corresponding to the left viewpoint in the first multi-viewpoint view; Φ(L_opt) is the pixel value of the pixel in the i-th row and j-th column of the right viewpoint in the first multi-view view, where i and j are both natural numbers. Φ(L_opt) is the image quality feature value of the left view in the second multi-view view, Φ(R_opt) is the image quality feature value of the right view in the second multi-view view, Φ(L_ideal) is the image quality feature value of the left view in the first multi-view view, and Φ(Rideal) is the image quality feature value of the right view in the first multi-view view.

9. The method according to claim 1, characterized in that, The step of calculating the model loss of the view optimization model based on the pixel data of the first multi-view view, the pixel data of the second multi-view view, and the pixel data of the predicted multi-view view includes: Substitute the pixel value of each pixel in the first multi-view view, the second multi-view view, and the predicted multi-view view into the following formula to calculate the total loss for each pixel. Then, sum the total losses of all pixels to obtain the model loss of the view optimization model: = α* + β* ; =α1*|| - ||+α2*|| - ||+……+αn*|| - ||; =β1*|| - ||+β2*|| - ||+……+βn*|| - ||; in, The total loss of the pixel in the i-th row and j-th column is... The first loss is for the pixel in the i-th row and j-th column. The second loss is for the pixel in the i-th row and j-th column, where α, β, α1 to αn, and β1 to βn are all weighting coefficients. The pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the predicted multi-view view. The pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the second multi-viewpoint view; is the pixel value of the pixel in the i-th row and j-th column of the k-th viewpoint in the first multi-view view, where k is a natural number between 1 and n, n is greater than 2, and i and j are both natural numbers.

10. The method according to claim 1, characterized in that, The step of calculating the model loss of the view optimization model based on the pixel data of the first multi-view view, the pixel data of the second multi-view view, and the pixel data of the predicted multi-view view includes: Substitute the pixel value of each pixel in the first multi-view view, the second multi-view view, and the predicted multi-view view into the following formula to calculate the total loss for each pixel. Then, sum the total losses of all pixels to obtain the model loss of the view optimization model: = a* + b* + c*Loss_perceptual; =α1*|| - ||+α2*|| - ||+……+αn*|| - ||; =β1*|| - ||+β2*|| - ||+……+βn*|| - ||; Loss_perceptual=γ1*||Φ(V1_opt)-Φ(V1_ideal)||+γ2*||Φ(V2_opt)-Φ(V2_ideal)||+…+γn*||Φ(Vn_opt) - Φ(Vn_ideal) ||; in, The total loss of the pixel in the i-th row and j-th column is... The first loss is for the pixel in the i-th row and j-th column. The second loss is for the pixel in the i-th row and j-th column, where Loss_perceptual is the feature-perceptual loss, and α, β, α1 to αn, β1 to βn, and γ1 to γn are all weight coefficients. The pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the predicted multi-view view. The pixel value of the pixel in the i-th row and j-th column corresponding to the k-th viewpoint in the second multi-viewpoint view; Φ(Vk_opt) is the pixel value of the pixel in the i-th row and j-th column of the k-th viewpoint in the first multi-view view, Φ(Vk_ideal) is the image quality feature value of the k-th view in the second multi-view view, and Φ(Vk_ideal) is the image quality feature value of the k-th view in the first multi-view view. k is a natural number between 1 and n, n is greater than 2, and i and j are both natural numbers.

11. A training device for a view optimization model, characterized in that, The system includes a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the training method for the view optimization model as described in any one of claims 1 to 10 based on the instructions stored in the memory.

12. A display control method, characterized in that, include: The pixel data of the initial multi-view view is input into a pre-trained view optimization model for crosstalk compensation to obtain the pixel data of the second multi-view view. The view optimization model is trained by the view optimization model training method as described in any one of claims 1 to 10. The pixel data of the second multi-viewpoint view is subjected to image interleaving processing, and the pixel data after image interleaving processing is input to the target display device for display. The target display device is a naked-eye 3D display.

13. A display control device, characterized in that, It includes a memory; and a processor connected to the memory, the memory being used to store instructions, the processor being configured to perform the steps of the display control method as claimed in claim 12 based on the instructions stored in the memory.

14. A computer-readable storage medium, characterized in that, The device stores computer-executable instructions for performing a training method for a view optimization model as described in any one of claims 1 to 10, or a display control method as described in claim 12.

15. A computer program product, characterized in that, The instructions include, when the computer program product is executed by a computer, the instructions performing a training method for a view optimization model as described in any one of claims 1 to 10, or a display control method as described in claim 12.

Citation Information

Cited By

  • Data processing method and device for naked eye 3D display

    CN122027779A