Stereoscopic parallax map generation method and device, head-mounted display device and storage medium

CN121348587BActive Publication Date: 2026-08-11HANGZHOU FOCUSIGHT INTELLIGENT TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]然而,相关技术的视差图只能呈现一些简单的图形,并且需要借助一些问答任务来了解用户的训练结果,训练过程枯燥,用户体验感低,最终导致用户在训练时的依从性降低,并且无法长时间进行,从而进一步降低单次训练时可能取得的效果

Benefits of technology

[0039]上述立体视差图生成方法、装置、头戴式显示设备和存储介质,通过提取二维视频的深度信息,并根据训练参数和深度信息进行视差转换,生成视差图,再基于视差图生成两幅存在偏移的随机点图像,最终构成RDS图像。其中,由于二维视频相比较于传统的视差图更复杂,更具有丰富的深度信息,因此,通过提取二维视频的深度信息并转换为与之匹配的视差图,就能突破传统视差图中只包含简单图像的限制。而且,基于二维视频可以提取连续的深度信息,最终整合生成连续的RDS图像帧,也即RDS视频,用户在训练过程中只需观看与理解RDS视频内容即可,无需再完成其他任务,可以让用户在轻松观看RDS视频的同时达到训练的效果,提升了用户双眼立体视训练的体验感。而且,该训练方法使得用户的依从性提高,有利于长时间沉浸式训练,从而进一步提升单次训练时取得的效果。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121348587B_ABST
    Figure CN121348587B_ABST
Patent Text Reader

Abstract

This application relates to a method, apparatus, head-mounted display device, and storage medium for generating stereo disparity maps. The method involves acquiring a two-dimensional video, extracting depth information from image frames in the video, and normalizing the depth information based on training parameters to obtain a disparity map. Random point coordinates are acquired, and a first sequence containing X-axis and Y-axis values ​​is generated based on these coordinates. The difference between the X-axis values ​​and the corresponding pixel values ​​in the disparity map is calculated to obtain a difference column. A second sequence is generated based on the difference column and the Y-axis column of the first sequence. A first random dot map is generated based on the first sequence, and a second random dot map is generated based on the second sequence. The first random dot map is displayed on a first display screen, and the second random dot map is displayed on a second display screen. This method enhances the user's experience during binocular stereo vision training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of stereo disparity map generation, and in particular to a method, apparatus, head-mounted display device, and storage medium for generating stereo disparity maps. Background Technology

[0002] Unlike anisometropic amblyopia, where monocular amblyopia presents with a significant decrease in processing ability, strabismic amblyopia is characterized by the fact that when the amblyopic eye is used alone, its information processing ability does not show a significant reduction, and may even be comparable to that of the healthy eye. However, when both eyes are used simultaneously, the amblyopic eye's processing ability is significantly suppressed. Therefore, unlike monocular input enhancement strategies that are effective for anisometropic amblyopia, finding a strategy that can improve binocular integration under binocular conditions, thereby increasing the proportion of input from the amblyopic eye in strabismic amblyopia, becomes a problem that needs further investigation.

[0003] Stereopsis is a crucial function of the visual system, achieved through both monocular cues (including relative size, occlusion, gradient, and perspective) and binocular cues (including binocular convergence and binocular disparity). Stereopsis via binocular disparity is one of the few functions in the visual system that absolutely requires binocular input. Therefore, training to improve stereopsis through binocular disparity becomes a superior option for enhancing binocular integration in strabismic amblyopia.

[0004] Related technologies typically employ disparity maps to train users' binocular stereoscopic vision. A disparity map consists of two images with the same subject matter, but certain parts of these two images are slightly offset. When viewing an object with both eyes, the offset areas appear as raised or recessed stereoscopic contours. Users perceive these contours to understand the content of the image. Training stereoscopic vision is achieved through designed question-and-answer tasks. For example, users can identify the orientation of objects and letters in an image, and whether the contours are raised or recessed.

[0005] However, the disparity maps of related technologies can only present some simple graphics and require some question-and-answer tasks to understand the user's training results. The training process is tedious and the user experience is poor, which ultimately leads to reduced user compliance during training and makes it impossible to train for a long time, thereby further reducing the effect that can be achieved in a single training session. Summary of the Invention

[0006] Therefore, it is necessary to provide a method, apparatus, head-mounted display device, and storage medium for generating stereo disparity maps that can enhance the user's stereo vision training experience, in response to the aforementioned technical problems.

[0007] Firstly, this application provides a method for generating a stereo disparity map. The method includes:

[0008] A two-dimensional video is acquired, the depth information of the image frames in the two-dimensional video is extracted, and the depth information is normalized based on the training parameters to obtain a disparity map.

[0009] Obtain the coordinates of a random point, and generate a first sequence based on the coordinates of the random point. The first sequence contains two columns of values: the X-axis and the Y-axis.

[0010] The difference between the values ​​in the X-axis column of the first sequence and the pixel values ​​at the corresponding coordinates in the disparity map is used to obtain the difference column. Then, a second sequence is generated based on the difference column and the Y-axis column of the first sequence.

[0011] A first random dot pattern is generated based on the first sequence of numbers, and a second random dot pattern is generated based on the second sequence of numbers. The first random dot pattern is displayed on the first display screen, and the second random dot pattern is displayed on the second display screen.

[0012] In one embodiment, extracting depth information from image frames in the two-dimensional video includes:

[0013] Calculate the distance between each pixel in the image frame and the observation point;

[0014] The depth information is obtained by assigning a value to each pixel in the range of 0 to 1 according to the distance; wherein, the larger the distance, the smaller the assigned value, and the smaller the distance, the larger the assigned value.

[0015] In one embodiment, before normalizing the depth information based on training parameters, the method further includes:

[0016] Obtain the horizontal viewing angle of the first display screen or the second display screen, obtain the horizontal resolution of the first display screen or the second display screen, and obtain the user's first stereoscopic vision detection result;

[0017] Based on the horizontal viewing angle, the horizontal resolution, and the first stereoscopic detection result, the number of image pixels that the user can observe based on the first display screen or the second display screen is calculated;

[0018] The training parameters are calculated based on the number of pixels in the image.

[0019] In one embodiment, the number of image pixels that the user can observe based on the horizontal viewing angle, the horizontal resolution, and the stereoscopic detection result is calculated, including:

[0020] ;

[0021] Wherein, N is the number of image pixels, VA is the horizontal viewing angle, nPix is ​​the horizontal resolution, and Titmus is the first stereoscopic detection result.

[0022] In one embodiment, the training parameters are calculated based on the number of image pixels, including:

[0023] Use the number of image pixels as the training parameter; or...

[0024] The number of pixels in the image is rounded down, and the resulting integer value is used as the training parameter.

[0025] In one embodiment, after displaying the first random dot pattern on the first display screen and the second random dot pattern on the second display screen, the method further includes:

[0026] The stereoscopic vision level of the test subject is re-detected to obtain a second stereoscopic vision detection result;

[0027] The training parameters are updated based on the second stereoscopic detection result;

[0028] The second random point map is recalculated based on the updated training parameters, and the newly generated second random point map is displayed on the second display screen.

[0029] In one embodiment, generating a first random point map based on the first sequence and generating a second random point map based on the second sequence includes:

[0030] Obtain a first black image. In the first black image, take each coordinate in the first sequence as the center and fill the pixels near the center with white to generate the first random dot map.

[0031] A second black image is obtained. In the second black image, the pixels near the center of each coordinate in the second sequence are filled with white to generate the second random dot map.

[0032] Secondly, this application provides a random point stereo disparity map generation device, comprising:

[0033] The disparity map generation module is used to acquire a two-dimensional video, extract the depth information of the image frames in the two-dimensional video, and normalize the depth information based on training parameters to obtain a disparity map.

[0034] The first data sequence acquisition module is used to acquire the coordinates of random points and generate a first data sequence based on the coordinates of the random points. The first data sequence contains two columns of values: X-axis and Y-axis.

[0035] The second data sequence acquisition module is used to subtract the values ​​of the X-axis column in the first data sequence from the pixel values ​​of the corresponding coordinates in the disparity map to obtain a difference column, and generate a second data sequence based on the difference column and the Y-axis column in the first data sequence.

[0036] The random dot plot generation module is used to generate a first random dot plot based on the first sequence of numbers, and to generate a second random dot plot based on the second sequence of numbers, and to display the first random dot plot on a first display screen, and the second random dot plot on a second display screen.

[0037] Thirdly, this application provides a head-mounted display device, including: a first display screen, a second display screen, and a control module, wherein the first display screen and the second display screen are respectively connected to the control module, and the control module is used to execute the stereo disparity map generation method described in the first aspect above.

[0038] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the stereo disparity map generation method described in the first aspect.

[0039] The aforementioned method, apparatus, head-mounted display device, and storage medium for generating stereo disparity maps extract depth information from 2D video, perform disparity transformation based on training parameters and depth information to generate a disparity map, and then generate two offset random point images based on the disparity map, ultimately forming an RDS image. Since 2D video is more complex and contains richer depth information than traditional disparity maps, extracting the depth information from 2D video and converting it into a matching disparity map overcomes the limitation of traditional disparity maps containing only simple images. Furthermore, continuous depth information can be extracted from 2D video and ultimately integrated to generate continuous RDS image frames, i.e., RDS video. During training, users only need to watch and understand the RDS video content without performing other tasks, allowing users to achieve training effects while easily watching the RDS video, thus improving the user's experience in binocular stereo vision training. Moreover, this training method improves user compliance, facilitating prolonged immersive training and further enhancing the effectiveness of a single training session. Attached Figure Description

[0040] Figure 1 This is a hardware structure block diagram of the terminal of the stereo disparity map generation method in one embodiment;

[0041] Figure 2 This is a structural block diagram of a head-mounted display device in one embodiment;

[0042] Figure 3 This is a flowchart of a method for generating a stereo disparity map in one embodiment;

[0043] Figure 4 This is a flowchart illustrating the method for calculating training parameters in one embodiment;

[0044] Figure 5 This is a schematic diagram illustrating an application scenario of a stereo disparity map generation method in one embodiment.

[0045] Figure 6 for Figure 5 A flowchart of a method for generating stereo disparity maps in application scenarios;

[0046] Figure 7 This is a structural block diagram of a random point stereo disparity map generation device in one embodiment. Detailed Implementation

[0047] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0048] Unless otherwise defined, the technical or scientific terms used in this application shall have the general meaning understood by one of ordinary skill in the art to which this application pertains. Words such as “a,” “an,” “an,” “the,” “the,” and “these” used in this application do not indicate quantitative limitation and may be singular or plural. The terms “comprising,” “including,” “having,” and any variations thereof used in this application are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or device that comprises a series of steps or modules (units) is not limited to the listed steps or modules (units) but may include steps or modules (units) not listed, or may include other steps or modules (units) inherent to these processes, methods, products, or devices. Words such as “connected,” “linked,” and “coupled” used in this application are not limited to physical or mechanical connections but may include electrical connections, whether direct or indirect. “Multiple” used in this application refers to two or more. “And / or” describes the relationship between related objects, indicating that three relationships may exist; for example, “A and / or B” can represent: A alone, A and B simultaneously, and B alone. Normally, the character " / " indicates that the objects before and after it are in an "or" relationship. The terms "first," "second," "third," etc., used in this application are merely to distinguish similar objects and do not represent a specific order of objects.

[0049] The method embodiments provided in this example can be executed on a terminal, computer, or similar computing device. For example, it can run on a terminal. Figure 1This is a hardware structure block diagram of a terminal for a stereo disparity map generation method according to an embodiment of this application. For example... Figure 1 As shown, a terminal may include one or more ( Figure 1 Only one is shown in the diagram. A processor 101 and a memory 102 for storing data are also included. The processor 101 may be, but is not limited to, a microprocessor (MCU) or a programmable logic device (FPGA). The terminal may also include a transmission device 103 for communication functions and an input / output device 104. Those skilled in the art will understand that… Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the terminal described above. For example, the terminal may also include components that are larger than... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown are illustrated.

[0050] The memory 102 can be used to store computer programs, such as application software programs and modules, like the computer program corresponding to the stereo disparity map generation method in this embodiment. The processor 101 executes various functional applications and data processing by running the computer program stored in the memory 102, thereby implementing the above-described method. The memory 102 may include high-speed random access memory, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 102 may further include memory remotely located relative to the processor 101, and these remote memories can be connected to the terminal via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0051] The transmission device 103 is used to receive or send data via a network. This network includes a wireless network provided by the terminal's communication provider. In one example, the transmission device 103 includes a Network Interface Controller (NIC), which can connect to other network devices via a base station to communicate with the Internet. In another example, the transmission device 103 can be a Radio Frequency (RF) module used for wireless communication with the Internet.

[0052] In one embodiment, a head-mounted display device is provided, which may be XR (Extended Reality) glasses or an XR helmet. The XR glasses may be AR (Augmented Reality) glasses, VR (Virtual Reality) glasses, or MR (Mixed Reality) glasses. This embodiment does not impose any limitations.

[0053] Figure 2 A structural block diagram of a head-mounted display device is provided, which includes: a first display screen, a second display screen, and a control module. The first display screen and the second display screen are respectively connected to the control module, and the control module is used to execute a stereo disparity map generation method.

[0054] Parallax refers to the spatial difference in position of the same object from different viewing angles. Dual-screen split-viewing utilizes the principle of parallax, projecting images of the same scene from different perspectives onto the left and right eyes respectively, achieving the effect of simulating realistic scenes of objects at different distances. This control module can generate left and right perspective images of the same scene (e.g., two random dot images) and display them on the first and second displays respectively, with a certain parallax angle between the two images. When the user wears the head-mounted display device, the left and right eyes receive images from the left and right displays respectively. Due to the parallax angle, the user's brain combines the two images into a stereoscopic image, thus creating a sense of depth.

[0055] Those skilled in the art will understand that Figure 2 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the head-mounted display device to which the present application is applied. A specific head-mounted display device may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.

[0056] For example, in some embodiments, the head-mounted display device may also include a camera device for capturing images or videos of the external natural scene, sending the images or videos to the control module for processing, and generating left and right view images.

[0057] Figure 3 This is a flowchart of a method for generating stereo disparity maps, using this method running on... Figure 2 Taking the head-mounted display device shown as an example, the steps include:

[0058] Step S101: Acquire a two-dimensional video, extract the depth information of the image frames in the two-dimensional video, and normalize the depth information based on the training parameters to obtain a disparity map.

[0059] In this step, a camera on a head-mounted display device can be used to capture images of the surrounding natural scene, obtaining a two-dimensional video. Alternatively, the head-mounted display device can receive two-dimensional video from other terminals. After obtaining the two-dimensional video, depth information can be extracted using neural network (CNN) technology. Specifically, the distance between each pixel in each frame of the image and the observation point (e.g., the lens in the camera device) is calculated. Each pixel is then assigned a value from 0 to 1 based on the distance, resulting in depth information (i.e., a depth map). The larger the distance, the smaller the value assigned; conversely, the smaller the distance, the larger the value assigned. The depth information includes pixel values; pixels with values ​​closer to 0 are farther from the observer (a pixel with a value of 0 is considered part of the background), while pixels with values ​​closer to 1 are closer to the observer.

[0060] Assuming the training parameter is valued at n, when normalizing depth information based on the training parameter, the training parameter is multiplied by the pixel value to expand the pixel value to the range of 0 to n, ultimately obtaining a disparity map. The value of each pixel in the disparity map indicates the size of the pixel to be offset. The training parameter is used to adjust the training difficulty; the larger the training parameter, the easier the training.

[0061] Step S102: Obtain the coordinates of a random point and generate a first sequence based on the coordinates of the random point. The first sequence contains two columns of values: the X-axis and the Y-axis.

[0062] Assuming the coordinates of the first random point are (x1, y1), then x1 ranges from 1 to m1, and y1 ranges from 1 to m2, where m1 is the number of horizontal pixels in the original image, and m2 is the number of vertical pixels in the original image. Assuming there are m first random points, then there are a total of m random point coordinates. Optionally, the number of random points can be greater than 10,000 to ensure as even a uniform coverage of the image as possible. Therefore, a 2×m first sequence is obtained, where the first column represents the X-axis coordinates, the second column represents the Y-axis coordinates, and the number of rows in the first sequence represents the number of random points.

[0063] Step S103: Subtract the values ​​of the X-axis column in the first sequence from the pixel values ​​of the corresponding coordinates in the disparity map to obtain the difference column, and generate the second sequence based on the difference column and the Y-axis column in the first sequence.

[0064] The second random point is obtained by offsetting the first random point. Assuming the coordinates of the second random point are (x2, y2), then x2 = x1 - d, y2 = y1.

[0065] Step S104: Generate a first random dot pattern based on the first sequence and a second random dot pattern based on the second sequence, and display the first random dot pattern on the first display screen and the second random dot pattern on the second display screen.

[0066] Specifically, a first black image is obtained. Within the first black image, a circular area is defined with each coordinate in the first sequence as the center and r as the radius. Within this circular area, pixels near the center are filled with white to generate a first random dot map. A second black image is obtained. Within the second black image, a circular area is defined with each coordinate in the second sequence as the center and r as the radius. Within this circular area, pixels near the center are filled with white to generate a second random dot map.

[0067] In this step, the first and second random dot maps constitute a RandomDots Stereo (RDS) image. Stereo vision in the RDS image is achieved through slight offsets (parallax) in the positions of random dots in certain regions of the two images. Specifically, the RDS image hides two-dimensional image information through random dots. For any single eye, the user sees a random dot image without any information, thus eliminating interference from single-eye cues. This allows the user to perceive the three-dimensional contour information of objects in the image only when simultaneously integrating binocular input. Therefore, RDS images can be used for training patients with binocular integration disorders. The offset region between the first and second random dot maps will exhibit a raised or recessed stereoscopic contour. When the user wears a head-mounted display device, the left and right eyes receive random dot maps from the left and right displays respectively. Due to the parallax angle, the user's brain combines the two images into a single stereoscopic image, generating a sense of depth, and identifying the content of objects based on their three-dimensional contour information.

[0068] In steps S101 to S104 above, due to the complexity of the offset region design, traditional disparity maps can often only represent simple graphics in three dimensions. This embodiment extracts depth information from two-dimensional video, performs disparity transformation based on training parameters and depth information to generate a disparity map, and then generates two offset random point images based on the disparity map, ultimately forming an RDS image. Since two-dimensional video is more complex and contains richer depth information than traditional disparity maps, extracting the depth information from the two-dimensional video and converting it into a matching disparity map overcomes the limitation of traditional disparity maps containing only simple images. Furthermore, continuous depth information can be extracted from the two-dimensional video, ultimately integrating it to generate continuous RDS image frames, i.e., RDS video. During training, users only need to watch and understand the RDS video content without completing other tasks, allowing users to easily achieve training results while watching the RDS video, improving the user's experience in binocular stereoscopic training. Moreover, this training method improves user compliance, facilitating long-term immersive training, thereby further enhancing the effectiveness of a single training session.

[0069] In one embodiment, Figure 4 A flowchart is provided for a method of calculating training parameters. Before normalizing the depth information based on the training parameters, the following steps are also included:

[0070] Step S201: Obtain the horizontal viewing angle of the first display screen or the second display screen, obtain the horizontal resolution of the first display screen or the second display screen, and obtain the user's first stereoscopic detection result.

[0071] In the Titmus (stereometry) test, a polarizing filter is used to achieve the effect of visual separation. The user's stereometry ability is tested based on images from different viewing angles, such as 80°, 40°, 20°, 14°, 10°, 8°, 6°, 5°, and 4° (in seconds). The user completes the test by sequentially judging whether the image has a stereoscopic effect, with the difficulty increasing as the viewing angle decreases. The better the user can achieve stereoscopic vision with a smaller parallax, the higher their stereoscopic vision level is considered, indicating better binocular integration. In this step, the user's stereometry test result is described as one of the following: 80°, 40°, 20°, 14°, 10°, 8°, 6°, 5°, or 4°.

[0072] Step S202: Based on the horizontal viewing angle, horizontal resolution, and the first stereoscopic detection result, calculate the number of image pixels that the user can observe from the first display screen or the second display screen. The specific calculation formula is as follows:

[0073] ;

[0074] Where N is the number of image pixels that the user can observe based on the first or second display screen, VA is the horizontal viewing angle, nPix is ​​the horizontal resolution, and Titmus is the stereoscopic detection result.

[0075] For example, assuming that the horizontal viewing angle of both the first and second displays is 50 degrees and the horizontal resolution is 1920 pixels, and taking a stereoscopic detection result of 800 seconds as an example, substituting into the above formula, we get:

[0076] ;

[0077] Taking two decimal places, we get N = 4 pixel differences. The corresponding number of image pixels N for each of the above viewpoints are 4, 2, 1, 0.7, 0.5, 0.4, 0.3, 0.25, and 0.2, respectively.

[0078] Step S203: Calculate the training parameters based on the number of image pixels.

[0079] Use the number of image pixels as a training parameter; or, round the number of image pixels and use the resulting integer value as a training parameter.

[0080] For example, since N is 4, 2, 1, 0.7, 0.5, 0.4, 0.3, 0.25, and 0.2 respectively, then n also takes the values ​​4, 2, 1, 0.7, 0.5, 0.4, 0.3, 0.25, and 0.2 respectively. In some embodiments, since the pixels are discretely distributed, their values ​​must be integers, and N reflects the maximum disparity in subsequent calculations, not the average level. Therefore, to make N rounded and reduce training difficulty, n is set to 5 × N, where n takes the values ​​20, 10, 5, 4, 3, 2, 2, 1, and 1 respectively.

[0081] In this embodiment, based on the relationship between parallax and depth, a parallax of 0 indicates that random points on both the left and right images are in the same position, which can be used as background for depth recognition. When the parallax is greater than 0, the offset random points will move as a whole towards the direction of binocular convergence, ultimately resulting in a convex depth perception. Conversely, when the parallax is less than 0, the offset random points will move as a whole in the opposite direction of binocular convergence, ultimately resulting in a concave depth perception. Simultaneously, the absolute value of the parallax reflects the degree of convexity or concavity; the larger the absolute value, the farther the perceived contour is from the background. For users with insufficient binocular integration ability, recognizing parallax with small absolute values ​​often presents some difficulty. Therefore, this embodiment combines the fixed parameters of the head-mounted display device (horizontal viewing angle and horizontal resolution) and the user's actual situation (stereoscopic vision detection results) to calculate the training parameter n. Based on the training parameter n, the parallax magnitude is controlled, thereby setting the optimal training difficulty for the user's actual situation to match the actual stereoscopic vision ability of different users.

[0082] In one embodiment, after displaying the first random dot pattern on the first display screen and the second random dot pattern on the second display screen, the method further includes:

[0083] The stereoscopic vision level of the test object is re-detected to obtain a second stereoscopic vision detection result; the training parameters are updated based on the second stereoscopic vision detection result; a second random dot map is recalculated and generated based on the updated training parameters, and the newly generated second random dot map is displayed on the second display screen.

[0084] For example, the training parameter 'n' should be set to gradually increase the training difficulty as the user's stereopsis level improves, in order to reduce the ceiling effect of training with constant parameters. For instance, for a patient with an initial stereopsis level of 800 seconds of visual field, the initial training parameter 'n' is 20. Subsequently, if the user's ability in the Titmus test improves to 400 seconds of visual field, the subsequent matching training parameter 'n' can be increased to 10. By updating the training parameters, the training difficulty can gradually increase as the user's stereopsis level improves, achieving a gradual training effect and helping to reduce the ceiling effect of training with constant parameters.

[0085] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0086] In one embodiment, Figure 5 A schematic diagram illustrating an application scenario for a method for generating stereo disparity maps is provided. Figure 6 A flowchart of the method for generating stereo disparity maps in this application scenario, such as... Figure 6 As shown, the process includes the following steps:

[0087] Step S301: Obtain the user's current stereoscopic detection result.

[0088] Step S302: Determine whether the current stereo vision detection result is less than or equal to the threshold. If yes, end the process; otherwise, proceed to step S303. Specifically, if the current stereo vision detection result is less than or equal to the threshold, it means the user's stereo vision level has reached the expected level, and training can stop. If the current stereo vision detection result is higher than the threshold, it means the user's stereo vision level is low and has not reached the expected level, requiring further training.

[0089] Step S303: Obtain the horizontal viewing angle of the first display screen or the second display screen, obtain the horizontal resolution of the first display screen or the second display screen, and calculate the number of image pixels that the user can observe based on the first display screen or the second display screen according to the horizontal viewing angle, the horizontal resolution and the current stereoscopic detection results.

[0090] Step S304: Calculate the training parameters based on the number of image pixels. The calculation formula is as follows:

[0091] ;

[0092] Where N is the number of image pixels that the user can observe from the first or second display screen, VA is the horizontal viewing angle, nPix is ​​the horizontal resolution, and Titmus is the first stereoscopic detection result. The number of image pixels can be used as a training parameter; alternatively, the number of image pixels can be rounded down, and the resulting integer value can be used as a training parameter.

[0093] Step S305: Acquire a two-dimensional video, extract the depth information of the image frames in the two-dimensional video, and normalize the depth information based on the training parameters to obtain a disparity map.

[0094] Calculate the distance between each pixel in the image frame and the observation point; assign a value to each pixel in the range of 0 to 1 according to the distance to obtain depth information; where the larger the distance, the smaller the value assigned, and the smaller the distance, the larger the value assigned.

[0095] Step S306: Obtain the coordinates of a random point and generate a first sequence based on the coordinates of the random point. The first sequence contains two columns of values: the X-axis and the Y-axis.

[0096] Step S307: Subtract the values ​​of the X-axis column in the first sequence from the pixel values ​​of the corresponding coordinates in the disparity map to obtain the difference column, and generate the second sequence based on the difference column and the Y-axis column in the first sequence.

[0097] Step S308: Generate a first random dot plot based on the first sequence and a second random dot plot based on the second sequence. Display the first random dot plot on the first display screen and the second random dot plot on the second display screen. After this round of training, return to step S301 to re-detect the stereo vision level of the test object. Update the training parameters based on the new stereo vision detection results, and then recalculate and generate the second random dot plot based on the updated training parameters. Display the newly generated second random dot plot on the second display screen. The random dot plot generation method is as follows:

[0098] Obtain a first black image. In the first black image, fill the pixels near the center of each coordinate in the first sequence with white to generate a first random dot map. Obtain a second black image. In the second black image, fill the pixels near the center of each coordinate in the second sequence with white to generate a second random dot map.

[0099] In steps S301 to 308 above, since 2D videos are more complex and contain richer depth information than traditional disparity maps, extracting the depth information from 2D videos and converting it into a matching disparity map can overcome the limitation of traditional disparity maps containing only simple images. Furthermore, continuous depth information can be extracted from 2D videos, ultimately integrating and generating continuous RDS image frames, i.e., RDS videos. During training, users only need to watch and understand the RDS video content without completing other tasks, allowing them to achieve training effects while easily watching the RDS video, thus improving the user's experience in binocular stereoscopic training. Moreover, this training method increases user compliance, facilitating prolonged immersive training and improving the effectiveness of single training sessions. Further, by updating training parameters and controlling the disparity magnitude based on these parameters, the optimal training difficulty can be set according to the user's actual situation to match the actual stereoscopic vision ability of different users. The training difficulty can gradually increase as the user's stereoscopic vision level improves, achieving a progressive training effect and reducing the ceiling effect under constant parameter training.

[0100] Based on the same inventive concept, this application also provides a random point stereo disparity map generation apparatus for implementing the above-described method. This apparatus is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below can refer to combinations of software and / or hardware that perform a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0101] Figure 7 This is a structural block diagram of the random point stereo disparity map generation device in this embodiment, as shown below. Figure 7 As shown, the device includes:

[0102] The disparity map generation module is used to acquire two-dimensional video, extract the depth information of image frames in the two-dimensional video, and normalize the depth information based on training parameters to obtain a disparity map.

[0103] The first data sequence acquisition module is used to acquire the coordinates of random points and generate the first data sequence based on the coordinates of random points. The first data sequence contains two columns of values: X-axis and Y-axis.

[0104] The second data acquisition module is used to subtract the values ​​of the X-axis column in the first data sequence from the pixel values ​​of the corresponding coordinates in the disparity map to obtain the difference column, and generate the second data sequence based on the difference column and the Y-axis column in the first data sequence.

[0105] The random dot plot generation module is used to generate a first random dot plot based on a first sequence of numbers, and to generate a second random dot plot based on a second sequence of numbers, and to display the first random dot plot on a first display screen, and the second random dot plot on a second display screen.

[0106] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0107] Furthermore, in conjunction with the random point stereo disparity map generation method provided in the above embodiments, this embodiment can also provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the random point stereo disparity map generation methods in the above embodiments.

[0108] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0109] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0110] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0111] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A method for generating a stereo disparity map, characterized in that, Applied to a head-mounted display device, the head-mounted display device including a first display screen and a second display screen, the method includes: A two-dimensional video is acquired, the depth information of the image frames in the two-dimensional video is extracted, and the depth information is normalized based on the training parameters to obtain a disparity map. Obtain the coordinates of a random point, and generate a first sequence based on the coordinates of the random point. The first sequence contains two columns of values: the X-axis and the Y-axis. The difference between the values ​​in the X-axis column of the first sequence and the pixel values ​​at the corresponding coordinates in the disparity map is used to obtain the difference column. Then, a second sequence is generated based on the difference column and the Y-axis column of the first sequence. A first random dot pattern is generated based on the first sequence of numbers, and a second random dot pattern is generated based on the second sequence of numbers. The first random dot pattern is displayed on the first display screen, and the second random dot pattern is displayed on the second display screen.

2. The method for generating a stereo disparity map according to claim 1, characterized in that, Extracting depth information from image frames in the two-dimensional video includes: Calculate the distance between each pixel in the image frame and the observation point; The depth information is obtained by assigning a value to each pixel in the range of 0 to 1 according to the distance; wherein, the larger the distance, the smaller the assigned value, and the smaller the distance, the larger the assigned value.

3. The method for generating a stereo disparity map according to claim 1, characterized in that, Before normalizing the depth information based on training parameters, the method further includes: Obtain the horizontal viewing angle of the first display screen or the second display screen, obtain the horizontal resolution of the first display screen or the second display screen, and obtain the user's first stereoscopic vision detection result; Based on the horizontal viewing angle, the horizontal resolution, and the first stereoscopic detection result, the number of image pixels that the user can observe based on the first display screen or the second display screen is calculated; The training parameters are calculated based on the number of pixels in the image.

4. The method for generating a stereo disparity map according to claim 3, characterized in that, Based on the horizontal viewing angle, the horizontal resolution, and the stereoscopic detection results, the number of image pixels that the user can observe based on the first display screen or the second display screen is calculated, including: ; Wherein, N is the number of image pixels, VA is the horizontal viewing angle, nPix is ​​the horizontal resolution, and Titmus is the first stereoscopic detection result.

5. The method for generating a stereo disparity map according to claim 3, characterized in that, The training parameters are calculated based on the number of pixels in the image, including: Use the number of image pixels as the training parameter; or... The number of pixels in the image is rounded down, and the resulting integer value is used as the training parameter.

6. The method for generating a stereo disparity map according to claim 3, characterized in that, After displaying the first random dot pattern on the first display screen and the second random dot pattern on the second display screen, the method further includes: The stereoscopic vision level of the test subject is re-detected to obtain a second stereoscopic vision detection result; The training parameters are updated based on the second stereoscopic detection result; The second random point map is recalculated based on the updated training parameters, and the newly generated second random point map is displayed on the second display screen.

7. The method for generating a stereo disparity map according to claim 1, characterized in that, Generating a first random point map based on the first sequence and generating a second random point map based on the second sequence include: Obtain a first black image. In the first black image, take each coordinate in the first sequence as the center and fill the pixels near the center with white to generate the first random dot map. A second black image is obtained. In the second black image, the pixels near the center of each coordinate in the second sequence are filled with white to generate the second random dot map.

8. A random point stereo disparity map generation device, characterized in that, include: The disparity map generation module is used to acquire a two-dimensional video, extract the depth information of the image frames in the two-dimensional video, and normalize the depth information based on training parameters to obtain a disparity map. The first data sequence acquisition module is used to acquire the coordinates of random points and generate a first data sequence based on the coordinates of the random points. The first data sequence contains two columns of values: X-axis and Y-axis. The second data sequence acquisition module is used to subtract the values ​​of the X-axis column in the first data sequence from the pixel values ​​of the corresponding coordinates in the disparity map to obtain a difference column, and generate a second data sequence based on the difference column and the Y-axis column in the first data sequence. The random dot plot generation module is used to generate a first random dot plot based on the first sequence of numbers, and to generate a second random dot plot based on the second sequence of numbers, and to display the first random dot plot on a first display screen, and the second random dot plot on a second display screen.

9. A head-mounted display device, characterized in that, include: The system comprises a first display screen, a second display screen, and a control module, wherein the first display screen and the second display screen are respectively connected to the control module, and the control module is used to execute the stereo disparity map generation method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Visual field auxiliary device cooperating with ARVR head-mounted typoscope

    CN215821381U

  • Imaging apparatus, image display apparatus and image recording and / or reproducing apparatus

    US6177952B1