Stereoscopic disparity map generation method and device, head-mounted display equipment and storage medium

By generating disparity maps based on 2D videos and displaying random point images, the problem of poor training effect of binocular stereoscopic vision function in strabismic amblyopia is solved, improving the user's training experience and compliance, and achieving efficient stereoscopic vision training.

CN121348587AActive Publication Date: 2026-01-16HANGZHOU FOCUSIGHT INTELLIGENT TECHNOLOGY CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511364987.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-22
Publication Date
2026-01-16
Estimated Expiration
2045-09-22

Smart Images

  • Figure CN121348587A_ABST
    Figure CN121348587A_ABST
Patent Text Reader

Abstract

The invention relates to a stereo disparity map generation method and device, head-mounted display equipment and a storage medium, and the method comprises the steps: obtaining a two-dimensional video, extracting the depth information of an image frame in the two-dimensional video, carrying out the normalization processing of the depth information based on a training parameter, and obtaining a disparity map; random point coordinates are obtained, a first sequence is generated according to the random point coordinates, and the first sequence comprises X-axis and Y-axis numerical values; subtracting the numerical value of the X-axis column in the first sequence from the pixel value of the corresponding coordinate in the disparity map to obtain a difference value column, and generating a second sequence according to the difference value column and the Y-axis column in the first sequence; generating a first random dot diagram according to the first sequence, generating a second random dot diagram according to the second sequence, displaying the first random dot diagram on the first display screen, and displaying the second random dot diagram on the second display screen; the experience feeling of binocular stereoscopic vision training of a user can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of stereoscopic disparity map generation, and in particular, to a stereoscopic disparity map generation method and device, a head-mounted display device, and a storage medium. BACKGROUND

[0002] Unlike anisometropic amblyopia with significantly reduced monocular ability, one of the characteristics of strabismic amblyopia is that the processing ability of the amblyopic eye is not significantly reduced when used alone, and may even be comparable to the ability of the healthy eye. However, when both eyes are used at the same time, the amblyopic eye is significantly inhibited. Therefore, unlike the monocular input improvement strategy that can effectively act on anisometropic amblyopia, how to find a strategy that can improve the integration ability between the two eyes under binocular conditions, thereby increasing the input proportion of the amblyopic eye of strabismic amblyopia, has become a problem that needs to be further solved.

[0003] Stereoscopic vision function is one of the important functions of the visual system, and the realization of stereoscopic vision function includes monocular cues (including relative size, occlusion, gradient, and perspective) and binocular cues (including binocular convergence and binocular disparity). Among them, the realization of stereoscopic vision through binocular disparity is one of the few functions in the visual system that must rely on binocular input to complete. Therefore, the training of stereoscopic vision through binocular disparity has become a better choice to improve the binocular integration ability of strabismic amblyopia.

[0004] The related art usually uses a disparity map to train the binocular stereoscopic vision function of a user. The disparity map is composed of two images with the same main content, and some parts in the two images have a slight offset. When viewing with both eyes, the offset area will show a convex or concave stereoscopic contour, and the user understands the picture content by perceiving these stereoscopic contours, and at the same time, some question and answer tasks are designed to achieve the purpose of training stereoscopic vision. For example, the user is asked to identify the orientation of the objects, letters, and contours displayed in the image.

[0005] However, the disparity map of the related art can only present some simple graphics, and some question and answer tasks are needed to understand the training results of the user, the training process is boring, the user experience is low, and ultimately leads to a decrease in user compliance during training, and cannot be performed for a long time, thereby further reducing the effect that can be achieved in a single training. SUMMARY

[0006] Therefore, it is necessary to provide a stereoscopic disparity map generation method and device, a head-mounted display device, and a storage medium that can improve the user's binocular stereoscopic vision training experience.

[0007] In a first aspect, the present application provides a stereoscopic disparity map generation method. The method comprises:

[0008] acquiring a two-dimensional video, extracting depth information of image frames in the two-dimensional video, and normalizing the depth information based on training parameters to obtain a parallax map;

[0009] acquiring random point coordinates, generating a first series based on the random point coordinates, the first series including two columns of values of X-axis and Y-axis;

[0010] subtracting values of the X-axis column in the first series from pixel values of corresponding coordinates in the parallax map to obtain a difference column, and generating a second series based on the difference column and the Y-axis column in the first series;

[0011] generating a first random point map based on the first series and a second random point map based on the second series, and displaying the first random point map on the first display screen and the second random point map on the second display screen.

[0012] In one embodiment, the extracting of the depth information of the image frames in the two-dimensional video includes:

[0013] calculating distances between each pixel point in the image frames and an observation point;

[0014] assigning values to each pixel point in the range of 0 to 1 based on the distances to obtain the depth information; wherein the greater the distance, the smaller the value, and the smaller the distance, the greater the value.

[0015] In one embodiment, before the normalizing of the depth information based on the training parameters, the method further includes:

[0016] acquiring a horizontal viewing angle of the first display screen or the second display screen, acquiring a horizontal resolution of the first display screen or the second display screen, and acquiring a first stereoscopic vision detection result of a user;

[0017] calculating the number of image pixels that the user can observe based on the first display screen or the second display screen based on the horizontal viewing angle, the horizontal resolution, and the first stereoscopic vision detection result;

[0018] calculating the training parameters based on the number of image pixels.

[0019] In one embodiment, the calculating of the number of image pixels that the user can observe based on the first display screen or the second display screen based on the horizontal viewing angle, the horizontal resolution, and the stereoscopic vision detection result includes:

[0020] ;

[0021] Wherein, N is the image pixel number, VA is the horizontal visual angle, nPix is the horizontal resolution, and Titmus is the first stereoscopic vision detection result.

[0022] In one embodiment, the training parameter is calculated according to the image pixel number, including:

[0023] The image pixel number is taken as the training parameter; or,

[0024] The image pixel number is rounded, and the obtained integer value is taken as the training parameter.

[0025] In one embodiment, after the first random dot pattern is displayed on the first display screen and the second random dot pattern is displayed on the second display screen, the method further includes:

[0026] Re-detecting the stereoscopic vision level of the test object to obtain a second stereoscopic vision detection result;

[0027] Updating the training parameter according to the second stereoscopic vision detection result;

[0028] According to the updated training parameter, a second random dot pattern is recalculated and generated, and the newly generated second random dot pattern is displayed on the second display screen.

[0029] In one embodiment, the first random dot pattern is generated according to the first number sequence, and the second random dot pattern is generated according to the second number sequence, including:

[0030] Obtaining a first black image, in the first black image, the pixels near the center are filled with white color with each coordinate in the first number sequence as the center, and the first random dot pattern is generated;

[0031] Obtaining a second black image, in the second black image, the pixels near the center are filled with white color with each coordinate in the second number sequence as the center, and the second random dot pattern is generated.

[0032] In a second aspect, the application provides a random dot stereoscopic disparity map generation device, including:

[0033] A disparity map generation module is configured to obtain a two-dimensional video, extract depth information of an image frame in the two-dimensional video, and perform normalization processing on the depth information based on a training parameter to obtain a disparity map.

[0034] A first number sequence acquisition module is configured to obtain random dot coordinates, generate a first number sequence according to the random dot coordinates, and the first number sequence includes two columns of number values of X-axis and Y-axis.

[0035] a second number sequence acquisition module, configured to subtract a value of an X-axis column in the first number sequence from a pixel value of a corresponding coordinate in the parallax map to obtain a difference value column, and generate a second number sequence according to the difference value column and a Y-axis column in the first number sequence;

[0036] a random dot map generation module, configured to generate a first random dot map according to the first number sequence and generate a second random dot map according to the second number sequence, and display the first random dot map on a first display screen and display the second random dot map on a second display screen.

[0037] In a third aspect, the present application provides a head-mounted display device, comprising: a first display screen, a second display screen and a control module, wherein the first display screen and the second display screen are respectively connected with the control module, and the control module is configured to execute the stereoscopic parallax map generation method in the first aspect.

[0038] In a fourth aspect, the present application further provides a computer readable storage medium, having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the stereoscopic parallax map generation method in the first aspect.

[0039] The stereoscopic parallax map generation method, device, head-mounted display device and storage medium described above extract depth information of a two-dimensional video, perform parallax conversion according to training parameters and the depth information, generate a parallax map, generate two random dot images with offsets based on the parallax map, and finally form an RDS image. Since the two-dimensional video is more complex and has more rich depth information than the traditional parallax map, the two-dimensional video is converted into a parallax map matched therewith by extracting the depth information of the two-dimensional video, thereby breaking through the limitation that the traditional parallax map only contains simple images. Moreover, continuous depth information can be extracted based on the two-dimensional video, and finally a continuous RDS image frame, i.e., an RDS video, is generated. In the training process, the user only needs to watch and understand the RDS video content, without the need to complete other tasks. The user can watch the RDS video easily while achieving the training effect, thereby improving the experience of the user's stereoscopic vision training. Moreover, the training method improves the user's compliance, is conducive to long-time immersive training, and further improves the effect achieved in a single training. BRIEF DESCRIPTION OF DRAWINGS

[0040] Figure 1 a hardware structure block diagram of a terminal for the stereoscopic parallax map generation method in an embodiment;

[0041] Figure 2 a structure block diagram of a head-mounted display device in an embodiment;

[0042] Figure 3 a flowchart of the stereoscopic parallax map generation method in an embodiment;

[0043] Figure 4 a flow chart of a method for calculating training parameters in one embodiment;

[0044] Figure 5 a schematic diagram of an application scenario of a method for generating a stereo disparity map in one embodiment;

[0045] Figure 6 a method for generating a stereo disparity map in an application scenario; Figure 5 a flow chart of a method for generating a stereo disparity map in an application scenario;

[0046] Figure 7 a structural block diagram of a random point stereo disparity map generation device in one embodiment. DETAILED DESCRIPTION

[0047] In order to make the purposes, technical solutions and advantages of the present application clearer, the present application is further described in detail below with reference to the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and do not limit the present application.

[0048] Unless otherwise defined, technical terms or scientific terms used in the present application shall have the general meaning understood by a person with ordinary skill in the art to which the present application belongs. In the present application, the terms "one", "a", "an", "the", "these", and similar words do not represent a quantitative limitation, but can be singular or plural. In the present application, the terms "include", "contain", "have" and any variants thereof have the purpose of covering non-exclusive inclusion; for example, a process, method and system, product or device containing a series of steps or modules (units) are not limited to the listed steps or modules (units), but can include steps or modules (units) not listed, or can include other steps or modules (units) inherent to the process, method, product or device. In the present application, the terms "connected", "connected", "coupled" and similar words are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. In the present application, "multiple" means two or more. The association between the associated objects is described by the term "and / or", which means that there can be three relationships, for example, "A and / or B" can mean that A exists alone, A and B exist together, and B exists alone. In general, the character " / " represents an "or" relationship between the associated objects. In the present application, the terms "first", "second", "third" and the like are only used to distinguish similar objects, and do not represent a specific order of the objects.

[0049] The method embodiments provided in the present embodiment can be executed in a terminal, a computer or a similar computing device. For example, the method embodiments can be run on a terminal, Figure 1is a hardware structure block diagram of a terminal of a stereoscopic parallax map generation method of an embodiment of the present application. As shown in Figure 1 , the terminal can include one or more (only one is shown in Figure 1 ) processors 101 and a memory 102 for storing data, wherein the processor 101 can include but is not limited to a processing device such as a microprocessor MCU or a programmable logic device FPGA. The terminal can also include a transmission device 103 for communication function and an input / output device 104. Those skilled in the art can understand that Figure 1 the structure shown is only schematic, which does not limit the structure of the terminal. For example, the terminal can include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.

[0050] The memory 102 can be used to store computer programs, such as software programs of application software and modules, such as the computer program corresponding to the stereoscopic parallax map generation method in the embodiment. The processor 101 performs various functional applications and data processing by running the computer program stored in the memory 102, that is, implements the method described above. The memory 102 can include a high-speed random access memory, and can also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some examples, the memory 102 can further include a memory remotely arranged with respect to the processor 101, which can be connected to the terminal through a network. Examples of the network include but are not limited to the Internet, an intranet, a local area network, a mobile communication network, and a combination thereof.

[0051] The transmission device 103 is used to receive or send data via a network. The network includes a wireless network provided by a communication provider of the terminal. In one example, the transmission device 103 includes a network adapter (Network Interface Controller, NIC) which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 103 can be a radio frequency (Radio Frequency, RF) module which is used to communicate with the Internet in a wireless manner.

[0052] In one embodiment, a head-mounted display device is provided, which can be an XR (Extended Reality) glasses or an XR helmet, wherein the XR glasses can be AR (Augmented Reality) glasses, VR (Virtual Reality) glasses or MR (Mixed Reality) glasses, which are not limited in the embodiment.

[0053] Figure 2 A structural diagram of a head-mounted display device is provided, which comprises a first display screen, a second display screen and a control module, the first display screen and the second display screen are connected with the control module respectively, and the control module is used for executing a stereoscopic parallax map generation method.

[0054] Parallax refers to the spatial position difference of the same object under different viewing angles. The double-screen split vision utilizes the parallax principle to project the images of different viewing angles of the same scene to the left and right eyes respectively, so as to realize the effect of simulating the real near and far object scene. The control module can generate the left and right viewing angle images (for example, two random dot maps) of the same scene and display them on the first display screen and the second display screen respectively, and there is a certain parallax angle between the two images. When the user wears the head-mounted display device, the left and right eyes receive the images from the left and right display screens respectively, and due to the existence of the parallax angle, the user's brain will synthesize two images into a stereoscopic image, thereby producing a depth perception.

[0055] Those skilled in the art can understand that, Figure 2 The structure shown in the figure is only a block diagram of part of the structure related to the scheme of the present application, and does not constitute a limitation on the head-mounted display device to which the scheme of the present application is applied. The specific head-mounted display device can include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0056] For example, in some embodiments, the head-mounted display device can further comprise a camera for collecting images or videos of natural scenes in the outside world, and sending the images or videos to the control module for processing to generate left and right viewing angle images.

[0057] Figure 3 A flowchart of a stereoscopic parallax map generation method is provided, and the method is run in the control module of the head-mounted display device. Figure 2 Taking the head-mounted display device shown in the figure as an example, the method comprises the following steps:

[0058] In step S101, a two-dimensional video is obtained, the depth information of the image frames in the two-dimensional video is extracted, and the depth information is normalized based on the training parameters to obtain a parallax map.

[0059] In this step, the natural scene outside can be photographed by a camera on the head-mounted display device to obtain a two-dimensional video. Alternatively, the head-mounted display device receives a two-dimensional video from another terminal. After obtaining the two-dimensional video, depth information can be extracted by a neural network (CNN) technology, specifically, the distance between each pixel point in each frame image and the observation point (such as the lens in the camera) is calculated; each pixel point is assigned a value in the range of 0 to 1 according to the distance to obtain depth information (i.e. depth information map); the greater the distance, the smaller the value, and the smaller the distance, the greater the value. The depth information contains pixel values, and the pixel value closer to 0 indicates that the pixel point is farther away from the observer (the pixel value of 0 is completely as background), and the pixel value closer to 1 indicates that the pixel point is closer to the observer.

[0060] Suppose the training parameter value is n, when normalizing the depth information based on the training parameter, the training parameter is multiplied by the pixel value to expand the pixel value to the range of 0 to n, and finally obtain the disparity map, the value of each pixel point in the disparity map indicates the pixel size that needs to be offset. The training parameter is used to adjust the training difficulty, and the larger the training parameter, the smaller the training difficulty.

[0061] Step S102, obtain random point coordinates, generate a first number sequence according to the random point coordinates, and the first number sequence contains two columns of values of X-axis and Y-axis.

[0062] Suppose the first random point coordinates are (x1, y1), the value range of x1 is 1 to m1, and the value range of y1 is 1 to m2, wherein m1 is the number of horizontal pixels in the original image, and m2 is the number of vertical pixels in the original image. Suppose that m first random points are generated, then there are m random point coordinates. Optionally, the number of random points is greater than 10,000, so as to cover as many corners of the image as possible. Therefore, a first number sequence of 2xm is obtained, wherein the first column is the X-axis coordinate, the second column is the Y-axis coordinate, and the number of rows of the first number sequence is the number of random points.

[0063] Step S103, subtract the value of the X-axis column in the first number sequence from the pixel value of the corresponding coordinate in the disparity map to obtain a difference value column, and generate a second number sequence according to the difference value column and the Y-axis column in the first number sequence.

[0064] The second random point is obtained by offsetting the first random point, and suppose the second random point coordinates are (x2, y2), then x2=x1-d, y2=y1.

[0065] Step S104, generate a first random point image according to the first number sequence, and generate a second random point image according to the second number sequence, and display the first random point image on the first display screen and the second random point image on the second display screen.

[0066] Specifically, a first black image is obtained. Within the first black image, a circular area is defined with each coordinate in the first sequence as the center and r as the radius. Within this circular area, pixels near the center are filled with white to generate a first random dot map. A second black image is obtained. Within the second black image, a circular area is defined with each coordinate in the second sequence as the center and r as the radius. Within this circular area, pixels near the center are filled with white to generate a second random dot map.

[0067] In this step, the first and second random dot maps constitute a RandomDots Stereo (RDS) image. Stereo vision in the RDS image is achieved through slight offsets (parallax) in the positions of random dots in certain regions of the two images. Specifically, the RDS image hides two-dimensional image information through random dots. For any single eye, the user sees a random dot image without any information, thus eliminating interference from single-eye cues. This allows the user to perceive the three-dimensional contour information of objects in the image only when simultaneously integrating binocular input. Therefore, RDS images can be used for training patients with binocular integration disorders. The offset region between the first and second random dot maps will exhibit a raised or recessed stereoscopic contour. When the user wears a head-mounted display device, the left and right eyes receive random dot maps from the left and right displays respectively. Due to the parallax angle, the user's brain combines the two images into a single stereoscopic image, generating a sense of depth, and identifying the content of objects based on their three-dimensional contour information.

[0068] In steps S101 to S104 above, due to the complexity of the offset region design, traditional disparity maps can often only represent simple graphics in three dimensions. This embodiment extracts depth information from two-dimensional video, performs disparity transformation based on training parameters and depth information to generate a disparity map, and then generates two offset random point images based on the disparity map, ultimately forming an RDS image. Since two-dimensional video is more complex and contains richer depth information than traditional disparity maps, extracting the depth information from the two-dimensional video and converting it into a matching disparity map overcomes the limitation of traditional disparity maps containing only simple images. Furthermore, continuous depth information can be extracted from the two-dimensional video, ultimately integrating it to generate continuous RDS image frames, i.e., RDS video. During training, users only need to watch and understand the RDS video content without completing other tasks, allowing users to easily achieve training results while watching the RDS video, improving the user's experience in binocular stereoscopic training. Moreover, this training method improves user compliance, facilitating long-term immersive training, thereby further enhancing the effectiveness of a single training session.

[0069] In one embodiment, Figure 4 A flowchart is provided for a method of calculating training parameters. Before normalizing the depth information based on the training parameters, the following steps are also included:

[0070] Step S201: Obtain the horizontal viewing angle of the first display screen or the second display screen, obtain the horizontal resolution of the first display screen or the second display screen, and obtain the user's first stereoscopic detection result.

[0071] In the Titmus (stereometry) test, a polarizing filter is used to achieve the effect of visual separation. The user's stereometry ability is tested based on images from different viewing angles, such as 80°, 40°, 20°, 14°, 10°, 8°, 6°, 5°, and 4° (in seconds). The user completes the test by sequentially judging whether the image has a stereoscopic effect, with the difficulty increasing as the viewing angle decreases. The better the user can achieve stereoscopic vision with a smaller parallax, the higher their stereoscopic vision level is considered, indicating better binocular integration. In this step, the user's stereometry test result is described as one of the following: 80°, 40°, 20°, 14°, 10°, 8°, 6°, 5°, or 4°.

[0072] Step S202: Based on the horizontal viewing angle, horizontal resolution, and the first stereoscopic detection result, calculate the number of image pixels that the user can observe from the first display screen or the second display screen. The specific calculation formula is as follows:

[0073] ;

[0074] Where N is the number of image pixels that the user can observe based on the first or second display screen, VA is the horizontal viewing angle, nPix is ​​the horizontal resolution, and Titmus is the stereoscopic detection result.

[0075] For example, assuming that the horizontal viewing angle of both the first and second displays is 50 degrees and the horizontal resolution is 1920 pixels, and taking a stereoscopic detection result of 800 seconds as an example, substituting into the above formula, we get:

[0076] ;

[0077] Taking two decimal places, we get N = 4 pixel differences. The corresponding number of image pixels N for each of the above viewpoints are 4, 2, 1, 0.7, 0.5, 0.4, 0.3, 0.25, and 0.2, respectively.

[0078] Step S203: Calculate the training parameters based on the number of image pixels.

[0079] Use the number of image pixels as a training parameter; or, round the number of image pixels and use the resulting integer value as a training parameter.

[0080] For example, since N is 4, 2, 1, 0.7, 0.5, 0.4, 0.3, 0.25, and 0.2 respectively, then n also takes the values ​​4, 2, 1, 0.7, 0.5, 0.4, 0.3, 0.25, and 0.2 respectively. In some embodiments, since the pixels are discretely distributed, their values ​​must be integers, and N reflects the maximum disparity in subsequent calculations, not the average level. Therefore, to make N rounded and reduce training difficulty, n is set to 5 × N, where n takes the values ​​20, 10, 5, 4, 3, 2, 2, 1, and 1 respectively.

[0081] In this embodiment, based on the relationship between parallax and depth, a parallax of 0 indicates that random points on both the left and right images are in the same position, which can be used as background for depth recognition. When the parallax is greater than 0, the offset random points will move as a whole towards the direction of binocular convergence, ultimately resulting in a convex depth perception. Conversely, when the parallax is less than 0, the offset random points will move as a whole in the opposite direction of binocular convergence, ultimately resulting in a concave depth perception. Simultaneously, the absolute value of the parallax reflects the degree of convexity or concavity; the larger the absolute value, the farther the perceived contour is from the background. For users with insufficient binocular integration ability, recognizing parallax with small absolute values ​​often presents some difficulty. Therefore, this embodiment combines the fixed parameters of the head-mounted display device (horizontal viewing angle and horizontal resolution) and the user's actual situation (stereoscopic vision detection results) to calculate the training parameter n. Based on the training parameter n, the parallax magnitude is controlled, thereby setting the optimal training difficulty for the user's actual situation to match the actual stereoscopic vision ability of different users.

[0082] In one embodiment, after displaying the first random dot pattern on the first display screen and the second random dot pattern on the second display screen, the method further includes:

[0083] The stereoscopic vision level of the test object is re-detected to obtain a second stereoscopic vision detection result; the training parameters are updated based on the second stereoscopic vision detection result; a second random dot map is recalculated and generated based on the updated training parameters, and the newly generated second random dot map is displayed on the second display screen.

[0084] For example, the training parameter 'n' should be set to gradually increase the training difficulty as the user's stereopsis level improves, in order to reduce the ceiling effect of training with constant parameters. For instance, for a patient with an initial stereopsis level of 800 seconds of visual field, the initial training parameter 'n' is 20. Subsequently, if the user's ability in the Titmus test improves to 400 seconds of visual field, the subsequent matching training parameter 'n' can be increased to 10. By updating the training parameters, the training difficulty can gradually increase as the user's stereopsis level improves, achieving a gradual training effect and helping to reduce the ceiling effect of training with constant parameters.

[0085] It should be understood that although the steps in the flowcharts of the embodiments described above are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the embodiments described above may include multiple steps or multiple stages. These steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the steps or stages of other steps.

[0086] In one embodiment, Figure 5 A schematic diagram illustrating an application scenario for a method for generating stereo disparity maps is provided. Figure 6 A flowchart of the method for generating stereo disparity maps in this application scenario, such as... Figure 6 As shown, the process includes the following steps:

[0087] Step S301: Obtain the user's current stereoscopic detection result.

[0088] Step S302: Determine whether the current stereo vision detection result is less than or equal to the threshold. If yes, end the process; otherwise, proceed to step S303. Specifically, if the current stereo vision detection result is less than or equal to the threshold, it means the user's stereo vision level has reached the expected level, and training can stop. If the current stereo vision detection result is higher than the threshold, it means the user's stereo vision level is low and has not reached the expected level, requiring further training.

[0089] Step S303: Obtain the horizontal viewing angle of the first display screen or the second display screen, obtain the horizontal resolution of the first display screen or the second display screen, and calculate the number of image pixels that the user can observe based on the first display screen or the second display screen according to the horizontal viewing angle, the horizontal resolution and the current stereoscopic detection results.

[0090] Step S304: Calculate the training parameters based on the number of image pixels. The calculation formula is as follows:

[0091] ;

[0092] Where N is the number of image pixels that the user can observe from the first or second display screen, VA is the horizontal viewing angle, nPix is ​​the horizontal resolution, and Titmus is the first stereoscopic detection result. The number of image pixels can be used as a training parameter; alternatively, the number of image pixels can be rounded down, and the resulting integer value can be used as a training parameter.

[0093] Step S305: Acquire a two-dimensional video, extract the depth information of the image frames in the two-dimensional video, and normalize the depth information based on the training parameters to obtain a disparity map.

[0094] Calculate the distance between each pixel in the image frame and the observation point; assign a value to each pixel in the range of 0 to 1 according to the distance to obtain depth information; where the larger the distance, the smaller the value assigned, and the smaller the distance, the larger the value assigned.

[0095] Step S306: Obtain the coordinates of a random point and generate a first sequence based on the coordinates of the random point. The first sequence contains two columns of values: the X-axis and the Y-axis.

[0096] Step S307: Subtract the values ​​of the X-axis column in the first sequence from the pixel values ​​of the corresponding coordinates in the disparity map to obtain the difference column, and generate the second sequence based on the difference column and the Y-axis column in the first sequence.

[0097] Step S308: Generate a first random dot plot based on the first sequence and a second random dot plot based on the second sequence. Display the first random dot plot on the first display screen and the second random dot plot on the second display screen. After this round of training, return to step S301 to re-detect the stereo vision level of the test object. Update the training parameters based on the new stereo vision detection results, and then recalculate and generate the second random dot plot based on the updated training parameters. Display the newly generated second random dot plot on the second display screen. The random dot plot generation method is as follows:

[0098] Obtain a first black image. In the first black image, fill the pixels near the center of each coordinate in the first sequence with white to generate a first random dot map. Obtain a second black image. In the second black image, fill the pixels near the center of each coordinate in the second sequence with white to generate a second random dot map.

[0099] In steps S301 to 308 above, since 2D videos are more complex and contain richer depth information than traditional disparity maps, extracting the depth information from 2D videos and converting it into a matching disparity map can overcome the limitation of traditional disparity maps containing only simple images. Furthermore, continuous depth information can be extracted from 2D videos, ultimately integrating and generating continuous RDS image frames, i.e., RDS videos. During training, users only need to watch and understand the RDS video content without completing other tasks, allowing them to achieve training effects while easily watching the RDS video, thus improving the user's experience in binocular stereoscopic training. Moreover, this training method increases user compliance, facilitating prolonged immersive training and improving the effectiveness of single training sessions. Further, by updating training parameters and controlling the disparity magnitude based on these parameters, the optimal training difficulty can be set according to the user's actual situation to match the actual stereoscopic vision ability of different users. The training difficulty can gradually increase as the user's stereoscopic vision level improves, achieving a progressive training effect and reducing the ceiling effect under constant parameter training.

[0100] Based on the same inventive concept, this application also provides a random point stereo disparity map generation apparatus for implementing the above-described method. This apparatus is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. The terms "module," "unit," "subunit," etc., used below can refer to combinations of software and / or hardware that perform a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0101] Figure 7 This is a structural block diagram of the random point stereo disparity map generation device in this embodiment, as shown below. Figure 7 As shown, the device includes:

[0102] The disparity map generation module is used to acquire two-dimensional video, extract the depth information of image frames in the two-dimensional video, and normalize the depth information based on training parameters to obtain a disparity map.

[0103] The first data sequence acquisition module is used to acquire the coordinates of random points and generate the first data sequence based on the coordinates of random points. The first data sequence contains two columns of values: X-axis and Y-axis.

[0104] The second data acquisition module is used to subtract the values ​​of the X-axis column in the first data sequence from the pixel values ​​of the corresponding coordinates in the disparity map to obtain the difference column, and generate the second data sequence based on the difference column and the Y-axis column in the first data sequence.

[0105] The random dot plot generation module is used to generate a first random dot plot based on a first sequence of numbers, and to generate a second random dot plot based on a second sequence of numbers, and to display the first random dot plot on a first display screen, and the second random dot plot on a second display screen.

[0106] It should be noted that the above modules can be functional modules or program modules, and can be implemented through software or hardware. For modules implemented through hardware, the above modules can reside in the same processor; or the above modules can be located in different processors in any combination.

[0107] Furthermore, in conjunction with the random point stereo disparity map generation method provided in the above embodiments, this embodiment can also provide a storage medium for implementation. This storage medium stores a computer program; when executed by a processor, the computer program implements any of the random point stereo disparity map generation methods in the above embodiments.

[0108] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties.

[0109] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM). The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.

[0110] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0111] The embodiments described above are merely illustrative of several implementation methods of this application, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of this patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.

Claims

1. A stereoscopic parallax map generation method characterized by comprising: The method is applied to a head-mounted display device including a first display screen and a second display screen, and comprises the following steps: obtaining a two-dimensional video, extracting depth information of image frames in the two-dimensional video, and normalizing the depth information based on a training parameter to obtain a parallax map; obtaining random point coordinates, generating a first sequence based on the random point coordinates, and the first sequence including two columns of values of X-axis and Y-axis; subtracting the values of the X-axis column in the first sequence from the pixel values of the corresponding coordinates in the parallax map to obtain a difference value column, and generating a second sequence based on the difference value column and the Y-axis column in the first sequence; generating a first random point map based on the first sequence, and generating a second random point map based on the second sequence, and displaying the first random point map on the first display screen and displaying the second random point map on the second display screen.

2. The method of generating a stereo disparity map according to claim 1, wherein, The method for extracting the depth information of the image frames in the two-dimensional video comprises the following steps: calculating the distance between each pixel point in the image frame and an observation point; assigning a value in the range of 0 to 1 to each pixel point according to the distance to obtain the depth information; wherein the greater the distance, the smaller the value, and the smaller the distance, the greater the value.

3. The method of generating a stereo disparity map according to claim 1, wherein, Before normalizing the depth information based on the training parameter, the method further comprises the following steps: obtaining the horizontal viewing angle of the first display screen or the second display screen, obtaining the horizontal resolution of the first display screen or the second display screen, and obtaining the first stereoscopic vision detection result of the user; calculating the number of image pixels that the user can observe based on the first display screen or the second display screen according to the horizontal viewing angle, the horizontal resolution and the first stereoscopic vision detection result; calculating the training parameter according to the number of image pixels.

4. The method of generating a stereo disparity map according to claim 3, wherein, The method for calculating the number of image pixels that the user can observe based on the first display screen or the second display screen according to the horizontal viewing angle, the horizontal resolution and the stereoscopic vision detection result comprises the following steps: ; wherein N is the number of image pixels, VA is the horizontal viewing angle, nPix is the horizontal resolution, and Titmus is the first stereoscopic vision detection result.

5. The method of generating a stereo disparity map according to claim 3, wherein, The method for calculating the training parameter according to the number of image pixels comprises the following steps: taking the number of image pixels as the training parameter; or taking the integer value obtained by rounding the number of image pixels as the training parameter.

6. The method of generating a stereo disparity map according to claim 3, wherein, After displaying the first random point map on the first display screen and displaying the second random point map on the second display screen, the method further comprises the following steps: re-detecting the stereoscopic vision level of the test object to obtain a second stereoscopic vision detection result; updating the training parameter according to the second stereoscopic vision detection result; re-calculating the second random point map based on the updated training parameter, and displaying the newly generated second random point map on the second display screen.

7. The method of generating a stereo disparity map according to claim 1, wherein, The method for generating a first random point map based on the first sequence and generating a second random point map based on the second sequence comprises the following steps: Obtaining a first black image, in the first black image, filling the pixel points near the center as white to generate the first random point image; Obtaining a second black image, in the second black image, filling the pixel points near the center as white to generate the second random point image.

8. A random dot stereoscopic disparity map generation apparatus characterized by comprising: It comprises: A parallax map generation module is configured to obtain a two-dimensional video, extract depth information of image frames in the two-dimensional video, and normalize the depth information based on training parameters to obtain a parallax map; A first sequence acquisition module is configured to obtain random point coordinates, generate a first sequence based on the random point coordinates, and the first sequence comprises two columns of values of X-axis and Y-axis; A second sequence acquisition module is configured to subtract the values of the X-axis column in the first sequence from the pixel values of the corresponding coordinates in the parallax map to obtain a difference column, and generate a second sequence based on the difference column and the Y-axis column in the first sequence; A random point image generation module is configured to generate a first random point image based on the first sequence, generate a second random point image based on the second sequence, display the first random point image on a first display screen, and display the second random point image on a second display screen.

9. A head-mounted display device, comprising: It comprises: A first display screen, a second display screen and a control module, the first display screen and the second display screen are connected with the control module respectively, and the control module is used for executing the stereoscopic parallax map generation method in any one of claims 1 to 7.

10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by a processor to realize the steps of the method in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Visual field auxiliary device cooperating with ARVR head-mounted typoscope

    CN215821381U

  • Imaging apparatus, image display apparatus and image recording and / or reproducing apparatus

    US6177952B1

  • Virtual-reality system and method for rehabilitating exotropia patients on basis of artificial intelligence, and computer-readable medium

    WO2021162207A1