Eyeball position recognition method and apparatus, device, and storage medium
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2026-08-13
Smart Images

Figure CN2025070602_13082026_PF_FP_ABST
Abstract
Description
Methods, devices, equipment, and storage media for identifying eye position Technical Field
[0001] This application relates to the field of display technology, and in particular to a method, apparatus, device, and storage medium for identifying the position of an eyeball. Background Technology
[0002] In the field of display technology, display devices achieve three-dimensional display functionality by recognizing the position of the user's eyeballs. For example, the display screen of the device displays two images from two viewpoints, with the image from the left eye viewpoint projected onto the left eye's position and the image from the right eye viewpoint projected onto the right eye's position, thereby achieving three-dimensional display. Summary of the Invention
[0003] This application provides a method, apparatus, device, and storage medium for identifying the position of a user's eyes as they gaze at a screen. The technical solution is as follows:
[0004] On the one hand, a method for identifying eye position is provided, the method comprising:
[0005] The position information of the target object is acquired by multiple first cameras at different angles. The multiple first cameras include a first main camera and a first auxiliary camera. The position information acquired by the first main camera has a greater weight than the position information acquired by the first auxiliary camera.
[0006] The location range of the target object's eyeball is determined based on the location information and weights acquired by the multiple first cameras;
[0007] The feature point information of the target object collected by multiple second cameras within the specified location range is acquired. The multiple second cameras include a second main camera and a second auxiliary camera. The weight of the feature point information collected by the second main camera is greater than the weight of the feature point information collected by the second auxiliary camera.
[0008] The eye position of the target object is determined based on the feature point information and weights acquired by the multiple second cameras.
[0009] In one possible implementation, the method further includes:
[0010] If the position information acquired by the first main camera indicates that the target object is occluded, at least one of the first auxiliary cameras is selected as the updated first main camera. The eye position is determined based on the position information acquired by the updated first main camera, wherein the position information acquired by the updated first main camera indicates that the target object is not occluded; and / or,
[0011] If the feature point information acquired by the second main camera indicates that the target object is occluded, at least one of the second auxiliary cameras is selected as the updated second main camera, and the eye position is determined based on the feature point information acquired by the updated second main camera. The feature point information acquired by the updated second main camera indicates that the target object is not occluded.
[0012] In one possible implementation, after determining the eye position of the target object based on the feature point information and weights acquired by the plurality of second cameras, the method further includes:
[0013] The resolution of the display screen is determined based on the eye position, and the display screen is controlled to display a three-dimensional image that matches the eye position according to the resolution.
[0014] In one possible implementation, the three-dimensional image includes a dual-viewpoint image determined based on eye position, and the method further includes:
[0015] If the eye position is not identified, the dual-viewpoint image is converted into a multi-viewpoint image;
[0016] Upon re-identifying the eye position, the display screen is controlled to gradually change the multi-viewpoint image into the corresponding dual-viewpoint image over a reference time period.
[0017] In one possible implementation, the method further includes:
[0018] Without recognizing the position of the eyeball, the display screen is controlled to display a two-dimensional image;
[0019] Once the eye position is re-identified, the display screen is controlled to gradually change the two-dimensional image into a corresponding three-dimensional image over a reference time period.
[0020] In one possible implementation, before controlling the display screen to gradually change the two-dimensional image into a corresponding three-dimensional image over a reference time, the method further includes:
[0021] Determine the distance between the eyeball and the display screen, and determine the linear adjustment parameters based on the distance;
[0022] The three-dimensional image is obtained by performing a linear depth transformation on the two-dimensional image as a whole using the linear adjustment parameters.
[0023] In one possible implementation, the two-dimensional image includes a display subject and a background, the display subject and the background having different depths of field, and the background being blurred; controlling the display screen to gradually change the two-dimensional image into a corresponding three-dimensional image over a reference time includes:
[0024] The depth of field between the displayed subject and the background is adjusted so that the background recovers to the depth of field before the blurring process after a reference time.
[0025] In one possible implementation, the method further includes:
[0026] If the eye position is not identified and the head position of the target object is within the acquisition range, the reference eye position range is determined based on the last determined eye position.
[0027] If the eye position is re-identified, the eye position of the target object is re-determined based on the reference eye position range.
[0028] In one possible implementation, the method further includes:
[0029] In the absence of eye position recognition, the eye position trajectory of the target object is determined based on the posture information acquired by the first camera, wherein the posture information includes at least one of the target object's body posture or head posture;
[0030] The location where the target object's eyeballs are re-collected is determined based on the eyeball position trajectory.
[0031] On the other hand, an eye position recognition device is also provided, the device comprising:
[0032] The acquisition module is used to acquire the position information of the target object collected by multiple first cameras at different angles. The multiple first cameras include a first main camera and a first auxiliary camera. The position information collected by the first main camera has a greater weight than the position information collected by the first auxiliary camera.
[0033] The determination module is used to determine the location range of the target object's eyeball based on the location information and weights acquired by the plurality of first cameras;
[0034] The acquisition module is further configured to acquire feature point information of the target object collected by multiple second cameras within the location range. The multiple second cameras include a second main camera and a second auxiliary camera. The weight of the feature point information collected by the second main camera is greater than the weight of the feature point information collected by the second auxiliary camera.
[0035] The determining module is further configured to determine the eye position of the target object based on the feature point information and weights acquired by the plurality of second cameras.
[0036] In one possible implementation, the determining module is further configured to, when the position information acquired by the first main camera indicates that the target object is occluded, select at least one from the first auxiliary cameras as an updated first main camera, determine the eye position based on the position information acquired by the updated first main camera, wherein the position information acquired by the updated first main camera indicates that the target object is not occluded; and / or,
[0037] If the feature point information acquired by the second main camera indicates that the target object is occluded, at least one of the second auxiliary cameras is selected as the updated second main camera, and the eye position is determined based on the feature point information acquired by the updated second main camera. The feature point information acquired by the updated second main camera indicates that the target object is not occluded.
[0038] In one possible implementation, the device further includes:
[0039] The first control module is used to determine the resolution of the display screen based on the eye position, and control the display screen to display a three-dimensional image matching the eye position according to the resolution.
[0040] In one possible implementation, the three-dimensional image includes a dual-viewpoint image determined based on the eye position. The first control module is further configured to convert the dual-viewpoint image into a multi-viewpoint image if the eye position is not identified; and to control the display screen to gradually change the multi-viewpoint image into the corresponding dual-viewpoint image over a reference time if the eye position is re-identified.
[0041] In one possible implementation, the device further includes:
[0042] The second control module is used to control the display screen to display a two-dimensional image when the eye position is not detected; and to control the display screen to gradually change the two-dimensional image into a corresponding three-dimensional image within a reference time when the eye position is re-detected.
[0043] In one possible implementation, the determining module is further configured to determine the distance between the eyeball and the display screen, determine a linear adjustment parameter based on the distance, and perform a depth-of-field linear transformation on the two-dimensional image as a whole using the linear adjustment parameter to obtain the three-dimensional image.
[0044] In one possible implementation, the two-dimensional image includes a display subject and a background, the display subject and the background having different depths of field, and the background being blurred; the second control module is used to adjust the depth of field of the display subject and the background so that the background recovers to the depth of field before the blurring process after a reference time.
[0045] In one possible implementation, the determining module is further configured to determine a reference eye position range based on the last determined eye position when no eye position is identified and the head position of the target object is within the acquisition range; and to re-determine the eye position of the target object based on the reference eye position range when an eye position is re-identified.
[0046] In one possible implementation, the determining module is further configured to determine the eye position trajectory of the target object based on the posture information acquired by the first camera when the eye position is not identified, wherein the posture information includes at least one of the target object's body posture or head posture; and determine the position where the target object's eye is re-acquired based on the eye position trajectory.
[0047] On the other hand, a computer device is also provided, the computer device including a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the computer device to implement the eye position recognition method described in any of the above aspects.
[0048] On the other hand, a computer-readable storage medium is also provided, wherein the non-transient computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable a computer to implement the eye position recognition method described in any of the above aspects.
[0049] On the other hand, a computer program product or computer program is also provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the eye position recognition method described in any of the preceding aspects.
[0050] The technical solution provided in this application has at least the following beneficial effects:
[0051] The technical solution provided in this application prioritizes determining the location range of the target object's eyeballs. The process of the second camera acquiring feature point information can then be performed within this location range, reducing the amount of data that needs to be processed and thus improving the efficiency of determining the eyeball position. Furthermore, by combining multiple first cameras and multiple second cameras, and determining the eyeball position according to primary and secondary cameras and their weights, the accuracy and reliability of the determination results are improved. Attached Figure Description
[0052] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0053] Figure 1 is a schematic diagram of the implementation environment of an eye position recognition method provided in an embodiment of this application;
[0054] Figure 2 is a flowchart of an eye position recognition method provided in an embodiment of this application;
[0055] Figure 3 is a schematic diagram showing the relationship between the average number of camera groups and accuracy jitter according to an embodiment of this application;
[0056] Figure 4 is a schematic diagram of a dual-camera combination method provided in an embodiment of this application;
[0057] Figure 5 is a schematic diagram of a camera position provided in an embodiment of this application;
[0058] Figure 6 is a schematic diagram of a scene for switching the main camera according to an embodiment of this application;
[0059] Figure 7 is a schematic diagram of a scene for switching the main camera according to an embodiment of this application;
[0060] Figure 8 is a schematic diagram of a camera position provided in an embodiment of this application;
[0061] Figure 9 is a schematic diagram of a feature point information determination process provided in an embodiment of this application;
[0062] Figure 10 is a schematic diagram illustrating the relationship between depth of field and distance according to an embodiment of this application;
[0063] Figure 11 is a schematic diagram of a display conversion process provided in an embodiment of this application;
[0064] Figure 12 is a schematic diagram of target object head movement provided in an embodiment of this application;
[0065] Figure 13 is a schematic diagram of a multi-viewpoint image transformation process provided in an embodiment of this application;
[0066] Figure 14 is a schematic diagram of the structure of an eye position recognition device provided in an embodiment of this application;
[0067] Figure 15 is a schematic diagram of the structure of a server provided in an embodiment of this application;
[0068] Figure 16 is a schematic diagram of the structure of a terminal provided in an embodiment of this application. Detailed Implementation
[0069] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0070] It should be noted that the terms "first," "second," etc. (if applicable) used in the specification of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. The implementations described in the following exemplary embodiments do not represent all implementations consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application.
[0071] In the field of display technology, by identifying the user's eye position, 3D interactive technology based on eye tracking can achieve 3D displays that match the user's eye position. However, due to the reflection and scattering of various light sources in natural environments or under artificial lighting conditions, the human eye may experience reflective points or shadow areas, affecting the accuracy of eye position detection and calculation. Furthermore, differences in eye color, pupil size, and shape also increase the complexity of detection. Moreover, a single camera can only determine two-dimensional image information and cannot directly obtain the precise position of the eye in three-dimensional space, affecting the continuity and accuracy of tracking.
[0072] In related technologies, eye position is identified using images captured by two types of cameras. For example, eye position is identified using images captured by a depth camera and a grayscale camera. The image captured by the depth camera determines the user's depth position in three-dimensional space, while the image captured by the grayscale camera determines the location of the user's eye feature points. However, in these technologies, the range of the images captured by the depth camera and the grayscale camera is the same, resulting in high computational resource consumption, slow eye position recognition speed, and low efficiency.
[0073] This application provides a method for identifying eye position, which improves the efficiency of eye position identification. Please refer to Figure 1, which illustrates a schematic diagram of the implementation environment for the eye position identification method provided in this application. This implementation environment may include a computer device 101 and a display screen 102, which are connected via wired or wireless means. Optionally, the display screen 102 may be a screen mounted on the computer device 101, in which case the display screen 102 is part of the computer device 101. The display screen 102 may integrate a first camera and a second camera, or the first camera and the second camera may be independent devices from the display screen 102, but the display screen 102 is connected to both the first camera and the second camera.
[0074] This application does not limit the type of computer device 101. For example, computer device 101 can be a terminal or a server. Optionally, the terminal can be any electronic product that can interact with the user through one or more methods such as a keyboard, touchpad, touch screen, remote control, voice interaction, or handwriting device, such as PC (Personal Computer), mobile phone, smartphone, PDA (Personal Digital Assistant), wearable device, PPC (Pocket PC), tablet computer, smart car system, smart TV, smart speaker, etc. The server can be a single server, a server cluster composed of multiple servers, or a cloud computing service center.
[0075] This application does not limit the type of display screen 102. For example, display screen 102 can be an LED (Light Emitting Diode) screen or an LCD (Liquid Crystal Display) screen, etc.
[0076] Those skilled in the art should understand that the computer device 101 and display screen 102 described above are merely examples, and other existing or future computer devices and display screens that are applicable to this application should also be included within the scope of protection of this application, and are hereby incorporated by reference.
[0077] Referring to Figure 2, which is a flowchart of an eye position recognition method provided in an embodiment of this application, the method is described using a computer device connected to a display screen. For example, the computer device can be the computer device 101 shown in Figure 1, and the display screen can be the display screen 102 shown in Figure 1. As shown in Figure 2, the eye position recognition method includes, but is not limited to, the following steps 201-204.
[0078] Step 201: Obtain the position information of the target object collected by multiple first cameras at different angles. The multiple first cameras include a first main camera and a first auxiliary camera. The position information collected by the first main camera has a greater weight than the position information collected by the first auxiliary camera.
[0079] This application does not limit the type of the first camera, as long as it can collect the position information of the target object. This application uses a depth camera as an example for illustration. A depth camera can also be called a 3D (3D) camera or a stereoscopic vision camera, capable of measuring the distance from objects in the scene to the camera. This application does not limit the method of collecting the distance from the object to the camera. For example, it can be done by emitting light (e.g., infrared or laser) and receiving the reflected light, or by comparing the differences between images from different viewpoints to calculate the depth information of each point in the scene, thereby constructing a 3D model of the target object in the scene, including the shape and size of the target object.
[0080] The depth camera can acquire the position information of the target object within a reference range corresponding to the display screen. This reference range can refer to a portion of the area in front of the display screen, i.e., the location range of the target object. Optionally, the first camera in this embodiment can be a single depth camera or a combination of multiple grayscale cameras to implement the function of a depth camera.
[0081] The location information acquired by a depth camera within a reference range can include depth maps or point cloud data of the scene, or it can be understood as the location information including the distance information from each point in the scene to the camera. For the depth maps or point cloud data acquired by the depth camera, computer equipment processes the location information of the target object acquired by the depth camera. The processing includes, but is not limited to, noise filtering, data smoothing, and edge detection steps to improve the accuracy and reliability of the data.
[0082] In the processed location information, the computer device needs to identify the target object. This application embodiment does not limit the method for identifying the target object; for example, it can be achieved through template matching, feature recognition, or machine learning algorithms. For the target object, the computer device can use depth data to calculate the object's precise location information in three-dimensional space. The location information may include the target object's center point coordinates, orientation, size, etc., and this application embodiment does not limit the location information.
[0083] In one possible implementation, since there are multiple first cameras, embodiments of this application employ multiple first cameras to construct a multi-view observation system. These multiple first cameras are placed at different locations within the scene to capture views of the target object from multiple angles. The number and location of the multiple first cameras depend on the requirements of the application scenario, the size and complexity of the scene, and the expected level of accuracy and robustness.
[0084] In a multi-view observation system constructed from multiple first-camera systems, positioning accuracy is improved by fusing data from multiple first-camera systems. The primary cause of positioning accuracy error is random noise, which in turn follows a normal distribution A ~ N(μ,σ). 2 As the number of cameras increases, the variance of the mean decreases by the inverse of the number of cameras. In other words, the error of a multi-camera system (relative to a single-camera system) will decrease by an inverse ratio. The proportion decreases (the amount of decrease in error e is) ).
[0085] Refer to Figure 3 for the relationship between the camera group average and accuracy jitter, where the camera group average y and accuracy jitter x satisfy y = x -0.5 And R 2 =1,R 2 R is used to measure the predictive power of a statistical model and can be understood as... 2 This indicates the interpretability of the regression model for the prediction results, with 1 indicating 100% interpretability. Based on Figure 3, it can be seen that achieving a 25% accuracy jitter for positioning with a single camera requires 15 cameras. If a dual-camera combination is used to acquire localization data, 15 dual-camera combinations can be achieved using only 6 cameras. For the dual-camera combination scenario, see Figure 4, which illustrates one possible dual-camera combination. Figure 4 uses three cameras A, B, and C as examples: the first combination is camera A and camera B; the second combination is camera B and camera C; and the third combination is camera A and camera C. In summary, the camera combinations satisfy the following formula 1.
[0086] Where m is the number of cameras. This indicates the different possible combinations of cameras that can be selected from m cameras, i.e., camera combinations.
[0087] Taking the first camera as an example of a depth camera, each depth camera can independently or in combination acquire depth images or point cloud data within its field of view. The positional information acquired by each depth camera refers to the target object's three-dimensional spatial coordinates, orientation, possible size, and shape from each camera's perspective. Due to the different camera positions, the positional information acquired from each perspective will differ, and the positional information acquired from different perspectives collectively describes the complete state of the target object in three-dimensional space. Furthermore, when processing data from multiple first cameras, it is necessary to ensure that the data is synchronized, i.e., corresponding to the scene state at the same point in time. Additionally, each first camera needs to be calibrated to determine its internal parameters (such as focal length) and external parameters (such as position and orientation).
[0088] In this embodiment, the position information of the target object acquired by each first camera can be considered as sub-position information of the target object. The position information acquired from different first cameras is fused to obtain information indicating the position of the target object. This embodiment does not limit the method of fusing the position information acquired by each first camera to obtain information indicating the position of the target object. For example, it can be done through multi-view geometry, stereo matching, optimization techniques, etc.
[0089] For example, this application does not limit the positions of different first cameras in its embodiments. For instance, referring to a camera position diagram in Figure 5, a first camera located at the top of the display screen can be called a top camera or top camera, and a first camera located at the bottom of the display screen can be called a bottom camera or bottom camera. In Figure 5(I), the top cameras include cameras A, B, and C, and the bottom cameras include cameras D, E, and F. In Figure 5(II), the top cameras include cameras A, B, C, and D, and the bottom cameras include cameras E and F. In Figure 5(III), the top cameras include cameras A and B, and the bottom cameras include cameras C, D, E, and F. Various arrangements can be made for the camera matrix arrangement scheme. To ensure noise consistency among multiple sets of first cameras and avoid issues with camera distance ranges, they can be arranged at equal intervals or around the top and bottom of the screen.
[0090] In the embodiments of this application, multiple first cameras are distinguished as a first main camera and a first auxiliary camera, that is, multiple first cameras include a first main camera and a first auxiliary camera. The position information acquired by the first main camera has a greater weight than the position information acquired by the first auxiliary camera. The position of the target object is determined based on the position information of the target object at different angles and the weight of each position information. In other words, the first main camera and the first auxiliary camera differ in function, which is reflected in the weight of the acquired position information when finally determining the position of the target object. The computer device assigns weights to different first cameras, and the weights of the position information acquired by different first cameras can be consistent with the weights of the first cameras. The position information acquired by the first main camera is given a higher weight, possibly because the main camera has a better field of view, that is, the target object within the field of view of the first main camera is clearer and not occluded. Therefore, when fusing multi-view data, the data of the first main camera will have a greater impact on the final result.
[0091] In this embodiment, the primary and secondary cameras can be determined based on the position information acquired by the cameras to indicate whether the target object is occluded. For example, if multiple primary cameras are not occluded, all of the primary cameras can be primary cameras. In practical applications, the target object may not be fully observed by a single camera due to occlusion by other objects. Therefore, the computer device has a dynamic camera role switching mechanism. When the position information acquired by the primary camera indicates that the target object is occluded, the computer device will select at least one of the secondary cameras as the updated primary camera.
[0092] If the first camera is composed of two grayscale cameras, taking a total of three grayscale cameras as an example (grayscale camera A, grayscale camera B, and grayscale camera C), then the first camera can be a combination of grayscale camera A and grayscale camera B, a combination of grayscale camera A and grayscale camera C, or a combination of grayscale camera B and grayscale camera C. The first main camera may be a combination of grayscale camera A and grayscale camera B. If grayscale camera B is obscured, then the combination of grayscale camera A and grayscale camera C needs to be switched to become the new first main camera.
[0093] When selecting the updated primary camera, the computer equipment needs to consider several factors, such as whether the secondary camera's field of view includes the unobstructed portion of the target object, and the secondary camera's performance (e.g., resolution, stability). The updated primary camera will continue to acquire the target object's location information, and this location information will have the same or higher weight as the original primary camera during the fusion process. Because the new primary camera can observe the unobstructed portion of the target object, the acquired data will be more helpful in accurately determining the target object's location.
[0094] Figure 6 illustrates a scenario of switching the primary camera. Taking the display screen as an example, it integrates three primary cameras: Camera 1, Camera 2, and Camera 3. When user A is using the device, Camera 1 acts as the primary camera to collect location information, while Cameras 2 and 3 act as secondary cameras to correct information accuracy. If user B then appears in the interaction scene and partially obscures user A, Camera 1 cannot detect user A's information, and the interaction process enters a new logical judgment. Secondary cameras 2 and 3 then confirm the information. If A is still in the interaction scene and the user has not actively switched to tracking a different person, then Cameras 2 and 3 become the primary cameras and collect information.
[0095] For example, referring to Figure 7, another scenario diagram of switching the main camera is shown. Since the head of the target object is occluded in the field of view of the original main camera D, camera A is switched to be the main camera. The head of the target object captured by main camera A is not occluded. Throughout the process, the computer device dynamically adjusts the role and weight allocation of the cameras according to the actual situation to ensure that the most accurate and comprehensive target object location information is always obtained. Furthermore, the computer device may continuously optimize performance through machine learning or optimization algorithms to improve the recognition accuracy and robustness under occlusion conditions.
[0096] Step 202: Determine the location range of the target object's eyeball based on the location information and weights collected by multiple first cameras.
[0097] In one possible implementation, the position of the target object is determined based on position information acquired by multiple first cameras. This position can refer to the global position of the target object in the image, which can be given in the form of a coordinate frame, indicating the approximate position and size of the target object in the image. Optionally, the image can be preprocessed before determining the eyeball position range. Preprocessing includes, but is not limited to, noise reduction, contrast enhancement, and color correction, to improve the accuracy and efficiency of subsequent processing. This application does not limit the process of determining the eyeball position range; for example, it can use the position range of the eyeball region on the head to determine the eyeball position range. Alternatively, it can use techniques such as geometric relationship inference from the acquired position information, and consider factors such as lighting conditions, occlusion issues, and individual differences to accurately determine the target object's eyeball position range.
[0098] Step 203: Obtain feature point information of the target object collected by multiple second cameras within the location range. The multiple second cameras include a second main camera and a second auxiliary camera. The weight of the feature point information collected by the second main camera is greater than the weight of the feature point information collected by the second auxiliary camera.
[0099] This application embodiment does not limit the type of the second camera; any camera capable of acquiring information for determining feature points is acceptable. This application embodiment uses a grayscale camera as an example, which can be called an RGB (Red, Green, Blue) camera. Grayscale cameras are used to capture the brightness information of an image. In visual processing, grayscale cameras are used to reduce computational load, improve processing speed, and reduce noise interference. Grayscale cameras have sufficiently high resolution and sensitivity to clearly capture subtle features of the eye. Furthermore, the grayscale camera used in step 201 to combine and implement the depth camera function can also be used independently as the grayscale camera in step 202 to implement the grayscale camera function.
[0100] In this embodiment, a second camera acquires grayscale images within a specified location range. These grayscale images contain brightness information. A computer device can perform a series of image processing operations on the acquired grayscale images, such as filtering, edge detection, and morphological operations, to highlight the features of the eyeball and reduce noise interference. Subsequently, feature point information of the target object is obtained based on the processed grayscale image. This embodiment does not limit the method of obtaining feature point information of the target object based on the grayscale image. The feature point information acquired by the second camera includes, but is not limited to, the position, shape, and pupil direction of the eyeball. Any method that can obtain feature point information based on a grayscale image is applicable to this application.
[0101] Furthermore, the term "multiple second cameras" refers to the use of multiple second cameras in a computer device to work collaboratively, rather than relying on a single camera to collect data. This allows for applications requiring a wider field of view, higher precision, or 3D reconstruction. Examples include augmented reality, virtual reality, robot vision, security monitoring, and 3D modeling. In this embodiment, each second camera is positioned at a different angle or location from which the target object can be observed. Each second camera can acquire unique information about the target object from its own perspective, particularly feature point information about the target object's surface or structure.
[0102] Furthermore, due to the different positions of the multiple second cameras, the feature point information acquired by each second camera will also differ. This diversity of feature point information provides a rich data source for subsequent data fusion and 3D reconstruction. Based on feature point information from different angles, after collecting feature point information from all the second cameras, the computer device can integrate the feature information to determine the overall feature point information of the target object. The processing includes, but is not limited to, aligning, transforming, and fusing images from multiple viewpoints to eliminate visual differences caused by different camera positions and angles.
[0103] This application does not limit the position of the second camera. For example, the arrangement shown in Figure 5 can be referred to, and will not be repeated here. Furthermore, for the simultaneous presence of a first camera and a second camera, a camera position diagram is shown in Figure 8, where cameras A, B, C, D, and E are grayscale cameras, and cameras X, Y, and Z are depth cameras. Optionally, the depth cameras and grayscale cameras can be freely combined according to user needs.
[0104] In one possible implementation, in a multi-camera computer device, multiple second cameras are explicitly distinguished as a second primary camera and a second auxiliary camera; that is, the multiple second cameras include a second primary camera and a second auxiliary camera. The feature point information acquired by the second primary camera has a greater weight than the feature point information acquired by the second auxiliary camera. The overall feature point information of the target object is determined based on the feature point information of the target object at different angles and the weights of each feature point. If the feature point information acquired by the second primary camera indicates that the target object is occluded, at least one of the second auxiliary cameras is selected as the updated second primary camera. The feature point information acquired by the updated second primary camera indicates that the target object is not occluded.
[0105] In other words, the feature point information captured by the second main camera is given greater weight in subsequent processing, and the position of the second main camera is more conducive to capturing the key features of the target object. The computer device determines the overall feature point information of the target object by comprehensively considering the feature point information of the target object from different angles and the weight of the feature point information. The determination process involves image processing algorithms and geometric transformations to ensure that the feature point information captured from different perspectives can be accurately and consistently fused.
[0106] During data acquisition, if the feature point information acquired by the second primary camera indicates that the target object is occluded (i.e., some or all feature points cannot be clearly identified), the computer device will detect this situation. Occlusion may be caused by other objects entering the field of view, changes in the target object's own posture, or limitations of the camera's viewing angle. Once occlusion is detected, the computer device will select at least one of the second auxiliary cameras as the updated second primary camera. The selection process may be based on various factors, such as the relative position of the auxiliary camera's viewing angle to the unoccluded portion of the target object. The updated second primary camera will take over the role of the original second primary camera, continuing to acquire and contribute feature point information with higher weight. Furthermore, if it is subsequently found that the occlusion in the field of view of the original second primary camera has been removed, the computer device can re-evaluate and restore its role as the primary camera, or maintain the updated configuration unchanged depending on the current situation.
[0107] By dynamically adjusting the camera's role, computer equipment can better cope with changes in the environment and the state of the target object, ensuring accurate and comprehensive feature point information is obtained under any circumstances. This improves the accuracy and reliability of the computer equipment, enhancing its adaptability and flexibility.
[0108] Based on the acquisition process of the depth camera and grayscale camera described above, please refer to the schematic diagram of the feature point information determination process shown in Figure 9. In Figure 9, A represents the depth image acquired by the depth camera. For the area where the eyeball is located (the boxed area in A of Figure 9) determined in the depth image, the grayscale camera acquires a grayscale image, resulting in B in Figure 9. Furthermore, in cases where the image contains occlusions, the depth image corresponding to C in Figure 9 lacks detail, while the grayscale image of D in Figure 9 can display details, such as occlusion by a user's finger.
[0109] Step 204: Determine the eye position of the target object based on the feature point information and weights collected by multiple second cameras.
[0110] In one possible implementation, when determining the target's eye position based on feature point information and weights acquired by multiple second cameras, the feature point information acquired by the multiple second cameras can be processed according to weights to obtain the overall feature point information of the target object. Based on this overall feature point information, by analyzing the position, shape, pupil direction, etc., of the eyeball, the three-dimensional spatial position of the eyeball, i.e., the eyeball position, can be inferred. This application does not limit the method of determining the eyeball position; this process can be implemented based on an eyeball model, gaze tracking algorithms, etc. For example, an eyeball model can be pre-trained using training samples of feature point information and eyeball positions. Then, the overall feature point information of the target object obtained is input into the eyeball model to obtain the eyeball position output by the eyeball model.
[0111] After identifying the eye position, a 3D display that matches the user's eye position can be achieved based on eye-tracking 3D interaction technology. In one possible implementation, after determining the target object's eye position, the method further includes: determining the resolution of the display screen based on the eye position, and controlling the display screen to display a 3D image matching the eye position according to the resolution.
[0112] For example, computer equipment collects, analyzes, and processes feature point information of the target object, determining the target object's eye position. A display screen exists in the scene, positioned in front of the target object, used to display visual content to it. For instance, this display screen could be one that supports 3D display. The determined eye position can then be used to display 3D images to the user; for example, the display screen can show two images from two different viewpoints, with the image from the left eye viewpoint projected onto the left eye's position and the image from the right eye viewpoint projected onto the right eye's position, thus achieving 3D display. The 3D image can be pre-made content or generated through real-time rendering, aiming to provide a more vivid and three-dimensional visual effect.
[0113] Furthermore, after determining the eye position, the resolution of the display screen can be determined based on that eye position, thereby controlling the display screen to display a 3D image matching the eye position at the determined resolution. This application does not limit the method of determining the display screen resolution based on the eye position. For example, the correspondence between position and resolution can be set in advance based on experience or by the user. After determining the eye position, the resolution is determined based on this correspondence. By controlling the display screen to display the 3D image at this resolution, it better meets the user's actual needs and improves the user experience.
[0114] Furthermore, in this embodiment, the computer device needs to determine whether it can capture the eye position; failure to capture the eye position is referred to as eye position loss. In a dynamic 3D interactive environment, every subtle user movement can become a key factor affecting tracking accuracy. For example, basic actions such as looking down or leaving the interaction area, or rapid head rotations, can pose challenges to computer devices used for eye tracking. Especially when a user suddenly changes their eye position or head movement speed, eye tracking algorithms in related technologies may struggle to adapt quickly, leading to data loss or misjudgment, which in turn affects subsequent 3D rendering and interactive response.
[0115] Once a user temporarily leaves the effective tracking range of the computer device, such as by looking down at a phone or simply turning to talk to someone, the eye-tracking computer device struggles to immediately re-lock onto the user's eye position. When the user re-enters the interaction area, the computer device needs to re-initialize or recalibrate, which not only prolongs the user's waiting time but may also introduce new errors due to inaccurate calibration. Due to the lack of continuous and accurate eye position information, the 3D display screen cannot precisely match the user's eye positions, leading to misalignment of the left and right eye images and causing visual conflict, i.e., visual backlash. This causes visual discomfort for the user and severely degrades the user experience. To address this, embodiments of this application provide the following methods to reduce user visual discomfort.
[0116] In one possible implementation, the method further includes: when the eye position is not detected, controlling the display screen to display a two-dimensional image; when the eye position is re-detected, controlling the display screen to gradually change the two-dimensional image into a corresponding three-dimensional image over a reference time period.
[0117] When a computer device fails to detect eye position, it may be because the user is not currently paying particular attention to the content on the screen. Therefore, to conserve computer resources and avoid unnecessary visual distractions, the computer controls the display screen to show a two-dimensional image, which can be a still image, video, or other flat content. When the computer device re-detects the eye position, the user may have become interested in the content on the screen. To respond to the user's gaze and provide a richer visual experience, the computer device begins preparing to convert the current two-dimensional image into a three-dimensional image.
[0118] During the transition, the computer controls the display screen to gradually change the 2D image to the corresponding 3D image within a reference time (i.e., a preset time period). The gradient effect helps to smooth the transition and reduce visual abruptness. The gradient effect can be switched gradually from low depth of field to high depth of field, gradually achieving normal display based on the characteristics of the human eye. The length of the reference time should be adjusted according to the actual situation; for example, the reference time could be 300ms (Millisecond).
[0119] This application does not limit the method of changing a two-dimensional image to a three-dimensional image. In one possible implementation, before controlling the display screen to gradually change the two-dimensional image to the corresponding three-dimensional image within a reference time, it also includes: determining the distance between the eyeball and the display screen, determining a linear adjustment parameter based on the distance, and performing a depth-of-field linear transformation on the two-dimensional image as a whole through the linear adjustment parameter to obtain a three-dimensional image.
[0120] In this method, the entire 2D image is treated as a whole, and a linear depth transformation is performed. First, the distance L between the eyeball and the display screen is determined, and the linear adjustment parameters are used to determine these distances. Then, the 2D image is treated as a whole and a linear depth transformation is performed using these parameters to obtain a 3D image. For example, the depth information of all points on the screen is shown below:
[0121] The depth-of-field output image, i.e., the 3D image, is as follows:
[0122] Where t is the time parameter detected by the human eye, and c is the linear equation of distance L. This application does not limit the equation c, as long as it reflects the relationship between distance L and depth of field. Furthermore, the 300 mentioned above is merely an example of one parameter, and this application does not limit it. For example, the relationship between depth of field and distance L can be shown in Figure 10. Moreover, the magnitudes of depth of field and distance can be shown in Table 1 below; the values in Table 1 are merely examples and are not intended to limit this application.
[0123] Table 1
[0124] In another possible implementation, the 2D image includes a subject and a background, with different depths of field for the subject and background, and the background is blurred. The display screen is controlled to gradually transform the 2D image into a corresponding 3D image over a reference time period. This includes adjusting the depth of field of the subject and background so that the background recovers to its depth of field before blurring after the reference time period. In this method, for example, if the 2D image displays a model (e.g., a cylinder in Figure 10), the subject and background can be differentiated and their depths of field adjusted during the 2D image display. The background is blurred, and once detected by the human eye, the depth of field of the subject can be quickly restored, while the background gradually recovers to its normal depth of field, i.e., the depth of field before blurring. The method of depth adjustment can be referenced in the above-described method of linear depth transformation, and will not be elaborated upon here.
[0125] For example, regarding the case where the user turns their head, see Figure 11 for a schematic diagram of the display transition process. When looking up directly, face tracking is normal. If the face disappears, it is determined whether it is obscured by a hand. If so, the hand position is determined using the gesture coordinates provided by the gesture recognition module. If the hand is near the line connecting the center of the face and the grayscale camera, it indicates that the gesture obscured the face, meaning the person has not left. The re-detection frequency is increased, and the coordinates before disappearance are locked. If the gesture is not in front of the face, it is determined whether the head is turned or lowered. If so, the distance point cloud data near the face's disappearance position is determined using the distance point cloud data provided by the gesture recognition module. The magnitude and area distribution of the distance values are statistically analyzed. If an area is close to the size of the face and the distance value is close to the face's distance before disappearance, it indicates that the face is lowered or turned. The re-detection frequency is increased, the coordinates before disappearance are locked, and the display is switched to 2D. Otherwise, it indicates that the face has left.
[0126] When the user looks up directly, the computer device maintains stable face tracking, updating the user's face location information in real time and displaying corresponding 3D content. If the computer device detects a sudden disappearance of the face, it immediately triggers the face disappearance event handling process. First, it checks for the possibility of hand occlusion. The gesture recognition module analyzes the image or depth data captured by the camera to identify the presence of a gesture and calculate its position, as shown by D in Figure 9. If the gesture is located near the line connecting the center of the face and the grayscale camera, and the size and shape of the gesture meet the occlusion criteria, it is determined that the face is obscured by a gesture, rather than the user leaving the interaction area. In this case, the frequency of face re-detection is increased, and the last known coordinates before the face disappeared are locked, maintaining continuous monitoring of that area.
[0127] If the gesture is not in front of the face, the computer device further analyzes whether the disappearance of the face was caused by the user turning or looking down. Using distance point cloud data of the area near the face provided by the gesture recognition module, the computer device performs depth scanning and data analysis on that area. By statistically analyzing the magnitude and area distribution of distance values, it searches for an area whose size roughly matches the face, and whose average distance value is close to the distance value recorded before the face disappeared. If such an area exists, the computer device determines that the user looked down or turned their head, rather than completely leaving the interaction area. In this case, the frequency of face re-detection is increased, and the coordinates before disappearance are locked, ready to quickly re-track when the user's head posture is restored. Simultaneously, to improve the user experience, the computer device can temporarily switch to 2D display mode to reduce visual confusion caused by missing depth information, until the face is successfully tracked again and 3D display is restored. If the aforementioned area does not exist—that is, the gesture is not in front of the face and no distance point cloud area matching the facial features can be found—the computer device confirms that the user has left the current interaction area and performs appropriate processing according to application requirements.
[0128] In one possible implementation, the method further includes: if no eye position is identified and the head position of the target object is within the acquisition range, determining a reference eye position range based on the last identified eye position; and if the eye position is re-identified, re-determining the eye position of the target object based on the reference eye position range.
[0129] The computer device can monitor the target object's eye position in real time, including its movement. Even when the target object's eye position cannot be captured, its head position remains within the computer device's line of sight to ensure that the eye position can be recaptured and tracked later. When the computer device fails to recognize the target object's eye position, the user temporarily shifts their attention away from the screen content. The computer device records the last valid eye position of the target object on the screen as a reference eye position range for subsequent operations.
[0130] During the above process, the computer equipment confirms that the target object's head position is still within the acquisition range. The acquisition range can be the same as or smaller than the reference range, sufficient to indicate whether the target object should lower or turn its head. When the target object's eye position is re-acquired, the computer equipment re-determines its eye position. Since directly re-determining the eye position from the new position across the entire image takes a considerable amount of time, the computer equipment utilizes the previously recorded reference eye position range as auxiliary information. The computer equipment combines the reference eye position range to define the range for re-determining the eye position. Based on this re-determined range and the currently captured eye position feature information, the computer equipment recalculates and determines the actual position of the target object's eyes using an algorithm model.
[0131] Based on the user's historical behavioral data (such as head movement speed and range), the system predicts the location where the user might re-enter the interactive area and proactively expands the tracking range and sensitivity in that area. When the user actually returns to the interactive area, the computer device is ready to re-establish the connection at a faster speed. In the event of eye position loss, the depth of field of the displayed image is limited to a fixed range (e.g., 5cm in / out of the screen). This introduces a smooth depth-of-field transition effect when the user temporarily loses their 3D visual experience due to tracking loss, mitigating visual discomfort that may be caused by sudden changes in depth of field. Simultaneously, after the user is tracked again, the original depth-of-field settings are gradually restored, enhancing the seamless transition of the user experience. The process of changing from a 3D image to a 2D image can be understood as gradually adjusting the depth of field of the image to 0.
[0132] In one possible implementation, the method further includes: determining the eye position trajectory of the target object based on the posture information acquired by the first camera when the eye position is not identified, the posture information including at least one of the target object's body posture or head posture; and determining the position where the target object's eye is re-acquired based on the eye position trajectory.
[0133] For example, a computer device can acquire the pose information of a target object in real time using a first camera. This pose information includes, but is not limited to, at least one of the target object's body pose or head pose. The computer device analyzes the pose information acquired by the first camera to extract key information related to changes in the target object's eye position. This key information includes, but is not limited to, head rotation angle, tilt degree, and translation distance. These key information collectively describe the dynamic changes in the target object's head or body. Based on these changes in pose information, the computer device uses mathematical models or machine learning algorithms to predict the trajectory of the target object's eye position change. The trajectory not only considers the current pose information but may also incorporate historical data, user behavior patterns, and other factors to improve the accuracy and reliability of the prediction.
[0134] Furthermore, the computer device calculates the possible intersections between the predicted eye position trajectory and the display screen boundary. These intersections represent possible locations where the target object's gaze re-enters the screen. If multiple intersections exist, the computer device can select the most likely new location for the target object's eye based on additional information (e.g., historical eye position data, screen content layout, etc.). Once the predicted re-entry location of the target object's eye position is determined, the computer device monitors the target object's actual eye position in real time and dynamically adjusts it as needed. When the actual eye position does indeed appear near the predicted location, the computer device considers the prediction successful and optimizes subsequent processing accordingly.
[0135] In one possible implementation, by predicting the reappearance position of the eyeball, the computer device can prepare the corresponding display content or perform the corresponding operation in advance, improving response speed and user experience. Accurately predicting changes in the user's gaze helps to achieve a more natural and fluid human-computer interaction experience, enhancing user immersion and satisfaction.
[0136] Referring to Figure 12, which illustrates the movement of a target object's head while the head is rotated, image processing algorithms (such as deep learning models) are used to accurately identify feature points on the top of the head from images captured by a camera. These feature points can be the boundary between hair and scalp, the vertices of the head's contour, etc. Stable tracking of the user's head is achieved through feature point matching or template matching techniques, continuing even during posture changes such as head tilting. Using the identified top-of-the-head feature points, combined with other possible facial feature points (such as ears, neck, etc., if visible), the head's attitude angles (pitch, yaw, roll) are predicted through geometric calculations or machine learning models. This attitude angle information is used to construct or update a 3D model of the user's head, which can be pre-established through a user calibration process.
[0137] Subsequently, using a 3D head model and the fixed positional relationship of the eyeballs relative to the head, combined with the head pose angle, the detailed coordinates of the eyeballs under occlusion were calculated. Considering the physiological range of eye movement, the calculated coordinates were validated for reasonableness to ensure the accuracy of the prediction results.
[0138] For human pose recognition, depth cameras or multi-camera computing devices are used in conjunction with pose recognition algorithms to capture real-time information about the user's torso, including the approximate positions of the limbs, torso, and head. This torso information is used to initially determine the head's position and pose, providing a rough but effective starting point for subsequent eye tracking. Building upon this torso information, the focus shifts to the head region, utilizing facial feature point recognition technology to obtain more refined head pose information. Based on the predicted eye coordinate range, the tracking area is expanded within the camera view, and the detection frequency is increased to ensure rapid eye tracking when the user looks up. Simultaneously, historical eye movement data is used for trajectory prediction, and potential 3D images are prepared or pre-rendered in advance to shorten the response time when the user re-enters the 3D interactive area.
[0139] This method analyzes a user's historical eye movement data to predict the time and location when the user's eyes are likely to enter the 3D interaction area. When the eye movement is predicted to enter the interaction area, i.e., when the eye position is re-captured, the computer device initiates the image rendering and display optimization process in advance to ensure that the correct 3D image is presented immediately upon re-capture of the eye position. This avoids the phenomenon of reverse gaze, improves the continuity of interaction and user experience, and allows users to experience a seamless and natural 3D interaction effect.
[0140] In one possible implementation, the three-dimensional image is a dual-viewpoint image based on eye position determination. The method further includes: converting the dual-viewpoint image into a multi-viewpoint image when the eye position is not identified; and controlling the display screen to gradually change the multi-viewpoint image into the corresponding dual-viewpoint image over a reference time when the eye position is re-identified.
[0141] Dual-viewpoint images are a common form of 3D display, simulating the parallax of the left and right sides of the human eye to create a sense of depth. Computer devices can monitor the eye position of a target object in real time. When the computer device cannot detect the eye position, to increase the user's observable range—that is, to allow the 3D image to be observed from multiple angles—the computer device converts the current dual-viewpoint image into a multi-viewpoint image. Compared to dual-viewpoint images, multi-viewpoint images provide more perspective information, enabling a more comprehensive display of the 3D scene.
[0142] In one possible implementation of this application, when the computer device detects the eye position of the target object, it switches the multi-view image back to a dual-view image. Optionally, the switching process may not be instantaneous, but rather gradually completed over a reference time. The computer device smoothly transitions the multi-view image to a dual-view image by gradually adjusting parameters such as the number of viewpoints and the magnitude of parallax. Displaying the multi-view image when the user temporarily leaves the screen expands the observation range of the three-dimensional stereoscopic image and provides a better visual experience when the user refocuses.
[0143] When a user looks down and temporarily leaves the 3D interactive area, the computer device takes a series of preprocessing and rendering measures to ensure a smooth transition back to the normal 3D display state when the user looks up again and re-enters the interactive area, while avoiding the phenomenon of reverse gaze. Sufficient adaptation time is provided for the user's re-entry by expanding the visible range and adjusting the rendering method. The computer device detects the user's head tilting down using eye tracking or facial recognition technology. After confirming that the user's head tilting down for a period of time, the computer device determines to enter multi-view ray tracing rendering mode. For example, if the computer device's SOC (System on a Chip) detects the loss of interactive user information, it sends a command to the FPGA (Field Programmable Gate Array) via a bus or dedicated interface to activate the multi-view rendering scheme.
[0144] Refer to Figure 13 for a schematic diagram of a multi-viewpoint image transformation process, including but not limited to the following steps 1301-1304.
[0145] Step 1301, Image preprocessing.
[0146] For example, after receiving a command, the FPGA first preprocesses the input binocular view to obtain a processed binocular view. This preprocessing includes, but is not limited to, camera calibration (determining the camera's intrinsic and extrinsic parameters) and binocular image correction (eliminating geometric distortion and parallax between cameras).
[0147] Step 1302, stereo matching.
[0148] In step 1302, the corrected binocular image can be processed using a binocular stereo matching algorithm to generate a disparity map, which is then converted into a depth map, thus obtaining the initial depth map. This application embodiment does not limit the binocular stereo matching algorithm; for example, a semi-global block matching (SGBM) algorithm can be used. The SGBM algorithm aims to estimate the disparity value of each pixel from the left and right view images and generate a high-quality disparity map. This process includes, but is not limited to, steps 13021-13024 (not shown in the figures).
[0149] Step 13021: Calculate the matching cost for each pixel of the corrected binocular image under different disparities.
[0150] In one possible implementation of this application, the matching cost calculation method includes, but is not limited to, using matching cost calculation based on SAD (Sum of Absolute Differences) or based on normalized Census transform. SAD is the sum of intensity differences of each pixel within the comparison window. SAD is simple and easy to calculate, but it is easily affected by changes in illumination. Census transform is a type of non-parametric image transform that can effectively detect local structural features in an image, such as edges and corner features. For example, Census transform converts the neighborhood of each pixel into a binary vector, and similarity is measured by comparing these vectors. Census transform is insensitive to changes in illumination and has strong robustness.
[0151] For example, for each pair of pixels (x, y) and disparity d, SGBM calculates the matching cost C(x, y, d) using the following formula.
[0152] Where W is the matching window, I L and I R These are the pixel values for the left and right views, respectively.
[0153] Step 13022: Approximate global optimization by accumulating matching costs along multiple directions.
[0154] In calculating the path cumulative cost, this application does not limit the number of directions, but calculates the cost accumulation in multiple directions, such as horizontal, vertical, and diagonal directions (45 degrees and 135 degrees). Typically, there are 8 or 16 directions (more directions can be added to improve accuracy). Then, for each path direction, the algorithm accumulates the cost within the disparity range, generating a path cost accumulation table. The path cost accumulation process can be represented as follows:
[0155] Where Lr(x, y, d) represents the cumulative cost in direction r, C(x, y, d) is the matching cost, and P1 and P2 are smoothing penalty parameters, used to penalize small disparity changes and large disparity changes, respectively.
[0156] Then, the cumulative costs in all directions are aggregated to obtain the total cost function S(x, y, d):
[0157] By aggregating the path costs from all directions through the above process, the impact of noise can be effectively reduced and a smooth disparity map can be generated.
[0158] Step 13023: After calculating the cumulative cost of each pixel under all disparities, select the disparity value with the minimum cost as the final disparity to obtain the initial disparity map.
[0159] Step 13023 generates an initial disparity map by minimizing the cost function.
[0160] Step 13024: Post-process the generated initial disparity map.
[0161] Since the initial disparity map generated in step 13023 may contain noise and incorrect matches, the method provided in this application embodiment further includes a post-processing step of the initial disparity map to further optimize the quality of the disparity map. This application embodiment does not limit the post-processing method; for example, the post-processing steps include, but are not limited to, at least one of the following.
[0162] Left-right consistency check: By comparing the disparity maps generated from the left and right views, inconsistent disparity points are eliminated, thereby removing incorrect matches.
[0163] Median filtering: Use a median filter to smooth the disparity map and remove isolated noise points.
[0164] Subpixel interpolation: By fitting the curve of the disparity cost function, the disparity value is accurately located to improve the accuracy of the disparity map.
[0165] The disparity map obtained through the above process, due to the advantages of combining local and global optimization through the SGBM algorithm, can generate a relatively smooth and accurate disparity map while retaining computational efficiency.
[0166] Step 1303, Depth Map Optimization.
[0167] After obtaining the initial depth map through step 1302 above, since the initial depth map may contain holes, that is, due to occlusion, weak texture, or other reasons, the depth information is missing, the parallax error is incorrect, etc. Therefore, the method provided in this application embodiment can optimize the initial depth map.
[0168] The depth map optimization process includes, but is not limited to, eliminating erroneous parallax, filling holes, and improving parallax accuracy, in order to ensure the accuracy and reliability of the depth map.
[0169] Step 1304, Virtual viewpoint reconstruction.
[0170] In step 1304, virtual viewpoint reconstruction is performed based on the optimized depth map to obtain a multi-view map. Virtual viewpoint reconstruction includes, but is not limited to, methods utilizing MBR (Model-Based Reconstruction) or IBR (Image-Based Rendering). Thus, images of multiple virtual viewpoints are reconstructed using the FPGA to obtain the multi-view map. Because the virtual viewpoint images cover a wider field of view, they can provide users with richer visual information.
[0171] Through steps 1301 to 1304 described above, a multi-view map is output, for example, the reconstructed multi-view map is output to a 3D display screen. Since the user may not have fully entered the interactive area at this time, the multi-view map can expand the visible range to a certain extent and reduce the visual discomfort of the user when re-entering.
[0172] After outputting the multi-view map, the computer device enters a waiting state, maintaining the current image until the user's eye position is detected again. When the user looks up and enters the interactive area, the computer device restarts the eye-tracking process to determine the exact position of the user's eyes. After successfully detecting the user's eye position, the SOC sends a recovery command to the FPGA. The FPGA then switches back to the normal 3D display mode and resumes precise rendering and display based on the user's eye position.
[0173] This solution utilizes multi-view ray tracing rendering technology to provide a smooth transition display when the user looks down, effectively avoiding the phenomenon of reverse gaze. Simultaneously, through the high-efficiency processing capabilities of the FPGA and preprocessing optimization techniques, it achieves rapid response and accurate recovery upon user return.
[0174] In summary, the eye position recognition method provided in this application prioritizes determining the location range of the target object's eyeball. The process of the second camera acquiring feature point information can then be performed within this location range, thereby narrowing the acquisition range of feature point information. This reduces the amount of data that needs to be processed for subsequent determination of the overall feature point information of the eyeball and improves the efficiency of determining the eyeball position.
[0175] Furthermore, if the target's eye position is not detected, the 3D display is switched to a 2D display. Upon re-detection of the eye position, the 2D display is gradually switched back to a 3D display, thus avoiding sudden 3D visual disturbances to the user. Additionally, the system can predict the re-detected eye position when the user tilts or turns their head, thereby reducing the time required to re-determine the eye position.
[0176] Referring to Figure 14, which is a schematic diagram of an eye position recognition device provided in an embodiment of this application, the device includes:
[0177] The acquisition module 1401 is used to acquire the position information of the target object collected by multiple first cameras at different angles. The multiple first cameras include a first main camera and a first auxiliary camera. The position information collected by the first main camera has a greater weight than the position information collected by the first auxiliary camera.
[0178] The determination module 1402 is used to determine the location range of the target object's eyeball based on the location information and weights collected by multiple first cameras;
[0179] The acquisition module 1401 is also used to acquire feature point information of the target object collected by multiple second cameras within the position range. The multiple second cameras include a second main camera and a second auxiliary camera. The weight of the feature point information collected by the second main camera is greater than the weight of the feature point information collected by the second auxiliary camera.
[0180] The determination module 1402 is also used to determine the eye position of the target object based on the feature point information and weights collected by multiple second cameras.
[0181] In one possible implementation, the determining module 1402 is further configured to, when the position information acquired by the first main camera indicates that the target object is occluded, select at least one from the first auxiliary cameras as an updated first main camera, determine the eye position based on the position information acquired by the updated first main camera, wherein the position information acquired by the updated first main camera indicates that the target object is not occluded; and / or,
[0182] If the feature point information acquired by the second main camera indicates that the target object is occluded, at least one of the second auxiliary cameras is selected as the updated second main camera. The eye position is determined based on the feature point information acquired by the updated second main camera. The feature point information acquired by the updated second main camera indicates that the target object is not occluded.
[0183] In one possible implementation, the device further includes:
[0184] The first control module is used to determine the resolution of the display screen based on the eye position and control the display screen to display a three-dimensional image that matches the eye position according to the resolution.
[0185] In one possible implementation, the three-dimensional image includes a dual-viewpoint image determined based on the eye position. The first control module is further configured to convert the dual-viewpoint image into a multi-viewpoint image if the eye position is not identified; and to control the display screen to gradually change the multi-viewpoint image into the corresponding dual-viewpoint image over a reference time if the eye position is re-identified.
[0186] In one possible implementation, the device further includes:
[0187] The second control module is used to control the display screen to display a two-dimensional image when the eye position is not detected; and to control the display screen to gradually change the two-dimensional image into the corresponding three-dimensional image within a reference time when the eye position is re-detected.
[0188] In one possible implementation, the determining module 1402 is further configured to determine the distance between the eyeball and the display screen, determine the linear adjustment parameter based on the distance, and perform a depth-of-field linear transformation on the two-dimensional image as a whole through the linear adjustment parameter to obtain a three-dimensional image.
[0189] In one possible implementation, the two-dimensional image includes a display subject and a background, the display subject and the background have different depths of field, and the background is blurred; the second control module is used to adjust the depth of field of the display subject and the background so that the background recovers to the depth of field before the blurring process after a reference time.
[0190] In one possible implementation, the determining module 1402 is further configured to determine a reference eye position range based on the last determined eye position when the eye position is not identified and the head position of the target object is within the acquisition range; and to re-determine the eye position of the target object based on the reference eye position range when the eye position is re-identified.
[0191] In one possible implementation, the determining module 1402 is further configured to determine the eye position trajectory of the target object based on the posture information acquired by the first camera when the eye position is not identified, the posture information including at least one of the target object's body posture or head posture; and determine the position where the target object's eye is re-acquired based on the eye position trajectory.
[0192] It should be noted that the eye position recognition device provided in the embodiment of Figure 14 above is only illustrated by the division of the above-described functional modules. In actual operation, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the device and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments.
[0193] Figure 15 is a schematic diagram of a server provided in an embodiment of this application. The server can vary considerably depending on its configuration or performance, and may include one or more processors 1501 and one or more memories 1502. Each memory 1502 stores at least one computer program, which is loaded and executed by the processors 1501 to enable the server to implement the eye position recognition method provided in the various method embodiments described above. Of course, the server may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The server may also include other components for implementing device functions, which will not be elaborated upon here.
[0194] Figure 16 is a schematic diagram of a terminal provided in an embodiment of this application, enabling the terminal to implement the eye position recognition methods provided in the above-described method embodiments. The terminal may be, for example, a smartphone, tablet computer, media player, laptop computer, or desktop computer. The terminal may also be referred to as a user device, portable terminal, laptop terminal, desktop terminal, or other names.
[0195] Typically, a terminal includes a processor 1601 and a memory 1602.
[0196] Processor 1601 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 1601 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Processor 1601 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 1601 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content required to be displayed on the screen. In some embodiments, processor 1601 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.
[0197] The memory 1602 may include one or more computer-readable storage media, which may be non-transitory. The memory 1602 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1602 is used to store at least one instruction, which is executed by the processor 1601 to cause the terminal to implement the eye position recognition method provided in the method embodiments of this application.
[0198] In some embodiments, the terminal may also optionally include: a peripheral device interface 1603 and at least one peripheral device. The processor 1601, memory 1602, and peripheral device interface 1603 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 1603 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of: a radio frequency circuit 1604, a display screen 1605, a camera assembly 1606, an audio circuit 1607, and a power supply 1608.
[0199] Peripheral interface 1603 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 1601 and memory 1602. In some embodiments, processor 1601, memory 1602 and peripheral interface 1603 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 1601, memory 1602 and peripheral interface 1603 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.
[0200] The radio frequency (RF) circuit 1604 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 1604 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 1604 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 1604 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 1604 can communicate with other terminals through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: metropolitan area networks (MANs), various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks (WLANs), and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 1604 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application.
[0201] Display screen 1605 is used to display a UI (User Interface). This UI may include graphics, text, icons, video, and any combination thereof. When display screen 1605 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 1601 for processing. In this case, display screen 1605 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1605, located on the front panel of the terminal; in other embodiments, there may be at least two display screens, respectively located on different surfaces of the terminal or in a folded design; in other embodiments, display screen 1605 may be a flexible display screen, located on a curved or folded surface of the terminal. Furthermore, display screen 1605 may be configured as a non-rectangular, irregular shape, i.e., a non-rectangular screen. Display screen 1605 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).
[0202] The camera assembly 1606 is used to acquire images or videos. Optionally, the camera assembly 1606 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the terminal, and the rear-facing camera is located on the back of the terminal. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, the camera assembly 1606 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.
[0203] The audio circuit 1607 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 1601 for processing, or input to the radio frequency circuit 1604 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each positioned at a different location on the terminal. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert electrical signals from the processor 1601 or the radio frequency circuit 1604 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 1607 may also include a headphone jack.
[0204] The power supply 1608 is used to power the various components in the terminal. The power supply 1608 can be AC power, DC power, a disposable battery, or a rechargeable battery. When the power supply 1608 includes a rechargeable battery, the rechargeable battery can support wired or wireless charging. The rechargeable battery can also be used to support fast charging technology.
[0205] In some embodiments, the terminal further includes one or more sensors 1609. The one or more sensors 1609 include, but are not limited to: an accelerometer 1610, a gyroscope 1611, a pressure sensor 1612, an optical sensor 1613, and a proximity sensor 1614.
[0206] Accelerometer 1610 can detect the magnitude of acceleration along the three coordinate axes of a coordinate system established by the terminal. For example, accelerometer 1610 can be used to detect the components of gravitational acceleration along the three coordinate axes. Processor 1601 can control display screen 1605 to display the user interface in either a landscape or portrait view based on the gravitational acceleration signal acquired by accelerometer 1610. Accelerometer 1610 can also be used for games or for acquiring user motion data.
[0207] The gyroscope sensor 1611 can detect the terminal's orientation and rotation angle. The gyroscope sensor 1611 can work in conjunction with the accelerometer sensor 1610 to collect the user's 3D movements on the terminal. Based on the data collected by the gyroscope sensor 1611, the processor 1601 can perform the following functions: motion sensing (e.g., changing the UI based on the user's tilt), image stabilization during shooting, game control, and inertial navigation.
[0208] The pressure sensor 1612 can be disposed on the side bezel of the terminal and / or the lower layer of the display screen 1605. When the pressure sensor 1612 is disposed on the side bezel of the terminal, it can detect the user's grip signal on the terminal, and the processor 1601 can perform left / right hand recognition or quick operation based on the grip signal collected by the pressure sensor 1612. When the pressure sensor 1612 is disposed on the lower layer of the display screen 1605, the processor 1601 can control the operable controls on the UI interface based on the user's pressure operation on the display screen 1605. The operable controls include at least one of button controls, scroll bar controls, icon controls, and menu controls.
[0209] Optical sensor 1613 is used to collect ambient light intensity. In one embodiment, processor 1601 can control the display brightness of display screen 1605 based on the ambient light intensity collected by optical sensor 1613. Specifically, when the ambient light intensity is high, the display brightness of display screen 1605 is increased; when the ambient light intensity is low, the display brightness of display screen 1605 is decreased. In another embodiment, processor 1601 can also dynamically adjust the shooting parameters of camera assembly 1606 based on the ambient light intensity collected by optical sensor 1613.
[0210] The proximity sensor 1614, also known as a distance sensor, is typically installed on the front panel of the terminal. The proximity sensor 1614 is used to detect the distance between the user and the front of the terminal. In one embodiment, when the proximity sensor 1614 detects that the distance between the user and the front of the terminal is gradually decreasing, the processor 1601 controls the display screen 1605 to switch from a screen-on state to a screen-off state; when the proximity sensor 1614 detects that the distance between the user and the front of the terminal is gradually increasing, the processor 1601 controls the display screen 1605 to switch from a screen-off state to a screen-on state.
[0211] Those skilled in the art will understand that the structure shown in Figure 16 does not constitute a limitation on the terminal, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.
[0212] In an exemplary embodiment, a computer device is also provided, comprising a processor and a memory storing at least one computer program. The at least one computer program is loaded and executed by one or more processors to enable the computer device to implement any of the aforementioned methods for recognizing eye positions.
[0213] In an exemplary embodiment, a computer-readable storage medium is also provided, which stores at least one computer program that is loaded and executed by a processor of a computer device to enable the computer to implement any of the above-described methods for recognizing eye positions.
[0214] In one possible implementation, the aforementioned computer-readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, or optical data storage device, etc. Alternatively, the computer-readable storage medium can be a non-transitory computer-readable storage medium.
[0215] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform any of the above-described methods for recognizing eye positions.
[0216] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the first and second information involved in this application were obtained with full authorization.
[0217] It should be understood that "multiple" as used in this article refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0218] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the principles of this application should be included within the protection scope of this application.
Claims
1. A method for identifying eyeball position, characterized in that, The method includes: The position information of the target object is acquired by multiple first cameras at different angles. The multiple first cameras include a first main camera and a first auxiliary camera. The position information acquired by the first main camera has a greater weight than the position information acquired by the first auxiliary camera. The location range of the target object's eyeball is determined based on the location information and weights acquired by the multiple first cameras; The feature point information of the target object collected by multiple second cameras within the specified location range is acquired. The multiple second cameras include a second main camera and a second auxiliary camera. The weight of the feature point information collected by the second main camera is greater than the weight of the feature point information collected by the second auxiliary camera. The eye position of the target object is determined based on the feature point information and weights acquired by the multiple second cameras.
2. The method according to claim 1, characterized in that, The method further includes: If the position information acquired by the first main camera indicates that the target object is occluded, at least one of the first auxiliary cameras is selected as the updated first main camera. The eye position is determined based on the position information acquired by the updated first main camera, wherein the position information acquired by the updated first main camera indicates that the target object is not occluded; and / or, If the feature point information acquired by the second main camera indicates that the target object is occluded, at least one of the second auxiliary cameras is selected as the updated second main camera, and the eye position is determined based on the feature point information acquired by the updated second main camera. The feature point information acquired by the updated second main camera indicates that the target object is not occluded.
3. The method according to claim 1 or 2, characterized in that, After determining the eye position of the target object based on the feature point information and weights acquired by the plurality of second cameras, the method further includes: The resolution of the display screen is determined based on the eye position, and the display screen is controlled to display a three-dimensional image that matches the eye position according to the resolution.
4. The method according to claim 3, characterized in that, The three-dimensional image includes a dual-viewpoint image determined based on eye position, and the method further includes: If the eye position is not identified, the dual-viewpoint image is converted into a multi-viewpoint image; Upon re-identifying the eye position, the display screen is controlled to gradually change the multi-viewpoint image into the corresponding dual-viewpoint image over a reference time period.
5. The method according to claim 1 or 2, characterized in that, The method further includes: Without recognizing the position of the eyeball, the display screen is controlled to display a two-dimensional image; Once the eye position is re-identified, the display screen is controlled to gradually change the two-dimensional image into a corresponding three-dimensional image over a reference time period.
6. The method according to claim 5, characterized in that, Before controlling the display screen to gradually change the two-dimensional image into the corresponding three-dimensional image within a reference time, the method further includes: Determine the distance between the eyeball and the display screen, and determine the linear adjustment parameters based on the distance; The three-dimensional image is obtained by performing a linear depth transformation on the two-dimensional image as a whole using the linear adjustment parameters.
7. The method according to claim 5, characterized in that, The two-dimensional image includes a display subject and a background, wherein the display subject and the background have different depths of field, and the background is blurred; controlling the display screen to gradually change the two-dimensional image into a corresponding three-dimensional image within a reference time includes: The depth of field between the displayed subject and the background is adjusted so that the background recovers to the depth of field before the blurring process after a reference time.
8. The method according to claim 1 or 2, characterized in that, The method further includes: If the eye position is not identified and the head position of the target object is within the acquisition range, the reference eye position range is determined based on the last determined eye position. If the eye position is re-identified, the eye position of the target object is re-determined based on the reference eye position range.
9. The method according to claim 1 or 2, characterized in that, The method further includes: In the absence of eye position recognition, the eye position trajectory of the target object is determined based on the posture information acquired by the first camera, wherein the posture information includes at least one of the target object's body posture or head posture; The location where the target object's eyeballs are re-collected is determined based on the eyeball position trajectory.
10. A device for recognizing the position of an eyeball, characterized in that, The device includes: The acquisition module is used to acquire the position information of the target object collected by multiple first cameras at different angles. The multiple first cameras include a first main camera and a first auxiliary camera. The position information collected by the first main camera has a greater weight than the position information collected by the first auxiliary camera. The determination module is used to determine the location range of the target object's eyeball based on the location information and weights acquired by the plurality of first cameras; The acquisition module is further configured to acquire feature point information of the target object collected by multiple second cameras within the location range. The multiple second cameras include a second main camera and a second auxiliary camera. The weight of the feature point information collected by the second main camera is greater than the weight of the feature point information collected by the second auxiliary camera. The determining module is further configured to determine the eye position of the target object based on the feature point information and weights acquired by the plurality of second cameras.
11. A computer device, characterized in that, The computer device includes a processor and a memory, the memory storing at least one computer program, the at least one computer program being loaded and executed by the processor to enable the computer device to implement the eye position recognition method as described in any one of claims 1 to 9.
12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer program, which is loaded and executed by a processor to enable the computer to implement the eye position recognition method as described in any one of claims 1 to 9.
13. A computer program product, characterized in that, The computer program product includes computer instructions stored in a computer-readable storage medium, a processor of a computer device reading the computer instructions from the computer-readable storage medium, and the processor executing the computer instructions to enable the computer device to implement the eye position recognition method as described in any one of claims 1 to 9.