Method, device and equipment for processing target object's sight line information
The camera component of the head-mounted device analyzes pupil information and display screen data, calculates the coordinates of the eye's rotation center, and determines the gaze point based on the preset trajectory. This solves the problem of inaccurate gaze point in the existing technology and achieves more accurate gaze point recognition.
Patent Information
- Application Number
- CN202311464177.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-06
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-11-06
AI Technical Summary
Existing gaze point recognition methods analyze the pupil position and display screen content, which leads to inaccurate gaze points.
A head-mounted device is used, equipped with a first camera component and a second camera component. By acquiring video data of the eyes and the display screen, analyzing information such as the pupil centroid and the direction of the maximum diameter of the ellipse, the coordinates and distance of the eye's rotation center are calculated, and the gaze point is determined based on the preset trajectory.
The accuracy of the gaze point is improved, the error caused by individual differences is reduced, the calibration process is simplified, the influence of head movement is eliminated, and more accurate gaze point determination is achieved.
Smart Images

Figure CN117475501B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer technology, and in particular to a method, device and apparatus for processing line of sight information of a target object. Background Art
[0002] Eye tracking is a technology that observes eye movements and changes, and can determine the gaze point by identifying the eyes.
[0003] Existing gaze point recognition analyzes the pupil position of the target object and the content on the display screen to determine the gaze point. However, the gaze point obtained by using the above method is inaccurate. Summary of the Invention
[0004] The embodiments of the present application provide a method, apparatus, and device for processing the gaze information of a target object, which can more accurately analyze the gaze point.
[0005] The technical solution is as follows:
[0006] In a first aspect, the present application provides a method for processing line of sight information of a target object, the method completing the processing based on a head-mounted device worn on the target object, the head-mounted device including a first camera component for photographing the target object's eyes and a second camera component for photographing a display screen located in front of the target object, the method comprising: obtaining first video data related to the target object's eyes captured by the first camera component; extracting first image data and second image data from the first video data, and determining pupil analysis information of the first image data and the second image data based on the first image data and the second image data, the pupil analysis information including the horizontal coordinate of the pupil centroid of the target object's eyes on the vertical plane and vertical coordinates, and the maximum diameter direction of the ellipse fitted to the pupil of the target object; calculating the horizontal coordinate and vertical coordinate of the eye rotation center based on the pupil analysis information of the first image data and the second image data as the first parameter information; judging whether the target object is observing the calibration point displayed on the display screen in front of the target object, and obtaining the third image data of the first camera component and the fourth image data of the second camera component when the target object observes the calibration point; analyzing based on the third image data, the fourth image data and the first parameter information, determining the distance between the target object's eye rotation center and the pupil as the second parameter information, so as to determine the target object's gaze point based on the first parameter information and the second parameter information of the two eyes.
[0007] Preferably, the method of determining whether the target object is at the calibration point displayed on the display screen in front of the target object includes: displaying guidance information of a preset trajectory on the display screen, the calibration point being located within the preset trajectory of the guidance information; extracting the movement trajectory of the pupil centroid from the first video data; and determining whether the target object is at the calibration point displayed on the display screen in front of the target object based on whether the movement trajectory of the pupil centroid conforms to the preset trajectory.
[0008] Preferably, the preset trajectory is a cross-shaped trajectory, and the calibration point is an extreme point of the cross-shaped trajectory, and the extreme point of the cross-shaped trajectory includes the end point or the center point of the cross.
[0009] Preferably, the analysis based on the third image data, the fourth image data and the first parameter information to determine the distance between the target object's eye rotation center and the pupil as the second parameter information includes: determining the first analysis data of the target object based on the third image data of the first camera component, the first analysis data including the maximum diameter direction of the ellipse fitted to the pupil of the target object and the ratio of the maximum diameter to the shortest diameter; determining the predicted coordinate information of the calibration point on the display screen based on the fourth image data of the second camera component, and forming the second analysis data in combination with the actual coordinate information of the calibration point; determining the distance between the target object's eye rotation center and the pupil as the second parameter information based on the first analysis data, the second analysis data and the first parameter information.
[0010] Preferably, the method of determining the distance between the target object's eye rotation center and the pupil as the second parameter information based on the first analysis data, the second analysis data and the first parameter information includes: obtaining first component parameters of the first camera component, the first component parameters including first pixel information, a first virtual working distance and first posture information of the first camera component; obtaining second component parameters of the second camera component, the second component parameters including second pixel information, a second virtual working distance and second posture information of the second camera component; and determining the distance between the target object's eye rotation center and the pupil as the second parameter information based on the first analysis data, the first component parameters, the second analysis data, the second component parameters and the first parameter information.
[0011] Preferably, two first camera assemblies are provided, the first camera assemblies are located in front of and below the eyes and facing directly rearward, the spacing between the two first camera assemblies is a first distance, the upward rotation angle of the first camera assembly is a first angle, and the first angle is between 30-45°; the second camera assembly is located above the center of the two first camera assemblies, the vertical spacing between the second camera assembly and the first camera assembly is a second distance, the horizontal spacing is a third distance, and it faces directly in front of the target object.
[0012] Preferably, determining the gaze point of the target object based on the first parameter information and the second parameter information of the two eyes includes: obtaining a first analysis image taken by two first camera components and a second analysis image taken by the second camera component; and performing analysis based on the two first analysis images, the second analysis image, the first parameter information and the second parameter information of the two eyes to determine the gaze point of the target object.
[0013] Preferably, the analysis based on the first analysis image, the second analysis image, the first parameter information and the second parameter information to determine the gaze point of the target object includes: determining the first gaze data of the target object based on the first analysis images of the two first camera components, the first gaze data including the maximum diameter direction of the ellipse fitted to the pupils of the target object's eyes and the ratio of the maximum diameter to the shortest diameter; determining the predicted gaze coordinates of each related point based on the second analysis image of the second camera component to form second gaze data; and determining the gaze point of the target object based on the first gaze data and the second gaze data.
[0014] In a second aspect, the present application provides a device for processing line of sight information of a target object, wherein the device completes processing based on a head-mounted device worn on the target object, wherein the head-mounted device includes a first camera component for photographing the eyes of the target object and a second camera component for photographing a display screen located in front of the target object, and the device includes: a first data acquisition module for acquiring first video data related to the eyes of the target object collected by the first camera component; a pupil information acquisition module for extracting first image data and second image data from the first video data, and determining pupil analysis information of the first image data and the second image data based on the first image data and the second image data, wherein the pupil analysis information includes the horizontal coordinate and vertical coordinate of the pupil centroid of the target object's eye on the vertical plane coordinates, and the maximum diameter direction of the ellipse fitted to the pupil of the target object; a first parameter acquisition module, used to calculate the horizontal coordinate and the vertical coordinate of the eye rotation center based on the pupil analysis information of the first image data and the second image data, as the first parameter information; a calibration image acquisition module, used to determine whether the target object is observing the calibration point displayed on the display screen in front of the target object, and obtain third image data of the first camera component and fourth image data of the second camera component when the target object observes the calibration point; a second parameter acquisition module, used to analyze based on the third image data and the fourth image data, and determine the distance between the target object's eye rotation center and the pupil as the second parameter information, so as to determine the target object's gaze point based on the first parameter information and the second parameter information of the two eyes.
[0015] In a third aspect, the present application provides a device comprising: a memory, a transceiver, and a processor; wherein the memory is used to store a computer program; the transceiver is used to send and receive data under the control of the processor; and the processor is used to read the computer program in the memory and execute the method described in the first aspect.
[0016] In a fourth aspect, the present application provides a storage medium having a computer program stored thereon, which implements the method described in the first aspect when the computer program is executed by a processor.
[0017] The beneficial effects of the technical solution provided by this application are:
[0018] The solution of the present application can be applied to scenarios involving identifying the gaze point of a target subject's eyes. A head-mounted device can be used to calibrate the target subject's gaze point, and then the target subject's gaze point can be determined based on the calibrated parameters and pupil-related data. This solution can be performed using a head-mounted device worn by the target subject, comprising a first camera assembly for capturing the target subject's eyes and a second camera assembly for capturing a display screen located in front of the target subject. This solution can analyze pupil analysis information of the target subject's eyes using multiple images from first video data captured by the first camera assembly, and determine the horizontal and vertical coordinates of the target subject's eye rotation center. The solution then analyzes whether the target subject is gazing at the calibration point. If so, images from the first and second camera assemblies can be extracted and analyzed to determine the distance between the target subject's eye rotation center and the target subject's pupil. The horizontal and vertical coordinates of the eye rotation center are then combined to form the three-dimensional coordinates of the target subject's eye rotation center. Subsequent analysis of the target subject's gaze point can be performed based on the three-dimensional coordinates of the target subject's eye rotation center, the first analyzed image from the first camera assembly, and the second analyzed image from the second camera assembly. This solution can determine the three-dimensional coordinates of the target object's eye rotation center during the calibration phase. In the subsequent gaze point analysis, the three-dimensional coordinates of the eye rotation center, pupil information (determined based on the first analysis image of the first camera component) and external environment information (determined based on the second analysis image of the second camera component) are considered to more accurately determine the gaze point.
[0019] Specifically, this solution can acquire first video data related to a target subject's eye, captured by a first camera assembly; extract first image data and second image data from the first video data, wherein the first and second image data have different, but not opposite, eye orientations. Based on the first and second image data, pupil analysis information for the first and second image data is determined. The pupil analysis information includes the horizontal and vertical coordinates of the target subject's pupil centroid on a vertical plane and the maximum diameter of an ellipse fitted to the target subject's pupil. The first and second image data correspond to one eye of the target subject; the calculation method for the other eye is similar and will not be further described here. Based on the pupil analysis information obtained from each of the first and second image data, the horizontal and vertical coordinates of the eye's rotation center are calculated as first parameter information. Subsequently, it can be determined whether the target subject is observing a calibration point displayed on a display screen in front of the target subject. Third image data from the first and second camera assembly and fourth image data from the second camera assembly are acquired when the target subject is observing the calibration point. Based on the third and fourth image data and the first parameter information, analysis is performed to determine the distance between the target subject's eye rotation center and the pupil as second parameter information. After determining the first parameter information and the second parameter information of one eye of the target object, a similar method can be used to determine the first parameter information and the second parameter information of the other eye of the target object, so as to determine the gaze point of the target object based on the first parameter information and the second parameter information of the two eyes. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments of the present application.
[0021] Figure 1 is a schematic structural diagram of a head-mounted device according to an embodiment of the present application;
[0022] Figure 2 1 is a flow chart of a method for processing sight line information of a target object according to an embodiment of the present application;
[0023] Figure 3 This is a schematic structural diagram of a target object's line of sight information processing device according to one embodiment of the present application;
[0024] Figure 4 This is a structural block diagram of a device according to an embodiment of the present application;
[0025] Figure 5 This is a structural block diagram of a user equipment according to an embodiment of the present application; DETAILED DESCRIPTION
[0026] The following describes in detail embodiments of the present application, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout identify the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application, and are not to be construed as limiting the present application.
[0027] Those skilled in the art will understand that, unless otherwise specified, the singular forms "a," "an," "the," and "the" used herein may also include the plural forms, and "a plurality" refers to two or more, and other quantifiers are similar. It should be further understood that the term "comprising" used in the specification of this application refers to the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof. It should be understood that when we refer to an element as being "connected" or "coupled" to another element, it can be directly connected or coupled to the other element, or there may be intermediate elements. In addition, "connected" or "coupled" as used herein may include wireless connection or wireless coupling. The term "and / or" used herein describes the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the associated objects are in an "or" relationship.
[0028] The solution of the present application can be applied to scenarios where the gaze point of a target object's eyes is identified. A head-mounted device can be used to calibrate the target object's gaze point, and then the target object's gaze point is determined based on the calibrated parameters and pupil-related data of the target object. This solution can be performed based on a head-mounted device worn on the target object, wherein the head-mounted device includes a first camera assembly for photographing the target object's eyes and a second camera assembly for photographing a display screen located in front of the target object. The target object can be an animal or a person.
[0029] The first camera assembly is also called the eye-tracking camera, and the second camera assembly is also called the central camera. Specifically, the first camera assembly is provided with two, such as Figure 1As shown, two first camera assemblies are located below and in front of the target subject's eyes, facing directly behind them. The distance between the two first camera assemblies is a first distance 2d, and the first camera assemblies can be rotated upwards by a first angle γ, which is between 30 and 45°. A second camera assembly is located above the center of the two first camera assemblies, with a vertical distance H and a horizontal distance su between the second camera assembly and the first camera assembly, facing directly in front of the target subject. The eye-tracking camera of the first camera assembly has a pixel width we, a height he, and a virtual working distance de. The pixel width of the second camera assembly is 1280, the height hu is 720, and the virtual working distance du. The display screen is located 0.5-0.75 meters in front of the target subject. This solution allows pre-recording of the first component parameters of the first camera assembly and the second component parameters of the second camera assembly, allowing data analysis based on the first and second component parameters. The first component parameters include the first pixel information, the first virtual working distance, and the first pose information of the first camera assembly. The second component parameters include the second pixel information, the second virtual working distance, and the second pose information of the second camera assembly.
[0030] This solution can analyze pupil analysis information of a target subject's eyes using multiple images from the first video data captured by the first camera assembly, and determine the horizontal and vertical coordinates of the target subject's eye rotation center. It also analyzes whether the target subject is gazing at a calibration point. If so, images from the first and second camera assemblies can be extracted and analyzed to determine the distance between the target subject's eye rotation center and the target subject's pupil. Combining the horizontal and vertical coordinates of the eye rotation center, the three-dimensional coordinates of the target subject's eye rotation center can be used to determine the target subject's gaze point. During subsequent analysis of the target subject's gaze point, the three-dimensional coordinates of the target subject's eye rotation center, the first analysis image from the first camera assembly, and the second analysis image from the second camera assembly can be used to determine the target subject's gaze point. This solution can determine the three-dimensional coordinates of the target subject's eye rotation center during the calibration phase. During subsequent gaze point analysis, the three-dimensional coordinates of the eye rotation center, pupil information (determined based on the first analysis image from the first camera assembly), and external environment information (determined based on the second analysis image from the second camera assembly) can be considered to more accurately determine the gaze point.
[0031] Specifically, this solution can acquire first video data related to a target subject's eyes, captured by a first camera assembly; and extract first image data and second image data from the first video data, where the eyes are oriented in different, but not opposite, directions. Based on the first and second image data, pupil analysis information for the first and second image data is determined, the pupil analysis information including the horizontal and vertical coordinates of the pupil centroid of the target subject's eyes on a vertical plane and the maximum diameter direction of an ellipse fitted to the target subject's pupil. The first and second image data represent images of one eye of the target subject, and the calculation method for the other eye is similar and will not be repeated here.
[0032] In an optional embodiment, the pupil analysis information may include the horizontal coordinate x and vertical coordinate y of the pupil centroid on the vertical plane, the maximum diameter d of the pupil-fitted ellipse of the target object, the direction α of the maximum diameter, and the ratio of the maximum diameter to the shortest diameter ra. The maximum diameter d can be regarded as the long diameter of the pupil-fitted ellipse. Because the ellipse fitting speed is slow and the error is large, the Feret diameter can actually be used. The direction α of the maximum diameter is the angle between the Feret long diameter (the long diameter of the ellipse) and the positive direction of the x-axis, both of which are positive and are angles rather than radians. The ratio ra is the Feret longest diameter / Feret shortest diameter (the long diameter of the ellipse / the short diameter of the ellipse), with a value ≥ 1.
[0033] This solution involves multiple coordinate systems, and data between these systems can be converted to each other. The first camera assembly can correspond to the first image coordinate system, or eye movement coordinate system e, with left and right cameras eL and eR, respectively. The second camera assembly can correspond to the second image coordinate system, or central camera coordinate system u. The image coordinates are centered at the upper left point, with the x-axis pointing right and the y-axis pointing downward. A standard coordinate system c can also be used to convert between the first and second image coordinate systems. The origin of c is at the midpoint between the left and right first camera assemblies, and its orientation is the same as u.
[0034] The abscissa and ordinate of the eye rotation center are calculated based on pupil analysis information obtained from analyzing the first and second image data, and used as first parameter information. The abscissa and ordinate of the eye rotation center for the left eye can be expressed as (xe0L, ye0L), where L represents the left eye and R represents the right eye. The eye rotation center can be calculated according to the following formulas 1 and 2.
[0035]
[0036]
[0037] Among them, x1, y1, α1 are pupil analysis information of the first image, and x2, y2, α2 are pupil analysis information of the second image.
[0038] This solution can display guidance information with a preset trajectory on a display screen and obtain the movement trajectory of the target object's pupil center of mass from the first video data. It can then determine whether the movement trajectory of the pupil center of mass conforms to the preset trajectory to determine whether the target object is observing a calibration point displayed on the display screen in front of the target object. The calibration point is located within the preset trajectory of the guidance information. The preset trajectory is a cross-shaped trajectory. The calibration point is the extreme point of the cross-shaped trajectory. The extreme point of the cross-shaped trajectory includes the endpoints or center point of the cross. When the target object observes the calibration point, third image data from the first camera component and fourth image data from the second camera component are obtained. The third image data and the fourth image data are data from the same time, which is the time when the target object is looking at the calibration point. After determining the third image data and the fourth image data, the third image data, the fourth image data, and the first parameter information can be analyzed to determine the distance between the rotation center of the target object's eye and the pupil as the second parameter information. The first analysis data of the target object can be determined based on the third image data of the first camera component, and the first analysis data includes the maximum diameter direction of the ellipse fitted to the pupil of the target object and the ratio of the maximum diameter to the shortest diameter; the predicted coordinate information of the calibration point on the display screen is determined based on the fourth image data of the second camera component, and combined with the actual coordinate information of the calibration point to form the second analysis data; based on the first analysis data, the second analysis data and the first parameter information, the distance between the rotation center of the target object's eye and the pupil is determined as the second parameter information.
[0039] After determining the first parameter information and the second parameter information of one eye of the target object, a similar method can be used to determine the first parameter information and the second parameter information of the other eye of the target object, so as to determine the gaze point of the target object based on the first parameter information and the second parameter information of the two eyes.
[0040] In an optional embodiment, the present solution can determine the distance between the target subject's eye rotation center and the pupil as the second parameter information based on the first analysis data, the first component parameter, the second analysis data, the second component parameter, and the first parameter information. The second parameter information can be determined according to the following formulas 4, 5, and 6.
[0041]
[0042]
[0043]
[0044] The spacing between the two first camera assemblies is 2d, where d is half the spacing between the two first camera assemblies. dp is the conversion parameter between pixel distance and real distance, used to convert pixel distance into real distance. dpe represents the eye movement coordinate system, and dpu represents the central camera coordinate system. x, y, and z are the spatial positions of the calibration points, using the standard coordinate system C. xu3 and yu3 are the coordinates of the calibration points captured by the second camera assembly. Other symbols represent the same information as in the above embodiment and are not repeated here.
[0045] After calculating the second parameter for the target subject's left eye using the above method, the second parameter for the right eye (ze0R') can be calculated using the same method, simply reversing the sign of d and substituting the corresponding value on the right. After determining the three-dimensional coordinates of the eye's rotation center for both eyes, the gaze point can be calculated. This can be calculated using the following formulas 7, 8, and 9. By combining formulas 7, 8, and 9, and substituting xe0L, ye0L, ze0L', xe0R, ye0R, and ze0R' calculated during the calibration step, we can calculate xu and yu.
[0046]
[0047]
[0048]
[0049] Where ra is the ratio of the major to minor diameters of the target object's current pupil fitting ellipse, α is the orientation of the target object's current pupil fitting ellipse, L represents the left eye, and R represents the right eye. The information represented by other symbols is the same as in the above embodiment and will not be repeated here. The obtained (xu, yu) is the target object's current gaze point.
[0050] Specifically, an embodiment of the present application provides a method for processing the sight line information of a target object, wherein the method performs processing based on a head-mounted device worn on the target object, wherein the head-mounted device includes a first camera component for photographing the eyes of the target object and a second camera component for photographing a display screen in front of the target object. Figure 2 As shown, the method includes:
[0051] Step 202: Acquire first video data related to the target object's eyes captured by the first camera assembly.
[0052] Step 204: extract the first image data and the second image data from the first video data, and determine pupil analysis information of the first image data and the second image data based on the first image data and the second image data, wherein the pupil analysis information includes the horizontal and vertical coordinates of the pupil centroid of the target object's eye on the vertical plane, and the maximum diameter direction of the ellipse fitted to the pupil of the target object.
[0053] Step 206: Calculate the horizontal coordinate and vertical coordinate of the eye rotation center based on the pupil analysis information of the first image data and the second image data as first parameter information.
[0054] Step 208 : Determine whether the target object is observing the calibration point displayed on the display screen in front of the target object, and acquire third image data of the first camera component and fourth image data of the second camera component when the target object observes the calibration point.
[0055] Step 210: Analyze the third image data, the fourth image data, and the first parameter information to determine the distance between the rotation center of the target object's eye and the pupil as the second parameter information, so as to determine the target object's gaze point based on the first parameter information and the second parameter information of the two eyes.
[0056] The solution of the present application can be applied to scenarios involving identifying the gaze point of a target subject's eyes. A head-mounted device can be used to calibrate the target subject's gaze point, and then the target subject's gaze point can be determined based on the calibrated parameters and pupil-related data. This solution can be performed using a head-mounted device worn by the target subject, comprising a first camera assembly for capturing the target subject's eyes and a second camera assembly for capturing a display screen located in front of the target subject. This solution can analyze pupil analysis information of the target subject's eyes using multiple images from first video data captured by the first camera assembly, and determine the horizontal and vertical coordinates of the target subject's eye rotation center. The solution then analyzes whether the target subject is gazing at the calibration point. If so, images from the first and second camera assemblies can be extracted and analyzed to determine the distance between the target subject's eye rotation center and the target subject's pupil. The horizontal and vertical coordinates of the eye rotation center are then combined to form the three-dimensional coordinates of the target subject's eye rotation center. Subsequent analysis of the target subject's gaze point can be performed based on the three-dimensional coordinates of the target subject's eye rotation center, the first analyzed image from the first camera assembly, and the second analyzed image from the second camera assembly. This solution can determine the three-dimensional coordinates of the target object's eye rotation center during the calibration phase. In the subsequent gaze point analysis, the three-dimensional coordinates of the eye rotation center, pupil information (determined based on the first analysis image of the first camera component) and external environment information (determined based on the second analysis image of the second camera component) are considered to more accurately determine the gaze point.
[0057] Guidance information for a preset trajectory can be displayed on a display screen to determine whether the target subject's line of sight aligns with the preset trajectory, thereby determining whether the data collected by the head-mounted device can be subsequently analyzed. Specifically, as an optional embodiment, determining whether the target subject is observing a calibration point displayed on a display screen in front of the target subject includes: displaying guidance information for a preset trajectory on the display screen, with the calibration point located within the preset trajectory of the guidance information; extracting the movement trajectory of the pupil centroid from the first video data; and determining whether the target subject is observing the calibration point displayed on the display screen in front of the target subject based on whether the movement trajectory of the pupil centroid aligns with the preset trajectory. Specifically, as an optional embodiment, the preset trajectory is a cross-shaped trajectory, and the calibration points are extreme points of the cross-shaped trajectory, including the endpoints or center points of the cross. Using extreme points for analysis simplifies calculations. By displaying guidance information for a preset trajectory to analyze the movement trajectory of the target subject's pupil centroid, the target subject does not need to perform specific calibration actions according to instructions, making calibration simpler and more convenient. This solution uses a head-mounted device for line of sight calibration and analysis, eliminating the influence of the target subject's head movement and making gaze point analysis simpler and more convenient.
[0058] This solution can combine the predicted coordinates and actual coordinates of the calibration point, as well as known parameters, to analyze the distance between the eye's rotation center and the pupil. Specifically, as an optional embodiment, the analysis based on the third image data, the fourth image data, and the first parameter information to determine the distance between the target object's eye's rotation center and the pupil as the second parameter information includes: determining first analysis data of the target object based on the third image data of the first camera component, the first analysis data including the maximum diameter direction of the ellipse fitted to the target object's pupil and the ratio of the maximum diameter to the shortest diameter; determining predicted coordinate information of the calibration point on the display screen based on the fourth image data of the second camera component, and combining the actual coordinate information of the calibration point to form second analysis data; and determining the distance between the target object's eye's rotation center and the pupil as the second parameter information based on the first analysis data, the second analysis data, and the first parameter information.
[0059] This solution can pre-record first component parameters of the first camera assembly and second component parameters of the second camera assembly to perform data analysis based on the first and second component parameters. Specifically, as an optional embodiment, determining the distance between the rotation center of the target subject's eye and the pupil as the second parameter information based on the first analysis data, the second analysis data, and the first parameter information includes: obtaining first component parameters of the first camera assembly, the first component parameters including first pixel information, a first virtual working distance, and first pose information of the first camera assembly; obtaining second component parameters of the second camera assembly, the second component parameters including second pixel information, a second virtual working distance, and second pose information of the second camera assembly; and determining the distance between the rotation center of the target subject's eye and the pupil as the second parameter information based on the first analysis data, the first component parameters, the second analysis data, the second component parameters, and the first parameter information. This solution does not require prior knowledge: parameters such as eyeball diameter do not need to be measured; rather, these are offset during the calculation process. This avoids errors caused by individual differences.
[0060] Specifically, as an optional embodiment, two first camera assemblies are provided, the first camera assemblies are located in front of and below the eyes and facing directly behind, the distance between the two first camera assemblies is a first distance, the upward rotation angle of the first camera assembly is a first angle, and the first angle is between 30-45°; the second camera assembly is located above the center of the two first camera assemblies, the vertical distance between the second camera assembly and the first camera assembly is a second distance, the horizontal distance is a third distance, and it faces directly in front of the target object. When the camera assemblies are set in the corresponding positions, the number of data conversion mappings can be reduced, thereby simplifying the overall calculation process. It should be noted that the first camera assembly and the second camera assembly in this solution can also be installed in other ways.
[0061] This solution can determine the three-dimensional coordinates of the target object's eye rotation center during the calibration phase, and in the subsequent gaze point analysis, the three-dimensional coordinates of the eye rotation center, pupil information (determined based on the first analysis image of the first camera component), and external environment information (determined based on the second analysis image of the second camera component) are considered to more accurately determine the gaze point. Specifically, as an optional embodiment, the determination of the target object's gaze point based on the first parameter information and the second parameter information of the two eyes includes: obtaining a first analysis image taken by the two first camera components and a second analysis image taken by the second camera component; and determining the target object's gaze point by analyzing the two first analysis images, the second analysis image, the first parameter information, and the second parameter information of the two eyes. This solution can analyze the pupil of the target object based on the first analysis image, and analyze the predicted gaze coordinates of relevant points in the target object's surrounding environment based on the second analysis image, thereby determining the gaze point. Specifically, as an optional embodiment, the analysis based on the first analysis image, the second analysis image, the first parameter information, and the second parameter information to determine the target object's gaze point includes: determining the target object's first gaze data based on the first analysis images of the two first camera components, the first gaze data including the maximum diameter direction of the ellipse fitted to the pupils of the target object's eyes and the ratio of the maximum diameter to the shortest diameter; determining the predicted gaze coordinates of each relevant point based on the second analysis image of the second camera component to form second gaze data; and determining the target object's gaze point based on the first gaze data and the second gaze data. This solution can first simply analyze the gaze direction, thereby reducing the analysis range of the image captured by the second camera component, and perform coordinate analysis on the relevant points within the analysis range, thereby further performing a more accurate analysis to determine the target object's gaze point.
[0062] On the basis of the above embodiments, the embodiment of the present application further provides a device for processing the sight line information of a target object, wherein the device performs processing based on a head-mounted device worn on the target object, wherein the head-mounted device includes a first camera component for photographing the eyes of the target object and a second camera component for photographing a display screen located in front of the target object, such as Figure 3 As shown, the device includes:
[0063] The first data acquisition module 302 is configured to acquire first video data related to the target object's eyes captured by the first camera assembly.
[0064] The pupil information acquisition module 304 is used to extract the first image data and the second image data from the first video data, and determine pupil analysis information of the first image data and the second image data based on the first image data and the second image data, wherein the pupil analysis information includes the horizontal and vertical coordinates of the pupil centroid of the target object's eye on the vertical plane, and the maximum diameter direction of the ellipse fitted to the pupil of the target object.
[0065] The first parameter acquisition module 306 is configured to calculate the horizontal coordinate and the vertical coordinate of the eye rotation center according to the pupil analysis information of the first image data and the second image data as first parameter information.
[0066] The calibration image acquisition module 308 is used to determine whether the target object is observing the calibration point displayed on the display screen in front of the target object, and to acquire third image data of the first camera component and fourth image data of the second camera component when the target object observes the calibration point.
[0067] The second parameter acquisition module 310 is used to analyze the third image data and the fourth image data to determine the distance between the rotation center of the target object's eye and the pupil as the second parameter information, so as to determine the target object's gaze point based on the first parameter information and the second parameter information of the two eyes.
[0068] The implementation of the embodiment of the present application is similar to the implementation of the above-mentioned method embodiment. The specific implementation can refer to the specific implementation of the above-mentioned method embodiment, which will not be repeated here.
[0069] The solution of the present application can be applied to scenarios involving identifying the gaze point of a target subject's eyes. A head-mounted device can be used to calibrate the target subject's gaze point, and then the target subject's gaze point can be determined based on the calibrated parameters and pupil-related data. This solution can be performed using a head-mounted device worn by the target subject, comprising a first camera assembly for capturing the target subject's eyes and a second camera assembly for capturing a display screen located in front of the target subject. This solution can analyze pupil analysis information of the target subject's eyes using multiple images from first video data captured by the first camera assembly, and determine the horizontal and vertical coordinates of the target subject's eye rotation center. The solution then analyzes whether the target subject is gazing at the calibration point. If so, images from the first and second camera assemblies can be extracted and analyzed to determine the distance between the target subject's eye rotation center and the target subject's pupil. The horizontal and vertical coordinates of the eye rotation center are then combined to form the three-dimensional coordinates of the target subject's eye rotation center. Subsequent analysis of the target subject's gaze point can be performed based on the three-dimensional coordinates of the target subject's eye rotation center, the first analyzed image from the first camera assembly, and the second analyzed image from the second camera assembly. This solution can determine the three-dimensional coordinates of the target object's eye rotation center during the calibration phase. In the subsequent gaze point analysis, the three-dimensional coordinates of the eye rotation center, pupil information (determined based on the first analysis image of the first camera component) and external environment information (determined based on the second analysis image of the second camera component) are considered to more accurately determine the gaze point.
[0070] It should be noted that the division of units and / or modules in the embodiments of the present application is schematic and is merely a logical functional division. In actual implementation, there may be other division methods. In addition, the functional units and / or modules in the various embodiments of the present application may be integrated into one processing unit and / or module, or each unit and / or module may exist physically alone, or two or more units and / or modules may be integrated into one unit and / or module. The above-mentioned integrated units and / or modules may be implemented in the form of hardware or in the form of software functional units and / or modules.
[0071] If the integrated units and / or modules are implemented in the form of software functional units and / or modules and sold or used as independent products, they can be stored in a processor-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the relevant technology or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0072] In addition, the data transmission device and data transmission method provided in the above embodiments are based on the same application concept. Since the principles of solving problems by the method and the device are similar, the implementation of the device and the method can refer to each other, and the repeated parts will not be repeated.
[0073] Figure 4 A structural block diagram of a device is shown according to an exemplary embodiment.
[0074] like Figure 4 As shown, the device 1100 includes at least: a processor 1110 , a memory 1120 , and a transceiver 1130 .
[0075] The transceiver 1130 is used to receive and send data under the control of the processor 1110 .
[0076] exist Figure 4In the embodiment of the present invention, the bus architecture may include any number of interconnected buses and bridges, specifically linking together various circuits of one or more processors represented by processor 1110 and memory represented by memory 1120. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are all well known in the art and therefore will not be further described herein. The bus interface provides an interface. The transceiver 1130 may be a plurality of components, i.e., a transmitter and a receiver, providing units and / or modules for communicating with various other devices over a transmission medium, such as a wireless channel, a wired channel, an optical cable, or the like.
[0077] The processor 1110 is responsible for managing the bus architecture and general processing, and the memory 1120 can store data used by the processor 1110 when performing operations.
[0078] Optionally, the processor 1110 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or a complex programmable logic device (CPLD). The processor 1110 may also employ a multi-core architecture. The processor 1110 and the memory 1120 may also be physically separated.
[0079] The processor 1110 calls the computer program stored in the memory 1120 to execute any one of the methods for allocating a cell radio network temporary identifier provided in the above embodiments of the present application according to the obtained executable instructions.
[0080] Figure 5 A structural block diagram of a user equipment is shown according to an exemplary embodiment.
[0081] like Figure 5 As shown, the user equipment 1300 includes at least: a processor 1310 , a memory 1320 and a transceiver 1330 .
[0082] The transceiver 1330 is used to receive and send data under the control of the processor 1310.
[0083] exist Figure 5In the embodiment, the bus architecture may include any number of interconnected buses and bridges, specifically linking together various circuits of one or more processors represented by processor 1310 and memory represented by memory 1320. The bus architecture may also link together various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are all well known in the art and therefore will not be described further herein. The bus interface provides an interface. The transceiver 1330 may be a plurality of components, i.e., a transmitter and a receiver, providing units and / or modules for communicating with various other devices on a transmission medium, such as wireless channels, wired channels, optical cables, and other transmission media. For different user devices, the user interface 1340 may also be an interface capable of connecting external or internal devices as required, and the connected devices include but are not limited to a keypad, a display, a speaker, a microphone, a joystick, and the like.
[0084] The processor 1310 is responsible for managing the bus architecture and general processing, and the memory 1320 can store data used by the processor 1310 when performing operations.
[0085] Optionally, the processor 1310 may be a CPU (central processing unit), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or a CPLD (Complex Programmable Logic Device). The processor 1310 may also employ a multi-core architecture. The processor 1310 and the memory 1320 may also be physically separated.
[0086] The processor 1310 calls the computer program stored in the memory 1320 to execute any one of the methods for allocating a cell radio network temporary identifier provided in the above embodiments of the present application according to the obtained executable instructions.
[0087] It should be noted here that the above-mentioned device provided in the embodiment of the present application can implement all the method steps implemented in the above-mentioned method embodiment and can achieve the same technical effect. The parts and beneficial effects of this embodiment that are the same as those in the method embodiment will not be described in detail here.
[0088] In addition, an embodiment of the present application provides a storage medium having a computer program stored thereon, which, when executed by a processor, implements the data transmission method of each of the above embodiments. The storage medium can be any available medium or data storage device that can be accessed by the processor, including but not limited to magnetic storage (such as a floppy disk, hard disk, magnetic tape, magneto-optical disk (MO), etc.), optical storage (such as CD, DVD, BD, HVD, etc.), and semiconductor storage (such as ROM, EPROM, EEPROM, non-volatile memory (NAND FLASH), solid-state drive (SSD)), etc.
[0089] In an embodiment of the present application, a program product is provided. For example, the program product is an FPGA chip or a DSP chip. The program product includes executable instructions stored in a storage medium. A processor reads the executable instructions from the storage medium, so that when the executable instructions are executed by the processor, the data transmission method described in each of the above embodiments is implemented.
[0090] The solution of the present application can be applied to the scenario of robot calibration. The robot includes the following components: a base, a first robotic arm base, a first robotic arm end, a second robotic arm base, a second robotic arm end and a depth camera arranged on the first robotic arm end. The first robotic arm base and the second robotic arm base are arranged on a movable base. The coordinate systems of each component are different. The calibration of the robot by this solution includes determining the conversion relationship between the coordinate systems corresponding to the components of the robot.
[0091] The solution of the present application can be applied to scenarios involving identifying the gaze point of a target subject's eyes. A head-mounted device can be used to calibrate the target subject's gaze point, and then the target subject's gaze point can be determined based on the calibrated parameters and pupil-related data. This solution can be performed using a head-mounted device worn by the target subject, comprising a first camera assembly for capturing the target subject's eyes and a second camera assembly for capturing a display screen located in front of the target subject. This solution can analyze pupil analysis information of the target subject's eyes using multiple images from first video data captured by the first camera assembly, and determine the horizontal and vertical coordinates of the target subject's eye rotation center. The solution then analyzes whether the target subject is gazing at the calibration point. If so, images from the first and second camera assemblies can be extracted and analyzed to determine the distance between the target subject's eye rotation center and the target subject's pupil. The horizontal and vertical coordinates of the eye rotation center are then combined to form the three-dimensional coordinates of the target subject's eye rotation center. Subsequent analysis of the target subject's gaze point can be performed based on the three-dimensional coordinates of the target subject's eye rotation center, the first analyzed image from the first camera assembly, and the second analyzed image from the second camera assembly. This solution can determine the three-dimensional coordinates of the target object's eye rotation center during the calibration phase. In the subsequent gaze point analysis, the three-dimensional coordinates of the eye rotation center, pupil information (determined based on the first analysis image of the first camera component) and external environment information (determined based on the second analysis image of the second camera component) are considered to more accurately determine the gaze point.
[0092] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage and optical storage, etc.) that contain computer-usable program code.
[0093] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer-executable instructions. These computer-executable instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the steps in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0094] These processor-executable instructions may also be stored in a processor-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the processor-readable memory produce an article of manufacture comprising an instruction device that implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 The function specified in one or more boxes.
[0095] These processor-executable instructions can also be loaded onto a computer or other programmable data processing device so that a series of operational steps are performed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for executing on the computer or other programmable device to implement the process. Figure 1 a process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0096] It should be understood that although the steps in the flowcharts of the accompanying drawings are shown in sequence as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some of the steps in the flowcharts of the accompanying drawings may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily executed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be executed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0097] The above description is only a partial implementation method of the present application. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present application. These improvements and modifications should also be regarded as the scope of protection of the present application.
Claims
1. A method for processing sight line information of a target object, characterized in that: The method is based on a head-mounted device worn on a target object, wherein the head-mounted device includes a first camera component for photographing the target object's eyes and a second camera component for photographing a display screen located in front of the target object. The method includes: Acquiring first video data related to the target object's eyes captured by the first camera assembly; Extracting first image data and second image data from the first video data, and determining pupil analysis information of the first image data and the second image data based on the first image data and the second image data, the pupil analysis information including the abscissa and ordinate of the pupil centroid of the target object's eye on a vertical plane, and the maximum diameter direction of an ellipse fitted to the pupil of the target object; Calculating the horizontal coordinate and the vertical coordinate of the eye rotation center as first parameter information based on pupil analysis information of the first image data and the second image data; determining whether the target object is observing a calibration point displayed on a display screen in front of the target object, and acquiring third image data of the first camera assembly and fourth image data of the second camera assembly when the target object observes the calibration point; Analyzing the third image data, the fourth image data, and the first parameter information to determine a distance between a rotation center of the target object's eye and the pupil as the second parameter information, and determining the target object's gaze point based on the first parameter information and the second parameter information of the two eyes; The horizontal coordinate xe0 and vertical coordinate ye0 of the eye rotation center are calculated according to the following formula: ; Among them, x1 is the horizontal coordinate of the pupil centroid of the target object's eye in the first image data on the vertical plane, y1 is the vertical coordinate of the pupil centroid of the target object's eye in the first image data on the vertical plane, and α1 is the maximum diameter direction of the ellipse fitted to the pupil of the target object in the first image data; x2 is the horizontal coordinate of the pupil centroid of the target object's eye in the second image data on the vertical plane, y2 is the vertical coordinate of the pupil centroid of the target object's eye in the second image data on the vertical plane, and α2 is the maximum diameter direction of the ellipse fitted to the pupil of the target object in the second image data.
2. The method according to claim 1, characterized in that The determining whether the target object is observing a calibration point displayed on a display screen in front of the target object includes: Displaying guidance information of a preset trajectory on a display screen, wherein the calibration point is located within the preset trajectory of the guidance information; Extracting a movement trajectory of the pupil centroid from the first video data; Whether the target object is observing the calibration point displayed on the display screen in front of the target object is determined based on whether the movement trajectory of the pupil mass center conforms to the preset trajectory.
3. The method according to claim 2, characterized in that The preset trajectory is a cross-shaped trajectory, and the calibration point is an extreme point of the cross-shaped trajectory. The extreme point of the cross-shaped trajectory includes the end point or the center point of the cross.
4. The method according to claim 1, wherein The analyzing the third image data, the fourth image data, and the first parameter information to determine the distance between the rotation center of the target object's eye and the pupil as the second parameter information includes: Determining first analysis data of the target object based on the third image data of the first camera assembly, the first analysis data including the maximum diameter direction of the ellipse fitted to the pupil of the target object and the ratio of the maximum diameter to the shortest diameter; Determining predicted coordinate information of a calibration point on the display screen based on the fourth image data of the second camera assembly, and combining the predicted coordinate information of the calibration point with the actual coordinate information of the calibration point to form second analysis data; The distance between the rotation center of the target object's eye and the pupil is determined based on the first analysis data, the second analysis data, and the first parameter information as the second parameter information.
5. The method according to any one of claims 1 to 4, characterized in that The determining, based on the first analysis data, the second analysis data, and the first parameter information, of the distance between the target object's eye rotation center and the pupil as the second parameter information includes: Acquire first component parameters of the first camera component, where the first component parameters include first pixel information, a first virtual working distance, and first pose information of the first camera component; Acquire second component parameters of the second camera assembly, where the second component parameters include second pixel information, a second virtual working distance, and second posture information of the second camera assembly; The distance between the target object's eye rotation center and the pupil is determined as the second parameter information based on the first analysis data, the first component parameter, the second analysis data, the second component parameter and the first parameter information.
6. The method according to claim 5, characterized in that Two first camera assemblies are provided, the first camera assemblies are located in front of and below the eyes and facing directly rearward, the distance between the two first camera assemblies is a first distance, the upward rotation angle of the first camera assembly is a first angle, and the first angle is between 30° and 45°; the second camera assembly is located above the center of the two first camera assemblies, the vertical distance between the second camera assembly and the first camera assembly is a second distance, the horizontal distance is a third distance, and it faces directly in front of the target object.
7. The method according to claim 1, characterized in that The determining the gaze point of the target object based on the first parameter information and the second parameter information of the two eyes includes: Acquire a first analysis image taken by the first camera assembly and a second analysis image taken by the second camera assembly; An analysis is performed based on the two first analysis images, the second analysis image, the first parameter information and the second parameter information of the two eyes to determine the gaze point of the target object.
8. The method according to claim 7, characterized in that Analyzing the first analysis image, the second analysis image, the first parameter information, and the second parameter information to determine the gaze point of the target object includes: Determining first gaze data of the target object based on the first analysis images of the two first camera assemblies, the first gaze data including the maximum diameter direction of the ellipse fitted to the pupils of the target object's eyes and the ratio of the maximum diameter to the shortest diameter; determining predicted gaze coordinates of each relevant point based on the second analysis image of the second camera assembly to form second gaze data; A gaze point of the target object is determined according to the first gaze data and the second gaze data.
9. A device for processing sight line information of a target object, characterized in that: The apparatus completes processing based on a head-mounted device worn on a target object, the head-mounted device including a first camera component for photographing the target object's eyes and a second camera component for photographing a display screen located in front of the target object, the apparatus comprising: A first data acquisition module is used to acquire first video data related to the target object's eyes collected by the first camera assembly; a pupil information acquisition module, configured to extract first image data and second image data from the first video data, and determine pupil analysis information of the first image data and the second image data based on the first image data and the second image data, wherein the pupil analysis information includes the horizontal and vertical coordinates of the pupil centroid of the target object's eye on a vertical plane, and the maximum diameter direction of the ellipse fitted to the pupil of the target object; A first parameter acquisition module is used to calculate the horizontal coordinate and the vertical coordinate of the eye rotation center according to the pupil analysis information of the first image data and the second image data as the first parameter information; a calibration image acquisition module, configured to determine whether the target object is observing a calibration point displayed on a display screen in front of the target object, and to acquire third image data of the first camera assembly and fourth image data of the second camera assembly when the target object observes the calibration point; a second parameter acquisition module, configured to analyze the third image data and the fourth image data to determine a distance between a rotation center of the target object's eye and the pupil as second parameter information, and to determine a gaze point of the target object based on the first parameter information and the second parameter information of the two eyes; The horizontal coordinate xe0 and vertical coordinate ye0 of the eye rotation center are calculated according to the following formula: ; Among them, x1 is the horizontal coordinate of the pupil centroid of the target object's eye in the first image data on the vertical plane, y1 is the vertical coordinate of the pupil centroid of the target object's eye in the first image data on the vertical plane, and α1 is the maximum diameter direction of the ellipse fitted to the pupil of the target object in the first image data; x2 is the horizontal coordinate of the pupil centroid of the target object's eye in the second image data on the vertical plane, y2 is the vertical coordinate of the pupil centroid of the target object's eye in the second image data on the vertical plane, and α2 is the maximum diameter direction of the ellipse fitted to the pupil of the target object in the second image data.
10. A device, characterized in that include: A memory, a transceiver, and a processor; wherein the memory is used to store computer programs; the transceiver is used to send and receive data under the control of the processor; The processor is configured to read the computer program in the memory and execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Systems and methods for performing eye gaze tracking
CN109690553A
Space staring tracking method and device based on human eyeball model
CN116434314A