Line-of-sight direction detection method and apparatus, electronic device, and storage medium

By acquiring binocular images in the terminal device and generating monocular calibration data to configure the monocular prediction model, the problem of inaccurate gaze direction detection caused by monocular occlusion or blinking is solved, achieving higher detection accuracy and interactive label stability.

CN116434315BActive Publication Date: 2026-05-22BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING ZITIAO NETWORK TECH CO LTD
Filing Date
2023-03-30
Publication Date
2026-05-22

Smart Images

  • Figure CN116434315B_ABST
    Figure CN116434315B_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide a line-of-sight direction detection method and device, electronic equipment and storage medium. The method comprises: obtaining a binocular image, processing the binocular image based on a binocular prediction model to obtain a binocular prediction result, the binocular prediction result representing a first line-of-sight direction corresponding to the binocular image; generating monocular calibration data according to the binocular prediction result, and configuring a monocular prediction model using the monocular calibration data; and when a monocular image is detected, processing the monocular image using the monocular prediction model to obtain a target line-of-sight direction corresponding to the monocular image. The monocular calibration data is generated using the binocular image collected in a normal state and the corresponding binocular prediction result, and the monocular prediction model is calibrated, so that the monocular prediction model has similar prediction ability to the binocular prediction model used to generate the binocular prediction result, and the target line-of-sight direction obtained based on the monocular prediction model has higher accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of virtual reality technology, and in particular to a gaze direction detection method, apparatus, electronic device, and storage medium. Background Technology

[0002] In virtual reality technology, interactive icons are displayed on the screen by detecting the user's gaze direction, thereby realizing human-computer interaction based on the user's gaze, which effectively improves the efficiency of human-computer interaction and the user experience.

[0003] In existing technologies, before performing gaze direction detection, terminal devices typically need to first perform parameter calibration and configure a binocular prediction model based on the calibration results, thereby reducing detection errors when using binocular images to predict gaze direction.

[0004] However, in practical applications, when there are situations such as monocular occlusion or blinking, the terminal device can only predict the gaze direction based on monocular images. In this case, it will lead to problems such as inaccurate gaze direction detection results and abrupt changes in interactive icons. Summary of the Invention

[0005] This disclosure provides a gaze direction detection method, apparatus, electronic device, and storage medium to overcome problems such as inaccurate gaze direction detection results and abrupt changes in interactive identifiers.

[0006] In a first aspect, embodiments of this disclosure provide a line-of-sight direction detection method, including:

[0007] A stereo image is acquired and processed based on a stereo prediction model to obtain a stereo prediction result, wherein the stereo prediction result represents the first line of sight corresponding to the stereo image; monocular calibration data is generated based on the stereo prediction result, and a monocular prediction model is configured using the monocular calibration data; when a monocular image is detected, the monocular image is processed using the monocular prediction model to obtain the target line of sight corresponding to the monocular image.

[0008] Secondly, embodiments of this disclosure provide a line-of-sight direction detection device, comprising:

[0009] An acquisition module is used to acquire binocular images and process the binocular images based on a binocular prediction model to obtain binocular prediction results, wherein the binocular prediction results characterize the first line of sight direction corresponding to the binocular images;

[0010] The generation module is used to generate monocular calibration data based on the binocular prediction results, and to configure a monocular prediction model using the monocular calibration data.

[0011] The processing module is used to process the monocular image using the monocular prediction model when a monocular image is detected, so as to obtain the target gaze direction corresponding to the monocular image.

[0012] Thirdly, embodiments of this disclosure provide an electronic device, including:

[0013] A processor, and a memory communicatively connected to the processor;

[0014] The memory stores computer-executed instructions;

[0015] The processor executes computer execution instructions stored in the memory to implement the line-of-sight direction detection method as described in the first aspect and various possible designs of the first aspect.

[0016] Fourthly, embodiments of this disclosure provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the gaze direction detection method described in the first aspect and various possible designs of the first aspect.

[0017] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the gaze direction detection method as described in the first aspect and various possible designs of the first aspect.

[0018] The gaze direction detection method, apparatus, electronic device, and storage medium provided in this embodiment acquire binocular images and process them based on a binocular prediction model to obtain a binocular prediction result, which represents the first gaze direction corresponding to the binocular image. Monocular calibration data is generated based on the binocular prediction result, and a monocular prediction model is configured using the monocular calibration data. When a monocular image is detected, the monocular prediction model is used to process the monocular image to obtain the target gaze direction corresponding to the monocular image. By using the normally acquired binocular image and the corresponding binocular prediction result to generate monocular calibration data and calibrate the monocular prediction model, the monocular prediction model achieves a similar predictive capability to the binocular prediction model that generated the binocular prediction result. This results in higher accuracy of the target gaze direction obtained based on the monocular prediction model. Simultaneously, it reduces the jump in detection results caused by switching between the monocular and binocular prediction models during continuous gaze direction detection, improving the stability of interactive markers. Attached Figure Description

[0019] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0020] Figure 1 This is an application scenario diagram of the gaze direction detection method provided in the embodiments of this disclosure;

[0021] Figure 2 Flowchart of the line-of-sight direction detection method provided in the embodiments of this disclosure Figure 1 ;

[0022] Figure 3 for Figure 2 A flowchart illustrating the specific implementation of step S102 in the illustrated embodiment;

[0023] Figure 4 This is a schematic diagram illustrating a process for detecting the target gaze direction in a monocular image, as provided in an embodiment of the present disclosure.

[0024] Figure 5 Flowchart of the line-of-sight direction detection method provided in the embodiments of this disclosure Figure 2 ;

[0025] Figure 6 A schematic diagram of a target visual region provided in an embodiment of this disclosure;

[0026] Figure 7 A schematic diagram of another target visual region provided in an embodiment of this disclosure;

[0027] Figure 8 for Figure 5 A flowchart illustrating the specific implementation of step S208 in the illustrated embodiment;

[0028] Figure 9 for Figure 5 A flowchart illustrating the specific implementation of step S210 in the illustrated embodiment;

[0029] Figure 10 This is a schematic diagram illustrating the generation of a target line of sight according to an embodiment of the present disclosure;

[0030] Figure 11 for Figure 5 A flowchart illustrating the specific implementation of step S211 in the illustrated embodiment;

[0031] Figure 12 This is a structural block diagram of the line-of-sight direction detection device provided in the embodiments of this disclosure;

[0032] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure;

[0033] Figure 14 This is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of this disclosure. Detailed Implementation

[0034] To make the objectives, technical solutions, and advantages of the embodiments of this disclosure clearer, the technical solutions of the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.

[0035] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.

[0036] The application scenarios of the embodiments of this disclosure are explained below:

[0037] Figure 1 This diagram illustrates an application scenario of the gaze direction detection method provided in this embodiment. The gaze direction detection method can be applied to scenarios such as Virtual Reality (VR) and Mixed Reality (MR). More specifically, it can be applied in virtual reality scenarios during human-computer interaction based on the user's gaze, such as... Figure 1 As shown in the figure, the method provided in this embodiment can be applied to a terminal device, such as a VR headset. The virtual reality headset is equipped with a camera for capturing eye images and a display screen for displaying image content. After the virtual reality headset captures the wearer's eye images through the camera, it analyzes the eye images to obtain the gaze direction. Based on the gaze direction, it generates a corresponding interactive symbol on the display screen, such as a hand icon as shown in the figure. Then, based on information such as the dwell time of the interactive symbol, the human-computer interaction process is completed.

[0038] In existing technologies, during human-computer interaction based on user gaze, the detection of gaze direction is typically achieved by converting the acquired eye images into the corresponding gaze direction using a pre-set image prediction model. However, this detection process is affected by factors such as the wearer's binocular distance and binocular size. To improve the accuracy of gaze direction detection, existing technologies usually calibrate the image prediction model before the start of the human-computer interaction process based on user gaze. This is done by displaying multiple calibration markers on the terminal device screen, guiding the user to observe each marker sequentially and acquire the corresponding eye images. During this process, the user's eyes are required to be unobstructed, allowing the terminal device to obtain images containing the user's two eyes, i.e., binocular images. Subsequently, based on these binocular images and the gaze directions corresponding to the calibration markers, the binocular prediction model is characterized, matching the binocular prediction model with the wearer's facial features to achieve accurate prediction of the gaze direction.

[0039] However, in practical applications, when the human-computer interaction process based on the user's gaze begins, if the eye image captured by the terminal device is a monocular image in an occluded state, the previously calibrated binocular prediction model cannot be used for gaze detection. Instead, a preset monocular prediction model must be used. Although the monocular prediction model can detect gaze direction using monocular images, it is uncalibrated. Therefore, on the one hand, using the monocular prediction model for gaze detection suffers from low accuracy; on the other hand, because the model parameters of the monocular prediction model and the calibrated binocular prediction model are different, the prediction results (gaze direction) output by the two are inconsistent. This leads to abrupt changes in the interaction identifiers generated based on the gaze direction, affecting interaction efficiency and smoothness. This disclosure provides a gaze direction detection method to solve the above problems.

[0040] refer to Figure 2 , Figure 2 Flowchart of the line-of-sight direction detection method provided in the embodiments of this disclosure Figure 1 The method of this embodiment can be applied in a terminal device. This gaze direction detection method includes:

[0041] Step S101: Acquire a stereo image and process the stereo image based on a stereo prediction model to obtain a stereo prediction result, wherein the stereo prediction result represents the first line of sight corresponding to the stereo image.

[0042] For example, refer to Figure 1The illustrated application scenario diagram shows a terminal device, such as a virtual reality headset. The terminal device is equipped with an image acquisition unit, such as a camera, for facing the wearer's eyes. The image acquisition unit acquires and detects images of the wearer's eyes. When the eye image contains the wearer's complete binoculars, i.e., the left and right eyes, the eye image is identified as a binocular image. Conversely, when the eye image only contains the wearer's single eye, i.e., the left or right eye, the eye image is identified as a monocular image.

[0043] Furthermore, the binocular prediction model is a pre-configured model used to predict the gaze direction based on the image features of binocular images. This binocular prediction model is a calibrated model that adapts to the wearer's eye features (such as interocular distance, eye contour size, etc.). Therefore, after the obtained binocular image is input into the binocular prediction model for processing, the first gaze direction with the accuracy corresponding to the binocular image is obtained, i.e., the binocular prediction result. The specific method of the binocular prediction model is existing technology and will not be elaborated here.

[0044] Step S102: Generate monocular calibration data based on the binocular prediction results, and configure the monocular prediction model using the monocular calibration data.

[0045] For example, after obtaining the binocular prediction results, corresponding monocular calibration data can be generated using the binocular prediction results as calibration data. For instance, the binocular prediction results can be cached in a preset storage location or queue, and processed asynchronously by a separate service. Specifically, the monocular calibration data is used to calibrate the monocular prediction model. The monocular calibration data generated from the binocular prediction results contains information characterizing or used to determine the logic by which the binocular prediction model processes images. Therefore, for example, after configuring the monocular prediction model with the monocular calibration data, the monocular prediction model also has the same or similar gaze detection capabilities as the binocular prediction model.

[0046] Specifically, in one possible implementation, the binocular prediction result includes radian information representing the first line of sight direction and corresponding binocular image features, wherein the radian information represents angle information in three-dimensional space. Further, the radian information may include heading angle, pitch angle, etc. The radian information represents the implementation of the binocular prediction model based on the above-mentioned binocular image prediction; while the binocular image features represent the image content features of the binocular image, which can be represented by feature rectangles. These features can be extracted using a pre-set feature extraction model, which can be composed of multiple convolutional layers, nonlinear layers, and downsampling layers. After inputting the binocular image into this pre-trained feature extraction model, the binocular image features can be obtained. The specific implementation steps are detailed below. The specific implementation form of the feature extraction model can be set as needed and will not be elaborated here.

[0047] Furthermore, in one possible implementation, such as Figure 3 As shown, the specific implementation of step S102 includes:

[0048] Step S1021: Obtain the calibration radian based on the radian information.

[0049] Step S1022: Obtain monocular model parameters based on binocular image features. The monocular model parameters are used to configure the monocular prediction model so that the monocular prediction model establishes a mapping relationship between monocular image features and corresponding viewing directions.

[0050] Step S1023: Obtain monocular calibration data based on the monocular model parameters and the calibration radian.

[0051] The radian information and the calibration radian represent the same meaning, namely, the angle in three-dimensional space used to represent the direction of vision. After obtaining the radian information, the calibration radian can be obtained through appropriate data format conversion. The two can differ only in data format or be completely identical, which will not be elaborated further. Binocular image features are the image features corresponding to binocular images. The parameters of the three-dimensional model of the eyeball are solved based on the spot-corneal reflection method, namely the monocular model parameters. For example, the monocular prediction model is a model that predicts the direction of vision based on the spot-corneal reflection method. After the monocular model parameters are configured in the monocular prediction model, the monocular prediction model can be calibrated, enabling the monocular prediction model to establish a mapping relationship between monocular image features and accurate direction of vision. Subsequently, the set ([F,L]) of the monocular model parameters (F) and the calibration radian (L) can be used as monocular calibration data for the subsequent configuration process of the monocular prediction model.

[0052] Furthermore, the binocular image features can include left-eye image features and right-eye image features, each being an independent feature map. Merging these features yields the binocular image features; conversely, splitting the binocular image features yields left-eye and right-eye image features. After processing the left-eye and right-eye image features separately using the corneal reflection method, the first monocular model parameters corresponding to the left-eye image features and the second monocular model parameters corresponding to the right-eye image features can be obtained. Then, based on the second monocular model parameters corresponding to the left and right eyes, monocular calibration data is obtained. Subsequently, when configuring the monocular prediction model using the monocular calibration data, the first monocular prediction model corresponding to the left eye and the second monocular prediction model corresponding to the right eye are configured separately, enabling the configured first and second monocular prediction models to accurately detect the monocular images corresponding to the left and right eyes, respectively.

[0053] Step S103: When a monocular image is detected, the monocular prediction model is used to process the monocular image to obtain the target gaze direction corresponding to the monocular image.

[0054] For example, when a wearer's monocular vision is obstructed due to blinking, foreign object blockage, or other reasons, the image captured by the terminal device will be detected as a monocular image. After detecting the monocular image, a monocular prediction model previously configured based on monocular calibration data is invoked to predict the monocular image, and the output result is the target gaze direction corresponding to the monocular image. The monocular prediction model can be a gaze detection model based on the pupil-corneal reflection method. After configuration, the specific implementation principle of this monocular prediction model is existing technology and will not be elaborated here.

[0055] Figure 4 This is a schematic diagram illustrating a process for detecting the target gaze direction in a monocular image, as provided in an embodiment of this disclosure. Figure 4 As shown, the terminal device continuously acquires eye images. When the acquired eye images are binocular images, they are input into a binocular prediction model for processing to obtain a binocular prediction result. Based on the binocular prediction result, monocular calibration data is generated, and the monocular prediction model is configured (updated) using this monocular calibration data. After the monocular calibration data is configured, the consumption process of the binocular image is completed. Afterward, the terminal device returns to continue acquiring eye images. When the acquired eye images are monocular images, they are input into a monocular prediction model for processing to obtain the target gaze direction. The above process can be executed asynchronously using different linear methods.

[0056] In this embodiment, binocular images are acquired and processed based on a binocular prediction model to obtain a binocular prediction result, which represents the first gaze direction corresponding to the binocular image. Monocular calibration data is generated based on the binocular prediction result, and a monocular prediction model is configured using the monocular calibration data. When a monocular image is detected, the monocular prediction model is used to process the monocular image to obtain the target gaze direction corresponding to the monocular image. By using the normally acquired binocular images and the corresponding binocular prediction results to generate monocular calibration data and calibrate the monocular prediction model, the monocular prediction model achieves a similar prediction capability to the binocular prediction model that generated the binocular prediction result. This results in higher accuracy of the target gaze direction obtained based on the monocular prediction model. Simultaneously, it reduces the jump in detection results caused by switching between the monocular and binocular prediction models during continuous gaze direction detection, improving the stability of the interactive identifier.

[0057] refer to Figure 5 , Figure 5Flowchart of the line-of-sight direction detection method provided in the embodiments of this disclosure Figure 2 This embodiment is in Figure 2 Based on the illustrated embodiment, steps S102 and S103 are further refined, and the line-of-sight direction detection method includes:

[0058] Step S200: Acquire an eye image. If the eye image is a binocular image, proceed to step S201; if the eye image is a monocular image, proceed to step S208.

[0059] Step S201: Process the binocular image based on the binocular prediction model to obtain a binocular prediction result. The binocular prediction result includes at least two initial data. The initial data includes radian coordinates and corresponding binocular image features. The radian coordinates are used to characterize the direction vector of the first line of sight.

[0060] For example, eye images are acquired and monocular / binocular identification is performed. If the eye image is determined to be a binocular image, it is processed based on a calibrated binocular prediction model to predict the gaze direction, i.e., the binocular prediction result. The specific implementation process of the above steps is described in... Figure 2 The embodiments shown have already been described and will not be repeated here. In one possible implementation, steps S200-S201 involve processing multiple frames of stereo images. For example, 10 frames of stereo images are input into the stereo prediction model for processing, resulting in a set of corresponding prediction results, which constitutes the stereo prediction result. Each frame of stereo image corresponds to a prediction result, which is considered initial data. The stereo prediction result can be a set of multiple initial data. Further, the initial data includes radian coordinates and corresponding stereo image features. The radian coordinates represent the direction vector of the first line of sight. For example, the initial data Data_i = [L_1, F_i], where L_1 is the radian coordinate representing the direction vector of the first line of sight, expressed as L_1 = (yaw, pitch). More specifically, yaw represents the heading angle of the first line of sight, and pitch represents the pitch angle of the first line of sight. L_1 = (yaw, pitch) represents the angle of the first line of sight in three-dimensional space. F_i can be the feature matrix representing the features of a stereo image, and will not be elaborated further here.

[0061] Step S202: Determine the gaze state corresponding to the binocular prediction result based on the radian coordinates of the initial data, wherein the gaze state includes a staring state and a non-staring state.

[0062] Step S203: If the binocular prediction result is a staring state, then proceed to step S204; if the binocular prediction result is a non-staring state, then proceed to step S200.

[0063] Furthermore, each initial data point corresponds to the detection result of the gaze direction in a single frame of binocular image. Subsequent calibration of the monocular prediction model is required based on the binocular prediction result composed of multiple initial data points. Therefore, before using the initial data, the validity and stability of each initial data point in the binocular prediction result must be ensured. Specifically, in this embodiment, the gaze state corresponding to the binocular prediction result is determined through the radian coordinates in the initial data, thereby achieving the above objective. For example, the specific implementation includes the following steps:

[0064] Multiple initial data points from the binocular prediction results are processed sequentially according to the acquisition sequence of their corresponding binocular images to determine the single-frame gaze state corresponding to each initial data point. Specifically, this includes: during the sequential processing of each initial data point, determining the change in gaze direction based on the difference between the radian coordinates of the current initial data point and the radian coordinates of the previous initial data point. If the change in gaze direction is greater than a threshold, the single-frame gaze state of the current initial data point is considered a single-frame non-gazing state; if the change in gaze direction is less than the threshold, the single-frame gaze state of the current initial data point is considered a single-frame gaze state. Subsequently, if the number of initial data points in the binocular prediction results that are in a single-frame gaze state is greater than a preset value or a preset proportion, the binocular prediction result is a gaze state; otherwise, the binocular prediction result is a non-gazing state. In this embodiment, the gaze state corresponding to the binocular prediction result is determined by the radian coordinates of each initial data corresponding to the binocular prediction result, thereby realizing data filtering based on gaze state, i.e., data stability filtering. The prediction results (initial data) corresponding to binocular images collected in an unstable state, such as when the wearer of the terminal device is moving their gaze, are filtered out, thereby improving the data quality of the subsequently generated monocular calibration data and improving the accuracy of gaze direction detection of the monocular prediction model.

[0065] Furthermore, the method in this embodiment also includes: after determining that the binocular prediction result is a staring state, the binocular image features of each initial data corresponding to the binocular prediction result are averaged and fused to obtain average image features, so that the binocular image features of each initial data are consistent (all are average image features), thereby achieving the purpose of removing noise and improving the data quality of the subsequently generated monocular calibration data.

[0066] Step S204: Based on the radian coordinates of the initial data, determine at least two first target data from the at least two initial data, wherein the radian coordinates corresponding to each first target data are evenly distributed within the target visual area, and the target visual area corresponds to a preset radian range that can be reached by the line of sight.

[0067] Furthermore, exemplarily, the monocular calibration data generated from the binocular prediction results is actually a mapping set representing the "image feature-viewing direction" mapping relationship. After being configured into the monocular prediction model, it enables the model to predict viewing angles based on this mapping set. Therefore, the monocular calibration data needs to cover as many different viewing angles as possible, allowing the monocular prediction model to predict different viewing angles. Consequently, the binocular prediction results used to generate the monocular prediction model need to have broad and uniform sample coverage. To achieve this objective, in this embodiment, after obtaining multiple initial data corresponding to the binocular prediction results, the initial data is further filtered to ensure uniform coverage of different viewing directions, thereby enabling the monocular calibration data generated based on the initial data to represent the image features corresponding to different viewing directions.

[0068] Specifically, from multiple initial data sets, several reference data sets, i.e., the first target data sets, are selected based on the radian coordinates of the initial data sets and are evenly distributed within the target visual region. The target visual region corresponds to a preset radian range reachable by the line of sight. Specifically, the target visual region can be a specific symmetrical shape region. After dividing the target visual region circumferentially based on a preset angle, multiple sub-regions are obtained, and each sub-region corresponds to one first target data set. Figure 6 A schematic diagram of a target visual region provided in an embodiment of this disclosure, such as... Figure 6 As shown, exemplarily, the target visual region is a circular region. The target visual region is divided into eight sub-regions of equal area (i.e., with the same angular spacing). Then, based on the radian coordinates of each initial data point, eight initial data points are obtained, namely data_1 to data_8. The radian coordinates of these eight initial data points are P1 to P8 (shown as P1, P2, P3, P4, P5, P6, P7, and P8 in the figure). Referring to the figure, P1 to P8 are each located within one sub-region of the target visual region. That is, the radian coordinates corresponding to the first target data are evenly distributed within multiple equally sized sub-regions of the target visual region. In other words, the radian coordinates corresponding to the first target data are evenly distributed within the target visual region. Through this method, the first target data can completely cover the target visual region, avoiding the problem of incomplete visual region coverage caused by randomly selecting initial data as the target data for generating monocular calibration data, thereby improving the data quality of the monocular calibration data.

[0069] Further, for example, the vector length corresponding to the radian coordinates of the first target data is greater than a first threshold and / or less than a second threshold, wherein the first threshold represents the lower limit of the radian in the line of sight and the second threshold represents the upper limit of the radian in the line of sight.

[0070] Figure 7 A schematic diagram of another target visual region provided in an embodiment of this disclosure, such as... Figure 7 As shown, the target visual region is a circular region, where the outer ring of the target visual region corresponds to the second threshold, and the inner ring of the target visual region corresponds to the first threshold. If the radian coordinate falls within the target visual region, then the vector length corresponding to the radian coordinate (i.e., the distance from the radian coordinate to the center of the target visual region) is greater than the first threshold and / or less than the second threshold. In the process of filtering the initial data based on the target visual region, the filtering conditions for the initial data are further refined by using the maximum and minimum values ​​corresponding to the target visual region. This is because, in the scenario of gaze direction detection, the problem of poor prediction accuracy of the monocular prediction model and the abrupt changes in interactive identifiers that this embodiment aims to solve is due to the fact that under a large gaze direction, the binocular prediction model is calibrated and can eliminate errors, while the monocular prediction model is not calibrated and cannot eliminate errors. The above-mentioned use of the initial data generated by the binocular prediction model to generate monocular calibration data is also to solve the above problems. However, in practical applications, when the line of sight is narrow, such as when the wearer of the terminal device is looking directly ahead, the radian coordinates corresponding to their line of sight fall within the center of the target visual region. In this case, the errors generated by the monocular and binocular prediction models are very small and do not affect the line-of-sight prediction results. Furthermore, the scenario of the wearer looking directly ahead is more common, resulting in the collection of a large amount of initial data with radian coordinates near the center of the first time region. Processing this data to generate monocular calibration data fails to calibrate the monocular prediction model and wastes computational resources. Similarly, for excessively wide line-of-sight directions, the prediction results (initial data) of the binocular prediction model may have significant errors, making it unsuitable for monocular calibration. Therefore, in this embodiment, the radian coordinates are restricted by the upper and / or lower limits of the target visual region to further filter the initial data, making the filtered initial data more effective and improving the data quality of the monocular calibration data.

[0071] Step S205: Obtain the remaining data from the at least two initial data sets, excluding the first target data.

[0072] Step S206: Generate second target data based on the evaluation value of the remaining data, wherein the evaluation value characterizes the accuracy of the line-of-sight direction obtained by the monocular prediction model from the remaining data.

[0073] Furthermore, after determining the first target data, the remaining data in the multiple initial data corresponding to the binocular prediction results are evaluated. Specifically, this includes: for each remaining data, evaluation is performed using a monocular calibration model, that is, the binocular image features in the remaining data are divided into left-eye image features and right-eye image features, which are then input into the monocular calibration model to obtain the viewing direction corresponding to the left-eye image features and the viewing direction corresponding to the right-eye image features. Then, based on the radian information corresponding to the left-eye image features in the remaining data, the viewing direction corresponding to the left-eye image features is compared to obtain the left-eye evaluation value; based on the radian information corresponding to the right-eye image features in the remaining data, the viewing direction corresponding to the right-eye image features is compared to obtain the right-eye evaluation value. Based on the left-eye evaluation value and the right-eye evaluation value, the evaluation value of the remaining data is obtained. For example, the higher the evaluation value, the higher the consistency between the label (radian information) of the remaining data and the result (predicted viewing direction) output by the monocular prediction model; conversely, the lower the consistency.

[0074] Furthermore, based on the evaluation value of the remaining data, if the evaluation value is greater than a preset value, the remaining data is identified as the second target data, thus achieving further filtering of the remaining data and improving the data quality of the subsequently generated monocular calibration data.

[0075] Step S207: Generate monocular calibration data based on the second target data and the at least two first target data, and configure a monocular prediction model based on the monocular calibration data.

[0076] For example, the second target data and the at least two first target data obtained in the above steps are the results of filtering the binocular prediction results. Therefore, merging the second target data and the at least two first target data results in the filtered binocular prediction result, which can be combined with... Figure 2 The data implementation method for the binocular prediction results in the illustrated embodiment is the same; therefore, the specific implementation method of step S207 is the same as... Figure 2 The specific implementation of step S102 in the illustrated embodiment is similar and can be referred to. Figure 2 The relevant descriptions in the illustrated embodiments will not be repeated here.

[0077] Step S208: Obtain the monocular image features corresponding to the monocular image.

[0078] For example, steps S201-S207 are the process of configuring the monocular prediction model. During the execution of steps S201-S207, steps S208-S211 are executed asynchronously. That is, after detecting that the eye image is a monocular image, the monocular prediction model is used to process the monocular image. This includes obtaining the monocular image features corresponding to the monocular image, i.e., extracting the features corresponding to the monocular image using a preset feature extraction model, such as... Figure 8 As shown, the specific implementation of step S208 includes:

[0079] Step S2081: The eye region in the monocular image is cropped to obtain the corresponding first eye crop.

[0080] Step S2082: Extract features from the first eye screenshot to obtain monocular image features.

[0081] For example, after obtaining a monocular image, the eye region in the monocular image is identified to obtain the eye region contour. Then, based on the eye region contour, the eye region is cropped to obtain a first eye screenshot. Subsequently, feature extraction is performed on the first eye screenshot to obtain the corresponding monocular image features. In one possible implementation, to address the problem of low eye region recognition accuracy in complex monocular images, this embodiment uses the following method to determine the eye region of the monocular image: First, the reference points of the eye region in the monocular image are identified, such as the corners of the eyes on both sides. Then, based on the coordinates of the corners of the eyes, the coordinates of the eye center point are calculated. Next, based on the coordinates of the eye center point and the size of the detected eye contour, or a preset size, a rectangular region centered on the eye center point is determined as the eye region. Compared to schemes that directly identify the eye region, this embodiment is simpler, more robust, avoids erroneous cropping of the eye region, and improves the information integrity of the first eye screenshot.

[0082] Step S209: Based on the monocular prediction model, obtain the approximate image features corresponding to the monocular image features, and the calibration radians corresponding to the approximate image features.

[0083] Step S210: Obtain the line-of-sight deviation radian based on the monocular image features and the approximate image features.

[0084] For example, the monocular prediction model configured with monocular calibration data contains discrete information representing multiple sets of "image feature-viewing direction" mapping relationships, such as [image feature F1, viewing direction L1], [image feature F2, viewing direction L2], [image feature F3, viewing direction L3], etc. The monocular prediction model maps discrete image features to corresponding viewing directions, thereby achieving viewing direction prediction. Since the image features in the monocular prediction model are discrete values, after receiving the monocular image features, it is necessary to first obtain the image feature in the monocular prediction model that is most similar to the monocular image features, i.e., the approximate image feature. This process can be achieved through image feature comparison, which will not be elaborated further. Then, based on the feature differences between the approximate image feature and the monocular image feature, the deviation radians are obtained.

[0085] For example, such as Figure 9 As shown, the specific implementation of step S210 includes:

[0086] Step S2101: Obtain the residual network model, which is used to output the feature residual between two input quantities.

[0087] Step S2102: Process the monocular image features and the approximate image features based on the residual network model to obtain the line-of-sight deviation radian.

[0088] For example, the residual network model is a residual network model used to calculate feature residuals. It can be implemented based on the model parameters of a monocular prediction model, or it can be implemented after separate pre-training. For example, the residual network model can be trained by using two image features with different labels (the radian of the gaze direction) as samples, so that the residual network model can map the two different image features to their corresponding label difference (the radian of the gaze deviation). The specific training process will not be described in detail.

[0089] Furthermore, by using the residual network model to process the monocular image features and the approximate image features, the deviation of their viewing directions, i.e., the viewing deviation in radians, can be obtained.

[0090] Step S211: Obtain the target line of sight direction based on the calibration radian and the line of sight deviation radian.

[0091] Furthermore, after obtaining the line-of-sight deviation radian, the target line-of-sight direction can be obtained by adding the calibration radian corresponding to the approximate image features obtained based on the monocular prediction model to the line-of-sight deviation radian. Figure 10 This is a schematic diagram illustrating the generation of a target viewing direction, as provided in an embodiment of this disclosure. Figure 10As shown, for example, after detecting a monocular image, corresponding monocular image features are generated. Then, the monocular image features are input into a monocular prediction model to obtain approximate image features and the calibration radians corresponding to the approximate image features. Then, the approximate image features and monocular image features are merged and input into a residual network model to obtain the line-of-sight deviation radians output by the residual network model. Finally, the target line-of-sight direction is obtained by using the sum of the line-of-sight deviation radians and the calibration radians.

[0092] For example, in another possible implementation, the obtained line-of-sight deviation radians can be further weighted, such as... Figure 11 As shown, the specific implementation of step S211 includes:

[0093] Step S2111: Obtain the first predicted direction based on the calibration radian and the line-of-sight deviation radian.

[0094] Step S2112: Obtain the deviation weight of the first predicted direction based on the line-of-sight deviation radian.

[0095] Step S2113: Obtain the target line-of-sight direction based on the product of the first predicted direction and the deviation weight.

[0096] For example, after obtaining the calibration radian and the line-of-sight deviation radian, the two are first added together to obtain a first predicted direction, which corresponds to the target implementation in the above implementation method. Then, a corresponding deviation weight is determined based on the line-of-sight deviation radian, wherein the deviation weight is determined based on the normalized value of the line-of-sight deviation radian. Specifically, for example, in the step of obtaining approximate image features corresponding to the monocular image features through a monocular prediction model, multiple approximate image features similar to the monocular image features are obtained. Then, corresponding line-of-sight deviation radians are generated based on the approximate image features. Next, the ratio of the target line-of-sight deviation radian (e.g., the minimum line-of-sight deviation radian) to the total line-of-sight deviation radians is calculated, and the deviation weight is determined based on this ratio. Then, the first predicted direction is adjusted based on the deviation weight, i.e., the product of the first predicted direction and the deviation weight is calculated to obtain the target line-of-sight direction. In this embodiment, the deviation weight is determined by the ratio of the target approximate image features to the sum of multiple approximate image features, and the first prediction direction is corrected based on the deviation weight to obtain the target line of sight direction. By averaging, the information utilization rate of the monocular prediction model is improved, and the prediction accuracy of the target line of sight direction of the monocular prediction model is further improved.

[0097] Corresponding to the gaze direction detection method in the above embodiment, Figure 12 This is a structural block diagram of a line-of-sight direction detection device provided in an embodiment of this disclosure. For ease of explanation, only the parts relevant to the embodiments of this disclosure are shown. (Refer to...) Figure 12The line-of-sight direction detection device 3 includes:

[0098] The acquisition module 31 is used to acquire a stereo image and process the stereo image based on a stereo prediction model to obtain a stereo prediction result, wherein the stereo prediction result represents the first line of sight direction corresponding to the stereo image.

[0099] The generation module 32 is used to generate monocular calibration data based on the binocular prediction results, and to configure a monocular prediction model using the monocular calibration data.

[0100] The processing module 33 is used to process the monocular image using the monocular prediction model when a monocular image is detected, so as to obtain the target gaze direction corresponding to the monocular image.

[0101] In one embodiment of this disclosure, the binocular prediction result includes radii information characterizing the first line of sight direction, and corresponding binocular image features;

[0102] The generation module 32 is specifically used for: obtaining a calibration radian based on the radian information; obtaining monocular model parameters based on binocular image features, wherein the monocular model parameters are configured in the monocular prediction model so that the monocular prediction model establishes a mapping relationship between monocular image features and the corresponding line of sight; and obtaining monocular calibration data based on the monocular model parameters and the calibration radian.

[0103] In one embodiment of this disclosure, the binocular image features include left-eye image features and right-eye image features. When the generation module 32 obtains monocular model parameters based on the binocular image features, it is specifically used to: process the left-eye image features and the right-eye image features respectively based on the corneal reflection method to obtain the corresponding first monocular model parameters and second monocular model parameters respectively; and obtain the monocular model parameters based on the combination of the first monocular model parameters and the second monocular model parameters.

[0104] In one embodiment of this disclosure, the binocular prediction result includes at least two initial data, the initial data including radian coordinates, the radian coordinates being used to characterize the direction vector of the first line of sight;

[0105] In one embodiment of this disclosure, when the generation module 32 generates monocular calibration data based on the binocular prediction result, it is specifically used to: determine at least two first target data from the at least two initial data based on the radian coordinates of the initial data, wherein the radian coordinates corresponding to each first target data are evenly distributed within the target visual area, and the target visual area corresponds to a preset radian range that can be reached by the line of sight; and generate monocular calibration data based on the at least two first target data.

[0106] In one embodiment of this disclosure, the vector length corresponding to the radian coordinates of the first target data is greater than a first threshold and / or less than a second threshold, wherein the first threshold represents the lower limit of the radian in the line of sight and the second threshold represents the upper limit of the radian in the line of sight.

[0107] In one embodiment of this disclosure, the generation module 32 is further configured to: acquire the remaining data other than the first target data from the at least two initial data; generate second target data based on the evaluation value of the remaining data, wherein the evaluation value characterizes the accuracy of the line-of-sight direction obtained by the monocular prediction model from the remaining data; when generating monocular calibration data based on the at least two first target data, the generation module 32 is specifically configured to: generate monocular calibration data based on the second target data and the at least two first target data.

[0108] In one embodiment of this disclosure, the generation module 32 is further configured to: determine the gaze state corresponding to the binocular prediction result based on the radian information in the binocular prediction result, wherein the gaze state includes a staring state and a non-staring state; when generating monocular calibration data based on the binocular prediction result, the generation module 32 is specifically configured to: if the binocular prediction result is a staring state, then generate monocular calibration data based on the binocular prediction result.

[0109] In one embodiment of this disclosure, when the processing module 33 processes the monocular image using the monocular prediction model to obtain the target gaze direction corresponding to the monocular image, it is specifically used to: acquire the monocular image features corresponding to the monocular image; based on the monocular prediction model, obtain the approximate image features corresponding to the monocular image features, and the calibration radian corresponding to the approximate image features; obtain the gaze deviation radian according to the monocular image features and the approximate image features; and obtain the target gaze direction according to the calibration radian and the gaze deviation radian.

[0110] In one embodiment of this disclosure, when the processing module 33 obtains the monocular image features corresponding to the monocular image, it is specifically used to: crop the eye region in the monocular image to obtain a corresponding first eye crop; and extract features from the first eye crop to obtain monocular image features.

[0111] In one embodiment of this disclosure, when the processing module 33 obtains the line-of-sight deviation radian based on the monocular image features and the approximate image features, it is specifically used to: obtain a residual network model, wherein the residual network model is used to output the feature residual between two input quantities; and process the monocular image features and the approximate image features based on the residual network model to obtain the line-of-sight deviation radian.

[0112] In one embodiment of this disclosure, when the processing module 33 obtains the target line of sight based on the calibration radian and the line of sight deviation radian, it is specifically configured to: obtain a first predicted direction based on the calibration radian and the line of sight deviation radian; obtain a deviation weight of the first predicted direction based on the line of sight deviation radian; and obtain the target line of sight based on the product of the first predicted direction and the deviation weight.

[0113] The acquisition module 31, generation module 32, and processing module 33 are connected in sequence. The gaze direction detection 3 provided in this embodiment can execute the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, so it will not be described again here.

[0114] Figure 13 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present disclosure, such as... Figure 13 As shown, the electronic device 4 includes:

[0115] Processor 41, and memory 42 communicatively connected to processor 41;

[0116] The memory 42 stores computer-executed instructions;

[0117] The processor 41 executes the computer execution instructions stored in the memory 42 to achieve, for example... Figures 2-11 The line-of-sight direction detection method in the illustrated embodiment.

[0118] Optionally, the processor 41 and the memory 42 are connected via a bus 43.

[0119] For relevant instructions, please refer to the corresponding text. Figures 2-11 The relevant descriptions and effects of the steps in the corresponding embodiments are understood, and will not be elaborated on here.

[0120] This disclosure provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement this disclosure. Figures 2-11 The line-of-sight direction detection method provided in any of the corresponding embodiments.

[0121] This disclosure provides a computer program product, including a computer program that, when executed by a processor, implements as follows: Figures 2-11 The gaze direction detection in the illustrated embodiment

[0122] refer to Figure 14The diagram illustrates a structural schematic of an electronic device 900 suitable for implementing embodiments of the present disclosure. The electronic device 900 can be a terminal device or a server. The terminal device can include, but is not limited to, mobile terminals such as mobile phones, laptops, digital radio receivers, personal digital assistants (PDAs), portable Android devices (PADs), portable media players (PMPs), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 14 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0123] like Figure 14 As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0124] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 14 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.

[0125] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.

[0126] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.

[0127] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.

[0128] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to perform the methods shown in the above embodiments.

[0129] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0130] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0131] The units described in the embodiments of this disclosure can be implemented in software or in hardware. The name of a unit does not necessarily limit the unit itself; for example, the first acquisition unit can also be described as "a unit that acquires at least two Internet Protocol addresses".

[0132] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.

[0133] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0134] In a first aspect, according to one or more embodiments of the present disclosure, a gaze direction detection method is provided, comprising:

[0135] A stereo image is acquired and processed based on a stereo prediction model to obtain a stereo prediction result, wherein the stereo prediction result represents the first line of sight corresponding to the stereo image; monocular calibration data is generated based on the stereo prediction result, and a monocular prediction model is configured using the monocular calibration data; when a monocular image is detected, the monocular image is processed using the monocular prediction model to obtain the target line of sight corresponding to the monocular image.

[0136] According to one or more embodiments of this disclosure, the binocular prediction result includes radian information characterizing the first line of sight direction and corresponding binocular image features; generating monocular calibration data based on the binocular prediction result includes: obtaining a calibration radian based on the radian information; obtaining monocular model parameters based on the binocular image features, the monocular model parameters being configured in the monocular prediction model to establish a mapping relationship between monocular image features and the corresponding line of sight direction; and obtaining monocular calibration data based on the monocular model parameters and the calibration radian.

[0137] According to one or more embodiments of this disclosure, the binocular image features include left-eye image features and right-eye image features. The step of obtaining monocular model parameters based on the binocular image features includes: processing the left-eye image features and the right-eye image features respectively based on the corneal reflection method to obtain corresponding first monocular model parameters and second monocular model parameters respectively; and obtaining monocular model parameters based on the combination of the first monocular model parameters and the second monocular model parameters.

[0138] According to one or more embodiments of this disclosure, the binocular prediction result includes at least two initial data, the initial data including radian coordinates, the radian coordinates being used to characterize the direction vector of the first line of sight; the step of generating monocular calibration data based on the binocular prediction result includes: determining at least two first target data from the at least two initial data based on the radian coordinates of the initial data, wherein the radian coordinates corresponding to each first target data are uniformly distributed within the target visual region, the target visual region corresponding to a preset radian range reachable by the line of sight; and generating monocular calibration data based on the at least two first target data.

[0139] According to one or more embodiments of this disclosure, the vector length corresponding to the radian coordinates of the first target data is greater than a first threshold and / or less than a second threshold, wherein the first threshold represents the lower limit of the radian in the line of sight and the second threshold represents the upper limit of the radian in the line of sight.

[0140] According to one or more embodiments of this disclosure, the method further includes: acquiring remaining data other than the first target data from the at least two initial data sets; generating second target data based on the evaluation value of the remaining data, wherein the evaluation value characterizes the accuracy of the line-of-sight direction obtained by the monocular prediction model from the remaining data; and generating monocular calibration data based on the at least two first target data sets, comprising: generating monocular calibration data based on the second target data and the at least two first target data sets.

[0141] According to one or more embodiments of this disclosure, the method further includes: determining the gaze state corresponding to the binocular prediction result based on the radian information in the binocular prediction result, wherein the gaze state includes a staring state and a non-staring state; generating monocular calibration data based on the binocular prediction result includes: if the binocular prediction result is a staring state, then generating monocular calibration data based on the binocular prediction result.

[0142] According to one or more embodiments of this disclosure, processing the monocular image using the monocular prediction model to obtain the target gaze direction corresponding to the monocular image includes: acquiring monocular image features corresponding to the monocular image; obtaining approximate image features corresponding to the monocular image features and calibration radians corresponding to the approximate image features based on the monocular prediction model; obtaining gaze deviation radians based on the monocular image features and the approximate image features; and obtaining the target gaze direction based on the calibration radians and the gaze deviation radians.

[0143] According to one or more embodiments of this disclosure, obtaining the monocular image features corresponding to the monocular image includes: cropping the eye region in the monocular image to obtain a corresponding first eye crop; and extracting features from the first eye crop to obtain monocular image features.

[0144] According to one or more embodiments of this disclosure, obtaining the line-of-sight deviation radian based on the monocular image features and the approximate image features includes: acquiring a residual network model, the residual network model being used to output the feature residual between two input quantities; and processing the monocular image features and the approximate image features based on the residual network model to obtain the line-of-sight deviation radian.

[0145] According to one or more embodiments of this disclosure, obtaining the target line of sight direction based on the calibration radian and the line of sight deviation radian includes: obtaining a first predicted direction based on the calibration radian and the line of sight deviation radian; obtaining a deviation weight of the first predicted direction based on the line of sight deviation radian; and obtaining the target line of sight direction based on the product of the first predicted direction and the deviation weight.

[0146] Secondly, according to one or more embodiments of this disclosure, a line-of-sight direction detection device is provided, comprising:

[0147] An acquisition module is used to acquire binocular images and process the binocular images based on a binocular prediction model to obtain binocular prediction results, wherein the binocular prediction results characterize the first line of sight direction corresponding to the binocular images;

[0148] The generation module is used to generate monocular calibration data based on the binocular prediction results, and to configure a monocular prediction model using the monocular calibration data.

[0149] The processing module is used to process the monocular image using the monocular prediction model when a monocular image is detected, so as to obtain the target gaze direction corresponding to the monocular image.

[0150] According to one or more embodiments of this disclosure, the binocular prediction result includes radian information characterizing the first line of sight direction and corresponding binocular image features; the generation module is specifically used for: obtaining a calibration radian based on the radian information; obtaining monocular model parameters based on the binocular image features, the monocular model parameters being configured in the monocular prediction model to enable the monocular prediction model to establish a mapping relationship between monocular image features and the corresponding line of sight direction; and obtaining monocular calibration data based on the monocular model parameters and the calibration radian.

[0151] According to one or more embodiments of this disclosure, the binocular image features include left-eye image features and right-eye image features. When the generation module obtains monocular model parameters based on the binocular image features, it is specifically used to: process the left-eye image features and the right-eye image features respectively based on the corneal reflection method to obtain corresponding first monocular model parameters and second monocular model parameters respectively; and obtain monocular model parameters based on the combination of the first monocular model parameters and the second monocular model parameters.

[0152] According to one or more embodiments of this disclosure, the binocular prediction result includes at least two initial data, the initial data including radian coordinates, the radian coordinates being used to characterize the direction vector of the first line of sight;

[0153] According to one or more embodiments of this disclosure, when the generation module generates monocular calibration data based on the binocular prediction result, it is specifically used to: determine at least two first target data from the at least two initial data based on the radian coordinates of the initial data, wherein the radian coordinates corresponding to each first target data are uniformly distributed within the target visual region, and the target visual region corresponds to a preset radian range that can be reached by the line of sight; and generate monocular calibration data based on the at least two first target data.

[0154] According to one or more embodiments of this disclosure, the vector length corresponding to the radian coordinates of the first target data is greater than a first threshold and / or less than a second threshold, wherein the first threshold represents the lower limit of the radian in the line of sight and the second threshold represents the upper limit of the radian in the line of sight.

[0155] According to one or more embodiments of this disclosure, the generation module is further configured to: obtain the remaining data other than the first target data from the at least two initial data; generate second target data based on the evaluation value of the remaining data, wherein the evaluation value characterizes the accuracy of the line-of-sight direction obtained by the monocular prediction model from the remaining data; and when the generation module generates monocular calibration data based on the at least two first target data, it is specifically configured to: generate monocular calibration data based on the second target data and the at least two first target data.

[0156] According to one or more embodiments of this disclosure, the generation module is further configured to: determine the gaze state corresponding to the binocular prediction result based on the radian information in the binocular prediction result, wherein the gaze state includes a staring state and a non-staring state; when generating monocular calibration data based on the binocular prediction result, the generation module is specifically configured to: if the binocular prediction result is a staring state, then generate monocular calibration data based on the binocular prediction result.

[0157] According to one or more embodiments of this disclosure, when the processing module processes the monocular image using the monocular prediction model to obtain the target gaze direction corresponding to the monocular image, it is specifically configured to: acquire monocular image features corresponding to the monocular image; obtain approximate image features corresponding to the monocular image features and calibration radians corresponding to the approximate image features based on the monocular prediction model; obtain gaze deviation radians based on the monocular image features and the approximate image features; and obtain the target gaze direction based on the calibration radians and the gaze deviation radians.

[0158] According to one or more embodiments of this disclosure, when the processing module obtains the monocular image features corresponding to the monocular image, it is specifically used to: crop the eye region in the monocular image to obtain a corresponding first eye crop; and extract features from the first eye crop to obtain monocular image features.

[0159] According to one or more embodiments of this disclosure, when the processing module obtains the line-of-sight deviation radian based on the monocular image features and the approximate image features, it is specifically used to: obtain a residual network model, wherein the residual network model is used to output the feature residual between two input quantities; and process the monocular image features and the approximate image features based on the residual network model to obtain the line-of-sight deviation radian.

[0160] According to one or more embodiments of this disclosure, when the processing module obtains the target line-of-sight direction based on the calibration radian and the line-of-sight deviation radian, it is specifically configured to: obtain a first predicted direction based on the calibration radian and the line-of-sight deviation radian; obtain a deviation weight of the first predicted direction based on the line-of-sight deviation radian; and obtain the target line-of-sight direction based on the product of the first predicted direction and the deviation weight.

[0161] Thirdly, according to one or more embodiments of the present disclosure, an electronic device is provided, including: a processor, and a memory communicatively connected to the processor;

[0162] The memory stores computer-executed instructions;

[0163] The processor executes computer execution instructions stored in the memory to implement the line-of-sight direction detection method as described in the first aspect and various possible designs of the first aspect.

[0164] Fourthly, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein computer-executable instructions are stored therein, and when a processor executes the computer-executable instructions, the gaze direction detection method described in the first aspect and various possible designs of the first aspect is implemented.

[0165] Fifthly, embodiments of this disclosure provide a computer program product, including a computer program that, when executed by a processor, implements the gaze direction detection method as described in the first aspect and various possible designs of the first aspect.

[0166] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.

[0167] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.

[0168] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative examples of implementing the claims.

Claims

1. A method for detecting gaze direction, characterized in that, include: A stereo image is acquired and processed based on a stereo prediction model to obtain a stereo prediction result, wherein the stereo prediction result represents the first line of sight corresponding to the stereo image; Based on the binocular prediction results, monocular calibration data is generated, and the monocular prediction model is configured using the monocular calibration data. When a monocular image is detected, the monocular prediction model is used to process the monocular image to obtain the target line of sight corresponding to the monocular image. The binocular prediction result includes radii information representing the first line of sight direction, and corresponding binocular image features, which include left eye image features and right eye image features. Based on the binocular prediction results, monocular calibration data is generated, including: Based on the radian information, the calibration radian is obtained; Based on the corneal reflection method, the features of the left eye image and the features of the right eye image are processed respectively to obtain the corresponding first monocular model parameters and second monocular model parameters; Based on the combination of the first monocular model parameters and the second monocular model parameters, monocular model parameters are obtained. These monocular model parameters are configured in the monocular prediction model so that the monocular prediction model establishes a mapping relationship between monocular image features and corresponding viewing directions. Monocular calibration data is obtained based on the monocular model parameters and the calibration radians.

2. The method according to claim 1, characterized in that, The binocular prediction result includes at least two initial data points, the initial data points including radian coordinates, the radian coordinates being used to characterize the direction vector of the first line of sight; The step of generating monocular calibration data based on the binocular prediction results includes: Based on the radian coordinates of the initial data, at least two first target data are determined from the at least two initial data, wherein the radian coordinates corresponding to each first target data are evenly distributed within the target visual area, and the target visual area corresponds to a preset radian range that can be reached by the line of sight. Monocular calibration data is generated based on the at least two first target data.

3. The method according to claim 2, characterized in that, The vector length corresponding to the radian coordinates of the first target data is greater than a first threshold and / or less than a second threshold, wherein the first threshold represents the lower limit of the radian in the line of sight and the second threshold represents the upper limit of the radian in the line of sight.

4. The method according to claim 2, characterized in that, The method further includes: Obtain the remaining data from the at least two initial data sets, excluding the first target data; Based on the evaluation value of the remaining data, second target data is generated, wherein the evaluation value characterizes the accuracy of the line-of-sight direction obtained by the monocular prediction model from the remaining data; The step of generating monocular calibration data based on the at least two first target data includes: Monocular calibration data is generated based on the second target data and the at least two first target data.

5. The method according to claim 1, characterized in that, The method further includes: Based on the radian information in the binocular prediction result, the gaze state corresponding to the binocular prediction result is determined, wherein the gaze state includes a staring state and a non-staring state. The step of generating monocular calibration data based on the binocular prediction results includes: If the binocular prediction result indicates a staring state, then monocular calibration data is generated based on the binocular prediction result.

6. The method according to claim 1, characterized in that, The step of processing the monocular image using the monocular prediction model to obtain the target gaze direction corresponding to the monocular image includes: Obtain the monocular image features corresponding to the monocular image; Based on the monocular prediction model, approximate image features corresponding to the monocular image features and calibration radians corresponding to the approximate image features are obtained. The line-of-sight deviation radians are obtained based on the monocular image features and the approximate image features. The target line of sight is obtained based on the calibration radian and the line of sight deviation radian.

7. The method according to claim 6, characterized in that, The step of obtaining the monocular image features corresponding to the monocular image includes: The eye region in the monocular image is cropped to obtain the corresponding first eye crop; Feature extraction is performed on the first eye screenshot to obtain monocular image features.

8. The method according to claim 6, characterized in that, The step of obtaining the line-of-sight deviation radian based on the monocular image features and the approximate image features includes: Obtain a residual network model, which is used to output the feature residual between two input quantities; The monocular image features and the approximate image features are processed based on the residual network model to obtain the line-of-sight deviation radian.

9. The method according to claim 6, characterized in that, The step of obtaining the target line of sight direction based on the calibration radian and the line of sight deviation radian includes: The first predicted direction is obtained based on the calibration radian and the line-of-sight deviation radian; The deviation weight of the first prediction direction is obtained based on the line-of-sight deviation radian. The target line-of-sight direction is obtained by multiplying the first predicted direction and the deviation weight.

10. A gaze direction detection device, characterized in that, include: An acquisition module is used to acquire binocular images and process the binocular images based on a binocular prediction model to obtain binocular prediction results, wherein the binocular prediction results characterize the first line of sight direction corresponding to the binocular images; The generation module is used to generate monocular calibration data based on the binocular prediction results, and to configure a monocular prediction model using the monocular calibration data. The processing module is used to process the monocular image using the monocular prediction model when a monocular image is detected, so as to obtain the target gaze direction corresponding to the monocular image. The binocular prediction result includes radian information representing the first line of sight direction, and corresponding binocular image features, including left-eye image features and right-eye image features. The generation module is specifically used to obtain a calibration radian based on the radian information; process the left-eye image features and the right-eye image features respectively using the corneal reflection method to obtain corresponding first monocular model parameters and second monocular model parameters; obtain monocular model parameters based on the combination of the first and second monocular model parameters, which are used to configure the monocular prediction model to establish a mapping relationship between monocular image features and the corresponding line of sight direction; and obtain monocular calibration data based on the monocular model parameters and the calibration radian.

11. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the line-of-sight direction detection method as described in any one of claims 1 to 9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by the processor, implement the line-of-sight direction detection method as described in any one of claims 1 to 9.

13. A computer program product, characterized in that, It includes a computer program that, when executed by a processor, implements the line-of-sight direction detection method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Trinocular rearview mirror and trinocular vision safe driving method and system

    CN110321877A

  • Binocular vision naked eye 3D image generation method

    CN111047709A