Interactive image labeling system and method based on facial posture and eye movement tracking and positioning

Through an interactive image annotation system combining facial posture detection and eye tracking technology, efficient and accurate image data annotation is achieved, solving the problems of low labeling accuracy and efficiency in the existing technology, and is suitable for the construction of image data sets in the field of computer vision.

CN120340030AActive Publication Date: 2025-07-18XIAN INST OF OPTICS & PRECISION MECHANICS CHINESE ACAD OF SCI

Patent Information

Application Number
CN202510796558.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-16
Publication Date
2025-07-18
Estimated Expiration
2045-06-16

AI Technical Summary

Technical Problem

The existing field of image data annotation has failed to effectively combine facial posture detection and eye tracking technologies, resulting in poor labeling accuracy and low efficiency.

Method used

An interactive image labeling system based on facial posture and eye movement tracking positioning is designed, including image data acquisition, preprocessing, facial posture detection, eye movement tracking, screen coordinate mapping and interactive control modules. Through secondary sliding mean filtering, combined with facial posture detection and eye movement tracking technology, rapid positioning and automatic labeling are achieved.

Benefits of technology

It significantly improves the efficiency and accuracy of image data labeling, reduces operator fatigue, simplifies the manual labeling process, and is suitable for the construction of large-scale image data sets and computer vision applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120340030A_ABST
    Figure CN120340030A_ABST
Patent Text Reader

Abstract

The invention relates to an interactive image annotation system and method based on facial posture and eye movement tracking and positioning, and solves the problems of poor annotation precision and low efficiency caused by the fact that two interactive technologies of facial posture detection and eye movement tracking are not organically combined with an automatic image segmentation algorithm in the existing image data annotation field. According to the invention, independent operation of two interaction modes of face posture detection and eye movement tracking is realized through modular design, a proper interaction mode is selected by judging the complexity of a scene, the subsequent processing time can be saved, and through secondary sliding mean filtering processing, the detection accuracy is improved. Errors caused by data noise of a depth camera or near-infrared eye movement tracking equipment based on iris reflection and micro movement of an operator are avoided, and the filtered stable coordinates are matched with a segmentation model of the interaction control module to carry out fine extraction on a target area. A traditional segmentation model and operator interaction information are organically combined in the interaction control module, and the manual labeling process is greatly simplified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an image annotation system and method, and particularly to an interactive image annotation system and method based on facial pose and eye movement tracking and positioning. Background Art

[0002] In recent years, with the rapid development of deep learning and computer vision technologies, high-quality and large-scale image annotation data has become the core element driving the continuous breakthrough of algorithm performance. Especially in tasks such as object detection, instance segmentation, and scene understanding, high-precision annotation data not only directly determines the generalization performance of the algorithm but also plays a key supporting role in the actual deployment effect. Therefore, how to efficiently complete the accurate annotation task of a large number of image targets has become an important technical problem that urgently needs to be solved in the current field of computer vision.

[0003] Currently, the widely adopted data annotation method still mainly relies on manual annotation. In a typical annotation process, annotators are required to use a mouse or a stylus to click on the edge of the target point by point to mark the accurate target contour. Although this traditional annotation method has certain advantages in terms of accuracy, the actual operation process is complex and inefficient. Especially in the case of a large number of targets, complex shapes, or irregular boundaries, manual annotation requires repeated and careful positioning and confirmation, consuming a large amount of time, and it is difficult to effectively improve the annotation efficiency. At the same time, due to the long-term and repetitive operation of the annotation task, annotators are prone to problems such as visual fatigue and decreased attention, which further affect the stability and consistency of the annotation data and reduce the overall annotation quality.

[0004] On the other hand, with the continuous progress of human-computer interaction technologies, facial pose detection and eye movement tracking technologies have shown excellent application prospects in fields such as virtual reality (VR), augmented reality (AR), and human-machine collaborative operation. Facial pose detection technology can intuitively and real-time reflect the user's head orientation and can efficiently capture the operator's focus of attention. Eye movement tracking technology can directly capture the user's eye movement trajectory and accurately reflect the user's current attention area and intention. These two technologies have been successfully applied in many fields such as attention analysis and interface interaction control.

[0005] However, at present, the applications of these two interactive technologies, namely facial pose detection and eye movement tracking, in the field of image data annotation are still in the blank stage. Most of the existing image annotation schemes have not effectively utilized the advantages of these two interactive technologies, and there is currently no complete technical system that organically combines them with an automatic image segmentation algorithm to achieve accurate, efficient, and real-time annotation. Summary of the Invention

[0006] The object of the present invention is to solve the technical problems in the existing image data annotation field that there is no organic combination of two interaction technologies, namely facial pose detection and eye movement tracking, with an automatic image segmentation algorithm, resulting in poor annotation accuracy and low efficiency, and to provide an interactive image annotation system and method based on facial pose and eye movement tracking positioning.

[0007] In order to achieve the above object of the invention, the present invention provides the following technical solutions: An interactive image annotation system based on facial pose and eye movement tracking positioning, which is characterized in that: It includes an image data acquisition module, an image preprocessing module, a facial pose detection module, an eye movement tracking module, a screen coordinate mapping module, a screen coordinate filtering module and an interactive control module; The image data acquisition module includes a depth camera and a near-infrared eye movement tracking device based on iris reflection. The depth camera is used to acquire continuous multi-frame color facial images and corresponding depth maps of the operator, and the near-infrared eye movement tracking device based on iris reflection is used to acquire continuous multi-frame near-infrared eye images of the operator. The output ends of both are electrically connected to the input end of the image preprocessing module, and the input end of the image preprocessing module is also used to input the image to be annotated; the image preprocessing module is used to preprocess the continuous multi-frame color facial images and near-infrared eye images to obtain continuous multi-frame preprocessed facial images and preprocessed eye images respectively; the output end of the image preprocessing module is electrically connected to the input ends of the facial pose detection module and the eye movement tracking module respectively; the facial pose detection module is used to obtain the two-dimensional coordinates of facial key points with the depth camera plane as the reference and the camera position of the depth camera as the origin according to the continuous multi-frame preprocessed facial images and the corresponding depth maps, and the eye movement tracking module is used to obtain the pupil center coordinates according to the continuous multi-frame preprocessed facial images; the output ends of the facial pose detection module and the eye movement tracking module are electrically connected to the input end of the screen coordinate mapping module respectively, and the screen coordinate mapping module is used to convert the two-dimensional coordinates of facial key points or pupil center coordinates with the depth camera plane as the reference and the camera position of the depth camera as the origin into two-dimensional pointing coordinates on the screen; the screen coordinate mapping module, the screen coordinate filtering module and the interactive control module are electrically connected in sequence, and the input end of the interactive control module is also used to input the image to be annotated; The screen coordinate filtering module is used to perform secondary moving average filtering on the two-dimensional pointing coordinates on the screen to obtain filtered stable coordinates, and the interactive control module is used to obtain an automatically annotated image according to the filtered stable coordinates.

[0008] Further, the facial posture detection module includes a face detection and key point extraction module, a key point mean filtering module and a head posture estimation module which are electrically connected in sequence, the input end of the face detection and key point extraction module is electrically connected to the output end of the image preprocessing module, and the output end of the head posture estimation module is electrically connected to the input end of the screen coordinate mapping module; The eye tracking module includes an iris detection module, a pupil detection module and a pupil center positioning module which are electrically connected in sequence, the input end of the iris detection module is electrically connected to the output end of the image preprocessing module, and the pupil center positioning module is electrically connected to the input end of the screen coordinate mapping module.

[0009] At the same time, the present invention also provides an interactive image annotation method based on facial gesture and eye tracking positioning, which adopts the above-mentioned interactive image annotation system based on facial gesture and eye tracking positioning, and its special feature is that it includes the following steps: S1. The depth camera of the image data acquisition module acquires multiple frames of color facial images and corresponding depth maps of the operator, and the near-infrared eye tracking device based on iris reflection acquires multiple frames of near-infrared eye images of the operator; S2, an image preprocessing module preprocesses a plurality of consecutive frames of color facial images and near-infrared eye images to obtain a plurality of consecutive frames of preprocessed facial images and preprocessed eye images, respectively; S3, input the image to be annotated into the image preprocessing module, and the image preprocessing module determines whether the scene of the image to be annotated is complex. If it is not complex, step S4 is executed; if it is complex, step S5 is executed; S4, the facial posture detection module processes the pre-processed facial images of the continuous multiple frames and the corresponding depth maps to obtain the two-dimensional coordinates of the facial key points with the depth camera plane as the reference and the camera position of the depth camera as the origin, and then executes step S6; S5, the eye tracking module processes the pre-processed eye images of the consecutive frames to obtain the pupil center coordinates; S6, passing the two-dimensional coordinates based on the depth camera plane as the reference and the camera position of the depth camera as the origin into the screen coordinate mapping module, and converting them into two-dimensional pointing coordinates on the screen through normalized mapping; Alternatively, the pupil center coordinates are passed into the screen coordinate mapping module and converted into two-dimensional pointing coordinates on the screen through normalized mapping; S7, passing the two-dimensional pointing coordinates into the screen coordinate filtering module for secondary sliding mean filtering to obtain filtered stable coordinates; S8. Transfer the filtered stable coordinates into the interactive control module. After receiving the filtered stable coordinates, input the image to be annotated. When the input of the image to be annotated is completed, activate the segmentation model built in the interactive control module to extract multi-scale features from the image to be annotated, and construct deep semantic features and detailed information. Subsequently, the interactive control module inputs the filtered stable coordinates into the feature map of the segmentation model to guide the segmentation model to focus on the target area of interest to the operator. Then, through the processing of the encoder-decoder structure of the segmentation model, a preliminary segmentation result of the target area is generated. Further, the preliminary segmentation result of the target area is refined and optimized through post-processing techniques, and the confirmation point corresponding to one of the filtered stable coordinates in the image to be annotated is marked to obtain an automatically annotated image. S9. Determine whether multi-point annotation is required. If so, return to step S8 to annotate another confirmation point until all confirmation points are annotated and then execute step S10; if not, directly execute step S10. S10. Output the automatically annotated image to complete the interactive image annotation based on facial pose and eye movement tracking and positioning.

[0010] Furthermore, step S4 is specifically as follows: S4.1. The face detection and key point extraction module extracts the two-dimensional image coordinates of the facial key points in consecutive frames of preprocessed facial images: ; Among them, represents the two-dimensional image coordinates of the th facial key point, and is the number of facial key points; S4.2. In the key point mean filtering module, combine the depth value of the corresponding pixel in the depth map to convert the two-dimensional coordinates of the facial key points into the three-dimensional coordinates of the facial key points: ; Among them: is the three-dimensional coordinate of the facial key point, is the internal parameter matrix of the depth camera, and is the transpose of the matrix; S4.3. Combine the obtained three-dimensional coordinates of the facial key points to form a local head point cloud : ; Subsequently, calculate the mean value of the local head point cloud to determine the initial reference position of the head pose: ; Wherein: is the initial reference position of the head pose; S4.4. Centralize the three-dimensional coordinates of the facial key points to obtain the three-dimensional coordinates of the facial key points after centralization processing , which constitute the centralized head local point cloud ; ; S4.5. Apply a moving average filter to the three-dimensional coordinates of the facial key points after centralization processing in the centralized head local point cloud to obtain the three-dimensional coordinates of the facial key points after filtering , which constitute the stable centralized head local point cloud : ; ; Wherein: represents the three-dimensional coordinates of the th facial key point at the th frame, is the window length of the moving average, is the frame number at the current moment; S4.6. Perform singular value decomposition on the stable centralized head local point cloud in the head pose estimation module: ; Wherein: is the left singular vector matrix, is the singular value diagonal matrix, is the right singular vector matrix, is the transpose of the matrix; S4.7. Take the third column of the right singular vector matrix as the principal normal vector, representing the orientation of the head pose: ; Wherein: is the principal normal vector; S4.8. Solve the vertical pitch angle and horizontal yaw angle of the head pose based on the principal normal vector : ; ; Wherein: is the vertical pitch angle of the head pose, is the horizontal yaw angle of the head pose, is the component of the principal normal vector in the X-axis direction, is the component of the principal normal vector in the Y-axis direction, is the component of the principal normal vector in the Z-axis direction; S4.9. Convert the vertical pitch angle and horizontal yaw angle of the head pose into two-dimensional coordinates with the plane of the depth camera as the reference and the position of the camera of the depth camera as the origin , and then execute step S6: ; ; where: is the th vertical distance between the facial key point and the plane of the depth camera (11).

[0011] Further, step S5 is specifically as follows: S5.1. The iris detection module performs eye key point positioning on consecutive preprocessed eye images, locates the area enclosed by the operator's eye key points, and obtains the gray mean value and standard deviation of the area enclosed by the eye key points. At the same time, the iris detection module collects the iris image of the operator; S5.2. Use the binarization method for consecutive preprocessed eye images to generate a binary image according to the adaptive threshold : ; ; where: is the abscissa of the binary image, is the ordinate of the binary image, is the iris image of the operator collected by the iris detection module, is the empirical coefficient, and ; S5.3. Detect the best center and radius in the binary image through the Hough circle transform, is the X-axis coordinate of the best center, is the Y-axis coordinate of the best center; S5.4. In the pupil detection module, define a local area with the center of the operator's iris image as the center and the radius of as the region of interest for pupil detection: ; where: is the region of interest for pupil detection, is the scaling coefficient, and ; S5.5. Construct a binary pupil candidate map based on the region of interest detected by pupil: ; Wherein: is the binary pupil candidate map; S5.6. Perform morphological opening operation on the binary pupil candidate map to obtain an optimized binary pupil map: ; Wherein: is the optimized binary pupil map, is a disk-shaped template with a radius of pixels in the morphological operation; S5.7. In the pupil center localization module, perform centroid localization on the optimized binary pupil map and finally output the pupil center coordinates : ; .

[0012] Furthermore, step S6 is specifically as follows: Input the two-dimensional coordinates with the depth camera plane as the reference and the camera position of the depth camera as the origin into the screen coordinate mapping module, and convert it into the two-dimensional pointing coordinates on the screen through normalized mapping : ; ; Wherein: and are respectively the abscissa and ordinate of the screen center, is the scale factor of the X axis, is the scale factor of the Y axis; Alternatively, input the pupil center coordinates into the screen coordinate mapping module and convert it into the two-dimensional pointing coordinates on the screen through normalized mapping : ; Wherein: is a 2×2 calibration matrix, is a translation vector.

[0013] Furthermore, step S7 is specifically as follows: Input the two-dimensional pointing coordinates into the screen coordinate filtering module for secondary sliding mean filtering to obtain the filtered stable coordinates : ; ; Wherein: is the window length of the moving average, is the frame number at the current moment, is at the original screen X-axis coordinate value obtained in the frame, is at the original screen Y-axis coordinate value obtained in the frame.

[0014] Further, in step S8, the segmentation model is static_edgeflow_cocolvis.

[0015] Compared with the prior art, the beneficial effects of the present invention are: (1) An interactive image annotation system based on facial pose and eye movement tracking positioning provided by the present invention realizes rapid positioning by combining facial pose detection and eye movement tracking technologies. This system does not rely on the traditional mouse point-by-point clicking annotation method, significantly shortens the data annotation cycle, and remarkably improves the processing efficiency of large-scale image data. At the same time, the operator does not need to distract to operate the mouse, and can keep both hands on the keyboard all the time, thus realizing efficient collaborative operation, reducing the fatigue caused by long-term repetitive labor, and further improving the sustainability and stability of the annotation work.

[0016] (2) An interactive image annotation method based on facial pose and eye movement tracking positioning provided by the present invention constructs a complete set of interactive and automatic annotation closed-loop methods. Through modular design, it realizes the independent operation of two interactive modes of facial pose detection and eye movement tracking. By judging the complexity of the scene, it selects a suitable interactive method, which helps to save subsequent processing time. And through secondary sliding mean filtering processing, it avoids the errors caused by data noise of depth cameras or near-infrared eye movement tracking devices based on iris reflection and the micro-movement of the operator. The stable coordinates after filtering cooperate with the segmentation model of the interactive control module itself to finely extract the target area. In the interactive control module, the traditional segmentation model and the operator's interactive information are organically combined, greatly simplifying the manual annotation process. This method has flexible scalability and compatibility, and is suitable for the construction of large-scale image data sets and various computer vision applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 is the architecture diagram of an embodiment of an interactive image annotation system based on facial pose and eye movement tracking positioning of the present invention; Figure 2 is the schematic diagram of an embodiment of an interactive image annotation system based on facial pose and eye movement tracking positioning of the present invention in a use state; Figure 3In an embodiment of an interactive image annotation method based on facial pose and eye movement tracking and positioning according to the present invention, it is a comparison graph of coordinates after being processed by step S7 and unprocessed coordinates, where a is the unprocessed coordinate and b is the coordinate processed by step S7; Figure 4 In an embodiment of an interactive image annotation method based on facial pose and eye movement tracking and positioning according to the present invention, when the scene is simple, it is a comparison graph of an automatically annotated image based on facial pose positioning and an image to be annotated, where a is the image to be annotated and b is the automatically annotated image; Figure 5 In an embodiment of an interactive image annotation method based on facial pose and eye movement tracking and positioning according to the present invention, when the scene is complex, it is a comparison graph of an automatically annotated image based on eye movement tracking and positioning and an image to be annotated, where a is the image to be annotated and b is the automatically annotated image.

[0018] The description of the reference numerals is as follows: 1 - Image data acquisition module, 11 - Depth camera, 12 - Near-infrared eye movement tracking device based on iris reflection; 2 - Image preprocessing module; 3 - Facial pose detection module, 31 - Face detection and key point extraction module, 32 - Key point mean filtering module, 33 - Head pose estimation module; 4 - Eye movement tracking module, 41 - Iris detection module, 42 - Pupil detection module, 43 - Pupil center positioning module; 5 - Screen coordinate mapping module, 6 - Screen coordinate filtering module, 7 - Interaction control module. Detailed implementation manners

[0019] The present invention will be further described below with reference to the accompanying drawings and exemplary embodiments.

[0020] Refer to Figure 1 、 Figure 2 An interactive image annotation system based on facial pose and eye movement tracking and positioning according to the present invention includes an image data acquisition module 1, an image preprocessing module 2, a facial pose detection module 3, an eye movement tracking module 4, a screen coordinate mapping module 5, a screen coordinate filtering module 6, and an interaction control module 7.

[0021] Among them, the image data acquisition module 1 includes a depth camera 11 and a near-infrared eye movement tracking device 12 based on iris reflection. As Figure 2 shown, when in use, the camera of the depth camera 11 is aimed at the face of the operator, and is responsible for collecting continuous multi-frame color facial images and corresponding depth maps of the operator. The depth map provides distance information for each pixel of the color facial image, enabling two-dimensional image features to be mapped into three-dimensional space and constructing a preliminary geometric structure of the scene.

[0022] The near-infrared eye movement tracking device 12 based on iris reflection continuously captures the eye region of the operator and collects continuous multi-frame near-infrared eye images of the operator.

[0023] The output ends of the depth camera 11 and the near-infrared eye movement tracking device 12 based on iris reflection are electrically connected to the input end of the image preprocessing module 2 respectively. In the image preprocessing module 2, preprocessing will be performed on the color facial image and the near-infrared eye image to obtain a continuous multi-frame preprocessed facial image and a preprocessed eye image. Common preprocessing operations for the color facial image include, for example, mirror flipping, grayscale conversion, noise suppression, etc., which can ensure the accuracy of subsequent key point extraction. Since the near-infrared eye image is itself a grayscale image, grayscale conversion is not required in its preprocessing, and the remaining preprocessing operations are the same as those of the color facial image.

[0024] The input end of the image preprocessing module 2 is also used to input the image to be annotated. In the image preprocessing module 2, the scene complexity of the image to be annotated is also judged to facilitate the selection of subsequent interactive annotation methods. The output end of the image preprocessing module 2 is electrically connected to the input ends of the facial pose detection module 3 and the eye movement tracking module 4 respectively. The facial pose detection module 3 includes a face detection and key point extraction module 31, a key point mean filtering module 32, and a head pose estimation module 33 that are electrically connected in sequence. The input end of the face detection and key point extraction module 31 is electrically connected to the output end of the image preprocessing module 2, and the output end of the head pose estimation module 33 is electrically connected to the input end of the screen coordinate mapping module 5. The facial pose detection module 3 is responsible for processing the continuous multi-frame preprocessed facial images obtained by the image preprocessing module 2 and the corresponding depth maps to obtain two-dimensional coordinates with the plane of the depth camera 11 as the reference and the camera position of the depth camera 11 as the origin.

[0025] The eye movement tracking module 4 includes an iris detection module 41, a pupil detection module 42, and a pupil center positioning module 43 that are electrically connected in sequence. The input end of the iris detection module 41 is electrically connected to the output end of the image preprocessing module 2, and the pupil center positioning module 43 is electrically connected to the input end of the screen coordinate mapping module 5. The entire eye movement tracking module 4 is responsible for processing the continuous multi-frame preprocessed eye images obtained by the image preprocessing module 2 to obtain the pupil center coordinates.

[0026] The screen coordinate mapping module 5, the screen coordinate filtering module 6, and the interaction control module 7 are electrically connected in sequence, and the input end of the interaction control module 7 is also used to input the image to be marked. Among them, the screen coordinate mapping module 5 is responsible for converting the two-dimensional coordinates or pupil center coordinates with the plane of the depth camera 11 as the reference and the camera position of the depth camera 11 as the origin into two-dimensional pointing coordinates on the screen through normalized mapping. The screen coordinate filtering module 6 is responsible for performing quadratic sliding mean filtering on the two-dimensional pointing coordinates on the screen to obtain the filtered stable coordinates. This processing avoids errors caused by data noise of the depth camera 11 or the near-infrared eye movement tracking device 12 based on iris reflection and the slight movement of the operator, making the marking result more accurate. The interaction control module 7 then obtains the automatically marked image according to the filtered stable coordinates.

[0027] Meanwhile, the present invention also provides an interactive image marking method based on facial pose and eye movement tracking positioning. Using the above-mentioned interactive image marking system based on facial pose and eye movement tracking positioning, it includes the following steps: S1. The depth camera 11 of the image data acquisition module 1 acquires multiple consecutive frames of color facial images and corresponding depth maps of the operator, and the near-infrared eye movement tracking device 12 based on iris reflection acquires multiple consecutive frames of near-infrared eye images of the operator; S2. The image preprocessing module 2 preprocesses the multiple consecutive frames of color facial images and near-infrared eye images to obtain multiple consecutive frames of preprocessed facial images and preprocessed eye images respectively; S3. Input the image to be marked into the image preprocessing module 2. The image preprocessing module 2 determines whether the scene of the image to be marked is complex. If it is not complex, then execute step S4. If it is complex, then execute step S5; The judgment of whether the scene is complex is comprehensively judged based on the following objective indicators: When the number of targets to be marked in the image to be marked is less than 3 targets, and the maximum distance between the target distributions does not exceed 30% of the image width, the boundary contour is clear, the boundary blur degree is less than 20% (evaluated based on pixel-level boundary clarity), the contrast between the background and the target to be marked is high, and the average gray difference between the target to be marked and the background is greater than 50%, the scene complexity is low, such as Figure 4 a; otherwise the scene complexity is high, such as Figure 5 a. Through scene judgment, it is possible to clearly select which interaction method to use for marking in advance, saving subsequent processing time to obtain a more efficient and accurate marking effect.

[0028] S4. The facial pose detection module 3 processes the multiple consecutive frames of preprocessed facial images and the corresponding depth maps to obtain two-dimensional coordinates with the plane of the depth camera 11 as the reference and the camera position of the depth camera 11 as the origin, and then execute step S6; And step S4 is specifically: S4.1. The face detection and key point extraction module 31 extracts the two-dimensional image coordinates of the facial key points in the preprocessed facial images of multiple consecutive frames: ; Among them, represents the two-dimensional image coordinates of the th facial key point, is the number of facial key points; S4.2. In the key point mean filtering module 32, in combination with the depth value of the corresponding pixel in the depth map, the two-dimensional coordinates of the facial key points are converted into the three-dimensional coordinates of the facial key points: ; Among them: is the three-dimensional coordinate of the facial key point, is the internal parameter matrix of the depth camera 11, is the transpose of the matrix; S4.3. The obtained three-dimensional coordinates of the facial key points are composed into a local head point cloud : ; Subsequently, calculate the mean value of the local head point cloud to determine the initial reference position of the head pose: ; Among them: is the initial reference position of the head pose; S4.4. Perform centering processing on the three-dimensional coordinates of the facial key points to obtain the three-dimensional coordinates of the facial key points after centering processing, which constitute the centered local head point cloud ; ; S4.5. Perform sliding mean filtering on the three-dimensional coordinates of the facial key points after centering processing in the centered local head point cloud to obtain the three-dimensional coordinates of the facial key points after filtering, which constitute the stable centered local head point cloud : ; Among them: represents the three-dimensional coordinates of the th facial key point at the th frame, is the window length of the sliding average, is the frame number at the current moment; S4.6. Perform singular value decomposition on the stable centered head local point cloud in the head pose estimation module 33 : ; where: is the left singular vector matrix, is the singular value diagonal matrix, is the right singular vector matrix, is the transpose of the matrix; S4.7. Take the third column of the right singular vector matrix as the principal normal vector, representing the orientation of the head pose: ; where: is the principal normal vector; S4.8. Solve the vertical pitch angle and horizontal yaw angle of the head pose based on the principal normal vector : ; ; where: is the vertical pitch angle of the head pose, is the horizontal yaw angle of the head pose, is the component of the principal normal vector in the X-axis direction, is the component of the principal normal vector in the Y-axis direction, is the component of the principal normal vector in the Z-axis direction; S4.9. Convert the vertical pitch angle and horizontal yaw angle of the head pose into two-dimensional coordinates with the plane of the depth camera 11 as the reference and the camera position of the depth camera 11 as the origin , and then execute step S6: ; ; where: is the th vertical distance between the facial key point and the plane of the depth camera 11.

[0029] S5. The eye movement tracking module 4 processes the preprocessed eye images of consecutive multiple frames to obtain the pupil center coordinates; Step S5 is specifically as follows: S5.1. The iris detection module 41 uses eye key point localization on the preprocessed eye images of consecutive multiple frames to locate the area enclosed by the operator's eye key points, and obtains the gray mean value and the standard deviation , meanwhile, the iris detection module 41 collects the iris image of the operator; S5.2. Apply the binarization method to the preprocessed eye images of multiple consecutive frames, and generate a binary image according to the adaptive threshold : ; ; Where: is the abscissa of the binary image, is the ordinate of the binary image, is the iris image of the operator collected by the iris detection module 41, is the empirical coefficient, and ; S5.3. Detect the optimal center and radius in the binary image through the Hough circle transform, is the X-axis coordinate of the optimal center, is the Y-axis coordinate of the optimal center; S5.4. In the pupil detection module 42, define a local area with the center of the operator's iris image as the center and the radius of as the region of interest for pupil detection: ; Where: is the region of interest for pupil detection, is the scaling coefficient, and ; S5.5. Construct a pupil candidate binary image according to the region of interest for pupil detection: ; Where: is the pupil candidate binary image; S5.6. Perform morphological opening operation on the pupil candidate binary image to obtain an optimized pupil binary image: ; Where: is the optimized pupil binary image, is a disk-shaped template with a radius of pixels in the morphological operation; S5.7. In the pupil center localization module 43, perform centroid localization on the optimized pupil binary image, and finally output the pupil center coordinates : ; . ​

[0030] S6. Input the two-dimensional coordinates with the plane of the depth camera 11 as the reference and the position of the camera of the depth camera 11 as the origin into the screen coordinate mapping module, and convert it into the two-dimensional pointing coordinates on the screen through normalized mapping : ; ; Wherein: and are respectively the abscissa and ordinate of the center of the screen, is the scale factor of the X-axis, is the scale factor of the Y-axis; Alternatively, input the pupil center coordinates into the screen coordinate mapping module, and convert it into the two-dimensional pointing coordinates on the screen through normalized mapping : ; Wherein: is a 2×2 calibration matrix, is a translation vector.

[0031] S7. Input the two-dimensional pointing coordinates into the screen coordinate filtering module 6 for quadratic moving average filtering to obtain the filtered stable coordinates : ; ; Wherein: is the window length of the moving average, is the frame number at the current moment, is the original screen X-axis coordinate value obtained in the th frame, is the original screen Y-axis coordinate value obtained in the th frame.

[0032] As Figure 3 a and Figure 3 b show, it can be seen that before the quadratic moving average filtering, the coordinate distribution is relatively scattered, while after the processing, the coordinate distribution is more concentrated and the fluctuation is smaller, significantly improving the pointing stability. This step effectively weakens the influence caused by the data noise of the depth camera 11, the data noise of the near-infrared eye movement tracking device 12 based on iris reflection, the slight head tremor or gaze drift, etc.

[0033] S8. Transmit the filtered stable coordinates into the interactive control module 7. After it receives the filtered stable coordinates, input the image to be annotated. When the input of the image to be annotated is completed, activate the segmentation model built in the interactive control module 7 to extract multi-scale features from the image to be annotated and construct deep semantic features and detailed information; Subsequently, the interactive control module 7 inputs the filtered stable coordinates into the feature map built in the segmentation model to guide the segmentation model to focus on the target area of interest to the operator; Then, through the processing of the encoder-decoder structure of the segmentation model, a preliminary segmentation result of the target area is generated. Further, the preliminary segmentation result of the target area is refined and optimized through post-processing techniques, and the confirmation points corresponding to one of the filtered stable coordinates in the image to be annotated are marked to obtain the automatically annotated image; The segmentation model adopts static_edgeflow_cocolvis, which organically combines the traditional segmentation model with the operator's interaction information in the interactive control module 7, greatly simplifying the manual annotation process. Among them, the post-processing technique is to optimize the pixel labels using conditional random field (CRF) to make the edges consistent with the color and gradient of the original image, that is, to refine and optimize the preliminary segmentation result of the target area.

[0034] S9. Determine whether multi-point annotation is required. If so, return to step S8 to annotate another confirmation point until all confirmation points are annotated and then execute step S10; if not, directly execute step S10; S10. Output the automatically annotated image to complete the interactive image annotation based on facial pose and eye movement tracking positioning.

[0035] As Figure 4 shown, this image to be annotated needs to annotate a fire hydrant. The number of objects to be annotated is small, and the boundary contour of the target is clear, with a high contrast with the background. Therefore, it can be defined as a simple scene, and it is suitable for the image annotation interaction method based on facial pose positioning, where Figure 4 a is the image to be annotated, Figure 4 b is the automatically annotated image, and it can be seen that the fire hydrant is clearly annotated.

[0036] As Figure 5 shown, this image to be annotated needs to annotate a herd of cows. The number of objects to be annotated in the whole image is large, and the poses of the cows are different and the boundary contours are not clear. Therefore, it is defined as a complex scene, and it is suitable for the image annotation interaction method based on eye movement tracking positioning, where Figure 5 a is the image to be annotated, Figure 5 b is the automatically annotated image, and it can be seen that all the cows in the herd are clearly annotated after image annotation by this method.

[0037] The embodiments described above are only descriptions of the specific implementation manners of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solutions of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An interactive image annotation system based on facial pose and eye movement tracking and positioning, characterized in that: It includes an image data acquisition module (1), an image preprocessing module (2), a facial pose detection module (3), an eye movement tracking module (4), a screen coordinate mapping module (5), a screen coordinate filtering module (6), and an interactive control module (7); The image data acquisition module (1) includes a depth camera (11) and a near-infrared eye movement tracking device (12) based on iris reflection. The depth camera (11) is used to collect continuous multi-frame color facial images and corresponding depth maps of the operator, and the near-infrared eye movement tracking device (12) based on iris reflection is used to collect continuous multi-frame near-infrared eye images of the operator. The output ends of both are electrically connected to the input end of the image preprocessing module (2), and the input end of the image preprocessing module (2) is also used to input the image to be annotated; the image preprocessing module (2) is used to preprocess continuous multi-frame color facial images and near-infrared eye images to obtain continuous multi-frame preprocessed facial images and preprocessed eye images respectively; the output end of the image preprocessing module (2) is electrically connected to the input ends of the facial pose detection module (3) and the eye movement tracking module (4) respectively; the facial pose detection module (3) is used to obtain the two-dimensional coordinates of facial key points with the plane of the depth camera (11) as the reference and the camera position of the depth camera (11) as the origin according to continuous multi-frame preprocessed facial images and corresponding depth maps, and the eye movement tracking module (4) is used to obtain the pupil center coordinates according to continuous multi-frame preprocessed facial images; the output ends of the facial pose detection module (3) and the eye movement tracking module (4) are electrically connected to the input end of the screen coordinate mapping module (5) respectively, and the screen coordinate mapping module (5) is used to convert the two-dimensional coordinates of facial key points or pupil center coordinates with the plane of the depth camera (11) as the reference and the camera position of the depth camera (11) as the origin into two-dimensional pointing coordinates on the screen; the screen coordinate mapping module (5), the screen coordinate filtering module (6), and the interactive control module (7) are electrically connected in sequence, and the input end of the interactive control module (7) is also used to input the image to be annotated; The screen coordinate filtering module (6) is used to perform secondary sliding mean filtering on the two-dimensional pointing coordinates on the screen to obtain filtered stable coordinates, and the interactive control module (7) is used to obtain an automatically annotated image according to the filtered stable coordinates.

2. The interactive image annotation system based on facial pose and eye movement tracking and positioning according to claim 1, characterized in that: The facial pose detection module (3) includes a face detection and key point extraction module (31), a key point mean filtering module (32), and a head pose estimation module (33) that are electrically connected in sequence. The input end of the face detection and key point extraction module (31) is electrically connected to the output end of the image preprocessing module (2), and the output end of the head pose estimation module (33) is electrically connected to the input end of the screen coordinate mapping module (5); The eye tracking module (4) comprises an iris detection module (41), a pupil detection module (42), and a pupil center positioning module (43) which are electrically connected in sequence, wherein an input end of the iris detection module (41) is electrically connected to an output end of the image preprocessing module (2), and the pupil center positioning module (43) is electrically connected to an input end of a screen coordinate mapping module (5).

3. An interactive image annotation method based on facial pose and eye movement tracking positioning, which uses an interactive image annotation system based on facial pose and eye movement tracking positioning as described in claim 1 or 2, characterized in that, The following steps are involved: S1, the depth camera (11) of the image data acquisition module (1) acquires a plurality of consecutive frames of color facial images and corresponding depth maps of the operator, and the near-infrared eye tracking device (12) based on iris reflection acquires a plurality of consecutive frames of near-infrared eye images of the operator; S2, an image preprocessing module (2) preprocesses the continuous multiple frames of color facial images and near-infrared eye images to obtain continuous multiple frames of preprocessed facial images and preprocessed eye images respectively; S3, inputting the image to be annotated into the image preprocessing module (2), the image preprocessing module (2) determines whether the scene of the image to be annotated is complex, if not complex, executing step S4, if complex, executing step S5; S4, the facial posture detection module (3) processes the pre-processed facial images of the continuous multiple frames and the corresponding depth maps to obtain the two-dimensional coordinates of the facial key points with the plane of the depth camera (11) as the reference and the camera position of the depth camera (11) as the origin, and then executes step S6; S5, the eye tracking module (4) processes the pre-processed eye images of the consecutive frames to obtain the pupil center coordinates; S6, passing the two-dimensional coordinates with the plane of the depth camera (11) as a reference and the camera position of the depth camera (11) as an origin into the screen coordinate mapping module (5), and converting them into two-dimensional pointing coordinates on the screen through normalized mapping; Alternatively, the pupil center coordinates are passed into the screen coordinate mapping module (5) and converted into two-dimensional pointing coordinates on the screen through normalized mapping; S7, passing the two-dimensional pointing coordinates into the screen coordinate filtering module (6) for secondary sliding mean filtering to obtain filtered stable coordinates; S8, passing the filtered stable coordinates to the interactive control module (7), and after receiving the filtered stable coordinates, inputting the image to be annotated. After the input of the image to be annotated is completed, activating the segmentation model of the interactive control module (7), extracting multi-scale features of the image to be annotated, and constructing deep semantic features and detail information; Subsequently, the interactive control module (7) inputs the filtered stable coordinates into the feature map of the segmentation model to guide the segmentation model to focus on the target area of interest to the operator; Then, after being processed by the encoder-decoder structure of the segmentation model, a preliminary segmentation result of the target area is generated, and the preliminary segmentation result of the target area is further refined and optimized through post-processing technology, and the confirmation point corresponding to one of the filtered stable coordinates in the image to be annotated is annotated to obtain an automatically annotated image; S9. Determine whether multi-point annotation is required. If so, return to step S8 to annotate another confirmation point until all confirmation points are annotated, and then execute step S10; if not, directly execute step S10; S10. Output the automatically annotated image to complete the interactive image annotation based on facial pose and eye movement tracking and positioning.

4. The interactive image annotation method based on facial gesture and eye movement tracking positioning according to claim 3, wherein Step S4 is specifically as follows: S4.

1. The face detection and key point extraction module (31) extracts the two-dimensional image coordinates of the facial key points in the preprocessed facial images of multiple consecutive frames: ; Among them, represents the two-dimensional image coordinates of the th facial key point, and is the number of facial key points; S4.

2. In the key point mean filtering module (32), combine the depth values of the corresponding pixels in the depth map , and convert the two-dimensional coordinates of the facial key points into three-dimensional coordinates of the facial key points : ; Wherein: is the three-dimensional coordinates of the facial key points, is the internal parameter matrix of the depth camera (11), is the transpose of the matrix; S4.

3. Compose the three-dimensional coordinates of the obtained facial key points into a local head point cloud : ; Subsequently, the local head point cloud is calculated and its mean value is used to determine the initial reference position of the head pose: ; Wherein: is the initial reference position of the head posture; S4.4 Three-dimensional coordinates of facial key points Perform centering processing to obtain the three-dimensional coordinates of the facial key points after centering processing , which constitute the centered head local point cloud ; ; S4.

5. Perform centering on the local point cloud of the centered head The three-dimensional coordinates of the facial key points that have been centered in the Perform sliding mean filtering on them to obtain the three-dimensional coordinates of the filtered facial key points , which form a stable local point cloud of the centered head : ; Wherein: represents the th facial key point's three-dimensional coordinates at the th frame, is the window length of the moving average, is the frame number at the current moment; S4.

6. Perform singular value decomposition on the stable centered head local point cloud in the head pose estimation module (33). to perform singular value decomposition: ; Wherein: is the left singular vector matrix, is the singular value diagonal matrix, is the right singular vector matrix, is the transpose of the matrix; S4.

7. Take the third column of the right singular vector matrix as the principal normal vector, representing the orientation of the head pose: ; Wherein: is the main normal vector; S4.

8. Solve the vertical pitch angle and horizontal yaw angle of the head pose based on the principal normal vector Solve the vertical pitch angle and horizontal yaw angle of the head pose: ; ; Wherein: is the vertical pitch angle of the head posture, is the horizontal yaw angle of the head posture, is the component of the principal normal vector in the X-axis direction, is the component of the principal normal vector in the Y-axis direction, is the component of the principal normal vector in the Z-axis direction; S4.

9. Convert the vertical pitch angle and horizontal yaw angle of the head pose into two-dimensional coordinates with the plane of the depth camera (11) as the reference and the position of the camera of the depth camera (11) as the origin. , and then perform step S6: ; ; Wherein: is the vertical distance between the -th facial key point and the plane of the depth camera (11).

5. The interactive image annotation method based on facial gesture and eye movement tracking positioning according to claim 4, wherein Step S5 is specifically as follows: S5.

1. The iris detection module (41) performs eye key-point localization on consecutive preprocessed eye images, locates the region enclosed by the operator's eye key points, and obtains the gray-scale mean of the region enclosed by the eye key points and standard deviation . At the same time, the iris detection module (41) acquires the iris image of the operator; S5.

2. Apply the binarization method to the preprocessed eye images of consecutive multiple frames, and generate a binary image according to the adaptive threshold Generate a binary image : ; ; Wherein: is the abscissa of the binary image, is the ordinate of the binary image, is the iris image of the operator collected by the iris detection module (41), is the empirical coefficient, and ; S5.

3. Detect the optimal center in the binary image through Hough circle transform and radius , is the X-axis coordinate of the optimal center, is the Y-axis coordinate of the optimal center; S5.

4. In the pupil detection module (42), define a local area with the center of the iris image of the operator as the center of the circle and the radius as as the region of interest for pupil detection: ​ ; Wherein: is the region of interest for pupil detection, is the scaling factor, and ; S5.

5. Construct a pupil candidate binary map according to the region of interest detected by the pupil: ; Wherein: is the pupil candidate binary image; S5.

6. Perform morphological opening operation on the pupil candidate binary map to obtain an optimized pupil binary map: ; Wherein: is the optimized binary pupil image, is a disk-shaped template with a radius of pixels in the morphological operation; S5.

7. In the pupil center positioning module (43), perform centroid positioning on the optimized binary pupil image, and finally output the pupil center coordinates : ; 。 6. The interactive image annotation method based on facial pose and eye movement tracking positioning according to claim 5, wherein Step S6 is specifically as follows: The two-dimensional coordinates with the plane of the depth camera (11) as the reference and the camera position of the depth camera (11) as the origin are input into the screen coordinate mapping module (5), and are converted into two-dimensional pointing coordinates on the screen through normalized mapping : ; ; Wherein: and are the abscissa and ordinate of the screen center respectively, is the scale factor of the X-axis, is the scale factor of the Y-axis; Alternatively, the pupil center coordinates are input into the screen coordinate mapping module (5), and are converted into two-dimensional pointing coordinates on the screen through normalized mapping. : ; Wherein: is a 2×2 calibration matrix, is a translation vector.

7. The interactive image annotation method based on facial gesture and eye movement tracking positioning according to claim 6, characterized in that, Step S7 is specifically as follows: Input the two-dimensional pointing coordinates into the screen coordinate filtering module (6) for secondary moving average filtering to obtain the stable coordinates after filtering : ; ; Wherein: is the window length of the moving average, is the frame number at the current moment, is at the frame, the original screen X-axis coordinate value obtained, is at the frame, the original screen Y-axis coordinate value obtained.

8. The interactive image annotation method based on facial pose and eye movement tracking and positioning according to claim 3, characterized in that In step S8, the segmentation model is static_edgeflow_cocolvis.

Citation Information

Patent Citations

  • Image target segmentation system combining eye-movement tracking

    CN106681484A

  • Sight line calibration, motion tracking and precision test method based on adaptive time sequence analysis and prediction

    CN116382473A

  • Focus recognition system based on eye movement information

    CN118172578A

  • Iris and pupil-based gaze estimation method for head-mounted device

    US20190121427A1

Cited By

  • Multi-modal data labeling and large model thinking chain training method based on visual interaction

    CN120562596A