An interactive image labeling system and method based on facial pose and eye movement tracking positioning
An interactive image annotation system that combines facial pose detection and eye tracking technologies solves the problems of low annotation accuracy and efficiency in existing technologies, achieving efficient and accurate image data annotation, and is suitable for image dataset construction in the field of computer vision.
Patent Information
- Application Number
- CN202510796558.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-16
- Publication Date
- 2025-10-24
- Estimated Expiration
- 2045-06-16
AI Technical Summary
The existing field of image data annotation has failed to effectively utilize facial posture detection and eye tracking technologies, resulting in poor annotation accuracy and low efficiency.
An interactive image annotation system based on facial pose and eye tracking is adopted, including image data acquisition, preprocessing, facial pose detection, eye tracking, screen coordinate mapping and interactive control modules. Through filtering of facial key points and pupil center coordinates, combined with an automatic segmentation model, accurate and efficient annotation is achieved.
It enables rapid target region localization, significantly improves image data processing efficiency, reduces operator fatigue, and enhances the continuity and stability of annotation work. It is suitable for the construction of large-scale image datasets and computer vision applications.
Smart Images

Figure CN120340030B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to an image labeling system and method, in particular to an interactive image labeling system and method based on facial pose and eye movement tracking positioning. BACKGROUND
[0002] In recent years, with the rapid development of deep learning and computer vision technology, high-quality and large-scale image labeling data has become the core element driving the continuous breakthrough of algorithm performance. Especially in tasks such as target detection, instance segmentation, and scene understanding, high-precision labeling data not only directly determines the generalization performance of the algorithm, but also plays a key supporting role in the actual deployment effect. Therefore, how to efficiently complete the precise labeling task of a large number of image targets has become an important technical problem that needs to be solved in the current computer vision field.
[0003] At present, the widely used data labeling method is still mainly manual labeling, and the typical labeling process requires the labeling personnel to click the edge of the target point by point using the mouse or stylus to mark the accurate target contour. Although this traditional labeling method has certain advantages in terms of precision, the actual operation process is complex and inefficient, especially in the case of a large number of targets, complex shapes, or irregular boundaries, manual labeling needs to be positioned and confirmed repeatedly and carefully, which consumes a lot of time, and the labeling efficiency is difficult to improve. At the same time, due to the long-term and repetitive operation of the labeling task, the labeling personnel is prone to visual fatigue and decreased attention, which further affects the stability and consistency of the labeled data and reduces the overall labeling quality.
[0004] On the other hand, with the continuous progress of human-computer interaction technology, facial pose detection and eye movement tracking technology have shown excellent application prospects in virtual reality (VR), augmented reality (AR), human-computer collaborative operation, and other fields. Facial pose detection technology can intuitively and in real time reflect the user's head orientation, and can efficiently capture the operator's attention focus. Eye movement tracking technology can directly capture the user's gaze trajectory and accurately reflect the user's current focus area and intention. These two technologies have been successfully applied in attention analysis, interface interaction control, and other fields.
[0005] However, the current facial pose detection and eye movement tracking technologies are still in the blank stage in the application of image data labeling. Most of the existing image labeling schemes fail to effectively utilize the advantages of these two interactive technologies, and there is currently no complete technical system that combines them with automatic image segmentation algorithms to achieve accurate, efficient, and real-time labeling. SUMMARY
[0006] The application aims to solve the technical problem that the existing image data labeling field does not organically combine the two interactive technologies of face posture detection and eye movement tracking with automatic image segmentation algorithms, resulting in poor labeling accuracy and low efficiency, and provides an interactive image labeling system and method based on face posture and eye movement tracking positioning.
[0007] In order to achieve the above application purposes, the application provides the following technical solutions:
[0008] An interactive image labeling system based on face posture and eye movement tracking positioning, characterized in that:
[0009] The interactive image labeling system comprises an image data acquisition module, an image preprocessing module, a face posture detection module, an eye movement tracking module, a screen coordinate mapping module, a screen coordinate filtering module and an interactive control module.
[0010] The image data acquisition module comprises a depth camera and an infrared eye movement tracking device based on iris reflection, the depth camera is used to acquire continuous multiple frames of color face images and corresponding depth maps of an operator, the infrared eye movement tracking device based on iris reflection is used to acquire continuous multiple frames of near-infrared eye images of the operator, the output ends of the two are electrically connected with the input end of the image preprocessing module, and the input end of the image preprocessing module is also used to input a to-be-labeled image; the image preprocessing module is used to preprocess the continuous multiple frames of color face images and near-infrared eye images to obtain continuous multiple frames of preprocessed face images and preprocessed eye images respectively; the output ends of the image preprocessing module are electrically connected with the input ends of the face posture detection module and the eye movement tracking module respectively; the face posture detection module is used to obtain two-dimensional coordinates of face key points with the depth camera plane as a reference and the camera position of the depth camera as an origin according to the continuous multiple frames of preprocessed face images and corresponding depth maps, and the eye movement tracking module is used to obtain pupil center coordinates according to the continuous multiple frames of preprocessed face images; the output ends of the face posture detection module and the eye movement tracking module are electrically connected with the input end of the screen coordinate mapping module, and the screen coordinate mapping module is used to convert the two-dimensional coordinates of the face key points or the pupil center coordinates with the depth camera plane as a reference and the camera position of the depth camera as an origin into two-dimensional pointing coordinates on a screen; the screen coordinate mapping module, the screen coordinate filtering module and the interactive control module are electrically connected in sequence, and the input end of the interactive control module is also used to input the to-be-labeled image.
[0011] The screen coordinate filtering module is used to perform secondary sliding mean filtering on the two-dimensional pointing coordinates on the screen to obtain filtered stable coordinates, and the interactive control module is used to obtain an automatic labeled image according to the filtered stable coordinates.
[0012] Further, the face posture detection module comprises, which are electrically connected in sequence, a face detection and key point extraction module, a key point mean filtering module and a head posture estimation module, an input end of the face detection and key point extraction module is electrically connected with an output end of the image preprocessing module, and an output end of the head posture estimation module is electrically connected with an input end of the screen coordinate mapping module.
[0013] The eye movement tracking module comprises, which are electrically connected in sequence, an iris detection module, a pupil detection module and a pupil center positioning module, an input end of the iris detection module is electrically connected with an output end of the image preprocessing module, and the pupil center positioning module is electrically connected with an input end of the screen coordinate mapping module.
[0014] Meanwhile, the application also provides an interactive image labeling method based on face posture and eye movement tracking positioning, which adopts the interactive image labeling system based on face posture and eye movement tracking positioning, and has the speciality that comprises the following steps.
[0015] S1, a depth camera of an image data acquisition module collects continuous multiple frames of color face images and corresponding depth maps of an operator, and a near-infrared eye movement tracking device based on iris reflection collects continuous multiple frames of near-infrared eye images of the operator;
[0016] S2, an image preprocessing module pre-processes the continuous multiple frames of color face images and near-infrared eye images, and respectively obtains continuous multiple frames of pre-processed face images and pre-processed eye images;
[0017] S3, input the image to be labeled into the image preprocessing module, and the image preprocessing module judges whether the scene of the image to be labeled is complex, if not, step S4 is executed, and if yes, step S5 is executed;
[0018] S4, a face posture detection module processes the continuous multiple frames of pre-processed face images and corresponding depth maps, and obtains two-dimensional coordinates of face key points with the depth camera plane as the reference and the camera position of the depth camera as the origin, and then step S6 is executed;
[0019] S5, an eye movement tracking module processes the continuous multiple frames of pre-processed eye images, and obtains pupil center coordinates;
[0020] S6, the two-dimensional coordinates with the depth camera plane as the reference and the camera position of the depth camera as the origin are transmitted into the screen coordinate mapping module, and are converted into two-dimensional pointing coordinates on the screen through normalization mapping;
[0021] Or, the pupil center coordinates are transmitted into the screen coordinate mapping module, and are converted into two-dimensional pointing coordinates on the screen through normalization mapping;
[0022] S7, the two-dimensional pointing coordinate is transmitted into a screen coordinate filtering module for secondary sliding mean filtering to obtain a filtered stable coordinate;
[0023] S8, the filtered stable coordinate is transmitted into an interactive control module, when the interactive control module receives the filtered stable coordinate, an image to be labeled is input, when the input of the image to be labeled is completed, a segmentation model built-in the interactive control module is activated to perform multi-scale feature extraction on the image to be labeled to construct deep semantic features and detailed information;
[0024] Subsequently, the interactive control module inputs the filtered stable coordinate into a feature map built-in the segmentation model to guide the segmentation model to focus on a target region concerned by the operator;
[0025] Then, a preliminary segmentation result of the target region is generated through processing of an encoder-decoder structure of the segmentation model, the preliminary segmentation result of the target region is further refined and optimized through post-processing technology, and a confirmation point corresponding to one of the filtered stable coordinates in the image to be labeled is labeled to obtain an automatically labeled image;
[0026] S9, whether multi-point labeling is needed is judged, if yes, step S8 is returned to label another confirmation point until all the confirmation points are labeled and step S10 is executed; if not, step S10 is directly executed;
[0027] S10, the automatically labeled image is output, and interactive image labeling based on facial pose and eye movement tracking positioning is completed.
[0028] Further, step S4 specifically includes:
[0029] S4.1, a face detection and key point extraction module extracts two-dimensional image coordinates of facial key points in continuous multiple frames of preprocessed facial images:
[0030] ;
[0031] wherein, represents the two-dimensional image coordinates of the i-th facial key point, is the number of facial key points;
[0032] S4.2, in a key point mean filtering module, the two-dimensional coordinates of the facial key points are converted into three-dimensional coordinates of the facial key points in combination with the depth values of corresponding pixels in a depth map :
[0033] ;
[0034] wherein: is the three-dimensional coordinates of the facial key points, is the intrinsic parameter matrix of the depth camera, is the transpose of the matrix;
[0035] S4.3. The three-dimensional coordinates of the obtained facial key points are combined into a head local point cloud :
[0036] ;
[0037] Then calculate the local point cloud of the head The mean of to determine the initial reference position of the head posture:
[0038] ;
[0039] in: is the initial reference position of the head posture;
[0040] S4.4. 3D coordinates of facial key points Perform centralization to obtain the three-dimensional coordinates of the facial key points after centralization , which constitutes the centralized head local point cloud ;
[0041] ;
[0042] S4.5. Centralized head local point cloud The 3D coordinates of the centered facial key points in Perform sliding mean filtering to obtain the three-dimensional coordinates of the filtered facial key points , which constitutes a stable centralized head local point cloud :
[0043] ;
[0044] in: Indicates the The facial landmarks are The three-dimensional coordinates of the frame, is the window length of the sliding average, is the frame number of the current moment;
[0045] S4.6. Stabilizing the centralized head local point cloud in the head pose estimation module Perform singular value decomposition:
[0046] ;
[0047] in: is the left singular vector matrix, is a diagonal matrix of singular values, is a right singular vector matrix, is a transpose of matrix;
[0048] S4.7, taking the third column of the right singular vector matrix as a principal normal vector, representing the orientation of the head pose:
[0049] ;
[0050] wherein: is the principal normal vector;
[0051] S4.8, solving the vertical pitch angle and the horizontal yaw angle of the head pose based on the principal normal vector :
[0052] ;
[0053] ;
[0054] wherein: is the vertical pitch angle of the head pose, is the horizontal yaw angle of the head pose, is the component of the principal normal vector in the X-axis direction, is the component of the principal normal vector in the Y-axis direction, is the component of the principal normal vector in the Z-axis direction;
[0055] S4.9, converting the vertical pitch angle and the horizontal yaw angle of the head pose into two-dimensional coordinates with the depth camera plane as the reference and the camera position of the depth camera as the origin , and then performing step S6:
[0056] ; ;
[0057] wherein: is the vertical distance between the th facial key point and the depth camera (11) plane.
[0058] Further, step S5 is specifically:
[0059] S5.1, the iris detection module adopts eye key point positioning on the preprocessed eye images of the continuous multiple frames, locates the region surrounded by the eye key points of the operator, and obtains the gray mean value of the region surrounded by the eye key points and the standard deviation , and the iris detection module collects the iris image of the operator;
[0060] S5.2, the preprocessed eye images of the continuous multiple frames are processed by a binaryzation method, and a binary image is generated according to the adaptive threshold value ; :
[0061] ;
[0062] ;
[0063] in: is the horizontal coordinate of the binary image, is the vertical coordinate of the binary image, The iris image of the operator collected by the iris detection module, is the empirical coefficient, and ;
[0064] S5.3. Detecting the optimal circle center in a binary image using Hough circle transform and radius , is the X-axis coordinate of the optimal circle center, is the Y-axis coordinate of the optimal circle center;
[0065] S5.4. In the pupil detection module, define the operator's iris image The center of the circle is the center, and the radius is The local area is used as the region of interest for pupil detection:
[0066] ;
[0067] in: is the region of interest for pupil detection, is the scaling factor, and ;
[0068] S5.5. Construct pupil candidate binary map based on the region of interest of pupil detection:
[0069] ;
[0070] in: is the pupil candidate binary map;
[0071] S5.6. Perform morphological opening on the pupil candidate binary image to obtain the optimized pupil binary image:
[0072] ;
[0073] in: is the optimized pupil binary image, The radius in the morphological operation is A disk-shaped template of pixels;
[0074] S5.7, in the pupil center positioning module, the optimized pupil binary graph is subjected to centroid positioning, and finally the pupil center coordinates are output :
[0075] ;
[0076] .
[0077] Further, step S6 is specifically:
[0078] The two-dimensional coordinates with the depth camera plane as the reference and the camera position of the depth camera as the origin are transmitted into the screen coordinate mapping module, and are converted into two-dimensional pointing coordinates on the screen through normalization mapping :
[0079] ;
[0080] ;
[0081] wherein: and are the horizontal and vertical coordinates of the screen center respectively, is the scale factor of the X axis, is the scale factor of the Y axis;
[0082] Alternatively, the pupil center coordinates are transmitted into the screen coordinate mapping module, and are converted into two-dimensional pointing coordinates on the screen through normalization mapping :
[0083] ;
[0084] wherein: is a 2x2 calibration matrix, is a translation vector.
[0085] Further, step S7 is specifically:
[0086] The two-dimensional pointing coordinates are transmitted into the screen coordinate filtering module for secondary sliding mean filtering, to obtain the filtered stable coordinates :
[0087] ;
[0088] ;
[0089] wherein: is the window length of the sliding average, is the frame number at the current time, is the original screen X axis coordinate value obtained at the frame, To the first The original screen Y-axis coordinate value obtained in the frame.
[0090] Further, in step S8, the segmentation model is static_edgeflow_cocolvis.
[0091] Compared with the prior art, the present application has the beneficial effects that:
[0092] (1) The interactive image labeling system based on facial pose and eye tracking positioning provided by the present application realizes rapid positioning by combining facial pose detection and eye tracking technology. This system does not rely on the traditional mouse point-by-point clicking labeling method, greatly shortens the data labeling period, and significantly improves the large-scale image data processing efficiency. At the same time, the operator does not need to be distracted by operating the mouse, so that the hands can always be left on the keyboard, thereby realizing efficient collaborative operation, reducing fatigue caused by long-term repetitive labor, and further improving the sustainability and stability of the labeling work.
[0093] (2) The interactive image labeling method based on facial pose and eye tracking positioning provided by the present application constructs a complete interactive and automatic labeling closed-loop method. Through modular design, the facial pose detection and eye tracking two interactive modes are independently run. By judging the complexity of the scene, the appropriate interactive mode is selected, which helps to save subsequent processing time. Through secondary sliding mean filtering processing, the error caused by the data noise of the depth camera or the near-infrared eye tracking device based on iris reflection and the micro-motion of the operator is avoided. The stable coordinates after filtering cooperate with the segmentation model of the interactive control module to finely extract the target region. In the interactive control module, the traditional segmentation model is organically combined with the operator interaction information, greatly simplifying the manual labeling process. This method has flexible expandability and compatibility, and is suitable for the construction of large-scale image data sets and various computer vision applications. BRIEF DESCRIPTION OF DRAWINGS
[0094] Figure 1 The figure is the architecture diagram of an embodiment of the interactive image labeling system based on facial pose and eye tracking positioning of the present application;
[0095] Figure 2 The figure is a schematic diagram of the interactive image labeling system based on facial pose and eye tracking positioning of the present application in a use state;
[0096] Figure 3 In an embodiment of the interactive image labeling method based on facial pose and eye tracking positioning of the present application, the comparison between the coordinates after step S7 processing and the unprocessed coordinates is shown in the figure, where a is the unprocessed coordinates and b is the coordinates after step S7 processing;
[0097] Figure 4 In the embodiment of the interactive image labeling method based on facial pose and eye movement tracking positioning of the present application, when the scene is simple, the automatic labeled image based on facial pose positioning is compared with the image to be labeled, where a is the image to be labeled and b is the automatic labeled image.
[0098] Figure 5 In the embodiment of the interactive image labeling method based on facial pose and eye movement tracking positioning of the present application, when the scene is complex, the automatic labeled image based on eye movement tracking positioning is compared with the image to be labeled, where a is the image to be labeled and b is the automatic labeled image.
[0099] The reference signs are explained as follows:
[0100] 1-image data acquisition module, 11-depth camera, 12-near-infrared eye movement tracking device based on iris reflection; 2-image preprocessing module; 3-facial pose detection module, 31-face detection and key point extraction module, 32-key point mean filtering module, 33-head pose estimation module; 4-eye movement tracking module, 41-iris detection module, 42-pupil detection module, 43-pupil center positioning module; 5-screen coordinate mapping module, 6-screen coordinate filtering module, 7-interactive control module. DETAILED DESCRIPTION
[0101] The present application will be further described below in combination with the drawings and exemplary embodiments.
[0102] Reference Figure 1 , Figure 2 The interactive image labeling system based on facial pose and eye movement tracking positioning of the present application comprises an image data acquisition module 1, an image preprocessing module 2, a facial pose detection module 3, an eye movement tracking module 4, a screen coordinate mapping module 5, a screen coordinate filtering module 6, and an interactive control module 7.
[0103] The image data acquisition module 1 comprises a depth camera 11 and a near-infrared eye movement tracking device 12 based on iris reflection, as shown in Figure 2 When in use, the camera of the depth camera 11 is aligned with the face of the operator, responsible for acquiring continuous multiple frames of color face images and corresponding depth maps of the operator, the depth map provides distance information for each pixel of the color face image, so that the two-dimensional image features can be mapped to three-dimensional space, and the preliminary geometric structure of the scene is constructed.
[0104] The near-infrared eye movement tracking device 12 based on iris reflection continuously captures the eye region of the operator, and acquires continuous multiple frames of near-infrared eye images of the operator.
[0105] The output ends of the depth camera 11 and the near-infrared eye movement tracking device 12 based on iris reflection are respectively electrically connected with the input end of the image preprocessing module 2. In the image preprocessing module 2, the color face image and the near-infrared eye image will be preprocessed to obtain a plurality of continuous frames of preprocessed face images and preprocessed eye images. The common preprocessing operations of the color face image include, for example, mirror flipping, grayscale conversion, noise suppression, etc., which can ensure the accuracy of subsequent key point extraction. The near-infrared eye image itself is a grayscale image, so it does not need to perform grayscale conversion in preprocessing, and the remaining preprocessing operations are the same as those of the color face image.
[0106] The input end of the image preprocessing module 2 is also used to input the image to be labeled. In the image preprocessing module 2, the scene complexity of the image to be labeled is also judged to facilitate the selection of the subsequent interactive labeling mode. The output end of the image preprocessing module 2 is respectively electrically connected with the input end of the face posture detection module 3 and the eye movement tracking module 4. The face posture detection module 3 includes a face detection and key point extraction module 31, a key point mean filtering module 32 and a head posture estimation module 33 which are electrically connected in sequence. The input end of the face detection and key point extraction module 31 is electrically connected with the output end of the image preprocessing module 2, and the output end of the head posture estimation module 33 is electrically connected with the input end of the screen coordinate mapping module 5. The face posture detection module 3 is responsible for processing the plurality of continuous frames of preprocessed face images and corresponding depth maps obtained by the image preprocessing module 2 to obtain a two-dimensional coordinate with the depth camera 11 plane as the reference and the camera position of the depth camera 11 as the origin.
[0107] The eye movement tracking module 4 includes an iris detection module 41, a pupil detection module 42 and a pupil center positioning module 43 which are electrically connected in sequence. The input end of the iris detection module 41 is electrically connected with the output end of the image preprocessing module 2, and the pupil center positioning module 43 is electrically connected with the input end of the screen coordinate mapping module 5. The entire eye movement tracking module 4 is responsible for processing the plurality of continuous frames of preprocessed eye images obtained by the image preprocessing module 2 to obtain the pupil center coordinates.
[0108] And screen coordinate mapping module 5, screen coordinate filtering module 6 and interactive control module 7 are electrically connected in turn, and the input end of interactive control module 7 is also used to input the image to be labeled. Screen coordinate mapping module 5 is responsible for converting the two-dimensional coordinates or pupil center coordinates with the depth camera 11 plane as the reference and the camera position of the depth camera 11 as the origin into two-dimensional pointing coordinates on the screen through normalization mapping. Screen coordinate filtering module 6 is responsible for performing secondary sliding mean filtering on the two-dimensional pointing coordinates on the screen to obtain stable coordinates after filtering. This processing avoids errors caused by data noise of the depth camera 11 or the near-infrared eye tracking device 12 based on iris reflection and micro-movement of the operator, so that the labeling result is more accurate, and interactive control module 7 obtains the automatic labeled image according to the stable coordinates after filtering.
[0109] Meanwhile, the application also provides an interactive image labeling method based on facial posture and eye tracking positioning, which adopts the above-mentioned interactive image labeling system based on facial posture and eye tracking positioning, and comprises the following steps:
[0110] S1, the depth camera 11 of the image data acquisition module 1 collects continuous multiple frames of color facial images and corresponding depth maps of the operator, and the near-infrared eye tracking device 12 based on iris reflection collects continuous multiple frames of near-infrared eye images of the operator;
[0111] S2, the image preprocessing module 2 pre-processes the continuous multiple frames of color facial images and near-infrared eye images to obtain continuous multiple frames of pre-processed facial images and pre-processed eye images respectively;
[0112] S3, the image to be labeled is input into the image preprocessing module 2, and the image preprocessing module 2 judges whether the scene of the image to be labeled is complex or not. If not, step S4 is executed, and if yes, step S5 is executed.
[0113] The judgment of whether the scene is complex or not is based on the following objective indicators for comprehensive judgment:
[0114] When the number of target objects to be labeled in the image to be labeled is less than 3, the maximum distance of target distribution is not more than 30% of the image width, the boundary contour is clear, the boundary blur degree is less than 20% (based on pixel-level boundary clarity evaluation), the background and the target object to be labeled have high contrast, and the average gray difference between the target object to be labeled and the background is greater than 50%, the scene complexity is low, such as Figure 4 a; otherwise, the scene complexity is high, such as Figure 5 a. Through scene judgment, it can be determined in advance which interactive mode to choose for labeling, saving subsequent processing time to obtain more efficient and accurate labeling effect.
[0115] S4, the face posture detection module 3 processes the pre-processed face images and the corresponding depth maps of the continuous multiple frames to obtain two-dimensional coordinates with the depth camera 11 plane as the reference and the camera position of the depth camera 11 as the origin, and then step S6 is executed;
[0116] And step S4 is specifically:
[0117] S4.1, the face detection and key point extraction module 31 extracts the two-dimensional image coordinates of the face key points in the pre-processed face images of the continuous multiple frames:
[0118]
[0119] Among them, represents the two-dimensional image coordinates of the i-th face key point, and n is the number of face key points;
[0120] S4.2, in the key point mean filtering module 32, the depth value of the corresponding pixel in the depth map is combined , the two-dimensional coordinates of the face key points are converted into three-dimensional coordinates of the face key points :
[0121]
[0122] Among them: is the three-dimensional coordinates of the face key points, is the intrinsic matrix of the depth camera 11, is the transpose of the matrix;
[0123] S4.3, the three-dimensional coordinates of the face key points obtained are composed into a head local point cloud :
[0124]
[0125] Then the mean value of the head local point cloud is calculated to determine the initial reference position of the head posture:
[0126]
[0127] Among them: is the initial reference position of the head posture;
[0128] S4.4, the three-dimensional coordinates of the face key points are processed to obtain the three-dimensional coordinates of the face key points after the centering processing , which constitute the centered head local point cloud ;
[0129] ;
[0130] S4.5. Centralized head local point cloud The 3D coordinates of the centered facial key points in Perform sliding mean filtering to obtain the three-dimensional coordinates of the filtered facial key points , which constitutes a stable centralized head local point cloud :
[0131] ;
[0132] in: Indicates the The facial landmarks are The three-dimensional coordinates of the frame, is the window length of the sliding average, is the frame number of the current moment;
[0133] S4.6, in the head pose estimation module 33, the stabilized centralized head local point cloud Perform singular value decomposition:
[0134] ;
[0135] in: is the left singular vector matrix, is a diagonal matrix of singular values, is the right singular vector matrix, is the transpose of the matrix;
[0136] S4.7. Take the right singular vector matrix The third column is the principal normal vector, indicating the orientation of the head pose:
[0137] ;
[0138] in: is the main normal vector;
[0139] S4.8, based on the principal normal vector Solve for the vertical pitch angle and horizontal yaw angle of the head posture:
[0140] ;
[0141] ;
[0142] in: is the vertical pitch angle of the head posture, is the horizontal yaw angle of the head posture, is the component of the principal normal vector in the X-axis direction, a component of the principal normal vector in the Y-axis direction, a component of the principal normal vector in the Z-axis direction;
[0143] S4.9, convert the vertical pitch angle and the horizontal yaw angle of the head posture into two-dimensional coordinates with the depth camera 11 plane as the reference and the camera position of the depth camera 11 as the origin , and then perform step S6:
[0144] ; ;
[0145] wherein: is the vertical distance between the i-th facial key point and the depth camera 11 plane.
[0146] S5, the eye tracking module 4 processes the preprocessed eye images of the continuous multiple frames to obtain the pupil center coordinates;
[0147] Step S5 is specifically:
[0148] S5.1, the iris detection module 41 uses eye key point positioning on the preprocessed eye images of the continuous multiple frames to locate the region surrounded by the eye key points of the operator, and obtains the gray mean value and the standard deviation of the region surrounded by the eye key points ; Meanwhile, the iris detection module 41 collects the iris image of the operator;
[0149] S5.2, a binary method is used on the preprocessed eye images of the continuous multiple frames to generate a binary image according to the adaptive threshold value :
[0150] ;
[0151] ;
[0152] wherein: is the horizontal coordinate of the binary image, is the vertical coordinate of the binary image, is the iris image of the operator collected by the iris detection module 41, is an empirical coefficient, and ;
[0153] S5.3, the best circle center and the radius are detected in the binary image through Hough circle transformation, is the X-axis coordinate of the best circle center, is the Y-axis coordinate of the best circle center;
[0154] S5.4, in the pupil detection module 42, define a local area with a center of the iris image of the operator and a radius of as the region of interest for pupil detection:
[0155] ;
[0156] wherein: is the region of interest for pupil detection, is a scaling factor, and ;
[0157] S5.5, construct a pupil candidate binary image according to the region of interest for pupil detection:
[0158] ;
[0159] wherein: is the pupil candidate binary image;
[0160] S5.6, perform a morphological opening operation on the pupil candidate binary image to obtain an optimized pupil binary image:
[0161] ;
[0162] wherein: is the optimized pupil binary image, is a disc-shaped template with a radius of pixels in the morphological operation;
[0163] S5.7, in the pupil center positioning module 43, perform centroid positioning on the optimized pupil binary image to finally output the pupil center coordinates :
[0164] ;
[0165] .
[0166] S6, transfer the two-dimensional coordinates with the depth camera 11 plane as the reference and the camera position of the depth camera 11 as the origin into the screen coordinate mapping module, and convert it into the two-dimensional pointing coordinates on the screen through normalized mapping :
[0167] ;
[0168] ;
[0169] wherein: and are the horizontal and vertical coordinates of the screen center, respectively, is the scale factor of the X axis, is the scale factor of the Y axis.
[0170] Or, the pupil center coordinates are transmitted into the screen coordinate mapping module, and are converted into two-dimensional pointing coordinates on the screen through normalization mapping :
[0171] ;
[0172] wherein: is a 2x2 calibration matrix, is a translation vector.
[0173] S7, the two-dimensional pointing coordinates are transmitted into the screen coordinate filtering module 6 for secondary sliding average filtering to obtain filtered stable coordinates :
[0174] ;
[0175] ;
[0176] wherein: is the window length of the sliding average, is the frame number at the current moment, is the original screen X axis coordinate value obtained at the frame, is the original screen Y axis coordinate value obtained at the frame.
[0177] As shown in Figure 3 a and Figure 3 b, it can be seen that before the secondary sliding average filtering, the coordinate distribution is relatively dispersed, and after the processing, the coordinate distribution is more concentrated and the fluctuation is smaller, which significantly improves the pointing stability. This step effectively weakens the influence caused by the data noise of the depth camera 11, the data noise of the near-infrared eye movement tracking device 12 based on iris reflection, the slight tremor of the head or the drift of the line of sight, etc.
[0178] S8, the filtered stable coordinates are transmitted into the interactive control module 7, which inputs the to-be-labeled image when receiving the filtered stable coordinates, and activates the segmentation model built-in the interactive control module 7 to perform multi-scale feature extraction on the to-be-labeled image to construct deep semantic features and detail information after the input of the to-be-labeled image is completed.
[0179] Subsequently, the interactive control module 7 inputs the filtered stable coordinates into the feature map built-in the segmentation model to guide the segmentation model to focus on the target region concerned by the operator.
[0180] Then, after being processed by the encoder-decoder structure of the segmentation model, a preliminary segmentation result of the target area is generated. The preliminary segmentation result of the target area is further refined and optimized through post-processing technology, and the confirmation point corresponding to one of the filtered stable coordinates in the image to be annotated is annotated to obtain an automatically annotated image.
[0181] The segmentation model uses static_edgeflow_cocolvis, which organically combines traditional segmentation models with operator interaction information in interactive control module 7, greatly simplifying the manual labeling process. Post-processing technology uses conditional random fields (CRFs) to optimize pixel labels, ensuring that edges are consistent with the color and gradient of the original image. This refines and optimizes the initial segmentation results of the target area.
[0182] S9, determine whether multiple points need to be marked. If yes, return to step S8 and mark another confirmation point. After all confirmation points are marked, execute step S10. If no, execute step S10 directly.
[0183] S10: Output the automatically annotated image, and complete the interactive image annotation based on facial posture and eye tracking positioning.
[0184] like Figure 4 As shown in the figure, this image needs to be labeled with fire hydrants. The number of objects to be labeled is small, and the boundary contours of the target are clear and the contrast with the background is high. Therefore, it can be defined as a simple scene. It is suitable for the image annotation interaction method based on facial posture positioning, in which Figure 4 a is the image to be labeled, Figure 4 b is the automatically annotated image, from which it can be seen that the fire hydrant is clearly labeled.
[0185] like Figure 5 As shown in the figure, this image needs to be labeled with cattle. There are many objects that need to be labeled in the entire image, and the postures of the cattle are different, and the boundary contours are not clear, so it is defined as a complex scene. It is suitable for the image annotation interaction method based on eye tracking positioning, in which Figure 5 a is the image to be labeled, Figure 5 b is an automatically annotated image. It can be seen that after the image is annotated using this method, all the cows in the herd are clearly marked.
[0186] The embodiments described above are merely descriptions of specific implementation methods of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should fall within the scope of protection determined by the claims of the present invention.
Claims
1. An interactive image annotation system based on facial pose and eye movement tracking positioning, characterized in that: it comprises an image data acquisition module (1), an image preprocessing module (2), a facial pose detection module (3), an eye movement tracking module (4), a screen coordinate mapping module (5), a screen coordinate filtering module (6), and an interactive control module (7); the image data acquisition module (1) comprises a depth camera (11) and an iris reflection-based near-infrared eye movement tracking device (12), the depth camera (11) is used to acquire continuous multiple frames of color facial images and corresponding depth maps of an operator, the iris reflection-based near-infrared eye movement tracking device (12) is used to acquire continuous multiple frames of near-infrared eye images of the operator, and the output ends of the two are respectively electrically connected with the input end of the image preprocessing module (2), and the input end of the image preprocessing module (2) is also used to input a to-be-labeled image; the image preprocessing module (2) is used to preprocess the continuous multiple frames of color facial images and near-infrared eye images to obtain continuous multiple frames of preprocessed facial images and preprocessed eye images respectively; the output ends of the image preprocessing module (2) are respectively electrically connected with the input ends of the facial pose detection module (3) and the eye movement tracking module (4); the facial pose detection module (3) is used to obtain two-dimensional coordinates of facial key points with the depth camera (11) plane as the reference and the camera position of the depth camera (11) as the origin according to the continuous multiple frames of preprocessed facial images and corresponding depth maps, and the eye movement tracking module (4) is used to obtain pupil center coordinates according to the continuous multiple frames of preprocessed facial images; the output ends of the facial pose detection module (3) and the eye movement tracking module (4) are respectively electrically connected with the input end of the screen coordinate mapping module (5), and the screen coordinate mapping module (5) is used to convert the two-dimensional coordinates of the facial key points or the pupil center coordinates with the depth camera (11) plane as the reference and the camera position of the depth camera (11) as the origin into two-dimensional pointing coordinates on the screen; the screen coordinate mapping module (5), the screen coordinate filtering module (6), and the interactive control module (7) are electrically connected in sequence, and the input end of the interactive control module (7) is also used to input the to-be-labeled image; the screen coordinate filtering module (6) is used to perform secondary sliding mean filtering on the two-dimensional pointing coordinates on the screen to obtain filtered stable coordinates, and the interactive control module (7) is used to obtain an automatically annotated image according to the filtered stable coordinates.
2. The interactive image annotation system based on facial pose and eye movement tracking positioning according to claim 1, characterized in that: the facial pose detection module (3) comprises a face detection and key point extraction module (31), a key point mean filtering module (32), and a head pose estimation module (33) which are electrically connected in sequence, the input end of the face detection and key point extraction module (31) is electrically connected with the output end of the image preprocessing module (2), and the output end of the head pose estimation module (33) is electrically connected with the input end of the screen coordinate mapping module (5). The eye movement tracking module (4) comprises an iris detection module (41), a pupil detection module (42) and a pupil center positioning module (43) connected in sequence, the input end of the iris detection module (41) is electrically connected with the output end of the image preprocessing module (2), and the pupil center positioning module (43) is electrically connected with the input end of the screen coordinate mapping module (5).
3. An interactive image annotation method based on facial pose and eye movement tracking positioning, using an interactive image annotation system based on facial pose and eye movement tracking positioning according to claim 1 or 2, characterized in that, It comprises the following steps: S1, the depth camera (11) of the image data acquisition module (1) collects continuous multiple frames of color face images and corresponding depth maps of the operator, and the near-infrared eye movement tracking device (12) collects continuous multiple frames of near-infrared eye images of the operator based on iris reflection; S2, the image preprocessing module (2) pre-processes the continuous multiple frames of color face images and near-infrared eye images to obtain continuous multiple frames of pre-processed face images and pre-processed eye images respectively; S3, input the image to be labeled into the image preprocessing module (2), and the image preprocessing module (2) judges whether the scene of the image to be labeled is complex, if not, execute step S4, if yes, execute step S5; S4, the face posture detection module (3) processes the continuous multiple frames of pre-processed face images and corresponding depth maps to obtain the two-dimensional coordinates of the face key points with the depth camera (11) plane as the reference and the camera position of the depth camera (11) as the origin, and then executes step S6; S5, the eye movement tracking module (4) processes the continuous multiple frames of pre-processed eye images to obtain the pupil center coordinates; S6, the two-dimensional coordinates with the depth camera (11) plane as the reference and the camera position of the depth camera (11) as the origin are transmitted into the screen coordinate mapping module (5), which is converted into two-dimensional pointing coordinates on the screen through normalization mapping; Or, the pupil center coordinates are transmitted into the screen coordinate mapping module (5), which is converted into two-dimensional pointing coordinates on the screen through normalization mapping; S7, the two-dimensional pointing coordinates are transmitted into the screen coordinate filtering module (6) for secondary sliding mean filtering to obtain the filtered stable coordinates; S8, the filtered stable coordinates are transmitted into the interactive control module (7), when the interactive control module (7) receives the filtered stable coordinates, the image to be labeled is input, when the image to be labeled is input, the segmentation model built-in the interactive control module (7) is activated, the multi-scale feature extraction of the image to be labeled is performed, and the deep semantic feature and the detail information are constructed; Then, the interactive control module (7) inputs the filtered stable coordinates into the feature map built-in the segmentation model to guide the segmentation model to focus on the target region concerned by the operator; Then, after the processing of the encoder-decoder structure of the segmentation model, the preliminary segmentation result of the target region is generated, the preliminary segmentation result of the target region is further refined and optimized through post-processing technology, and the confirmation point corresponding to one of the filtered stable coordinates in the image to be labeled is labeled to obtain the automatic labeled image; S9, judging whether multi-point labeling is needed, if yes, returning to step S8 to label another confirmation point until all confirmation points are labeled and step S10 is executed; if no, directly executing step S10; S10, outputting the automatic labeling image, completing the interactive image labeling based on the facial pose and eye movement tracking positioning.
4. The method of claim 3, wherein, Step S4 is specifically: S4.1, the face detection and key point extraction module (31) extracts the two-dimensional image coordinates of the facial key points in the preprocessed facial images of continuous multiple frames: p i = (u i ,v i ), i = 1, 2,... N; wherein p i represents the two-dimensional image coordinates of the i-th facial key point, and N is the number of facial key points; S4.2, in the key point mean filtering module (32), combine the depth value d of the corresponding pixel in the depth map i , convert the two-dimensional coordinates of the face key points into three-dimensional coordinates P of the face key points i : P i = d i · K -1 [u i , v i , 1] T = (X i , Y i , Z i ); wherein: (X i ,Y i ,Z i ) are the three-dimensional coordinates of the facial key points, K is the intrinsic matrix of the depth camera (11), and T is the transpose of the matrix; S4.3, the obtained three-dimensional coordinates of the facial key points are composed into a head local point cloud P: Then, the mean value of the head local point cloud P is calculated to determine the initial reference position of the head pose: wherein: is the initial reference position for the head pose; S4.4, the three-dimensional coordinates P of the facial key points i centralization processing is performed to obtain the three-dimensional coordinates P of the facial key points after the centralization processing i , which constitutes the centralization head local point cloud P'; S4.5, the three-dimensional coordinates P of the facial key points in the centralized head local point cloud P' after the centralization processing i sliding mean filtering is performed to obtain the three-dimensional coordinates of the filtered facial key points which constitutes a stable centralized head local point cloud wherein: P i (j) represents the three-dimensional coordinates of the i-th facial landmark at the j-th frame, M is the window length of the sliding average, and t is the frame number of the current time. S4.
6. Stabilized centralized head local point cloud in the head pose estimation module (33) Perform singular value decomposition: Wherein: U is a left singular vector matrix, S is a singular value diagonal matrix, V is a right singular vector matrix, and T is a transpose of the matrix; S4.7, taking the third column of the right singular vector matrix V as the principal normal vector, representing the orientation of the head pose: n=V(:,3); Wherein: n is the principal normal vector; S4.8, solving the vertical pitch angle and the horizontal yaw angle of the head pose based on the principal normal vector n: where: θ is the vertical pitch angle of the head pose, is the horizontal yaw angle of the head pose, n x is the component of the principal normal vector in the X-axis direction, n y is the component of the principal normal vector in the Y-axis direction, n z is the component of the principal normal vector in the Z-axis direction; S4.
9. Convert the vertical pitch angle and the horizontal yaw angle of the head pose into two-dimensional coordinates (x head , y head ) with respect to the depth camera (11) plane and with the camera position of the depth camera (11) as the origin, followed by step S6: y head = L i tan(θ); where: L i is the vertical distance between the i-th facial landmark and the depth camera (11) plane.
5. The method of claim 4, wherein, Step S5 is specifically: S5.1, the iris detection module (41) adopts eye key point positioning on the preprocessed eye images of the continuous multiple frames, locates the area surrounded by the eye key points of the operator, and obtains the gray mean μ of the area surrounded by the eye key points roi and the standard deviation σ roi At the same time, the iris detection module (41) collects the iris image of the operator; S5.2, the pre-processed eye images of continuous multiple frames are binarized, and the adaptive threshold T is determined according to the maximum value and the minimum value of the image iris generating a binary image B iris (u, v): T iris = μ roi - αs roi ; wherein: u is the horizontal coordinate of the binary image, v is the vertical coordinate of the binary image, I iris (u,v) is the iris image of the operator collected by the iris detection module (41), and a is an empirical coefficient, and a ∈ [0.8, 1.2]. S5.3, detecting the best circle center (u0, v0) and radius R in the binary image by the Hough circle transform iris u0 is the X-axis coordinate of the best circle center, and v0 is the Y-axis coordinate of the best circle center; S5.
4. In the pupil detection module (42), define the local region of the image I iris (u,v) as the center of the circle, and the radius kR iris as the region of interest for pupil detection: Ω pupil = {(u, v) | (u - u0) 2 + (v - v0) 2 ≤ (kR iris ) 2}; where: Ω pupil is the region of interest for pupil detection, k is a scaling factor, and k e [0.1, 0.8]. S5.5, constructing a pupil candidate binary graph according to the region of interest of the pupil detection: wherein: B pupil (u, v) is a pupil candidate binary map; S5.6, performing morphological opening operation on the pupil candidate binary graph to obtain an optimized pupil binary graph: B′ pupil (u,v) = Opening(B pupil (u,v), disk(r)); wherein: B' pupil (u, v) is the optimized pupil binary image, and disk(r) is a circular disk-shaped template with a radius of r pixels in the morphological operation. S5.
7. In the pupil center positioning module (43), the optimized pupil binary image is subjected to centroid positioning, and the pupil center coordinates (x pupil , y pupil ) are finally output:
6. The method of claim 5, wherein the method further comprises: Step S6 is specifically: The two-dimensional coordinates with the depth camera (11) plane as the reference and the camera position of the depth camera (11) as the origin are transmitted into the screen coordinate mapping module (5), and are converted into two-dimensional pointing coordinates (s creen ,y screen ) on the screen through normalization mapping. x screen = x0+ K x · x head ; y screen = y0+ K y · y head ; wherein: x0 and y0 are the horizontal and vertical coordinates of the center of the screen, K x is the scale factor for the X axis, K y is the scale factor for the Y axis; Alternatively, the pupil center coordinates are passed into a screen coordinate mapping module (5) which converts them into a two-dimensional pointing coordinate (x screen ,y screen ) on the screen by a normalizing mapping: Wherein: A is a 2*2 calibration matrix, and b is a translation vector.
7. The method of claim 6, wherein the method further comprises: Step S7 is specifically: The two-dimensional pointing coordinates (x screen ,y screen ) are transmitted into the screen coordinate filtering module (6) for secondary sliding mean filtering, and the filtered stable coordinates wherein: M is the window length of the moving average, t is the frame number of the current time, x screen (j) is the original screen X-axis coordinate value obtained in the jth frame screen (j) is the original screen Y-axis coordinate value obtained in the jth frame.
8. The interactive image labeling method based on facial pose and eye movement tracking positioning according to claim 3, characterized in that, In step S8, the segmentation model is static_edgeflow_cocolvis.
Citation Information
Patent Citations
Image target segmentation system combining eye-movement tracking
CN106681484A
Focus recognition system based on eye movement information
CN118172578A