Eye key point detection method and device and electronic equipment
By combining the face and eye key point detection models and utilizing the global features of the face image and the local features of the eye area, the problem of inaccurate eye key point positioning in the AR-HUD system is solved, achieving higher positioning accuracy and better alignment of the HUD image with the actual road.
Patent Information
- Application Number
- CN202510021710.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-09-12
AI Technical Summary
In the existing technology, the problem of inaccurate positioning of eye key points in AR-HUD systems causes the HUD image to not fit the actual road.
By acquiring a facial image, using the facial key point detection model and the eye key point detection model, and combining the global features of the facial image and the local features of the eye area, the target coordinates of the eye key points are determined to improve positioning accuracy.
The effective combination of the global features of the face and the local features of the eyes improves the positioning accuracy of the key points of the eyes and ensures that the HUD image fits the actual road.
Smart Images

Figure CN120635972A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of automotive technology, and in particular to a method, device and electronic equipment for detecting key points of an eye. Background Art
[0002] In the field of automotive technology, head-up display (HUD) is the main function to assist drivers in driving. Conventional head-up displays can only provide drivers with a single flat digital information, which can easily distract the driver's attention. In order to alleviate the distraction of driver attention, AR-HUD has developed rapidly.
[0003] AR-HUD combines AR augmented reality technology and HUD head-up display technology, which can directly superimpose display effects on the real road surface; in related technologies, the driver's eye key points are tracked and the HUD image is calibrated based on the located eye key points to make the HUD image fit the actual road. However, there is a problem of inaccurate positioning of the eye key points. Summary of the Invention
[0004] In view of this, the embodiments of the present application propose an eye key point detection method, device and electronic device to solve the problem of inaccurate positioning of eye key points in related technologies.
[0005] The embodiments of the present application are implemented using the following technical solutions:
[0006] In a first aspect, an embodiment of the present application provides an eye key point detection method, the method comprising: acquiring a facial image; performing key point detection on the facial image based on a facial key point detection model to obtain key point coordinates of multiple facial key points in the facial image, wherein the multiple facial key points include multiple eye key points; based on the key point coordinates of the multiple eye key points, intercepting an eye area image in the facial image; performing eye key point detection on the eye area image based on the eye key point detection model to obtain reference coordinates of the multiple eye key points; and determining target coordinates of the multiple eye key points based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points.
[0007] In second aspect, an embodiment of the present application provides an eye key point detection device, which includes: an acquisition module for acquiring a face image; a face detection module for performing key point detection on the face image based on a face key point detection model to obtain key point coordinates of multiple face key points in the face image, wherein the multiple face key points include multiple eye key points; a capture module for capturing an eye area image in the face image based on the key point coordinates of the multiple eye key points; an eye detection module for performing eye key point detection on the eye area image based on the eye key point detection model to obtain reference coordinates of the multiple eye key points; and an output module for determining the target coordinates of the multiple eye key points based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points.
[0008] In a third aspect, an embodiment of the present application provides an electronic device, which includes a processor; a memory, wherein computer-readable instructions are stored in the memory, and when the computer-readable instructions are executed by the processor, the method described above is implemented.
[0009] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by a processor, the method described above is implemented.
[0010] The embodiments of the present application provide an eye key point detection method, device and electronic device, the method comprising: obtaining a face image; performing key point detection on the face image based on a face key point detection model to obtain key point coordinates of multiple face key points in the face image, wherein the multiple face key points include multiple eye key points; obviously, since the key point coordinates of the eye key points are obtained based on the detection of the face image, the key point coordinates of the eye key points are determined based on the global features of the face image; based on the key point coordinates of the multiple eye key points, an eye area image is intercepted in the face image; based on the eye key point coordinates .... The key point detection model performs eye key point detection on the eye area image to obtain the reference coordinates of multiple eye key points; obviously, since the reference coordinates of the eye key points are obtained based on the detection of the eye area image, the reference coordinates of the eye key points are more determined based on the local features of the eye area image; therefore, based on the key point coordinates of multiple eye key points and the reference coordinates of multiple eye key points, the target coordinates of multiple eye key points are determined; the target coordinates of the eye key points can effectively combine the global features of the face and the local features of the eyes, thereby improving the accuracy of eye key point positioning.
[0011] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.
[0013] Figure 1 A flow chart of a method for detecting key eye points according to an embodiment of the present application is shown.
[0014] Figure 2 An embodiment of the present application provides Figure 1 Flow chart of step S130 in FIG.
[0015] Figure 3 A schematic diagram of a detection frame captured in a face image according to an embodiment of the present application is shown.
[0016] Figure 4 A flow chart of a method for detecting key points of the eyes provided in another embodiment of the present application is shown.
[0017] Figure 5 An embodiment of the present application provides Figure 4 Flow chart of step S220 in FIG.
[0018] Figure 6 A schematic diagram of image changes of a sample face image involved in an embodiment of the present application is shown.
[0019] Figure 7 An embodiment of the present application provides Figure 4 Flow chart of step S240 in FIG.
[0020] Figure 8 An embodiment of the present application provides Figure 7 Flow chart of step S242 in FIG.
[0021] Figure 9 An embodiment of the present application provides Figure 8 Flow chart of step S330 in FIG.
[0022] Figure 10 A flow chart of a method for detecting key eye points provided in another embodiment of the present application is shown.
[0023] Figure 11 The figure shows a structural block diagram of an eye key point detection device provided in one embodiment of the present application.
[0024] Figure 12 A structural block diagram of an electronic device provided in one embodiment of the present application is shown. DETAILED DESCRIPTION
[0025] The embodiments of the present application are described in detail below. Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present application and are not to be construed as limiting the present application.
[0026] In order to enable those skilled in the art to better understand the present invention, the following will provide a clear and complete description of the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present invention.
[0027] In the following description, the terms "first\second" and the like are merely used to distinguish similar objects and do not represent a specific ordering of the objects. It is understood that "first\second" can be interchanged with a specific order or precedence where permitted, so that the embodiments of the present application described herein can be implemented in an order other than that illustrated or described herein. In the following description, references to "some embodiments or some embodiment modes" describe a subset of all possible embodiments, but it is understood that "some embodiments" can be the same subset or different subsets of all possible embodiments, and can be combined with each other without conflict.
[0028] In this document, "plurality" refers to two or more. "And / or" describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates that the related objects are in an "or" relationship.
[0029] In the field of automotive technology, head-up display (HUD) is the main function to assist drivers in driving. Conventional head-up displays can only provide drivers with a single flat digital information, which can easily distract the driver's attention. In order to alleviate the distraction of driver attention, AR-HUD has developed rapidly.
[0030] AR-HUD combines AR augmented reality technology and HUD head-up display technology, which can directly superimpose display effects on the real road surface; in related technologies, the driver's facial image is collected, the driver's eye key points are tracked through the facial image, and the HUD image is calibrated according to the located eye key points to make the HUD image fit the actual road. However, there is a problem of inaccurate positioning of the eye key points.
[0031] In order to solve the above problems, the present application provides an eye key point detection method, device and electronic device, the method including: obtaining a face image; performing key point detection on the face image based on a face key point detection model to obtain key point coordinates of multiple face key points in the face image, the multiple face key points including multiple eye key points; based on the key point coordinates of the multiple eye key points, intercepting an eye area image in the face image; performing eye key point detection on the eye area image based on the eye key point detection model to obtain reference coordinates of multiple eye key points; determining the target coordinates of the multiple eye key points based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points.
[0032] Through the method provided by the present application, since the key point coordinates of the eye key points are obtained based on the detection of the face image, the key point coordinates of the eye key points are determined based on the global features of the face image; similarly, since the reference coordinates of the eye key points are obtained based on the detection of the eye area image, the reference coordinates of the eye key points are more determined based on the local features of the eye area image; by obtaining the target coordinates of the eye key points through the key point coordinates and reference coordinates of the eye key points, the target coordinates of the eye key points can effectively combine the global features of the face and the local features of the eyes, thereby improving the accuracy of eye key point positioning.
[0033] The embodiments provided in this application will be described below with reference to the accompanying drawings.
[0034] See also Figure 1 , Figure 1 A flow chart of the eye key point detection method provided in an embodiment of the present application is provided. The eye key point detection method includes steps S110-S150:
[0035] S110: Acquire a facial image.
[0036] The face image is an image showing the face area to be detected for eye key point detection. It can be a face image collected in real time or a pre-given face image.
[0037] In some embodiments, the facial image may be a driver's facial image captured in real time by a camera in the vehicle facing the driver.
[0038] S120 , performing key point detection on the face image based on a face key point detection model to obtain key point coordinates of a plurality of face key points in the face image, where the plurality of face key points include a plurality of eye key points.
[0039] The facial key point detection model is pre-trained, and its training process will be described in subsequent embodiments. The facial key point detection model is a neural network model constructed using one or more neural networks. The facial key point detection model is used to identify the coordinates of multiple facial key points located in the facial region of a facial image. In some embodiments, a set of facial key points that the facial key point detection model is suitable for detecting can be pre-set. After training, the facial key point detection model can identify the coordinates of each facial key point in the set of facial key points in a facial image.
[0040] Facial key points refer to the key feature points on the face, such as the key feature points located in key parts such as the eyes, nose, mouth, eyebrows, and facial contours.
[0041] Multiple facial key points may include basic facial contour key points and key part key points; among them, basic facial contour key points are key points used to describe the outer contour of the face in the face image, and key part key points are used to describe the key points of the eyes, mouth, nose, eyebrows and other parts in the face image.
[0042] S130 , based on the key point coordinates of the multiple eye key points, intercepting an eye region image from the face image.
[0043] The captured eye area image includes multiple eye key points.
[0044] In some embodiments, the boundary of the eye pixel area in the face image can be determined based on the key point coordinates of multiple eye key points. Then, the eye pixel area is cut out from the face image according to the boundary of the eye pixel area as an eye area image.
[0045] In some embodiments, the boundary of the eye pixel area may be a boundary formed by a plurality of eye key points and distinguished from other facial key points except the eye key points.
[0046] In some embodiments, after determining the key point coordinates of multiple eye key points, target pixel points that meet the requirements are determined based on the relationship between the coordinates of each pixel point in the facial image and the key point coordinates of the eye key points, and the image formed by the target pixel points is determined as the eye area image; for example, the smallest vertical coordinate is determined from the key point coordinates of multiple eye key points as the reference vertical coordinate, and then the pixel points in the facial image whose vertical coordinates are greater than the reference vertical coordinates are determined as target pixel points, and the image formed by the target pixel points is determined as the eye area image; wherein the vertical coordinates of each pixel point in the facial image gradually decrease from the eyes to the mouth in the vertical direction.
[0047] In other embodiments, see Figure 2 , Figure 2The embodiment of this application provides Figure 1 Schematic diagram of the process of step S130, step S130 may include steps S131-S132:
[0048] S131. Determine position information of a detection frame surrounding the multiple eye key points in the face image based on the key point coordinates of the multiple eye key points.
[0049] S132 . According to the position information of the detection frame, intercept the pixel area surrounded by the detection frame in the face image to obtain an eye area image.
[0050] In some embodiments, the detection frame may be a rectangular frame, and the position information of the detection frame may be defined by the position information of a reference point of the detection frame and the size of the detection frame. For example, the reference point of the detection frame may be the geometric center of the detection frame or any vertex of the detection frame. The size of the detection frame may be determined based on the maximum horizontal span and the maximum vertical span involved in the key point coordinates of multiple eye key points. For example, the width of the detection frame may be equal to the maximum horizontal span, and the height of the detection frame may be equal to the maximum vertical span. In some embodiments, a portion of excess size may be added to the maximum horizontal span and the maximum vertical span to form the corresponding width and height of the detection frame.
[0051] In other embodiments, the position information of the detection frame may include the position information of each boundary of the detection frame; specifically, the position information of multiple boundaries may be determined by multiple eye key points, the position information of the detection frame may be determined by the position information of multiple boundaries, and then the eye area image may be captured in the face image through the detection frame.
[0052] In some embodiments, the plurality of eye key points include a left outer corner key point, a right outer corner key point, an eyebrow key point, a left lower eyelid key point, and a right lower eyelid key point; the position information of the detection frame includes position information of the upper boundary, position information of the lower boundary, position information of the left vertical boundary, and position information of the right vertical boundary of the detection frame. Step S131 may include steps 1 to 4:
[0053] Step 1: Determine the position information of the left vertical boundary based on the horizontal coordinate of the left outer corner key point.
[0054] Step 2: Determine the position information of the right vertical boundary based on the horizontal coordinate of the right outer corner key point.
[0055] It can be understood that the horizontal span of the eye area can be well determined by the left outer corner key point and the right outer corner key point. Therefore, the horizontal coordinate of the left outer corner key point can be used as the position information of the left vertical boundary, and the horizontal coordinate of the right outer corner key point can be used to determine the position information of the right vertical boundary, so that the detection frame can cover all the eye key points in the horizontal direction.
[0056] Step 3: Determine the position information of the upper boundary based on the vertical coordinate of the eyebrow key point.
[0057] There may be multiple eyebrow key points. In this case, the ordinate of the highest eyebrow key point in the vertical direction can be used as the position information of the upper boundary, so that the detection frame can include all eye key points.
[0058] Step 4: Determine the position information of the lower boundary based on the vertical coordinates of the left lower eyelid key point and the right lower eyelid key point.
[0059] In some implementations, the minimum value between the vertical coordinate of the first lower eyelid key point and the vertical coordinate of the second lower eyelid key point may be used as the position information of the lower boundary.
[0060] To understand the above steps 1 to 4, please refer to Figure 3 , Figure 3 The following is an example of how to capture the detection frame in a face image. Figure 3 In the figure, point a is the left outer corner key point, point b is the right outer corner key point, point c is the eyebrow key point, point d is the left lower eyelid key point, point e is the right lower eyelid key point, and area D is the constructed detection frame.
[0061] For example, the position information of each boundary can be calculated as follows:
[0062] If the horizontal coordinate of the left outer corner key point a is x0, the position information of the left vertical boundary can be x l =x0 / 2.
[0063] If the horizontal coordinate of the right outer corner key point b is x1, the position information of the right vertical boundary can be x r =(w+x1) / 2, where w is the width of the face image; obviously, x l will be smaller than x0, x r will be greater than x1, so that the left vertical boundary and the right vertical boundary can include all eye key points in the horizontal direction.
[0064] If the vertical coordinate of the eyebrow key point c is y0, then the position information of the upper boundary y u =y0.
[0065] If the vertical coordinate of the left lower eyelid key point d is y1, and the vertical coordinate of the right lower eyelid key point e is y2, then the position information of the lower boundary can be y d =min(y1,y2).
[0066] According to the above x l 、x r 、y u and y d Four boundaries can uniquely determine the position of the detection frame, such as Figure 3 As shown in the middle area D, it can be seen that the detection box D constructed based on the four boundaries can include all the eye key points.
[0067] S140 , performing eye key point detection on the eye region image based on the eye key point detection model to obtain reference coordinates of multiple eye key points.
[0068] The eye keypoint detection model is a neural network model constructed using one or more neural networks. It is used to identify the locations of eye keypoints in an image. To ensure the accuracy of the reference coordinates of the eye keypoints identified by the eye keypoint detection model, the model must be trained in advance. The specific training process is described below.
[0069] It can be understood that the key point coordinates of the eye key points are determined by taking the face image as input. During the process of face key point detection, the global features of the face image are integrated to predict the key point coordinates of the face key points (including the eye key points). The reference coordinates of the eye key points are determined by taking the eye area image as input. The reference coordinates of the eye key points are determined in combination with the local features of the eye area image.
[0070] S150 : Determine target coordinates of the multiple eye key points based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points.
[0071] In some implementations, the key point coordinates and reference coordinates of the same eye key point may be weighted averaged to obtain the target coordinates of each eye key point.
[0072] For example, for the eye key point e, its key point coordinates are (x e1 ,y e1 ), the reference coordinate is (x e2 ,y e2 ), then the target coordinates of the eye key point e can be (x e0 ,y e0 ), where x e0 =(x e1 +x e2 ) / 2,y e0 =(ye1 +y e2 ) / 2.
[0073] Through the method provided by the present application, since the key point coordinates of the eye key points are identified by using the face image as input, the key point coordinates of the eye key points are determined by integrating the global face features of the face image; since the reference coordinates of the eye key points are identified by using the eye area image as input, the reference coordinates of the eye key points are more determined based on the local features of the eye; by determining the target coordinates of the eye key points through the key point coordinates and reference coordinates of the eye key points, the target coordinates of the eye key points can effectively combine the global features of the face and the local features of the eyes, thereby improving the accuracy of eye key point positioning.
[0074] In some embodiments, see Figure 4 , Figure 4 Another flowchart of the eye key point detection method provided by the embodiment of the present application is provided. Before step S120, the eye key point detection method further includes steps S210-S240:
[0075] S210: Obtain a first sample face image and face annotation information of the first sample face image, where the face annotation information includes the annotation coordinates of a plurality of face key points in the first sample face image.
[0076] In some implementations, the first sample face image and the face annotation information of the first sample face image may be derived from an existing face database.
[0077] S220: Add a spotlight illumination effect to the first sample facial image to obtain a second sample facial image.
[0078] Adding the spotlight illumination effect to the first sample face image refers to changing the brightness of the first sample face image so that a scattered light point appears in the image and the brightness of pixels around the scattered light point in the image gradually decreases.
[0079] In actual driving scenarios, when there is insufficient lighting (such as at night, in a basement, etc.), an auxiliary infrared light source is required to capture facial images. Since the auxiliary infrared light source is a divergent light source, the facial image captured under the auxiliary infrared light source will have a focal point. Near this focal point, as the distance from the focal point increases, the brightness of the facial image gradually decreases, that is, the facial image captured in a scene with insufficient lighting presents a lighting effect of concentrated illumination. Therefore, the second sample facial image can simulate the facial image captured under insufficient lighting and illumination by the auxiliary infrared light source.
[0080] S230: Reuse the face annotation information of the first sample face image as the face annotation information of the second sample face image.
[0081] It is understandable that adding a spotlight lighting effect to the first sample face image does not change the positions of the facial key points in the first sample face image. Therefore, the face annotation information of the first sample face image can be reused as the face annotation information of the second sample face image.
[0082] S240: Train a facial key point detection model using the first sample facial image, the facial annotation information of the first sample facial image, the second sample facial image, and the facial annotation information of the second sample facial image.
[0083] In the above embodiment, a spotlighting lighting effect is added to the first sample facial image to obtain a second sample facial image, so that the second sample facial image can be used to better simulate the facial image collected under insufficient lighting conditions; on the basis of training the facial key point detection model with the first sample facial image collected under normal lighting conditions and the facial annotation information of the first sample facial image, the facial key point detection model is also trained with the second sample facial image and the facial annotation information of the second sample facial image, so that the trained facial key point detection model can not only accurately detect facial key points of facial images collected under normal lighting conditions, but also accurately detect facial key points of facial images collected under insufficient lighting conditions, thereby improving the facial key point detection capability of the facial key point detection model in different scenarios.
[0084] In some embodiments, see Figure 5 , Figure 5 The embodiment of this application provides Figure 4 Flow diagram of step S220, step S220 includes steps S221-S226:
[0085] S221 . Perform grayscale conversion on the first sample face image to obtain a grayscale image of the first sample face image.
[0086] Grayscale conversion of an image refers to converting a color image into an image with only grayscale. In this process, the image loses its original color information and only retains the brightness information, forming a black and white image or an image with different grayscale levels. In a grayscale image, the color of each pixel is determined by its grayscale value. The lower the grayscale value, the darker the color; the higher the grayscale value, the brighter the color.
[0087] It is worth mentioning that in actual driving scenarios, facial images collected under insufficient lighting conditions (such as at night, in basements, etc.) often appear in a single color. Therefore, by performing grayscale conversion on the first sample facial image to obtain a grayscale image of the first sample facial image, the facial image collected under insufficient lighting conditions can be better simulated.
[0088] S222: Determine a reference brightness value based on the brightness value of each pixel in the facial pixel area in the grayscale image.
[0089] The facial pixel area refers to the area in the face image that is surrounded by the key points of the basic facial contour.
[0090] In some embodiments, the reference brightness value may be obtained by taking a weighted average of the brightness values of each pixel in the facial pixel region; or by taking a weighted average of the brightness values of facial landmarks contained in the facial pixel region, without limitation. Alternatively, the reference brightness value may be obtained by taking the maximum brightness value in the facial pixel region of the grayscale image, or by taking the brightness value at a specified percentile in the facial pixel region of the grayscale image, such as 85%, 90%, or 95%.
[0091] S223: Determine a brightness difference value based on the target brightness value and the reference brightness value.
[0092] In some embodiments, the target brightness value may be a spotlight brightness value, which may be randomly obtained within a preset spotlight brightness value range. In practice, different spotlight illumination effects may be added to a first sample face image by changing different target brightness values.
[0093] In some implementations, a random function may be used to randomly select a brightness value from a range of spotlight brightness values as the target brightness value; the random function may be, for example, a random function.
[0094] The brightness difference value may be the difference between the target brightness value and the reference brightness value. Generally speaking, the target brightness value is greater than the reference brightness value, and therefore, the brightness difference value may be regarded as the brightness that needs to be compensated.
[0095] S224: Select a target facial key point in the first sample facial image, and use the marked coordinates of the target facial key point as focus coordinates for simulated spotlight illumination.
[0096] Without considering the initial brightness of the first sample face image, the focus coordinate is the coordinate position with the highest brightness in the first sample face image under simulated spotlight illumination. The farther from the focus coordinate, the lower the brightness.
[0097] The focal coordinates of the simulated spotlight illumination are, in other words, the focal point of the auxiliary infrared light source in the captured facial image in a simulated low-light scenario. Considering that the brightness of the focal point of the auxiliary infrared light source in actual low-light scenarios can fluctuate due to environmental factors such as the distance between the driver and the auxiliary infrared light source and external lighting, in some embodiments, a facial key point can be randomly selected from the first sample facial image as the target facial key point. It is understood that selecting a different facial key point as the target facial key point is equivalent to changing the focus illumination position. Thus, the same first sample facial image can be used to generate multiple second sample facial images exhibiting different spotlight illumination effects by changing the focus illumination position.
[0098] In some embodiments, a second brightness range can be preset. The brightness value within the second brightness range is also the brightness value that may appear at the focusing point of the auxiliary infrared light source. A brightness value is randomly selected from the second brightness range as the brightness value of the simulated spotlight illumination at the focusing coordinate.
[0099] S225 , focusing on the target face key points in the grayscale image based on the focus coordinates and the brightness difference, adding a spotlight illumination effect, and obtaining a spotlight illumination image.
[0100] In some embodiments, the overall brightness of the grayscale image can be adjusted first by using the brightness difference value (overall brightness increase or overall brightness decrease) to make the overall brightness of the grayscale image as close as possible to the brightness of the facial image captured in an actual low-light scene; then, the local brightness value of the grayscale image is adjusted according to the brightness value of the simulated concentrated illumination at the focusing coordinates, thereby simulating the illumination effect of the auxiliary infrared light source in an actual low-light scene; specifically, for the target facial key points in the grayscale image, the brightness value of the simulated concentrated illumination at the focusing coordinates needs to be added to simulate the focusing point of the auxiliary infrared light source; for other areas in the grayscale image, the increased brightness value gradually decreases as the distance from the target facial key points increases, thereby simulating the divergence of the auxiliary infrared light source.
[0101] In other embodiments, a blank image of the same size as the grayscale image can be initialized. A virtual spotlight source is then controlled to focus on the position indicated by the focus coordinates in the blank image according to the focus coordinates, and spotlight illumination is performed on the blank image according to the brightness difference value to generate an illumination image. In the illumination image, the brightness value at the position indicated by the focus coordinates is the largest, and the brightness value at the position indicated by the focus coordinates can be approximately equal to the brightness difference value. The farther the pixel from the position indicated by the focus coordinates in the illumination image, the smaller the brightness value. Subsequently, the brightness values of the pixels at the same pixel position in the illumination image and the grayscale image can be weighted averaged to obtain a spotlight illumination image. For example, the brightness value of a pixel in the spotlight illumination image can be the average of the brightness value of the pixel at the corresponding position in the illumination image and the brightness value of the pixel at the corresponding position in the grayscale image.
[0102] According to the above process, since the brightness values in the grayscale image are basically evenly distributed, and the brightness value of the position indicated by the focus coordinates in the illumination image is the largest and the brightness value of the pixel point farther away from the position indicated by the focus coordinates in the illumination image is smaller, then, after taking a weighted average of the brightness values of the corresponding pixel points in the two images, in the resulting focused illumination image, the brightness value of the position indicated by the focus coordinates is the largest and the brightness value of the pixel point farther away from the position indicated by the focus coordinates in the focused illumination image is smaller, which presents the lighting effect of focused illumination.
[0103] S226: Smoothing the spotlight image to obtain a second sample face image.
[0104] Among them, smoothing processing is also called filtering processing or blurring processing, which refers to reducing mutations or noise in image data through algorithms.
[0105] In some implementations, the smoothing process may be mean filtering, Gaussian filtering, median filtering, or the like.
[0106] Through the method provided in the embodiment of the present application, a grayscale image is obtained by grayscale conversion of the first sample facial image to simulate the monotonous characteristics of the facial image collected in the insufficient lighting scene; the overall brightness of the grayscale image is adjusted by the brightness difference to make the brightness of the grayscale image close to the brightness of the facial image collected in the insufficient lighting scene; by adding the lighting effect of spotlight illumination, the second sample facial image obtained can simulate the image collected under auxiliary infrared light, so that the facial key point detection model trained based on the second sample facial image can also accurately identify facial key points when facing the facial image collected under spotlight illumination.
[0107] To understand the above steps S221-S226, please refer to Figure 6 , Figure 6A schematic diagram of image changes of sample facial images provided in an embodiment of the present application is given. Figure 6 Where a is the first sample face image, b is the grayscale image of the first sample face image a, and c is the second sample face image obtained based on the grayscale image b.
[0108] For the first sample face image a, which is a color face image under normal lighting (the color of the face image is not shown in the figure), after grayscale conversion, a grayscale image b is obtained, that is, the color in a is removed and only the brightness information is retained. The intensity of the brightness is expressed by the grayscale value. The lower the grayscale value, the darker the color; the higher the grayscale value, the brighter the color.
[0109] After obtaining the grayscale image b, the brightness of the grayscale image b is changed through the above steps S222-S226, and the spotlight illumination effect is added to obtain the second sample face image c. It can be seen that an obvious highlight area appears in the second sample face image c, that is, the increased brightness of the highlight area is also the added spotlight illumination effect.
[0110] It should be noted that Figure 6 The sample face images shown in FIG are face images automatically generated by a computer and do not involve actual people.
[0111] See also Figure 7 , Figure 7 The embodiment of this application provides Figure 4 Flow diagram of step S240, step S240 includes steps S241-S243:
[0112] S241. Perform key point detection on a sample face image using a face key point detection model to obtain predicted coordinates of multiple face key points in the sample face image, where the sample face image is the first sample face image or the second sample face image.
[0113] S242. Calculate a first loss based on the predicted coordinates of multiple facial key points in the sample facial image and the facial annotation information of the sample facial image.
[0114] Among them, the first loss can be calculated by a loss function; the loss function can be a mean square error loss function, a cross entropy loss function, a KL divergence loss function, a quadratic regression loss function, etc., and no specific restrictions are made here.
[0115] Taking the quadratic regression loss function as an example, the first loss L can be calculated using the following formula:
[0116] Calculation formula 1:
[0117] Among them, M represents the number of sample face images (including the first sample face image and the second sample face image) used for the same batch training, N represents the number of facial key points detected in each sample face image, and represents the distance between the predicted coordinates of the nth facial key point of the mth sample face image and the annotated coordinates in the face annotation information, which is a pre-given parameter.
[0118] S243. Based on the first loss, adjust the parameters of the facial key point detection model until the first end condition is met.
[0119] The first termination condition may be that the first loss is less than a loss threshold, or that the number of training times reaches a number threshold.
[0120] In some embodiments, the face annotation information further includes the face annotation posture of the face in the corresponding face image. On this basis, the eye key point detection method may further include:
[0121] The pose prediction branch network performs pose prediction on the sample face image to obtain the predicted pose of the face in the sample face image.
[0122] Among them, the posture prediction branch network can be a branch network structure of a face key point detection model.
[0123] On this basis, see Figure 8 , Figure 8 The embodiment of this application provides Figure 7 Flow diagram of step S242, step S242 may include steps S310-S330:
[0124] S310 : Determine a key point prediction loss based on the predicted coordinates and corresponding labeled coordinates of multiple facial key points in the sample face image.
[0125] S320: Determine a pose prediction loss based on the predicted pose and the annotated face pose.
[0126] S330: Determine a first loss based on the key point prediction loss and the posture prediction loss.
[0127] In some embodiments, the posture of a face can be represented by parameters in three dimensions: pitch, yaw, and roll; pitch refers to the angle of lowering or raising the head, yaw refers to the angle of turning the head to the left or right, that is, the rotation that changes the frontal orientation of the face, and roll refers to the angle of tilting the head to the left or right, that is, the rotation that does not change the frontal orientation of the face.
[0128] Through the method provided in the embodiments of the present application, when calculating the first loss, the key point prediction loss of the predicted coordinates and the labeled coordinates, as well as the posture prediction loss of the predicted posture and the facial labeled posture are taken into account at the same time. Since the posture prediction loss can guide the facial key point detection model to learn and internalize the high-level collective features of the face, the first loss is made more comprehensive and accurate, thereby improving the training effect of the facial key point detection model.
[0129] See also Figure 9 , Figure 9 The embodiment of this application provides Figure 8 Schematic diagram of the process of step S330, step S330 may include steps S331-S332:
[0130] S331. Using the posture prediction loss of the sample face image as the sample weight of the key point prediction loss of the sample face image.
[0131] S332. According to the corresponding sample weights, the key point prediction losses of multiple sample face images in the same batch are weighted to obtain a first loss.
[0132] In some embodiments, the first loss L can be calculated using the following formula 2:
[0133] Calculation formula 2:
[0134] Among them, M represents the number of sample face images (including the first sample face image and the second sample face image) trained in the same batch, N represents the number of facial key points detected in each sample face image, represents the distance between the predicted coordinates and the labeled coordinates of the nth facial key point of the mth sample face image, represents the angular difference between the predicted posture of the nth facial key point of the sample face image and the labeled posture of the face, K represents the dimension of the facial posture, which is usually 3, i.e., including the three dimensions of pitch, yaw and roll, and represents the sample weight of the mth sample face image.
[0135] In some embodiments, the training of the eye key point detection model is similar to the training of the face key point detection model. The training of the eye key point detection model may include the following steps 5 to 8:
[0136] Step 5: Obtain a sample eye area image.
[0137] Among them, the sample eye area image can be obtained by performing key point detection on the sample face image (including the first sample face image and the second sample face image) by the trained face key point detection model, and after obtaining the eye key points, the eye area image is intercepted in the sample face image according to the key point coordinates of the eye key points.
[0138] In other implementations, the sample eye region image may be obtained by manually capturing the eye region of the sample face image.
[0139] Step 6: The eye key point detection model performs key point detection on the sample eye area image to obtain the predicted reference coordinates of multiple eye key points in the sample eye area image.
[0140] Step 7: Calculate the second loss based on the predicted reference coordinates of multiple eye key points in the sample eye area image and the face annotation information of the sample face image.
[0141] In some embodiments, the second loss l can be calculated using the following formula 3:
[0142] Formula 3:
[0143] Among them, V represents the number of sample eye area images, U represents twice the number of eye key points, and y' u is the predicted reference coordinate of the eye key point, and y u It is the labeled coordinates of the eye key points.
[0144] Step 8: Based on the second loss, adjust the parameters of the eye key point detection model until the second end condition is met.
[0145] In some embodiments, the eye key point detection model may be a CNN model (Convolutional Neural Networks), such as a lightweight framework MobileNetV1, whose structural framework is shown in the following table:
[0146]
[0147]
[0148] In the above embodiment, the CNN model can achieve model lightweighting through multi-layer separable convolution, that is, achieve lightweighting of the eye key point detection model, thereby improving the detection speed; at the same time, through the combination of the human eye key point detection model and the eye key point detection model, the detection accuracy is improved, thereby balancing the detection speed and detection accuracy.
[0149] In some embodiments, see Figure 10 , Figure 10 Another flowchart of the eye key point detection method provided in an embodiment of the present application is provided. The eye key point detection method further includes steps S410-S430:
[0150] S410: Perform eye tracking according to the target coordinates of multiple eye key points to obtain the driver's line of sight information.
[0151] Among them, eye tracking is to determine the direction and amplitude of eye movement by calculating the target coordinates of multiple eye key points after determining the eye key points, and calculating the relative positions between the eye key points and the changing trajectory of the positions of the eye key points over time. Then, the driver's line of sight information can be determined based on the direction and amplitude of the eye movement.
[0152] S420: Determine the driver's focused road target based on the driver's line of sight information.
[0153] The driver's focused road target is the focus position of the driver's line of sight on the road. For example, the driver's line of sight may be focused on the road directly in front or on the road to the left.
[0154] S430: If it is determined based on the display position information of the augmented reality head-up display image that the virtual object in the augmented reality head-up display image is not aligned with the focused road target, adjust the display position information of the augmented reality head-up display image.
[0155] Among them, the virtual objects in the augmented reality head-up display image refer to road sign images, text images, etc. simulated by the augmented reality head-up display image.
[0156] It is understandable that the virtual object in the augmented reality head-up display image is not aligned with the focused road target, which means that the virtual object in the augmented reality head-up display image is not on the driver's line of sight, and the driver needs to shift his current line of sight to pay attention to the virtual object in the augmented reality head-up display image; therefore, it is necessary to adjust the display position information of the augmented reality head-up display image so that the virtual object in the augmented reality head-up display image is located on the driver's line of sight. In this way, the driver can pay attention to the virtual object in the augmented reality head-up display image without shifting his current line of sight, which effectively avoids distraction of the driver's attention and improves driving safety.
[0157] In some embodiments, see Figure 11 , Figure 11 A schematic diagram of an eye key point detection device provided in an embodiment of the present application is provided. The eye key point detection device 500 includes:
[0158] The acquisition module 510 is used to acquire a face image.
[0159] The face detection module 520 is used to perform key point detection on the face image based on the face key point detection model to obtain key point coordinates of multiple face key points in the face image, where the multiple face key points include multiple eye key points.
[0160] The interception module 530 is used to intercept the eye area image in the face image based on the key point coordinates of multiple eye key points.
[0161] The eye detection module 540 is used to perform eye key point detection on the eye region image based on the eye key point detection model to obtain reference coordinates of multiple eye key points.
[0162] The output module 550 is configured to determine target coordinates of the multiple eye key points based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points.
[0163] In some embodiments, the eye key point detection device 500 also includes a sample acquisition module for acquiring a first sample face image and face annotation information of the first sample face image, the face annotation information including the annotation coordinates of multiple face key points in the first sample face image; an image processing module for adding a spotlight lighting effect to the first sample face image to obtain a second sample face image; and, reusing the face annotation information of the first sample face image as the face annotation information of the second sample face image; a training module for training the face key point detection model through the first sample face image, the face annotation information of the first sample face image, the second sample face image and the face annotation information of the second sample face image.
[0164] In some embodiments, the image processing module includes a graying unit for performing grayscale conversion on the first sample facial image to obtain a grayscale image of the first sample facial image; a brightness calculation unit for determining a reference brightness value based on the brightness value of each pixel point located in the facial pixel area in the grayscale image; and determining a brightness difference value based on the target brightness value and the reference brightness value; a selection unit for selecting a target facial key point in the first sample facial image and using the marked coordinates of the target facial key point as the focusing coordinates for simulating spotlight illumination; a spotlight illumination unit for focusing the target facial key point in the grayscale image based on the focusing coordinates and the brightness difference value, adding a spotlight illumination effect to obtain a spotlight illumination image; and a smoothing processing unit for smoothing the spotlight illumination image to obtain a second sample facial image.
[0165] In some embodiments, the training module includes a loss calculation unit for calculating a first loss based on the predicted coordinates of multiple facial key points in the sample facial image and the facial annotation information of the sample facial image; and a parameter adjustment unit for adjusting the parameters of the facial key point detection model based on the first loss until a first end condition is reached.
[0166] In some embodiments, the face annotation information also includes the facial annotation posture of the face in the corresponding face image; the eye key point detection device 500 also includes a posture prediction module, which is used to perform posture prediction on the sample face image by the posture prediction branch network to obtain the predicted posture of the face in the sample face image; the loss calculation unit is also used to determine the key point prediction loss based on the predicted coordinates and corresponding annotation coordinates of multiple facial key points in the sample face image; determine the posture prediction loss based on the predicted posture and the face annotation posture; and determine the first loss based on the key point prediction loss and the posture prediction loss.
[0167] In some embodiments, the loss calculation unit is specifically used to use the posture prediction loss of the sample face image as the sample weight of the key point prediction loss of the sample face image; and, according to the corresponding sample weights, weightedly process the key point prediction losses of multiple sample face images of the same batch to obtain a first loss.
[0168] In some embodiments, the interception module 530 includes a detection frame determination unit for determining position information of a detection frame surrounding multiple eye key points in a facial image based on the key point coordinates of the multiple eye key points; and a detection frame interception unit for intercepting a pixel area surrounded by the detection frame in the facial image according to the position information of the detection frame to obtain an eye area image.
[0169] In some embodiments, multiple eye key points include a left outer corner key point, a right outer corner key point, an eyebrow key point, a left lower eyelid key point, and a right lower eyelid key point; the position information of the detection frame includes the position information of the upper boundary of the detection frame, the position information of the lower boundary, the position information of the left vertical boundary, and the position information of the right vertical boundary; the detection frame determination unit is specifically used to determine the position information of the left vertical boundary according to the horizontal coordinate of the left outer corner key point; determine the position information of the right vertical boundary according to the horizontal coordinate of the right outer corner key point; determine the position information of the upper boundary according to the vertical coordinate of the eyebrow key point; determine the position information of the lower boundary according to the vertical coordinate of the left lower eyelid key point and the vertical coordinate of the right lower eyelid key point.
[0170] In some embodiments, the output module 550 is specifically configured to perform weighted averaging on the key point coordinates and reference coordinates of the same eye key point to obtain the target coordinates of each eye key point.
[0171] In some embodiments, the facial image is collected facing the driver; the eye key point detection device 500 also includes an eye tracking module, which is used to perform eye tracking based on the target coordinates of multiple eye key points to obtain the driver's line of sight information; a target confirmation module, which is used to determine the driver's focused road target based on the driver's line of sight information; and a display adjustment module, which is used to adjust the display position information of the augmented reality head-up display image if it is determined that the virtual object in the augmented reality head-up display image is not aligned with the focused road target based on the display position information of the augmented reality head-up display image.
[0172] In some embodiments, according to the eye key point detection method provided in the above embodiment, the embodiment of the present application also provides an electronic device, such as Figure 12 , Figure 12 A structural block diagram of an electronic device provided in an embodiment of the present application is given. The electronic device 600 includes a processor 610; a memory 620; and computer-readable instructions are stored on the memory 620. When the computer-readable instructions are executed by the processor 610, the above method is implemented.
[0173] The electronic device 600 may be a terminal device, and the terminal device may be a vehicle-mounted terminal, etc.
[0174] The processor 610 may include one or more processing cores. The processor 610 utilizes various interfaces and circuits to connect the various components within the wearable device. It executes instructions, programs, code sets, or instruction sets stored in the memory 620, as well as accesses data stored in the memory 620, to perform various functions of the wearable device and process data. Optionally, the processor 610 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 610 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing displayed content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor and may be implemented separately via a communication chip.
[0175] The memory 620 may include a random access memory (RAM) or a read-only memory (ROM). The memory 620 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 620 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created by the electronic device during use.
[0176] In some embodiments, the present application further provides a computer-readable storage medium having computer-readable instructions stored thereon. When the computer-readable instructions are executed by the processor 610, the above method is implemented.
[0177] The computer-readable storage medium may be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium includes a non-transitory computer-readable storage medium. The computer-readable storage medium has storage space for computer-readable instructions for executing any of the method steps described above. These computer-readable instructions can be read from or written to one or more computer program products. The computer-readable instructions can be compressed in a suitable form.
[0178] In particular, according to embodiments of the present application, the processes described above may be implemented as computer software programs. For example, embodiments of the present application include a computer program product comprising computer instructions. When the computer instructions are executed by a central processing unit (CPU), various functions defined in the system of the present application are performed.
[0179] In the embodiments of the present application, the term "module" or "unit" refers to a computer program or a part of a computer program that has a predetermined function and works together with other related parts to achieve a predetermined goal. It can be implemented in whole or in part by using software, hardware (such as processing circuits or memories), or a combination thereof. Similarly, a processor (or multiple processors or memories) can be used to implement one or more modules or units. In addition, each module or unit can be part of an overall module or unit of the module or unit function.
[0180] The above is only a preferred embodiment of the present application and does not constitute any form of limitation to the present application. Although the present application has been disclosed as above with preferred embodiments, it is not intended to limit the present application. Any person skilled in the art can make some changes or modifications to equivalent embodiments using the technical contents disclosed above without departing from the scope of the technical solution of the present application. However, any brief modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present application without departing from the content of the technical solution of the present application are still within the scope of the technical solution of the present application.
[0181] In this application, when the facial image data involved is applied to specific products or technologies in the above embodiments of this application, the relevant data collection, use and processing processes should comply with national laws and regulations. Before collecting facial image data, the information processing rules should be informed and the target object's separate consent should be obtained. The relevant data should be processed strictly in accordance with the requirements of laws and regulations and personal information processing rules, and technical measures should be taken to ensure the security of the relevant data.
Claims
1. A method for detecting key points of an eye, characterized in that: include: Get face image; Performing key point detection on the facial image based on a facial key point detection model to obtain key point coordinates of a plurality of facial key points in the facial image, wherein the plurality of facial key points include a plurality of eye key points; Based on the key point coordinates of the multiple eye key points, intercepting an eye area image in the face image; Performing eye key point detection on the eye region image based on an eye key point detection model to obtain reference coordinates of the multiple eye key points; The target coordinates of the multiple eye key points are determined based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points.
2. The method according to claim 1, characterized in that Before performing key point detection on the facial image based on the facial key point detection model to obtain key point coordinates of multiple facial key points in the facial image, the method further includes: Acquire a first sample face image and face annotation information of the first sample face image, wherein the face annotation information includes the annotated coordinates of a plurality of facial key points in the first sample face image; adding a spotlight illumination effect to the first sample face image to obtain a second sample face image; Reusing the face annotation information of the first sample face image as the face annotation information of the second sample face image; The facial key point detection model is trained using the first sample facial image, the facial annotation information of the first sample facial image, the second sample facial image, and the facial annotation information of the second sample facial image.
3. The method according to claim 2, characterized in that Adding a spotlight illumination effect to the first sample face image to obtain a second sample face image includes: Performing grayscale conversion on the first sample face image to obtain a grayscale image of the first sample face image; determining a reference brightness value based on the brightness value of each pixel in the facial pixel area in the grayscale image; Determining a brightness difference value based on the target brightness value and the reference brightness value; Selecting a target facial key point in the first sample facial image, and using the marked coordinates of the target facial key point as the focus coordinates of the simulated spotlight illumination; Focusing the target face key points in the grayscale image based on the focus coordinates and the brightness difference, adding a spotlight illumination effect, and obtaining a spotlight illumination image; The spotlight illumination image is smoothed to obtain the second sample face image.
4. The method according to claim 2, characterized in that The training of the facial key point detection model using the first sample face image, the face annotation information of the first sample face image, the second sample face image, and the face annotation information of the second sample face image includes: performing key point detection on a sample facial image using the facial key point detection model to obtain predicted coordinates of a plurality of facial key points in the sample facial image, where the sample facial image is the first sample facial image or the second sample facial image; Calculating a first loss based on the predicted coordinates of a plurality of facial key points in the sample facial image and the facial annotation information of the sample facial image; Based on the first loss, parameters of the facial key point detection model are adjusted until a first end condition is met.
5. The method according to claim 4, characterized in that The face annotation information also includes the face annotation posture of the face in the corresponding face image; The method further comprises: Performing posture prediction on the sample face image by a posture prediction branch network to obtain a predicted posture of the face in the sample face image; The calculating of the first loss based on the predicted coordinates of the plurality of facial key points in the sample facial image and the facial annotation information of the sample facial image comprises: Determining a key point prediction loss based on the predicted coordinates and corresponding annotated coordinates of a plurality of facial key points in the sample face image; determining a pose prediction loss based on the predicted pose and the annotated face pose; The first loss is determined based on the key point prediction loss and the pose prediction loss.
6. The method according to claim 5, characterized in that The determining the first loss based on the key point prediction loss and the posture prediction loss includes: Using the posture prediction loss of the sample face image as the sample weight of the key point prediction loss of the sample face image; According to the corresponding sample weights, the key point prediction losses of the multiple sample face images in the same batch are weighted to obtain the first loss.
7. The method according to any one of claims 1 to 6, characterized in that The step of intercepting an eye region image from the face image based on the key point coordinates of the plurality of eye key points includes: Determining position information of a detection frame surrounding the multiple eye key points in the face image based on the key point coordinates of the multiple eye key points; According to the position information of the detection frame, a pixel area surrounded by the detection frame is intercepted in the face image to obtain the eye area image.
8. The method according to claim 7, characterized in that The multiple eye key points include a left outer corner key point, a right outer corner key point, an eyebrow key point, a left lower eyelid key point, and a right lower eyelid key point; the position information of the detection frame includes position information of the upper boundary, position information of the lower boundary, position information of the left vertical boundary, and position information of the right vertical boundary of the detection frame; The determining, based on the key point coordinates of the multiple eye key points, position information of a detection frame surrounding the multiple eye key points in the face image includes: Determine the position information of the left vertical boundary according to the horizontal coordinate of the left outer corner key point; Determine the position information of the right vertical boundary according to the horizontal coordinate of the right outer corner key point; Determining the position information of the upper boundary according to the vertical coordinate of the eyebrow key point; The position information of the lower boundary is determined according to the vertical coordinate of the left lower eyelid key point and the vertical coordinate of the right lower eyelid key point.
9. The method according to any one of claims 1 to 6, characterized in that The determining the target coordinates of the multiple eye key points based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points includes: The key point coordinates and reference coordinates of the same eye key point are weighted averaged to obtain the target coordinates of each eye key point.
10. The method according to any one of claims 1 to 6, characterized in that The facial image is collected facing the driver; after determining the target coordinates of the multiple eye key points based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points, the method further includes: performing eye tracking according to the target coordinates of the multiple eye key points to obtain the driver's line of sight information; determining a road target focused on by the driver based on the driver's line of sight information; If it is determined based on the display position information of the augmented reality head-up display image that the virtual object in the augmented reality head-up display image is not aligned with the focused road target, the display position information of the augmented reality head-up display image is adjusted.
11. An eye key point detection device, characterized in that: include: An acquisition module is used to acquire a face image; A face detection module, configured to perform key point detection on the face image based on a face key point detection model to obtain key point coordinates of a plurality of face key points in the face image, wherein the plurality of face key points include a plurality of eye key points; A capture module, configured to capture an eye region image in the face image based on the key point coordinates of the plurality of eye key points; An eye detection module, configured to perform eye key point detection on the eye region image based on an eye key point detection model, and obtain reference coordinates of the plurality of eye key points; An output module is configured to determine target coordinates of the multiple eye key points based on the key point coordinates of the multiple eye key points and the reference coordinates of the multiple eye key points.
12. An electronic device, characterized in that: include: processor; A memory having computer-readable instructions stored thereon, wherein when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 10 is implemented.
13. A computer-readable storage medium, characterized in that Computer-readable instructions are stored thereon, and when the computer-readable instructions are executed by a processor, the method according to any one of claims 1 to 10 is implemented.
Citation Information
Patent Citations
Image processing method, apparatus, electronic device, and computer-readable storage medium
CN109242794A
Face key point prediction method, APP, terminal device and storage medium
CN116152896A
Key point detection method and device and electronic equipment
CN116168438A
Living body detection method and device, equipment and storage medium
CN116206374A